YouTube2Text

AI & ML Full Course 2026 | Complete Artificial Intelligence and Machine Learning Tutorial | Edureka — Transcript

by edureka! · 130,062 words · 20,074 segments · language en · Watch on YouTube

Full transcript

  1. 0:09Hello everyone and welcome to the AI
  2. 0:11[music] and machine learning full
  3. 0:12course. Artificial intelligence and
  4. 0:15machine learning are rapidly changing
  5. 0:17the way organization [music] analyze
  6. 0:18data, automate processes, and build
  7. 0:21intelligent applications. [music]
  8. 0:23This course provides a complete
  9. 0:25introduction to a key ideas, techniques,
  10. 0:27and tools
  11. 0:28>> [music]
  12. 0:28>> used in modern AI and machine learning.
  13. 0:31Throughout the course, you will explore
  14. 0:33how machines [music] learn from data,
  15. 0:35understand the role of algorithms and
  16. 0:37models, and see how AI systems are
  17. 0:40applied to solve real-world problems.
  18. 0:41[music]
  19. 0:42The course is designed to build your
  20. 0:44understanding step-by-step, making
  21. 0:46complex concepts [music] easier to
  22. 0:48understand. And by the end of this
  23. 0:50course, you will have a strong
  24. 0:51foundation in AI and machine learning
  25. 0:54and a clear view [music] of how these
  26. 0:56technologies are driving innovation
  27. 0:57across industries. So, before we begin,
  28. 1:00please like, share, and subscribe to
  29. 1:01Edureka's YouTube channel and hit the
  30. 1:03bell icon to stay updated on the latest
  31. 1:05tech content [music] from Edureka. Also,
  32. 1:08check out Edureka's postgraduate program
  33. 1:10in generative [music] AI and machine
  34. 1:12learning in collaboration with Illinois
  35. 1:14Tech. It offers a unique opportunity to
  36. 1:16explore [music] the cutting-edge world
  37. 1:18of generative AI and develop advanced
  38. 1:20AI-powered solutions. This program
  39. 1:23[music] covers in-demand topics
  40. 1:24including machine learning, deep
  41. 1:26learning, natural language processing,
  42. 1:28[music] prompt engineering, generative
  43. 1:30AI, LLMs, RAG, agentic AI, and much
  44. 1:33more. Learn from industry [music]
  45. 1:35experts through a curriculum built
  46. 1:37around real-world hands-on use cases
  47. 1:39designed to equip you with a practical
  48. 1:41and job-ready skills. So, check out the
  49. 1:43course link given in the description box
  50. 1:45[music] below. Now, let us get started
  51. 1:48by understanding what artificial
  52. 1:49intelligence is.
  53. 1:51AI or artificial intelligence is a
  54. 1:54branch of computer science where focused
  55. 1:56on creating systems that were performed
  56. 1:59task and that would be normally required
  57. 2:01by human intelligence. These tasks can
  58. 2:04range understanding of natural language.
  59. 2:07Secondly, recognizing patterns, then
  60. 2:10making decisions, and lastly, learning
  61. 2:12from the experiences.
  62. 2:14AI is a collaboration of ideas, methods,
  63. 2:18and knowledge where from the multiple
  64. 2:20academic disciplines work on a different
  65. 2:22problem-solving and share their
  66. 2:24knowledge to a better understanding and
  67. 2:26come up with a good solution.
  68. 2:28But, it can also be rule-based and
  69. 2:31operate under a set of rules and
  70. 2:33conditions only.
  71. 2:34So, I would like to tell you all guys a
  72. 2:37small real-time experience or an
  73. 2:39experiment done by Alan Turing to
  74. 2:42propagate and to establish the
  75. 2:44artificial intelligence.
  76. 2:46Alan Turing was the first person to
  77. 2:48conduct the sustainable research in the
  78. 2:50field that he called machine
  79. 2:52intelligence.
  80. 2:53The Turing test was conducted to explore
  81. 2:55whether machines could exhibit
  82. 2:57human-like intelligence.
  83. 2:59Proposed by Alan Turing in 1950, it
  84. 3:02involves a human evaluator communicating
  85. 3:05with both a human and a machine through
  86. 3:08a text interface.
  87. 3:09If a evaluator cannot distinguish
  88. 3:11between the two based on their
  89. 3:13responses, the machine is said to have
  90. 3:15passed the test. It serves as a
  91. 3:17benchmark of the assessing the progress
  92. 3:20of AI and discussions of the nature and
  93. 3:23intelligence and consciousness.
  94. 3:25So, this is to evolve and to establish
  95. 3:28that even machines can work as humans.
  96. 3:32And that is how it is made to bring up
  97. 3:35the machines' knowledge and the humans'
  98. 3:37knowledge into machines.
  99. 3:39There are two types of artificial
  100. 3:40intelligence.
  101. 3:42First, let's talk about weak AI. Before
  102. 3:44getting into what is weak AI, I would
  103. 3:46like to tell you with an example that is
  104. 3:49performed in a real world.
  105. 3:51I know you're all guys will be knowing
  106. 3:53about Alexa, Apple Siri, and also
  107. 3:55self-driving vehicles.
  108. 3:57So, all of these considered as in weak
  109. 4:00AI. It is also known as a narrow AI or
  110. 4:04an AI narrow intelligent that is trained
  111. 4:07by the AI and focused to perform a
  112. 4:09specific task. Weak AI drives most of
  113. 4:12the AI that surrounds by us today.
  114. 4:14Narrow might be a more apt descriptor
  115. 4:17for this type of AI as it is anything
  116. 4:19but weak.
  117. 4:20Next, let us learn about strong AI.
  118. 4:23Strong AI is made up of artificial
  119. 4:25general intelligence or artificial super
  120. 4:28intelligence.
  121. 4:30Simply, it could be told as where a
  122. 4:31machine would have a intelligence equal
  123. 4:34to humans.
  124. 4:35Where humans will track what AI needs to
  125. 4:37be done.
  126. 4:39It would be able to self-aware with the
  127. 4:41consciousness that would be this ability
  128. 4:43to solve problems, learn, and also plan
  129. 4:45for the future.
  130. 4:46Have you ever thought what does AI do at
  131. 4:49its core?
  132. 4:50I'm here to tell you what. AI is
  133. 4:52essential. It works by analyzing a lot
  134. 4:55of data
  135. 4:57to find patterns and a useful
  136. 4:58information.
  137. 4:59It even learns from this data to get
  138. 5:02better at the task over time. With this
  139. 5:04learning, AI can make decisions, predict
  140. 5:06future events, and do tasks
  141. 5:08automatically that would normally need
  142. 5:10human intelligence.
  143. 5:11This help business and other
  144. 5:13organizations work more efficiently.
  145. 5:16So, AI is about making computers smarter
  146. 5:19and more helpful in everyday life. So,
  147. 5:22let me tell you some use cases that is
  148. 5:24happening in daily uses or daily real
  149. 5:27life based.
  150. 5:28Firstly, I have taken is about
  151. 5:30cybersecurity. As it is a very important
  152. 5:33and a vital role for many platforms here
  153. 5:35after.
  154. 5:36Cybersecurity is a critical concern for
  155. 5:38individual businesses and also
  156. 5:40governments as cyber threats continue to
  157. 5:43evolve in a complexity and
  158. 5:44sophistication. It would play a very
  159. 5:47role for augmenting cybersecurity
  160. 5:49defenses.
  161. 5:50It is all based on the false positive
  162. 5:53effects that is made by the false
  163. 5:55information given by any criteria.
  164. 5:58Leading the alert of fatigue and reduced
  165. 6:01operational efficiencies.
  166. 6:03Cyber attacks can exploit
  167. 6:05vulnerabilities in AI models
  168. 6:08by invading detection and compromising
  169. 6:10security defenses.
  170. 6:12They have some private data where it is
  171. 6:15trained on the basis of incomplete data
  172. 6:18sets may produce the outcomes of private
  173. 6:21data. This will raise an ethical concern
  174. 6:24and also regulatory compliances issues.
  175. 6:27Having a very complicated problems, too.
  176. 6:29So, the next one will be your
  177. 6:31entertainment. As people know,
  178. 6:33entertainment is taking a very huge part
  179. 6:36in everyone's life.
  180. 6:37For example, it would be your social
  181. 6:39media, too.
  182. 6:40So, AI is transforming the entertainment
  183. 6:42industry by revolutionizing content
  184. 6:45creation, personalization, and audience
  185. 6:47engagement. So, this can be such as
  186. 6:50television, gaming, music, and also
  187. 6:52digital media platforms.
  188. 6:54One significant use of AI in
  189. 6:56entertainment is personalized content
  190. 6:59recommendation.
  191. 7:01So, personalized entertainment
  192. 7:02experiences are enhanced through AI
  193. 7:04driven. It is recommended through
  194. 7:06systems.
  195. 7:07They are even having some platforms like
  196. 7:10Netflix,
  197. 7:11Prime Amazon, and also Spotify
  198. 7:13that is making people engaged in a very
  199. 7:15hype as of now.
  200. 7:17These recommendations improve over time.
  201. 7:20So, I would like to conclude by telling
  202. 7:22artificial intelligence can do amazing
  203. 7:24things like analyzing data and making
  204. 7:27task easier.
  205. 7:28But, it also brings up huge question
  206. 7:31about fairness, jobs, and who controls
  207. 7:33it.
  208. 7:34We need to be careful in how we develop
  209. 7:36and use the AI.
  210. 7:37Making sure it helps everyone and
  211. 7:39doesn't cause harm at all.
  212. 7:41So, this makes easier for people to get
  213. 7:43into creativity and make rules and
  214. 7:45guidelines.
  215. 7:49>> [music]
  216. 7:52>> So, now let's get started with the first
  217. 7:53topic, which is history of artificial
  218. 7:56intelligence.
  219. 7:57The concept of AI goes back to the
  220. 7:59classical ages. Under Greek mythology,
  221. 8:02the concept of machines and mechanical
  222. 8:04men were well thought of. An example is
  223. 8:07Talos. Talos was supposedly a giant
  224. 8:10animated bronze warrior who was
  225. 8:12programmed to guard the island of Crete.
  226. 8:15Now, let's get back to the 19th century.
  227. 8:17In 1950, Alan Turing proposed the Turing
  228. 8:20test. The Turing test basically
  229. 8:22determines whether or not a computer can
  230. 8:25intelligently think like a human being.
  231. 8:27The Turing test was the first serious
  232. 8:29proposal in the philosophy of artificial
  233. 8:31intelligence.
  234. 8:331951 marked the era for game artificial
  235. 8:36intelligence. This period was called
  236. 8:38game AI because here a lot of computer
  237. 8:40scientists developed programs for
  238. 8:42checkers and for chess. However, these
  239. 8:45programs were later rewritten and redone
  240. 8:47in a better way.
  241. 8:491956 marked the most important year for
  242. 8:52artificial intelligence.
  243. 8:54During this year, John McCarthy first
  244. 8:56coined the term artificial intelligence.
  245. 8:58This was followed by the first AI
  246. 9:00laboratory, which was set up in 1959.
  247. 9:03MIT AI Lab was the first setup, which
  248. 9:06was basically dedicated to the research
  249. 9:08of AI.
  250. 9:09In 1960, the first robot was introduced
  251. 9:12to the General Motors assembly line. In
  252. 9:151961, the first AI chatbot called Eliza
  253. 9:19was introduced. In 1997, IBM's Deep Blue
  254. 9:23beats the world champion Garry Kasparov
  255. 9:25in the game of chess.
  256. 9:272005 marks for the year when an
  257. 9:29autonomous robotic car called Stanley
  258. 9:32won the DARPA Grand Challenge.
  259. 9:34In 2011, IBM's question-answering
  260. 9:37machine Watson defeated the two greatest
  261. 9:40Jeopardy champions Brad Rutter and Ken
  262. 9:42Jennings. So, that was a brief history
  263. 9:45of AI. Now guys, since the emergence of
  264. 9:47artificial intelligence in 1950s, we
  265. 9:50have seen an exponential growth in its
  266. 9:52potential. AI covers domains such as
  267. 9:55machine learning, deep learning, neural
  268. 9:57networks, natural language processing,
  269. 9:59knowledge-based expert systems, and so
  270. 10:02on.
  271. 10:02Now that you know a brief history of
  272. 10:04artificial intelligence, let's move on
  273. 10:06and understand what exactly artificial
  274. 10:08intelligence is. So, the term artificial
  275. 10:11intelligence was first coined by John
  276. 10:13McCarthy, like I mentioned earlier. He
  277. 10:15defined AI as a science and engineering
  278. 10:18of making intelligent machines. In other
  279. 10:20words, artificial intelligence can also
  280. 10:23be defined as a development of computer
  281. 10:25systems that are capable of performing
  282. 10:28tasks that require human intelligence
  283. 10:30such as decision-making, object
  284. 10:32detection, solving complex problems, and
  285. 10:34so on. So, like I mentioned, artificial
  286. 10:37intelligence helps in decision-making,
  287. 10:39solving complex problems, it performs
  288. 10:42high-level computations, and also
  289. 10:44increases the accuracy of your
  290. 10:46predictions. Right? These are the main
  291. 10:48features of AI.
  292. 10:49So, now let's understand the different
  293. 10:51stages of artificial intelligence. So,
  294. 10:54basically, when I was doing my research,
  295. 10:55I found a lot of videos and a lot of
  296. 10:58articles that stated that artificial
  297. 11:00general intelligence, artificial narrow
  298. 11:03intelligence, and artificial super
  299. 11:04intelligence are the different types of
  300. 11:07AI. If I have to be more precise with
  301. 11:09you, then artificial intelligence has
  302. 11:11three different stages. Right? The types
  303. 11:13of AI are completely different from the
  304. 11:15stages of AI. So, under the stages of
  305. 11:17artificial intelligence, we have
  306. 11:19artificial narrow intelligence,
  307. 11:21artificial general intelligence, and
  308. 11:23artificial super intelligence.
  309. 11:25So, what is artificial narrow
  310. 11:26intelligence? Artificial narrow
  311. 11:28intelligence, also known as weak AI, is
  312. 11:31a stage of artificial intelligence that
  313. 11:33involves machines that can perform only
  314. 11:36a narrowly defined set of specific
  315. 11:38tasks. Right? At this stage, the
  316. 11:40machines don't possess any thinking
  317. 11:42ability. They just perform a set of
  318. 11:44predefined functions. Examples of weak
  319. 11:47AI include Siri, Alexa, AlphaGo, Sophia,
  320. 11:50the self-driving cars, and so on. Almost
  321. 11:53all the AI-based systems that are built
  322. 11:55till this date fall under the category
  323. 11:57of weak AI or artificial narrow
  324. 11:59intelligence.
  325. 12:01Next, we have something known as
  326. 12:02artificial general intelligence.
  327. 12:04Artificial general intelligence is also
  328. 12:06known as strong AI. This stage is the
  329. 12:09evolution of artificial intelligence,
  330. 12:11wherein machines will possess the
  331. 12:13ability to think and make decisions just
  332. 12:16like human beings. There are currently
  333. 12:18no existing examples of strong AI, but
  334. 12:21it's believed that we will soon be able
  335. 12:23to create machines that are as smart as
  336. 12:25human beings. Strong AI is actually
  337. 12:28considered a threat to human existence
  338. 12:30by many scientists. This includes
  339. 12:32Stephen Hawking. Stephen Hawking quoted
  340. 12:35that the development of full artificial
  341. 12:37intelligence could spell the end of
  342. 12:40human race. Moving on to our last stage,
  343. 12:42which is artificial super intelligence.
  344. 12:45Artificial super intelligence is that
  345. 12:47stage of AI when the capability of
  346. 12:49computers will surpass human beings.
  347. 12:52Artificial super intelligence is
  348. 12:54currently seen as a hypothetical
  349. 12:56situation as depicted in movies and
  350. 12:58science fiction books. You see a lot of
  351. 13:00movies which show that machines are
  352. 13:02taking over the world. All of that is
  353. 13:04artificial super intelligence. Now, I
  354. 13:06believe that machines are not very far
  355. 13:08from reaching the stage taking into
  356. 13:10consideration our current pace.
  357. 13:12However, such systems don't currently
  358. 13:14exist, right? We don't have any machine
  359. 13:16that is capable of thinking better than
  360. 13:19a human being or reasoning in a better
  361. 13:21way than a human. Artificial super
  362. 13:23intelligence, basically any robot that
  363. 13:24is much smarter than humans. Now, moving
  364. 13:27on to the different types of artificial
  365. 13:29intelligence. Based on the functionality
  366. 13:32of AI-based systems, artificial
  367. 13:34intelligence can be categorized into
  368. 13:36four types. The first type is reactive
  369. 13:38machines AI. This type of AI includes
  370. 13:41machines that operate solely based on
  371. 13:44the present data and take into
  372. 13:46consideration only the current
  373. 13:48situation.
  374. 13:49Reactive AI machines cannot form
  375. 13:51inferences from the data to evaluate any
  376. 13:54future actions. They can perform a
  377. 13:56narrowed range of predefined tasks.
  378. 13:59An example of reactive AI is the famous
  379. 14:02IBM chess program that beat the world
  380. 14:04champion Garry Kasparov.
  381. 14:06This is one of the most impressive AI
  382. 14:08machines built so far.
  383. 14:10Next, we have limited memory AI. Now,
  384. 14:13like the name suggests, limited memory
  385. 14:15AI can make informed and improved
  386. 14:17decisions by studying the past data from
  387. 14:20its memory. So, such an AI has a
  388. 14:22short-lived or you can say a temporary
  389. 14:25memory that can be used to store past
  390. 14:27experiences and hence evaluate your
  391. 14:29future actions.
  392. 14:31Self-driving cars are limited memory AI
  393. 14:33that use the data collected in the
  394. 14:35recent past to make immediate decisions.
  395. 14:38For example, self-driving cars use
  396. 14:40sensors to identify civilians that are
  397. 14:43crossing the road. They identify any
  398. 14:45steep roads or traffic signals and they
  399. 14:48use this to make better driving
  400. 14:49decisions. This also helps in preventing
  401. 14:52any future accidents. Next, we have
  402. 14:54something known as theory of mind
  403. 14:56artificial intelligence. The theory of
  404. 14:58mind AI is a more advanced type of
  405. 15:00artificial intelligence. This category
  406. 15:03is speculated to play a very important
  407. 15:06role in psychology. This type of AI will
  408. 15:08mainly focus on emotional intelligence
  409. 15:11so that human beliefs and thoughts can
  410. 15:13be better comprehended. The theory of
  411. 15:15mind AI has not been fully developed
  412. 15:18yet, but rigorous research is happening
  413. 15:20in this area.
  414. 15:21Moving on to our last type of artificial
  415. 15:23intelligence is the self-aware
  416. 15:25artificial intelligence.
  417. 15:27So guys, let us fold hands and pray that
  418. 15:29we don't reach the state of AI where
  419. 15:32machines have their own consciousness
  420. 15:34and become self-aware. This type of AI
  421. 15:37is a little far-fetched, but in the
  422. 15:38future achieving a stage of
  423. 15:40superintelligence might be possible.
  424. 15:43Geniuses like Elon Musk and Stephen
  425. 15:45Hawking have constantly warned us about
  426. 15:47evolution of AI.
  427. 15:49So guys, let me know your thoughts in
  428. 15:50the comment section. Do you ever think
  429. 15:52we'll reach the stage of artificial
  430. 15:54superintelligence?
  431. 15:56Moving on to the last topic of today's
  432. 15:58session is the different domains or the
  433. 16:00different branches of artificial
  434. 16:02intelligence. So artificial intelligence
  435. 16:04can be used to solve real-world problems
  436. 16:06by implementing machine learning, deep
  437. 16:08learning, natural language processing,
  438. 16:10robotics, expert systems, and fuzzy
  439. 16:13logic. Now guys, these are the different
  440. 16:15domains or you can say the different
  441. 16:16branches that AI uses in order to solve
  442. 16:19any problem. Recently, AI has also been
  443. 16:22used as an application in computer
  444. 16:24vision and image processing. Right, for
  445. 16:26now let me tell you briefly about each
  446. 16:28of these domains. Machine learning is
  447. 16:30basically the science of getting
  448. 16:32machines to interpret, process, and
  449. 16:34analyze data in order to solve
  450. 16:36real-world problems. Right, under
  451. 16:38machine learning there's supervised,
  452. 16:40unsupervised, and reinforcement
  453. 16:41learning. If any of you are interested
  454. 16:43in learning about these technologies,
  455. 16:45I'll leave a link in the description
  456. 16:46box. You all can go through that
  457. 16:47content. Next, we have deep learning or
  458. 16:50neural networks. So deep learning is a
  459. 16:52process of implementing neural networks
  460. 16:54on high-dimensional data to gain
  461. 16:57insights and form solutions. It is
  462. 16:59basically the logic behind the face
  463. 17:01verification algorithm on Facebook. It
  464. 17:04is the logic behind the self-driving
  465. 17:06cars, virtual assistants like Siri and
  466. 17:08Alexa. Then we have natural language
  467. 17:10processing. Natural language processing
  468. 17:12refers to the science of drawing
  469. 17:14insights from natural human language in
  470. 17:16order to communicate with machines and
  471. 17:19grow businesses. So, an example of NLP
  472. 17:22is Twitter and Amazon. Twitter uses NLP
  473. 17:25to filter out terroristic language in
  474. 17:27their tweets. Amazon uses NLP to
  475. 17:30understand customer reviews and improve
  476. 17:32user experience. Then we have robotics.
  477. 17:35Robotics is a branch of artificial
  478. 17:37intelligence which focuses on the
  479. 17:39different branches and applications of
  480. 17:41robots. AI robots are artificial agents
  481. 17:44which act in the real world environment
  482. 17:47to produce results by taking some
  483. 17:49accountable actions.
  484. 17:51So, I'm sure all of you have heard of
  485. 17:52Sophia. Sophia the humanoid is a very
  486. 17:55good example of AI in robotics. Then we
  487. 17:58have fuzzy logic. So, fuzzy logic is a
  488. 18:00computing approach that is based on the
  489. 18:02principle of degree of truth instead of
  490. 18:05the usual modern logic that we use which
  491. 18:08is basically the Boolean logic. Fuzzy
  492. 18:10logic is used in medical fields to solve
  493. 18:12complex problems which involve decision
  494. 18:15making. It is also used in automating
  495. 18:18gear systems in your cars and all of
  496. 18:20that. Then we have expert systems. An
  497. 18:22expert system is an AI-based computer
  498. 18:24system that learns and reciprocates the
  499. 18:27decision-making ability of a human
  500. 18:29expert. Expert systems use if-then logic
  501. 18:32notions in order to solve any complex
  502. 18:34problem. They do not rely on
  503. 18:37conventional procedural programming.
  504. 18:39Expert systems are mainly used in
  505. 18:41information management. They're seen to
  506. 18:43be used in fraud detection, virus
  507. 18:46detection, also in managing medical and
  508. 18:48hospital records, and so on. So, guys,
  509. 18:50to sum it up, these were the different
  510. 18:52branches of artificial intelligence.
  511. 18:55>> [music]
  512. 19:00>> These are the term which have confused a
  513. 19:02lot of people. And if you too are one
  514. 19:04among them, let me resolve it for you.
  515. 19:07Well, artificial intelligence is a
  516. 19:09broader umbrella under which machine
  517. 19:11learning and deep learning come. You can
  518. 19:13also see in the diagram that even deep
  519. 19:15learning is a subset of machine
  520. 19:17learning. So, you can say that all three
  521. 19:19of them, the AI, the machine learning,
  522. 19:21and deep learning, are just the subset
  523. 19:24of each other. So, let's move on and
  524. 19:26understand how exactly they differ from
  525. 19:28each other. So, let's start with
  526. 19:30artificial intelligence. The term
  527. 19:32artificial intelligence was first coined
  528. 19:35in the year 1956.
  529. 19:37The concept is pretty old, but it has
  530. 19:39gained its popularity recently. But why?
  531. 19:42Well, the reason is earlier we had very
  532. 19:45small amount of data. The data we had
  533. 19:48was not enough to predict the accurate
  534. 19:50result. But now, there's a tremendous
  535. 19:52increase in the amount of data.
  536. 19:54Statistics suggest that by 2020, the
  537. 19:57accumulated volume of data will increase
  538. 20:00from 4.4 zettabytes to roughly around 44
  539. 20:03zettabytes, or 44 trillion GBs of data.
  540. 20:06Along with such enormous amount of data,
  541. 20:09now we have more advanced algorithm and
  542. 20:12high-end computing power and storage
  543. 20:14that can deal with such large amount of
  544. 20:15data.
  545. 20:16As a result, it is expected that 70% of
  546. 20:19enterprise will implement AI over the
  547. 20:21next 12 months, which is up from 40% in
  548. 20:242016 and 51% in 2017.
  549. 20:28Just for your understanding, what is AI?
  550. 20:31Well, it's nothing but a technique that
  551. 20:33enables the machine to act like humans
  552. 20:35by replicating the behavior and nature.
  553. 20:38With AI, it is possible for machine to
  554. 20:40learn from the experience. The machines
  555. 20:43adjust their responses based on new
  556. 20:45input, thereby performing human-like
  557. 20:47tasks.
  558. 20:48Artificial intelligence can be trained
  559. 20:50to accomplish specific tasks by
  560. 20:51processing large amount of data and
  561. 20:53recognizing pattern in them.
  562. 20:55You can consider that building an
  563. 20:57artificial intelligence is like building
  564. 20:59a church. The first church took
  565. 21:01generations to finish. So, most of the
  566. 21:04workers who were working on it never saw
  567. 21:06the final outcome. Those working on it
  568. 21:08took pride in their crafts, building
  569. 21:10bricks and chiseling stone that was
  570. 21:12going to be placed into the great
  571. 21:13structure. So, as AI researchers, we
  572. 21:16should think of ourselves as humble
  573. 21:18brick makers whose job is to study how
  574. 21:21to build components, example parsers,
  575. 21:23planners, or learning algorithm, or
  576. 21:25etc., anything that someday someone and
  577. 21:27somewhere will integrate into the
  578. 21:29intelligent systems. Some of the
  579. 21:31examples of artificial intelligence from
  580. 21:33our day-to-day life are Apple series,
  581. 21:36chess-playing computer, Tesla's
  582. 21:38self-driving car, and many more. These
  583. 21:40examples are based on deep learning and
  584. 21:42natural language processing.
  585. 21:44Well, this was about what is AI and how
  586. 21:46it gained its height. So, moving on
  587. 21:48ahead, let's discuss about machine
  588. 21:50learning and see what it is and why it
  589. 21:53was even introduced. Well, machine
  590. 21:55learning came into existence in the late
  591. 21:57'80s and the early '90s. But, what were
  592. 21:59the issues with the people which made
  593. 22:01the machine learning come into
  594. 22:02existence? Let us discuss them one by
  595. 22:05one.
  596. 22:05In the field of statistics, the problem
  597. 22:08was how to efficiently train large
  598. 22:10complex model. In the field of computer
  599. 22:12science and artificial intelligence, the
  600. 22:14problem was how to train more robust
  601. 22:16version of AI system. While in the case
  602. 22:18of neuroscience, problem faced by the
  603. 22:20researchers was how to design
  604. 22:22operational model of the brain.
  605. 22:24So, these were some of the issues which
  606. 22:26had the largest influence and led to the
  607. 22:28existence of the machine learning.
  608. 22:30Now, this machine learning shifted its
  609. 22:32focus from the symbolic approaches it
  610. 22:33had inherited from the AI and moved
  611. 22:36towards the methods and model it had
  612. 22:38bought from statistics and probability
  613. 22:40theory.
  614. 22:41So, let's proceed and see what exactly
  615. 22:43is machine learning. Well, machine
  616. 22:45learning is a subset of AI which enables
  617. 22:48the computer to act and make data-driven
  618. 22:50decisions to carry out a certain task.
  619. 22:52These programs or algorithms are
  620. 22:54designed in a way that they can learn
  621. 22:56and improve over time when exposed to
  622. 22:58new data. Let's see an example of
  623. 23:00machine learning. Let's say you want to
  624. 23:02create a system which tells the expected
  625. 23:04weight of a person based on its height.
  626. 23:07The first thing you do is you collect
  627. 23:08the data. Let's see, this how your data
  628. 23:10looks like. Now, each point on the graph
  629. 23:13represent one data point. To start with,
  630. 23:16we can draw a simple line to predict the
  631. 23:18weight based on the height. For example,
  632. 23:20a simple line W equal H minus 100, where
  633. 23:23W is weight in kg and H is height in cm.
  634. 23:27This line can help us to make the
  635. 23:28prediction. Our main goal is to reduce
  636. 23:31the difference between the estimated
  637. 23:32value and the actual value. So, in order
  638. 23:35to achieve it, we try to draw a straight
  639. 23:37line that fits through all these
  640. 23:39different points and minimize the error.
  641. 23:41So, our main goal is to minimize the
  642. 23:43error and make them as small as
  643. 23:45possible. Decreasing the error or the
  644. 23:47difference between the actual value and
  645. 23:49estimated value increases the
  646. 23:50performance of the model. Further on,
  647. 23:53the more data points we collect, the
  648. 23:55better our model will become. We can
  649. 23:56also improve our model by adding more
  650. 23:58variables and creating different
  651. 24:00prediction lines for them. Once the line
  652. 24:02is created, so from the next time if we
  653. 24:04feed a new data, for example, height of
  654. 24:06a person to the model, it would easily
  655. 24:08predict the data for you and it will
  656. 24:10tell you what its predicted weight could
  657. 24:12be. I hope you got a clear understanding
  658. 24:14of machine learning. So, moving on
  659. 24:16ahead, let's learn about deep learning.
  660. 24:18Now, what is deep learning? You can
  661. 24:20consider deep learning model as a rocket
  662. 24:22engine and its fuel is its huge amount
  663. 24:25of data that we feed to these
  664. 24:26algorithms.
  665. 24:27The concept of deep learning is not new.
  666. 24:30But recently, it's hype has increased
  667. 24:32and deep learning is getting more
  668. 24:33attention.
  669. 24:34This field is a particular kind of
  670. 24:36machine learning that is inspired by the
  671. 24:38functionality of our brain cells called
  672. 24:39neuron, which led to the concept of
  673. 24:42artificial neural network.
  674. 24:44It simply takes the data connection
  675. 24:45between all the artificial neurons and
  676. 24:47adjust them according to the data
  677. 24:49pattern. More neurons are added if the
  678. 24:51size of the data is large. It
  679. 24:53automatically features learning at
  680. 24:55multiple levels of abstraction, thereby
  681. 24:57allowing a system to learn complex
  682. 24:59function mapping without depending on
  683. 25:01any specific algorithm. You know what?
  684. 25:04No one actually knows what happens
  685. 25:06inside a neural network and why it works
  686. 25:08so well. So, currently you can call it
  687. 25:10as a black box. Let [snorts] us discuss
  688. 25:12some of the example of deep learning and
  689. 25:14understand it in a better way. Let me
  690. 25:16start with a simple example and explain
  691. 25:18you how things happen at a conceptual
  692. 25:21level. Let us try and understand how you
  693. 25:23recognize a square from other shapes.
  694. 25:26The first thing you do is you check
  695. 25:28whether there are four lines associated
  696. 25:30with the figure or not. Simple concept,
  697. 25:32right? If yes, we further check if they
  698. 25:35are connected and closed. Again, if yes,
  699. 25:37we finally check whether it is
  700. 25:39perpendicular and all its sides are
  701. 25:41equal. Correct? If everything fulfills,
  702. 25:44yes, it is a square.
  703. 25:46Well, it is nothing but a nested
  704. 25:47hierarchy of concepts.
  705. 25:50What we did here, we took a complex task
  706. 25:52of identifying a square in this case and
  707. 25:54broke it into simpler task. Now, this
  708. 25:56deep learning also does the same thing
  709. 25:58but at a larger scale. Let's take an
  710. 26:01example of machine which recognizes the
  711. 26:03animal. The task of the machine is to
  712. 26:05recognize whether the given image is of
  713. 26:07a cat or of a dog.
  714. 26:09What if we were asked to resolve the
  715. 26:10same issue using the concept of machine
  716. 26:12learning? What we would do? First, we
  717. 26:15would define the features such as check
  718. 26:17whether the animal has whiskers or not
  719. 26:19or check if the animal has pointed ears
  720. 26:21or not or whether its tail is straight
  721. 26:23or curved. In short, we will define the
  722. 26:25facial features and let the system
  723. 26:27identify which features are more
  724. 26:29important in classifying a particular
  725. 26:31animal. Now, when it comes to deep
  726. 26:34learning, it takes this to one step
  727. 26:35ahead. Deep learning automatically finds
  728. 26:38out the feature which are most important
  729. 26:40for classification compared to machine
  730. 26:42learning where we had to manually give
  731. 26:44out that features.
  732. 26:46By now, I guess you have understood that
  733. 26:48AI is a bigger picture and machine
  734. 26:50learning and deep learning are its
  735. 26:51subpart. So, let's move on and focus our
  736. 26:53discussion on machine learning and deep
  737. 26:55learning.
  738. 26:56The easiest way to understand the
  739. 26:58difference between the machine learning
  740. 26:59and deep learning is to know that deep
  741. 27:01learning is machine learning. More
  742. 27:03specifically, it is the next evolution
  743. 27:05of machine learning. Let's take few
  744. 27:07important parameter and compare machine
  745. 27:09learning with deep learning. So,
  746. 27:11starting with data dependencies. The
  747. 27:13most important difference between deep
  748. 27:15learning and machine learning is its
  749. 27:17performance as the volume of the data
  750. 27:19gets increased. From the below graph,
  751. 27:21you can see that when the size of the
  752. 27:23data is small, deep learning algorithm
  753. 27:25doesn't perform that well. But, why?
  754. 27:28Well, this is because deep learning
  755. 27:30algorithm needs a large amount of data
  756. 27:32to understand it perfectly.
  757. 27:34On the other hand, the machine learning
  758. 27:36algorithm can easily work with smaller
  759. 27:38data set. Fine?
  760. 27:40Next comes the hardware dependencies.
  761. 27:42Deep learning algorithms are heavily
  762. 27:44dependent on high-end machines, while
  763. 27:46the machine learning algorithm can work
  764. 27:48on low-end machines as well.
  765. 27:50This is because the requirement of deep
  766. 27:52learning algorithm include GPUs, which
  767. 27:55is an integral part of its working.
  768. 27:57The deep learning algorithm require GPUs
  769. 27:59as they do a large amount of matrix
  770. 28:01multiplication operations, and these
  771. 28:03operations can only be efficiently
  772. 28:06optimized using a GPU as it is built for
  773. 28:09this purpose only.
  774. 28:10Our third parameter will be feature
  775. 28:12engineering. Well, feature engineering
  776. 28:15is a process of putting the domain
  777. 28:17knowledge to reduce the complexity of
  778. 28:19the data and make patterns more visible
  779. 28:21to learning algorithms.
  780. 28:23This process is difficult and expensive
  781. 28:25in terms of time and expertise.
  782. 28:28In case of machine learning, most of the
  783. 28:29features are needed to be identified by
  784. 28:31an expert and then hand coded as per the
  785. 28:34domain and the data type. For example,
  786. 28:37the features can be a pixel value,
  787. 28:38shapes, texture, position, orientation,
  788. 28:41or anything. Fine? The performance of
  789. 28:44most of the machine learning algorithm
  790. 28:46depends on how accurately the features
  791. 28:48are identified and extracted.
  792. 28:50Whereas in case of deep learning
  793. 28:52algorithms, it try to learn high-level
  794. 28:54features from the data. This is a very
  795. 28:56distinctive part of deep learning, which
  796. 28:57makes it way ahead of traditional
  797. 28:59machine learning.
  798. 29:01Deep learning reduces the task of
  799. 29:03developing new feature extractor for
  800. 29:04every problem. Like in the case of CNN
  801. 29:07algorithm, it first try to learn the
  802. 29:09low-level features of the image, such as
  803. 29:11edges and lines, and then it proceeds to
  804. 29:13the parts of faces of people, and then
  805. 29:16finally to the high-level representation
  806. 29:17of the face. I hope the things are
  807. 29:19getting clear to you.
  808. 29:21So, let's move on ahead and see the next
  809. 29:23parameter. So, our next parameter is
  810. 29:25problem-solving approach.
  811. 29:27When we are solving a problem using
  812. 29:29traditional machine learning algorithm,
  813. 29:30it is generally recommended that we
  814. 29:33first break down the problem into
  815. 29:34different sub parts, solve them
  816. 29:36individually, and then finally combine
  817. 29:38them to get the desired result. This is
  818. 29:41how the machine learning algorithm
  819. 29:42handles the problem. On the other hand,
  820. 29:45the deep learning algorithm solves the
  821. 29:46problem from end to end.
  822. 29:48Let's take an example to understand
  823. 29:50this.
  824. 29:51Suppose you have a task of multiple
  825. 29:52object detection, and your task is to
  826. 29:54identify what is the object and where it
  827. 29:57is present in the image. So, let's see
  828. 29:59and compare how will you tackle this
  829. 30:01issue using the concept of machine
  830. 30:03learning and deep learning.
  831. 30:04Starting with machine learning, in a
  832. 30:06typical machine learning approach, you
  833. 30:08would first divide the problem into two
  834. 30:10step. First, object detection and then
  835. 30:13object recognition. First of all, you'd
  836. 30:16use a bounding box detection algorithm
  837. 30:18like GrabCut for example,
  838. 30:20to scan through the image and find out
  839. 30:22all the possible objects. Now, once the
  840. 30:25objects are recognized, you'd use object
  841. 30:27recognition algorithm like SVM with HOG,
  842. 30:31to recognize relevant objects.
  843. 30:33Now, finally when you combine the
  844. 30:35result, you would be able to identify
  845. 30:37what is the object and where it is
  846. 30:38present in the image.
  847. 30:40On the other hand, in deep learning
  848. 30:42approach, you would do the process from
  849. 30:44end to end. For example, in a YOLO net,
  850. 30:46which is a type of deep learning
  851. 30:48algorithm, you would pass an image and
  852. 30:50it would give out the location along
  853. 30:52with the name of the object. Now, let's
  854. 30:54move on to our fifth comparison
  855. 30:56parameter.
  856. 30:57It's execution time.
  857. 30:59Usually, a deep learning algorithm takes
  858. 31:01a long time to train. This is because
  859. 31:03there are so many parameter in a deep
  860. 31:05learning algorithm that makes the
  861. 31:06training longer than usual. The training
  862. 31:09might even last for 2 weeks or more than
  863. 31:11that if you're training completely from
  864. 31:13the scratch. Whereas in the case of
  865. 31:15machine learning, it relatively takes
  866. 31:17much less time to train, ranging from a
  867. 31:19few weeks to few hours.
  868. 31:21Now, the execution time is completely
  869. 31:23reversed when it comes to the testing of
  870. 31:25data. During testing, the deep learning
  871. 31:28algorithm takes much less time to run.
  872. 31:30Whereas if you compare it with a KNN
  873. 31:32algorithm, which is a type of machine
  874. 31:33learning algorithm, the test time
  875. 31:35increases as the size of the data
  876. 31:36increase.
  877. 31:38Last but not the least, we have
  878. 31:39interpretability as a factor for
  879. 31:41comparison of machine learning and deep
  880. 31:43learning. This factor is the main reason
  881. 31:46why deep learning is still thought 10
  882. 31:48times before anyone uses it in the
  883. 31:50industry. Let's take an example. Suppose
  884. 31:53we use deep learning to give automated
  885. 31:56scoring to essays. The performance it
  886. 31:58gives in scoring is quite excellent and
  887. 32:00is near to the human performance. But
  888. 32:02there's an issue with it. It does not
  889. 32:04reveal why it has given that score.
  890. 32:06Indeed, mathematically it is possible to
  891. 32:09find out that which node of a deep
  892. 32:11neural network were activated, but we
  893. 32:13don't know what the neurons are supposed
  894. 32:15to model and what these layers of neuron
  895. 32:17were doing collectively.
  896. 32:19So, we failed to interpret the result.
  897. 32:21On the other hand, machine learning
  898. 32:22algorithm like decision tree gives us a
  899. 32:25crisp rule for why it chose and what it
  900. 32:27chose. So, it is particularly easy to
  901. 32:30interpret the reasoning behind it.
  902. 32:32Therefore, the algorithms like decision
  903. 32:33tree and linear or logistic regression
  904. 32:36are primarily used in industry for
  905. 32:38interpretability.
  906. 32:45So, why exactly are we using Python for
  907. 32:47artificial intelligence? Why aren't we
  908. 32:49using any other language? Right? Now,
  909. 32:52there are a couple of reasons as to why
  910. 32:54Python is so popular when it comes to
  911. 32:56AI, machine learning, and deep learning.
  912. 32:58The first reason is less coding is
  913. 33:00required. Now, artificial intelligence
  914. 33:02has a lot of algorithms. If you have to
  915. 33:04implement AI in any code or in any
  916. 33:07problem, then there are going to be tons
  917. 33:09and tons of machine learning algorithms
  918. 33:11involved, deep learning algorithms
  919. 33:13involved, right? Now, testing all of
  920. 33:15these can become a very tiresome task.
  921. 33:18That's where Python usually comes in
  922. 33:20handy. Now, the language has something
  923. 33:23known as check as you code methodology,
  924. 33:26which eases the process of testing,
  925. 33:28right? You can check your program as you
  926. 33:30code it. Basically, as you're typing
  927. 33:32each sentence, your errors or your any
  928. 33:34sort of mistakes in your code will be
  929. 33:36given to you. Right? So, testing becomes
  930. 33:38much easier when it comes to Python.
  931. 33:40The next important reason why we're
  932. 33:42choosing Python is it has support for
  933. 33:44pre-built libraries. Right? Python is
  934. 33:47very convenient for AI developers
  935. 33:50because all of the algorithms, machine
  936. 33:52learning algorithms, and deep learning
  937. 33:53algorithms are already predefined in
  938. 33:56libraries, right? So, you don't have to
  939. 33:57actually sit down and code each and
  940. 33:59every algorithm. That would take a lot
  941. 34:01of time. And that's a very
  942. 34:03time-consuming task. And thanks to
  943. 34:05Python, you don't have to do that
  944. 34:06because they have libraries and packages
  945. 34:09that have all the algorithms built in
  946. 34:11them, right? So, if you want to run any
  947. 34:14algorithm, all you have to do is you
  948. 34:15have to call the function and load the
  949. 34:17library. That's all. It's as simple as
  950. 34:19that. Now, the next reason is ease of
  951. 34:21learning. So, guys, Python is actually
  952. 34:24the most simplest programming language,
  953. 34:26right? If you ask me, I think it is is
  954. 34:28the most easiest programming language.
  955. 34:30It's very similar to English language,
  956. 34:32right? If you read a couple of lines in
  957. 34:34Python, you'll understand what exactly
  958. 34:36the code is doing. It has a very simple
  959. 34:38syntax, and this simple syntax can be
  960. 34:41implemented to solve simple problems
  961. 34:43like addition of two strings, and it can
  962. 34:45also be used to solve complex problems
  963. 34:48like building machine learning models
  964. 34:50and deep learning models. So, ease of
  965. 34:52learning is a major factor when it comes
  966. 34:54to why Python is chosen for artificial
  967. 34:57intelligence, right? Next, we have
  968. 34:59platform independent.
  969. 35:01So, good thing about Python is that you
  970. 35:02can get your project running on
  971. 35:04different operating systems, right? And
  972. 35:07what happens when you transfer your code
  973. 35:09from one operating system to another
  974. 35:11operating system is we find a lot of
  975. 35:13dependency issues. To solve that, Python
  976. 35:16has a couple of packages such as there
  977. 35:18is a package known as PyInstaller,
  978. 35:20right? This PyInstaller will take care
  979. 35:22of all the dependency issues when you're
  980. 35:24transferring your code from one platform
  981. 35:26to the other platform. So, all of this
  982. 35:29support is provided by Python. The last
  983. 35:31reason is massive community support.
  984. 35:33This is a very important point because
  985. 35:36it is important that you have a large
  986. 35:38community that will help you out with
  987. 35:40any errors or with any sort of problems
  988. 35:43in your code, right? So, Python has
  989. 35:46several communities and several forums
  990. 35:48and groups on Facebook. So, if you have
  991. 35:50any doubts regarding any error, you can
  992. 35:53just post those errors in these groups,
  993. 35:55and you'll have like a bunch of people
  994. 35:56helping you out. Right? So, guys, these
  995. 35:59are a couple of reasons as to why Python
  996. 36:01is chosen for artificial intelligence.
  997. 36:03It's actually considered the most
  998. 36:04popular and the most used language for
  999. 36:07data science, AI, machine learning, and
  1000. 36:09deep learning. To prove that to you,
  1001. 36:11here is a stat from Stack Overflow.
  1002. 36:13Stack Overflow recently stated that
  1003. 36:16Python is the fastest growing
  1004. 36:17programming language. If you look at the
  1005. 36:19graph, you can see that it has taken
  1006. 36:21over JavaScript and Java and C#, C++,
  1007. 36:25and PHP, right? So, Python is actually
  1008. 36:28growing at an exponential rate,
  1009. 36:30especially when it comes to data science
  1010. 36:32and artificial intelligence. A lot of
  1011. 36:34developers are very comfortable with the
  1012. 36:35Python language because, you know, it's
  1013. 36:37a general-purpose language, first of
  1014. 36:38all. So, most of the developers are
  1015. 36:40already aware of Python. And then, using
  1016. 36:43the same language in order to solve
  1017. 36:45complex problems like artificial
  1018. 36:47intelligence, machine learning, and deep
  1019. 36:49learning is something every developer
  1020. 36:51wants, right? They want a simple
  1021. 36:52language in order to code all the
  1022. 36:54complex algorithms or the complex
  1023. 36:56models. Right? So, that's why Python is
  1024. 36:59the best choice for artificial
  1025. 37:00intelligence.
  1026. 37:02For those of you who are not aware of
  1027. 37:03Python programming and don't know much
  1028. 37:05about Python, I'm going to leave a
  1029. 37:07couple of links in the description box.
  1030. 37:09Right? You can go through those links
  1031. 37:11and study a little bit more about how
  1032. 37:12Python works or how the coding part
  1033. 37:15works. Right? I'm going to be focusing
  1034. 37:17mainly on artificial intelligence, and
  1035. 37:19I'll be showing you a lot of demos. So,
  1036. 37:21those of you are not aware of Python,
  1037. 37:23make sure you check the description box.
  1038. 37:25Right?
  1039. 37:26Next, I'm going to discuss the different
  1040. 37:28Python packages for artificial
  1041. 37:30intelligence. Now, these are the
  1042. 37:31packages that are specifically for
  1043. 37:33machine learning, deep learning, natural
  1044. 37:35language processing, and so on. So,
  1045. 37:37let's take a look at all these packages.
  1046. 37:39So, first, we have TensorFlow. If you
  1047. 37:42are currently working on a machine
  1048. 37:44learning project in Python, then you
  1049. 37:46must have heard of this popular
  1050. 37:48open-source library known as TensorFlow.
  1051. 37:51Right? This library was developed by
  1052. 37:52Google in collaboration with Brain team.
  1053. 37:55TensorFlow is used in almost every
  1054. 37:57Google application for machine learning.
  1055. 37:59Now, let me just discuss a few features
  1056. 38:01of TensorFlow. It has a responsive
  1057. 38:03construct, meaning that with TensorFlow,
  1058. 38:06we can easily visualize each and every
  1059. 38:08part of the graph, which is not an
  1060. 38:10option when you're using other packages
  1061. 38:12such as NumPy or scikit. Right? Another
  1062. 38:15feature is that it's very flexible. Now,
  1063. 38:17one of the most important TensorFlow
  1064. 38:19features is that it is flexible in
  1065. 38:22operability. Meaning that it has
  1066. 38:24modularity and the parts of which you
  1067. 38:27want to make standalone, it offers you
  1068. 38:29that option. Right? It's very flexible
  1069. 38:31in that way. It'll give you exactly what
  1070. 38:33you want. Now, good feature about
  1071. 38:35TensorFlow is that you can train it on
  1072. 38:37both CPU and GPU. Right? So, for
  1073. 38:39distributed computing, you can have both
  1074. 38:42these options. Also, it supports
  1075. 38:44parallel neural network training. So,
  1076. 38:46TensorFlow offers pipelining in the
  1077. 38:49sense that you can train multiple neural
  1078. 38:51networks and multiple GPUs, which makes
  1079. 38:54the models very efficient on any
  1080. 38:56large-scale system. Right? So, parallel
  1081. 38:58neural network training is supported by
  1082. 39:00TensorFlow. Right? This is one of the
  1083. 39:03most important features of TensorFlow.
  1084. 39:05Apart from this, it has a very large
  1085. 39:07community. And needless to say, if it
  1086. 39:09has been developed by Google, then
  1087. 39:11there's already a large team of software
  1088. 39:13engineers who work on stability,
  1089. 39:16improvements, and all of that. Right?
  1090. 39:19The next library I'm going to talk about
  1091. 39:20is scikit-learn.
  1092. 39:22Now, scikit-learn is a Python library
  1093. 39:24that is associated with NumPy and SciPy.
  1094. 39:26Right? That's why it has the name
  1095. 39:28scikit-learn. Now, this is considered to
  1096. 39:30be one of the best uh libraries for
  1097. 39:32working with complex data. And there are
  1098. 39:34a lot of changes that are being made in
  1099. 39:36this library. And one modification is
  1100. 39:39the cross-validation feature, which
  1101. 39:41provides the ability to use more than
  1102. 39:43one metric. Right? Cross-validation is
  1103. 39:46one of the most important and one of the
  1104. 39:47most easiest methods for checking the
  1105. 39:50accuracy of a model. Right? So,
  1106. 39:51cross-validation is being implemented in
  1107. 39:53scikit-learn. And apart from that,
  1108. 39:56again, there are a large spread of
  1109. 39:57algorithms that you can implement by
  1110. 39:59using scikit-learn. Right? These include
  1111. 40:01unsupervised learning algorithms,
  1112. 40:03starting from clustering, factor
  1113. 40:05analysis, principal component analysis,
  1114. 40:08to all the unsupervised neural networks.
  1115. 40:11Scikit-learn is also very essential uh
  1116. 40:13for feature extracting in images and
  1117. 40:16text.
  1118. 40:17So, mainly scikit-learn is used for
  1119. 40:19implementing all the standard machine
  1120. 40:21learning and data mining tasks like
  1121. 40:24reducing dimensionality, classification,
  1122. 40:26regression, clustering, and model
  1123. 40:28selection. Next up, we have NumPy. Now,
  1124. 40:30NumPy is considered as one of the most
  1125. 40:33popular machine learning libraries in
  1126. 40:35Python. Now, let me tell you that
  1127. 40:36TensorFlow and other libraries, they
  1128. 40:38make use of NumPy internally for
  1129. 40:41performing multiple operations on
  1130. 40:43tensors. The most important feature of
  1131. 40:47NumPy is the array interface. It
  1132. 40:49supports multi-dimensional arrays.
  1133. 40:51Right? That's one of the most important
  1134. 40:53features of NumPy. Another feature is uh
  1135. 40:55it makes complex mathematical
  1136. 40:57implementations very simple. Right? It's
  1137. 41:00mainly known for computing mathematical
  1138. 41:03data. So, NumPy is a package that you
  1139. 41:05should be using for any sort of
  1140. 41:07statistical analysis or data analysis
  1141. 41:10that involves a lot of math. Apart from
  1142. 41:12that, it makes coding very easy and
  1143. 41:15grasping the concept is extremely easy
  1144. 41:16with NumPy. Now, NumPy is mainly used
  1145. 41:20for expressing images, sound waves, and
  1146. 41:22other mathematical computations.
  1147. 41:24All right? Moving on to our next
  1148. 41:26library, we have Theano. Theano is a
  1149. 41:29computational framework which is used
  1150. 41:32for computing multi-dimensional arrays.
  1151. 41:34Right? Theano actually works very
  1152. 41:36similar to TensorFlow, but the only
  1153. 41:38drawback is that you can't fit Theano
  1154. 41:41into production environments. But apart
  1155. 41:43from that, Theano allows you to define,
  1156. 41:46optimize, and evaluate mathematical
  1157. 41:48expressions that involve
  1158. 41:49multi-dimensional arrays. Right? This is
  1159. 41:52another library that lets you implement
  1160. 41:54multi-dimensional arrays. Features of
  1161. 41:56Theano include tight integration with
  1162. 41:58NumPy. An advantage of Theano is that
  1163. 42:01you can easily implement NumPy arrays in
  1164. 42:03Theano. Right? That's why there's a
  1165. 42:05connection between Theano and NumPy
  1166. 42:07because both of them effectively use
  1167. 42:09multi-dimensional arrays. Transparent
  1168. 42:11use of GPU. Now, performing data
  1169. 42:14intensive computations are much faster
  1170. 42:16when it comes to uh Theano because of
  1171. 42:18its use of GPU, right? Theano also lets
  1172. 42:21you detect and diagnose multiple types
  1173. 42:24of errors and any sort of ambiguity in
  1174. 42:27the model. So, guys, Theano was actually
  1175. 42:29designed to handle the types of
  1176. 42:31computations required for large neural
  1177. 42:34network algorithms, right? It was mainly
  1178. 42:36built for deep learning and neural
  1179. 42:38networks. It was one of the first
  1180. 42:41libraries of its kind and it is
  1181. 42:43considered as an industry standard for
  1182. 42:46deep learning research and development.
  1183. 42:48Theano is being used in multiple neural
  1184. 42:50networks projects and the popularity of
  1185. 42:52Theano is only going to grow with time,
  1186. 42:54right? A lot of people actually haven't
  1187. 42:56heard of Theano, but let me tell you
  1188. 42:58that this is one of the best ways to
  1189. 42:59implement deep learning and neural
  1190. 43:01network models.
  1191. 43:03Moving on, uh we have Keras. Now, Keras
  1192. 43:05is considered to be the most popular
  1193. 43:08Python package. It provides some of the
  1194. 43:10best functionalities for compiling
  1195. 43:12models, processing your data sets, and
  1196. 43:15visualizing graphs. It is also popular
  1197. 43:17in the implementation of neural
  1198. 43:19networks, right? It is considered to be
  1199. 43:21the simplest package uh with which you
  1200. 43:23can implement neural networks. In fact,
  1201. 43:25in our today's demo for deep learning,
  1202. 43:27we'll be implementing Keras in order to
  1203. 43:29understand how neural networks work. Few
  1204. 43:32of the features of Keras include that it
  1205. 43:34runs very smoothly on both CPU and GPU.
  1206. 43:37It supports almost all the models of the
  1207. 43:40neural network, right? From fully
  1208. 43:42connected, convolutional, pooling,
  1209. 43:44recurrent, embedding, all of these
  1210. 43:46models are supported by Keras.
  1211. 43:48And not only that, you can combine these
  1212. 43:50models to build more complex models.
  1213. 43:53Keras is completely Python-based, which
  1214. 43:55makes it very easy to debug and explore,
  1215. 43:58right? Since Python has a huge community
  1216. 44:00of followers, it's very simple in order
  1217. 44:03to debug any sort of error that you find
  1218. 44:05while implementing Keras. So, the
  1219. 44:07libraries that I discussed so far were
  1220. 44:09dedicated to machine learning and deep
  1221. 44:11learning. For natural language
  1222. 44:13processing, we have the most famous
  1223. 44:15library known as the Natural Language
  1224. 44:17Toolkit, which is an open-source Python
  1225. 44:19library, mainly used for natural
  1226. 44:21language processing, text analysis, and
  1227. 44:24text mining. The main features include
  1228. 44:26that it studies and analyzes natural
  1229. 44:28language text in order to draw useful
  1230. 44:31information from all this natural
  1231. 44:32language text. It performs text analysis
  1232. 44:36and sentimental analysis by performing
  1233. 44:38tasks such as stemming, lemmatization,
  1234. 44:41tokenization, and so on. Now, don't
  1235. 44:44worry if you don't know what any of
  1236. 44:45those terms mean. I'll be discussing all
  1237. 44:47of those terms with you by the end of
  1238. 44:49today's session. So guys, these were a
  1239. 44:51couple of Python-based libraries, which
  1240. 44:53are very essential for implementing
  1241. 44:55machine learning and deep learning and
  1242. 44:58artificial intelligence when you're
  1243. 45:00using Python, right? These libraries are
  1244. 45:02perfect for implementing AI. So guys, if
  1245. 45:05any of you have any doubts regarding the
  1246. 45:07libraries or if you want to learn more
  1247. 45:09about the libraries, I will leave a
  1248. 45:11couple of links in the description box.
  1249. 45:13You can go through those videos as well.
  1250. 45:15So now, let's move on to the main topic
  1251. 45:17of discussion, which is artificial
  1252. 45:18intelligence.
  1253. 45:20Now, before we get started with the
  1254. 45:22demand of artificial intelligence, let
  1255. 45:25me tell you that AI was invented long
  1256. 45:27ago. AI goes back to the 19th century.
  1257. 45:29It was not something that was recently
  1258. 45:31invented, even though AI has recently
  1259. 45:33gained a lot of popularity. We can say
  1260. 45:36that in the past decade, AI has gained
  1261. 45:39the maximum popularity. But, it was
  1262. 45:41actually invented in the 19th century.
  1263. 45:44Now, especially in the year 1950, there
  1264. 45:46was somebody known as Alan Turing. I'm
  1265. 45:48sure a lot of you have heard about the
  1266. 45:50Turing test. The Turing test is
  1267. 45:52basically used to determine whether or
  1268. 45:54not a machine is artificially
  1269. 45:57intelligent, meaning that whether a
  1270. 45:59machine can think intelligently like a
  1271. 46:01human being.
  1272. 46:02Right? This was the first proposition
  1273. 46:04and this was one of the most important
  1274. 46:07breakthroughs in artificial
  1275. 46:08intelligence. Right? Somebody known as
  1276. 46:10Alan Turing, he published a landmark
  1277. 46:13paper in which he speculated about the
  1278. 46:15possibility of creating machines that
  1279. 46:17think. Right? So, the Turing test was
  1280. 46:20the first serious proposal in the
  1281. 46:22philosophy of artificial intelligence.
  1282. 46:24This was done in 1950. Right? After
  1283. 46:27this, we had eras of AI. We had the game
  1284. 46:30AI which was in 1951. Now, since the
  1285. 46:33emergence of AI in 1950s, we have seen
  1286. 46:37an exponential growth in its potential.
  1287. 46:39Right? AI covers domains like machine
  1288. 46:41learning, deep learning, neural
  1289. 46:43networks, natural language processing,
  1290. 46:45knowledge base, and so on. It's also
  1291. 46:47made its way into computer vision and
  1292. 46:49image processing. But, the question is
  1293. 46:51if AI has been here for over half a
  1294. 46:54century,
  1295. 46:55why has it suddenly gained so much
  1296. 46:57importance? Right? Why are we talking
  1297. 47:00about artificial intelligence now? The
  1298. 47:02main reasons for the vast popularity of
  1299. 47:04AI are the following. Right? The first
  1300. 47:06reason is more computational power. Now,
  1301. 47:09AI requires a lot of computing power.
  1302. 47:12Recently, many advances have been made
  1303. 47:14and complex deep learning models can be
  1304. 47:16deployed. And one of the greatest
  1305. 47:18technology that made this possible are
  1306. 47:20GPUs. Since the invention of GPUs, we
  1307. 47:23can compute much more with our
  1308. 47:25computers. Initially, we could barely
  1309. 47:27process 1 GB of data. Right? We only had
  1310. 47:30hard disk to store additional memory and
  1311. 47:32all of that. Now, our computers can
  1312. 47:34process tons and tons of data. So, now
  1313. 47:37we have more computational power, which
  1314. 47:39is one of the main reasons behind why AI
  1315. 47:41became so popular. So, by having more
  1316. 47:44computational power, it becomes much
  1317. 47:46easier to implement artificial
  1318. 47:47intelligence. Next reason is more data.
  1319. 47:50Now, big data is one of the most
  1320. 47:52important reasons behind the development
  1321. 47:55of artificial intelligence. Now, AI and
  1322. 47:58data science and machine learning, deep
  1323. 48:00learning, all of these processes are
  1324. 48:03here only because we have a lot of data
  1325. 48:05at present. Now, the main idea behind
  1326. 48:08all these technologies is to draw useful
  1327. 48:10insights from data. Now, since we start
  1328. 48:12generating a lot of data, we need to
  1329. 48:14find a method that can process this much
  1330. 48:17data and draw useful insights from data
  1331. 48:20such that it benefits an organization or
  1332. 48:23it grows a business. That's why
  1333. 48:25artificial intelligence and machine
  1334. 48:27learning comes into the picture. Right?
  1335. 48:28So, more data led to the demand of
  1336. 48:31artificial intelligence. Apart from
  1337. 48:33this, we also have better algorithms
  1338. 48:35now, right? We have state-of-the-art
  1339. 48:37algorithms. Most of them are based on
  1340. 48:40the idea of neural networks and these
  1341. 48:42are constantly getting better. Neural
  1342. 48:43networks are actually one of the most
  1343. 48:45significant discoveries in artificial
  1344. 48:48intelligence because with neural
  1345. 48:50networks, you can take in thousand
  1346. 48:52layers of input data. Right? You can
  1347. 48:54take in a lot of input data to perform
  1348. 48:56computations. So, through neural
  1349. 48:58networks, we are actually able to solve
  1350. 49:00a lot of problems including healthcare
  1351. 49:02problems, fraud detection problems, and
  1352. 49:04so on. Another reason is broad
  1353. 49:06investment. So, our universities and
  1354. 49:09governments and startups and any tech
  1355. 49:12giants like Google, Amazon, and
  1356. 49:14Facebook, they are all investing heavily
  1357. 49:16in artificial intelligence, which also
  1358. 49:18led to the demand of AI. So, AI is
  1359. 49:21rapidly growing both as a field of study
  1360. 49:23and also as an economy. Right? It's
  1361. 49:26adding a lot to the economy and I think
  1362. 49:29this is the perfect time for you to get
  1363. 49:30into the field of artificial
  1364. 49:31intelligence because right now AI is in
  1365. 49:34a really high demand. AI, machine
  1366. 49:36learning, data science, all of this are
  1367. 49:38of really high demand at present. All
  1368. 49:41right. So, this is the perfect time for
  1369. 49:42you to get started with artificial
  1370. 49:43intelligence. Now, let me tell you that
  1371. 49:45the term artificial intelligence was
  1372. 49:47first coined in the year 1956
  1373. 49:50by a scientist known as John McCarthy.
  1374. 49:53Now, John McCarthy defined artificial
  1375. 49:56intelligence as the science and
  1376. 49:57engineering of making intelligent
  1377. 50:00machines. So, now let's move on and talk
  1378. 50:02about how artificial intelligence is
  1379. 50:05different from machine learning and deep
  1380. 50:06learning. A lot of people uh tend to
  1381. 50:08assume that artificial intelligence,
  1382. 50:10machine learning, and deep learning are
  1383. 50:12the same because they have common
  1384. 50:14applications, right? For example, Siri
  1385. 50:17is an application of AI, machine
  1386. 50:19learning, and deep learning. So, how are
  1387. 50:22these technologies uh related, right? Or
  1388. 50:24how are they different from each other?
  1389. 50:26Now, artificial intelligence is the
  1390. 50:28science of getting machines to mimic the
  1391. 50:30behavior of human beings. Machine
  1392. 50:33learning is the subset of artificial
  1393. 50:36intelligence that focuses on getting
  1394. 50:38machines to make decisions by feeding
  1395. 50:40them data. Deep learning, on the other
  1396. 50:43hand, is a subset of machine learning
  1397. 50:45that uses the concept of neural networks
  1398. 50:48to solve complex problems.
  1399. 50:50So, to sum it up to you, artificial
  1400. 50:52intelligence, machine learning, and deep
  1401. 50:54learning are heavily interconnected
  1402. 50:56fields, right? Machine learning and deep
  1403. 50:58learning aids artificial intelligence by
  1404. 51:00providing a set of algorithms and neural
  1405. 51:03networks to solve data-driven problems.
  1406. 51:05However, AI is not restricted to only
  1407. 51:08machine learning and deep learning,
  1408. 51:09right? It covers a vast domain of fields
  1409. 51:12which include natural language
  1410. 51:13processing, object detection, computer
  1411. 51:15vision, robotics, expert systems, and so
  1412. 51:18on, right? So, AI is a very vast field.
  1413. 51:20Guys, I hope I cleared the difference
  1414. 51:22between AI, machine learning, and deep
  1415. 51:24learning. Also, a lot of you might be
  1416. 51:26confused about data science. Data
  1417. 51:28science is now an umbrella term, right?
  1418. 51:31Data science basically means to derive
  1419. 51:33useful insights from data. So, data
  1420. 51:35science actually uh uses AI, machine
  1421. 51:38learning, and deep learning, right? So,
  1422. 51:40it implements all of these three
  1423. 51:42technologies in order to derive useful
  1424. 51:44insights from data, right? Now, let's
  1425. 51:47move on to the most interesting topic in
  1426. 51:49artificial intelligence, which is
  1427. 51:51machine learning. Now guys, the term
  1428. 51:53machine learning was first coined by a
  1429. 51:55scientist known as Arthur Samuel in the
  1430. 51:58year 1959.
  1431. 51:59Looking back, that year was probably the
  1432. 52:02most significant in terms of
  1433. 52:03technological advancements.
  1434. 52:05In order to define machine learning, if
  1435. 52:08you browse the internet for what is
  1436. 52:10machine learning, you'll get at least
  1437. 52:11100 different definitions.
  1438. 52:13In simple terms, machine learning is a
  1439. 52:16subset of artificial intelligence, which
  1440. 52:18provides machines the ability to learn
  1441. 52:21automatically and improve from
  1442. 52:23experience without being explicitly
  1443. 52:26programmed to do so. In a sense, it is
  1444. 52:29the practice of getting machines to
  1445. 52:31solve problems by gaining the ability to
  1446. 52:33think. Now the question here is, can a
  1447. 52:36machine think or can a machine make
  1448. 52:38decisions? Well, if you feed a machine a
  1449. 52:41good amount of data, it will learn how
  1450. 52:43to interpret, process, and analyze this
  1451. 52:46data by using something known as machine
  1452. 52:48learning algorithms. To give you a basic
  1453. 52:51idea of how the machine learning process
  1454. 52:53works, look at the figure on this slide.
  1455. 52:56A machine learning process always begins
  1456. 52:58by feeding the machine lots and lots of
  1457. 53:00data. Now by using this data, the
  1458. 53:03machine is trained to detect any hidden
  1459. 53:05insights and trends in the data. These
  1460. 53:08insights are then used to build a
  1461. 53:10machine learning model by using a
  1462. 53:12machine learning algorithm in order to
  1463. 53:14solve a problem. The basic aim of
  1464. 53:17machine learning is to solve a problem
  1465. 53:19or find a solution by using data. Now
  1466. 53:22moving ahead, I'll be discussing the
  1467. 53:23machine learning process in depth,
  1468. 53:25right? So don't worry if you haven't got
  1469. 53:27the exact idea of what machine learning
  1470. 53:29is.
  1471. 53:30Now the machine learning process
  1472. 53:32involves building a predictive model
  1473. 53:34that can be used to find a solution for
  1474. 53:37a particular problem. A well-defined
  1475. 53:39machine learning process will have
  1476. 53:41around seven steps. It always begins
  1477. 53:44with defining the objective followed by
  1478. 53:46data gathering or data collection. Then
  1479. 53:49we have something known as preparing
  1480. 53:51data, which is also called data
  1481. 53:53pre-processing. Then we have data
  1482. 53:55exploration or exploratory data
  1483. 53:57analysis. This is followed by building a
  1484. 54:00machine learning model.
  1485. 54:02Then we have model evaluation and
  1486. 54:04finally predictions. This is how the
  1487. 54:06process of machine learning works. To
  1488. 54:08understand the machine learning process,
  1489. 54:10let's assume that you've been given a
  1490. 54:12problem that needs to be solved by using
  1491. 54:14machine learning. Let's say that the
  1492. 54:16problem is to predict the occurrence of
  1493. 54:18rain in your local area by using machine
  1494. 54:21learning. Now, the first step is to
  1495. 54:23define the objective of the problem.
  1496. 54:25Right? At this step we must understand
  1497. 54:27what exactly needs to be predicted. In
  1498. 54:29our case, the objective is to predict
  1499. 54:31the possibility of rain by studying the
  1500. 54:34weather conditions. So, at this stage it
  1501. 54:36is essential to take mental notes on
  1502. 54:39what kind of data can be used to solve
  1503. 54:41this problem or the type of approach
  1504. 54:43that you must follow to get to the
  1505. 54:45solution.
  1506. 54:46The questions you should be asking
  1507. 54:48yourself is what are we trying to
  1508. 54:50predict? Right? Here we're trying to
  1509. 54:51predict whether it'll rain or not.
  1510. 54:54Right? You need to understand what are
  1511. 54:56the target features. Target features are
  1512. 54:59basically the variable that you need to
  1513. 55:01predict. Here we need to predict a
  1514. 55:03variable that'll show us whether it's
  1515. 55:04going to rain tomorrow or not. Then you
  1516. 55:07must also understand what kind of data
  1517. 55:09you'll need to solve this problem. Apart
  1518. 55:11from that, you need to know what kind of
  1519. 55:13problem you're facing. Is it a binary
  1520. 55:15classification problem or is it a
  1521. 55:17clustering problem? Now, if you don't
  1522. 55:19know what classification and clustering
  1523. 55:21is, don't worry. I'll be talking about
  1524. 55:23all of these things in the upcoming
  1525. 55:24slides. So, your first step is to define
  1526. 55:27the objective of your problem. You need
  1527. 55:30to understand what exactly needs to be
  1528. 55:32done here. Right? How can you solve this
  1529. 55:34problem?
  1530. 55:35Moving on, your next step is to gather
  1531. 55:37the data that you need. At this stage,
  1532. 55:39you must be asking questions such as
  1533. 55:42what kind of data is needed to solve
  1534. 55:44this problem. Is the data available to
  1535. 55:46me? And if it's not available, how can I
  1536. 55:48get the data? Right? Once you know the
  1537. 55:51type of data that is required, you must
  1538. 55:53understand how you can derive this data.
  1539. 55:56Data collection can be either done
  1540. 55:58manually or it can be done by web
  1541. 56:00scraping. But don't worry if you're a
  1542. 56:02beginner and you're just looking to
  1543. 56:04learn machine learning, you don't have
  1544. 56:05to worry about getting the data.
  1545. 56:08There are thousands of data resources on
  1546. 56:10the web. You can just download the data
  1547. 56:12set and you can get going.
  1548. 56:14Coming back to the problem at hand, the
  1549. 56:16data needed for weather forecasting
  1550. 56:18includes measures such as humidity
  1551. 56:20level, your temperature, the pressure,
  1552. 56:23the locality, whether or not you live in
  1553. 56:26a hill station, and so on.
  1554. 56:28Such data must be collected and it has
  1555. 56:30to be stored for analysis. This is where
  1556. 56:33you collect all the data. Now, moving on
  1557. 56:35to step number three is data
  1558. 56:37preparation. The data that you collected
  1559. 56:40is almost never in the right format. All
  1560. 56:43right, even if you collect it from a
  1561. 56:45internet resource, if you download it
  1562. 56:47from some website, even then your data
  1563. 56:50is not going to be clean. Right? It's
  1564. 56:52not going to be in the correct format.
  1565. 56:53There's always going to be some sort of
  1566. 56:55inconsistencies in your data.
  1567. 56:58Inconsistencies include any missing
  1568. 57:00values or any redundant variables,
  1569. 57:03duplicate values. All of these are
  1570. 57:05inconsistencies.
  1571. 57:07Removing all of this is very essential
  1572. 57:09because they might lead to any wrongful
  1573. 57:11computation. Therefore, at this stage,
  1574. 57:13you can scan the entire data set for any
  1575. 57:16missing values and you have to fix them
  1576. 57:18here itself.
  1577. 57:20Now, actually, this is one of the most
  1578. 57:21time-consuming steps in a machine
  1579. 57:23learning process. If you ask a data
  1580. 57:26scientist which step he hates the most
  1581. 57:28or which step is, you know, the most
  1582. 57:30time-consuming, they're probably going
  1583. 57:32to tell you data processing and data
  1584. 57:33cleaning. Right? It's one of the most
  1585. 57:36tiresome task because you need to look
  1586. 57:38at all the values that are there. You
  1587. 57:39need to find any missing values, any
  1588. 57:41data that is not relevant to you. Right?
  1589. 57:43All of this has to be removed so that
  1590. 57:45you can analyze the data in a better
  1591. 57:47way.
  1592. 57:48Now, step number four is exploratory
  1593. 57:50data analysis.
  1594. 57:52So, guys, this stage is all about
  1595. 57:54getting deep into your data and finding
  1596. 57:57all the hidden data mysteries.
  1597. 58:00EDA or exploratory data analysis is like
  1598. 58:03the brainstorming stage of machine
  1599. 58:05learning.
  1600. 58:06Data exploration involves understanding
  1601. 58:08the patterns and the trends in your
  1602. 58:09data. So, at this stage all the useful
  1603. 58:12insights are drawn and any correlations
  1604. 58:15between the variables are understood.
  1605. 58:17For example, in the case of predicting
  1606. 58:19rainfall, we know that there is a strong
  1607. 58:22possibility of rain if the temperature
  1608. 58:24has fallen low. Such correlations have
  1609. 58:27to be understood and mapped at this
  1610. 58:29stage.
  1611. 58:30EDA is actually the most important step
  1612. 58:32in a machine learning process because
  1613. 58:34here is where you understand your data.
  1614. 58:37You understand how your data is going to
  1615. 58:39help you predict the outcome.
  1616. 58:41Moving on to step number five, we have
  1617. 58:44building a machine learning model. So,
  1618. 58:46all the insights and all the patterns
  1619. 58:49that you got from your data exploration
  1620. 58:51stage, those insights are used to build
  1621. 58:53the machine learning model. So, this
  1622. 58:55stage always begins by splitting the
  1623. 58:57data set into two parts, that is
  1624. 59:00training and testing data. Now, remember
  1625. 59:02that the training data will be used to
  1626. 59:05build and analyze the model.
  1627. 59:07The model is basically the machine
  1628. 59:09learning algorithm that predicts the
  1629. 59:11output by using the data that you feed
  1630. 59:13to it. An example of machine learning
  1631. 59:15algorithm is logistic regression and
  1632. 59:18linear regression. All of these are
  1633. 59:20machine learning algorithms.
  1634. 59:22Now, don't worry about choosing the
  1635. 59:23right algorithm. Right? First, we'll
  1636. 59:25focus on what the machine learning
  1637. 59:26process is.
  1638. 59:28But anyway, choosing the right algorithm
  1639. 59:30will depend on several factors, right?
  1640. 59:32It depends on the type of problem you're
  1641. 59:34trying to solve, the data set, and the
  1642. 59:36level of complexity of the problem.
  1643. 59:39In the upcoming sections, we'll discuss
  1644. 59:40all the different types of problems that
  1645. 59:42can be solved by using machine learning.
  1646. 59:44Moving on to step number six, we have
  1647. 59:46model evaluation and optimization.
  1648. 59:49Now, after you build a model by using
  1649. 59:51the training data set, it is finally
  1650. 59:53time to put the model to a test. The
  1651. 59:56testing data set is used to check the
  1652. 59:58efficiency of the model and how
  1653. 1:00:00accurately it can predict the outcome.
  1654. 1:00:03Now, once the accuracy is calculated and
  1655. 1:00:06any further improvements in the model,
  1656. 1:00:08they have to be implemented at this
  1657. 1:00:10stage. Methods like parameter tuning and
  1658. 1:00:13cross-validation can be used to improve
  1659. 1:00:15the performance of the model.
  1660. 1:00:17Before I move any further, I don't know
  1661. 1:00:19if all of you know what training and
  1662. 1:00:21testing data set means. In machine
  1663. 1:00:23learning, the input data is always
  1664. 1:00:25divided into two sets. We have something
  1665. 1:00:28known as the training data set, and we
  1666. 1:00:29have something known as the testing data
  1667. 1:00:31set.
  1668. 1:00:32So, in machine learning, you always
  1669. 1:00:34split the data into two parts, right?
  1670. 1:00:36This process is known as a data
  1671. 1:00:38splicing. Now, the training data set
  1672. 1:00:40will be used to build the machine
  1673. 1:00:42learning model, and the testing data set
  1674. 1:00:44will be used to test the efficiency of
  1675. 1:00:47the model that you built. This is what
  1676. 1:00:49training and testing data set is.
  1677. 1:00:51They're not any different data that you
  1678. 1:00:53derive. They're the same as the input
  1679. 1:00:55data set. The only thing is you are
  1680. 1:00:57splitting the data set so that you can
  1681. 1:00:59train the model on one data and test the
  1682. 1:01:01model on another data.
  1683. 1:01:03Now, remember that the training data set
  1684. 1:01:05is always larger in size when compared
  1685. 1:01:08to the testing data set. Because
  1686. 1:01:10obviously, you are training and building
  1687. 1:01:12the model by using the training data
  1688. 1:01:13set. The testing data set is just for
  1689. 1:01:16evaluating the performance of your
  1690. 1:01:18model.
  1691. 1:01:19Now, let's move on and understand step
  1692. 1:01:21number seven, which is predictions. Now,
  1693. 1:01:24once a model is evaluated and you've
  1694. 1:01:26improved the model, it is finally used
  1695. 1:01:29to make predictions. The final output
  1696. 1:01:31can be a categorical variable or it can
  1697. 1:01:34be a continuous quantity. Right? All of
  1698. 1:01:36this depends on the type of problem
  1699. 1:01:38you're trying to solve. Don't worry,
  1700. 1:01:39I'll be discussing the type of problems
  1701. 1:01:41that can be solved using machine
  1702. 1:01:43learning in the upcoming slides. In our
  1703. 1:01:45case for predicting the occurrence of
  1704. 1:01:47rainfall, the output will be a
  1705. 1:01:49categorical variable.
  1706. 1:01:51Categorical variable is anything that
  1707. 1:01:53has some categorical value. For example,
  1708. 1:01:56gender is a categorical variable. Gender
  1709. 1:01:59has either male, female, or other.
  1710. 1:02:02It has a defined set of values. That is
  1711. 1:02:04a categorical variable. So guys, that
  1712. 1:02:06was the entire machine learning process.
  1713. 1:02:09Now, as we continue with this tutorial,
  1714. 1:02:11in the upcoming sections, I will be
  1715. 1:02:14running a demo in Python, in which we
  1716. 1:02:16will be performing weather forecasting.
  1717. 1:02:18So, make sure you remember all these
  1718. 1:02:20steps that I spoke about because I'll be
  1719. 1:02:21going through all these steps by using
  1720. 1:02:24Python. We'll be coding all of this that
  1721. 1:02:26we just spoke about.
  1722. 1:02:28Now, the next topic we're going to
  1723. 1:02:29discuss is the types of machine
  1724. 1:02:31learning.
  1725. 1:02:32A machine can learn to solve a problem
  1726. 1:02:34by following any one of the three
  1727. 1:02:37approaches.
  1728. 1:02:38You can say that there are three ways in
  1729. 1:02:40which a machine learns. The three ways
  1730. 1:02:43are supervised learning, unsupervised
  1731. 1:02:45learning, and reinforcement learning.
  1732. 1:02:48These are the three methods in which you
  1733. 1:02:50can train a machine to learn.
  1734. 1:02:52So first, let's discuss supervised
  1735. 1:02:54learning.
  1736. 1:02:55So, what is supervised learning?
  1737. 1:02:57Supervised learning is a technique in
  1738. 1:02:59which we teach or train the machine by
  1739. 1:03:01using data which is labeled. To
  1740. 1:03:04understand this better, let's consider
  1741. 1:03:06an analogy.
  1742. 1:03:07As kids, we all needed guidance to solve
  1743. 1:03:10math problems. At least I had a really
  1744. 1:03:12tough time solving math problems. Yeah,
  1745. 1:03:15so our teachers always helped us
  1746. 1:03:17understand what addition is and how it
  1747. 1:03:19is done. Similarly, you can think of
  1748. 1:03:21supervised learning as a type of machine
  1749. 1:03:24learning that involves a guide. The
  1750. 1:03:26label data set is a teacher that will
  1751. 1:03:28train the machine to understand the
  1752. 1:03:30patterns in the data. The label data set
  1753. 1:03:33is nothing but the training data set.
  1754. 1:03:35So, to better understand this, consider
  1755. 1:03:37the figure. Right here, we're feeding
  1756. 1:03:40the machine images of Tom and Jerry, and
  1757. 1:03:42the goal is for the machine to identify
  1758. 1:03:45and classify the images into two
  1759. 1:03:47separate groups. Basically, one group
  1760. 1:03:49will contain Tom images, and the other
  1761. 1:03:51group will contain images of Jerry. Now,
  1762. 1:03:54pay attention to the training data set.
  1763. 1:03:56The training data set that is fed to a
  1764. 1:03:58model is labeled. As in, we're telling
  1765. 1:04:01the machine, "Listen, this is how Tom
  1766. 1:04:03looks, and this is how Jerry looks."
  1767. 1:04:05But, basically, labeling each data point
  1768. 1:04:08that we're feeding to the machine.
  1769. 1:04:10Right? If the image is of Tom's, we've
  1770. 1:04:12labeled it as Tom, and if the image is a
  1771. 1:04:15Jerry image, then we're going to label
  1772. 1:04:17it as Jerry. By doing this, you're
  1773. 1:04:19training the machine by using labeled
  1774. 1:04:22data. So, to sum it up, in supervised
  1775. 1:04:24learning, there is a well-defined
  1776. 1:04:26training phase done with the help of
  1777. 1:04:28labeled data. Right? The rest of the
  1778. 1:04:30process is the same. After you feed the
  1779. 1:04:33machine labeled data, you're going to
  1780. 1:04:34perform data cleaning, then exploratory
  1781. 1:04:37data analysis, followed by building the
  1782. 1:04:39machine learning model, and then model
  1783. 1:04:41evaluation, and finally, your
  1784. 1:04:43predictions. Also, one more point to
  1785. 1:04:45remember is that the output that you're
  1786. 1:04:47going to get in a supervised learning
  1787. 1:04:49algorithm is a labeled output. This
  1788. 1:04:52Jerry will be labeled as Jerry, and this
  1789. 1:04:54Tom will be labeled as Tom. Basically,
  1790. 1:04:56you'll get a labeled output. Now, let's
  1791. 1:04:59understand what is unsupervised
  1792. 1:05:01learning.
  1793. 1:05:02Unsupervised learning involves training
  1794. 1:05:05by using unlabeled data and allowing the
  1795. 1:05:07model to act on that information without
  1796. 1:05:10any guidance. So, think of unsupervised
  1797. 1:05:13learning as a smart kid that learns
  1798. 1:05:15without any guidance. In this type of
  1799. 1:05:17machine learning, the model is not fed
  1800. 1:05:20with any label data. As in, the model
  1801. 1:05:22has no clue that this image is Tom and
  1802. 1:05:25this image is Jerry. It figures out
  1803. 1:05:28patterns and the differences between Tom
  1804. 1:05:30and Jerry on its own by taking in tons
  1805. 1:05:32of data. For example, it identifies
  1806. 1:05:35prominent features of Tom such as pointy
  1807. 1:05:38ears, bigger in size, and so on to
  1808. 1:05:40understand that this image is of type
  1809. 1:05:42one.
  1810. 1:05:43Similarly, it finds such features in
  1811. 1:05:45Jerry and knows that this is another
  1812. 1:05:47type of image, [clears throat] maybe
  1813. 1:05:48type two. Right? Therefore, it
  1814. 1:05:50classifies the images into two different
  1815. 1:05:52clusters without knowing who is Tom and
  1816. 1:05:55who is Jerry. Now, the main idea behind
  1817. 1:05:57unsupervised learning is to understand
  1818. 1:05:59the patterns in your data set and form
  1819. 1:06:01clusters based on feature similarity.
  1820. 1:06:04Basically, it'll feature similar images
  1821. 1:06:06or similar data points into one cluster,
  1822. 1:06:09and it'll form another cluster which is
  1823. 1:06:11totally different from the first
  1824. 1:06:13cluster. So, look at the output over
  1825. 1:06:15here. The unlabeled output is basically
  1826. 1:06:17clusters or groups of two different
  1827. 1:06:20data.
  1828. 1:06:21Next, we have something known as the
  1829. 1:06:22reinforcement learning. Now,
  1830. 1:06:24reinforcement learning is comparatively
  1831. 1:06:26different, right? It's pretty different
  1832. 1:06:28from supervised and unsupervised.
  1833. 1:06:31It is basically a part of machine
  1834. 1:06:32learning where you put an agent in an
  1835. 1:06:35environment, and this agent learns to
  1836. 1:06:37behave in the environment by performing
  1837. 1:06:40certain actions and observing the
  1838. 1:06:42rewards which it gets from these
  1839. 1:06:44actions. To understand reinforcement
  1840. 1:06:46learning, imagine that you were dropped
  1841. 1:06:49off at an isolated island. What would
  1842. 1:06:51you do? Initially, we'd all panic,
  1843. 1:06:54right? But, as time passes by, you will
  1844. 1:06:56learn how to live on the island. You
  1845. 1:06:59will explore the environment. You will
  1846. 1:07:01understand the climate conditions.
  1847. 1:07:03You'll understand the type of food that
  1848. 1:07:05grows there. You'll know what is
  1849. 1:07:07dangerous to you and what is not. You'll
  1850. 1:07:09understand which food is good for you
  1851. 1:07:11and which is not. This is exactly how
  1852. 1:07:13reinforcement learning works. It
  1853. 1:07:15involves an agent, which is basically
  1854. 1:07:17you stuck on the island, that is put in
  1855. 1:07:20an unknown environment, which is the
  1856. 1:07:22island, where the agent must learn by
  1857. 1:07:24observing and performing actions that
  1858. 1:07:27result in rewards.
  1859. 1:07:28Reinforcement learning is mainly used in
  1860. 1:07:30advanced machine learning areas such as
  1861. 1:07:33self-driving cars, AlphaGo, and so on.
  1862. 1:07:35So, guys, that sums up the types of
  1863. 1:07:37machine learning. Before we go any
  1864. 1:07:39further, I'd like to discuss the
  1865. 1:07:40difference between supervised,
  1866. 1:07:42unsupervised, and reinforcement
  1867. 1:07:43learning. Now, first of all, we have the
  1868. 1:07:45definition. Supervised learning is all
  1869. 1:07:48about teaching a machine by using
  1870. 1:07:50labeled data. Unsupervised learning,
  1871. 1:07:53like the name suggests, there is no
  1872. 1:07:55supervision over here. The machine is
  1873. 1:07:57trained on unlabeled data without any
  1874. 1:08:00guidance.
  1875. 1:08:01Reinforcement learning is totally
  1876. 1:08:02different. Here, you have an agent who
  1877. 1:08:04interact with the environment by
  1878. 1:08:06producing actions and discover some
  1879. 1:08:09errors and rewards.
  1880. 1:08:10Now, the type of problem that is solved
  1881. 1:08:12using supervised learning is regression
  1882. 1:08:14and classification problems. We'll
  1883. 1:08:16discuss what regression, classification,
  1884. 1:08:18and clustering is in the upcoming slide,
  1885. 1:08:21right? So, don't worry if you don't know
  1886. 1:08:22what it is. Unsupervised learning is
  1887. 1:08:24mainly to solve association and
  1888. 1:08:26clustering problems.
  1889. 1:08:28Reinforcement learning is for
  1890. 1:08:29reward-based problems.
  1891. 1:08:31Now, what is the type of data in
  1892. 1:08:33supervised learning? It is labeled data.
  1893. 1:08:35That is the main difference between
  1894. 1:08:37supervised and any other type of machine
  1895. 1:08:39learning.
  1896. 1:08:40In supervised, you have labeled data. In
  1897. 1:08:42unsupervised, we have unlabeled data.
  1898. 1:08:44Whereas, in reinforcement learning, we
  1899. 1:08:46have no predefined data at all. The
  1900. 1:08:49machine has to perform everything from
  1901. 1:08:51scratch. It has to collect data,
  1902. 1:08:53analyze, do everything on its own. Now,
  1903. 1:08:55the training in supervised learning is
  1904. 1:08:57external supervision, meaning that we
  1905. 1:09:00have external supervision in the form of
  1906. 1:09:02the labeled training data set. In
  1907. 1:09:05unsupervised, there is obviously no
  1908. 1:09:07supervision. There is an unlabeled data
  1909. 1:09:09set, therefore there's no supervision.
  1910. 1:09:11In reinforcement learning, there is no
  1911. 1:09:13supervision at all. Now, the approach to
  1912. 1:09:15solving supervised learning problem is
  1913. 1:09:17basically you're going to map your
  1914. 1:09:19labeled input to your known output. In
  1915. 1:09:22unsupervised learning, the machine is
  1916. 1:09:23going to understand the patterns and
  1917. 1:09:25discover the output on its own.
  1918. 1:09:27Reinforcement learning, here the agent
  1919. 1:09:30will follow something known as a trial
  1920. 1:09:32and error method. Right? It's totally
  1921. 1:09:34based on the concept of trial and error.
  1922. 1:09:37Popular algorithms under supervised
  1923. 1:09:39learning are linear regression, logistic
  1924. 1:09:42regression, support vector machines, and
  1925. 1:09:44so on. Under unsupervised learning, we
  1926. 1:09:46have the famous K-means clustering
  1927. 1:09:48algorithm. Under reinforcement learning,
  1928. 1:09:50we have the Q-learning algorithm, which
  1929. 1:09:52is one of the most important algorithms.
  1930. 1:09:55It is basically the logic behind the
  1931. 1:09:57famous AlphaGo game. I'm sure all of you
  1932. 1:09:59have heard of that. So, guys, these were
  1933. 1:10:01the differences between supervised,
  1934. 1:10:03unsupervised, and reinforcement
  1935. 1:10:04learning. Now, let's move on and discuss
  1936. 1:10:07the type of problems that you can solve
  1937. 1:10:09by using machine learning.
  1938. 1:10:11Now, there are three types of problems
  1939. 1:10:13in machine learning. Now, any problem
  1940. 1:10:15that needs to be solved in machine
  1941. 1:10:17learning can fall into one of these
  1942. 1:10:19three categories. Now, what is a
  1943. 1:10:22regression? In this type of problem, the
  1944. 1:10:24output is a continuous quantity. For
  1945. 1:10:27example, if you want to predict the
  1946. 1:10:29speed of a car given the distance. That
  1947. 1:10:32means it is a regression problem.
  1948. 1:10:34First of all, what is a continuous
  1949. 1:10:36quantity? A continuous quantity is any
  1950. 1:10:38variable that can hold a continuous
  1951. 1:10:40value. A continuous variable is any
  1952. 1:10:43variable that can have infinite number
  1953. 1:10:45of values.
  1954. 1:10:47For example, the height of a person or
  1955. 1:10:49the weight of a person is a continuous
  1956. 1:10:51quantity. Right? I can have a weight of
  1957. 1:10:5350.1 kgs or 50.12 or 50.112 kg.
  1958. 1:10:59This is a continuous quantity.
  1959. 1:11:01Regression problems can be solved by
  1960. 1:11:03using supervised learning algorithms.
  1961. 1:11:06Another type of problem is a
  1962. 1:11:07classification problem. Here, the output
  1963. 1:11:10is always a categorical value.
  1964. 1:11:13Classifying emails into two classes, for
  1965. 1:11:15example, classifying your email as spam
  1966. 1:11:17and non-spam is a classification
  1967. 1:11:19problem.
  1968. 1:11:20Here again, you'll be using supervised
  1969. 1:11:22learning classification algorithms such
  1970. 1:11:24as support vector machines, naive bias,
  1971. 1:11:27logistic regression, and so on.
  1972. 1:11:29Then we have clustering problem. And
  1973. 1:11:31this type of problem involves assigning
  1974. 1:11:34the input into two or more clusters
  1975. 1:11:36based on feature similarity. For
  1976. 1:11:39example, clustering the viewers into
  1977. 1:11:41similar groups based on their interest
  1978. 1:11:44or based on their age or geography can
  1979. 1:11:46be done by using unsupervised learning
  1980. 1:11:48algorithms like K-means clustering.
  1981. 1:11:51One thing you need to understand is
  1982. 1:11:53under supervised learning, you can solve
  1983. 1:11:55regression and classification problems.
  1984. 1:11:58Under unsupervised learning, you can
  1985. 1:11:59solve clustering problems. Reinforcement
  1986. 1:12:02learning is something else altogether,
  1987. 1:12:05right? You can solve reward-based
  1988. 1:12:06problems and more complex and deep
  1989. 1:12:08problems. So, now let's move on and
  1990. 1:12:12understand the different machine
  1991. 1:12:13learning algorithms. Now, I will not be
  1992. 1:12:16going into depth for machine learning
  1993. 1:12:18algorithms because there are a lot of
  1994. 1:12:19algorithms to cover, but we have content
  1995. 1:12:22around almost every machine learning
  1996. 1:12:24algorithm out there. So, I'll be leaving
  1997. 1:12:26a couple of links in the description
  1998. 1:12:28box, right? You can check out all these
  1999. 1:12:30links and understand how each of these
  2000. 1:12:32machine learning algorithms work in
  2001. 1:12:34depth. So, I'm just going to show you a
  2002. 1:12:36hierarchical diagram of how the
  2003. 1:12:38algorithms are structured. So, under
  2004. 1:12:40machine learning, we have three types of
  2005. 1:12:42learning. We have supervised,
  2006. 1:12:43unsupervised, and reinforcement. Under
  2007. 1:12:45supervised learning, we have regression
  2008. 1:12:47and classification problems. And under
  2009. 1:12:50unsupervised learning, we have
  2010. 1:12:51clustering problems. Reinforcement
  2011. 1:12:54learning is completely different. I'll
  2012. 1:12:55be leaving a link in the description
  2013. 1:12:57specifically for reinforcement learning.
  2014. 1:12:59You can check out the entire content of
  2015. 1:13:01reinforcement learning there.
  2016. 1:13:03Now, regression problems can be solved
  2017. 1:13:05by using linear regression algorithm
  2018. 1:13:07such as linear regression, decision
  2019. 1:13:10trees, and random forest can also be
  2020. 1:13:11used in regression problems. But,
  2021. 1:13:14usually decision trees and random
  2022. 1:13:16forest, all of these are used to solve
  2023. 1:13:18classification problems. Famous
  2024. 1:13:20classification algorithms include
  2025. 1:13:22K-nearest neighbor, which is basically
  2026. 1:13:24KNN, decision trees and random forest,
  2027. 1:13:27logistic regression, naive bias, support
  2028. 1:13:30vector machines. All of these are
  2029. 1:13:32classification algorithms. Coming to
  2030. 1:13:34unsupervised learning, we have
  2031. 1:13:36clustering and association analysis. And
  2032. 1:13:39clustering problems can be solved by
  2033. 1:13:40using K-means. And association analysis
  2034. 1:13:43can be solved by using a priori
  2035. 1:13:45algorithm. A priori algorithm is mainly
  2036. 1:13:48used in market basket analysis. Right?
  2037. 1:13:51For this algorithm as well, I'll be
  2038. 1:13:52leaving a link in the description. We've
  2039. 1:13:55performed a very excellent demo where in
  2040. 1:13:57we've shown how market basket analysis
  2041. 1:13:59can be done by using a priori algorithm.
  2042. 1:14:02Markov model is also explained in one of
  2043. 1:14:05the videos. I'll be leaving that link in
  2044. 1:14:07the description box.
  2045. 1:14:08Now, to sum up machine learning to you,
  2046. 1:14:10I'll be running a small demonstration in
  2047. 1:14:13Python. Right? Like I promised earlier,
  2048. 1:14:16I'll be using Python to understand the
  2049. 1:14:18whole machine learning process. All
  2050. 1:14:20right. So, let's get started with that
  2051. 1:14:21demo.
  2052. 1:14:22So guys, for those of you who don't know
  2053. 1:14:24Python, I will leave a couple of links
  2054. 1:14:26in the description box so that you
  2055. 1:14:27understand Python. But, apart from that,
  2056. 1:14:30Python is pretty understandable. If you
  2057. 1:14:31just look at the code, you'll know what
  2058. 1:14:33exactly I'm talking about. Right? So,
  2059. 1:14:35don't worry. And also, I'll be
  2060. 1:14:36explaining everything in the code.
  2061. 1:14:39So, I'm using PyCharm in order to run
  2062. 1:14:42the demo.
  2063. 1:14:43Right? So guys, like I said, if you
  2064. 1:14:44don't know Python, I'll leave a couple
  2065. 1:14:46of links in the description box. You can
  2066. 1:14:48go through those videos as well. The
  2067. 1:14:51main aim of our demo is to build a
  2068. 1:14:53machine learning model that will predict
  2069. 1:14:55whether or not it will rain tomorrow by
  2070. 1:14:58studying the past data set. Now, this
  2071. 1:15:00data set contains around 145,000
  2072. 1:15:03observations on the daily weather
  2073. 1:15:05conditions as observed in Australia.
  2074. 1:15:08Right, the data set has around 24
  2075. 1:15:10features and we will be using 23
  2076. 1:15:13features out of that to predict the
  2077. 1:15:15target variable which is rain tomorrow.
  2078. 1:15:18So, this data set I collected from
  2079. 1:15:19Kaggle. Right, for those of you don't
  2080. 1:15:21know Kaggle is a online platform where
  2081. 1:15:23you can find hundreds of data sets and
  2082. 1:15:26you know, there are a lot of
  2083. 1:15:26competitions held by machine learning
  2084. 1:15:28engineers and all of that. It's an
  2085. 1:15:31interesting website.
  2086. 1:15:33Now, the problem statement itself is to
  2087. 1:15:35build a machine learning model that will
  2088. 1:15:37predict whether or not it will rain
  2089. 1:15:39tomorrow.
  2090. 1:15:40This is clearly a classification
  2091. 1:15:42problem. The machine learning model has
  2092. 1:15:44to classify the output into two classes,
  2093. 1:15:47that is either yes or no. Yes will stand
  2094. 1:15:50for it will rain tomorrow and no will
  2095. 1:15:52basically denote that it will not rain
  2096. 1:15:54tomorrow. Right, this is a
  2097. 1:15:55classification problem.
  2098. 1:15:57So, I hope the objective is clear.
  2099. 1:15:59Right, so we'll begin the demonstration
  2100. 1:16:01by importing the required libraries. So,
  2101. 1:16:04first of all, for mathematical
  2102. 1:16:06computations, we'll be importing the
  2103. 1:16:07NumPy library. We'll also be importing
  2104. 1:16:10the Pandas library for data processing.
  2105. 1:16:13Next, we will load the CSV file.
  2106. 1:16:15Basically, my data is stored in a CSV
  2107. 1:16:18format in this file. weatherAUS.csv is
  2108. 1:16:21my data set.
  2109. 1:16:22So, basically I've saved this file in
  2110. 1:16:24this path. Right, so that's what I'm
  2111. 1:16:26doing here. I'm loading my data set and
  2112. 1:16:28I'm storing it in a variable known as
  2113. 1:16:31DF.
  2114. 1:16:32Next, what we'll do is we'll see the
  2115. 1:16:33size of our data frame.
  2116. 1:16:36Let's print the size of the data frame.
  2117. 1:16:38We'll also display the first five
  2118. 1:16:40observations in our data frame.
  2119. 1:16:42Let's look at the output.
  2120. 1:16:45Basically, around 145,000
  2121. 1:16:47observations and 24 features. Now, 24
  2122. 1:16:51features are basically the variables
  2123. 1:16:53that are there in my data set. You know,
  2124. 1:16:55for example, date is a variable,
  2125. 1:16:57location is a variable, minimum
  2126. 1:16:59temperature till rain tomorrow. All of
  2127. 1:17:01these are variables. So, I have around
  2128. 1:17:0224 features in my data set, right? Now,
  2129. 1:17:05the variable that I have to predict is
  2130. 1:17:07rain tomorrow. Okay? If the value of
  2131. 1:17:10rain tomorrow is no, it denotes that it
  2132. 1:17:12will not rain tomorrow. But, if the
  2133. 1:17:14value is yes, then it will denote that
  2134. 1:17:16it will rain tomorrow.
  2135. 1:17:18So, rain tomorrow is basically my target
  2136. 1:17:20variable.
  2137. 1:17:21Right? I'll be finding out whether it's
  2138. 1:17:23going to rain tomorrow or not. So, this
  2139. 1:17:25is my target variable, also known as
  2140. 1:17:27your output variable.
  2141. 1:17:28My input variables will be the other 23
  2142. 1:17:31variables. Date, location, minimum
  2143. 1:17:33temperature, rain today, risk, all of
  2144. 1:17:36this will be my input variables. Now,
  2145. 1:17:38these variables are also known as
  2146. 1:17:40predictor variables. Basically, they're
  2147. 1:17:42used to predict your outcome. So, these
  2148. 1:17:44are also known as predictor variables.
  2149. 1:17:47Now, the next step is checking for null
  2150. 1:17:49values.
  2151. 1:17:50This is basically data pre-processing.
  2152. 1:17:53Let me just comment it for you.
  2153. 1:17:55This is data preparation or data
  2154. 1:18:01Right. So, this stage is data
  2155. 1:18:03pre-processing. Right here, we start
  2156. 1:18:05checking for any null values or any
  2157. 1:18:07missing values. This is exactly what I'm
  2158. 1:18:09doing over here. I am checking for any
  2159. 1:18:11missing or null values in my data set.
  2160. 1:18:14If you notice the output, it shows that
  2161. 1:18:16the first four columns have more than
  2162. 1:18:1840% null values. Right? So, it's always
  2163. 1:18:21best for us to remove features or such
  2164. 1:18:23variables because they will not help us
  2165. 1:18:25in our prediction.
  2166. 1:18:26Now, during data pre-processing, it is
  2167. 1:18:28always necessary to remove the variables
  2168. 1:18:31that are not significant. Unnecessary
  2169. 1:18:33data will just increase our
  2170. 1:18:35computations. That's why it's always
  2171. 1:18:37best if you remove the unwanted or
  2172. 1:18:39unnecessary variables.
  2173. 1:18:41Now, apart from removing these four
  2174. 1:18:42variables, we'll also remove the
  2175. 1:18:44location variable, and we will remove
  2176. 1:18:47the date variable. Right, I'll come to
  2177. 1:18:49this variable in a minute. We'll also be
  2178. 1:18:51removing location and date variable
  2179. 1:18:54because both of these variables are not
  2180. 1:18:56needed in order to predict whether it'll
  2181. 1:18:57rain tomorrow. Right, we do not need to
  2182. 1:19:00know the location and the date.
  2183. 1:19:02Now, we'll also be removing this a risk
  2184. 1:19:04mm variable. Risk mm variable basically
  2185. 1:19:07tells us the amount of rain that might
  2186. 1:19:09occur the next day.
  2187. 1:19:11Right, now this is a very informative
  2188. 1:19:13variable, and it might actually leak
  2189. 1:19:15some information to our model.
  2190. 1:19:17By using this variable, we'll easily be
  2191. 1:19:19able to predict, and there's no point of
  2192. 1:19:21doing that. So, this variable will give
  2193. 1:19:23us too much information. And so, that's
  2194. 1:19:25why we're going to remove this variable
  2195. 1:19:26as well. It'll leak a lot of
  2196. 1:19:28information. So, after that, if you
  2197. 1:19:30print the shape of your data frame,
  2198. 1:19:33we have only 17 variables and so many
  2199. 1:19:37observations.
  2200. 1:19:38Now, after this, we'll just uh look at
  2201. 1:19:40any null values, and we'll remove them.
  2202. 1:19:43This drop.any function will just remove
  2203. 1:19:45all the null values. Right, then if you
  2204. 1:19:47print the shape of your data frame,
  2205. 1:19:49we'll have around 112,000 rows with 17
  2206. 1:19:53variables. This is the shape of the data
  2207. 1:19:55set after removing all the null values
  2208. 1:19:57and all the redundant or unnecessary
  2209. 1:20:00variables.
  2210. 1:20:01Now, it's time to remove the outliers in
  2211. 1:20:03the data. So, after you remove any null
  2212. 1:20:05values, we should also check our data
  2213. 1:20:07set for any outliers. An outlier is a
  2214. 1:20:11data point that is very different from
  2215. 1:20:13your other observations.
  2216. 1:20:15Outliers usually occur because of
  2217. 1:20:17miscalculations while collecting the
  2218. 1:20:19data.
  2219. 1:20:20These are some sort of errors in your
  2220. 1:20:22data set.
  2221. 1:20:23So, in this whole code snippet, we're
  2222. 1:20:25just getting rid of outliers.
  2223. 1:20:29This is the output that we get. All our
  2224. 1:20:31outliers.
  2225. 1:20:32Next, what we'll be doing is we will be
  2226. 1:20:35assigning zeros and ones in the place of
  2227. 1:20:38yes and no.
  2228. 1:20:39The only thing is we're going to change
  2229. 1:20:40the categorical variables from yes and
  2230. 1:20:42no to zero and one. Right, that's
  2231. 1:20:44exactly what we're doing over here.
  2232. 1:20:46Now, if there are any unique values such
  2233. 1:20:48as any character values which are not
  2234. 1:20:50supposed to be there, we'll be changing
  2235. 1:20:52them into integer values.
  2236. 1:20:54That's all we're doing over here.
  2237. 1:20:56After this, we'll be normalizing our
  2238. 1:20:58data set.
  2239. 1:20:59This is a very important step because in
  2240. 1:21:02order to avoid any biasness in your
  2241. 1:21:04output, you have to normalize your input
  2242. 1:21:06variables.
  2243. 1:21:08Right, to do this, we can make use of
  2244. 1:21:09the min-max scalar function which Python
  2245. 1:21:11provides in a package known as
  2246. 1:21:13scikit-learn. You can use that package
  2247. 1:21:16in order to normalize your data set.
  2248. 1:21:18So, after normalizing our data set, this
  2249. 1:21:20is what our data set looks like.
  2250. 1:21:23This is before normalization. You can
  2251. 1:21:25see that these are in two digits,
  2252. 1:21:27whereas these values are in single
  2253. 1:21:29digits. Right? This causes a lot of
  2254. 1:21:30biasness. But once we normalize the
  2255. 1:21:33values, we know that all of the values
  2256. 1:21:35are in a similar range. We have
  2257. 1:21:37everything in decimals. Right, so
  2258. 1:21:39normalization is something that has to
  2259. 1:21:41be performed because if you have a data
  2260. 1:21:43set like this, your output is not going
  2261. 1:21:44to be correct. And that's why we perform
  2262. 1:21:46normalization.
  2263. 1:21:48So, now that we are done with
  2264. 1:21:50pre-processing, what we're going to do
  2265. 1:21:51is it's time for exploratory data
  2266. 1:21:54analysis.
  2267. 1:21:55Let me just comment it for you. This is
  2268. 1:21:58exploratory data analysis.
  2269. 1:22:01So, basically here, what we're going to
  2270. 1:22:02do is we're going to analyze and
  2271. 1:22:04identify the significant variables that
  2272. 1:22:07will help us predict the outcome. To do
  2273. 1:22:09this, we'll be using the select key best
  2274. 1:22:12function which is present in the
  2275. 1:22:13scikit-learn library.
  2276. 1:22:15There's a predefined function in Python
  2277. 1:22:17called select key best, which will
  2278. 1:22:19basically select the most significant
  2279. 1:22:21predictor variables in our data set.
  2280. 1:22:24When we run that line of code,
  2281. 1:22:26we get these three variables to be the
  2282. 1:22:29most significant variables in our data
  2283. 1:22:31set. Right, the main aim of this demo is
  2284. 1:22:33to make you understand how machine
  2285. 1:22:35learning works. That's why to simplify
  2286. 1:22:37the competition, we'll assign only one
  2287. 1:22:39of these significant variables as the
  2288. 1:22:42input. Instead of taking all three
  2289. 1:22:44variables as input, we'll select one
  2290. 1:22:46variable and we'll take that as the
  2291. 1:22:48input and the output is the rain
  2292. 1:22:50tomorrow variable.
  2293. 1:22:52So, basically, we are creating a data
  2294. 1:22:54frame of all the significant variables.
  2295. 1:22:56Basically, we're choosing this variable
  2296. 1:22:58in order to predict our outcome.
  2297. 1:23:00Obviously, our outcome is rain tomorrow
  2298. 1:23:02variable.
  2299. 1:23:03So, our input is humidity level and our
  2300. 1:23:06output is to detect whether it'll rain
  2301. 1:23:08tomorrow.
  2302. 1:23:09The next step is data modeling. All of
  2303. 1:23:11you are aware of what data modeling is.
  2304. 1:23:13To solve this, we'll be using
  2305. 1:23:15classification algorithms over here.
  2306. 1:23:17We'll use logistic regression. We will
  2307. 1:23:20use random forest classifier, which is
  2308. 1:23:23another machine learning algorithm.
  2309. 1:23:25We'll also use the decision tree
  2310. 1:23:26classifier and support vector machine.
  2311. 1:23:30Right, we'll be using all of these
  2312. 1:23:31algorithms in order to predict the
  2313. 1:23:33outcome. We'll also check which
  2314. 1:23:35algorithm gives us the best accuracy.
  2315. 1:23:38So guys, we're just using multiple
  2316. 1:23:40algorithms or multiple classification
  2317. 1:23:42algorithms on the same data set. We're
  2318. 1:23:44not doing anything very complex over
  2319. 1:23:46here.
  2320. 1:23:47So, we start by importing all the
  2321. 1:23:48necessary libraries for the logistic
  2322. 1:23:50regression algorithm.
  2323. 1:23:52We're also going to import time because
  2324. 1:23:54we'll be calculating the accuracy and
  2325. 1:23:56the time taken by the algorithm to get
  2326. 1:23:58the output.
  2327. 1:23:59So, the first step is data splicing.
  2328. 1:24:01I've already mentioned data splicing is
  2329. 1:24:03splitting your data set into your
  2330. 1:24:05testing data set and into your training
  2331. 1:24:07data set. That's exactly what we're
  2332. 1:24:09doing over here. So, 25% of your data is
  2333. 1:24:12assigned for the testing data and the
  2334. 1:24:14remaining 75% is your training data.
  2335. 1:24:17Here, you're creating the instance of
  2336. 1:24:19the logistic regression algorithm. This
  2337. 1:24:21is an instance that you created. Then
  2338. 1:24:23you'll fit the model by using your
  2339. 1:24:25training data set. So basically, to
  2340. 1:24:27build your machine learning algorithm,
  2341. 1:24:29you'll be fitting your training data
  2342. 1:24:30set. So X_train and Y_train variables
  2343. 1:24:33have your training data set.
  2344. 1:24:35After that, you will be evaluating the
  2345. 1:24:37model by using your testing data set.
  2346. 1:24:40Then you'll calculate the accuracy
  2347. 1:24:42score. Right? I'll also be printing the
  2348. 1:24:44accuracy using logistic regression and
  2349. 1:24:47the time taken using logistic
  2350. 1:24:48regression. Let's look at the accuracy.
  2351. 1:24:51Don't worry about these warnings. They
  2352. 1:24:53are not important. So accuracy using
  2353. 1:24:55logistic regression is around 0.83%,
  2354. 1:24:59which is 83% accuracy, approximately
  2355. 1:25:0184%. And this is the time taken. So the
  2356. 1:25:05accuracy is actually pretty good, right?
  2357. 1:25:0684% is a good number. Then we have
  2358. 1:25:09random forest classifier. Here again,
  2359. 1:25:11we'll import the libraries that are
  2360. 1:25:13needed to run random forest classifier.
  2361. 1:25:16Then we're again calculating the
  2362. 1:25:17accuracy and the time taken by the
  2363. 1:25:19classifier.
  2364. 1:25:20Data splicing, like I mentioned,
  2365. 1:25:22splitting the data into testing and
  2366. 1:25:24training data set. Then you're just
  2367. 1:25:26building the model by using the training
  2368. 1:25:28data set. After that, you'll evaluate
  2369. 1:25:31the model by using the testing data set
  2370. 1:25:33and you'll finally calculate the
  2371. 1:25:35accuracy. The accuracy using random
  2372. 1:25:37forest is again approximately 84%, which
  2373. 1:25:40is a really good number. Then we have
  2374. 1:25:42decision tree classifier. Here again,
  2375. 1:25:45we'll be importing the libraries needed
  2376. 1:25:47for this classifier. We'll be
  2377. 1:25:49calculating the accuracy and the time
  2378. 1:25:51taken by this classifier. Data splicing
  2379. 1:25:54followed by building the model by using
  2380. 1:25:56the training data set, evaluating the
  2381. 1:25:58model by using the testing data set, and
  2382. 1:26:00finally calculating the accuracy and
  2383. 1:26:02printing the accuracy.
  2384. 1:26:04So let's see the accuracy using decision
  2385. 1:26:06tree classifier. Again, we have an
  2386. 1:26:09accuracy of around 83 to 84%.
  2387. 1:26:12This is a pretty good number. And last,
  2388. 1:26:14we're going to do this by using another
  2389. 1:26:16classification algorithm known as
  2390. 1:26:18support vector machine.
  2391. 1:26:20Here again, we're importing the needed
  2392. 1:26:22libraries. Then we're calculating the
  2393. 1:26:24accuracy and the time, performing data
  2394. 1:26:26splicing.
  2395. 1:26:27Then we're building the model by using
  2396. 1:26:29the training data set, testing the model
  2397. 1:26:31using the testing data set, and finally
  2398. 1:26:33printing the accuracy.
  2399. 1:26:35So guys, all the classification models
  2400. 1:26:38gave us an accuracy score of
  2401. 1:26:39approximately 84% to 83%.
  2402. 1:26:43So this is exactly how a machine
  2403. 1:26:44learning process works. Right? You begin
  2404. 1:26:47by importing all your data, then you
  2405. 1:26:49perform data pre-processing or data
  2406. 1:26:51cleaning. After that, you perform
  2407. 1:26:53exploratory data analysis, where you
  2408. 1:26:55understand the important patterns or the
  2409. 1:26:58important variables in your data set.
  2410. 1:27:00After that, you build a model, then you
  2411. 1:27:03will evaluate the model by using the
  2412. 1:27:05testing data set, and finally calculate
  2413. 1:27:07the accuracy. I showed you all the steps
  2414. 1:27:09in the machine learning process by using
  2415. 1:27:11a practical demonstration in Python.
  2416. 1:27:14So guys, give yourself a pat on the back
  2417. 1:27:16because we just understood the whole
  2418. 1:27:17machine learning process with a small
  2419. 1:27:19implementation in Python.
  2420. 1:27:22Now let's move on to our next topic,
  2421. 1:27:24which is limitations of machine
  2422. 1:27:26learning.
  2423. 1:27:27Before we understand what deep learning
  2424. 1:27:29is, it's important to know the
  2425. 1:27:30limitations of machine learning, and why
  2426. 1:27:33these limitations gave rise to the
  2427. 1:27:35concept of deep learning. One major
  2428. 1:27:37problem in machine learning is machine
  2429. 1:27:40learning algorithms and models are not
  2430. 1:27:42capable of handling high-dimensional
  2431. 1:27:45data.
  2432. 1:27:46Right? We can take in data with 20 to 30
  2433. 1:27:49feature variables, but when it comes to
  2434. 1:27:51data sets which have thousands of
  2435. 1:27:53variables, machine learning does not
  2436. 1:27:55work. Machine learning is not capable
  2437. 1:27:58enough to process that much data.
  2438. 1:28:01So high-dimensional data cannot be
  2439. 1:28:03analyzed, processed, and modeled by
  2440. 1:28:05using machine learning.
  2441. 1:28:07Another limitation is that it cannot be
  2442. 1:28:09used in image recognition and object
  2443. 1:28:12detection because these applications
  2444. 1:28:14require the implementation of
  2445. 1:28:16high-dimensional data. Another major
  2446. 1:28:19challenge in machine learning is to tell
  2447. 1:28:21the machine what are the important
  2448. 1:28:23features it should look for in order to
  2449. 1:28:25precisely predict the outcome. So,
  2450. 1:28:28basically you're selecting the important
  2451. 1:28:29features for the machine learning model
  2452. 1:28:31and you're telling them like these are
  2453. 1:28:32the important features and this is what
  2454. 1:28:35you should use in order to build the
  2455. 1:28:36model.
  2456. 1:28:37This process is known as feature
  2457. 1:28:39extraction.
  2458. 1:28:40Now, in machine learning this is a
  2459. 1:28:41manual process. You're going to manually
  2460. 1:28:43input as a programmer, you're going to
  2461. 1:28:45tell that these are the important
  2462. 1:28:47predictor variables. But, what happens
  2463. 1:28:49when your data set has hundreds of
  2464. 1:28:51variables?
  2465. 1:28:52How are you going to sit and choose
  2466. 1:28:54every variable and perform analysis on
  2467. 1:28:56each variable to understand which is a
  2468. 1:28:58really significant variable? That's
  2469. 1:29:01going to become a very tedious task,
  2470. 1:29:03right? It's not possible for you to
  2471. 1:29:04manually sit down with 100 variables,
  2472. 1:29:06check the correlation with each variable
  2473. 1:29:08and understand which variable is
  2474. 1:29:09significant in predicting the output.
  2475. 1:29:12So, performing feature extraction
  2476. 1:29:14manually is very tedious and that is one
  2477. 1:29:16of the major limitations of machine
  2478. 1:29:17learning. Now, deep learning comes to
  2479. 1:29:20the rescue to all of these problems.
  2480. 1:29:22So, let's understand what deep learning
  2481. 1:29:25is and why we have deep learning in the
  2482. 1:29:27first place.
  2483. 1:29:28So, deep learning is actually one of the
  2484. 1:29:30only methods by which we can overcome
  2485. 1:29:33the challenge of feature extraction.
  2486. 1:29:35This is because deep learning models are
  2487. 1:29:37capable of learning to focus on the
  2488. 1:29:39right features by themselves requiring
  2489. 1:29:42minimal human intervention. Meaning that
  2490. 1:29:45feature extraction will be performed by
  2491. 1:29:47the deep learning model itself. You
  2492. 1:29:49don't have to manually tell that this
  2493. 1:29:50feature is important, that feature is
  2494. 1:29:52important, choose this feature for
  2495. 1:29:54predicting the output. All of this is
  2496. 1:29:56not needed in deep learning. The model
  2497. 1:29:58itself will learn which features are
  2498. 1:30:00most significant in predicting the
  2499. 1:30:02output.
  2500. 1:30:03Also, deep learning is mainly used to
  2501. 1:30:05deal with high-dimensional data, right?
  2502. 1:30:08It is based on the concept of neural
  2503. 1:30:10networks and is often used in object
  2504. 1:30:13detection and image processing. This is
  2505. 1:30:15exactly why we need deep learning. It
  2506. 1:30:17solves the problem of processing
  2507. 1:30:19high-dimensional data and manual feature
  2508. 1:30:22extraction.
  2509. 1:30:23Now, how exactly does deep learning
  2510. 1:30:25work? Now, deep learning mimics the
  2511. 1:30:27basic component of the human brain
  2512. 1:30:29called the brain cell. The brain cell is
  2513. 1:30:32also known as a neuron.
  2514. 1:30:34So, inspired from a neuron, an
  2515. 1:30:37artificial neuron was developed. Deep
  2516. 1:30:39learning is based on the functionality
  2517. 1:30:41of a biological neuron. So, let's
  2518. 1:30:44understand how we mimic this
  2519. 1:30:46functionality in an artificial neuron.
  2520. 1:30:49Now guys, an artificial neuron is also
  2521. 1:30:51known as a perceptron.
  2522. 1:30:53Let's understand what this biological
  2523. 1:30:55neuron does and how deep learning is
  2524. 1:30:57based on this concept.
  2525. 1:30:59In a biological neuron, you can see
  2526. 1:31:01these dendrites, right? In this image,
  2527. 1:31:03you see something known as dendrites.
  2528. 1:31:06These dendrites are used to receive any
  2529. 1:31:08input. These inputs are summed in the
  2530. 1:31:11cell body and through the axon, it is
  2531. 1:31:13passed on to the next neuron.
  2532. 1:31:16So, similar to the biological neuron, a
  2533. 1:31:18perceptron or a artificial neuron
  2534. 1:31:20receives multiple inputs, applies
  2535. 1:31:23various transformations and functions,
  2536. 1:31:25and provides an output.
  2537. 1:31:27Right? So, that's how artificial neural
  2538. 1:31:29networks or that's how deep learning
  2539. 1:31:31works.
  2540. 1:31:32Now guys, the human brain consists of
  2541. 1:31:34multiple connected neurons called a
  2542. 1:31:36neural network. Similarly, by combining
  2543. 1:31:39multiple perceptrons, we've developed
  2544. 1:31:41what is known as deep neural networks.
  2545. 1:31:44The main idea behind deep learning is
  2546. 1:31:46neural networks and that's what we're
  2547. 1:31:48going to learn about. So now, let's
  2548. 1:31:50understand what exactly deep learning
  2549. 1:31:52is.
  2550. 1:31:53Deep learning is a collection of
  2551. 1:31:55statistical machine learning techniques
  2552. 1:31:57used to learn feature hierarchies based
  2553. 1:32:00on the concept of artificial neural
  2554. 1:32:02networks. So, the main idea behind deep
  2555. 1:32:05learning is to use the concept of neural
  2556. 1:32:08networks.
  2557. 1:32:09A deep neural network will have three
  2558. 1:32:11layers. Okay, there's something known as
  2559. 1:32:13the input layer followed by the hidden
  2560. 1:32:15layers and then we have the output
  2561. 1:32:17layer. The input layer is basically the
  2562. 1:32:19first layer and it receives all the
  2563. 1:32:21inputs. So, all the inputs are fed into
  2564. 1:32:24this input layer.
  2565. 1:32:25The last layer is obviously the output
  2566. 1:32:27layer. This layer will provide your
  2567. 1:32:30desired output. Now, all the layers
  2568. 1:32:32between the input and your output layer
  2569. 1:32:34are known as the hidden layers.
  2570. 1:32:36Now, the number of hidden layers in a
  2571. 1:32:39deep learning network will depend on the
  2572. 1:32:41type of problem you're trying to solve
  2573. 1:32:42and the data that you have.
  2574. 1:32:44We'll get into depth of what exactly a
  2575. 1:32:46hidden layer does, but for now this is
  2576. 1:32:48how a neural network is structured in
  2577. 1:32:51deep learning. So, guys uh deep learning
  2578. 1:32:53is used in highly computational use
  2579. 1:32:56cases such as face verification,
  2580. 1:32:58self-driving cars, and so on. Right? So,
  2581. 1:33:00let's understand the importance of deep
  2582. 1:33:02learning by looking at a real-world use
  2583. 1:33:05case.
  2584. 1:33:06So, I'm sure all of you have heard of
  2585. 1:33:07the company PayPal. Now, PayPal makes
  2586. 1:33:10use of deep learning to identify any
  2587. 1:33:13possible fraudulent activities.
  2588. 1:33:16So, the company makes use of deep
  2589. 1:33:18learning for fraud detection. Now,
  2590. 1:33:20PayPal recently processed over 235
  2591. 1:33:24billion dollars in payments from 4
  2592. 1:33:27billion transactions by its more than
  2593. 1:33:30170 million customers. So, basically it
  2594. 1:33:33processed this much data by using deep
  2595. 1:33:35learning. PayPal uses machine learning
  2596. 1:33:38and deep learning algorithms to mine
  2597. 1:33:40data from the customers purchasing
  2598. 1:33:42history in addition to reviewing
  2599. 1:33:44patterns of any sort of fraud stored in
  2600. 1:33:47the database and it will do this to
  2601. 1:33:49predict whether a particular transaction
  2602. 1:33:51is fraudulent or not. Now, the company
  2603. 1:33:53has been relying on deep learning and
  2604. 1:33:55machine learning technology for around
  2605. 1:33:5710 years.
  2606. 1:33:59Initially, the fraud monitoring team
  2607. 1:34:01used simple linear models, right? They
  2608. 1:34:03used machine learning, but over the
  2609. 1:34:05years the company switched to more
  2610. 1:34:07advanced machine learning technology
  2611. 1:34:09called deep learning. This shows how
  2612. 1:34:11deep learning is used in more advanced
  2613. 1:34:14and more complicated use cases.
  2614. 1:34:17The fraud risk manager and the data
  2615. 1:34:19scientist at PayPal, he quoted that what
  2616. 1:34:23we enjoy from more modern advanced
  2617. 1:34:25machine learning is its ability to
  2618. 1:34:27consume a lot more data, handle layers
  2619. 1:34:30and layers of abstraction, and be able
  2620. 1:34:32to see things that a simpler technology
  2621. 1:34:35would not be able to see. Even human
  2622. 1:34:37beings might not able to see. This is
  2623. 1:34:40exactly what he quoted. He said that a
  2624. 1:34:42simple linear model is capable of
  2625. 1:34:44consuming around 20 variables, but with
  2626. 1:34:48deep learning technology, you can run
  2627. 1:34:50thousands of data points.
  2628. 1:34:52He also quoted that there is a magnitude
  2629. 1:34:55of difference. You'll be able to analyze
  2630. 1:34:57a lot more information and identify
  2631. 1:35:00patterns that are a lot more
  2632. 1:35:02sophisticated.
  2633. 1:35:03So, by implementing deep learning
  2634. 1:35:05technology, PayPal can finally analyze
  2635. 1:35:07millions of transactions to identify any
  2636. 1:35:10fraudulent activity.
  2637. 1:35:12This is how PayPal makes use of deep
  2638. 1:35:14learning.
  2639. 1:35:15Not only PayPal, we also have Facebook,
  2640. 1:35:17right? Facebook makes use of deep
  2641. 1:35:19learning technology for face
  2642. 1:35:21verification.
  2643. 1:35:22You've all seen the tagging feature at
  2644. 1:35:24Facebook where we tag our friends in
  2645. 1:35:26photos. All of that is based on deep
  2646. 1:35:29learning and machine learning.
  2647. 1:35:31So guys, that was a real-world use case
  2648. 1:35:33to make you understand how important
  2649. 1:35:35deep learning is.
  2650. 1:35:36Now, let's move on and look at what
  2651. 1:35:39exactly a perceptron is, right? We'll be
  2652. 1:35:41going in depth about deep learning.
  2653. 1:35:44A perceptron is basically a single-layer
  2654. 1:35:47neural network that is used to classify
  2655. 1:35:49linear data.
  2656. 1:35:51It is the most basic component of a
  2657. 1:35:53neural network. Now, a perceptron has
  2658. 1:35:55four important components. It has
  2659. 1:35:58something known as inputs, weights, and
  2660. 1:36:00bias, summation functions, activation
  2661. 1:36:03and transformation functions.
  2662. 1:36:05These are four important parts of a
  2663. 1:36:07perceptron.
  2664. 1:36:09Now, before I discuss this diagram with
  2665. 1:36:11you, let me tell you the basic logic
  2666. 1:36:13behind a perceptron.
  2667. 1:36:15There is something known as inputs,
  2668. 1:36:17right? The input X here, you can see X1,
  2669. 1:36:19X2 till Xn.
  2670. 1:36:21So, let me explain the structure of a
  2671. 1:36:22perceptron. What you're going to do is
  2672. 1:36:24you're going to input variables into the
  2673. 1:36:26perceptron, right? This X1, X2 till Xn
  2674. 1:36:29basically stands for input. W1, W2 till
  2675. 1:36:33Wn stands for the weight assigned to
  2676. 1:36:36each of these inputs.
  2677. 1:36:38Right? There is a specific weight
  2678. 1:36:40that'll be randomly initialized in the
  2679. 1:36:42beginning for each of your input.
  2680. 1:36:45Next, you have something known as the
  2681. 1:36:46summation element. Here, what you do is
  2682. 1:36:48you multiply the respective input with
  2683. 1:36:51the respective weight, and you add all
  2684. 1:36:54these products. Right? That is basically
  2685. 1:36:57your summation function. After this is
  2686. 1:36:59what is your transfer function, also
  2687. 1:37:01known as activation function.
  2688. 1:37:03Right? The activation function map your
  2689. 1:37:05input to your desired output. So, your
  2690. 1:37:08input will go through these processes.
  2691. 1:37:11It'll go through summation and
  2692. 1:37:12activation function in order to get to
  2693. 1:37:14the output. So, guys, remember that the
  2694. 1:37:17neural networks work the same way as a
  2695. 1:37:19perceptron. So, if you want to
  2696. 1:37:21understand how deep neural networks
  2697. 1:37:23work, you need to understand what a
  2698. 1:37:24perceptron does.
  2699. 1:37:26A deep neural network is nothing but
  2700. 1:37:27multiple perceptrons.
  2701. 1:37:29So, let me tell you how the entire works
  2702. 1:37:31once again.
  2703. 1:37:32So, basically, all your inputs are
  2704. 1:37:34multiplied with their respective
  2705. 1:37:36weights. Now, you add all the multiplied
  2706. 1:37:39values, and you call them as a weight
  2707. 1:37:41sum. You use the summation function to
  2708. 1:37:43add all of this. After that, you apply
  2709. 1:37:46the weighted sum to the correct
  2710. 1:37:48activation transfer function. Activation
  2711. 1:37:50function is very similar to a function
  2712. 1:37:53in our brain. The neurons become active
  2713. 1:37:56in our brain after a certain potential
  2714. 1:37:59is reached. That threshold is known as
  2715. 1:38:01the activation potential.
  2716. 1:38:03So, mathematically, there are a few
  2717. 1:38:05functions which represent the activation
  2718. 1:38:07function. Basically, the signum, the
  2719. 1:38:09sigmoid, the tan h, all of these are
  2720. 1:38:11activation functions. You can think of
  2721. 1:38:13activation function as a function that
  2722. 1:38:15maps the input to the respective output.
  2723. 1:38:19Then, I spoke about something known as
  2724. 1:38:20weights and biases. All right, now you
  2725. 1:38:23must be wondering, why do we have to
  2726. 1:38:25assign weights to each of our input?
  2727. 1:38:28Weights basically show the strength of a
  2728. 1:38:30particular input or how important a
  2729. 1:38:33particular input is for predicting the
  2730. 1:38:35output. In simple words, the weightage
  2731. 1:38:38denotes the importance of an input. Bias
  2732. 1:38:41is basically a value which allows you to
  2733. 1:38:44shift the activation function curve in
  2734. 1:38:47order to get a precise output.
  2735. 1:38:50All right, so that's exactly what
  2736. 1:38:51weights are.
  2737. 1:38:52I hope all of you are clear with inputs,
  2738. 1:38:54weights, summation, and activation
  2739. 1:38:57function. Also, one important thing I
  2740. 1:38:59forgot to mention in a perceptron is a
  2741. 1:39:01single layer perceptron will have no
  2742. 1:39:04hidden layers. All right, there'll only
  2743. 1:39:05be an input layer, an output layer, and
  2744. 1:39:07a couple of transformation function in
  2745. 1:39:09between. That's all will be there in a
  2746. 1:39:11perceptron. Now, perceptron, like I
  2747. 1:39:13mentioned, is used to solve only linear
  2748. 1:39:15problems. If you look at this data
  2749. 1:39:17distribution, how do you think we can
  2750. 1:39:19solve this? This data is not linearly
  2751. 1:39:22separable. So, you cannot use a single
  2752. 1:39:24layer perceptron to separate this data.
  2753. 1:39:27All right, that's why we need something
  2754. 1:39:28known as a multi-layer perceptron with
  2755. 1:39:31backpropagation.
  2756. 1:39:32I'll be explaining this in the next
  2757. 1:39:34slide. So, complex problems that involve
  2758. 1:39:38a lot of parameters and high-dimensional
  2759. 1:39:41data can be solved by using multiple
  2760. 1:39:43layer perceptron. Now, a multi-layer
  2761. 1:39:46perceptron is the same as a single layer
  2762. 1:39:48perceptron. The only difference is that
  2763. 1:39:50a multi-layer perceptron will have
  2764. 1:39:52hidden layers.
  2765. 1:39:53So, the number of hidden layers in a
  2766. 1:39:56model depends upon various factors. I
  2767. 1:39:58told you it depends on the complexity of
  2768. 1:40:00the problem you're trying to solve. It
  2769. 1:40:01depends on the number of inputs in your
  2770. 1:40:03data and so on. So, it works in the same
  2771. 1:40:05way. All your inputs are multiplied with
  2772. 1:40:08your weights and then you do the
  2773. 1:40:09summation and then there is a
  2774. 1:40:11transformation function or a activation
  2775. 1:40:14function.
  2776. 1:40:15While designing a neural network, in the
  2777. 1:40:17beginning itself I told you we
  2778. 1:40:19initialize weights with some random
  2779. 1:40:21values. We do not have some specific
  2780. 1:40:23[clears throat] value for each
  2781. 1:40:24weightage. Initially, we've selected
  2782. 1:40:27random values.
  2783. 1:40:28It is always important that whatever
  2784. 1:40:31weight values we have selected will be
  2785. 1:40:33correct.
  2786. 1:40:34Now, whatever weight values we've
  2787. 1:40:36assigned to each input, it denotes the
  2788. 1:40:38importance of that input variable. So,
  2789. 1:40:41we need to assign the weights in such a
  2790. 1:40:43way or we need to update the weights in
  2791. 1:40:45such a way that it denotes the
  2792. 1:40:47significance of that particular input.
  2793. 1:40:50So, initially we're selecting some
  2794. 1:40:52random value for weight. And let's say
  2795. 1:40:54that we use this weight value to get our
  2796. 1:40:57output.
  2797. 1:40:58Now, what happens is the output is
  2798. 1:41:00actually very different or it is not
  2799. 1:41:02precise when compared to our actual
  2800. 1:41:04output. Basically, the error value is
  2801. 1:41:07very huge.
  2802. 1:41:08So, how will you reduce the error?
  2803. 1:41:10The main thing in a neural network is
  2804. 1:41:13the weightage that you give to a input
  2805. 1:41:15variable, right? Depending on the
  2806. 1:41:16weightage that you give to a input
  2807. 1:41:18variable, you're telling the neural
  2808. 1:41:19network how important that variable is.
  2809. 1:41:22Now, what if you randomly give some
  2810. 1:41:24weightage and your output is wrong? The
  2811. 1:41:28first thing that comes into your mind is
  2812. 1:41:29that you need to change the weight
  2813. 1:41:31because the weight signifies the
  2814. 1:41:33importance of a variable.
  2815. 1:41:35So, basically what we need to do is we
  2816. 1:41:37need to somehow explain to the model to
  2817. 1:41:39change the weight in such a way that the
  2818. 1:41:41error becomes minimum. Let's put it in
  2819. 1:41:44another way. So, basically, we need to
  2820. 1:41:46train a model. One way to train a model
  2821. 1:41:48is called as backpropagation.
  2822. 1:41:51So, in backpropagation, what happens is
  2823. 1:41:54once you've initialized a weight to each
  2824. 1:41:56of the input, you calculate the output.
  2825. 1:41:59Right? You get an output, and let's say
  2826. 1:42:01you have a very high error value in that
  2827. 1:42:03output. What you do is you'll
  2828. 1:42:05backpropagate as in you'll go back to
  2829. 1:42:08the weight, and you'll keep updating the
  2830. 1:42:10weight in such a way that your error
  2831. 1:42:12becomes minimum.
  2832. 1:42:14This is exactly what backpropagation is.
  2833. 1:42:16You'll be going back to the first layer.
  2834. 1:42:18You'll be updating each of the weights
  2835. 1:42:20in such a way that your output is more
  2836. 1:42:23precise. So, guys, basically, the weight
  2837. 1:42:26and the error in a neural network is
  2838. 1:42:29highly related. By updating the weight
  2839. 1:42:32in a particular way, your error will
  2840. 1:42:33decrease. So, you need to figure out how
  2841. 1:42:36you need to update the weight. Do you
  2842. 1:42:38have to increase the weight or decrease
  2843. 1:42:40the weight? Once you figure out whether
  2844. 1:42:41you have to increase or decrease the
  2845. 1:42:43weight, you have to just follow that
  2846. 1:42:45direction in such a way that your error
  2847. 1:42:47is minimized. And that's exactly what
  2848. 1:42:50backpropagation is. So, the final output
  2849. 1:42:53of backpropagation is you're going to
  2850. 1:42:55select the weight that minimizes the
  2851. 1:42:57error function. And then you're going to
  2852. 1:42:59use that weight to solve the whole
  2853. 1:43:01problem.
  2854. 1:43:02Right? This is what backpropagation is
  2855. 1:43:04about.
  2856. 1:43:05Now, in order to make you understand
  2857. 1:43:07deep neural networks, let's look at a
  2858. 1:43:09practical implementation.
  2859. 1:43:12So, again, guys, I'll be using Python to
  2860. 1:43:14run the demo. If you don't have a good
  2861. 1:43:16idea about Python, check the
  2862. 1:43:17description. I'll leave a couple of
  2863. 1:43:19links about Python programming. Now, in
  2864. 1:43:21this demo, I'll be walking you through
  2865. 1:43:23one of the most important applications
  2866. 1:43:25of deep learning. I will demonstrate how
  2867. 1:43:28you can construct a high-performance
  2868. 1:43:30model to detect credit card fraud.
  2869. 1:43:32Right? We'll be using deep learning
  2870. 1:43:34models to do this. Now, before that, let
  2871. 1:43:36me just tell you something about our
  2872. 1:43:37data set. Right? The data set contains
  2873. 1:43:40transactions made by credit cards in the
  2874. 1:43:43year September 2013 by European card
  2875. 1:43:46holders. This data set presents
  2876. 1:43:48transactions that occurred in 2 days,
  2877. 1:43:51where we have 492 frauds out of 285,000
  2878. 1:43:56transactions. Approximately 285,000
  2879. 1:44:00transactions. Out of these transactions,
  2880. 1:44:02492 were frauds, and the data set is
  2881. 1:44:05quite unbalanced. Right? The positive
  2882. 1:44:07class accounts for 0.172%.
  2883. 1:44:11So, the positive class, basically the
  2884. 1:44:13fraudulent class. So, again, we're going
  2885. 1:44:15to start by importing the required
  2886. 1:44:17packages. We're going to import Keras,
  2887. 1:44:19Matplotlib library, Seaborn library, and
  2888. 1:44:22scikit-learn for preprocessing. Right?
  2889. 1:44:24Again, min-max scalar, which is for
  2890. 1:44:26normalization. We're going to import our
  2891. 1:44:29data set and store it in this variable.
  2892. 1:44:31Right? This is the path to my data set.
  2893. 1:44:34My data set is in the CSV format, or
  2894. 1:44:36also known as comma-separated version.
  2895. 1:44:39Now, we're going to print out the first
  2896. 1:44:41five rows of our data set. Right? Let's
  2897. 1:44:44take a look at the output.
  2898. 1:44:47So, here is the time of the transaction.
  2899. 1:44:49V1, V2, V3, etc. These are all the
  2900. 1:44:52features of our data set. I'm not going
  2901. 1:44:55to go into depth of what these features
  2902. 1:44:57stand for, because this demo is all
  2903. 1:44:59about understanding deep learning. Now,
  2904. 1:45:01these V1, V2, V3, these are all
  2905. 1:45:03predictor variables, which will help us
  2906. 1:45:05predict our class.
  2907. 1:45:07So, guys, don't worry about what these
  2908. 1:45:09features are. These features are just
  2909. 1:45:11information and details about your
  2910. 1:45:13transaction, such as the amount you
  2911. 1:45:15spend, or the time of transaction, and
  2912. 1:45:18so on.
  2913. 1:45:19So, here we have the amount variable,
  2914. 1:45:20which denotes the amount spent. After
  2915. 1:45:23that, we have the class variable. Now,
  2916. 1:45:25this class variable is your output
  2917. 1:45:27variable or your target variable. So,
  2918. 1:45:30your class is basically your output
  2919. 1:45:32variable. Value zero denotes that there
  2920. 1:45:34has been no fraudulent activity, but if
  2921. 1:45:36you get a class of one, it means that
  2922. 1:45:39this transaction is a fraudulent
  2923. 1:45:41transaction. For example, this
  2924. 1:45:43transaction is not fraudulent, and
  2925. 1:45:45that's why we have a value of zero over
  2926. 1:45:47here. All right, so this is our data
  2927. 1:45:49set.
  2928. 1:45:51Next what we're doing is we're counting
  2929. 1:45:53the number of samples for each class.
  2930. 1:45:55Right, we have class zero and class one,
  2931. 1:45:57where in class zero denotes the normal
  2932. 1:45:59transaction, which is non-fraudulent
  2933. 1:46:02transaction, and class one will denote
  2934. 1:46:04the fraudulent transactions. Right, so
  2935. 1:46:06we have around 492 fraudulent
  2936. 1:46:08transactions and around 284,315
  2937. 1:46:13non-fraudulent transactions.
  2938. 1:46:15So, when you see this, you know that our
  2939. 1:46:17data set is highly unbalanced. Highly
  2940. 1:46:19unbalanced means that one class has a
  2941. 1:46:22really small number when compared to the
  2942. 1:46:24other class. Right, there's no balance
  2943. 1:46:25between the two classes.
  2944. 1:46:27So, here what we're doing is we are
  2945. 1:46:29starting the data set by class for
  2946. 1:46:31stratified sampling. Stratified sampling
  2947. 1:46:34is a statistical technique for sampling
  2948. 1:46:37your data set. Now, this type of
  2949. 1:46:39sampling is always good if you have an
  2950. 1:46:41unbalanced data set.
  2951. 1:46:43Next what we're going to do is we're
  2952. 1:46:44going to perform data pre-processing.
  2953. 1:46:47Data pre-processing in deep learning
  2954. 1:46:49mainly has a method known as dropout
  2955. 1:46:52method.
  2956. 1:46:53Next what we're going to do is we're
  2957. 1:46:54going to uh drop out the entire time
  2958. 1:46:57column. We do not need the time of the
  2959. 1:47:00transaction in order to understand if
  2960. 1:47:02the transaction was fraudulent or not.
  2961. 1:47:05Right, so that's why we're getting rid
  2962. 1:47:06of unnecessary variables. Right, so
  2963. 1:47:08we're dropping out that variable.
  2964. 1:47:12So, after dropping out the time
  2965. 1:47:14variable, we're going to assign the
  2966. 1:47:16first 3,000 samples to our new data
  2967. 1:47:18frame. Right, this DF sample will have
  2968. 1:47:20our first 3,000 samples, and we're going
  2969. 1:47:24to use those 3,000 samples.
  2970. 1:47:26So, here we're just counting the number
  2971. 1:47:27of class for each of these samples.
  2972. 1:47:30After that, we're just counting the
  2973. 1:47:31number of samples for each of the class.
  2974. 1:47:33Like we're doing the same thing again
  2975. 1:47:34and here we get class zero has 2,508
  2976. 1:47:38samples and class one has 492 samples.
  2977. 1:47:41Now, this makes the data set quite
  2978. 1:47:43balanced, right? It's very balanced when
  2979. 1:47:45compared to our old data set.
  2980. 1:47:48Next, we'll just randomly shuffle our
  2981. 1:47:49data set, right? In order to remove any
  2982. 1:47:52sort of biasness in the data.
  2983. 1:47:54After that, we'll split our data set
  2984. 1:47:56into two parts. One is for training and
  2985. 1:47:59your other data set is for testing,
  2986. 1:48:01right? This is also known as data
  2987. 1:48:03splicing.
  2988. 1:48:05Then, we'll be splitting each data frame
  2989. 1:48:06into feature and label, meaning that
  2990. 1:48:09your input and your output.
  2991. 1:48:12We'll be doing this for your training
  2992. 1:48:13data and for your testing data, right?
  2993. 1:48:15All you're doing is you're separating
  2994. 1:48:17your input from your output.
  2995. 1:48:19Next, we're looking at our training data
  2996. 1:48:21set, right? We're printing the shape of
  2997. 1:48:23our training data set. The training data
  2998. 1:48:25set has around 2,400
  2999. 1:48:27observations and 29 variables or 29
  3000. 1:48:31features.
  3001. 1:48:32Similarly, we'll be printing out the
  3002. 1:48:34size of our test data frame, right?
  3003. 1:48:36That's exactly what we're doing over
  3004. 1:48:37here.
  3005. 1:48:38After that, we'll perform normalization,
  3006. 1:48:41right? For this, we'll be using the
  3007. 1:48:42min-max scaler.
  3008. 1:48:44So, in normalization, we'll basically be
  3009. 1:48:46scaling all our predictor variables
  3010. 1:48:48around the same range so that there is
  3011. 1:48:50no biasness in our prediction.
  3012. 1:48:53After this, we'll be plotting a function
  3013. 1:48:55for each of the learning curves. For
  3014. 1:48:57your training phase and for your testing
  3015. 1:48:59phase, you'll be plotting a learning
  3016. 1:49:01curve. Now, I'll show you the output of
  3017. 1:49:03this in a couple of minutes. For now,
  3018. 1:49:05let's move on to the main part, which is
  3019. 1:49:07model creation, right? In this demo,
  3020. 1:49:10we'll use three fully connected layers.
  3021. 1:49:13We'll also use dropout technique. Now,
  3022. 1:49:15dropout is a type of regularization
  3023. 1:49:18technique that is used to avoid any sort
  3024. 1:49:20of overfitting in a neural network. It
  3025. 1:49:23is a technique where you select neurons
  3026. 1:49:25and you drop them during the training
  3027. 1:49:26phase.
  3028. 1:49:27We'll be using the ReLU as the
  3029. 1:49:29activation function, which is a type of
  3030. 1:49:31activation function just like sigmoid
  3031. 1:49:33and tanh.
  3032. 1:49:35So, the type of model that we'll be
  3033. 1:49:36using is the sequential model. Right,
  3034. 1:49:39sequential is the easiest way to build a
  3035. 1:49:41model in Keras. Right, we're using the
  3036. 1:49:43Keras library over here. If you
  3037. 1:49:45remember, I imported that in the
  3038. 1:49:47beginning. Right, it allows you to build
  3039. 1:49:49a model layer by layer. So, each layer
  3040. 1:49:52has weights that correspond to the layer
  3041. 1:49:54that follows it. After this, you'll use
  3042. 1:49:57the add function to add the dense
  3043. 1:50:00layers. Basically, your hidden layers
  3044. 1:50:01you're going to add over here. So, in
  3045. 1:50:03our model, we'll be adding two dense
  3046. 1:50:05layers or hidden layers, you can say.
  3047. 1:50:08So, here what we're doing is we're
  3048. 1:50:10adding the first dense layer. Now guys,
  3049. 1:50:12a dense layer is standard layer type
  3050. 1:50:15that works for most cases. Right, in a
  3051. 1:50:18dense layer, all the nodes in the
  3052. 1:50:19previous layer connect to the nodes in
  3053. 1:50:21the current layer.
  3054. 1:50:23So guys, don't get too involved into
  3055. 1:50:25what exactly is happening here. All I'm
  3056. 1:50:27doing is I'm creating a sequential
  3057. 1:50:29model, and what is happening is I'm just
  3058. 1:50:32assigning the number of inputs for each
  3059. 1:50:34of the dense layer or for each of the
  3060. 1:50:36hidden layer. I'm also assigning dropout
  3061. 1:50:39value. Dropout is basically to prevent
  3062. 1:50:41overfitting. Overfitting might occur
  3063. 1:50:43when your model memorizes the training
  3064. 1:50:45data set. Overfitting basically reduces
  3065. 1:50:48the accuracy of a model. That's why
  3066. 1:50:50we're using the dropout method to
  3067. 1:50:52prevent overfitting. So, in the first
  3068. 1:50:54hidden layer, we have around 200 units.
  3069. 1:50:56Right, we have the activation function
  3070. 1:50:58ReLU. Then, we're adding the second
  3071. 1:51:01dense layer with again 200 neurons and
  3072. 1:51:03the ReLU activation function. Kernel
  3073. 1:51:06initializer is uniform, meaning that
  3074. 1:51:08it's just sequential and normal. Then,
  3075. 1:51:10we're again adding a dropout layer of
  3076. 1:51:120.5. The dropout value of a network has
  3077. 1:51:16to be chosen very wisely. Okay, a value
  3078. 1:51:19that is too low will result in a minimal
  3079. 1:51:21effect and a value that is too high will
  3080. 1:51:23result in under learning by the network.
  3081. 1:51:26So, 0.5 is a standard dropout value.
  3082. 1:51:30Now, this last layer is our output
  3083. 1:51:32layer. In the output layer, we'll
  3084. 1:51:33obviously have only one neuron. We'll
  3085. 1:51:35have one neuron that will show us the
  3086. 1:51:37output class, either zero or one. Zero
  3087. 1:51:41will show us non-fraudulent transactions
  3088. 1:51:43and one will denote fraudulent
  3089. 1:51:45transaction. Right, that's why we have
  3090. 1:51:46only one neuron over here. And the
  3091. 1:51:49activation function here is sigmoid.
  3092. 1:51:51Right, since the number of neurons is
  3093. 1:51:52only one.
  3094. 1:51:54After that, we're printing the model
  3095. 1:51:55summary. Now, I'll show you the summary
  3096. 1:51:58and everything. Before that, let us
  3097. 1:51:59understand what exactly optimization
  3098. 1:52:02functions are. We'll understand what
  3099. 1:52:04this optimization function does. Now, an
  3100. 1:52:06optimizer takes care of the necessary
  3101. 1:52:09computations that are used to change the
  3102. 1:52:12network's weights and bias. So,
  3103. 1:52:14basically, your optimizers will take
  3104. 1:52:16care of all your computations such as
  3105. 1:52:19changing the weight or updating the
  3106. 1:52:21weight. If you all remember, I spoke
  3107. 1:52:22about backpropagation, right? Where
  3108. 1:52:24you'll update the weight and all of
  3109. 1:52:25that. That is done by using optimizers.
  3110. 1:52:28Here, we're selecting an optimizer known
  3111. 1:52:30as the Adam optimizer.
  3112. 1:52:33So, Adam optimizer is one of the current
  3113. 1:52:35default optimizers in deep learning.
  3114. 1:52:38Right, it stands for adaptive moment
  3115. 1:52:40estimation. We don't have to get into
  3116. 1:52:42the depth of all of this. Right, all of
  3117. 1:52:44these are predefined optimizers in our
  3118. 1:52:46Keras package itself. After this, we're
  3119. 1:52:48going to fit our model by using the
  3120. 1:52:50training features. We're also setting
  3121. 1:52:53200 epochs and also there's something
  3122. 1:52:56known as epochs and batch size. Right,
  3123. 1:52:58we're setting epochs as 200 and batch
  3124. 1:53:00size as 500. I'll tell you what exactly
  3125. 1:53:03this means. Now, batch sizes are
  3126. 1:53:05basically used so that we don't overfit
  3127. 1:53:08our model, right? We're going to
  3128. 1:53:09basically split our data set into 500
  3129. 1:53:12batches. So, our input will be going in
  3130. 1:53:14the form of batches, right? And our
  3131. 1:53:17batch size is 500 inputs per batch, and
  3132. 1:53:20we'll be going through 200 epochs.
  3133. 1:53:23Meaning that our training will iterate
  3134. 1:53:25200 times. This is basically the number
  3135. 1:53:27of times that training our model. All
  3136. 1:53:29right, that's what epoch and batch size
  3137. 1:53:31is. After that, we're just showing our
  3138. 1:53:34training history. I mean, just printing
  3139. 1:53:35the accuracy curve for our training
  3140. 1:53:37phase. We're also going to print our
  3141. 1:53:39loss curves for our training phase,
  3142. 1:53:41basically the error curves. And then
  3143. 1:53:43finally, we have the evaluation. Here,
  3144. 1:53:45we'll be testing our model by using our
  3145. 1:53:47testing data set.
  3146. 1:53:49Then we're finally printing the accuracy
  3147. 1:53:51on our testing data set. After that,
  3148. 1:53:54we're just going to plot a heat map,
  3149. 1:53:56which I'll be showing y'all. Let me just
  3150. 1:53:58show you the output.
  3151. 1:54:00So guys, in this entire line of code,
  3152. 1:54:01all we're doing is we're printing an
  3153. 1:54:03accuracy plot. All right, basically
  3154. 1:54:05we're printing a heat map. I'll show you
  3155. 1:54:07what the heat map looks like.
  3156. 1:54:10This is just to check the accuracy.
  3157. 1:54:12We're comparing all the correctly
  3158. 1:54:14predicted values to our incorrectly
  3159. 1:54:16predicted values. So, this is our
  3160. 1:54:19training history. Here, blue stands for
  3161. 1:54:21our training phase, and this is our
  3162. 1:54:22validation or our prediction stage.
  3163. 1:54:26That was our training curve, and this is
  3164. 1:54:28our loss curve. Now, when you compare it
  3165. 1:54:30to the actual validation stage, it's
  3166. 1:54:32quite similar, right? Meaning that our
  3167. 1:54:34model is doing pretty well. So guys,
  3168. 1:54:37this is the heat map that I was talking
  3169. 1:54:38about. This is basically going to give
  3170. 1:54:40us the class for each of our
  3171. 1:54:41predictions. All right, it basically
  3172. 1:54:43plots the classes that we correctly
  3173. 1:54:46predicted, right? Basically, for each
  3174. 1:54:48data point, it's just going to tell us
  3175. 1:54:49whether we predicted it correctly or
  3176. 1:54:51not. It's sort of a confusion matrix in
  3177. 1:54:53the form of a heat map. So guys, these
  3178. 1:54:56are all our epochs, basically the 200
  3179. 1:54:59iterations that we went through, right?
  3180. 1:55:01This is the 50th iteration is showing us
  3181. 1:55:03our loss, it's showing us our accuracy
  3182. 1:55:05as well. And here we have 88%, 90%, 92%.
  3183. 1:55:10Now, if you carefully look at the epoch
  3184. 1:55:12accuracy values, you see that as we
  3185. 1:55:15train our model even more, our accuracy
  3186. 1:55:17keeps increasing. Initially, our
  3187. 1:55:19accuracy was around 83, right? At epoch
  3188. 1:55:22number 15, our accuracy was around 83%.
  3189. 1:55:26But as we kept training our model a
  3190. 1:55:29little bit more, our accuracy kept
  3191. 1:55:31increasing. We have 90, we have 91, 94,
  3192. 1:55:3595, 96, and so on. All right, so
  3193. 1:55:38basically, the more you train your
  3194. 1:55:39model, the better it's going to be. So
  3195. 1:55:42guys, this was our entire demo. Now, in
  3196. 1:55:44the end, I'm printing out the false
  3197. 1:55:46positive rate and the false negative
  3198. 1:55:48rate. All of this basically denotes how
  3199. 1:55:50many of the data points was I correctly
  3200. 1:55:53able to predict as fraudulent and how
  3201. 1:55:55many did I predict wrongly. That's all
  3202. 1:55:58the false negative and the false
  3203. 1:56:00positive rate denotes. So guys, this was
  3204. 1:56:02the entire demo on deep learning. Now,
  3205. 1:56:06if you have any doubts regarding the
  3206. 1:56:07deep learning demo, please mention them
  3207. 1:56:09in the comment section and I will solve
  3208. 1:56:11your queries. All right, now let's look
  3209. 1:56:13at our last topic for the day, which is
  3210. 1:56:15natural language processing. Now, before
  3211. 1:56:18we understand what is natural language
  3212. 1:56:20processing, let's understand the need
  3213. 1:56:21for natural language processing and a
  3214. 1:56:24process known as text mining. Text
  3215. 1:56:26mining and natural language processing
  3216. 1:56:28are heavily correlated. All right, I'll
  3217. 1:56:29talk about both of these in the upcoming
  3218. 1:56:32slides. For now, let me tell you why we
  3219. 1:56:34need natural language processing or text
  3220. 1:56:36mining. So guys, the amount of data that
  3221. 1:56:38we're generating these days is
  3222. 1:56:40unbelievable. It is a known fact that
  3223. 1:56:42we're creating 2.5 quintillion bytes of
  3224. 1:56:46data every day, and this number is only
  3225. 1:56:48going to grow. With the evolution of
  3226. 1:56:50communication through social media, we
  3227. 1:56:52generate tons and tons of data. All
  3228. 1:56:55right, the numbers are on your screen.
  3229. 1:56:57So, basically, we post around 1.7
  3230. 1:56:59million pictures on Instagram per
  3231. 1:57:01minute. Right, I'm talking about post
  3232. 1:57:04per minute. All of these numbers are per
  3233. 1:57:06minute values. These are the amount of
  3234. 1:57:08tweets, 347,000
  3235. 1:57:10tweets per minute. Right, this is a lot
  3236. 1:57:13of data. We're generating data while
  3237. 1:57:15we're watching YouTube videos, when
  3238. 1:57:17we're sending emails, when we are
  3239. 1:57:19chatting, and all of that. Right, even
  3240. 1:57:21the IoT devices at our house, right, we
  3241. 1:57:24have Alexa all of This is generating a
  3242. 1:57:25lot of data. A single click on your
  3243. 1:57:27phone is generating a lot of data. Now,
  3244. 1:57:30not only that, out of all the data that
  3245. 1:57:32we generate, only 21% of the data is
  3246. 1:57:35structured and well formatted. Right,
  3247. 1:57:38the remaining of the data is
  3248. 1:57:39unstructured. And the major sources of
  3249. 1:57:42unstructured data include text messages
  3250. 1:57:44from WhatsApp, Facebook likes, comments
  3251. 1:57:46on Instagram, the bulk emails, and all
  3252. 1:57:49of this. Right, all of this accounts for
  3253. 1:57:51the unstructured data that we have
  3254. 1:57:53today. Now, the data we generate is used
  3255. 1:57:56to grow a business. So, by analyzing and
  3256. 1:57:58mining the data, we can add more value
  3257. 1:58:01to a business. This is exactly what
  3258. 1:58:04natural language processing and text
  3259. 1:58:06mining is all about. Text mining and NLP
  3260. 1:58:09is a subset of artificial intelligence,
  3261. 1:58:11wherein we try and understand the
  3262. 1:58:14natural language text that we get from
  3263. 1:58:16text messages and so on, in order to
  3264. 1:58:19derive useful insights and grow
  3265. 1:58:21businesses by using these insights. So,
  3266. 1:58:23what exactly is text mining? Text mining
  3267. 1:58:26is a process of deriving meaningful
  3268. 1:58:28insights or information from natural
  3269. 1:58:31language text. So, all the data that we
  3270. 1:58:34generate through text messages, emails,
  3271. 1:58:36and documents are written in natural
  3272. 1:58:38language text. Right, and we're going to
  3273. 1:58:40use text mining and natural language
  3274. 1:58:42processing to draw useful insights or
  3275. 1:58:44patterns from such data in order to grow
  3276. 1:58:47a business.
  3277. 1:58:48Now, let's understand where exactly do
  3278. 1:58:50we make use of natural language
  3279. 1:58:51processing and text mining? Now, have
  3280. 1:58:53you ever noticed that if you start
  3281. 1:58:55typing a word on Google, you immediately
  3282. 1:58:58get suggestions, right? This feature is
  3283. 1:59:00known as auto complete. It will
  3284. 1:59:02basically suggest the rest of the word
  3285. 1:59:04to you. We also have something known as
  3286. 1:59:05spam detection, right? Here's an example
  3287. 1:59:08of how Google recognizes this
  3288. 1:59:10misspelling Netflix and shows results
  3289. 1:59:13for the keyword that matches your
  3290. 1:59:14misspelling. Let me show you a couple of
  3291. 1:59:17more examples. We also have predictive
  3292. 1:59:20typing and spell checkers and features
  3293. 1:59:22like auto correct, email classification.
  3294. 1:59:25So, predictive typing and spell
  3295. 1:59:27checkers, all of these are applications
  3296. 1:59:30of natural language processing. All of
  3297. 1:59:32this basically involves processing the
  3298. 1:59:34natural language that we use and
  3299. 1:59:36deriving some useful information from
  3300. 1:59:38it, right? Or running businesses from
  3301. 1:59:40it. Netflix uses natural language
  3302. 1:59:42processing in a really good fashioned
  3303. 1:59:45way, right? It basically studies the
  3304. 1:59:47reviews that customer gives for a
  3305. 1:59:49particular movie and it tries to figure
  3306. 1:59:51out if that movie is good or bad
  3307. 1:59:53depending on the review. So, Netflix
  3308. 1:59:55actually uses NLP in a very interesting
  3309. 1:59:58manner. It tries to understand the type
  3310. 2:00:00of movies that a person likes by the way
  3311. 2:00:03a person has rated the movie or by the
  3312. 2:00:05way the person has reviewed a movie. So,
  3313. 2:00:08by understanding what type of review a
  3314. 2:00:10person is giving to a movie, Netflix
  3315. 2:00:12will recommend more movies that you
  3316. 2:00:15like. That's how important NLP has
  3317. 2:00:17become. Now, let's look at what exactly
  3318. 2:00:19NLP is. NLP, which also stands for
  3319. 2:00:22natural language processing, is a part
  3320. 2:00:24of computer science and artificial
  3321. 2:00:26intelligence, which deals with human
  3322. 2:00:28language. Right? It's basically the
  3323. 2:00:30process of processing natural language
  3324. 2:00:32in order to derive some useful
  3325. 2:00:34information from it. For those of you
  3326. 2:00:36who have studied natural language
  3327. 2:00:38processing or have heard of natural
  3328. 2:00:40language processing, there is a huge
  3329. 2:00:42confusion between text mining and
  3330. 2:00:44natural language processing. So, text
  3331. 2:00:46mining is the process of deriving high
  3332. 2:00:48quality information from text. But, the
  3333. 2:00:51overall goal is to turn the text into
  3334. 2:00:53data for analysis by using natural
  3335. 2:00:56language processing. So, basically text
  3336. 2:00:58mining is implemented by using natural
  3337. 2:01:01language processing techniques. Right?
  3338. 2:01:03There are various techniques in natural
  3339. 2:01:04language processing that can help us
  3340. 2:01:06perform text mining. That's how text
  3341. 2:01:08mining and natural language processing
  3342. 2:01:10are related. Natural language processing
  3343. 2:01:12is the techniques that are used to solve
  3344. 2:01:15the problem of text mining, text
  3345. 2:01:17analysis, and all of that. Let's look at
  3346. 2:01:19a couple more applications. Sentimental
  3347. 2:01:22analysis is one of the major
  3348. 2:01:23applications of natural language
  3349. 2:01:25processing. You see Twitter performs
  3350. 2:01:27sentimental analysis, Facebook, Google,
  3351. 2:01:30all of these perform sentimental
  3352. 2:01:31analysis. Sentimental analysis mainly
  3353. 2:01:34used to analyze social media content
  3354. 2:01:36that can help us determine the public
  3355. 2:01:38opinion on a certain topic. Then we have
  3356. 2:01:41chatbots. Now, chatbots use natural
  3357. 2:01:43language processing to convert human
  3358. 2:01:45language into desirable actions. We also
  3359. 2:01:48have machine translation. NLP is used in
  3360. 2:01:51machine translation by studying the
  3361. 2:01:52morphological analysis of each word and
  3362. 2:01:55translating it to another language.
  3363. 2:01:58Advertisement matching is also done
  3364. 2:01:59using NLP in order to recommend ads
  3365. 2:02:02based on your history. Right? These are
  3366. 2:02:04few of the applications of NLP. Now, let
  3367. 2:02:07me tell you the basic terminologies
  3368. 2:02:08under natural language processing. So,
  3369. 2:02:10tokenization is the most basic step in
  3370. 2:02:13natural language processing.
  3371. 2:02:14Tokenization means breaking down the
  3372. 2:02:17data into smaller chunks or tokens so
  3373. 2:02:20that they can be easily analyzed. So,
  3374. 2:02:22the first step is you'll break a complex
  3375. 2:02:24sentence into words, then you'll
  3376. 2:02:26understand the importance of each of the
  3377. 2:02:28word with respect to that sentence in
  3378. 2:02:31order to produce a structural
  3379. 2:02:33description on an input sentence. So,
  3380. 2:02:35for example, take this sentence. How
  3381. 2:02:37would I perform tokenizations on the
  3382. 2:02:39sentence?
  3383. 2:02:40Let's say that tokens are simple is a
  3384. 2:02:43sentence and I want to perform
  3385. 2:02:44tokenization on the sentence. This is
  3386. 2:02:47what I'm going to do. I'm going to split
  3387. 2:02:48the sentence into different words. I'm
  3388. 2:02:50going to understand each word with
  3389. 2:02:52respect to that sentence. Right? This is
  3390. 2:02:55done to simplify operations in natural
  3391. 2:02:57language processing. Right? It's always
  3392. 2:02:59simpler to analyze a single token
  3393. 2:03:02instead of analyzing an entire sentence.
  3394. 2:03:04Then we have something known as
  3395. 2:03:05stemming. Now look at this example.
  3396. 2:03:08Right here we have words such as
  3397. 2:03:10detection, detecting, detected, and
  3398. 2:03:12detections. We all know that the root
  3399. 2:03:15word for all of these words is detect.
  3400. 2:03:17So stemming algorithm basically does
  3401. 2:03:20that. It works by cutting off the end or
  3402. 2:03:23the beginning of the word and taking
  3403. 2:03:25into account a list of common prefixes
  3404. 2:03:28and suffixes that can be found in an
  3405. 2:03:30inflicted word. Stemming basically helps
  3406. 2:03:33us in analyzing a lot of words. We know
  3407. 2:03:36that detections, detected, and detection
  3408. 2:03:38basically mean the same thing. So all
  3409. 2:03:40we're doing is we're going to ease our
  3410. 2:03:42analysis by removing prefixes and
  3411. 2:03:44suffixes which not make sense. Right? We
  3412. 2:03:47just need to understand the
  3413. 2:03:48morphological analysis of the word.
  3414. 2:03:50Right? So that's why we're randomly
  3415. 2:03:51cutting the prefixes and suffixes in
  3416. 2:03:53such a way that we only get the
  3417. 2:03:55important part of the word. This is
  3418. 2:03:57called stemming. Now this cutting of
  3419. 2:03:59words can be successful in some
  3420. 2:04:02occasions, but not always. That is why
  3421. 2:04:04we say that stemming approach has a few
  3422. 2:04:08limitations. In order to get over these
  3423. 2:04:11limitations, we have a process known as
  3424. 2:04:13lemmatization. Right? Lemmatization on
  3425. 2:04:16the other hand takes into consideration
  3426. 2:04:18the morphological analysis of the words.
  3427. 2:04:21It does not randomly cut the word in the
  3428. 2:04:23beginning and the ending. It understands
  3429. 2:04:25what the word means and only then it
  3430. 2:04:27cuts the word. For example, let's
  3431. 2:04:29consider the word recap. If we perform
  3432. 2:04:32stemming on the word recap, we'll get
  3433. 2:04:35cap. Right? The output will be cap. But,
  3434. 2:04:38cap and recap do not have the same
  3435. 2:04:40meaning, do they? They have absolutely
  3436. 2:04:41different meanings. That's why stemming
  3437. 2:04:43is sometimes not considered to be the
  3438. 2:04:45right thing to do. But, when it comes to
  3439. 2:04:47lemmatization, it's going to understand
  3440. 2:04:49the meaning of recap. Only then will it
  3441. 2:04:52perform any sort of change in the word,
  3442. 2:04:54or it'll cut down the word. So,
  3443. 2:04:56basically, it groups together different
  3444. 2:04:58inflected forms of a word called lemma.
  3445. 2:05:01Lemmatization is similar to stemming
  3446. 2:05:03because it maps several words into one
  3447. 2:05:06common root. But, the output of a
  3448. 2:05:09lemmatization process is always a proper
  3449. 2:05:11word. An example of lemmatization is to
  3450. 2:05:15map gone, going, and went into go. Gone,
  3451. 2:05:18going, went, all of them mean go. So,
  3452. 2:05:21basically, by lemmatization, you can
  3453. 2:05:22just output the words as go. That is
  3454. 2:05:25what lemmatization is. Next, we have
  3455. 2:05:27something known as stop words, right?
  3456. 2:05:29Stop words are basically a set of
  3457. 2:05:31commonly used words in any language,
  3458. 2:05:34right? Not just English, any language.
  3459. 2:05:36The reason why stop words are critical
  3460. 2:05:38to many applications is that if we
  3461. 2:05:41remove the words that are very commonly
  3462. 2:05:44used in a given language, we can finally
  3463. 2:05:46focus on the important words. For
  3464. 2:05:48example, in the context of Let's say you
  3465. 2:05:51open up Google and you look for
  3466. 2:05:53strawberry milkshake recipe. Instead of
  3467. 2:05:55typing strawberry milkshake recipe,
  3468. 2:05:57let's say you type how to make
  3469. 2:05:59strawberry milkshake. Now, here, what
  3470. 2:06:02Google will do is it'll find results for
  3471. 2:06:04how, to, and make. Instead, if you just
  3472. 2:06:07type strawberry milkshake recipe, you'll
  3473. 2:06:10get the most desired output. That's why
  3474. 2:06:13it's always considered a good practice
  3475. 2:06:15in natural language processing to get
  3476. 2:06:17rid of stop words, right? Stop words
  3477. 2:06:19will just increase our computation, and
  3478. 2:06:21it'll just add additional work to us.
  3479. 2:06:23They are not very helpful when we're
  3480. 2:06:25analyzing important documents, right? We
  3481. 2:06:27need to focus on the important keywords
  3482. 2:06:29in the documents instead of all of these
  3483. 2:06:31commonly used words. Example of stop
  3484. 2:06:34words include the, how, when, why, not,
  3485. 2:06:38yes, no. All of these are stop words,
  3486. 2:06:41right? So, in order to better analyze
  3487. 2:06:43our data, we need to get rid of stop
  3488. 2:06:45words. Now, the last terminology I'm
  3489. 2:06:47going to discuss is document term
  3490. 2:06:49matrix. It is important to create
  3491. 2:06:51something known as the document term
  3492. 2:06:53matrix in natural language processing. A
  3493. 2:06:56DTM or a document term matrix is
  3494. 2:06:59basically a matrix that shows the
  3495. 2:07:01frequency of words in a particular
  3496. 2:07:03document. Let's say that we're trying to
  3497. 2:07:05understand if the sentence this is fun
  3498. 2:07:09is available in one of my documents.
  3499. 2:07:12So, if it is there in my document one,
  3500. 2:07:14I'm going to put a one corresponding to
  3501. 2:07:16each of the words that is available in
  3502. 2:07:18my document. For example, in document
  3503. 2:07:20two, I have this is, but I do not have
  3504. 2:07:23the word fun.
  3505. 2:07:24Similarly, in document four, I have the
  3506. 2:07:27word this, but I do not have the word is
  3507. 2:07:29and fun. So, basically, a document term
  3508. 2:07:31matrix is like the frequency matrix of a
  3509. 2:07:34document. So, during text analysis, you
  3510. 2:07:37always begin by building a document term
  3511. 2:07:39matrix, right? Here, you try to
  3512. 2:07:40understand which words frequently occur
  3513. 2:07:43and which words are important and not
  3514. 2:07:45important in the document. So, guys,
  3515. 2:07:47these were a couple of terminologies in
  3516. 2:07:49natural language processing.
  3517. 2:07:57Artificial intelligence and machine
  3518. 2:07:59learning are not just trending
  3519. 2:08:00technologies anymore.
  3520. 2:08:02They're becoming the backbone of every
  3521. 2:08:04industry. And right now, the demand for
  3522. 2:08:06AI and ML engineers is exploding
  3523. 2:08:09worldwide. In India, AI engineers earn
  3524. 2:08:12anywhere from 8 lakhs to 36 lakhs per
  3525. 2:08:15year, depending on skills and
  3526. 2:08:17experience. In the United States, the
  3527. 2:08:19same roles can start from $120,000 and
  3528. 2:08:23can go all the way up to $250,000 for
  3529. 2:08:26senior and specialized positions.
  3530. 2:08:28So, if you have been thinking about
  3531. 2:08:30getting into AI and ML, switching
  3532. 2:08:32careers, or upskilling for higher-paying
  3533. 2:08:35opportunities, there has never been a
  3534. 2:08:37better time. And in this video, I am
  3535. 2:08:39giving you a complete,
  3536. 2:08:41beginner-friendly, and deeply practical
  3537. 2:08:43AI and ML engineer roadmap that shows
  3538. 2:08:46you exactly what to learn and how to
  3539. 2:08:48grow in this booming field.
  3540. 2:08:50And now, the first step of your AI
  3541. 2:08:53journey starts with a strong
  3542. 2:08:55foundations.
  3543. 2:08:56And no, you don't need to be a
  3544. 2:08:57mathematician. You simply need the
  3545. 2:09:00essentials. So, start with Python,
  3546. 2:09:02because Python is the language that
  3547. 2:09:04powers almost every modern AI system.
  3548. 2:09:07So, focus on the basics like variables,
  3549. 2:09:10loops, functions, lists, dictionaries,
  3550. 2:09:14file handling, and how to work with
  3551. 2:09:16APIs.
  3552. 2:09:17So, these skills are enough to write
  3553. 2:09:18simple programs and understand AI code.
  3554. 2:09:21Next, learn data handling, because AI is
  3555. 2:09:25built on data. Use Pandas to clean and
  3556. 2:09:27organize data, NumPy to perform
  3557. 2:09:30calculation, and Matplotlib or Seaborn
  3558. 2:09:33to visualize patterns. Even simple tasks
  3559. 2:09:36like removing missing values or
  3560. 2:09:38analyzing sales trends will prepare you
  3561. 2:09:40for the real AI projects.
  3562. 2:09:42Then, learn the essential math behind
  3563. 2:09:44AI. Not heavy equations, just the basic
  3564. 2:09:47understanding. Understand mean, median,
  3565. 2:09:50variance, probability basic,
  3566. 2:09:52correlations, and what vectors and
  3567. 2:09:54matrices are. Learn what gradient
  3568. 2:09:57descent means conceptually, so you
  3569. 2:09:59understand how models learn without
  3570. 2:10:01getting buried in complex math. So, once
  3571. 2:10:03your foundations are ready, move to
  3572. 2:10:05machine learning. ML is simply teaching
  3573. 2:10:08computers to learn from examples.
  3574. 2:10:10Instead of writing instructions, show
  3575. 2:10:12the model real-world data and let it
  3576. 2:10:14find patterns.
  3577. 2:10:16So, learn key ML concepts like training
  3578. 2:10:18and testing, accuracy and precision,
  3579. 2:10:21underfitting and overfitting, cross
  3580. 2:10:23validation, and feature engineering. So,
  3581. 2:10:26these concepts helps you understand how
  3582. 2:10:28to build, tune, and improve models.
  3583. 2:10:31Then, learn the core email algorithms
  3584. 2:10:33that companies use every single day. So,
  3585. 2:10:36you need to start with linear
  3586. 2:10:38regression, logistic regression,
  3587. 2:10:40decision trees, random forest, SVM,
  3588. 2:10:44naive Bayes, K-means clustering, and
  3589. 2:10:46PCA.
  3590. 2:10:47These algorithms cover most practical
  3591. 2:10:49business problems like predicting sales,
  3592. 2:10:52detecting fraud, segmenting customers,
  3593. 2:10:55and identifying patterns in large data
  3594. 2:10:57sets.
  3595. 2:10:58Then, build small email projects such as
  3596. 2:11:00house price predictor, spam email
  3597. 2:11:03classifier, credit score predictor, or
  3598. 2:11:06customer segmentation model. So, these
  3599. 2:11:08projects give you confidence and make
  3600. 2:11:10your portfolio job ready. After ML, move
  3601. 2:11:13into deep learning, the technology
  3602. 2:11:15behind ChatGPT, self-driving cars, and
  3603. 2:11:18medical AI.
  3604. 2:11:19Start by understanding how neural
  3605. 2:11:21networks work. Learn what neurons,
  3606. 2:11:24layers, activation functions, and loss
  3607. 2:11:26functions are.
  3608. 2:11:28You don't need to memorize formulas,
  3609. 2:11:30just understand how the network adjusts
  3610. 2:11:32itself to improve predictions.
  3611. 2:11:34Pick either TensorFlow or PyTorch as
  3612. 2:11:37your deep learning framework because
  3613. 2:11:39both are used by companies in
  3614. 2:11:40production, so you only need to choose
  3615. 2:11:42one. And then, build deep learning
  3616. 2:11:44projects like digit recognition, image
  3617. 2:11:46classification, sentiment analysis, or
  3618. 2:11:49go for fake news classification. So,
  3619. 2:11:51these projects teach you how to use
  3620. 2:11:53neural networks in real scenarios.
  3621. 2:11:56So, AI is no longer about learning
  3622. 2:11:59everything. It's about choosing your
  3623. 2:12:01specialization.
  3624. 2:12:02So, here we have options. So, option A
  3625. 2:12:05is NLP and LLMs, which has the highest
  3626. 2:12:08demand. So, if you want to work with
  3627. 2:12:10chatbots, smart assistants, or language
  3628. 2:12:13models like ChatGPT, choose NLP and
  3629. 2:12:16LLMs. And all you need to learn is
  3630. 2:12:19tokenization, embeddings, transformers,
  3631. 2:12:22BERT, GPT models, prompt engineering,
  3632. 2:12:25fine-tuning, RAG, and agentic
  3633. 2:12:28architectures.
  3634. 2:12:29And you can build projects like AI
  3635. 2:12:31chatbots, document search tools,
  3636. 2:12:34question and answer systems, or customer
  3637. 2:12:36support bots.
  3638. 2:12:37Option B is computer vision. If you like
  3639. 2:12:40working with images and videos, choose
  3640. 2:12:42computer vision. Learn CNNs, YOLO,
  3641. 2:12:45object detection, and segmentation. And
  3642. 2:12:48you can build real-world projects like
  3643. 2:12:50face detection, medical image analysis,
  3644. 2:12:53CCTV monitoring systems, or vehicle
  3645. 2:12:55counting tools. Option C is generative
  3646. 2:12:58AI. If you enjoy creativity, choose
  3647. 2:13:01generative AI. Learn GANs, VAEs, and
  3648. 2:13:05diffusion models. Build applications
  3649. 2:13:07like AI art generators, product design
  3650. 2:13:10tools, image-to-image systems, or video
  3651. 2:13:12generation models. Option D is MLOps. If
  3652. 2:13:16you prefer infrastructure and
  3653. 2:13:18deployment, then choose MLOps. Learn
  3654. 2:13:21Docker, Kubernetes, MLflow, CI/CD
  3655. 2:13:23pipelines, cloud deployment, and model
  3656. 2:13:26monitoring. And you can build projects
  3657. 2:13:28that focus on deploying ML and LLM
  3658. 2:13:31models into real environments. All
  3659. 2:13:33right. So, agentic AI is the biggest
  3660. 2:13:36trend of 2026.
  3661. 2:13:38These are not just models, these are the
  3662. 2:13:41intelligent agents that can reason,
  3663. 2:13:43plan, use tools, and take actions. So,
  3664. 2:13:46learn how agents work with frameworks
  3665. 2:13:48like LangChain, LangGraph, and
  3666. 2:13:51crew-based agent architectures.
  3667. 2:13:53Understand tool calling, memory systems,
  3668. 2:13:56planning, and multi-agent collaboration.
  3669. 2:13:59Also, build agentic projects like an AI
  3670. 2:14:02research assistant, an autonomous email
  3671. 2:14:04automation agent, a financial analysis
  3672. 2:14:07agent, or a customer service automation
  3673. 2:14:10agent. So, these projects stand out in
  3674. 2:14:13the interviews because companies want
  3675. 2:14:15people who can build intelligent
  3676. 2:14:17workflows and not just models.
  3677. 2:14:19So, now that you have skills, you need
  3678. 2:14:21projects that prove it. So, your
  3679. 2:14:23portfolio should have two machine
  3680. 2:14:25learning projects, two deep learning
  3681. 2:14:27projects, two specialization projects,
  3682. 2:14:29and one real end-to-end AI system. And
  3683. 2:14:32this final project could be a chatbot
  3684. 2:14:34with rack, a vision-based attendance
  3685. 2:14:36system, an AI assistant with memory, or
  3686. 2:14:40a complete ML pipeline deployed on
  3687. 2:14:43cloud. And finally, upload your work on
  3688. 2:14:46GitHub, write clear documentation, and
  3689. 2:14:48add deployment links so employers can
  3690. 2:14:51test your work instantly. So, with this
  3691. 2:14:53roadmap, you can apply for the most
  3692. 2:14:55in-demand roles in 2026,
  3693. 2:14:58such as AI engineer, machine learning
  3694. 2:15:00engineer, LLM engineer, NLP engineer,
  3695. 2:15:04generative AI engineer, computer vision
  3696. 2:15:06engineer, MLOps engineer, or AI
  3697. 2:15:09automation specialist.
  3698. 2:15:11AI is not just the future, it's the
  3699. 2:15:14career shift of today. If you follow
  3700. 2:15:16this roadmap step by step, you will
  3701. 2:15:18build the skills, the projects, and the
  3702. 2:15:20confidence to enter the world of AI and
  3703. 2:15:23machine learning.
  3704. 2:15:27>> [music]
  3705. 2:15:30>> So guys, let's see what we are going to
  3706. 2:15:32explore today. So today, we are going to
  3707. 2:15:34explore real-world examples of machine
  3708. 2:15:36learning, starting your journey, how you
  3709. 2:15:38can start your journey to machine
  3710. 2:15:40learning, and key concepts and impacts
  3711. 2:15:42of machine learning in your day-to-day
  3712. 2:15:44life. So guys, to help you navigate
  3713. 2:15:47through this video, here is a quick
  3714. 2:15:49rundown for you guys. Table of contents.
  3715. 2:15:51What is machine learning? How does
  3716. 2:15:53machine learning works? Five features of
  3717. 2:15:55machine learning, types of machine
  3718. 2:15:57learning, what skills one should have to
  3719. 2:15:59learn machine learning, machine learning
  3720. 2:16:01applications, and last but not the least
  3721. 2:16:04guys, that is future of machine
  3722. 2:16:06learning.
  3723. 2:16:07So guys, I'm going to amaze you with
  3724. 2:16:10this best example of Google Translator
  3725. 2:16:12and AI which converts from one language
  3726. 2:16:15to the another. I have chosen one
  3727. 2:16:17language for you guys since many of you
  3728. 2:16:19watch animes and all. So, I'm going to
  3729. 2:16:21convert from Japanese language to the
  3730. 2:16:24English language. Buckle up. Let's get
  3731. 2:16:26started. So guys, I'm going to write a
  3732. 2:16:29word in Japanese that is "Ohayo
  3733. 2:16:32gozaimasu".
  3734. 2:16:34If you know what "Ohayo gozaimasu"
  3735. 2:16:36means, then please comment down below.
  3736. 2:16:38So, I'm going to convert "Ohayo
  3737. 2:16:40gozaimasu" from Japanese to English. So,
  3738. 2:16:43let's copy this and paste this in here.
  3739. 2:16:46So, "Ohayo gozaimasu" in Japanese, but
  3740. 2:16:50in English it means good morning. A
  3741. 2:16:52Google Translator works on a neural
  3742. 2:16:54network. And a neural network is a
  3743. 2:16:57machine learning algorithm which learns
  3744. 2:16:59from Japanese language, keeps on
  3745. 2:17:01improving itself, and gives you the
  3746. 2:17:04optimal output that is good morning in
  3747. 2:17:06English. This is how a language
  3748. 2:17:08translator works. Let's see one more
  3749. 2:17:11example. If I write here one more thing
  3750. 2:17:14that is "Hajimemashite".
  3751. 2:17:18"Hajimemashite" in Japanese, if you
  3752. 2:17:20know, then please comment down below.
  3753. 2:17:22Let me see what does it mean.
  3754. 2:17:25So, "Hajimemashite" in Japanese, but in
  3755. 2:17:27English it means nice to meet you. Here
  3756. 2:17:30also the same. The neural network is
  3757. 2:17:32learning from Japanese language and
  3758. 2:17:34showing you what does it mean in English
  3759. 2:17:37language. So, let me properly explain
  3760. 2:17:40you what it actually does. So guys, a
  3761. 2:17:43translator is nothing, but it's an AI
  3762. 2:17:45machine that converts from one language
  3763. 2:17:47to the another using a machine learning
  3764. 2:17:50algorithm called neural network. Now,
  3765. 2:17:52what neural network does is it learns
  3766. 2:17:54from one language, makes some mistake or
  3767. 2:17:56errors, then keeps on improving itself
  3768. 2:17:58so that it can convert to to language
  3769. 2:18:00such as from Japanese to English or from
  3770. 2:18:03English to Japanese.
  3771. 2:18:04So guys, the translator which you are
  3772. 2:18:06using works on neural networks and this
  3773. 2:18:09is how a translation works.
  3774. 2:18:13So guys, you must be thinking then what
  3775. 2:18:15is machine learning? So moving on to
  3776. 2:18:17what is machine learning we have a
  3777. 2:18:19machine learning is nothing but a
  3778. 2:18:21training from historical data or
  3779. 2:18:23experiences or in layman terms you can
  3780. 2:18:25say learning from data to predict future
  3781. 2:18:28or required output is called machine
  3782. 2:18:30learning.
  3783. 2:18:31Guys, a simple machine learning
  3784. 2:18:33algorithm works in a way that you have a
  3785. 2:18:35data of any kind and you give it to a
  3786. 2:18:37machine. Now what does machine does is
  3787. 2:18:40it learns from it in different ways
  3788. 2:18:42using some different algorithms and
  3789. 2:18:44gives you the required amount of future
  3790. 2:18:47or predicts the required amount of
  3791. 2:18:48output you wanted.
  3792. 2:18:51Now guys, you must have understood what
  3793. 2:18:53is machine learning. Let's deep dive a
  3794. 2:18:55bit. Let's see what are the features of
  3795. 2:18:57machine learning. We have five features
  3796. 2:18:59of machine learning that is predictive
  3797. 2:19:01modeling, automation, scalability,
  3798. 2:19:05generalization, and adaptiveness. These
  3799. 2:19:07are the five main features of any
  3800. 2:19:09machine learning.
  3801. 2:19:11So guys, tighten your seat belts. Let's
  3802. 2:19:13move to these one by one. Predictive
  3803. 2:19:15model. In the predictive model, what it
  3804. 2:19:17does it it uses some mathematical
  3805. 2:19:19functions and statistical techniques on
  3806. 2:19:21the historical data and gives you the
  3807. 2:19:24future predictions. The best example I
  3808. 2:19:26can give you guys is the stock market
  3809. 2:19:28prediction app where it uses some
  3810. 2:19:29graphs, straight line graphs, and charts
  3811. 2:19:32to show you the prediction based on the
  3812. 2:19:34historical data using some statistical
  3813. 2:19:36and mathematical functions. This is how
  3814. 2:19:39a predictive model works.
  3815. 2:19:42Moving on to our next topic that is
  3816. 2:19:44automation. Automation is one of the
  3817. 2:19:46best feature to save money. You know
  3818. 2:19:48why? Because the companies which are
  3819. 2:19:51having the less domains and cannot hire
  3820. 2:19:53the employees, they can automate
  3821. 2:19:55different machines to do the same work
  3822. 2:19:58as the employee does. Such as if you
  3823. 2:20:00want a developer, but you don't have the
  3824. 2:20:02cost to pay, then you can automate a
  3825. 2:20:04machine that can develop for you. This
  3826. 2:20:06is the best example I can give you for
  3827. 2:20:08the automation feature.
  3828. 2:20:11So guys, moving on to our next feature,
  3829. 2:20:13that is scalability. Scalability has its
  3830. 2:20:15own importance because a machine
  3831. 2:20:17learning algorithm, if it is not
  3832. 2:20:19scalable, then it cannot handle larger
  3833. 2:20:21amount of data sets or bigger data sets.
  3834. 2:20:24I can give you the best example, that is
  3835. 2:20:26Amazon.
  3836. 2:20:28So guys, here I am at the Amazon
  3837. 2:20:30website, and you can see a lot of
  3838. 2:20:32product, and not only you can see, but
  3839. 2:20:34whole world can see who are using Amazon
  3840. 2:20:36app. Now, this system is a scalable
  3841. 2:20:38system. No matter how many customers are
  3842. 2:20:41here, and they are buying, the system
  3843. 2:20:43will never crash. It can handle that
  3844. 2:20:45amount of larger data sets. So, this is
  3845. 2:20:48the best example of a scalable system.
  3846. 2:20:51So guys, moving on to our next feature,
  3847. 2:20:53that is generalization. In
  3848. 2:20:55generalization, what it does is it's the
  3849. 2:20:58ability of the model to generalize
  3850. 2:21:01things, to forecast new data. Suppose
  3851. 2:21:03your model is trained on a data set, and
  3852. 2:21:06you're going to test it on some another
  3853. 2:21:08data set, which is not there in the
  3854. 2:21:09training, but still your model is giving
  3855. 2:21:1295% of accuracy. Means your model is
  3856. 2:21:15generalizing, your model is summarizing,
  3857. 2:21:17and giving you the best and optimal
  3858. 2:21:19result.
  3859. 2:21:20Now guys, moving on to our last feature,
  3860. 2:21:22that is adaptiveness. You can take it as
  3861. 2:21:25a survival of the fittest thing, because
  3862. 2:21:26if your model is not surviving the
  3863. 2:21:28real-time environments or the new
  3864. 2:21:30problems, then your model is not
  3865. 2:21:32adaptive or good. It will going to
  3866. 2:21:33extinct. Suppose there is a model which
  3867. 2:21:36is built on traditional model, and still
  3868. 2:21:38giving you best and advanced solutions
  3869. 2:21:40on the real-time problems, then your
  3870. 2:21:42model is adaptive and is the optimal
  3871. 2:21:44model you can have.
  3872. 2:21:46So guys, clear your mind because we are
  3873. 2:21:48going to go in the types of machine
  3874. 2:21:50learning. There are four different types
  3875. 2:21:52of machine learning. First, we have is
  3876. 2:21:54supervised or guided machine learning.
  3877. 2:21:57In the supervised or guided machine
  3878. 2:21:58learning, what it does is if there is a
  3879. 2:22:00data set and having some values and you
  3880. 2:22:03are labeling it as a specifying the
  3881. 2:22:05value to the machine, then it will
  3882. 2:22:07recognize those data through the labels
  3883. 2:22:10and giving you the optimal
  3884. 2:22:11classification. For example, if you have
  3885. 2:22:14the pictures of cats and dogs and you
  3886. 2:22:16have to classify it, then you will label
  3887. 2:22:18it as cats and dogs. Then your machine
  3888. 2:22:20will recognize those labels and classify
  3889. 2:22:23and give you the optimal result.
  3890. 2:22:26So, in supervised learning, what we have
  3891. 2:22:28is a supervised algorithm. We have some
  3892. 2:22:30labels and we are putting those labels
  3893. 2:22:32to a data set. And after this, the
  3894. 2:22:35machine easily recognizes and giving you
  3895. 2:22:37the optimal results.
  3896. 2:22:39So guys, moving on to our next type,
  3897. 2:22:41that is unsupervised or unguided
  3898. 2:22:44learning. In unsupervised or unguided
  3899. 2:22:46learning, what machine does is you are
  3900. 2:22:48giving the data which is not labeled.
  3901. 2:22:50And after few trainings and making some
  3902. 2:22:52errors, the machine easily recognizes
  3903. 2:22:54this. So, what we have is a unsupervised
  3904. 2:22:57algorithm. We have some unlabeled data
  3905. 2:23:00and in those unlabeled data, your
  3906. 2:23:02machine is trying to find patterns.
  3907. 2:23:03After few errors, it will give you the
  3908. 2:23:06optimal results.
  3909. 2:23:07Moving on to our next type, that is
  3910. 2:23:09semi-supervised learning. In
  3911. 2:23:11semi-supervised learning, the algorithm
  3912. 2:23:13uses both unsupervised and supervised in
  3913. 2:23:16combined form giving you the optimal
  3914. 2:23:18result. For example, we have a
  3915. 2:23:20semi-supervised algorithm and we are
  3916. 2:23:22trying to find out patterns from the
  3917. 2:23:24data sets which are both labeled and
  3918. 2:23:26unlabeled. So, this is how a
  3919. 2:23:28semi-supervised learning works.
  3920. 2:23:30Moving on to our last type, that is
  3921. 2:23:33reinforcement learning. The
  3922. 2:23:34reinforcement learning you can
  3923. 2:23:36understand in a way like when you are
  3924. 2:23:37playing a game, you make some mistake,
  3925. 2:23:39then learn from them, then again make
  3926. 2:23:41some mistake in a level, learn from them
  3927. 2:23:43and reach your goal. This This how
  3928. 2:23:45reinforcement learning works in a
  3929. 2:23:47software. The software is being trained
  3930. 2:23:49multiple times making some errors and
  3931. 2:23:51learning from them and giving you the
  3932. 2:23:53optimal results. This is the best
  3933. 2:23:55example I can give you for the
  3934. 2:23:57reinforcement learning that is the
  3935. 2:23:58gaming system. So guys, what we have is
  3936. 2:24:01a reinforcement algorithm and an
  3937. 2:24:03environment. We are taking some actions
  3938. 2:24:05in that environment, making some errors
  3939. 2:24:07and getting some results. Then again
  3940. 2:24:09making some errors and getting some
  3941. 2:24:11results. This is how a reinforcement
  3942. 2:24:13model works in a machine learning.
  3943. 2:24:16So guys, moving on to the examples of
  3944. 2:24:18machine learning, we have a voice
  3945. 2:24:19recognition system and a image
  3946. 2:24:22recognition system. These both you are
  3947. 2:24:24using in your phone, in your laptop
  3948. 2:24:26every day and machine learning is being
  3949. 2:24:28used. Now guys, you have learned so
  3950. 2:24:30much. Now you must be thinking that how
  3951. 2:24:33should I start my journey? What skills I
  3952. 2:24:35should have? So these are the skills
  3953. 2:24:38required to learn machine learning. That
  3954. 2:24:40is SQL, structured query language,
  3955. 2:24:42JavaScript, C++, R programming for those
  3956. 2:24:46who are moving with machine learning to
  3957. 2:24:47the data scientist and Python, one of
  3958. 2:24:50the best programming language for the
  3959. 2:24:51machine learning. And last but not the
  3960. 2:24:53least, if you are moving a bit deep down
  3961. 2:24:56in the machine learning towards the deep
  3962. 2:24:57learning, then you need NLP or natural
  3963. 2:25:01language processing. These skills are
  3964. 2:25:03required for the one who want to learn
  3965. 2:25:05machine learning.
  3966. 2:25:07So guys, moving on to the real-life
  3967. 2:25:09examples of machine learning I have,
  3968. 2:25:11that is Google searches. Google searches
  3969. 2:25:14uses our history, track it down, then
  3970. 2:25:16giving you the prediction based on your
  3971. 2:25:18histories. For example, let me show you.
  3972. 2:25:21So guys, I came here at Google. Now I'm
  3973. 2:25:24going to type Amazon and it's giving me
  3974. 2:25:28that results which are already there in
  3975. 2:25:30my history. I must have searched before
  3976. 2:25:32a month ago, uh 2 months ago, then 4
  3977. 2:25:34months ago and all that history combined
  3978. 2:25:37form is being predicted in here. Now
  3979. 2:25:39Amazon Prime is there, Amazon videos are
  3980. 2:25:42there. So, this is how a predictive
  3981. 2:25:44model is working behind this Google
  3982. 2:25:45searches.
  3983. 2:25:47Moving on to our next real-life example,
  3984. 2:25:49guys, that is Instagram. You swipe
  3985. 2:25:52Instagram every day, every night. So,
  3986. 2:25:54all that is based on your past data
  3987. 2:25:57only. If you are seeing some videos of
  3988. 2:25:58cats, then further swipes will be the
  3989. 2:26:00cats only. So, this is how your history
  3990. 2:26:03is being tracked down and watched by the
  3991. 2:26:05Instagram machine learning algorithms
  3992. 2:26:07and giving you the results of the same.
  3993. 2:26:10So, this is how a real feeder in
  3994. 2:26:13Instagram works. Moving on to our next
  3995. 2:26:15example, that is movie recommendation
  3996. 2:26:17system on any movie watching website
  3997. 2:26:19such as Netflix or anime websites such
  3998. 2:26:22as Watch Anime, Anycon, etc.
  3999. 2:26:25So, guys, if I take you to Any watch and
  4000. 2:26:28show you the anime recommendation
  4001. 2:26:30trending ones are based on the people's
  4002. 2:26:32choices they are watching more based on
  4003. 2:26:34your histories only. If I watch any of
  4004. 2:26:36these animes and I watch them regularly,
  4005. 2:26:39then the recommendations will show me
  4006. 2:26:41the same. So, this is how a
  4007. 2:26:43recommendation system works in Any watch
  4008. 2:26:46or you can say in Netflix based on your
  4009. 2:26:48choices in past data.
  4010. 2:26:51So, guys, moving on to the future of
  4011. 2:26:53machine learning, it will be using
  4012. 2:26:55everywhere. The advanced techniques in
  4013. 2:26:57medical field, architecture field,
  4014. 2:26:59electrical field, and yes, in the
  4015. 2:27:01computer science field. Machine learning
  4016. 2:27:03will be everywhere. It will be
  4017. 2:27:04revolutionizing the world.
  4018. 2:27:12Now, let's explore different types of ML
  4019. 2:27:14models.
  4020. 2:27:16So, not all data is structured the same
  4021. 2:27:18way. And different problems require
  4022. 2:27:19different approaches.
  4023. 2:27:21So, for example, predicting stock prices
  4024. 2:27:23requires the models that learn from
  4025. 2:27:25historical trends.
  4026. 2:27:26And then identifying objects in images
  4027. 2:27:29needs models that recognize patterns in
  4028. 2:27:31visual data.
  4029. 2:27:32Next, the chatbots and voice assistants
  4030. 2:27:35rely on the models trained to understand
  4031. 2:27:37and generate human language.
  4032. 2:27:39So, to tackle these challenges, as I
  4033. 2:27:41discussed previously, that ML is divided
  4034. 2:27:43into different learning models, such as
  4035. 2:27:45supervised, unsupervised, and
  4036. 2:27:48reinforcement learning. And each has its
  4037. 2:27:50own strengths, and it is used depending
  4038. 2:27:52on the problem at hand.
  4039. 2:27:54Since we know why different ML models
  4040. 2:27:56are needed, let's see how they play a
  4041. 2:27:58crucial role in generative AI.
  4042. 2:28:00Well, generative AI is one of the most
  4043. 2:28:03exciting applications of machine
  4044. 2:28:04learning. And unlike traditional ML
  4045. 2:28:07models that make predictions or
  4046. 2:28:08classifications, generative models
  4047. 2:28:10create entirely new content. And here's
  4048. 2:28:13how ML enables AI to generate.
  4049. 2:28:15So, first here we have text. A language
  4050. 2:28:18models, like GPT, generate human-like
  4051. 2:28:20text for chatbots, content writing, and
  4052. 2:28:22coding.
  4053. 2:28:23Next is the image.
  4054. 2:28:25So, AI-powered tools, like DALL-E, can
  4055. 2:28:27create realistic images from textual
  4056. 2:28:30descriptions.
  4057. 2:28:31Next is videos. So, advanced ML models
  4058. 2:28:34synthesize lifelike video content,
  4059. 2:28:37transforming media, marketing, and even
  4060. 2:28:39filmmaking.
  4061. 2:28:40So, these advancements in generative AI
  4062. 2:28:42are reshaping creativity and automation,
  4063. 2:28:45proving that machine learning is not
  4064. 2:28:46just about making decision, it's about
  4065. 2:28:48creating new possibilities.
  4066. 2:28:50So, now that we have seen how ML models
  4067. 2:28:52enable AI to create new content. So, now
  4068. 2:28:55let us briefly understand the different
  4069. 2:28:56types of machine learning models.
  4070. 2:28:58So, here, the first type of machine
  4071. 2:29:00learning model is supervised learning.
  4072. 2:29:03Supervised learning trains a model using
  4073. 2:29:05labeled data, where each input has a
  4074. 2:29:07corresponding correct output. And this
  4075. 2:29:09makes it ideal for tasks where
  4076. 2:29:10historical data can be used to predict
  4077. 2:29:12future outcomes.
  4078. 2:29:14For example, let's say spam detection.
  4079. 2:29:17Email services, like Gmail, use a
  4080. 2:29:19supervised learning to classify emails
  4081. 2:29:21as spam or not spam by learning from
  4082. 2:29:24past labeled examples.
  4083. 2:29:26The next example is the price
  4084. 2:29:27predictions.
  4085. 2:29:29So, real estate platforms use regression
  4086. 2:29:31models to predict house prices based on
  4087. 2:29:33the features like location, size, and
  4088. 2:29:36amenities.
  4089. 2:29:37Now, let us see some of the popular
  4090. 2:29:39algorithms.
  4091. 2:29:40So, first let's discuss on decision
  4092. 2:29:42trees. These models break down the data
  4093. 2:29:45into a tree-like structure, where each
  4094. 2:29:47node represent a decision based on a
  4095. 2:29:49feature.
  4096. 2:29:50So, they are easy to interpret and work
  4097. 2:29:52well for both classification. For
  4098. 2:29:54example, deciding if an email is a spam
  4099. 2:29:57or not. And regression example,
  4100. 2:29:59predicting house price.
  4101. 2:30:01However, they can become overly complex.
  4102. 2:30:04Next is the support vector machines.
  4103. 2:30:07So, SVMs are powerful for classification
  4104. 2:30:09task, as they find the optimal boundary,
  4105. 2:30:11also called a hyperplane. And that best
  4106. 2:30:14separates different classes in the data.
  4107. 2:30:17They work well for high-dimensional
  4108. 2:30:19spaces and cases where the distinction
  4109. 2:30:21between categories is clear, such as
  4110. 2:30:23handwriting, facial recognition, or
  4111. 2:30:26medical diagnosis.
  4112. 2:30:27So, now that we have seen how labeled
  4113. 2:30:29data is used. So, now let's explore how
  4114. 2:30:31unsupervised learning finds patterns
  4115. 2:30:33without labels.
  4116. 2:30:35Well, unsupervised learning works with
  4117. 2:30:37unlabeled data, identifying hidden
  4118. 2:30:39patterns and relationships without
  4119. 2:30:41predefined categories.
  4120. 2:30:42So, here we have some of the popular
  4121. 2:30:44algorithms. So, first is the K-means
  4122. 2:30:47clustering. This algorithm partitions
  4123. 2:30:49data into a predefined number of
  4124. 2:30:51clusters.
  4125. 2:30:52By grouping similar data points based on
  4126. 2:30:54their attributes. It works well for
  4127. 2:30:57tasks like customer segmentation. Where
  4128. 2:30:59businesses can group customer based on
  4129. 2:31:02purchasing behavior. However, it assumes
  4130. 2:31:04clusters are spherical and may struggle
  4131. 2:31:07with irregular shaped data. Next we have
  4132. 2:31:10autoencoders.
  4133. 2:31:12So, these are specialized neural
  4134. 2:31:13networks designed to learn efficient
  4135. 2:31:15data representations by encoding and
  4136. 2:31:18reconstructing input data. Let us see
  4137. 2:31:20some of the examples.
  4138. 2:31:22So, first example here we have is
  4139. 2:31:24customer segmentation.
  4140. 2:31:26Where e-commerce platforms group
  4141. 2:31:28customer based on their shopping
  4142. 2:31:29behavior to offer personalized
  4143. 2:31:31recommendations.
  4144. 2:31:32The next example is market analysis.
  4145. 2:31:35Businesses analyze purchasing trends to
  4146. 2:31:37find associations such as which products
  4147. 2:31:40are frequently brought together.
  4148. 2:31:42Now we have covered both labeled and
  4149. 2:31:43unlabeled learning. So let's see how
  4150. 2:31:45semi-supervised learning combines the
  4151. 2:31:47best of both worlds.
  4152. 2:31:49So semi-supervised learning bridges the
  4153. 2:31:52gap between the supervised and
  4154. 2:31:53unsupervised learning by using a small
  4155. 2:31:55amount of data along with large amount
  4156. 2:31:58of unlabeled data.
  4157. 2:31:59So for example, let's say AI assistant
  4158. 2:32:01medical diagnosis.
  4159. 2:32:03Labeled medical images such as x-rays
  4160. 2:32:06with diagnosis are scarce, but large
  4161. 2:32:08amounts of unlabeled images exist.
  4162. 2:32:11Semi-supervised learning help AI learn
  4163. 2:32:14patterns from both labeled and unlabeled
  4164. 2:32:16data improving accuracy in disease
  4165. 2:32:18detection. All right. Now let's explore
  4166. 2:32:21the reinforcement learning where AI
  4167. 2:32:23learns through trial and error.
  4168. 2:32:25Well, reinforcement learning is inspired
  4169. 2:32:28by the concept of learning through trial
  4170. 2:32:30and error. So models interact with an
  4171. 2:32:32environment, receive rewards or
  4172. 2:32:34penalties for actions and refine their
  4173. 2:32:37strength over time. For example, let's
  4174. 2:32:39say gaming.
  4175. 2:32:40Mario AI developed using reinforcement
  4176. 2:32:42learning learns to navigate levels by
  4177. 2:32:45optimizing actions through trial and
  4178. 2:32:47error. The next example is robotics.
  4179. 2:32:50Where robots learn to walk, balance, or
  4180. 2:32:52perform tasks through reinforcement
  4181. 2:32:54learning by maximizing positive
  4182. 2:32:56outcomes.
  4183. 2:32:57Also, reinforcement learning uses
  4184. 2:32:59agents, actions, and rewards to improve
  4185. 2:33:02decision-making.
  4186. 2:33:03Making it ideal for tasks requiring
  4187. 2:33:05continuous learning and adaptation.
  4188. 2:33:08So now that we have covered all the
  4189. 2:33:09types of machine learning models. So
  4190. 2:33:11let's go over some of the key tips to
  4191. 2:33:13help you choose the right one for your
  4192. 2:33:14needs.
  4193. 2:33:16So here are the tips.
  4194. 2:33:17When it comes to supervised learning
  4195. 2:33:19classifying emails as spam or not and
  4196. 2:33:22diagnosing diseases from patient data.
  4197. 2:33:25Next is the unsupervised learning. So,
  4198. 2:33:27unsupervised learning is best when
  4199. 2:33:29you're grouping shoppers by behavior and
  4200. 2:33:31detecting fraud in banking. Next, we
  4201. 2:33:33have semi-supervised learning.
  4202. 2:33:35And this is best when you're improving
  4203. 2:33:37speech recognition with limited label
  4204. 2:33:39data and identifying fake news.
  4205. 2:33:42And finally, the reinforcement learning.
  4206. 2:33:45This will be best when you're training
  4207. 2:33:46self-driving cars to navigate,
  4208. 2:33:48optimizing AI in video games like Mario.
  4209. 2:33:51So, whether it's supervised,
  4210. 2:33:53unsupervised, semi-supervised, or
  4211. 2:33:55reinforcement learning, each model plays
  4212. 2:33:57a crucial role in shaping AI's future.
  4213. 2:34:00So, as generative AI continues to
  4214. 2:34:02evolve, these models are driving
  4215. 2:34:03innovation in text, images, and video
  4216. 2:34:06generation.
  4217. 2:34:07So, which machine learning model do you
  4218. 2:34:09find the most fascinating? Let me know
  4219. 2:34:11in the comments below.
  4220. 2:34:15>> [music]
  4221. 2:34:18>> Let me connect you to the real life and
  4222. 2:34:20tell you what all are the things which
  4223. 2:34:22you can easily do using the concepts of
  4224. 2:34:23machine learning.
  4225. 2:34:25So, you can easily get answer to the
  4226. 2:34:26questions like which types of house lies
  4227. 2:34:28in this segment or what is the market
  4228. 2:34:30value of this house? Or is this a mail a
  4229. 2:34:33spam or not a spam? Is there any fraud?
  4230. 2:34:36Well, these are some of the question you
  4231. 2:34:37could ask to the machine. But for
  4232. 2:34:38getting an answer to these, you need
  4233. 2:34:40some algorithm. The machine need to
  4234. 2:34:42train on the basis of some algorithm.
  4235. 2:34:44Okay, but how will you decide which
  4236. 2:34:46algorithm to choose and when?
  4237. 2:34:48Okay, so the best option for us is to
  4238. 2:34:50explore them one by one.
  4239. 2:34:53So, the first is classification
  4240. 2:34:54algorithm where the category is
  4241. 2:34:56predicted using the data. If you have
  4242. 2:34:58some question like is this person a male
  4243. 2:35:01or a female? Or is this a mail a spam or
  4244. 2:35:04not a spam? Then these category of
  4245. 2:35:06question would fall under the
  4246. 2:35:07classification algorithm.
  4247. 2:35:09Classification is a supervised learning
  4248. 2:35:10approach in which the computer program
  4249. 2:35:13learns from the input given to it and
  4250. 2:35:14then uses this learning to classify new
  4251. 2:35:17observation. Some examples of
  4252. 2:35:19classification problems are speech
  4253. 2:35:21organization, handwriting recognition,
  4254. 2:35:23biometric identification, document
  4255. 2:35:25classification, etc.
  4256. 2:35:28Shall we move ahead?
  4257. 2:35:30Okay.
  4258. 2:35:32So, next is the anomaly detection
  4259. 2:35:34algorithm where you identify the unusual
  4260. 2:35:37data point. So, what is anomaly
  4261. 2:35:38detection? Well, it's a technique that
  4262. 2:35:40is used to identify unusual pattern that
  4263. 2:35:43do not conform to expected behavior. Or
  4264. 2:35:45you can say the outliers.
  4265. 2:35:47It has many application in business like
  4266. 2:35:49intrusion detection, like identifying
  4267. 2:35:51strange patterns in the network traffic
  4268. 2:35:53that could signal a hack, or system
  4269. 2:35:55health monitoring, that is spotting a
  4270. 2:35:56deadly tumor in the MRI scan.
  4271. 2:35:59Or you can even use it for fraud
  4272. 2:36:01detection in credit card transaction, or
  4273. 2:36:03to deal with fault detection in
  4274. 2:36:04operating environment.
  4275. 2:36:06So, next comes the clustering algorithm.
  4276. 2:36:08You can use this clustering algorithm to
  4277. 2:36:10group the data based on some similar
  4278. 2:36:12condition. Now, you can get answer to
  4279. 2:36:14which type of houses lies in this
  4280. 2:36:16segment, or what type of customer buys
  4281. 2:36:18this product. The clustering is a task
  4282. 2:36:20of dividing the population or data
  4283. 2:36:22points into a number of groups such that
  4284. 2:36:24the data point in the same groups are
  4285. 2:36:26more similar to other data points in the
  4286. 2:36:28same group than those in the other
  4287. 2:36:30groups. In simple words, the aim is to
  4288. 2:36:33segregate groups with similar trait and
  4289. 2:36:35assign them into cluster.
  4290. 2:36:37Now, this clustering is a task of
  4291. 2:36:38dividing the population or data points
  4292. 2:36:40into a number of groups such that the
  4293. 2:36:42data points in the X group is more
  4294. 2:36:44similar to the other data points in the
  4295. 2:36:46same group rather than those in the
  4296. 2:36:48other group. In other words, the aim is
  4297. 2:36:50to segregate the groups with similar
  4298. 2:36:52traits and assign them into different
  4299. 2:36:54clusters. Let's understand this with an
  4300. 2:36:56example. Suppose you're the head of a
  4301. 2:36:58rental store and you wish to understand
  4302. 2:37:00the preference of your customer to scale
  4303. 2:37:02up your business. So, is it possible for
  4304. 2:37:04you to look at the detail of each
  4305. 2:37:05customer and design a unique business
  4306. 2:37:08strategy for each of them?
  4307. 2:37:09Definitely not. Right?
  4308. 2:37:12But what you can do is to cluster all
  4309. 2:37:14your customer saying to 10 different
  4310. 2:37:16groups based on their purchasing habit
  4311. 2:37:18and you can use a separate strategy for
  4312. 2:37:20customers in each of these 10 different
  4313. 2:37:22groups. And this is what we call
  4314. 2:37:24clustering.
  4315. 2:37:26Next we have regression algorithm where
  4316. 2:37:28the data itself is predicted. Question
  4317. 2:37:30you may ask to this type of model is
  4318. 2:37:32like what is the market value of this
  4319. 2:37:34house or is it going to rain tomorrow or
  4320. 2:37:36not?
  4321. 2:37:37So regression is one of the most
  4322. 2:37:39important and broadly used machine
  4323. 2:37:40learning and statistics tool.
  4324. 2:37:43It allows you to make prediction from
  4325. 2:37:44data by learning the relationship
  4326. 2:37:46between the features of your data and
  4327. 2:37:48some observed continuous valued
  4328. 2:37:49response. Regression is used in a
  4329. 2:37:52massive number of application.
  4330. 2:37:54You know what? Stock prices prediction
  4331. 2:37:55can be done using regression.
  4332. 2:37:57Now you know about different machine
  4333. 2:37:59learning algorithm. How will you decide
  4334. 2:38:01which algorithm to choose and when?
  4335. 2:38:03So let's cover this part using a demo.
  4336. 2:38:05So in this demo part, what we'll do,
  4337. 2:38:07we'll create six different machine
  4338. 2:38:08learning model and pick the best model
  4339. 2:38:11and build the confidence such that it
  4340. 2:38:12has the most reliable accuracy.
  4341. 2:38:16So for our demo part, we'll be using the
  4342. 2:38:17Iris data set. This data set is quite
  4343. 2:38:20very famous and is considered one of the
  4344. 2:38:22best small project to start with.
  4345. 2:38:24You can consider this as a hello world
  4346. 2:38:25data set for machine learning. So this
  4347. 2:38:27data set consists of 150 observation of
  4348. 2:38:30Iris flower.
  4349. 2:38:31There are four columns of measurement of
  4350. 2:38:33flowers in centimeters. The fifth column
  4351. 2:38:35being the species of the flower
  4352. 2:38:36observed. All the observed flowers
  4353. 2:38:38belong to one of the three species of
  4354. 2:38:40Iris setosa, Iris virginica and Iris
  4355. 2:38:43versicolor.
  4356. 2:38:44Well, this is a good project because it
  4357. 2:38:46is so well to understand. The attributes
  4358. 2:38:48are numeric so you have to figure out
  4359. 2:38:49how to load and handle the data. It is a
  4360. 2:38:51classification problem thereby allowing
  4361. 2:38:53you to practice with perhaps an easier
  4362. 2:38:55type of supervised learning algorithm.
  4363. 2:38:57It has only four attributes and 150 rows
  4364. 2:38:59meaning it is very small and can easily
  4365. 2:39:01fit into the memory.
  4366. 2:39:02And even all of the numeric attributes
  4367. 2:39:04are in same unit and the same scale. It
  4368. 2:39:07means you do not require any special
  4369. 2:39:08scaling or transformation to get
  4370. 2:39:10started.
  4371. 2:39:12So, let's start coding and as I told
  4372. 2:39:14earlier for the demo part, I'll be using
  4373. 2:39:16Anaconda with Python 3.0 installed on
  4374. 2:39:18it. So, when you install Anaconda, how
  4375. 2:39:20your navigator would look like. So,
  4376. 2:39:22there's my home page of my Anaconda
  4377. 2:39:23navigator. On this I'll be using the
  4378. 2:39:26Jupiter notebook, which is a web-based
  4379. 2:39:27interactive computing notebook
  4380. 2:39:29environment, which will help me to write
  4381. 2:39:30and execute my Python codes on it. So,
  4382. 2:39:32let's hit the launch button and execute
  4383. 2:39:34our Jupiter notebook.
  4384. 2:39:36So, as you can see that my Jupiter
  4385. 2:39:37notebook is starting on localhost 8890.
  4386. 2:39:41Okay? So, this is my Jupiter notebook.
  4387. 2:39:42What I'll do here, I'll select new
  4388. 2:39:44notebook Python 3.
  4389. 2:39:48There's my environment where I can write
  4390. 2:39:50and execute all my Python codes on it.
  4391. 2:39:52So, let's start by checking the version
  4392. 2:39:54of the libraries. In order to make this
  4393. 2:39:56video short and more interactive and
  4394. 2:39:57more informative, I've already done the
  4395. 2:39:59set of code. So, let me just copy and
  4396. 2:40:01paste it down. I'll explain you then one
  4397. 2:40:03by one.
  4398. 2:40:04So, let's start by checking the version
  4399. 2:40:05of the Python libraries.
  4400. 2:40:07Okay? So, there's the code. Let's just
  4401. 2:40:10copy it.
  4402. 2:40:11Copied and let's paste it. Okay. First,
  4403. 2:40:14let me summarize things for you. What we
  4404. 2:40:16are doing here, we are just checking the
  4405. 2:40:17version of the different libraries.
  4406. 2:40:19Starting with Python, we'll first check
  4407. 2:40:20what version of Python we are working
  4408. 2:40:22on, then we'll check what are the
  4409. 2:40:23version of SciPy we are using, then
  4410. 2:40:25NumPy, Matplotlib, then Pandas, then
  4411. 2:40:27scikit-learn. Okay? So, let's execute
  4412. 2:40:29the run button and see what are the
  4413. 2:40:30various version of libraries which we
  4414. 2:40:32are using. Hit the run. So, we are
  4415. 2:40:33working on Python 3.6.4, SciPy 1.0,
  4416. 2:40:37NumPy 1.14, Matplotlib 2.12, Pandas
  4417. 2:40:400.22, and scikit-learn of version 0.19.
  4418. 2:40:44Okay?
  4419. 2:40:45So, these are the version which I'm
  4420. 2:40:46using. Ideally, your version should be
  4421. 2:40:48more recent or it should match. But,
  4422. 2:40:50don't worry if you lag few versions
  4423. 2:40:52behind as the APIs do not change so
  4424. 2:40:54quickly. Everything in this tutorial
  4425. 2:40:56will very likely still work for you.
  4426. 2:40:58Okay? But, in case you're getting an
  4427. 2:41:00error, stop and try to fix that error.
  4428. 2:41:03In case you're unable to find the
  4429. 2:41:04solution for the error, feel free to
  4430. 2:41:06reach out Edureka even after this class.
  4431. 2:41:08Let me tell you this, if you're not able
  4432. 2:41:09to run the script properly, you will not
  4433. 2:41:11be able to complete this tutorial, okay?
  4434. 2:41:13So, whenever you get a doubt, reach out
  4435. 2:41:15to Edureka and just resolve it.
  4436. 2:41:17Now, if everything is working smoothly,
  4437. 2:41:19then now it's the time to load the data
  4438. 2:41:21set. So, as I said, I'll be using the
  4439. 2:41:23Iris flower data set for this tutorial.
  4440. 2:41:25But, before loading the data set, let's
  4441. 2:41:27import all the modules, function, and
  4442. 2:41:29the object which we are going to use in
  4443. 2:41:31this tutorial. Same, I've already
  4444. 2:41:32written the set of code, so let's just
  4445. 2:41:34copy and paste them. Let's load all the
  4446. 2:41:36libraries.
  4447. 2:41:38So, these are the various libraries
  4448. 2:41:39which we'll be using in our tutorial.
  4449. 2:41:42So, everything should work fine without
  4450. 2:41:43an error. If you get an error, just
  4451. 2:41:45stop. You need to work on your SciPy
  4452. 2:41:46environment before you continue any
  4453. 2:41:48further. So, I guess everything should
  4454. 2:41:50work fine. Let's hit the run button and
  4455. 2:41:51see.
  4456. 2:41:53Okay, it worked. So, let's now move
  4457. 2:41:56ahead and load the data. We can load the
  4458. 2:41:58data direct from the UCI machine
  4459. 2:41:59learning repository. First of all, let
  4460. 2:42:01me tell you, we are using Panda to load
  4461. 2:42:03the data.
  4462. 2:42:04Okay?
  4463. 2:42:05So, let's say my URL is this. So, this
  4464. 2:42:08is my URL for the UCI machine learning
  4465. 2:42:09repository from where I'll be
  4466. 2:42:10downloading the data set, okay?
  4467. 2:42:13Now, what I'll do, I'll specify the name
  4468. 2:42:14of each column when loading the data.
  4469. 2:42:16This will help me later to explore the
  4470. 2:42:18data, okay?
  4471. 2:42:19So, I'll just copy and paste it down.
  4472. 2:42:22Okay?
  4473. 2:42:23So, I'm defining a variable names which
  4474. 2:42:25consists of various parameters including
  4475. 2:42:27sepal length, sepal width, petal length,
  4476. 2:42:29petal width, and class. So, these are
  4477. 2:42:31just the name of column from the data
  4478. 2:42:32set, okay? Now, let's define the data
  4479. 2:42:35set. So, data set equals panda.read_csv.
  4480. 2:42:39Inside that, we are defining URL and the
  4481. 2:42:41names, that is equal to name.
  4482. 2:42:44As I already said, we'll be using Panda
  4483. 2:42:46to load the data, all right?
  4484. 2:42:49So, we are using panda.read_csv, so we
  4485. 2:42:51are reading the CSV file, and inside
  4486. 2:42:53that, from where that CSV is coming?
  4487. 2:42:54From the the Which URL? So, this is my
  4488. 2:42:56URL. Okay?
  4489. 2:42:58And names equal names. It's just
  4490. 2:43:00specifying the names of the various
  4491. 2:43:01columns in that particular CSV file.
  4492. 2:43:03Okay?
  4493. 2:43:04So, let's move forward and execute it.
  4494. 2:43:06So, even our data set is loaded.
  4495. 2:43:09In case you have some network issues,
  4496. 2:43:11just go ahead and download the Iris data
  4497. 2:43:13file into your working directory and
  4498. 2:43:14load it using the same method. But yeah,
  4499. 2:43:16make sure that you change the URL to the
  4500. 2:43:18local name, or else you might get an
  4501. 2:43:19error. Okay. Yeah, our data set is
  4502. 2:43:22loaded. So, let's move ahead and check
  4503. 2:43:23our data set. Let's see how many columns
  4504. 2:43:25or rows we have in our data set. Okay.
  4505. 2:43:28So, let's print the number of rows and
  4506. 2:43:30columns in our data set. So, our data
  4507. 2:43:32set is data set.shape.
  4508. 2:43:35What this will do, it will just give you
  4509. 2:43:37the numbers of total number of rows and
  4510. 2:43:39total number of column, or you can say
  4511. 2:43:40the total number of instances or
  4512. 2:43:42attributes in your data set. Fine?
  4513. 2:43:44So, print data set.shape. What are you
  4514. 2:43:46getting? 150 and 5. So, 150 is the total
  4515. 2:43:49number of rows in your data set, and 5
  4516. 2:43:50is the total number of columns. Fine?
  4517. 2:43:53So, moving on ahead, what if I want to
  4518. 2:43:55see the sample data set? Okay. So, let
  4519. 2:43:58me just print the first 30 instances of
  4520. 2:43:59the data set. Okay? So, print
  4521. 2:44:03data set.head.
  4522. 2:44:07What I want is the first 30 instances.
  4523. 2:44:09Fine? This will give me the first 30
  4524. 2:44:11result of my data set. Okay? So, when I
  4525. 2:44:13hit the run button, what I'm getting is
  4526. 2:44:15the first 30 result. Okay?
  4527. 2:44:180
  4528. 2:44:20to 29. So, this is how my sample data
  4529. 2:44:22set looks like.
  4530. 2:44:23Sepal length, sepal width, petal length,
  4531. 2:44:25petal width, and the class. Okay?
  4532. 2:44:28So, this is how our data set looks like.
  4533. 2:44:31Now, let's move on and look at the
  4534. 2:44:32summary of each attribute. What if I
  4535. 2:44:34want to find out the count, mean, the
  4536. 2:44:37minimum and the maximum values, and some
  4537. 2:44:39other percentiles as well. So, what
  4538. 2:44:40should I do then?
  4539. 2:44:41For that, print
  4540. 2:44:43data set.describe.
  4541. 2:44:47What it will give,
  4542. 2:44:48let's see.
  4543. 2:44:50So, you can see that all the numbers are
  4544. 2:44:52the same scales of similar range between
  4545. 2:44:540 to 8 cm, right? The mean value, the
  4546. 2:44:57standard deviation, the minimum value,
  4547. 2:44:59the 25th percentile, 50th percentile,
  4548. 2:45:0175th percentile, the maximum value, all
  4549. 2:45:03these values lies in the range between 0
  4550. 2:45:05to 8 cm.
  4551. 2:45:07Okay.
  4552. 2:45:08So what we just did is we just took a
  4553. 2:45:10summary of each attribute. Now let's
  4554. 2:45:13look at the number of instances that
  4555. 2:45:14belong to each class. So for that, what
  4556. 2:45:17we'll do print data set first of all.
  4557. 2:45:21So let's print data set and I want to
  4558. 2:45:24group it
  4559. 2:45:25group by using
  4560. 2:45:28class
  4561. 2:45:30and I want the size of it, size of each
  4562. 2:45:32class. Fine?
  4563. 2:45:34And let's hit the run.
  4564. 2:45:47Okay. So what I want to do, I want to
  4565. 2:45:49print print what? Data set. How I want
  4566. 2:45:52to get it? I want it by class. So group
  4567. 2:45:54by class.
  4568. 2:45:56Okay. Now I want the size of each class.
  4569. 2:45:59Find the size of each class. So group by
  4570. 2:46:01class.size. Execute the run.
  4571. 2:46:04So you can see that I have 15 instances
  4572. 2:46:06of Iris setosa, 15 instances of Iris
  4573. 2:46:08versicolor, and 15 instances of Iris
  4574. 2:46:10virginica. Okay? All are of data type
  4575. 2:46:13integer of base 64. Fine? So now we have
  4576. 2:46:16a basic idea of our data. Now let's move
  4577. 2:46:18ahead and create some visualization for
  4578. 2:46:20it. So for this we are going to create
  4579. 2:46:22two different types of plot. First would
  4580. 2:46:23be the univariate plot and the next
  4581. 2:46:25would be the multivariate plot. So we'll
  4582. 2:46:26be creating univariate plots to better
  4583. 2:46:28understand about each attribute. And the
  4584. 2:46:30next we'll be creating the multivariate
  4585. 2:46:32plot to better understand the
  4586. 2:46:33relationship between different
  4587. 2:46:34attributes. Okay? So we start with some
  4588. 2:46:36univariate plot. That is plot of each
  4589. 2:46:38individual variable. So given that the
  4590. 2:46:40input variables are numeric, we can
  4591. 2:46:41create box and whiskers plot for it.
  4592. 2:46:43Okay? So let's move ahead and create a
  4593. 2:46:44box and whiskers plot. So data set.plot.
  4594. 2:46:47What kind I want? It's a box.
  4595. 2:46:50Okay. And do I need a subplot? Yeah, I
  4596. 2:46:53need subplots for that. So, subplots
  4597. 2:46:55equal true. What type of layout do I
  4598. 2:46:57want? So, my layout structure is 2 cross
  4599. 2:47:012.
  4600. 2:47:02Next, do I want to share my coordinates,
  4601. 2:47:04X and Y coordinates? No, I don't want to
  4602. 2:47:06share it. So, share X equal false.
  4603. 2:47:09And even share Y, that too equals false.
  4604. 2:47:13Okay. So, we have our dataset.plot kind
  4605. 2:47:16equal box. My subplots is true, layout 2
  4606. 2:47:19cross 2. And then what I want to do it,
  4607. 2:47:21I want to see it. So, plot.show.
  4608. 2:47:23Whatever I created, show it. Okay.
  4609. 2:47:26Execute it.
  4610. 2:47:29Now, this gives us a much clearer idea
  4611. 2:47:30about the distribution of the input
  4612. 2:47:32attribute. Now, what if I had given the
  4613. 2:47:33layout to 2 cross 2 instead of that, I'd
  4614. 2:47:36have given it 4 cross 4. So, what it
  4615. 2:47:39will result? Just see. Fine. Everything
  4616. 2:47:41would be printed in just one single row.
  4617. 2:47:43Hold on, guys. Arya has a doubt. He's
  4618. 2:47:44asking that why we are using the share X
  4619. 2:47:46and share Y values. What are these? Why
  4620. 2:47:48we have assigned false values to it?
  4621. 2:47:50Okay, Arya. So, in order to resolve this
  4622. 2:47:52query, I need to show you what will
  4623. 2:47:54happen if I give true values to them.
  4624. 2:47:55Okay. So, be with me. So, share X equal
  4625. 2:47:58true and share Y, that equals true. So,
  4626. 2:48:01let's see what result we'll get.
  4627. 2:48:04You're getting it. The X and Y
  4628. 2:48:05coordinates are just shared among all
  4629. 2:48:07the four visualization, right? So, Arya,
  4630. 2:48:09you can see that the sepal length and
  4631. 2:48:10sepal width has Y values ranging from
  4632. 2:48:130.0 to 7.5 which are being shared among
  4633. 2:48:15both the visualization. So, is with the
  4634. 2:48:17petal length, it has shared value
  4635. 2:48:19between 0.0 to 7.5. Okay. So, that is
  4636. 2:48:22why I don't want to share the value of X
  4637. 2:48:24and Y. It's just giving us a cluttered
  4638. 2:48:26visualization. So, Arya, why I'm doing
  4639. 2:48:28this? I'm just doing it cuz I don't want
  4640. 2:48:31my X and Y coordinates to be shared
  4641. 2:48:33among any visualization. Okay. That is
  4642. 2:48:35why my share X and share Y value are
  4643. 2:48:37false. Okay. Let's execute it.
  4644. 2:48:40So, this is a
  4645. 2:48:41pretty much clear visualization which
  4646. 2:48:43gives a clear idea about the
  4647. 2:48:44distribution of the input attributes.
  4648. 2:48:46Now, if you want, you can also create a
  4649. 2:48:48histogram of each input variable to get
  4650. 2:48:50a clear idea of the distribution. So,
  4651. 2:48:52let's create a histogram for it. So,
  4652. 2:48:53dataset.hist, okay? I would need to
  4653. 2:48:56proceed. So, plot.show. Let's see. So,
  4654. 2:48:58this is my histogram and it seems that
  4655. 2:49:00we have two input variables that have a
  4656. 2:49:02Gaussian distribution. So, this is
  4657. 2:49:04useful to note as we can use the
  4658. 2:49:06algorithms that can exploit this
  4659. 2:49:07assumption, okay? So, next comes the
  4660. 2:49:09multivariate plot. Now that we have
  4661. 2:49:10created the univariate plot to
  4662. 2:49:12understand about each attribute, let's
  4663. 2:49:14move on and look at the multivariate
  4664. 2:49:16plot and see the interaction between the
  4665. 2:49:18different variables. So, first let's
  4666. 2:49:20look at the scatter plot of all the
  4667. 2:49:21attribute. This can be helpful to spot
  4668. 2:49:23structured relationship between input
  4669. 2:49:24variables, okay? So, let's create a
  4670. 2:49:26scatter matrix. So, for creating a
  4671. 2:49:27scatter plot, we need scatter matrix and
  4672. 2:49:31we need to pass our dataset into it,
  4673. 2:49:33okay? And then, what I want, I want to
  4674. 2:49:35see it. So, plot.show. So, this is how
  4675. 2:49:37my scatter matrix looks like. It's like
  4676. 2:49:39that the diagonal grouping of some pair,
  4677. 2:49:41right? So, this suggests a high
  4678. 2:49:42correlation and a predictable
  4679. 2:49:44relationship, all right? This was our
  4680. 2:49:45multivariate plot. Now, let's move on
  4681. 2:49:47and evaluate some algorithm. Now, it's
  4682. 2:49:49time to create some model of the data
  4683. 2:49:51and estimate the accuracy on the base of
  4684. 2:49:53unseen data, okay? So, now we know all
  4685. 2:49:56about our dataset, right? We know how
  4686. 2:49:58many instances and attributes are there
  4687. 2:49:59in our dataset. We know the summary of
  4688. 2:50:01each attribute. Now, I guess we have
  4689. 2:50:03seen much about our dataset. Now, let's
  4690. 2:50:05move on and create some algorithm and
  4691. 2:50:07estimate their accuracy based on the
  4692. 2:50:09unseen data. Okay. Now, what we'll do,
  4693. 2:50:11we'll create some model of the data and
  4694. 2:50:13estimate the accuracy based on the some
  4695. 2:50:15unseen data, okay? So, for that, first
  4696. 2:50:17of all, let's create a validation
  4697. 2:50:18dataset. What is a validation dataset?
  4698. 2:50:20Validation dataset is your training
  4699. 2:50:22dataset that will be using it to train
  4700. 2:50:24our model, fine? All right. So, how
  4701. 2:50:26we'll create a validation dataset? For
  4702. 2:50:28creating a validation dataset, what we
  4703. 2:50:30are going to do is we are going to split
  4704. 2:50:31our dataset into two part, okay? So, the
  4705. 2:50:33very first thing we'll do is to create a
  4706. 2:50:35validation dataset. So, why do we even
  4707. 2:50:37need a validation data set? So, we need
  4708. 2:50:39a validation data set to know that the
  4709. 2:50:41model we created is any good. Later,
  4710. 2:50:43what we'll do, we'll use the statistical
  4711. 2:50:45method to estimate the accuracy of the
  4712. 2:50:47model that we create on the unseen data.
  4713. 2:50:49We also want a more concrete estimate of
  4714. 2:50:51the accuracy of the best model on unseen
  4715. 2:50:53data by evaluating it on the actual
  4716. 2:50:55unseen data. Okay? Confused? Let me
  4717. 2:50:57simplify this for you. What we'll do,
  4718. 2:50:59we'll split the loaded data into two
  4719. 2:51:00parts. The first 80% of the data we'll
  4720. 2:51:03use it to train our model. And the rest
  4721. 2:51:0520% we'll hold back as the validation
  4722. 2:51:07data set that we'll use it to verify our
  4723. 2:51:09trained model. Okay? Fine. So, let's
  4724. 2:51:11define an array. This is my array. What
  4725. 2:51:14it will consist of? It will consist of
  4726. 2:51:16all the values from the data set. So,
  4727. 2:51:17data set.values.
  4728. 2:51:19Okay? Next, I'll define a variable X
  4729. 2:51:22which will consist of all the column
  4730. 2:51:25from the array from zero to four.
  4731. 2:51:28Starting from zero to four. And the next
  4732. 2:51:30variable Y which would consist of the
  4733. 2:51:33array starting from this. So, first of
  4734. 2:51:37all, we'll define a variable X that will
  4735. 2:51:39consist of the values in the array
  4736. 2:51:41starting from the beginning zero till
  4737. 2:51:43four. Okay? So, these are the column
  4738. 2:51:45which we'll include in the X variable.
  4739. 2:51:47And for a Y variable, I'll define it as
  4740. 2:51:49a class or the output. So, what I need,
  4741. 2:51:51I just need the fourth column that is my
  4742. 2:51:53class column. So, I'll start it from the
  4743. 2:51:55beginning and I just want the fourth
  4744. 2:51:57column. Okay? Now, I'll define the my
  4745. 2:51:59validation size.
  4746. 2:52:01validation_size.
  4747. 2:52:04I'll define it as 0.20 and I'll use a
  4748. 2:52:07seed.
  4749. 2:52:08I'll define seed equals six.
  4750. 2:52:11So, this method seed sets the integer
  4751. 2:52:13starting value used in generating random
  4752. 2:52:15number. Okay? I'll define the value of
  4753. 2:52:17seed equals six. I'll tell you what is
  4754. 2:52:19the importance of it later on. Okay? So,
  4755. 2:52:21let me define first few variables such
  4756. 2:52:23as X_train, test, Y_train,
  4757. 2:52:27and Y_test.
  4758. 2:52:30Okay? So, what we want to do is select
  4759. 2:52:32some model. Okay. So, model underscore
  4760. 2:52:34selection. But, before doing that, what
  4761. 2:52:36we have to do is split our training data
  4762. 2:52:37set into two halves. Okay. So, dot train
  4763. 2:52:39underscore test underscore split. What
  4764. 2:52:42we want to split is the value of X and
  4765. 2:52:45Y. Okay. And my test size is
  4766. 2:52:49equals to validation size.
  4767. 2:52:52Which is a 0.20. Correct? And my random
  4768. 2:52:55state
  4769. 2:52:57is equal to seed. So, what the seed is
  4770. 2:52:59doing here, it's helping me to keep the
  4771. 2:53:01same randomness in the training and
  4772. 2:53:03testing data set. Fine. So, let's
  4773. 2:53:05execute it and see what is our result.
  4774. 2:53:08Let's execute it. Next, we'll create a
  4775. 2:53:10test harness. For this, we'll use
  4776. 2:53:1210-fold cross-validation to estimate the
  4777. 2:53:14accuracy.
  4778. 2:53:16So, what it will do, it will split our
  4779. 2:53:17data set into 10 parts. Train on the
  4780. 2:53:20nine part and test on the one part. And
  4781. 2:53:22this will repeat for all combination of
  4782. 2:53:24train and test splits. Okay. So, for
  4783. 2:53:26that, let's define again
  4784. 2:53:29my seed that was six, already defined,
  4785. 2:53:32and scoring
  4786. 2:53:34equals accuracy.
  4787. 2:53:37Fine.
  4788. 2:53:37So, we are using the metric of accuracy
  4789. 2:53:39to evaluate the model. So, what is this?
  4790. 2:53:42This is a ratio of number of correctly
  4791. 2:53:44predicted instances divided by the total
  4792. 2:53:46number of instances in the data set
  4793. 2:53:48multiplied by 100, giving a percentage.
  4794. 2:53:50Example, it's 98% accurate or 99%
  4795. 2:53:54accurate, things like that. Okay. So,
  4796. 2:53:55we'll be using the scoring variable when
  4797. 2:53:57we run the build and evaluate each model
  4798. 2:54:00in the next step. So, next part is
  4799. 2:54:02building model.
  4800. 2:54:04Till now, we don't know which algorithm
  4801. 2:54:05would be good for this problem or what
  4802. 2:54:07configuration to use. So, let's begin
  4803. 2:54:09with six different algorithm. I'll be
  4804. 2:54:11using logistic regression, linear
  4805. 2:54:13discriminant analysis, K-nearest
  4806. 2:54:15neighbor, classification and regression
  4807. 2:54:17trees, Naive Bayes, and support vector
  4808. 2:54:19machine. Well, these algorithms which
  4809. 2:54:21I'm using is a good mixture of simple
  4810. 2:54:23linear or non-linear algorithms. In
  4811. 2:54:25simple linear which included the
  4812. 2:54:26logistic regression and the linear
  4813. 2:54:28discriminant analysis or the non-linear
  4814. 2:54:30part which included the KNN algorithm,
  4815. 2:54:32the CART algorithm, the Naive Bayes, and
  4816. 2:54:34the support vector machines. Okay. So,
  4817. 2:54:36we reset the random number seed before
  4818. 2:54:38each run to ensure that evaluation of
  4819. 2:54:40each algorithm is performed using
  4820. 2:54:42exactly the same data splits. It ensures
  4821. 2:54:44the result are directly comparable.
  4822. 2:54:46Okay. So, let me just copy and paste it.
  4823. 2:54:49Okay.
  4824. 2:54:53So, what we're doing here, we're
  4825. 2:54:54building five different types of model.
  4826. 2:54:56We're building a logistic regression,
  4827. 2:54:58linear discriminant analysis, K-nearest
  4828. 2:55:00neighbor, decision tree, Gaussian Naive
  4829. 2:55:02Bayes, and the support vector machine.
  4830. 2:55:04Okay. Next, what we'll do, we'll
  4831. 2:55:05evaluate model in each turn. Okay.
  4832. 2:55:08So, what is this? So, we have six
  4833. 2:55:10different model and accuracy estimation
  4834. 2:55:12for each one of them. Now, we need to
  4835. 2:55:14compare the model to each other and
  4836. 2:55:15select the most accurate of them all.
  4837. 2:55:17So, running this script, we saw the
  4838. 2:55:19following result. So, we can see some of
  4839. 2:55:21the result on the screen. What is this?
  4840. 2:55:23It is just the accuracy score using
  4841. 2:55:24different set of algorithms. Okay. When
  4842. 2:55:27we are using logistic regression, what
  4843. 2:55:28is the accuracy rate? When we are using
  4844. 2:55:30linear discriminant algorithm, what is
  4845. 2:55:32the accuracy? And so on and so. Okay.
  4846. 2:55:34So, from the output, it seems that LD
  4847. 2:55:36algorithm was the most accurate model
  4848. 2:55:38that we tested. Now, we want to get an
  4849. 2:55:40idea of the accuracy of the model on our
  4850. 2:55:42validation set or the testing data set.
  4851. 2:55:44So, this will give us a independent
  4852. 2:55:46final check on the accuracy of the best
  4853. 2:55:47model. It is always valuable to keep a
  4854. 2:55:50testing data set for just in case you
  4855. 2:55:52made a over-fitting to the testing data
  4856. 2:55:54set or you made a data leak. Both will
  4857. 2:55:56result in a overly optimistic result.
  4858. 2:55:58Okay.
  4859. 2:55:59You can run the LD model directly on the
  4860. 2:56:01validation set and summarize the result
  4861. 2:56:03as a final score, a confusion matrix,
  4862. 2:56:06and a classification report.
  4863. 2:56:09>> [music]
  4864. 2:56:13>> Let us understand what regression in
  4865. 2:56:15machine learning is.
  4866. 2:56:16So, what exactly is regression?
  4867. 2:56:18The main goal of regression is the
  4868. 2:56:20construction of an efficient model to
  4869. 2:56:22predict the dependent attributes from a
  4870. 2:56:24bunch of attribute variables.
  4871. 2:56:26A regression problem is where the output
  4872. 2:56:28variable is either real or a continuous
  4873. 2:56:30value like salary, weight, area, etc.
  4874. 2:56:33We can also define regression as a
  4875. 2:56:35statistical means that is used in
  4876. 2:56:36applications like housing, investing,
  4877. 2:56:38etc. to predict the relationship between
  4878. 2:56:40a dependent variable and a bunch of
  4879. 2:56:42independent variables.
  4880. 2:56:44For example, let's say in the finance
  4881. 2:56:46application or investing, we can
  4882. 2:56:48actually predict the values of certain
  4883. 2:56:50stock prices or you know those values
  4884. 2:56:52depending on the independent variables
  4885. 2:56:55like how many years it takes for a stock
  4886. 2:56:57to you know actually mature or how many
  4887. 2:56:59days will it take to grow or those
  4888. 2:57:01variables that you have in investing and
  4889. 2:57:04depending upon that we can make a
  4890. 2:57:05possible outcome or a possible
  4891. 2:57:06prediction of how a stock is going to be
  4892. 2:57:09invested in a profit state or a loss
  4893. 2:57:11state or all those things or we can take
  4894. 2:57:13another example like housing. We can
  4895. 2:57:15take different parameters like number of
  4896. 2:57:17years it's been there, how many people
  4897. 2:57:19have used it or what is the area of the
  4898. 2:57:22house depending on all these factors or
  4899. 2:57:24how many rooms does the house have, we
  4900. 2:57:26can predict the price of a house.
  4901. 2:57:28So this is basically what regression
  4902. 2:57:30really is.
  4903. 2:57:31So let us take a look at the various
  4904. 2:57:32types of regression techniques that we
  4905. 2:57:34have.
  4906. 2:57:35We have simple linear regression, then
  4907. 2:57:36we have polynomial regression, support
  4908. 2:57:38vector regression, decision tree
  4909. 2:57:40regression, we have random forest
  4910. 2:57:42regression and we have logistic
  4911. 2:57:43regression as well. That is also a type
  4912. 2:57:45of regression that we have. But for now
  4913. 2:57:47we'll be focusing on simple linear
  4914. 2:57:49regression.
  4915. 2:57:50So let's talk about how or what exactly
  4916. 2:57:52is simple linear regression first. So
  4917. 2:57:54one of the most interesting and common
  4918. 2:57:56regression technique is a simple linear
  4919. 2:57:57regression. In this we predict the
  4920. 2:57:59outcome of a dependent variable Y based
  4921. 2:58:02on the independent variables X. So the
  4922. 2:58:04relationship between the variables is
  4923. 2:58:06linear, hence the word linear
  4924. 2:58:08regression.
  4925. 2:58:09Then comes the polynomial regression.
  4926. 2:58:11So in this regression technique, we
  4927. 2:58:13transform the original features into a
  4928. 2:58:15polynomial feature of a given degree and
  4929. 2:58:17then perform regression on it. So, this
  4930. 2:58:19is basically polynomial regression.
  4931. 2:58:22After this, we have support vector
  4932. 2:58:23machine regression or we can also call
  4933. 2:58:25it SVR. We identify a hyperplane with
  4934. 2:58:28maximum margin such that the maximum
  4935. 2:58:31number of data points are within those
  4936. 2:58:33margins.
  4937. 2:58:34It is also quite similar to the support
  4938. 2:58:35vector machine classification algorithm.
  4939. 2:58:38Then we have decision tree regression.
  4940. 2:58:41A decision tree can be used for both
  4941. 2:58:42regression and classification. But, in
  4942. 2:58:44this case of regression, we use the ID3
  4943. 2:58:47algorithm, which is iterative
  4944. 2:58:48dichotomizer 3, to identify the
  4945. 2:58:51splitting node by reducing the standard
  4946. 2:58:53deviation.
  4947. 2:58:54After this, we have a random forest
  4948. 2:58:55regression, which is basically an
  4949. 2:58:57ensemble of predictions of several
  4950. 2:58:59decision tree regressions.
  4951. 2:59:01So, this is all about the types of
  4952. 2:59:02regressions for now. We're going to
  4953. 2:59:04focus on simple linear regression.
  4954. 2:59:06So, let's take a look at what exactly is
  4955. 2:59:08a simple linear regression.
  4956. 2:59:10Simple linear regression is a regression
  4957. 2:59:12technique in which the independent
  4958. 2:59:14variable has a linear relationship with
  4959. 2:59:16the dependent variable.
  4960. 2:59:18The straight line in the diagram is the
  4961. 2:59:19best fit line, and the main goal of the
  4962. 2:59:21simple linear regression is to consider
  4963. 2:59:24the given data points and plot the best
  4964. 2:59:25fit line to fit the model in the best
  4965. 2:59:27way possible.
  4966. 2:59:28So, if you talk about a real-life
  4967. 2:59:30analogy to explain linear regression, we
  4968. 2:59:32can take an example of a car resale
  4969. 2:59:34value. So, we have different parameters,
  4970. 2:59:36you know, when we are talking about
  4971. 2:59:37resale value of a car. Like how many
  4972. 2:59:40years the car has been there in the
  4973. 2:59:41market, and how many kilometers it has
  4974. 2:59:44been driven,
  4975. 2:59:45the kind of mileage the car gives, and
  4976. 2:59:47then we have different parameters we can
  4977. 2:59:49focus upon. And all these independent
  4978. 2:59:51variables somehow are linearly connected
  4979. 2:59:53or interconnected to the price of the
  4980. 2:59:55car.
  4981. 2:59:56So, that is one example to understand
  4982. 2:59:58linear regression. We'll be doing that
  4983. 2:59:59in the use case. I'll be telling you
  4984. 3:00:01about how you can predict the price of
  4985. 3:00:02car.
  4986. 3:00:03Now, talking about linear regression
  4987. 3:00:05terminologies, there are a few
  4988. 3:00:07terminologies that you have to be
  4989. 3:00:08thorough with to begin with linear
  4990. 3:00:10regression.
  4991. 3:00:11So, first of all, we have to talk about
  4992. 3:00:13cost function.
  4993. 3:00:14So, the best fit line can be based on
  4994. 3:00:16the linear equation that is given here.
  4995. 3:00:18So, in this, the dependent variable that
  4996. 3:00:20is to be predicted is denoted by Y.
  4997. 3:00:22A line that touches the Y axis is
  4998. 3:00:24denoted by the intercept B0. The B1 is
  4999. 3:00:27the slope of the line, and X represents
  5000. 3:00:29the independent variables that determine
  5001. 3:00:31the prediction of Y.
  5002. 3:00:33The error in the resultant prediction is
  5003. 3:00:35denoted by E.
  5004. 3:00:36Now, talking about cost function, the
  5005. 3:00:38cost function provides the best possible
  5006. 3:00:40values for B0 and B1 to make the best
  5007. 3:00:43fit line for the data points.
  5008. 3:00:45We do this by converting this problem
  5009. 3:00:46into a minimization problem to get the
  5010. 3:00:49best values for B0 and B1.
  5011. 3:00:51So, with this, the error is minimized in
  5012. 3:00:53this problem between the actual value
  5013. 3:00:55and the predicted value, and we choose
  5014. 3:00:57the function above to minimize.
  5015. 3:00:59Now, we square the error difference and
  5016. 3:01:01sum the error over all the data points.
  5017. 3:01:04The division between the total number of
  5018. 3:01:05data points and the produced value
  5019. 3:01:07provides the average square error for
  5020. 3:01:09all the data points.
  5021. 3:01:11It is also known as mean squared error,
  5022. 3:01:13and we can change the values of B0 and
  5023. 3:01:15B1 so that the MSE or the mean squared
  5024. 3:01:17error value is settled at the minimum.
  5025. 3:01:20So, this is one terminology that is cost
  5026. 3:01:22function that we use in linear
  5027. 3:01:23regression.
  5028. 3:01:24Then, we have the gradient descent.
  5029. 3:01:27So, the next important terminology to
  5030. 3:01:28understand linear regression is gradient
  5031. 3:01:30descent, of course, and it is a method
  5032. 3:01:32of updating B0 and B1 value to reduce
  5033. 3:01:35the MSE, which is the mean squared
  5034. 3:01:36error.
  5035. 3:01:37The idea behind this is to keep
  5036. 3:01:39iterating the B0 and B1 values until we
  5037. 3:01:41reduce the MSE to the minimum.
  5038. 3:01:44Now, to update B0 and B1, we take the
  5039. 3:01:45gradients from the cost function, and to
  5040. 3:01:48find these gradients, we take partial
  5041. 3:01:50derivatives with respect to B0 and B1.
  5042. 3:01:53And these partial derivatives are the
  5043. 3:01:55gradients and are used to update the
  5044. 3:01:57values of B0 and B1.
  5045. 3:01:59I'm sure guys, this is might be a little
  5046. 3:02:00confusing for you guys if you are new to
  5047. 3:02:03this, like gradient descent and cost
  5048. 3:02:04function, but you don't have to worry
  5049. 3:02:06about this because in Python when we're
  5050. 3:02:08using linear regression, we're going to
  5051. 3:02:09be using the scikit-learn or the
  5052. 3:02:11scikit-learn library, so you don't have
  5053. 3:02:12to worry about this. You just have to
  5054. 3:02:14integrate your model with the linear
  5055. 3:02:15regression model that we have already
  5056. 3:02:17over there, and you'll be done with it.
  5057. 3:02:19And when I'm implementing the linear
  5058. 3:02:21regression model, you'll see how easy it
  5059. 3:02:23is to actually implement linear
  5060. 3:02:25regression in Python.
  5061. 3:02:26So, after this, let's talk about a few
  5062. 3:02:28advantages and disadvantages of linear
  5063. 3:02:30regression.
  5064. 3:02:32So, talking about the advantages first,
  5065. 3:02:34linear regression performs exceptionally
  5066. 3:02:36well for linearly separable data. And it
  5067. 3:02:38is actually very easy to implement,
  5068. 3:02:40interpret, and very efficient to train
  5069. 3:02:43as well.
  5070. 3:02:44And even though the linear regression is
  5071. 3:02:45prone to overfitting, it handles it
  5072. 3:02:48pretty well using dimension reduction
  5073. 3:02:49techniques, regularization, and
  5074. 3:02:51cross-validation. And one more advantage
  5075. 3:02:54is that the extrapolation beyond a
  5076. 3:02:56specific data set.
  5077. 3:02:58So, these are all the advantages that we
  5078. 3:02:59have with linear regression. Let's talk
  5079. 3:03:01about a few disadvantages as well.
  5080. 3:03:03So, one of the most common disadvantage
  5081. 3:03:05with linear regression is that it takes
  5082. 3:03:07the assumption of linearity between
  5083. 3:03:09dependent and independent variables. The
  5084. 3:03:11next disadvantage is it is often very
  5085. 3:03:14prone to noise and overfitting as well,
  5086. 3:03:16which is not a very good sign for any
  5087. 3:03:18model if you're doing regression or
  5088. 3:03:19classification in machine learning.
  5089. 3:03:22The next disadvantage is it is very
  5090. 3:03:24quite sensitive to outliers as well.
  5091. 3:03:27And the last one is that it is very
  5092. 3:03:28prone to multicollinearity.
  5093. 3:03:30So, these are all the advantages and
  5094. 3:03:32disadvantages of linear regression.
  5095. 3:03:35>> [music]
  5096. 3:03:39>> So, let's understand the what and why of
  5097. 3:03:41logistic regression. Now, this algorithm
  5098. 3:03:44is most widely used when the dependent
  5099. 3:03:46variable, or you can say the output, is
  5100. 3:03:47in the binary format. So, here you need
  5101. 3:03:50to predict the outcome of a categorical
  5102. 3:03:52dependent variable. So, the outcome
  5103. 3:03:54should be always discrete or categorical
  5104. 3:03:56in nature. Now, by discrete, I mean the
  5105. 3:03:58value should be binary, or you can say
  5106. 3:04:00you just have two values. It can either
  5107. 3:04:02be zero or one. It can either be yes or
  5108. 3:04:05a no. Either be true or false. Or high
  5109. 3:04:07or low. So only these can be the
  5110. 3:04:09outcomes. So the value which you need to
  5111. 3:04:12predict should be discrete or you can
  5112. 3:04:13say categorical in nature. Whereas in
  5113. 3:04:16linear regression we have the value of Y
  5114. 3:04:18or you can say the value you need to
  5115. 3:04:19predict is in a range. So that is how
  5116. 3:04:21there's a difference between linear
  5117. 3:04:22regression and logistic regression. Now
  5118. 3:04:24you must be having a question, why not
  5119. 3:04:26linear regression? Now guys, in linear
  5120. 3:04:28regression the value of Y or the value
  5121. 3:04:30which you need to predict is in a range.
  5122. 3:04:32But in our case, as in the logistic
  5123. 3:04:34regression, we just have two values. It
  5124. 3:04:36can be either zero or it can be one. It
  5125. 3:04:39should not entertain the values which is
  5126. 3:04:40below zero or above one. But in linear
  5127. 3:04:43regression we have the value of Y in the
  5128. 3:04:45range. So here, in order to implement
  5129. 3:04:47logistic regression, we need to clip
  5130. 3:04:48this part. So we don't need the value
  5131. 3:04:51that is below zero or we don't need the
  5132. 3:04:52value which is above one. So since the
  5133. 3:04:54value of Y will be between only zero and
  5134. 3:04:57one, that is the main rule of logistic
  5135. 3:04:58regression, the linear line has to be
  5136. 3:05:00clipped at zero and one. Now once we
  5137. 3:05:02clip this graph, it would look somewhat
  5138. 3:05:04like this. So here you're getting a
  5139. 3:05:06curve which is nothing but three
  5140. 3:05:07different straight lines. So here we
  5141. 3:05:09need to make a new way to solve this
  5142. 3:05:11problem. So this has to be formulated
  5143. 3:05:13into equation and hence we come up with
  5144. 3:05:15logistic regression. So here the outcome
  5145. 3:05:17is either zero or one, which is the main
  5146. 3:05:20rule of logistic regression. So with
  5147. 3:05:21this our resulting curve cannot be
  5148. 3:05:23formulated. So hence our main aim to
  5149. 3:05:25bring the values to zero and one is
  5150. 3:05:26fulfilled. So that is how we came up
  5151. 3:05:28with logistic regression. Now here, once
  5152. 3:05:31it gets formulated into an equation, it
  5153. 3:05:33looks somewhat like this.
  5154. 3:05:35So guys, this is nothing but a S curve
  5155. 3:05:36or you can say the sigmoid curve or
  5156. 3:05:38sigmoid function curve. So this sigmoid
  5157. 3:05:41function basically converts any value
  5158. 3:05:43from minus infinity to infinity to your
  5159. 3:05:45discrete values which a logistic
  5160. 3:05:47regression wants or you can say the
  5161. 3:05:48values which are in binary format,
  5162. 3:05:50either zero or one. So if you see here
  5163. 3:05:53the values are either zero or one. And
  5164. 3:05:55this is nothing but just a transition of
  5165. 3:05:57it. But guys, there's a catch over here.
  5166. 3:05:59So, let's say I have a data point that
  5167. 3:06:01is 0.8. Now, how can you decide whether
  5168. 3:06:04your value is zero or one? Now, here you
  5169. 3:06:07have the concept of threshold, which
  5170. 3:06:09basically divides your line. So, here
  5171. 3:06:11threshold value basically indicates the
  5172. 3:06:13probability of either winning or losing.
  5173. 3:06:16So, here by winning I mean the values
  5174. 3:06:18equals to one, and by losing I mean the
  5175. 3:06:20values equals to zero. But how does it
  5176. 3:06:22do that? Let's say I have data point
  5177. 3:06:24which is over here. Let's say my cursor
  5178. 3:06:26is at 0.8. So, here I'll check whether
  5179. 3:06:28this value is less than my threshold
  5180. 3:06:30value or not. Let's say if it is more
  5181. 3:06:33than my threshold value, it should give
  5182. 3:06:34me the result as one. If it is less than
  5183. 3:06:36that, then it should give me the result
  5184. 3:06:38as zero. So, here my threshold value is
  5185. 3:06:400.5. Now, I need to define that if my
  5186. 3:06:43value, let's say 0.8, it is more than
  5187. 3:06:450.5, then the value shall be rounded off
  5188. 3:06:48to one. And let's say if it is less than
  5189. 3:06:500.5, let's say I have a value 0.2, then
  5190. 3:06:52it should reduce it to zero. So, here
  5191. 3:06:55you can use the concept of threshold
  5192. 3:06:56value to find the output. So, here it
  5193. 3:06:59should be discrete, it should be either
  5194. 3:07:00zero or it should be one.
  5195. 3:07:02So, I hope you caught this curve of
  5196. 3:07:03logistic regression. So, the guys, this
  5197. 3:07:05is the sigmoid S curve.
  5198. 3:07:08So, to make this curve, we need to make
  5199. 3:07:10an equation. So, let me address that
  5200. 3:07:11part as well.
  5201. 3:07:13So, let's see how an equation is formed
  5202. 3:07:14to imitate this functionality. So, over
  5203. 3:07:17here we have an equation of a straight
  5204. 3:07:18line, which is Y is equals to MX + C.
  5205. 3:07:21So, in this case, I just have only one
  5206. 3:07:23independent variable. But let's say if
  5207. 3:07:25we have many independent variable, then
  5208. 3:07:27the equation becomes M1 X1 + M2 X2 + M3
  5209. 3:07:30X3 and so on till MN XN. Now, let us put
  5210. 3:07:34in B and X. So, here the equation
  5211. 3:07:36becomes Y is equals to B1 X1 + B2 X2 +
  5212. 3:07:39B3 X3 and so on till BN XN + C.
  5213. 3:07:44So, guys, the equation of the straight
  5214. 3:07:45line has a range from minus infinity to
  5215. 3:07:47infinity. But in our case, or you can
  5216. 3:07:50say in logistic equation, the value
  5217. 3:07:52which we need to predict or you can say
  5218. 3:07:53the Y value, it can have the range only
  5219. 3:07:55from zero to one. So, in that case, we
  5220. 3:07:57need to transform this equation. So, to
  5221. 3:08:00do that, what we had done, we had just
  5222. 3:08:02divide the equation by 1 - Y. So, now
  5223. 3:08:04when Y is equals to zero, so zero over 1
  5224. 3:08:07- 0 which is equals to 1. So, zero over
  5225. 3:08:091 is again zero. And if we take Y is
  5226. 3:08:12equals to 1, then 1 over 1 - 1 which is
  5227. 3:08:15zero. So, 1 over zero is infinity. So,
  5228. 3:08:17here my range is now between zero to
  5229. 3:08:19infinity. But, again we want the range
  5230. 3:08:21from minus infinity to infinity. So, for
  5231. 3:08:24that, what we'll do, we'll have the log
  5232. 3:08:25of this equation. So, let's go ahead and
  5233. 3:08:27have the logarithmic of this equation.
  5234. 3:08:29So, here we have just transform it
  5235. 3:08:31further to get the range between minus
  5236. 3:08:33infinity to infinity. So, over here we
  5237. 3:08:35have log of Y over 1 - 1 and this is
  5238. 3:08:38your final logistic regression equation.
  5239. 3:08:40So, guys, don't worry, you don't have to
  5240. 3:08:42write this formula or memorize this
  5241. 3:08:44formula. In Python, you just need to
  5242. 3:08:46call this function which is logistic
  5243. 3:08:47regression and everything will be
  5244. 3:08:49automatically for you. So, I don't want
  5245. 3:08:51to scare you with the maths and the
  5246. 3:08:52formulas behind it, but it's always good
  5247. 3:08:54to know how the formula was generated.
  5248. 3:08:57Moving ahead, let us see the various use
  5249. 3:08:58cases wherein logistic regression is
  5250. 3:09:00implemented in real life.
  5251. 3:09:03So, the very first is weather
  5252. 3:09:04prediction.
  5253. 3:09:05Now, logistic regression helps you to
  5254. 3:09:06predict your weather. For example, it is
  5255. 3:09:09used to predict whether it is raining or
  5256. 3:09:10not, whether it is sunny, is it cloudy
  5257. 3:09:13or not. So, all these things can be
  5258. 3:09:15predicted using logistic regression.
  5259. 3:09:17Whereas, you need to keep in mind that
  5260. 3:09:19both linear regression and logistic
  5261. 3:09:20regression can be used in predicting
  5262. 3:09:22weather. So, in that case, linear
  5263. 3:09:24regression helps you to predict what
  5264. 3:09:25will be the temperature tomorrow.
  5265. 3:09:27Whereas, logistic regression will only
  5266. 3:09:29tell you whether it's going to rain or
  5267. 3:09:30not or whether it's cloudy or not,
  5268. 3:09:32whether it's going to snow or not. So,
  5269. 3:09:34these values are discrete. Whereas, if
  5270. 3:09:36you apply linear regression, you're
  5271. 3:09:37predicting things like what is the
  5272. 3:09:39temperature tomorrow or what is the
  5273. 3:09:41temperature day after tomorrow and all
  5274. 3:09:43those things. So, these are the slight
  5275. 3:09:44differences between linear regression
  5276. 3:09:46and logistic regression. Now moving
  5277. 3:09:47ahead, we have classification problem.
  5278. 3:09:50So Python performs multi-class
  5279. 3:09:51classification. So here it can help you
  5280. 3:09:53tell whether it's a bird or it's not a
  5281. 3:09:55bird. Then you classify different kind
  5282. 3:09:57of mammals. Let's say whether it's a dog
  5283. 3:09:59or it's not a dog. Similarly, you can
  5284. 3:10:01check it for reptile whether it's a
  5285. 3:10:03reptile or not a reptile. So in logistic
  5286. 3:10:05regression, it can perform multi-class
  5287. 3:10:07classification. So this point I've
  5288. 3:10:09already discussed that it is used in
  5289. 3:10:10classification problems. Next, it also
  5290. 3:10:13helps you to determine the illness as
  5291. 3:10:14well. So let me take an example. Let's
  5292. 3:10:17say a patient goes for a routine checkup
  5293. 3:10:19in hospital. So what doctor will do it
  5294. 3:10:21it will perform various tests on the
  5295. 3:10:22patient and will check whether the
  5296. 3:10:24patient is actually ill or not. So what
  5297. 3:10:26will be the features? So doctor can
  5298. 3:10:29check the sugar level, the blood
  5299. 3:10:30pressure, then what is the age of the
  5300. 3:10:32patient? Is it very small or is it a old
  5301. 3:10:34person? Then what is the previous
  5302. 3:10:36medical history of the patient? And all
  5303. 3:10:38of these features will be recorded by
  5304. 3:10:40the doctor. And finally, doctor checks
  5305. 3:10:42the patient data and determines the
  5306. 3:10:44outcome of the illness and the severity
  5307. 3:10:46of illness. So using all the data, a
  5308. 3:10:48doctor can identify whether a patient is
  5309. 3:10:51ill or not. So these are the various use
  5310. 3:10:53cases in which you can use logistic
  5311. 3:10:54regression. Now I guess enough of theory
  5312. 3:10:57part, so let's move ahead and see some
  5313. 3:10:59of the practical implementation of
  5314. 3:11:00logistic regression.
  5315. 3:11:02So over here I'll be implementing two
  5316. 3:11:04projects wherein I have the data set of
  5317. 3:11:06a Titanic. So over here we'll predict
  5318. 3:11:08what factors made people more likely to
  5319. 3:11:10survive the sinking of the Titanic ship.
  5320. 3:11:12And in my second project, we'll see the
  5321. 3:11:14data analysis on the SUV cars. So over
  5322. 3:11:16here we have the data of the SUV cars,
  5323. 3:11:18who can purchase it, and what factors
  5324. 3:11:21made people more interested in buying
  5325. 3:11:23SUV.
  5326. 3:11:24So these will be the major questions as
  5327. 3:11:25to why you should implement logistic
  5328. 3:11:27regression and what output will you get
  5329. 3:11:29by it. So let's start by the very first
  5330. 3:11:31project that is Titanic data analysis.
  5331. 3:11:33So some of you might know that there was
  5332. 3:11:35a ship called as Titanic which basically
  5333. 3:11:37hit an iceberg and it sank to the bottom
  5334. 3:11:39of the ocean. And it was a big disaster
  5335. 3:11:42at that time because it was the first
  5336. 3:11:44voyage of the ship and it was supposed
  5337. 3:11:45to be really, really strongly built and
  5338. 3:11:47one of the best ships of that time. So,
  5339. 3:11:49it was a big disaster of that time and
  5340. 3:11:51of course there's a movie about this as
  5341. 3:11:53well. So, many of you might have watched
  5342. 3:11:55it. So, what we have we have data of the
  5343. 3:11:57passengers, those who survived and those
  5344. 3:11:59who did not survive in this particular
  5345. 3:12:00tragedy. So, what you have to do you
  5346. 3:12:02have to look at this data and analyze
  5347. 3:12:04which factors would have been
  5348. 3:12:05contributed the most to the chances of a
  5349. 3:12:08person's survival on the ship or not.
  5350. 3:12:10So, using the logistic regression we can
  5351. 3:12:12predict whether the person survived or
  5352. 3:12:14the person died. Now, apart from this we
  5353. 3:12:16also have a look with the various
  5354. 3:12:17features along with that. So, first let
  5355. 3:12:19us explore the data set. So, over here
  5356. 3:12:21we have the index value. Then the first
  5357. 3:12:24column is passenger ID. Then my next
  5358. 3:12:26column is survived. So, over here we
  5359. 3:12:28have two values, a zero and a one. So,
  5360. 3:12:31zero stands for did not survive and one
  5361. 3:12:33stands for survived. So, this column is
  5362. 3:12:35categorical where the values are
  5363. 3:12:37discrete. Next we have passenger class.
  5364. 3:12:39So, over here we have three values, one,
  5365. 3:12:41two, and three. So, this basically tells
  5366. 3:12:43you that whether a passenger is
  5367. 3:12:45traveling in the first class, second
  5368. 3:12:47class, or third class. Then we have the
  5369. 3:12:49name of the passenger, we have the sex
  5370. 3:12:51or you can say the gender of the
  5371. 3:12:52passenger, whether passenger is a male
  5372. 3:12:54or female. Then we have the age, we have
  5373. 3:12:56the sib SP. So, this basically means the
  5374. 3:12:59number of siblings or the spouses aboard
  5375. 3:13:01the Titanic. So, over here we have
  5376. 3:13:03values such as 1, 0, and so on. Then we
  5377. 3:13:06have parch. So, parch is basically the
  5378. 3:13:09number of parents or children aboard the
  5379. 3:13:11Titanic. So, over here we also have some
  5380. 3:13:13values.
  5381. 3:13:14Then we have the ticket number, we have
  5382. 3:13:16the fare, we have the cabin number, and
  5383. 3:13:18we have the embarked column. So, in my
  5384. 3:13:20embarked column we have three values, we
  5385. 3:13:22have S, C, and Q. So, S basically stands
  5386. 3:13:25for Southampton, C stands for Cherbourg,
  5387. 3:13:27and Q stands for Queenstown.
  5388. 3:13:30So, these are the features that we'll be
  5389. 3:13:31applying our model on. So, here we'll
  5390. 3:13:33perform various steps and then we'll be
  5391. 3:13:35implementing logistic regression. So,
  5392. 3:13:37now these are the various steps which
  5393. 3:13:39are required to implement any algorithm.
  5394. 3:13:41So now in our case we are implementing
  5395. 3:13:43logistic regression. So very first step
  5396. 3:13:45is to collect your data or to import the
  5397. 3:13:47libraries that are used for collecting
  5398. 3:13:49your data and then taking it forward.
  5399. 3:13:51Then my second step is to analyze your
  5400. 3:13:53data. So over here I can go through the
  5401. 3:13:55various fields and then I can analyze
  5402. 3:13:57the data. I can check did the females or
  5403. 3:13:59children survive better than the males
  5404. 3:14:01or did the rich passengers survive more
  5405. 3:14:03than the poor passenger or did the money
  5406. 3:14:05matter as in who paid more to get into
  5407. 3:14:08the ship were they evacuated first and
  5408. 3:14:10what about the workers? Does the worker
  5409. 3:14:12survive or what is the survival rate if
  5410. 3:14:15you were the worker in the ship and not
  5411. 3:14:16just a traveling passenger? So all of
  5412. 3:14:18these are very very interesting
  5413. 3:14:20questions and you would be going through
  5414. 3:14:21all of them one by one. So in this stage
  5415. 3:14:24you need to analyze your data and
  5416. 3:14:25explore your data as much as you can.
  5417. 3:14:28Then my third step is to wrangle your
  5418. 3:14:29data. Now data wrangling basically means
  5419. 3:14:32cleaning your data. So over here you can
  5420. 3:14:34simply remove the unnecessary items or
  5421. 3:14:36if you have a null values in the data
  5422. 3:14:38set you can just clear that data and
  5423. 3:14:40then you can take it forward. So in this
  5424. 3:14:42step you can build your model using the
  5425. 3:14:44train data set and then you can test it
  5426. 3:14:46using the test. So over here you will be
  5427. 3:14:48performing a split which basically split
  5428. 3:14:50your data set into training and testing
  5429. 3:14:52data set and finally you will check the
  5430. 3:14:54accuracy so as to ensure how much
  5431. 3:14:56accurate your values are. So I hope you
  5432. 3:14:58guys got these five steps that we're
  5433. 3:15:00going to implement in logistic
  5434. 3:15:01regression. So now let's go into all
  5435. 3:15:03these steps in detail. So number one we
  5436. 3:15:05have to collect your data or you can say
  5437. 3:15:07import the libraries. So let me show you
  5438. 3:15:09the implementation part as well. So I'll
  5439. 3:15:11just open my Jupiter notebook and I'll
  5440. 3:15:13just implement all of these steps side
  5441. 3:15:15by side.
  5442. 3:15:17So guys this is my Jupiter notebook. So
  5443. 3:15:19first let me just rename Jupiter
  5444. 3:15:21notebook to let's say Titanic data
  5445. 3:15:23analysis.
  5446. 3:15:27Now our first step was to import all the
  5447. 3:15:29libraries and collect the data. So let
  5448. 3:15:31me just import all the libraries first.
  5449. 3:15:33So, first of all, I'll import pandas.
  5450. 3:15:35So, pandas is used for data analysis.
  5451. 3:15:38So, I'll say import pandas as pd. Then,
  5452. 3:15:40I'll be importing NumPy. So, I'll say
  5453. 3:15:42import NumPy as np. So, NumPy is a
  5454. 3:15:45library in Python which basically stands
  5455. 3:15:47for numerical Python. And it is widely
  5456. 3:15:49used to perform any scientific
  5457. 3:15:51computation. Next, we'll be importing
  5458. 3:15:53seaborn. So, seaborn is a library for
  5459. 3:15:55statistical plotting. So, I'll say
  5460. 3:15:57import seaborn as sns. I'll also import
  5461. 3:16:00matplotlib.
  5462. 3:16:01So, matplotlib library is again for
  5463. 3:16:03plotting. So, I'll say import
  5464. 3:16:05matplotlib.pyplot
  5465. 3:16:07as pld.
  5466. 3:16:09Now, to run this library in Jupyter
  5467. 3:16:10Notebook, all I have to write in is
  5468. 3:16:12percentage matplotlib inline.
  5469. 3:16:15Next, I'll be importing one module as
  5470. 3:16:17well. So, as to calculate the basic
  5471. 3:16:20mathematical functions. So, I'll say
  5472. 3:16:22import maths. So, these are the
  5473. 3:16:23libraries that I'll be needing in this
  5474. 3:16:25Titanic data analysis. So, now let me
  5475. 3:16:27just import my dataset. So, I'll take a
  5476. 3:16:29variable, let's say Titanic data. And
  5477. 3:16:32using the pandas, I will just read my
  5478. 3:16:34CSV. Or you can say the dataset.
  5479. 3:16:37I'll write the name of my dataset, that
  5480. 3:16:38is titanic.csv.
  5481. 3:16:40Now, I have already showed you the
  5482. 3:16:42dataset. So, over here, let me just
  5483. 3:16:43print the top 10 rows. So, for that,
  5484. 3:16:45I'll just say I'll take the variable
  5485. 3:16:47Titanic data. head and I'll say the top
  5486. 3:16:5010 rows. Now, I'll just run this. So, to
  5487. 3:16:52run this, I just have to press shift
  5488. 3:16:54plus enter. Or else, you can just
  5489. 3:16:56directly click on the cell.
  5490. 3:16:58So, over here, I have the index. We have
  5491. 3:17:00the passenger ID, which is nothing but
  5492. 3:17:02again the index which is starting from
  5493. 3:17:03one. Then, we have the survived column
  5494. 3:17:05which has the categorical values or you
  5495. 3:17:07can say the discrete values, which is in
  5496. 3:17:09the form of zero or one. Then, we have
  5497. 3:17:11the passenger class. We have the name of
  5498. 3:17:13the passenger, sex, age, and so on. So,
  5499. 3:17:15this is the dataset that I'll be going
  5500. 3:17:17forward with. Next, let us print the
  5501. 3:17:18number of passengers which are there in
  5502. 3:17:20this original dataset. So, for that,
  5503. 3:17:22I'll just simply type in print. I'll say
  5504. 3:17:25number of passengers.
  5505. 3:17:31And using the length function, I can
  5506. 3:17:32calculate the total length. So, I'll say
  5507. 3:17:34length and inside this I'll be passing
  5508. 3:17:36this variable which is Titanic data. So,
  5509. 3:17:38I'll just copy it from here. I'll just
  5510. 3:17:40paste it {dot} index.
  5511. 3:17:42And next, let me just print this one.
  5512. 3:17:45So, here the number of passengers which
  5513. 3:17:46are there in the original data set we
  5514. 3:17:48have is 891. So, around this number were
  5515. 3:17:51traveling in the Titanic ship. So, over
  5516. 3:17:54here my first step is done. We have just
  5517. 3:17:56collected data, imported all the
  5518. 3:17:57libraries, and find out the total number
  5519. 3:17:59of passengers which are traveling in
  5520. 3:18:01Titanic. So, let me just go back to
  5521. 3:18:03presentation and let's see what is my
  5522. 3:18:04next step.
  5523. 3:18:05So, we're done with the collecting data.
  5524. 3:18:07Next step is to analyze your data. So,
  5525. 3:18:09over here we'll be creating different
  5526. 3:18:11plots to check the relationship between
  5527. 3:18:13variables as in how one variable is
  5528. 3:18:15affecting the other. So, you can simply
  5529. 3:18:17explore your data set by making use of
  5530. 3:18:19various columns and then you can plot a
  5531. 3:18:21graph between them. So, you can either
  5532. 3:18:23plot a correlation graph, you can plot a
  5533. 3:18:25distribution graph. It's up to you guys.
  5534. 3:18:27So, let me just go back to my Jupiter
  5535. 3:18:29notebook and let me analyze some of the
  5536. 3:18:30data. Over here my second part is to
  5537. 3:18:32analyze data. So, I'll just put this in
  5538. 3:18:34header two.
  5539. 3:18:36Now, to put this in header two, I just
  5540. 3:18:37have to go on code, click on markdown,
  5541. 3:18:39and I'll just run this.
  5542. 3:18:41So, first let us plot a count plot where
  5543. 3:18:43you can compare between the passengers
  5544. 3:18:44who survived and who did not survive.
  5545. 3:18:46So, for that I'll be using the seaborn
  5546. 3:18:47library. So, over here I have imported
  5547. 3:18:50seaborn as sns. So, I don't have to
  5548. 3:18:52write the whole name. I'll simply say
  5549. 3:18:53sns.count plot.
  5550. 3:18:58I'll say x is equal to survive and the
  5551. 3:19:00data that I'll be using is the Titanic
  5552. 3:19:01data. Or you can say the name of
  5553. 3:19:03variable in which you have stored your
  5554. 3:19:04data set. So, now let me just run this.
  5555. 3:19:07So, over here as you can see I have
  5556. 3:19:09survived column on my x-axis and on the
  5557. 3:19:11y-axis I have the count. So, zero
  5558. 3:19:13basically stands for did not survive and
  5559. 3:19:15one stands for the passengers who did
  5560. 3:19:17survive. So, over here you can see that
  5561. 3:19:19around 550 of the passengers who did not
  5562. 3:19:22survive and there were around 350
  5563. 3:19:24passengers who only survived. So here
  5564. 3:19:26you can basically conclude that there
  5565. 3:19:28are very less survivors than
  5566. 3:19:29non-survivors. So this was the very
  5567. 3:19:32first plot. Now let us plot another plot
  5568. 3:19:34to compare the sex as to whether out of
  5569. 3:19:36all the passengers who survived and who
  5570. 3:19:38did not survive, how many were men and
  5571. 3:19:40how many were female. So to do that I'll
  5572. 3:19:42simply say sns.countplot.
  5573. 3:19:47I'll add the hue as sex.
  5574. 3:19:49So I want to know how many females and
  5575. 3:19:51how many males survived.
  5576. 3:19:53Then I'll be specifying the data. So I'm
  5577. 3:19:54using Titanic data set.
  5578. 3:19:57And let me just run this.
  5579. 3:19:59Okay, I've done a mistake over here.
  5580. 3:20:01So over here you can see I have survived
  5581. 3:20:02column on the x-axis and I have the
  5582. 3:20:04count on the y.
  5583. 3:20:06Now so here your blue color stands for
  5584. 3:20:07your male passengers and orange stands
  5585. 3:20:09for your female.
  5586. 3:20:11So as you can see here the passengers
  5587. 3:20:13who did not survive, that has a value
  5588. 3:20:14zero.
  5589. 3:20:15So we can see that majority of males did
  5590. 3:20:18not survive. And if we see the people
  5591. 3:20:20who survived, here we can see the
  5592. 3:20:22majority of females survived. So this
  5593. 3:20:24basically concludes the gender of the
  5594. 3:20:25survival rate. So it appears on average
  5595. 3:20:28women were more than three times more
  5596. 3:20:30likely to survive than men. Next let us
  5597. 3:20:32plot another plot where we have the hue
  5598. 3:20:34as the passenger class. So over here we
  5599. 3:20:36can see which class that the passenger
  5600. 3:20:38was traveling in, whether it was
  5601. 3:20:39traveling in class one, two or three.
  5602. 3:20:42So for that I'll just write the same
  5603. 3:20:44command. I'll say
  5604. 3:20:45sns.countplot.
  5605. 3:20:49I'll keep my x-axis as only. I'll change
  5606. 3:20:52my hue to passenger class.
  5607. 3:20:54So my variable named as pclass.
  5608. 3:20:57And the data set that I'll be using is
  5609. 3:20:58Titanic data. So this is my result. So
  5610. 3:21:01over here you can see I have blue for
  5611. 3:21:03first class, orange for second class and
  5612. 3:21:05green for the third class.
  5613. 3:21:07So here the passengers who did not
  5614. 3:21:09survive were majorly of the third class
  5615. 3:21:11or you can see the lowest class or the
  5616. 3:21:12cheapest class to get into the Titanic.
  5617. 3:21:15And the people who did survive majorly
  5618. 3:21:16belong to the higher classes. So here
  5619. 3:21:18one and two has more rise than the
  5620. 3:21:20passenger who were traveling in the
  5621. 3:21:21third class. So here we have concluded
  5622. 3:21:24that the passengers who did not survive
  5623. 3:21:26are majorly of third class or you can
  5624. 3:21:27say the lowest class. And the passengers
  5625. 3:21:30who were traveling in first and second
  5626. 3:21:31class would tend to survive more. Next
  5627. 3:21:34let us plot a graph for the age
  5628. 3:21:35distribution. Over here I can simply use
  5629. 3:21:37my data. So we'll be using pandas
  5630. 3:21:39library for this. I'll declare a array
  5631. 3:21:42and I'll pass in the column that is age.
  5632. 3:21:45So I plot and I want a histogram so I'll
  5633. 3:21:47say plot.hist.
  5634. 3:21:51So you can notice over here that we have
  5635. 3:21:53more of young passengers or you can see
  5636. 3:21:55the children between the ages zero to
  5637. 3:21:5710. And then we have the average age
  5638. 3:21:59people. And if you go ahead lesser would
  5639. 3:22:01be the population. So this is the
  5640. 3:22:03analysis on the age column. So we saw
  5641. 3:22:06that we have more young passengers and
  5642. 3:22:08more mediocre age passengers who are
  5643. 3:22:10traveling in the Titanic.
  5644. 3:22:11So next let me plot a graph of fare as
  5645. 3:22:13well. So I'll say Titanic data.
  5646. 3:22:17I'll say fare.
  5647. 3:22:18And again I'll plot a histogram so I'll
  5648. 3:22:20say hist.
  5649. 3:22:23So here you can see the fare size is
  5650. 3:22:25between zero to 100. Now let me add the
  5651. 3:22:27bin size so as to make it more clear.
  5652. 3:22:30So over here I'll say bin is equals to
  5653. 3:22:32let's say 20 and I'll increase the
  5654. 3:22:34figure size as well. So I'll say fix
  5655. 3:22:36size. Let's say I'll give the dimensions
  5656. 3:22:39as 10 by 5.
  5657. 3:22:41So it is bins. So this is more clear
  5658. 3:22:43now. Next let us analyze the other
  5659. 3:22:45columns as well.
  5660. 3:22:47So I'll just type in Titanic data. And I
  5661. 3:22:50want the information as to what all
  5662. 3:22:51columns are left.
  5663. 3:22:54So here we have passenger ID which I
  5664. 3:22:56guess it's of no use. Then we have to
  5665. 3:22:58see how many passengers survived and how
  5666. 3:23:00many did not. We also see the analysis
  5667. 3:23:02on the gender basis. We saw whether
  5668. 3:23:03female tend to survive more or the men
  5669. 3:23:05tend to survive more. Then we saw the
  5670. 3:23:07passenger class where the passenger is
  5671. 3:23:09traveling in in first class, second
  5672. 3:23:10class or third class. Then we have the
  5673. 3:23:12name, so in name we cannot do any
  5674. 3:23:14analysis. We saw the sex, we saw the age
  5675. 3:23:17as well. Then we have sibsp. So this
  5676. 3:23:20stands for the number of siblings or the
  5677. 3:23:22spouses which are aboard the Titanic. So
  5678. 3:23:24let us do this as well. So I'll say
  5679. 3:23:26sns.countplot.
  5680. 3:23:30I'll mention x as sibsp.
  5681. 3:23:33And I'll be using the Titanic data.
  5682. 3:23:36So you can see the plot over here. So
  5683. 3:23:38over here you can conclude that it has
  5684. 3:23:40the maximum value on zero. So you can
  5685. 3:23:42conclude that neither a children nor a
  5686. 3:23:44spouse was on board the Titanic. The
  5687. 3:23:46second most highest value is one. And
  5688. 3:23:49then we have very less values for two,
  5689. 3:23:51three, four, and so on.
  5690. 3:23:53Next, if I go above, we saw this column
  5691. 3:23:55as well. Similarly, you can do for
  5692. 3:23:56parch.
  5693. 3:23:57So next we have parch, or you can say
  5694. 3:23:59the number of parents or children which
  5695. 3:24:00were aboard the Titanic. So you
  5696. 3:24:02similarly can do this as well. Then we
  5697. 3:24:04have the ticket number. So I don't think
  5698. 3:24:06so any analysis is required for ticket.
  5699. 3:24:08Then we have fare. So fare we have
  5700. 3:24:10already discussed as in the people who
  5701. 3:24:12tend to travel in the first class
  5702. 3:24:13usually pay the highest fare. Then we
  5703. 3:24:15have the cabin number, and we have
  5704. 3:24:17embarked. So these are the columns that
  5705. 3:24:19we'll be doing data wrangling on.
  5706. 3:24:21So we have analyzed the data, and we
  5707. 3:24:22have seen quite a few graphs in which we
  5708. 3:24:24can conclude which variable is better
  5709. 3:24:27than the another or or what are the
  5710. 3:24:28relationship they hold.
  5711. 3:24:29So third step is my data wrangling. So
  5712. 3:24:31data wrangling basically means cleaning
  5713. 3:24:33your data. So if you have a large data
  5714. 3:24:36set, you might be having some null
  5715. 3:24:37values, or you can say NaN values. So
  5716. 3:24:39it's very important that you remove all
  5717. 3:24:41the unnecessary items that are present
  5718. 3:24:43in your data set. So removing this
  5719. 3:24:45directly affects your accuracy. So I'll
  5720. 3:24:47just go ahead and clean my data by
  5721. 3:24:49removing all the NaN values and
  5722. 3:24:51unnecessary columns which has a null
  5723. 3:24:52value in the data set. So next I'll be
  5724. 3:24:54performing data wrangling.
  5725. 3:25:02So first of all, I'll check whether my
  5726. 3:25:03data set is null or not. So I'll say
  5727. 3:25:06Titanic data, which is the name of my
  5728. 3:25:07data set, and I'll say is null. So, this
  5729. 3:25:10will basically tell me what all values
  5730. 3:25:12are null, and it will return me a
  5731. 3:25:13Boolean result. So, this basically
  5732. 3:25:15checks the missing data, and your result
  5733. 3:25:17will be in Boolean format, as in the
  5734. 3:25:19result will be in true or false. So,
  5735. 3:25:20false mean if it is not null, and true
  5736. 3:25:23means if it is null. So, let me just run
  5737. 3:25:25this.
  5738. 3:25:27Over here, you can see the values as
  5739. 3:25:29false or true. So, false is where the
  5740. 3:25:31value is not null, and true is where the
  5741. 3:25:34value is null. So, over here, you can
  5742. 3:25:35see in the cabin column, we have the
  5743. 3:25:37very first value, which is null. So, we
  5744. 3:25:39have to do something on this.
  5745. 3:25:41So, you can see that we have a large
  5746. 3:25:43data set.
  5747. 3:25:44The counting does not stop, and we can
  5748. 3:25:47actually see the sum of it. We can
  5749. 3:25:48actually print the number of passengers
  5750. 3:25:50who have the NaN value in each column.
  5751. 3:25:52So, I'll say Titanic_data
  5752. 3:25:55is null, and I want the sum of it. So,
  5753. 3:25:57I'll say dot sum. So, this will
  5754. 3:25:59basically print the number of passengers
  5755. 3:26:01who have the NaN values in each column.
  5756. 3:26:03So, we can see that we have missing
  5757. 3:26:05values in age column, that is 177. Then,
  5758. 3:26:07we have the maximum value in the cabin
  5759. 3:26:09column, and we have very less in the
  5760. 3:26:11embarked column, that is two.
  5761. 3:26:13So, here, if you don't want to see these
  5762. 3:26:15numbers, you can also plot a heat map,
  5763. 3:26:17and then you can visually analyze it.
  5764. 3:26:19So, let me just do that as well. So,
  5765. 3:26:20I'll say sns.heatmap
  5766. 3:26:26and say yticklabels
  5767. 3:26:31to false. So, I'll just run this. So, as
  5768. 3:26:33we have already seen that there were
  5769. 3:26:35three columns in which missing data
  5770. 3:26:36value was present. So, this might be
  5771. 3:26:38age. So, over here, almost 20% of age
  5772. 3:26:41column has a missing value. Then, we
  5773. 3:26:43have the cabin columns. So, this is
  5774. 3:26:44quite a large value, and then we have
  5775. 3:26:46two values for embarked column as well.
  5776. 3:26:49Add a cmap for color coding. So, I'll
  5777. 3:26:51say cmap
  5778. 3:26:55So, if I do this, so the graph becomes
  5779. 3:26:57more attractive. So, over here, your
  5780. 3:26:59yellow stands for true, or you can say
  5781. 3:27:01the values are null.
  5782. 3:27:03So, here we have concluded that we have
  5783. 3:27:05the missing value of age. We have a lot
  5784. 3:27:07of missing values in the cabin column
  5785. 3:27:09and we have very less value, which is
  5786. 3:27:11not even visible in the embark column as
  5787. 3:27:13well.
  5788. 3:27:14So, to remove these missing values, you
  5789. 3:27:16can either replace the values and you
  5790. 3:27:18can put in some dummy values to it or
  5791. 3:27:20you can simply drop the column.
  5792. 3:27:22So, here let us first pick the age
  5793. 3:27:24column. So, first let me just plot a box
  5794. 3:27:26plot and then we analyze with having a
  5795. 3:27:28column as age. So, I'll say sns.
  5796. 3:27:31boxplot
  5797. 3:27:33I'll say x is equals to passenger class.
  5798. 3:27:36So, it's pclass. I'll say y is equals to
  5799. 3:27:38age.
  5800. 3:27:39And the data set that I'll be using is
  5801. 3:27:41Titanic set. So, I'll say data is equals
  5802. 3:27:43to Titanic data.
  5803. 3:27:45You can see the age in first class and
  5804. 3:27:47second class tends to be more older
  5805. 3:27:49rather than we have it in the third
  5806. 3:27:50class. Well, that depends on the
  5807. 3:27:52experience, how much you earn or might
  5808. 3:27:54be there any number of reasons.
  5809. 3:27:56So, here we concluded that passengers
  5810. 3:27:57who were traveling in class one and
  5811. 3:27:59class two attend to be older than what
  5812. 3:28:01we have in the class three.
  5813. 3:28:03So, we have found that we have some
  5814. 3:28:05missing values in M.
  5815. 3:28:06Now, one way is to either just drop the
  5816. 3:28:08column or you can just simply fill in
  5817. 3:28:10some values to that. So, this method is
  5818. 3:28:12called as imputation.
  5819. 3:28:14Now, to perform data wrangling or
  5820. 3:28:16cleaning, let us first print the head of
  5821. 3:28:17the data set. So, I'll say Titanic.head.
  5822. 3:28:20Sorry, it's Titanic_data.
  5823. 3:28:23Let's say I just want the five rows.
  5824. 3:28:25So, here we have survived, which is
  5825. 3:28:27again categorical. So, in this
  5826. 3:28:28particular column, I can apply logistic
  5827. 3:28:30regression. So, this can be my y value
  5828. 3:28:33or the value that you need to predict.
  5829. 3:28:35Then, we have the passenger class. We
  5830. 3:28:36have the name. Then, we have ticket
  5831. 3:28:38number, fare, cabin. So, over here we
  5832. 3:28:41have seen that in cabin we have a lot of
  5833. 3:28:43null values or you can say the NaN
  5834. 3:28:44values, which is quite visible as well.
  5835. 3:28:46So, first of all, we'll just drop this
  5836. 3:28:48column. So, for dropping it, I'll just
  5837. 3:28:50say Titanic_data.
  5838. 3:28:52And I'll simply type in drop and the
  5839. 3:28:53column which I need to drop. So, I have
  5840. 3:28:56to drop the cabin column.
  5841. 3:28:58I mention the axis equals to one and
  5842. 3:29:00I'll say in place also to true.
  5843. 3:29:04So, now again I'll just print the head
  5844. 3:29:05and let us see whether this column has
  5845. 3:29:07been removed from the data set or not.
  5846. 3:29:09So, I'll say Titanic .head.
  5847. 3:29:12So, as you can see here we don't have
  5848. 3:29:14cabin column anymore.
  5849. 3:29:16Now, you can also drop the NA values.
  5850. 3:29:18So, I'll say Titanic data
  5851. 3:29:20.drop all the NA values or you can say
  5852. 3:29:22NaN which is not a number and I'll say
  5853. 3:29:25in place is equals to true.
  5854. 3:29:27Let's Titanic
  5855. 3:29:29So, over here let me again plot the heat
  5856. 3:29:31map and let's see all the values which
  5857. 3:29:33were before showing a lot of null values
  5858. 3:29:35has it been removed or not. So, I'll say
  5859. 3:29:37sns.heatmap I'll pass in the data set.
  5860. 3:29:41I'll check if this null.
  5861. 3:29:43I'll say white labels is equals to
  5862. 3:29:45false.
  5863. 3:29:47And I don't want color coding, so again
  5864. 3:29:49I'll say false.
  5865. 3:29:51So, this will basically help me to check
  5866. 3:29:53whether my values has been removed from
  5867. 3:29:55the data set or not. So, as you can see
  5868. 3:29:56here I don't have any null values. So,
  5869. 3:29:59it's entirely black.
  5870. 3:30:01Now, you can actually know the sum as
  5871. 3:30:02well, so I'll just go above.
  5872. 3:30:05So, I'll just copy this part and I'll
  5873. 3:30:07just use the sum function to calculate
  5874. 3:30:09the sum.
  5875. 3:30:10So, here that tells me the data set is
  5876. 3:30:12clean as in the data set does not
  5877. 3:30:14contain any null value or any NaN value.
  5878. 3:30:18So, now we have wrangled our data or you
  5879. 3:30:20can say clean our data.
  5880. 3:30:21So, here we have done just one step in
  5881. 3:30:23data wrangling that is just removing one
  5882. 3:30:25column out of it. Now, you can do a lot
  5883. 3:30:27of things. You can actually fill in the
  5884. 3:30:28values with some other values or you can
  5885. 3:30:31just calculate the mean and then you can
  5886. 3:30:32just fit in the null values.
  5887. 3:30:34But, now if I see my data set So, I'll
  5888. 3:30:37say Titanic data.head.
  5889. 3:30:39But, now if I see over here I have a lot
  5890. 3:30:41of string values. So, this has to be
  5891. 3:30:43converted to a categorical variables in
  5892. 3:30:45order to implement logistic regression.
  5893. 3:30:47So, what we will do we will convert this
  5894. 3:30:49to categorical variable into some dummy
  5895. 3:30:51variables and this can be done using
  5896. 3:30:53pandas because logistic regression just
  5897. 3:30:55take two values.
  5898. 3:30:57So whenever you apply machine learning,
  5899. 3:30:58you need to make sure that there are no
  5900. 3:31:00string values present because it won't
  5901. 3:31:02be taking these as your input variables.
  5902. 3:31:05So using string, you don't have to
  5903. 3:31:06predict anything. But in my case, I have
  5904. 3:31:08the survived columns, so I need to
  5905. 3:31:09predict how many people tend to survive
  5906. 3:31:11and how many did not. So zero stands for
  5907. 3:31:13did not survive and one stands for
  5908. 3:31:15survive.
  5909. 3:31:16So now let me just convert these
  5910. 3:31:17variables into dummy variables.
  5911. 3:31:20So let's use pandas and I'll say
  5912. 3:31:22pd.get_dummies.
  5913. 3:31:24You can simply press tab to auto
  5914. 3:31:26complete. I'll say Titanic data.
  5915. 3:31:29And I'll pass the sex.
  5916. 3:31:30So you can just simply click on shift
  5917. 3:31:32plus tab to get more information on
  5918. 3:31:34this.
  5919. 3:31:35So here we have the type data frame and
  5920. 3:31:37we have the passenger ID survived and
  5921. 3:31:39passenger class.
  5922. 3:31:40So if you run this, you'll see that zero
  5923. 3:31:42basically stands for not a female and
  5924. 3:31:44one stand for it is a female. Similarly
  5925. 3:31:46for male, zero stands for it's not male
  5926. 3:31:48and one stand for it's male. Now we
  5927. 3:31:50don't require both these columns because
  5928. 3:31:52one column itself is enough to tell us
  5929. 3:31:55whether it's male or you can say female
  5930. 3:31:56or not. So let's say if I want to keep
  5931. 3:31:58only male, I'll say if the value of male
  5932. 3:32:01is one, so it is definitely a male and
  5933. 3:32:03it is not a female. So that is how it
  5934. 3:32:05you don't need both of these values. So
  5935. 3:32:07for that, I'll just remove the first
  5936. 3:32:09column, let's say a female.
  5937. 3:32:10So I'll say drop first
  5938. 3:32:13and true.
  5939. 3:32:15So over here it has given me just one
  5940. 3:32:17column which is male and has the value
  5941. 3:32:19zero and one. Let me just set this as a
  5942. 3:32:22variable, let's say sex. So over here I
  5943. 3:32:24can say sex.head.
  5944. 3:32:26I just want to see the first five rows.
  5945. 3:32:29Sorry, it's dot.
  5946. 3:32:31So this is how my data looks like.
  5947. 3:32:34Now here we have done it for sex, then
  5948. 3:32:35we have the numerical values in age, we
  5949. 3:32:37have the numerical values in spouses,
  5950. 3:32:39then we have the ticket number, we have
  5951. 3:32:41the fare and we have embarked as well.
  5952. 3:32:43So in embarked, the values are in S, C
  5953. 3:32:45and Q. So here we can apply this get
  5954. 3:32:48dummy function.
  5955. 3:32:49So, let's say I'll take a variable,
  5956. 3:32:51let's say embarked.
  5957. 3:32:53I'll use the pandas library.
  5958. 3:32:57I'll enter the column name, that is
  5959. 3:32:59embarked.
  5960. 3:33:03So, let me just print the head of it.
  5961. 3:33:04So, I'll say embarked.head.
  5962. 3:33:07So, over here we have C, Q, and S. Now,
  5963. 3:33:09here also we can drop the first column
  5964. 3:33:11because these two values are enough
  5965. 3:33:13whether the passenger is either
  5966. 3:33:15traveling from Q, that is Queenstown, S
  5967. 3:33:17for Southampton. And if both the values
  5968. 3:33:18are zero, then definitely the passenger
  5969. 3:33:20is from Cherbourg, that is the third
  5970. 3:33:22value. So, you can again drop the first
  5971. 3:33:24value. So, I'll say drop
  5972. 3:33:27and true.
  5973. 3:33:28Let me just run this. So, this is how my
  5974. 3:33:29output looks like. Now, similarly you
  5975. 3:33:31can do it for passenger class as well.
  5976. 3:33:33So, here also we have three classes,
  5977. 3:33:35one, two, and three.
  5978. 3:33:36So, I'll just copy the whole statement.
  5979. 3:33:42So, let's say I want the variable name,
  5980. 3:33:44let's say PCL.
  5981. 3:33:47I'll pass in the column name, that is P
  5982. 3:33:48class, and I'll just drop the first
  5983. 3:33:51column. So, here also the values would
  5984. 3:33:53be one, two, or three, and I'll just
  5985. 3:33:55remove the first column. So, here we
  5986. 3:33:57just left with two and three. So, if
  5987. 3:33:59both the values are zero, then
  5988. 3:34:00definitely the passenger is traveling in
  5989. 3:34:02the first class. Now, we have made the
  5990. 3:34:04values as categorical. Now, my next step
  5991. 3:34:06would be to concatenate all these new
  5992. 3:34:09rows into a data set.
  5993. 3:34:11I can say Titanic data. Using the
  5994. 3:34:13pandas, we'll just concatenate all these
  5995. 3:34:15columns. So, I'll say pd.concat.
  5996. 3:34:18And I'll say we have to concatenate sex,
  5997. 3:34:21we have to concatenate embarked and PCL.
  5998. 3:34:24And then I'll mention the axis to one.
  5999. 3:34:26I'll just run this.
  6000. 3:34:28Okay, I need to print the head.
  6001. 3:34:30So, over here you can see that these
  6002. 3:34:32columns have been added over here. So,
  6003. 3:34:34we have the male column which basically
  6004. 3:34:36tells whether a person is male or it's a
  6005. 3:34:38female. Then we have the embarked which
  6006. 3:34:40is basically Q and S. So, if it's
  6007. 3:34:43traveling from Queenstown, value would
  6008. 3:34:44be one, else it would be zero. And if
  6009. 3:34:46both of these values are zero, it is
  6010. 3:34:48definitely traveling from Cherbourg.
  6011. 3:34:50Then we have the passenger class as two
  6012. 3:34:52and three. So, if the value of both
  6013. 3:34:54these is zero, then the passenger is
  6014. 3:34:56traveling in class one.
  6015. 3:34:58So, I hope you got this till now.
  6016. 3:35:00Now, these are the irrelevant columns
  6017. 3:35:02that we have it over here. So, we can
  6018. 3:35:04just drop these columns. We're dropping
  6019. 3:35:05P class,
  6020. 3:35:07the embarked column,
  6021. 3:35:09and the sex column.
  6022. 3:35:10So, I'll just type in Titanic data
  6023. 3:35:13{dot} drop. I'll mention the columns
  6024. 3:35:15that I want to drop. So, I'll say
  6025. 3:35:22I'll even delete the passenger ID
  6026. 3:35:24because it's nothing but just the index
  6027. 3:35:25value, which is starting from one.
  6028. 3:35:27So, I'll drop this as well.
  6029. 3:35:29Then I don't want name as well, so I'll
  6030. 3:35:31delete name as well.
  6031. 3:35:32Then what else we can drop? We can drop
  6032. 3:35:34the ticket as well.
  6033. 3:35:37And then I'll just mention the axis.
  6034. 3:35:40I'll say in place is equals to true.
  6035. 3:35:43Okay, so the my column name starts from
  6036. 3:35:45upper case.
  6037. 3:35:47So, these have been dropped. Now, let me
  6038. 3:35:48just print my data set again.
  6039. 3:35:51So, this is my final data set, guys. We
  6040. 3:35:52have the survived column, which has the
  6041. 3:35:54value zero and one. Then we have the
  6042. 3:35:56passenger class. Oh, we forgot to drop
  6043. 3:35:58this as well. So, no worries. I'll drop
  6044. 3:36:00this again.
  6045. 3:36:05So, now let me just run this.
  6046. 3:36:08So, over here we have the survived, we
  6047. 3:36:10have the age, we have the sibsp, we have
  6048. 3:36:12the parch, we have fare, male, and these
  6049. 3:36:15we have just converted.
  6050. 3:36:17So, here we have just performed data
  6051. 3:36:18wrangling, or you can say clean the
  6052. 3:36:20data. And then we have just converted
  6053. 3:36:22the values of gender to male, then
  6054. 3:36:25embarked to Q and S, and the passenger
  6055. 3:36:27class to two and three. So, this was all
  6056. 3:36:29about my data wrangling, or just
  6057. 3:36:30cleaning the data.
  6058. 3:36:32Then my next step is training and
  6059. 3:36:33testing your data. So, here we will
  6060. 3:36:35split the data set into train subset and
  6061. 3:36:37test subset. And then what we'll do
  6062. 3:36:39we'll build a model on the train data
  6063. 3:36:41and then predict the output on your test
  6064. 3:36:43data set. So let me just go back to
  6065. 3:36:45Jupiter and let us implement this as
  6066. 3:36:46well.
  6067. 3:36:47Over here I need to train my data set.
  6068. 3:36:49So I'll just put this in date heading
  6069. 3:36:51three.
  6070. 3:36:54So over here you need to define your
  6071. 3:36:55dependent variable and independent
  6072. 3:36:57variable.
  6073. 3:36:58So here my Y is the output or you can
  6074. 3:37:00say the value that I need to predict.
  6075. 3:37:03So over here I'll write Titanic data.
  6076. 3:37:06I'll take the column which is survive.
  6077. 3:37:09So basically I have to predict this
  6078. 3:37:11column whether the passenger survived or
  6079. 3:37:12not. And as you can see we have the
  6080. 3:37:14discrete outcome which is in the form of
  6081. 3:37:16zero and one. And rest all the things we
  6082. 3:37:19can take it as a features or you can say
  6083. 3:37:21independent variable. So I'll say
  6084. 3:37:22Titanic data
  6085. 3:37:24.drop.
  6086. 3:37:26So we'll just simply drop the survive
  6087. 3:37:28and all the other columns will be my
  6088. 3:37:29independent variable.
  6089. 3:37:31So everything else are the features
  6090. 3:37:32which leads to the survival rate. So
  6091. 3:37:34once we have defined the independent
  6092. 3:37:36variable and the dependent variable,
  6093. 3:37:38next step is to split your data into
  6094. 3:37:39training and testing subset. So for that
  6095. 3:37:42we'll be using SK learn. I'll just type
  6096. 3:37:44in from SK learn.cross_validation
  6097. 3:37:48import train_test_split.
  6098. 3:37:54Now here if you just click on shift and
  6099. 3:37:56tab, you can go to the documentation and
  6100. 3:37:59you can just see the examples over here.
  6101. 3:38:02I'll click on plus to open it. And then
  6102. 3:38:04I'll just go to examples and see how you
  6103. 3:38:06can split your data. So over here you
  6104. 3:38:08have X_train, X_test, Y_train, Y_test.
  6105. 3:38:12And then using this train_test_split you
  6106. 3:38:13can just pass in your independent
  6107. 3:38:15variable and dependent variable and just
  6108. 3:38:17define a size and a random state to it.
  6109. 3:38:19So let me just copy this.
  6110. 3:38:21And I'll just paste it over here.
  6111. 3:38:24Over here we'll train_test. Then we have
  6112. 3:38:27the dependent variable train and test.
  6113. 3:38:29And using the split function we'll pass
  6114. 3:38:30in the independent and dependent
  6115. 3:38:32variable and then we'll set a split
  6116. 3:38:34size. So, let's say I'll put it at 0.3.
  6117. 3:38:37So, this basically means that your data
  6118. 3:38:38set is divided in 0.3. That is in 70/30
  6119. 3:38:41ratio. And then I can add any random
  6120. 3:38:43state to it. So, let's say I'm applying
  6121. 3:38:45one. This is not necessary. If you want
  6122. 3:38:48the same result as that of mine, you can
  6123. 3:38:49add the random state. So, this will
  6124. 3:38:51basically take exactly the same sample
  6125. 3:38:53every time.
  6126. 3:38:55Next, I have to train and predict by
  6127. 3:38:57creating a model. So, here logistic
  6128. 3:38:59regression will grab from the linear
  6129. 3:39:01regression. So, next I'll just type in
  6130. 3:39:03from sklearn.linear_model
  6131. 3:39:07import LogisticRegression.
  6132. 3:39:09Next, I'll just create the instance of
  6133. 3:39:11this logistic regression model. So, I'll
  6134. 3:39:13say log model
  6135. 3:39:15is equals to LogisticRegression.
  6136. 3:39:17Now, I just need to fit my model. So,
  6137. 3:39:19I'll say log model.fit and I'll just
  6138. 3:39:22pass in my X_train
  6139. 3:39:25and Y_train.
  6140. 3:39:28All right. So, here it gives me all the
  6141. 3:39:30details of logistic regression.
  6142. 3:39:32So, here it gives me the class weight,
  6143. 3:39:34dual, fit intercept and all those
  6144. 3:39:35things.
  6145. 3:39:36Then, what I need to do, I need to make
  6146. 3:39:38prediction. So, I'll take a variable,
  6147. 3:39:40let's say predictions, and I'll pass on
  6148. 3:39:42the model to it. So, I'll say log
  6149. 3:39:44model.predict
  6150. 3:39:46and I'll pass in the value that is
  6151. 3:39:48X_test. So, here we have just created a
  6152. 3:39:50model, fit that model, and then we have
  6153. 3:39:52made predictions. So, now to evaluate
  6154. 3:39:54how my model has been performing, so you
  6155. 3:39:56can simply calculate the accuracy or you
  6156. 3:39:58can also calculate the classification
  6157. 3:40:00report. So, don't worry, guys. I'll be
  6158. 3:40:02showing both of these methods. So, I'll
  6159. 3:40:04say from sklearn.metrics
  6160. 3:40:08import classification_report.
  6161. 3:40:11So, over here I'll use
  6162. 3:40:12classification_report and inside this
  6163. 3:40:14I'll be passing in Y_test and the
  6164. 3:40:16predictions.
  6165. 3:40:21So, guys, this is my classification
  6166. 3:40:22report. So, over here I have the
  6167. 3:40:24precision, I have the recall, we have
  6168. 3:40:26the F1 score, and then we have support.
  6169. 3:40:29So, here we have the value of precision
  6170. 3:40:31as 75, 72, and 73, which is not that
  6171. 3:40:34bad. Now, in order to calculate the
  6172. 3:40:36accuracy as well, you can also use the
  6173. 3:40:38concept of confusion matrix. So, if you
  6174. 3:40:40want to print the confusion matrix, I'll
  6175. 3:40:42simply say from SK learn {dot} metrics
  6176. 3:40:46import confusion matrix first of all,
  6177. 3:40:48and then we'll just print this.
  6178. 3:40:51So, here my function has been imported
  6179. 3:40:52successfully, so I'll say confusion
  6180. 3:40:54matrix.
  6181. 3:40:56And I'll again pass in the same
  6182. 3:40:57variables, which is Y test and
  6183. 3:40:59predictions.
  6184. 3:41:01So, I hope you guys already know the
  6185. 3:41:03concept of confusion matrix. So, can you
  6186. 3:41:05guys give me a quick confirmation as to
  6187. 3:41:07whether you guys remember this confusion
  6188. 3:41:09matrix concept or not? So, if not, I can
  6189. 3:41:11just quickly summarize this as well.
  6190. 3:41:14Okay, Jagriti says a yes.
  6191. 3:41:16Okay, Swati is not clear with this. So,
  6192. 3:41:18I'll just tell you in a brief what
  6193. 3:41:19confusion matrix is all about.
  6194. 3:41:22So, confusion matrix is nothing but a 2
  6195. 3:41:24by 2 matrix, which has a four outcomes.
  6196. 3:41:27This basically tells us that how
  6197. 3:41:28accurate your values are. So, here we
  6198. 3:41:30have the column as predicted no,
  6199. 3:41:32predicted yes,
  6200. 3:41:34and we have actual no and an actual yes.
  6201. 3:41:38So, this is the concept of confusion
  6202. 3:41:40matrix. So, here let me just feed in
  6203. 3:41:42these values which we have just
  6204. 3:41:43calculated. So, here we have 105,
  6205. 3:41:48105, 21, 25, and 63.
  6206. 3:41:53So, as you can see here, we have got
  6207. 3:41:55four outcomes. Now, 105 is the value
  6208. 3:41:58where our model has predicted no, and in
  6209. 3:42:00reality it was also a no. So, here we
  6210. 3:42:03have predicted no and an actual no.
  6211. 3:42:05Similarly, we have 63 as a predicted
  6212. 3:42:07yes, so here the model predicted yes,
  6213. 3:42:09and actually also it was a yes. So, in
  6214. 3:42:12order to calculate the accuracy, you
  6215. 3:42:13just need to add the sum of these two
  6216. 3:42:15values and just divide the whole by the
  6217. 3:42:18sum. So, here these two values tells me
  6218. 3:42:20where the model has actually predicted
  6219. 3:42:22the correct output. So, this value is
  6220. 3:42:24also called as true negative. This is
  6221. 3:42:26called as false positive. This is called
  6222. 3:42:28as true positive and this is called as
  6223. 3:42:30false negative. Now, in order to
  6224. 3:42:31calculate the accuracy, you don't have
  6225. 3:42:33to do it manually. So, in Python, you
  6226. 3:42:35can just import accuracy score function
  6227. 3:42:38and you can get the results from that.
  6228. 3:42:39So, I'll just do that as well. So, I'll
  6229. 3:42:41say from sklearn.metrics
  6230. 3:42:44import accuracy score
  6231. 3:42:46and I'll simply print the accuracy.
  6232. 3:42:49I'll pass in the same variables, that is
  6233. 3:42:51Y test and predictions.
  6234. 3:42:53So, over here it tells me the accuracy
  6235. 3:42:54as 78, which is quite good. So, over
  6236. 3:42:57here if you want to do it manually, you
  6237. 3:42:59have to plus these two numbers, which is
  6238. 3:43:01105 + 63. So, this comes out to almost
  6239. 3:43:04168.
  6240. 3:43:06And then you have to divide it by the
  6241. 3:43:07sum of all the four numbers. So, 105 +
  6242. 3:43:1063 + 21 + 25. So, this gives me a result
  6243. 3:43:13of 214. So, now if you divide these two
  6244. 3:43:16number, you'll get the same accuracy
  6245. 3:43:18that is 78% or you can say 0.78.
  6246. 3:43:21So, that is how you can calculate the
  6247. 3:43:23accuracy. So, now let me just go back to
  6248. 3:43:25my presentation and let's see what all
  6249. 3:43:27we have covered till now.
  6250. 3:43:28So, here we have first split our data
  6251. 3:43:30into train and test subset. Then we have
  6252. 3:43:32built our model on the train data and
  6253. 3:43:34then predicted the output on the test
  6254. 3:43:36data set. And then my fifth step is to
  6255. 3:43:38check the accuracy. So, here we have
  6256. 3:43:40calculated accuracy to almost 78%, which
  6257. 3:43:43is quite good. You cannot say that
  6258. 3:43:45accuracy is bad.
  6259. 3:43:46So, here it tells me how accurate your
  6260. 3:43:48results are. So, here my accuracy score
  6261. 3:43:50defines that and hence we got a good
  6262. 3:43:52accuracy.
  6263. 3:43:54So, now moving ahead, let us see the
  6264. 3:43:55second project, that is SUV data
  6265. 3:43:57analysis.
  6266. 3:43:58So, in this a car company has released
  6267. 3:44:00new SUV in the market. And using the
  6268. 3:44:03previous data about the sales of their
  6269. 3:44:05SUV, they want to predict the category
  6270. 3:44:07of people who might be interested in
  6271. 3:44:08buying this. So, using the logistic
  6272. 3:44:10regression, you need to find what
  6273. 3:44:12factors make people more interested in
  6274. 3:44:14buying this SUV.
  6275. 3:44:15So, for this let us see our data set
  6276. 3:44:17where I have user ID, I have gender as
  6277. 3:44:19male and female.
  6278. 3:44:21Then we have the age, we have the
  6279. 3:44:22estimated salary.
  6280. 3:44:24And then we have the purchased column.
  6281. 3:44:26So this is my discrete column, or you
  6282. 3:44:28can say the categorical column. So here
  6283. 3:44:30we just have the value that is zero and
  6284. 3:44:31one. And this column we need to predict
  6285. 3:44:33whether a person can actually purchase a
  6286. 3:44:35SUV or not. So based on these factors,
  6287. 3:44:38we will be deciding whether a person can
  6288. 3:44:40actually purchase a SUV or not.
  6289. 3:44:42So we know the salary of a person, we
  6290. 3:44:43know the age. And using these, we can
  6291. 3:44:46predict whether a person can actually
  6292. 3:44:47purchase a SUV or not. So let me just go
  6293. 3:44:49to my Jupiter notebook and let's
  6294. 3:44:51implement logistic regression. So guys,
  6295. 3:44:53I will not be going through all the
  6296. 3:44:54details of data cleaning and analyzing
  6297. 3:44:56the part. So that part, I'll just leave
  6298. 3:44:58it on you. So just go ahead and practice
  6299. 3:45:00as much as you can.
  6300. 3:45:02All right. So my second project is SUV
  6301. 3:45:04predictions.
  6302. 3:45:07All right. So first of all, I have to
  6303. 3:45:08import all the libraries. So I say
  6304. 3:45:10import NumPy as NP.
  6305. 3:45:13And similarly, I'll do the rest of it.
  6306. 3:45:21All right.
  6307. 3:45:22So now let me just print the head of
  6308. 3:45:24this data set.
  6309. 3:45:26So this we have already seen that we
  6310. 3:45:27have columns as user ID, we have gender,
  6311. 3:45:30we have the age, we have the salary, and
  6312. 3:45:32then we have to calculate whether a
  6313. 3:45:33person can actually purchase a SUV or
  6314. 3:45:35not.
  6315. 3:45:36So now let us just simply go on to the
  6316. 3:45:38algorithm part. So we'll directly start
  6317. 3:45:40off with the logistic regression and how
  6318. 3:45:42you can train a model. So for doing all
  6319. 3:45:44those things, we first need to define
  6320. 3:45:46your independent variable and dependent
  6321. 3:45:47variable. So in this case, I want my X,
  6322. 3:45:50that is our independent variable. I say
  6323. 3:45:52dataset.iloc.
  6324. 3:45:54So here I'll be specifying all the rows.
  6325. 3:45:56So colon basically stands for that. And
  6326. 3:45:58in the columns, I want only two and
  6327. 3:46:01three. dot values. So here it should
  6328. 3:46:04fetch me all the rows and only the
  6329. 3:46:05second and third column, which is age
  6330. 3:46:07and estimated salary. So these are the
  6331. 3:46:09factors which will be used to predict
  6332. 3:46:11the dependent variable that is purchase.
  6333. 3:46:13So, here my dependent variable is
  6334. 3:46:15purchase.
  6335. 3:46:16And the dependent variable is of age and
  6336. 3:46:17salary. So, I'll say
  6337. 3:46:20dataset.iloc
  6338. 3:46:21I'll have all the rows and I just want
  6339. 3:46:24fourth column that is my purchase
  6340. 3:46:25column. dot values. All right, so I just
  6341. 3:46:28forgot one
  6342. 3:46:29one square bracket over here. All right.
  6343. 3:46:31So, over here I have defined my
  6344. 3:46:33independent variable and dependent
  6345. 3:46:34variable. So, here my independent
  6346. 3:46:37variable is age and salary and dependent
  6347. 3:46:39variable is the column purchase.
  6348. 3:46:41Now, you must be wondering what is this
  6349. 3:46:42iloc function. So, iloc function is
  6350. 3:46:44basically an indexer for pandas data
  6351. 3:46:46frame and it is used for integer-based
  6352. 3:46:48indexing or you can also say selection
  6353. 3:46:50by index.
  6354. 3:46:52Now, let me just print these independent
  6355. 3:46:53variables and dependent variables. So,
  6356. 3:46:56if I print the independent variable, I
  6357. 3:46:57have the age as well as the salary.
  6358. 3:47:00Next, let me print the dependent
  6359. 3:47:01variable as well. So, over here you can
  6360. 3:47:03see I just have the values in zero and
  6361. 3:47:05one. So, zero stands for did not
  6362. 3:47:07purchase. Next, let me just divide my
  6363. 3:47:10dataset into training and test subset.
  6364. 3:47:12So, I'll simply write in from
  6365. 3:47:13sklearn.cross_split
  6366. 3:47:15dot cross_validation
  6367. 3:47:18import train_test.
  6368. 3:47:20Next, I'll just press shift and tab. And
  6369. 3:47:23over here I'll go to the examples and
  6370. 3:47:25just copy the same line.
  6371. 3:47:27So, I'll just copy this.
  6372. 3:47:30I'll remove the points.
  6373. 3:47:31Now, I want the test size to be let's
  6374. 3:47:33say 25. So, I have divided the train and
  6375. 3:47:35test split in 75 25 ratio.
  6376. 3:47:38Now, let's say I'll take the random
  6377. 3:47:40state as zero. So, random state
  6378. 3:47:42basically ensures the same result or you
  6379. 3:47:44can say the same samples taken whenever
  6380. 3:47:46you run the code. So, let me just run
  6381. 3:47:48this. Now, you can also scale your input
  6382. 3:47:50values for better performing and this
  6383. 3:47:52can be done using standard scalar. So,
  6384. 3:47:54let me do that as well. So, I'll say
  6385. 3:47:56from sklearn.preprocessing
  6386. 3:48:00import standard scalar.
  6387. 3:48:02Now, why do we scale it? Now, if you see
  6388. 3:48:04our dataset we are dealing with large
  6389. 3:48:07numbers.
  6390. 3:48:08Well, although we are using a very small
  6391. 3:48:10data set, so whenever you're working in
  6392. 3:48:12a broad environment, you'll be working
  6393. 3:48:13with large data set where you'll be
  6394. 3:48:15using thousands and hundred thousands of
  6395. 3:48:17tuples. So, there scaling down will
  6396. 3:48:19definitely affect the performance by a
  6397. 3:48:20large extent. So, here let me just show
  6398. 3:48:22you how you can scale down these input
  6399. 3:48:24values. And then the pre-processing
  6400. 3:48:26contains all your methods and
  6401. 3:48:27functionality which is required to
  6402. 3:48:29transform your data.
  6403. 3:48:30So, now let us scale down for test as
  6404. 3:48:32well as our training data set. So, I'll
  6405. 3:48:34first make an instance of it. So, I'll
  6406. 3:48:36say
  6407. 3:48:37standard scalar.
  6408. 3:48:39Then I'll have X_train. I'll say sc.fit
  6409. 3:48:43fit_transform.
  6410. 3:48:46I'll pass in my X_train variable.
  6411. 3:48:52And similarly, I can do it for test
  6412. 3:48:53wherein I'll pass the X_test.
  6413. 3:48:57All right.
  6414. 3:48:58Now, my next step is to import logistic
  6415. 3:49:00regression. So, I'll simply apply
  6416. 3:49:01logistic regression by first importing
  6417. 3:49:03it. So, I'll say from sklearn
  6418. 3:49:06from sklearn.linear_model
  6419. 3:49:09import logistic regression.
  6420. 3:49:12Now, over here I'll be using classifier.
  6421. 3:49:14So, I'll say classifier. is equals to
  6422. 3:49:16logistic regression.
  6423. 3:49:18So, over here I'll just make an instance
  6424. 3:49:19of it. So, I'll say logistic regression.
  6425. 3:49:22And over here I'll just pass in the
  6426. 3:49:23random state which is zero.
  6427. 3:49:26And now I'll simply fit the model.
  6428. 3:49:31And I'll simply pass in X_train and
  6429. 3:49:32Y_train.
  6430. 3:49:34So, here it tells me all the details of
  6431. 3:49:36logistic regression.
  6432. 3:49:39Then I have to predict the values. So,
  6433. 3:49:41I'll say Y_pred
  6434. 3:49:42is equals to classifier
  6435. 3:49:45then predict function
  6436. 3:49:47and then I'll just pass in X_test.
  6437. 3:49:50So, now we have created the model, we
  6438. 3:49:52have scaled down our input values, then
  6439. 3:49:53we have applied logistic regression, we
  6440. 3:49:55have predicted the values, and now we
  6441. 3:49:57want to know the accuracy. So, to know
  6442. 3:49:59the accuracy, first we need to import
  6443. 3:50:01accuracy score. So, I'll say from
  6444. 3:50:03sklearn.metrics
  6445. 3:50:06import accuracy score.
  6446. 3:50:08And using this function we can calculate
  6447. 3:50:09the accuracy. Or you can manually do
  6448. 3:50:12that by creating a confusion matrix.
  6449. 3:50:14So, I'll just pass in my Y_test and my
  6450. 3:50:16Y_predicted.
  6451. 3:50:19All right. So, over here I get the
  6452. 3:50:20accuracy as 89%. So, if we want to know
  6453. 3:50:23the accuracy in percentage, so I just
  6454. 3:50:24have to multiply it by 100. And if I run
  6455. 3:50:26this,
  6456. 3:50:27so it gives me 89%. So, I hope you guys
  6457. 3:50:30are clear with whatever I have taught
  6458. 3:50:31you today. So, here I have taken my
  6459. 3:50:33independent variables as age and salary.
  6460. 3:50:35And then we have calculated that how
  6461. 3:50:37many people can purchase the SUV. And
  6462. 3:50:40then we have calculated our model by
  6463. 3:50:42checking the accuracy. So, over here we
  6464. 3:50:44get the accuracy as 89, which is great.
  6465. 3:50:51Let's compare the two models.
  6466. 3:50:54So, first of all, let's look at the
  6467. 3:50:55definition of linear regression and
  6468. 3:50:57logistic regression. So, the main aim of
  6469. 3:51:00linear regression is to predict a
  6470. 3:51:01continuous dependent variable based on
  6471. 3:51:04the values of the independent variables.
  6472. 3:51:07But when it comes to logistic
  6473. 3:51:08regression, the aim is to predict a
  6474. 3:51:10categorical dependent variable based on
  6475. 3:51:13the values of independent variables.
  6476. 3:51:15These are the main aim of each of these
  6477. 3:51:17models. Now, let's look at the variable
  6478. 3:51:20type. Now, in linear regression, the
  6479. 3:51:22dependent variable is always continuous.
  6480. 3:51:25All right. This is very important to
  6481. 3:51:26remember, because this is the main
  6482. 3:51:28objective of linear regression. Okay, it
  6483. 3:51:30makes use of continuous dependent
  6484. 3:51:32variables to predict continuous values.
  6485. 3:51:34Similarly, when it comes to logistic
  6486. 3:51:36regression, you're going to use
  6487. 3:51:38categorical dependent variable to
  6488. 3:51:40predict a categorical value. All right.
  6489. 3:51:43Now, let's look at the estimation
  6490. 3:51:44method. So guys, linear regression is
  6491. 3:51:47based on the least square estimation,
  6492. 3:51:49which basically says that the regression
  6493. 3:51:51coefficients should be chosen in such a
  6494. 3:51:54way that it minimizes the sum of the
  6495. 3:51:56squared distance of each observed
  6496. 3:51:58response. Okay? Now, this is in-depth
  6497. 3:52:01about linear regression, so that's why
  6498. 3:52:02I'm going to leave a link in the
  6499. 3:52:03description. Now, logistic regression on
  6500. 3:52:06the other hand is based on maximum
  6501. 3:52:08likelihood estimation. Okay? This
  6502. 3:52:10basically says that the coefficients
  6503. 3:52:12should be chosen in such a way that it
  6504. 3:52:15maximizes the probability of Y given
  6505. 3:52:18some value of X. Next is the equation.
  6506. 3:52:21Earlier, we discussed this equation
  6507. 3:52:22where for linear regression, we have Y
  6508. 3:52:24is equal to B not plus B1 into X plus E.
  6509. 3:52:28Okay? Similarly, this is the equation
  6510. 3:52:30for logistic regression.
  6511. 3:52:32So, the next difference is a best fit
  6512. 3:52:33line. So, guys, linear regression aims
  6513. 3:52:36at finding the best fitting straight
  6514. 3:52:38line, which is also called the
  6515. 3:52:40regression line. All right? But, when it
  6516. 3:52:42comes to logistic uh regression, if you
  6517. 3:52:44try and uh map the relationship between
  6518. 3:52:46the dependent and independent variable,
  6519. 3:52:48you're going to get a curve, which is
  6520. 3:52:50also known as the sigmoid curve. All
  6521. 3:52:52right? So, in linear regression, the
  6522. 3:52:53relationship between the dependent and
  6523. 3:52:55independent variable is represented
  6524. 3:52:58using a straight line, but when it comes
  6525. 3:52:59to logistic regression, the relationship
  6526. 3:53:01between the dependent and independent
  6527. 3:53:03variable is represented using a sigmoid
  6528. 3:53:06curve. Now, let's look at the
  6529. 3:53:07relationship between dependent and
  6530. 3:53:09independent variable. Now, when it comes
  6531. 3:53:11to linear regression, there has to be a
  6532. 3:53:13linear relationship between the two. So,
  6533. 3:53:16when I say linear, I mean that the
  6534. 3:53:17variables have to vary linearly. Okay?
  6535. 3:53:20That's how the straight line is formed
  6536. 3:53:22in the first place. All right? When it
  6537. 3:53:24comes to logistic regression, it's not
  6538. 3:53:26necessary to have a linear relationship.
  6539. 3:53:29Now, the output of linear regression is
  6540. 3:53:31always going to be a predicted integer
  6541. 3:53:33value or basically a continuous value.
  6542. 3:53:36So, when it comes to logistic
  6543. 3:53:37regression, the output has to be a
  6544. 3:53:39binary value. Okay? So, it should either
  6545. 3:53:41be class A or class B or it should be
  6546. 3:53:43zero or one, something like that.
  6547. 3:53:45Finally, we have applications. Now,
  6548. 3:53:47linear regression is mainly used to
  6549. 3:53:49predict outcomes like the expected
  6550. 3:53:52number of sales. And you know, it's
  6551. 3:53:54always used to predict some continuous
  6552. 3:53:55value. All right, but when it comes to
  6553. 3:53:57logistic regression, it's mainly used in
  6554. 3:53:59classification. So, when you want to
  6555. 3:54:01classify a data set into two different
  6556. 3:54:03classes, then you use logistic
  6557. 3:54:05regression. You can find a lot of
  6558. 3:54:07applications of linear regression in the
  6559. 3:54:09business domain. And logistic regression
  6560. 3:54:12is mainly used in the cybersecurity,
  6561. 3:54:14image processing, and classification
  6562. 3:54:16domain.
  6563. 3:54:17>> [music]
  6564. 3:54:22>> What is classification?
  6565. 3:54:25And if I have to tell you about
  6566. 3:54:26classification,
  6567. 3:54:29like for example, what happens is like
  6568. 3:54:32we have two type of when we talk about
  6569. 3:54:34machine learning, machine learning is
  6570. 3:54:36nothing but you know, like a series of
  6571. 3:54:38instructions you give it to the computer
  6572. 3:54:41so that it can learn the patterns from
  6573. 3:54:42your data set, right? To give you an
  6574. 3:54:44example, imagine that there is a
  6575. 3:54:46trending topic, for example, you found
  6576. 3:54:48it you want to find it out whether
  6577. 3:54:51Prime Minister Modi will be the second
  6578. 3:54:53Prime Minister once again the Prime
  6579. 3:54:55Minister for the country or not, okay?
  6580. 3:54:57So, now what you will do is you will
  6581. 3:54:59collect the data set from multiple
  6582. 3:55:00different sources. And you will you will
  6583. 3:55:04actually build it a like a algorithm
  6584. 3:55:06where you you will get a label as yes or
  6585. 3:55:09no. Yes, he will continue as a next
  6586. 3:55:11Prime Minister or no, he will not
  6587. 3:55:12continue as a next Prime Minister. So,
  6588. 3:55:14you will collect the data set and you
  6589. 3:55:16will feed this data set to the computer
  6590. 3:55:17and this process is called as a machine
  6591. 3:55:20learning, right? So, now in this case
  6592. 3:55:22what happens is um
  6593. 3:55:24So, this this is about the
  6594. 3:55:26classification, right? So, now machine
  6595. 3:55:29learning is basically of two type. One
  6596. 3:55:31is called as a supervised machine
  6597. 3:55:33learning, another is called as a
  6598. 3:55:34unsupervised machine learning, and the
  6599. 3:55:36third one is a reinforcement machine
  6600. 3:55:38learning. So, when we speak about
  6601. 3:55:40supervised machine learning, as the name
  6602. 3:55:42suggests, it provides some supervision,
  6603. 3:55:44right? For example, the teacher teaching
  6604. 3:55:46the kid is a supervised machine
  6605. 3:55:48learning. Right? So, we will give the
  6606. 3:55:50trained examples. We will give the
  6607. 3:55:52trained data set with the pure label on
  6608. 3:55:54top of that. Uh this is called as a
  6609. 3:55:57supervised machine learning. So, if I
  6610. 3:55:58draw in front of you, this type of
  6611. 3:56:00machine learning look like this,
  6612. 3:56:02supervised machine learning. Where what
  6613. 3:56:04happens is like you would have the data
  6614. 3:56:06set, which is a structured data set, and
  6615. 3:56:08you would have one column, which is
  6616. 3:56:10called as a label. What you want to
  6617. 3:56:12predict, okay? And you would have a uh
  6618. 3:56:15various predictors by which you want to
  6619. 3:56:17predict. To give you an example, imagine
  6620. 3:56:19that you want to predict the pricing of
  6621. 3:56:21a community. Okay? You want to predict
  6622. 3:56:23what would be the pricing of apartment
  6623. 3:56:25in a particular community, right? Now,
  6624. 3:56:27these can be the variable like you can
  6625. 3:56:29see that uh what would be the number of
  6626. 3:56:31how many floors it has. You can have a
  6627. 3:56:33variable like what is a pollution level,
  6628. 3:56:35how how many educational institution are
  6629. 3:56:37nearby, right? So, based upon that, the
  6630. 3:56:39pricing will change. But this type of
  6631. 3:56:42supervised machine learning, why it is
  6632. 3:56:43called supervised machine learning
  6633. 3:56:45because we provide the independent
  6634. 3:56:47variable or we provide the predictors.
  6635. 3:56:50Also, we provide the label data set.
  6636. 3:56:52Okay? Now, the supervised machine
  6637. 3:56:54learning is basically of two type. This
  6638. 3:56:57supervised machine learning, one type
  6639. 3:56:59one first is called as the, you know,
  6640. 3:57:02regression based supervised machine
  6641. 3:57:03learning.
  6642. 3:57:05Okay? And the second one is called as a
  6643. 3:57:07classification based supervised machine
  6644. 3:57:09learning.
  6645. 3:57:10Now, what is a difference between
  6646. 3:57:11regression based and a classification
  6647. 3:57:13based supervised machine learning?
  6648. 3:57:15Regression based supervised machine
  6649. 3:57:17learning is that machine learning where
  6650. 3:57:19what you want to predict
  6651. 3:57:22is continuous in nature. Okay? Imagine
  6652. 3:57:25that you want to predict the
  6653. 3:57:27community prices, right? Which is a
  6654. 3:57:29continuous value. If it is a continuous
  6655. 3:57:32value, then we will go ahead with
  6656. 3:57:33regression based supervised machine
  6657. 3:57:35learning. Okay? Whereas, if you want to
  6658. 3:57:38predict something which is the discrete
  6659. 3:57:40outcome, to give you an example, you
  6660. 3:57:42want to predict that whether I will win
  6661. 3:57:44the match or not.
  6662. 3:57:46Okay? I want to predict whether the
  6663. 3:57:48particular employee will turn out from
  6664. 3:57:50the company or not. You want to predict,
  6665. 3:57:53you know, like whether the person will
  6666. 3:57:55have a cancer or not.
  6667. 3:57:57You're getting my point, right? So, if
  6668. 3:57:59you have the output which you want to
  6669. 3:58:01predict is in the form of yes or no or
  6670. 3:58:04true or false, right? This is called as
  6671. 3:58:07the supervised machine learning, but a
  6672. 3:58:10classification-based supervised machine
  6673. 3:58:12learning. Okay? So, classification-based
  6674. 3:58:15supervised machine learning is the
  6675. 3:58:16process of dividing the data set into
  6676. 3:58:19different categories or group by adding
  6677. 3:58:21a label. Okay, so always remember that
  6678. 3:58:24whenever you guys want to predict the
  6679. 3:58:27classes in the data set, whenever you
  6680. 3:58:29want to predict, you know, like whether
  6681. 3:58:31this will happen or not, whether the
  6682. 3:58:33whether a person will do a credit card
  6683. 3:58:36fraud or not. You're getting my point,
  6684. 3:58:38right? Whether the employee will turn
  6685. 3:58:39out from the company or not, right?
  6686. 3:58:42Whether the particular person will have
  6687. 3:58:43a diabetes as a disease or not. All
  6688. 3:58:45these questions, wherever you want to
  6689. 3:58:47find it out yes or no or true or false
  6690. 3:58:50or you want to predict classes in the
  6691. 3:58:51data set, this is called as a
  6692. 3:58:53classification-based supervised machine
  6693. 3:58:56learning. Okay? So, this is what is
  6694. 3:58:58called as a classification-based
  6695. 3:59:00supervised machine learning and
  6696. 3:59:02today I will teach you, you know, I will
  6697. 3:59:05tell you about various form of
  6698. 3:59:06classification-based supervised machine
  6699. 3:59:08learning. Although we will do a deep
  6700. 3:59:10dive into decision tree. Right? Now, you
  6701. 3:59:13will be able to understand that decision
  6702. 3:59:15tree how decision tree is connected to
  6703. 3:59:17the network what we are learning today
  6704. 3:59:19with classification-based supervised
  6705. 3:59:22machine learning. Okay? So, now
  6706. 3:59:25we have various algorithms. Algorithm is
  6707. 3:59:28nothing but a set of mathematical
  6708. 3:59:30equations for classification-based
  6709. 3:59:33supervised machine learning. We first of
  6710. 3:59:36all, we have something called as
  6711. 3:59:37decision tree. Then we would have
  6712. 3:59:39something called as random forest, then
  6713. 3:59:41naive base, and then KNN, which is
  6714. 3:59:43called as a K nearest neighbors. Okay?
  6715. 3:59:46So, let me give you few statements about
  6716. 3:59:48this algorithm, and we have many others,
  6717. 3:59:50but today we will focus on one of them,
  6718. 3:59:52which is decision tree.
  6719. 3:59:54So, now first let's start with decision
  6720. 3:59:56tree. Now, what is a decision tree? What
  6721. 3:59:58I was telling you is, believe me or not,
  6722. 4:00:00decision tree is something you use every
  6723. 4:00:02day in your um in your daily life. For
  6724. 4:00:05example, you take decisions, and for ex-
  6725. 4:00:08today also you took a decision to attend
  6726. 4:00:10this webinar, right? But, how do you
  6727. 4:00:13decide a decision based on various
  6728. 4:00:15further decisions, right? For example,
  6729. 4:00:18for today joining the webinar, you have
  6730. 4:00:20seen that, okay,
  6731. 4:00:21uh when this webinar is about, okay? So,
  6732. 4:00:24you said it is weekday or weekend. Then
  6733. 4:00:26you might have said check it out, right?
  6734. 4:00:28What is a time of the webinar? Then you
  6735. 4:00:30might have checked it out, what is a
  6736. 4:00:31topic of the webinar, right? So, and who
  6737. 4:00:34is conducting this webinar? So, based
  6738. 4:00:36upon this, you took a decision, shall I
  6739. 4:00:38go ahead or not go ahead? You're getting
  6740. 4:00:40my point, right? So, this is what is
  6741. 4:00:42called as a decision tree. We call it as
  6742. 4:00:44decision tree because it is a graphical
  6743. 4:00:46representation of all the possible
  6744. 4:00:48solution to a decision. It's like a
  6745. 4:00:49tree. Like, decision tree is like a
  6746. 4:00:52tree. Why? Because a tree also start
  6747. 4:00:55with a root, and then it emerge into
  6748. 4:00:57various branches. Similarly, you have a
  6749. 4:00:59decision tree, which I'm showing to you
  6750. 4:01:01here as a simplest algorithm, which is
  6751. 4:01:03being used for machine learning
  6752. 4:01:05purposes.
  6753. 4:01:06Now, it is something like this. Imagine
  6754. 4:01:08that you want to find it out that, you
  6755. 4:01:11know, like, do you want to go to a
  6756. 4:01:13restaurant, or do you want to buy an
  6757. 4:01:15hamburger? Okay? So, you have two
  6758. 4:01:17choices. Either you can go for a
  6759. 4:01:19restaurant, or either you can buy a
  6760. 4:01:21hamburger. Now, how would you decide
  6761. 4:01:23which one you would follow? So, you will
  6762. 4:01:25start with what is called as a root
  6763. 4:01:27node. You will start with the root node
  6764. 4:01:28that whether I'm hungry or not, right?
  6765. 4:01:32If I am hungry, right? Then only I will
  6766. 4:01:35go for all these activity. If I am not
  6767. 4:01:38hungry at all, then simply go and sleep.
  6768. 4:01:41You got my point, right? So, this is how
  6769. 4:01:43you will start with the first node,
  6770. 4:01:45which is to find it out whether I'm
  6771. 4:01:47feeling hungry or not. If I'm feeling
  6772. 4:01:48hungry, then I will decide that, "Well,
  6773. 4:01:51I Do I have money, which is around $25
  6774. 4:01:54worth?" If I have money, then I will go
  6775. 4:01:56for a restaurant. If I don't have a
  6776. 4:01:58money, then I will buy a hamburger.
  6777. 4:02:00You understood this? This is simplest
  6778. 4:02:03representation of decision tree.
  6779. 4:02:05Basically, you decide something on the
  6780. 4:02:07basis of the previous outcomes. And you
  6781. 4:02:10can imagine any sort of example here.
  6782. 4:02:12Imagine that you want to find it out
  6783. 4:02:14whether the person will do a credit card
  6784. 4:02:16fraud or not. Again, it will depend upon
  6785. 4:02:18the previous circumstances that, you
  6786. 4:02:20know, for example, it will depend upon
  6787. 4:02:23how much is the salary of the person.
  6788. 4:02:24You It will depend upon what is a job
  6789. 4:02:26profile of the person. It will depend
  6790. 4:02:28upon the fact that, you know, like, for
  6791. 4:02:29example, how many fraud or how many card
  6792. 4:02:32this person have, right? So, on [snorts]
  6793. 4:02:34basis of you will decide that whether
  6794. 4:02:35this will do a credit card fraud or not.
  6795. 4:02:39So, this is the simplest form of
  6796. 4:02:41classification-based algorithm.
  6797. 4:02:43Then we have the next algorithm, which
  6798. 4:02:45is called as a random forest. Now, what
  6799. 4:02:48is a random forest? As the name
  6800. 4:02:50suggests, you built one decision tree in
  6801. 4:02:52the last example, right? But
  6802. 4:02:55one decision tree can sometime be
  6803. 4:02:57over-fitting as a case, right? So,
  6804. 4:02:59people said that, you know, like, "Why
  6805. 4:03:01should I only trust one decision tree?"
  6806. 4:03:03For example, whenever you take a bold
  6807. 4:03:05decision in your life, you don't trust
  6808. 4:03:08only a single voice, right? You want to
  6809. 4:03:10hear [clears throat] it from multiple
  6810. 4:03:12different people to make your decision
  6811. 4:03:14more stronger, right? Go to a doctor. If
  6812. 4:03:16the doctor says to you that you have
  6813. 4:03:18this type of a disease, you don't
  6814. 4:03:21believe in that. What you do is you also
  6815. 4:03:23talk to second doctor, third doctor,
  6816. 4:03:25fourth doctor to confirm that this is
  6817. 4:03:27true or not. Okay? So, this is what is
  6818. 4:03:30called as a random forest. As the name
  6819. 4:03:32suggests, why it is called random
  6820. 4:03:33forest? It is called random forest
  6821. 4:03:36because now you're building the various
  6822. 4:03:38number of decision trees here. It's like
  6823. 4:03:40a forest of all the trees, right? It is
  6824. 4:03:43no longer a single tree, it is a forest
  6825. 4:03:45of complete trees.
  6826. 4:03:47So, for example, you can imagine that if
  6827. 4:03:49it is my training data set, I will I
  6828. 4:03:51will split my training data set into
  6829. 4:03:53multiple examples, multiple decision
  6830. 4:03:55trees will be built, and then based upon
  6831. 4:03:58the majority, I will decide whether
  6832. 4:04:00should I do this or not. Okay? So, it is
  6833. 4:04:03also called as a bagging sort of
  6834. 4:04:04methodology where we bring the outcome
  6835. 4:04:07of various models, or we bring the
  6836. 4:04:09outcome of various trees all together to
  6837. 4:04:12make the
  6838. 4:04:13uh to make a powerful decision. Okay?
  6839. 4:04:16So, this is another example This is
  6840. 4:04:18another algorithm which is called as a
  6841. 4:04:20random forest.
  6842. 4:04:23Then, after random forest, we have
  6843. 4:04:25something called as naive base. Okay?
  6844. 4:04:28Now, what is a naive base algorithm?
  6845. 4:04:31Naive base is a also a simplest
  6846. 4:04:33algorithm, but naive base is basically
  6847. 4:04:36based on the base theorem. Okay? So,
  6848. 4:04:38this algorithm is basically based on the
  6849. 4:04:40base theorem, and base theorem is based
  6850. 4:04:42on the conditional probability.
  6851. 4:04:45Right? So, it is based on the
  6852. 4:04:47conditional probability, which is your
  6853. 4:04:49naive base. Now, in naive base, what
  6854. 4:04:52happens is like we decide that whether
  6855. 4:04:56something will happen or not on the
  6856. 4:04:58basis of probability. So, let me
  6857. 4:05:01illustrate you with this example.
  6858. 4:05:03Imagine that you want to find it out
  6859. 4:05:06whether I would have a disease or not.
  6860. 4:05:09Okay? So, first of all, the probability
  6861. 4:05:12of having a disease is 0.10. And the
  6862. 4:05:16probability of not having a disease is
  6863. 4:05:180.90. Okay? So, if there is a
  6864. 4:05:21probability of having a disease is 0.10,
  6865. 4:05:24you will further find it out that what
  6866. 4:05:26is a probability that my test will be
  6867. 4:05:30positive given I'm diseased.
  6868. 4:05:33What is the probability that my test to
  6869. 4:05:35diagnose the disease is negative given I
  6870. 4:05:37have a disease? Right? Similarly, if you
  6871. 4:05:40go into this direction, if there is a
  6872. 4:05:42probability of not having a disease is
  6873. 4:05:440.90, you will find it out that what is
  6874. 4:05:47a probability of [snorts] having a
  6875. 4:05:49disease or test being positive with
  6876. 4:05:51having with no disease. And what is a
  6877. 4:05:54probability of no disease given I don't
  6878. 4:05:57have the disease, which is 0.90. So,
  6879. 4:05:59basically, here you check the outcomes.
  6880. 4:06:03Here you check the outcome of all the
  6881. 4:06:06all the possible combinations. This is
  6882. 4:06:08what is called as the naive base
  6883. 4:06:10algorithm or base theorem or
  6884. 4:06:13you know, conditional probability base
  6885. 4:06:15theorem. Okay?
  6886. 4:06:17Then we also have something which is
  6887. 4:06:19called as a K nearest neighbor, which is
  6888. 4:06:22the one of the finest algorithm
  6889. 4:06:25which also helps you in deciding the
  6890. 4:06:28classification, right? Now, what is K
  6891. 4:06:31nearest neighbor? K nearest neighbor, as
  6892. 4:06:33the name suggests, what happens in this
  6893. 4:06:35case is we try to build, you know, we
  6894. 4:06:38try to build
  6895. 4:06:40basically, it's like a neighbor. So,
  6896. 4:06:42it's something like this. Like if I give
  6897. 4:06:44you the data set, say I give you the
  6898. 4:06:46customer data set. Okay?
  6899. 4:06:49And I have It is a transaction data set.
  6900. 4:06:51So, customer number one
  6901. 4:06:53has bought a product number one, right?
  6902. 4:06:56From a particular vendor at a particular
  6903. 4:06:58rate, and this is a profit this customer
  6904. 4:07:00has given to us. And this is the
  6905. 4:07:02revenues. Okay? This is my class
  6906. 4:07:04customer number one. Similarly, you
  6907. 4:07:07would have customer number two, customer
  6908. 4:07:09number three, customer number four,
  6909. 4:07:11right? Now, if I ask you that which
  6910. 4:07:13customer profiling is same
  6911. 4:07:15is nearby same, right? So, what you can
  6912. 4:07:18do is you can group your customers or
  6913. 4:07:20you will get to know that which customer
  6914. 4:07:23behave in the similar way. So, you can
  6915. 4:07:25say that customer one, customer three,
  6916. 4:07:27and customer four, they behave in a
  6917. 4:07:29similar way because they are giving us
  6918. 4:07:30the high profit and high revenue
  6919. 4:07:32margins.
  6920. 4:07:34You understood? So, this is what happens
  6921. 4:07:37in the case of K nearest neighbor. So, K
  6922. 4:07:39nearest neighbor, what you generally do
  6923. 4:07:41is you find it out that what uh I would
  6924. 4:07:45be able to, you know,
  6925. 4:07:47uh find it out who is my nearest
  6926. 4:07:49neighbor, like what are the similarities
  6927. 4:07:51in the patterns we have, right? This is
  6928. 4:07:53what we use for K nearest neighbors. And
  6929. 4:07:56this you can see that, for example, the
  6930. 4:07:58algorithm based on the distance-based
  6931. 4:08:00mechanism find it out that how many
  6932. 4:08:02people or how what type of audience is
  6933. 4:08:05similar to the outcomes, right?
  6934. 4:08:08So, then moving further here
  6935. 4:08:11the with the second topic, let's go into
  6936. 4:08:13detail about what is decision tree all
  6937. 4:08:15about, right? So, so far I have touched
  6938. 4:08:18base on uh classification-based
  6939. 4:08:20algorithm, right? And I was teaching you
  6940. 4:08:22different type of classification-based
  6941. 4:08:24algorithm, but since the focus of
  6942. 4:08:26today's class is decision tree, let's
  6943. 4:08:28take a deep dive into the decision tree.
  6944. 4:08:31Okay? So, let's get started with
  6945. 4:08:33decision tree.
  6946. 4:08:34A decision tree is a graphical
  6947. 4:08:36representation of all possible solution
  6948. 4:08:39to a decision based on a certain
  6949. 4:08:40condition. What does it mean? It means
  6950. 4:08:43that it is as simple as that. Imagine
  6951. 4:08:45that this is a tree, right? So, it is
  6952. 4:08:47like a tree where
  6953. 4:08:49you have a problem statement that should
  6954. 4:08:51I accept a job offer or not. Imagine
  6955. 4:08:54that you want to find it out that should
  6956. 4:08:56I accept a job offer or not. This is a
  6957. 4:08:58problem statement which is there in your
  6958. 4:09:00mind. Now, how would you solve this
  6959. 4:09:01problem statement with the help of a
  6960. 4:09:03decision tree?
  6961. 4:09:04First of all, you will start with what
  6962. 4:09:06we call as a root node. We will start
  6963. 4:09:08with a root node starting with, you
  6964. 4:09:11know, to find it out what is the salary.
  6965. 4:09:13Okay, what I'm what I what is the salary
  6966. 4:09:15I'm getting. So if my salary is equal to
  6967. 4:09:19or greater than equal to 50,000, I'll go
  6968. 4:09:22here. If my salary is not greater than
  6969. 4:09:25equal to 50,000, I will say I will not
  6970. 4:09:28accept this offer.
  6971. 4:09:29Okay? Now imagine that you say that your
  6972. 4:09:32salary is greater than 50,000, then you
  6973. 4:09:34will check another
  6974. 4:09:36another variable here. You will check it
  6975. 4:09:38out whether I have to commute more than
  6976. 4:09:401 hour. If I have to commute more than 1
  6977. 4:09:43hour, I will decline the offer.
  6978. 4:09:46You're getting my point, right? Then you
  6979. 4:09:48will if you don't have to commute more
  6980. 4:09:50than 1 hour, then you will still
  6981. 4:09:51consider this option and then you will
  6982. 4:09:53check it out. For example, in this case
  6983. 4:09:54we're checking it out whether you're
  6984. 4:09:56getting the free offers also like coffee
  6985. 4:09:59or some snacks or other things like
  6986. 4:10:00that. If yes, then you will accept
  6987. 4:10:03finally the offer, otherwise you will
  6988. 4:10:05decline the offer.
  6989. 4:10:07You got my point. This is how our
  6990. 4:10:09decision tree works. Basically, decision
  6991. 4:10:11tree will keep on splitting, keep on
  6992. 4:10:13splitting unless and until you are able
  6993. 4:10:16to find it out your decision.
  6994. 4:10:18Okay? So here the decision was shall I
  6995. 4:10:21accept this or not, right? So it will
  6996. 4:10:23keep on splitting, keep on splitting
  6997. 4:10:25unless and until you get your decision
  6998. 4:10:27whether you do this or not. Like this
  6999. 4:10:30example which I have to I have explained
  7000. 4:10:32you. Okay?
  7001. 4:10:34Now with this what happens is like let
  7002. 4:10:37me if I go further and explain you
  7003. 4:10:39further on this decision tree, let's
  7004. 4:10:41understand this more importantly, right?
  7005. 4:10:43Let's understand this
  7006. 4:10:45one by one. So imagine that this is my
  7007. 4:10:48data set. Okay? Now in my data set you
  7008. 4:10:51can see that
  7009. 4:10:53I have various [clears throat] colors
  7010. 4:10:54given to you. It's like a green color,
  7011. 4:10:55yellow color, red color, red color and a
  7012. 4:10:58yellow color. I'm saying if this is a if
  7013. 4:11:01there is a green color fruit with a
  7014. 4:11:03diameter of three, right? And I can say
  7015. 4:11:06it is a mango.
  7016. 4:11:08Whereas I'm saying if it is a yellow
  7017. 4:11:10color and a diameter three, still will
  7018. 4:11:12be called as a mango. If it is a red
  7019. 4:11:14color with one diameter, then it is a
  7020. 4:11:17grape.
  7021. 4:11:18And it can be a red color with a
  7022. 4:11:20diameter of one still can be a grape,
  7023. 4:11:22but even a yellow with a diameter of
  7024. 4:11:25three can be lemon. Now, what is
  7025. 4:11:27happening in this case is if you have
  7026. 4:11:29this type of a data set, where you want
  7027. 4:11:31to predict the label of the fruit,
  7028. 4:11:33right? You want to predict whether a
  7029. 4:11:35fruit will be mango, whether a fruit
  7030. 4:11:37will be lemon, right? Or whether a fruit
  7031. 4:11:40in this case is mango and lemon or
  7032. 4:11:42grape.
  7033. 4:11:43I have to build a classifier, which is a
  7034. 4:11:45decision tree classifier on the basis of
  7035. 4:11:47this data set. Imagine this is a problem
  7036. 4:11:50statement given to you. Now, first of
  7037. 4:11:52all, what you will do here is you will
  7038. 4:11:54take this data set, right? You will take
  7039. 4:11:56this data set and you will start with a
  7040. 4:11:59root node. Root node imagine here is
  7041. 4:12:01like, is my diameter of the fruit
  7042. 4:12:03greater than equal to three or not?
  7043. 4:12:06Okay? Is my diameter of the fruit
  7044. 4:12:08greater than equal to three or not? If
  7045. 4:12:11it is greater than equal to three,
  7046. 4:12:13right? Now, if it is greater than equal
  7047. 4:12:16to three, you can see that you have
  7048. 4:12:18three fruits here, three rows of the
  7049. 4:12:19data set here. One is green with three,
  7050. 4:12:22mango. Yellow with three, lemon. Yellow
  7051. 4:12:25with three, mango. And wherever it fails
  7052. 4:12:28in the condition, you are left with
  7053. 4:12:30where the diameter is not greater than
  7054. 4:12:32equal to three, it is less than equal to
  7055. 4:12:34three, then what which are the data set
  7056. 4:12:36we have? Red one, grape and red one,
  7057. 4:12:39grape.
  7058. 4:12:41Okay? So, I hope you are understanding
  7059. 4:12:43this, right? How we have starting with
  7060. 4:12:45the decision tree with the base on one
  7061. 4:12:47condition here, right? Which is
  7062. 4:12:49diameter. Now, based upon this, what I
  7063. 4:12:52will do is I have to split it further.
  7064. 4:12:54Here in this case, I don't have to split
  7065. 4:12:56it because I already got the result. So,
  7066. 4:12:58if my diameter is greater than equal to
  7067. 4:13:00three in my data set, if the diameter is
  7068. 4:13:03not equal to greater than equal to
  7069. 4:13:05three, I know it is a grape. Right?
  7070. 4:13:08Whereas, if the diameter is greater than
  7071. 4:13:10equal to three, then it can be a mango
  7072. 4:13:13or it can be a lemon. I'm not sure on my
  7073. 4:13:15decision. So, what I will do here is I
  7074. 4:13:17will split it further. So, I have to
  7075. 4:13:20split this further.
  7076. 4:13:22Here, the splitting is not required.
  7077. 4:13:24But, if I split this further, I may
  7078. 4:13:26check it out now. Is the color equal to
  7079. 4:13:29yellow or not? Now, if the color is
  7080. 4:13:31equal to yellow, then I have two fruit
  7081. 4:13:34which is
  7082. 4:13:35your this row, which is, you know, um
  7083. 4:13:38your mango or lemon. And if the color is
  7084. 4:13:42not equal to yellow, then you will have
  7085. 4:13:44the another row of the data set, which
  7086. 4:13:46is, you know, which you're left with,
  7087. 4:13:48right? So, this is how you have done
  7088. 4:13:50this work
  7089. 4:13:52in the case of your decision tree.
  7090. 4:13:55Okay? And how you will find it out at
  7091. 4:13:57which type of the criteria or which type
  7092. 4:14:00of the algorithm I should or which type
  7093. 4:14:02of criteria or variable I should choose
  7094. 4:14:04here to split the tree is on the basis
  7095. 4:14:06of Gini index and the information gain,
  7096. 4:14:09which I will illustrate and show you in
  7097. 4:14:11the few slides from now. Okay? So, I
  7098. 4:14:14hope with this you understand the how
  7099. 4:14:17your decision tree works. Basically, on
  7100. 4:14:18the basis of the condition, and if you
  7101. 4:14:20get a pure subset, then no need to split
  7102. 4:14:23it further. If there is no pure subset,
  7103. 4:14:25keep on splitting, keep on splitting
  7104. 4:14:27unless and until you get a pure subset.
  7105. 4:14:30Okay?
  7106. 4:14:32So, in this case, what has happened is
  7107. 4:14:34you got 100% mango. Here, you got 50%
  7108. 4:14:37mango and 50% lemon. Okay?
  7109. 4:14:41So, now
  7110. 4:14:43this is what I have told you in the
  7111. 4:14:45previous slide, like how does it work,
  7112. 4:14:48right?
  7113. 4:14:50Now, this all depend upon the Gini index
  7114. 4:14:53basis, right? Now, in the next
  7115. 4:14:55subsequent slide, let me tell you and
  7116. 4:14:58explain you about Gini index and
  7117. 4:15:00information gain. How does it work?
  7118. 4:15:02Right? How does that something works on
  7119. 4:15:05the basis of uh your Gini index?
  7120. 4:15:08So, imagine that, you know, like imagine
  7121. 4:15:11that I have a feature here is the color
  7122. 4:15:14green or not, right? Basis of whether
  7123. 4:15:16the feature whether you have a color
  7124. 4:15:18green or not, what would happen here is
  7125. 4:15:21if the color is green,
  7126. 4:15:22you will get this row here. If it is
  7127. 4:15:24not, you will get this these two rows
  7128. 4:15:26here, right? In the next, it may decide
  7129. 4:15:29on the other basis, right? It can be is
  7130. 4:15:31the diameter which is greater than equal
  7131. 4:15:33to three or not and many other things.
  7132. 4:15:36Okay?
  7133. 4:15:37Now, let's go ahead the example and the
  7134. 4:15:40questions which is coming to your mind
  7135. 4:15:42which is related to, you know, decision
  7136. 4:15:44tree terminologies. I'm pretty sure you
  7137. 4:15:47would have these questions in your mind
  7138. 4:15:49and you would be thinking that how
  7139. 4:15:50should I decided which feature I should
  7140. 4:15:52use and which feature shouldn't I be
  7141. 4:15:54using? This is the question which you
  7142. 4:15:55had. And this is a excellent question
  7143. 4:15:57for understanding the decision tree. But
  7144. 4:16:00before I go on to that, I have to tell
  7145. 4:16:02you about some of the terminologies
  7146. 4:16:05which we commonly use while we build the
  7147. 4:16:07decision tree. So, first of all,
  7148. 4:16:09decision tree looks like this type of a
  7149. 4:16:11tree-based structure, okay? So, every
  7150. 4:16:13decision tree will have its root node.
  7151. 4:16:16So, root node is where the decision tree
  7152. 4:16:18will start. So, it represent the entire
  7153. 4:16:20population or sample and it is further
  7154. 4:16:23get divided or into two or more
  7155. 4:16:25homogeneous sets. So, as you know that
  7156. 4:16:26this will be the first feature on the
  7157. 4:16:28basis of your tree will start. Like tree
  7158. 4:16:31start with a root, your in this case,
  7159. 4:16:33your decision tree will also start with
  7160. 4:16:35a root node.
  7161. 4:16:37Okay? Now, once you have the root, after
  7162. 4:16:39that in a tree, what happens is you will
  7163. 4:16:41start getting the branches, right? I
  7164. 4:16:43hope you understand what is branches. In
  7165. 4:16:45this case, we will keep on splitting,
  7166. 4:16:47keep on splitting unless and until we
  7167. 4:16:49get a decision whether this will happen
  7168. 4:16:50or not. So, it's like a branches.
  7169. 4:16:53Okay? Then, we would also have a parent
  7170. 4:16:55or a child node. So, child node is
  7171. 4:16:57nothing but when you have a branches,
  7172. 4:17:00then the branches can also have the
  7173. 4:17:01outcomes. So, in the previous example,
  7174. 4:17:04like for example, we have whether the
  7175. 4:17:06diameter is greater than equal to three
  7176. 4:17:09or not. You remember? Whether the
  7177. 4:17:10diameter is greater than equal to three
  7178. 4:17:12or not. If it is true, what is
  7179. 4:17:14happening? If it is a false, what is
  7180. 4:17:15happening? So, this is a child or child
  7181. 4:17:18node. Basically, this is a intermediate
  7182. 4:17:20node. This is not the final decision
  7183. 4:17:22which is being made. So, this is called
  7184. 4:17:24as a parent or the child node.
  7185. 4:17:27Then, we also have other terminology
  7186. 4:17:29here, which is splitting, which you know
  7187. 4:17:31that we will keep on splitting unless
  7188. 4:17:33and until you get a desired node. And
  7189. 4:17:35finally, the tree end with a leaf node.
  7190. 4:17:37So, always remember one thing, you will
  7191. 4:17:39start your tree with a root node. Root
  7192. 4:17:42node is a node with which we will start
  7193. 4:17:44the decision tree. And you will end your
  7194. 4:17:46decision as decision tree at a leaf node
  7195. 4:17:49where you will get a decision that you
  7196. 4:17:51should do this or not.
  7197. 4:17:53Okay? And now,
  7198. 4:17:56uh pruning is a activity where
  7199. 4:17:58uh you you will cut down the decision
  7200. 4:18:00tree if it is pruned lot amount of
  7201. 4:18:03times. I will even explain you this. It
  7202. 4:18:05is a case of overfitting. You shouldn't
  7203. 4:18:08build thousand You can shouldn't build
  7204. 4:18:09thousand uh branches of the tree when it
  7205. 4:18:12is not even required. So, I'll I'll
  7206. 4:18:13explain you this point.
  7207. 4:18:17Now, let's move further here and let's
  7208. 4:18:20see this with the help of an example,
  7209. 4:18:23which was your
  7210. 4:18:24uh Gini index and your information gain.
  7211. 4:18:28Right?
  7212. 4:18:29So, this is where you were asking me
  7213. 4:18:32that which question to ask and when,
  7214. 4:18:34right? How would you decide that which
  7215. 4:18:37feature has to be taken first? Right?
  7216. 4:18:39Let's take up an example here and I
  7217. 4:18:42would encourage everyone of you to hear
  7218. 4:18:45me really fine here because this is on
  7219. 4:18:47the basis how would you decide or how
  7220. 4:18:49the algorithm decide to break this
  7221. 4:18:51further. And and I'm explaining you with
  7222. 4:18:53the help of one
  7223. 4:18:55simplest example, and we will do some
  7224. 4:18:57maths over it.
  7225. 4:18:58So, imagine that
  7226. 4:19:00there is a data set where I want to find
  7227. 4:19:03it out whether I will play the match or
  7228. 4:19:06not.
  7229. 4:19:08Okay, there is a cricket match or there
  7230. 4:19:10is a football match or whatever it is. I
  7231. 4:19:12want to find it out whether I will play
  7232. 4:19:14the match or not. Okay? Now,
  7233. 4:19:17how do the decision tree look like?
  7234. 4:19:19Decision tree look like this. If the
  7235. 4:19:21outlook if the outlook is humid
  7236. 4:19:25if the outlook is humid
  7237. 4:19:27and if the outlook is humid and humidity
  7238. 4:19:30is very high, then I will not play the
  7239. 4:19:33match. On the other hand, if the outlook
  7240. 4:19:36is humid and humidity is normal, I may
  7241. 4:19:38play the match.
  7242. 4:19:40Okay?
  7243. 4:19:41If the outlook is
  7244. 4:19:42absolutely clear, then I will always
  7245. 4:19:44play. On the other hand, if outlook is
  7246. 4:19:47windy and the winds are very strong, I
  7247. 4:19:49may not play the match. On the other
  7248. 4:19:51hand, if the outlook is windy and the
  7249. 4:19:54winds are weak, I may play the match.
  7250. 4:19:56So, now what is happening again is that
  7251. 4:19:58you are not sure how you decided with
  7252. 4:20:00outlook, how you decided with these
  7253. 4:20:03nodes here, how you decided with these
  7254. 4:20:05features here. Let me try to show you
  7255. 4:20:07this entire data set. So, this entire
  7256. 4:20:09data set look like this.
  7257. 4:20:11Okay? So, this is basically your 14 days
  7258. 4:20:14data set. And this exactly happens in
  7259. 4:20:17the case of your classification based
  7260. 4:20:18algorithm, you will get a data set like
  7261. 4:20:20this.
  7262. 4:20:21So, you can imagine that you want to
  7263. 4:20:23find it out I would like whether I will
  7264. 4:20:26play or not. Right? This is what you
  7265. 4:20:29want to essentially find it want to find
  7266. 4:20:31it out whether I would play or not. On
  7267. 4:20:33basis of what? On basis of four
  7268. 4:20:35different variables you have. Four
  7269. 4:20:37different variables or features in the
  7270. 4:20:39data set is outlook
  7271. 4:20:41temperature, humidity, and wind. Right?
  7272. 4:20:45So, now
  7273. 4:20:46it can be like if the outlook is sunny,
  7274. 4:20:49temperature is hot, humidity is high,
  7275. 4:20:52wind is not there, I will not play the
  7276. 4:20:55match. This is how you will read this
  7277. 4:20:57one data point, right? Similarly, I have
  7278. 4:20:59multiple data point. Now, I have to
  7279. 4:21:01decide how can I build a decision tree
  7280. 4:21:04and out of these four features, which
  7281. 4:21:07feature should I use first as my root
  7282. 4:21:09node? Right? So, this is what I have to
  7283. 4:21:12decide and
  7284. 4:21:14let's do it
  7285. 4:21:16accordingly, right? So, now what I'll do
  7286. 4:21:18from here is
  7287. 4:21:21uh we I will illustrate you what is Gini
  7288. 4:21:24index, what is information gain, and how
  7289. 4:21:27does this happens, right? And this is
  7290. 4:21:29basically used to build your decision
  7291. 4:21:31tree.
  7292. 4:21:32So, before I make you understand about
  7293. 4:21:35Gini index or information gain, one
  7294. 4:21:38thing which you should always remember
  7295. 4:21:39is to understand the concept of
  7296. 4:21:42impurity. What is impurity? Impurity is
  7297. 4:21:44nothing but you can see there is a
  7298. 4:21:46basket where you have apples, right?
  7299. 4:21:49Now, if you have a basket which is of
  7300. 4:21:50apple and another
  7301. 4:21:52tray, it is written the label as apple.
  7302. 4:21:55Now, in this case, you will never make a
  7303. 4:21:57mistake. You will never make a mistake
  7304. 4:22:00because
  7305. 4:22:01here you have an apple, here you have
  7306. 4:22:03only one label. So, everything will be
  7307. 4:22:05perfect. It will be 100%. Basically,
  7308. 4:22:08there is no impurity, there is no
  7309. 4:22:10problem in your data set, right? On the
  7310. 4:22:13other hand, let me flip the story. In
  7311. 4:22:16this case, imagine that I have different
  7312. 4:22:18fruits in the basket, which is like
  7313. 4:22:20apple, you have a banana, you have a
  7314. 4:22:22grapes, you have a you know, cherries,
  7315. 4:22:25and many other things, right? And you
  7316. 4:22:26have many apples, many labels here. In
  7317. 4:22:29this case, try imagining you have to
  7318. 4:22:32match each fruit with its label.
  7319. 4:22:36Right? Now, in this case, the impurity
  7320. 4:22:38cannot be equal to zero. The impurity
  7321. 4:22:41will not be equal to zero in this case
  7322. 4:22:43because what would happen here is that
  7323. 4:22:45you have a chances of misclassification.
  7324. 4:22:48Right? This is a very important concept,
  7325. 4:22:50right? When you have perfect thing, the
  7326. 4:22:53misclassification will not happen. But,
  7327. 4:22:55whereas, if you have a multiple labels
  7328. 4:22:57with multiple fruit, misclassification
  7329. 4:22:59or impurities will not be equal to zero.
  7330. 4:23:03Right? So, this is associated with a
  7331. 4:23:06term called as entropy. Maybe in your
  7332. 4:23:08childhood days, you have learned about
  7333. 4:23:10entropy in your chemistry class. In a
  7334. 4:23:12simple sense, what is entropy? Entropy
  7335. 4:23:15is a randomness of the space sample
  7336. 4:23:18space. Whenever you are not sure on your
  7337. 4:23:21decision, then entropy will be more.
  7338. 4:23:24Imagine that I'm giving you the data set
  7339. 4:23:26where it is like 51% of doing this
  7340. 4:23:29thing, 49% of not doing this thing. 51%
  7341. 4:23:33chance that the employee may leave the
  7342. 4:23:35organization, 49% chance that employee
  7343. 4:23:37may not leave the organization. So,
  7344. 4:23:39basically, you are not sure on your
  7345. 4:23:40decision. If you are not sure on your
  7346. 4:23:42decision, then the entropy will be very
  7347. 4:23:45high. On the other hand, if you are very
  7348. 4:23:47sure on your decision, then the entropy
  7349. 4:23:50will be very low. So, basically, we need
  7350. 4:23:53that feature which can provide us lowest
  7351. 4:23:56entropy rather than the highest entropy
  7352. 4:24:00uh to select as that as a good feature.
  7353. 4:24:04So, we generally find it out entropy by
  7354. 4:24:06the help of this formula. What we simply
  7355. 4:24:09do is don't get scared with this formula
  7356. 4:24:11because everything happens automatically
  7357. 4:24:13in R or Python, right? Imagine just look
  7358. 4:24:17at this formula. What is this formula?
  7359. 4:24:18This formula says that what is the
  7360. 4:24:21probability that something will happen
  7361. 4:24:23multiplied by log base two probability
  7362. 4:24:27that something will happen subtract this
  7363. 4:24:29with probability that something will not
  7364. 4:24:31happen into log base two probability
  7365. 4:24:34that something will not happen. Okay,
  7366. 4:24:36let's take up an example. Don't worry
  7367. 4:24:38about it. Let's take up an example that
  7368. 4:24:41probability that I will win the match
  7369. 4:24:43you will apply here and probability that
  7370. 4:24:46I will not play the match or win the
  7371. 4:24:47match you will apply here and then you
  7372. 4:24:49will calculate the entropy. Let's take
  7373. 4:24:53you know an example how you can do this
  7374. 4:24:56in our case.
  7375. 4:24:58So, in our case let me show you.
  7376. 4:25:05Yeah, how we will build the decision
  7377. 4:25:07tree in our case. In our case you just
  7378. 4:25:10see here what is happening is that we
  7379. 4:25:12have
  7380. 4:25:1414 instances
  7381. 4:25:16or 14 [clears throat] rows of the data
  7382. 4:25:17set where nine times if you see it I
  7383. 4:25:20will play the match and five times I
  7384. 4:25:23will not play the match. Okay, so if you
  7385. 4:25:25see carefully there are nine labels
  7386. 4:25:27where I'm playing the match and there
  7387. 4:25:29are five labels where I'm not playing
  7388. 4:25:31the match. Right? So, first of all I
  7389. 4:25:33have to find it out the total entropy.
  7390. 4:25:36How I will find the total entropy? This
  7391. 4:25:38is being determined by probability that
  7392. 4:25:41I will play the match sub multiply this
  7393. 4:25:44with log base power two probability that
  7394. 4:25:47I will play the match. What are the
  7395. 4:25:49chances that I will play the match? Nine
  7396. 4:25:51out of 14.
  7397. 4:25:53All of you will be with me, right? This
  7398. 4:25:55is nine out of 14. Multiply with log
  7399. 4:25:58base two nine out of 14.
  7400. 4:26:00Subtract this with what is the
  7401. 4:26:01probability that I will not play the
  7402. 4:26:03match? Five out of 14. Multiply this
  7403. 4:26:06with log base two five out of 14.
  7404. 4:26:09So, once you calculate this you will
  7405. 4:26:11find it out the entropy of this entire
  7406. 4:26:14system, entropy of this entire system is
  7407. 4:26:170.94. Okay, so this is the entropy of
  7408. 4:26:21your entire sample space. This is the
  7409. 4:26:23first thing. Now, how this entropy will
  7410. 4:26:26help you in selecting which features you
  7411. 4:26:28will take or not? So, let's go further.
  7412. 4:26:32So, now we will take each feature one by
  7413. 4:26:34one. Whether I should take outlook,
  7414. 4:26:36whether I should take temperature,
  7415. 4:26:38whether I should pick up humidity, or
  7416. 4:26:40whether should I pick up windy, right?
  7417. 4:26:42Let's go one by one. Now, first I'm
  7418. 4:26:44plotting for outlook. Imagine for
  7419. 4:26:47outlook, how many times is what are the
  7420. 4:26:50distinct value of outlook? Outlook can
  7421. 4:26:51be sunny,
  7422. 4:26:53outlook can be overcast, or outlook can
  7423. 4:26:55be rainy. Right? These are the three
  7424. 4:26:57different combinations you can have for
  7425. 4:26:59outlook. Now, if the outlook is sunny,
  7426. 4:27:02two times I'm playing, three times I'm
  7427. 4:27:04not playing the match. If the outlook is
  7428. 4:27:06overcast, I'm always playing the match.
  7429. 4:27:08If the outlook is rainy, three times I'm
  7430. 4:27:11playing, two times I'm not playing the
  7431. 4:27:12match.
  7432. 4:27:13Right? This is how I have bifurcated it.
  7433. 4:27:16What I will do in the next iteration is,
  7434. 4:27:18let me find it out the entropy of
  7435. 4:27:21outlook.
  7436. 4:27:22Okay? So, if we start with outlook,
  7437. 4:27:26remember this formula which is, you
  7438. 4:27:28know, probability that I will play
  7439. 4:27:30multiply with the probability that I
  7440. 4:27:32will play, right? And subtract with
  7441. 4:27:35probability which I will not play, and
  7442. 4:27:37log base two of not playing. So, 2 by 5
  7443. 4:27:40is a chances that I will not I will play
  7444. 4:27:42into log base two 2 by 5. This is a
  7445. 4:27:45subtraction here, right? With I will
  7446. 4:27:48play and log 3 by 3 by 5 I will play.
  7447. 4:27:52Right? So, you will calculate this
  7448. 4:27:54entropy when outlook is sunny.
  7449. 4:27:57Yeah? So, you got this entropy when the
  7450. 4:27:59outlook is sunny is 0.971.
  7451. 4:28:02Accordingly, you will proceed with
  7452. 4:28:04calculating the entropy when the outlook
  7453. 4:28:07is overcast. Outlook is overcast, always
  7454. 4:28:10you are playing.
  7455. 4:28:11Right? If the outlook is overcast, every
  7456. 4:28:13time you are playing, so it means that
  7457. 4:28:16you will get a probability of zero. If
  7458. 4:28:18you apply in that formula, you will get
  7459. 4:28:20zero.
  7460. 4:28:20And third, what what what would happen
  7461. 4:28:23if the
  7462. 4:28:24if the outlook is sunny? In that case,
  7463. 4:28:27you will again multiply and you will put
  7464. 4:28:29the formula and you will get it out
  7465. 4:28:310.971, right? So, you will find it out
  7466. 4:28:34entropy for each and every distinct
  7467. 4:28:37combination of your feature.
  7468. 4:28:40You got my point, right? Outlook being
  7469. 4:28:42sunny, outlook being overcast, outlook
  7470. 4:28:44being sunny
  7471. 4:28:46uh you know,
  7472. 4:28:47overcast, sunny, and rainy. This should
  7473. 4:28:49be replaced here, right? And then
  7474. 4:28:52what you will do is you will finally
  7475. 4:28:54calculate the information gain.
  7476. 4:28:56Information gain is nothing but what you
  7477. 4:28:58will do is you will pick it up the
  7478. 4:29:00chances, you will pick it up the entire
  7479. 4:29:03chances when you are playing, right?
  7480. 4:29:05Which is five out of 14. You remember
  7481. 4:29:07five out of 14 were the total chances
  7482. 4:29:09that I will play into if it is sunny,
  7483. 4:29:13plus four out of 14 if it is overcast,
  7484. 4:29:16five out of 14 if it is rainy, right?
  7485. 4:29:20And you will calculate the information
  7486. 4:29:22from this outlook. And once you
  7487. 4:29:24calculate the information, you will
  7488. 4:29:25subtract this information from your
  7489. 4:29:28total entropy which I found it out in
  7490. 4:29:30the last slide which was 0.94. You will
  7491. 4:29:33subtract this and you will get the
  7492. 4:29:35information gain. Or this is also called
  7493. 4:29:37as a information gain from a particular
  7494. 4:29:40feature. So, basically these type of a
  7495. 4:29:43calculation first of all, don't get
  7496. 4:29:44scared away that you have to do this
  7497. 4:29:46calculation. But what I'm trying to
  7498. 4:29:48explain you is this is how your
  7499. 4:29:50algorithm will work for each and every
  7500. 4:29:53feature. It will calculate the
  7501. 4:29:55information gain from your feature.
  7502. 4:29:59Right? If you have the more information
  7503. 4:30:01gain, it means that this variable is of
  7504. 4:30:04very very important in predicting that
  7505. 4:30:07something will happen or not. Okay? So,
  7506. 4:30:10for outlook, I got the information gain
  7507. 4:30:13as 0.247 with all the calculation.
  7508. 4:30:16Remember then we will proceed with wind
  7509. 4:30:19if the wind is there or not, right? And
  7510. 4:30:23then I will proceed with wind and I will
  7511. 4:30:25calculate the information gain and say
  7512. 4:30:27information gain I found it out is
  7513. 4:30:290.048, right? Similarly, we will
  7514. 4:30:31calculate for all the four of them. Let
  7515. 4:30:34me put it together for all of you. So,
  7516. 4:30:37this is what happens here. Now, if I put
  7517. 4:30:40in front of all of you, these were the
  7518. 4:30:42four different variable I have, outlook,
  7519. 4:30:45temperature, humidity and wind, right?
  7520. 4:30:47I'm calculating the information gain for
  7521. 4:30:50each one of them and the information
  7522. 4:30:52gain I got for outlook is 0.247.
  7523. 4:30:56So, if the information gain is highest
  7524. 4:30:58in a particular feature, that feature
  7525. 4:31:01will become your root node.
  7526. 4:31:04You got my point? So, therefore, we will
  7527. 4:31:06pick it up outlook as our root node.
  7528. 4:31:09Similarly, for when the tree get
  7529. 4:31:11started, later on as a branch node also,
  7530. 4:31:14your information gain will be calculated
  7531. 4:31:17and wherever whichever feature is giving
  7532. 4:31:19you more information gain, that will be
  7533. 4:31:20picked up later in the
  7534. 4:31:23your tree also.
  7535. 4:31:25So, this is how you will build your
  7536. 4:31:26decision tree and you will finally get a
  7537. 4:31:29decision tree like this, okay? So, now
  7538. 4:31:32what I will do is
  7539. 4:31:34you know, like I will quickly show you
  7540. 4:31:36how do we do the decision tree
  7541. 4:31:39in your Python, okay? So, I'll show you
  7542. 4:31:42how do you do this in Python and then I
  7543. 4:31:44will summarize for all of you that why
  7544. 4:31:47decision trees or tree based algorithms
  7545. 4:31:50are better than your
  7546. 4:31:52other algorithm, okay? So, how do you
  7547. 4:31:55choose
  7548. 4:31:56basically that which algorithm you will
  7549. 4:31:58select and when, okay? So, I'll I'll
  7550. 4:32:00describe you this, but before let's jump
  7551. 4:32:03on to Python and go there.
  7552. 4:32:06So, what I have done here is like I had
  7553. 4:32:08built the decision tree in front like
  7554. 4:32:12you know already. So, quickly I will
  7555. 4:32:15walk you through the commands. Okay? So,
  7556. 4:32:17what happens here is that
  7557. 4:32:19in Python
  7558. 4:32:21as you would be well versed with this
  7559. 4:32:22that we generally import packages in
  7560. 4:32:25Python. So, I'm importing NumPy. I'm
  7561. 4:32:27importing Matplotlib for plotting the
  7562. 4:32:29chart. I'm importing various packages
  7563. 4:32:32from your scikit-learn which is
  7564. 4:32:35for machine learning purposes, right?
  7565. 4:32:36So, I'm importing your label encoder,
  7566. 4:32:39your decision tree classifier,
  7567. 4:32:40classification report, and I also I'm
  7568. 4:32:43importing your
  7569. 4:32:44tree, right? So, we have all this which
  7570. 4:32:48I'm importing, right? After I import,
  7571. 4:32:50what I will do is I'm reading my my data
  7572. 4:32:53set. So, I'm showing you this with the
  7573. 4:32:55help of a Iris data set which is one of
  7574. 4:32:58the very popular data set for building
  7575. 4:33:00the a decision tree, right? For any
  7576. 4:33:03particular source. So, imagine that this
  7577. 4:33:04is my data set and I'm just showing you
  7578. 4:33:06six rows of the data set where you know
  7579. 4:33:08like
  7580. 4:33:09I want to find it out whether a
  7581. 4:33:11particular flower species will be
  7582. 4:33:13setosa, versicolor, or virginica. I have
  7583. 4:33:16three different flowers which I want to
  7584. 4:33:18predict and on the basis of sepal
  7585. 4:33:20length, petal length, sepal width, and
  7586. 4:33:22petal width. So, basically I have
  7587. 4:33:24different dimensions of flowers length
  7588. 4:33:27and width and based on that I want to
  7589. 4:33:29find it out whether the particular
  7590. 4:33:31species will be setosa, versicolor, or
  7591. 4:33:33virginica. You got my point, right? So,
  7592. 4:33:35this is a data set. So, what I will do
  7593. 4:33:37is I as you know with every machine
  7594. 4:33:39learning data set we play with the data
  7595. 4:33:41set. So, I'm checking the information
  7596. 4:33:43here
  7597. 4:33:44like what type of the data type it is.
  7598. 4:33:46So, sepal length it is a float, petal
  7599. 4:33:48length it is a float, and species is a
  7600. 4:33:51object.
  7601. 4:33:52Why object? Because this is a
  7602. 4:33:54categorical column.
  7603. 4:33:56So, after this I'm also checking whether
  7604. 4:33:58there is any null value present in the
  7605. 4:34:00data set or not because if there is any
  7606. 4:34:02null value, then we have to get rid of
  7607. 4:34:04that null value or we have to replace
  7608. 4:34:06that null value with some imputed value,
  7609. 4:34:09right? This is what I'm checking here.
  7610. 4:34:11Once I do that, I'm also plotting this
  7611. 4:34:13because we usually do visualization,
  7612. 4:34:15right? To understand the data set
  7613. 4:34:17better. So, what I'm doing here in the
  7614. 4:34:18with the help of SNS, which is your
  7615. 4:34:20SNS.pairplot, I'm plotting all the
  7616. 4:34:23possible plots. So, basically, sepal
  7617. 4:34:25length with sepal width, what type of
  7618. 4:34:27the combination look like. So, setosa,
  7619. 4:34:29versicolor, and virginica, there are
  7620. 4:34:31three different species you can see in
  7621. 4:34:32the data set. And this is what I'm
  7622. 4:34:35getting a trend between your sepal
  7623. 4:34:36length and sepal width.
  7624. 4:34:38Similarly, this is basically from sepal
  7625. 4:34:40width to petal length.
  7626. 4:34:42Right? So, this is how I'm understanding
  7627. 4:34:45the patterns or I'm understanding the
  7628. 4:34:46relationship between the variables in
  7629. 4:34:48the data set. So, this is what I'm
  7630. 4:34:50understanding here. I'm also checking
  7631. 4:34:52whether there's a correlation or not.
  7632. 4:34:54Higher the shade, it means there will be
  7633. 4:34:56a strong correlation. So, all of these
  7634. 4:34:58things is being done as a part of
  7635. 4:35:00exploratory data analysis before even
  7636. 4:35:02you start your machine learning model,
  7637. 4:35:04right? Once you do this, after that what
  7638. 4:35:06I have to do is after that what I'm
  7639. 4:35:08trying to do is I'm taking your species
  7640. 4:35:11column, what I want to predict as my
  7641. 4:35:13target variable, which is your dependent
  7642. 4:35:15variable. And what I So, this is my the
  7643. 4:35:18dependent variable and with the help of
  7644. 4:35:20what I want to predict will be my
  7645. 4:35:22independent variable. So, I'm calling X
  7646. 4:35:25all my independent variable and I'm
  7647. 4:35:27calling target as my dependent variable.
  7648. 4:35:30Once I do this, then you will also think
  7649. 4:35:32about it that your
  7650. 4:35:35uh the variable which I want to predict
  7651. 4:35:36is the flower species, right? But I want
  7652. 4:35:39to convert this into zero and one class,
  7653. 4:35:42zero, one, two class because your
  7654. 4:35:44computer cannot understand text, right?
  7655. 4:35:46Computer can only understand the
  7656. 4:35:48numbers. So, what I'm doing here is I'm
  7657. 4:35:51uh changing it to the
  7658. 4:35:52uh I'm changing it to the class.
  7659. 4:35:55Right? So, what I'm trying to do here is
  7660. 4:35:57I'm saying, "Wherever it is setosa, it
  7661. 4:35:59will will zero. Wherever it is
  7662. 4:36:01virginica, it will become one. And it is
  7663. 4:36:04versicolor as a third category, it will
  7664. 4:36:06become two. So, now imagine my data set,
  7665. 4:36:08my label to predict becomes zero, one,
  7666. 4:36:11or two instead of three flower which was
  7667. 4:36:13setosa, versicolor, and virginica. This
  7668. 4:36:15is what I'm doing with the encoder here.
  7669. 4:36:18Once I convert this, and this become my
  7670. 4:36:20target, what I will do is as I told you
  7671. 4:36:22that I will split this data set into
  7672. 4:36:24training and test data set. So,
  7673. 4:36:26basically 80% of the data is going into
  7674. 4:36:29the training data set and remaining 20%
  7675. 4:36:32I'm taking as a test data set.
  7676. 4:36:34Okay?
  7677. 4:36:36Now, I will call decision tree
  7678. 4:36:37classifier.
  7679. 4:36:38You know, like I want to make a decision
  7680. 4:36:40tree, and I want to fit this on my
  7681. 4:36:42training data set and the test data set.
  7682. 4:36:45And I will start building the decision
  7683. 4:36:48tree and with the help of, you know,
  7684. 4:36:50from So, here is when I have created the
  7685. 4:36:53decision tree, and here is when I'm
  7686. 4:36:55checking the prediction, how accurate my
  7687. 4:36:57decision tree is. So, I got my
  7688. 4:37:00precision, which is good. I got my
  7689. 4:37:01recall. I got my F1 score, and I also
  7690. 4:37:04get the support, right? So, all these
  7691. 4:37:06matrices are being used to calculate
  7692. 4:37:08your
  7693. 4:37:09how accurate your predictions are. So,
  7694. 4:37:12higher the precision, better the results
  7695. 4:37:14would be. Right? And finally, I'm
  7696. 4:37:16showing to you that how does the tree
  7697. 4:37:18look like? So, I had tried to plot this
  7698. 4:37:21decision tree in front of you. So,
  7699. 4:37:22basically it start with petal length,
  7700. 4:37:24and you can see the gini index coming up
  7701. 4:37:26here or the information gain, which is
  7702. 4:37:28information gain and gini index are
  7703. 4:37:30reciprocal to each other. If you want to
  7704. 4:37:31use information gain, you will get that
  7705. 4:37:33score. If you don't want to use
  7706. 4:37:35information gain, you will get gini
  7707. 4:37:36index. So, they are both reciprocal of
  7708. 4:37:38each other. It is one in the same thing.
  7709. 4:37:40You use gini index or you use
  7710. 4:37:42information gain. They are the two
  7711. 4:37:44different metrics to build your decision
  7712. 4:37:46tree.
  7713. 4:37:47So, you can see that
  7714. 4:37:49based on a petal length, the tree gets
  7715. 4:37:51splitted like this. Then based on the
  7716. 4:37:53petal length
  7717. 4:37:54of different dimensions, your tree
  7718. 4:37:56further split, then it further split,
  7719. 4:37:59then it further split, and finally you
  7720. 4:38:01get to know that whether the particular
  7721. 4:38:03species will be versicolor or virginica.
  7722. 4:38:07Okay? So, this is how your entire
  7723. 4:38:10decision tree is built in Python. Okay?
  7724. 4:38:14So, this is how it's so simple. Uh I
  7725. 4:38:16know it takes time to build this thing,
  7726. 4:38:19but once you are a good data scientist
  7727. 4:38:21and you understand all these things, it
  7728. 4:38:23is very simple to build all these things
  7729. 4:38:26very easily in Python or R.
  7730. 4:38:29Right? So, now
  7731. 4:38:31uh finally going back how would you
  7732. 4:38:34decide how would you decide that which
  7733. 4:38:36algorithm, you know, which algorithm
  7734. 4:38:39will be taken when?
  7735. 4:38:41So, uh what happens is this is on the
  7736. 4:38:44basis of scikit-learn. So, it it starts
  7737. 4:38:48something like this, okay? So, it is on
  7738. 4:38:50the basis of like this that first of all
  7739. 4:38:52you see here that uh
  7740. 4:38:56whether how many samples you have, how
  7741. 4:38:58many data points you have in the data
  7742. 4:39:00set, right? If you have more than 50
  7743. 4:39:02samples, then you will go here. If you
  7744. 4:39:05have less than 50 data point, then you
  7745. 4:39:07know, you will go here. So, you will get
  7746. 4:39:10more data set, right? So, if you have
  7747. 4:39:12more data If you have
  7748. 4:39:13greater than 50 data point, then you
  7749. 4:39:15will go further. You split on your data
  7750. 4:39:17set. If it is not, then you will you
  7751. 4:39:19just the kind of advice that get more
  7752. 4:39:21data set, okay? So, then you will decide
  7753. 4:39:24that what you want to predict. If you
  7754. 4:39:25have a labeled data set, then you will
  7755. 4:39:27go here, right? And do clustering. If
  7756. 4:39:30you don't have the labeled data set,
  7757. 4:39:31then you will see whether you want to
  7758. 4:39:33predict a quantity. If yes, you will go
  7759. 4:39:35in regression. If you want to predict uh
  7760. 4:39:37if you just want to do exploratory
  7761. 4:39:39analysis, you can do dimensional
  7762. 4:39:40reduction. If you have a labeled data
  7763. 4:39:42set, uh you know, for classification,
  7764. 4:39:45you can do all this classification. So,
  7765. 4:39:47basically, this is a cheat sheet which
  7766. 4:39:49we generally use to decide that what we
  7767. 4:39:51have to do with the data set and when.
  7768. 4:39:59Let's understand what is a random
  7769. 4:40:01forest.
  7770. 4:40:02A random forest is constructed by using
  7771. 4:40:05multiple decision trees and the final
  7772. 4:40:08decision is obtained by majority votes
  7773. 4:40:11of these decision trees. So, let me make
  7774. 4:40:13things very simple for you by taking an
  7775. 4:40:16example. Now, suppose we have got three
  7776. 4:40:18independent decision trees. Here we are
  7777. 4:40:20just taking three decision trees and
  7778. 4:40:22I've got an unknown fruit and I want
  7779. 4:40:24that these trees would give me a result
  7780. 4:40:27of what exactly this fruit is. So, I
  7781. 4:40:29pass this fruit to the first decision
  7782. 4:40:31tree, the second decision tree, and the
  7783. 4:40:33third decision tree. Now, a random
  7784. 4:40:35forest is nothing but a combination of
  7785. 4:40:37these decision trees. So, the results
  7786. 4:40:40are being fed into the random forest
  7787. 4:40:42algorithm. So, what it sees is that,
  7788. 4:40:45okay, the first decision tree classifies
  7789. 4:40:47it as peach, the second decision tree
  7790. 4:40:49says that it is an apple, and the third
  7791. 4:40:51one says that it is a peach. So, random
  7792. 4:40:54forest classifier says that, okay, I've
  7793. 4:40:56got the result as two peach and one for
  7794. 4:41:01an apple. So, I would say that the
  7795. 4:41:03unknown fruit is an peach.
  7796. 4:41:06All right, so this is based on the
  7797. 4:41:08majority voting of the decision trees
  7798. 4:41:10and that is how a random forest
  7799. 4:41:12classifier comes to a decision of
  7800. 4:41:14predicting the unknown value.
  7801. 4:41:16Okay, so this was a classification
  7802. 4:41:18problem, so it took the majority vote.
  7803. 4:41:20Now, suppose if it was in regression
  7804. 4:41:22problem, it would have taken mean of it,
  7805. 4:41:24okay? So, now let's move on further to
  7806. 4:41:27understanding what is a decision tree.
  7807. 4:41:29But before that, we should understand
  7808. 4:41:30that random forest the building blocks
  7809. 4:41:33are decision trees and that's why
  7810. 4:41:35studying decision tree becomes important
  7811. 4:41:37because if we understand one decision
  7812. 4:41:40tree, we can apply the same concept to
  7813. 4:41:42random forest, okay? So, now let's move
  7814. 4:41:45on forward and understand the important
  7815. 4:41:47terms in random forest. And this will
  7816. 4:41:49also help us consolidate whatever we
  7817. 4:41:51have learned so far. So, we have taken
  7818. 4:41:53the same small decision tree of the
  7819. 4:41:55previous example, and let's understand
  7820. 4:41:57these are also the important terms which
  7821. 4:41:59will be relevant to random forest also.
  7822. 4:42:01So, the first is the root node. Now,
  7823. 4:42:03here what happens is that the entire
  7824. 4:42:05training data has been fed to the root
  7825. 4:42:07node. And then we've got here that each
  7826. 4:42:10node will ask either true or false
  7827. 4:42:12question with respect to one of the
  7828. 4:42:14feature. And then in response to that
  7829. 4:42:16question, it will partition the data set
  7830. 4:42:18into different subsets. That's what it
  7831. 4:42:20is it is doing here based on the
  7832. 4:42:23condition that it if the mass body mass
  7833. 4:42:25is greater than equal to 2500, it ask a
  7834. 4:42:27question either yes or no. And based on
  7835. 4:42:30that, again further partition is done.
  7836. 4:42:32And if not, then it just classifies the
  7837. 4:42:34species. And then again, what happens is
  7838. 4:42:37that the splitting Now, this is very
  7839. 4:42:39important here. The splitting takes
  7840. 4:42:41place either with the help of a genie or
  7841. 4:42:43entropy methods. And these helps to
  7842. 4:42:45decide the optimal split.
  7843. 4:42:48And we will be discussing about
  7844. 4:42:49splitting methods very soon, right?
  7845. 4:42:51Okay. And then we've got the decision
  7846. 4:42:53nodes which provide the link to the leaf
  7847. 4:42:56nodes. And these are really important
  7848. 4:42:57because then only the leaf nodes will
  7849. 4:43:00tell us what actually the real
  7850. 4:43:02predictions are to which class does the
  7851. 4:43:05species belong. So, now coming to the
  7852. 4:43:08leaf node and these are the end points
  7853. 4:43:10where no further division will take
  7854. 4:43:11place and we will obtain our
  7855. 4:43:13predictions. Okay? So, now coming up to
  7856. 4:43:17another important thing here is working
  7857. 4:43:19of random forest. So, now for working of
  7858. 4:43:22random forest, we will have to
  7859. 4:43:23understand a few important concepts like
  7860. 4:43:26random sampling with replacement,
  7861. 4:43:28feature selection, and also the ensemble
  7862. 4:43:30technique which is used in random forest
  7863. 4:43:33and that is bootstrap aggregation which
  7864. 4:43:35is also known as bagging. So, we will
  7865. 4:43:37understand this with the help of an
  7866. 4:43:39example which will be very simple. And
  7867. 4:43:42then we will go on understanding how
  7868. 4:43:44feature selection is done in both the
  7869. 4:43:46classification and the regression
  7870. 4:43:48problem. Actually, how random forest
  7871. 4:43:50select features for the construction of
  7872. 4:43:53decision trees. Well, in random forest,
  7873. 4:43:55the best split is chosen based on Gini
  7874. 4:43:57impurity or information gain methods.
  7875. 4:44:00So, this also we will understand. Now,
  7876. 4:44:03let us first understand random sampling
  7877. 4:44:05with replacement. Now, what happens here
  7878. 4:44:07is that we have got a small subset of
  7879. 4:44:09the same penguin data set, wherein we
  7880. 4:44:11have got some six rows and four
  7881. 4:44:14features, that means four columns. And
  7882. 4:44:16the arrows that you can see is that now
  7883. 4:44:18we will be creating three subsets from
  7884. 4:44:21this small subset, right? And these
  7885. 4:44:23three subsets will become our decision
  7886. 4:44:25trees. And then we'll be constructing
  7887. 4:44:27decision trees from these subsets. So,
  7888. 4:44:29let us create our first subset. And you
  7889. 4:44:32can see here that the subset is randomly
  7890. 4:44:34being created. And for convenience'
  7891. 4:44:36sake, let me just also show you the
  7892. 4:44:39different subsets here. Okay. So, now
  7893. 4:44:41for better understanding, let us
  7894. 4:44:42understand this that in the first
  7895. 4:44:44subset, if we focus, we've got certain
  7896. 4:44:47random rows here, and we have got
  7897. 4:44:49certain feature. But we do not know how
  7898. 4:44:52this feature has been selected. We got
  7899. 4:44:54island and we got body mass. But in the
  7900. 4:44:56second subset, we got island and flipper
  7901. 4:44:59length. And in the third subset, we got
  7902. 4:45:02body mass and flipper length, right?
  7903. 4:45:04Now, let's look at the rows. Now, when I
  7904. 4:45:07am talking about these features, I will
  7905. 4:45:09say this is feature selection, and
  7906. 4:45:11remember this term. Now, coming to the
  7907. 4:45:13second concept, that is random sampling.
  7908. 4:45:16Now, random sampling is nothing but
  7909. 4:45:17selecting randomly from your subset. So,
  7910. 4:45:21I'm selecting randomly certain rows from
  7911. 4:45:24my subset and creating further subset,
  7912. 4:45:27okay? So, what is replacement here?
  7913. 4:45:30Replacement is can be seen here and can
  7914. 4:45:32be understood with the second subset. We
  7915. 4:45:35see here that the Gentoo species, this
  7916. 4:45:37is being repeated again. And this is
  7917. 4:45:40replacement. That means that when we are
  7918. 4:45:43working with repeated rows, and this row
  7919. 4:45:46can be repeated again in the second or
  7920. 4:45:49the third subset, then this is random
  7921. 4:45:51sampling with replacement. That means my
  7922. 4:45:53random forest can use a row multiple
  7923. 4:45:56times in multiple decision trees, right?
  7924. 4:45:59So, this is the basic concept of random
  7925. 4:46:01sampling with replacement and feature
  7926. 4:46:03selection in random forest. Another
  7927. 4:46:06important term which I would like to
  7928. 4:46:08bring into the notice is that
  7929. 4:46:10when we are working with these type of
  7930. 4:46:13small subsets, these are also known as a
  7931. 4:46:16bootstrap data sets. And when we
  7932. 4:46:18aggregate the results of all these data
  7933. 4:46:20set, it becomes bootstrap aggregation.
  7934. 4:46:23So, just filling in the gaps so that
  7935. 4:46:25later on the concepts become more clear.
  7936. 4:46:28So, now let's move on to drawing
  7937. 4:46:30decision trees of these subsets, okay?
  7938. 4:46:33So, let's draw the decision tree of the
  7939. 4:46:35first subset. Again, we are taking body
  7940. 4:46:38mass as the first root node, and then
  7941. 4:46:40based on a decision like if the mass is
  7942. 4:46:43greater than equal to 3,500, then take a
  7943. 4:46:45decision either yes or no. If it is no,
  7944. 4:46:47then the specie is chinstrap. And if it
  7945. 4:46:50is yes, then again you partition based
  7946. 4:46:52on island. And if it is Torgersen, then
  7947. 4:46:55it is Adélie. And if it is Biscoe, then
  7948. 4:46:58it is Gentoo species. Okay? So, this is
  7949. 4:47:01how we will construct two more decision
  7950. 4:47:03trees of the remaining subsets. So, in
  7951. 4:47:05the second subset, let us just again
  7952. 4:47:07create decision tree. And here now we
  7953. 4:47:10are taking flipper length, and then
  7954. 4:47:12based on a condition that if the flipper
  7955. 4:47:15length is greater than equal to 190,
  7956. 4:47:17then make a split. If it is yes, then
  7957. 4:47:19the specie become Gentoo. And if it is
  7958. 4:47:22no, that means again make a decision
  7959. 4:47:25based on island. And if it is Torgersen,
  7960. 4:47:27it is Adélie. And if the island is Dream
  7961. 4:47:31Island, then it is a chinstrap species.
  7962. 4:47:32So, this is how the decision tree of the
  7963. 4:47:35second subset has been created and this
  7964. 4:47:37is how it will take decisions, right?
  7965. 4:47:40Based on the tree length, depth, and
  7966. 4:47:42also the features it is selecting, okay?
  7967. 4:47:46So, now let's create the third decision
  7968. 4:47:47tree of the third subset and we get a
  7969. 4:47:49decision tree something like this
  7970. 4:47:51wherein body mass if it is greater than
  7971. 4:47:534,000 and if it is yes, then clearly it
  7972. 4:47:56is a Gentoo species and if it is no,
  7973. 4:47:59then again make a partition with the
  7974. 4:48:01with respect to flipper length, another
  7975. 4:48:03feature here, and then if it is again
  7976. 4:48:06greater than equal to 190, then the
  7977. 4:48:08species would be Adélie, else it would
  7978. 4:48:10be chinstrap. So, this is how decision
  7979. 4:48:13tree three will make a decision.
  7980. 4:48:15Now, let's just keep these decision
  7981. 4:48:17trees with us, okay? And we will make
  7982. 4:48:20sense of these trees just in a while.
  7983. 4:48:23Okay?
  7984. 4:48:24But before that, let us understand how
  7985. 4:48:26feature selection is done in a random
  7986. 4:48:28forest. How am I selecting the columns?
  7987. 4:48:31So, for classification, by default the
  7988. 4:48:33feature selection is taken as the square
  7989. 4:48:35root of total number of all the
  7990. 4:48:36features. Now, suppose I've got here
  7991. 4:48:39four features, so it is a classification
  7992. 4:48:41problem, I will take the square root of
  7993. 4:48:43these four features, which becomes two.
  7994. 4:48:45So, decision tree would be constructed
  7995. 4:48:46based on two features each. If suppose I
  7996. 4:48:49had 16 features, then it would be square
  7997. 4:48:51root of 16, that would be four. So, four
  7998. 4:48:53features would be taken in each decision
  7999. 4:48:55tree. All right? And suppose if this
  8000. 4:48:58would have been a regression problem,
  8001. 4:49:00then by default what would happen? The
  8002. 4:49:01features would be selected by taking the
  8003. 4:49:04total number of features and dividing
  8004. 4:49:06them by three, okay? So, this is how by
  8005. 4:49:08default the feature selection is being
  8006. 4:49:10done by a random forest. Okay, now let
  8007. 4:49:13us move on forward to consolidating our
  8008. 4:49:15learning.
  8009. 4:49:16So, now we are coming to ensemble
  8010. 4:49:18techniques, that is also known as
  8011. 4:49:20bootstrap aggregation.
  8012. 4:49:22Random forest uses ensemble techniques.
  8013. 4:49:25And what is ensembling? It just means
  8014. 4:49:27that you're aggregating the result of
  8015. 4:49:29the decision trees and taking the
  8016. 4:49:31majority vote in case of classification
  8017. 4:49:34and the mean in case of regression
  8018. 4:49:35problems and giving the output. Okay. So
  8019. 4:49:38now we have again plotted all our
  8020. 4:49:41decision trees here. And below we can
  8021. 4:49:44see that there's an unknown data and I
  8022. 4:49:47want to predict the species of this
  8023. 4:49:48data. So what will happen is that again
  8024. 4:49:51let us just feed this problem to each of
  8025. 4:49:54the decision trees. And let's see what
  8026. 4:49:57each decision tree makes the prediction.
  8027. 4:49:59So I just feed this unknown data to
  8028. 4:50:01decision tree one and it says that okay
  8029. 4:50:03the species seems to be chinstrap. Okay.
  8030. 4:50:06And then decision tree two says that
  8031. 4:50:08based on the data it has been found that
  8032. 4:50:11the species Adélie. And then decision
  8033. 4:50:14tree three says that no I I with my
  8034. 4:50:16decision tree this species is chinstrap.
  8035. 4:50:19Okay. Now all these data is being fed to
  8036. 4:50:23random forest classifier. And it says
  8037. 4:50:25that okay for chinstrap I've got two
  8038. 4:50:28votes for Adélie it's got one vote. So
  8039. 4:50:31the new species would be chinstrap,
  8040. 4:50:33right? So this is how the bootstrap
  8041. 4:50:35aggregation is done based on the
  8042. 4:50:38majority voting and the decisions taken
  8043. 4:50:41by different decision trees they have
  8044. 4:50:43been combined together aggregated and we
  8045. 4:50:46get an ensemble result in the random
  8046. 4:50:49forest. Okay. So this was very simple
  8047. 4:50:51concept of ensemble techniques which has
  8048. 4:50:53been used in random forest.
  8049. 4:50:56Okay. So now let's move on forward to
  8050. 4:50:58splitting methods. So what are the
  8051. 4:51:00splitting methods that we use in random
  8052. 4:51:02forest? So splitting methods are many
  8053. 4:51:04like Gini impurity, information gain or
  8054. 4:51:07chi-square. So let's discuss about Gini
  8055. 4:51:09impurity. So Gini impurity is nothing
  8056. 4:51:12but it is used to predict the likelihood
  8057. 4:51:14that a randomly selected example would
  8058. 4:51:16be incorrectly classified by a specific
  8059. 4:51:19node. And it is called impurity metric
  8060. 4:51:21because it shows how the model differs
  8061. 4:51:24from a pure division, right? And another
  8062. 4:51:26interesting fact about Gini impurity is
  8063. 4:51:28that the impurity ranges from zero to
  8064. 4:51:31one with zero indicating that all of the
  8065. 4:51:33elements belong to a single class and
  8066. 4:51:36one indicates that only one class exist.
  8067. 4:51:39Now value which is like 0.5, this
  8068. 4:51:42indicates that the elements they are
  8069. 4:51:44uniformly distributed across some
  8070. 4:51:46classes, right? Now moving on forward to
  8071. 4:51:49information gain. Now this is another
  8072. 4:51:51method which random forest can use and
  8073. 4:51:54information gain utilizes entropy. So
  8074. 4:51:57entropy is nothing but it is a measure
  8075. 4:51:59of uncertainty. So information gain
  8076. 4:52:01let's talk about that first. So the
  8077. 4:52:04features they are selected that provide
  8078. 4:52:06most of the information about a class,
  8079. 4:52:08right? And this utilizes the entropy
  8080. 4:52:10concept. So let's see what is entropy.
  8081. 4:52:14This is a measure of randomness or
  8082. 4:52:16uncertainty in the data, right? So we
  8083. 4:52:19will understand this entropy with the
  8084. 4:52:20help of a small example. So don't worry
  8085. 4:52:22about it. So let's understand this
  8086. 4:52:24entropy. Now suppose there's a fruit
  8087. 4:52:26fruit tray with four different fruits,
  8088. 4:52:28right? And what do you feel about the
  8089. 4:52:31entropy here? That means the randomness
  8090. 4:52:33of the data. Is it really easy to
  8091. 4:52:36classify these fruits into the
  8092. 4:52:38respective class? So this becomes really
  8093. 4:52:40uncertain and the data looks messy here.
  8094. 4:52:43But what if we just split here these
  8095. 4:52:45into two trays where in the first tray
  8096. 4:52:48would have peaches and oranges and the
  8097. 4:52:50second tray will have apples and lemons.
  8098. 4:52:53So now this becomes a little more
  8099. 4:52:55certain. We get low randomness here and
  8100. 4:52:58this is called as low entropy. So when
  8101. 4:53:01we move down from tree, that means from
  8102. 4:53:03root node to the leaf nodes, the entropy
  8103. 4:53:06reduces and we can also calculate
  8104. 4:53:09information gain from this entropy. That
  8105. 4:53:12is the difference in entropy before and
  8106. 4:53:14after to split that is known as
  8107. 4:53:16information gain. Okay? So, once we move
  8108. 4:53:19down the tree and start reducing the
  8109. 4:53:21randomness from the data, the entropy
  8110. 4:53:23becomes lower and that is what we want
  8111. 4:53:26in our data. If there's low entropy,
  8112. 4:53:28that means we are likely that the
  8113. 4:53:30predictions would be more accurate and
  8114. 4:53:33we can make predictions very easily as
  8115. 4:53:35compared to very messy data which has
  8116. 4:53:38high entropy. Okay? So, that was about
  8117. 4:53:40entropy and now let us just move on to
  8118. 4:53:44the practical demonstration or a
  8119. 4:53:45hands-on on random forest.
  8120. 4:53:48Okay. So, now it's time for a hands-on
  8121. 4:53:50on random forest. So, let us just import
  8122. 4:53:53a few basic libraries of Python in our
  8123. 4:53:55Jupyter notebook. And we will run this.
  8124. 4:53:58We will import pandas as pd, numpy as
  8125. 4:54:00np, and seaborn as sns. Now, seaborn is
  8126. 4:54:04needed here because we want to load a
  8127. 4:54:05data set, that is a penguins data set
  8128. 4:54:08with the help of seaborn. And this has
  8129. 4:54:10already been preloaded in seaborn. This
  8130. 4:54:12is already loaded data set and seaborn
  8131. 4:54:15has got multiple data sets, you know,
  8132. 4:54:16for practice for beginners. So, it is a
  8133. 4:54:18good way to practice for data sets. Now,
  8134. 4:54:21we can see this asterisk sign that means
  8135. 4:54:23it is telling us to wait. So, let us
  8136. 4:54:25just let it get loaded. So, we get got
  8137. 4:54:27our data in an object called df and we
  8138. 4:54:29can see the first five entries here. And
  8139. 4:54:32this uh data frame is shown in the form
  8140. 4:54:34of a table, rows and columns. And we see
  8141. 4:54:37here some species, island, bill length,
  8142. 4:54:39bill depth, flipper length, body mass,
  8143. 4:54:41and the sex of the penguin. So, our task
  8144. 4:54:43is to specify or to classify these
  8145. 4:54:46species of penguins into their
  8146. 4:54:48respective correct species, right? So,
  8147. 4:54:51we see the shape of our data and we see
  8148. 4:54:53that it is like 344 rows and seven
  8149. 4:54:55columns.
  8150. 4:54:56And we will see the info. So, we see
  8151. 4:54:58df.info and this gives us, along with
  8152. 4:55:01the non-null count, we also get the data
  8153. 4:55:04type of the values. So, we have got
  8154. 4:55:07species, island as the object data type,
  8155. 4:55:09whereas the bill length, bill depth,
  8156. 4:55:11flipper length, and body mass are in
  8157. 4:55:13floating point, or or you can say
  8158. 4:55:15floating data type. And the sex is in
  8159. 4:55:18object data type, right? So, now moving
  8160. 4:55:20on forward to calculating how many null
  8161. 4:55:22values are there with the help of
  8162. 4:55:24df.isnull.sum.
  8163. 4:55:26So, we get certain like some around two
  8164. 4:55:29null values in all these columns, as you
  8165. 4:55:32can see the features like bill length,
  8166. 4:55:34bill depth, flipper length, and body
  8167. 4:55:36mass. Whereas there are 11 null values
  8168. 4:55:38in sex feature, right? So, what we do is
  8169. 4:55:40since they are very small null values,
  8170. 4:55:42we can just drop it, or you can also
  8171. 4:55:44ignore them. So, here in this data
  8172. 4:55:46frame, what I'm doing is I'm just
  8173. 4:55:47dropping these null values, and let us
  8174. 4:55:50just check whether they are they are
  8175. 4:55:51being dropped or not with the help of
  8176. 4:55:53again the same function {dot} is null
  8177. 4:55:55{dot} sum. And then we see that yes,
  8178. 4:55:58they are being dropped from our data
  8179. 4:55:59frame. Now, let us do some feature
  8180. 4:56:01engineering with our data. Now, we have
  8181. 4:56:03seen that we have got some object data
  8182. 4:56:05type in our data frame. And before
  8183. 4:56:07feeding it into algorithm that is random
  8184. 4:56:10forest, we have to transform the
  8185. 4:56:13categorical data or the object data type
  8186. 4:56:15into the numeric. So, we are using here
  8187. 4:56:17one-hot encoding to convert the
  8188. 4:56:19categorical data into numeric. Now,
  8189. 4:56:21there are various ways in Python which
  8190. 4:56:22we can do that, like one-hot encoding or
  8191. 4:56:25you can also use mapping function in
  8192. 4:56:27Python, but here we are using one-hot
  8193. 4:56:29encoding. So, let us just do that.
  8194. 4:56:31And we find here first of all, let us
  8195. 4:56:33apply it on the sex column. And here we
  8196. 4:56:36see that we have got two unique values
  8197. 4:56:37in sex, that is male and female. And we
  8198. 4:56:40use pandas here to get dummies, that is
  8199. 4:56:42how we will apply this one-hot encoding
  8200. 4:56:45because this is how get dummies work.
  8201. 4:56:47So, what happens is here is that the new
  8202. 4:56:50unique values are converted into the
  8203. 4:56:52respective columns in the data frame.
  8204. 4:56:54So, we see here we have got two unique
  8205. 4:56:56values, males and females, and they are
  8206. 4:56:58being converted into the columns. Okay?
  8207. 4:57:00So, one thing to note here is that we
  8208. 4:57:03also get a problem of dummy trap because
  8209. 4:57:06here we see only two unique values. Now,
  8210. 4:57:08suppose if I had six or seven unique
  8211. 4:57:10values and I do this one hot encoding, I
  8212. 4:57:13would have lots of features in my data
  8213. 4:57:16frame and that would lead to several
  8214. 4:57:18complexities. So, what I do is
  8215. 4:57:21to keep things simple, I can use one hot
  8216. 4:57:23encoding when my data frame or my unique
  8217. 4:57:25counts are low, when my unique values
  8218. 4:57:27are less. So, since I had just two or
  8219. 4:57:30three, I can use it. So, I'm using here.
  8220. 4:57:33So, what I do is again, now one row, one
  8221. 4:57:35column as we can see here that it is
  8222. 4:57:37redundant, giving me extra information,
  8223. 4:57:39so I will just drop it. So, I drop this
  8224. 4:57:42first column and what I get in this data
  8225. 4:57:44frame is only male. So, let us just
  8226. 4:57:47infer whether I can also infer females
  8227. 4:57:49from this or not. So, if the value is
  8228. 4:57:51one, that means the penguin is a male
  8229. 4:57:53and if the value is zero, that means the
  8230. 4:57:55penguin is a female. Okay? So, only one
  8231. 4:57:58column is needed for this data frame.
  8232. 4:58:01So, I just kept one and dropped the
  8233. 4:58:02other one. Okay, now apply again one hot
  8234. 4:58:05encoding to the island feature. So, in
  8235. 4:58:08island if we check the unique values,
  8236. 4:58:10we've got three unique values here.
  8237. 4:58:12Torgersen, Biscoe and Dream Island and
  8238. 4:58:14the object is the data type, right? So,
  8239. 4:58:17again we will use pandas, pd.get_dummies
  8240. 4:58:21and we will use apply it on the feature
  8241. 4:58:24island and let's get the head of it. So,
  8242. 4:58:27we get here again the unique values were
  8243. 4:58:29converted into columns and we get here
  8244. 4:58:31expected three columns. And then again
  8245. 4:58:33we will just drop the first column to
  8246. 4:58:35get the remaining two columns. So, here
  8247. 4:58:38also we can infer that if the island is
  8248. 4:58:40Torgersen, if it is one, then it is not
  8249. 4:58:43Dream, neither Biscoe, right? So, this
  8250. 4:58:45is how we can read it from the data
  8251. 4:58:46frame and understand that. Now, remember
  8252. 4:58:48this thing that these two island and
  8253. 4:58:50here sex, these are two independent data
  8254. 4:58:53frames. These are not yet included in
  8255. 4:58:55the main data frame. So, what we will do
  8256. 4:58:57now is we will concatenate the above two
  8257. 4:58:59data frames into the original data
  8258. 4:59:01frame. So, what we do, we again create a
  8259. 4:59:03new data frame that is new data and let
  8260. 4:59:05us just concat with the help of
  8261. 4:59:06pd.concat function and we will concat
  8262. 4:59:09what? df.island and sex. And axis is one
  8263. 4:59:13that means in the column. Okay. So, when
  8264. 4:59:15we will run this, let's see the head of
  8265. 4:59:17it. So, everything gets concatenated in
  8266. 4:59:20a single data frame which is good for
  8267. 4:59:22the feeding this data into or splitting
  8268. 4:59:24the data into test and train data. So,
  8269. 4:59:26now we have this new data frame and
  8270. 4:59:29we've got some repeated columns here
  8271. 4:59:31which needs to be deleted. So, what we
  8272. 4:59:33do is we will delete sex and island here
  8273. 4:59:36which are just repeating because we've
  8274. 4:59:37got here male and we have also got here
  8275. 4:59:40dream and togerson. So, we do not
  8276. 4:59:42require this island column neither this
  8277. 4:59:44sex. So, we just drop it with the help
  8278. 4:59:46of new data.drop and the column names x
  8279. 4:59:49is one in place equals to true, right?
  8280. 4:59:51And let's see the head of this data
  8281. 4:59:53frame. Head of the data frame gives me
  8282. 4:59:55five unique values, right? And now it is
  8283. 4:59:58time to create a separate target
  8284. 5:00:00variable. And what we'll do is we will
  8285. 5:00:03store in a variable called y only
  8286. 5:00:05species. So, what we do is from this new
  8287. 5:00:08data.species, we will just store the
  8288. 5:00:10species in this y. And we see this
  8289. 5:00:12y.head that is the first five species
  8290. 5:00:15and we got the values here. That means
  8291. 5:00:17another target variable has been created
  8292. 5:00:19now.
  8293. 5:00:20So, and you can also see the y.unique
  8294. 5:00:23values as Adelie, Chinstrap and Gentoo.
  8295. 5:00:25So, now we see here three unique values
  8296. 5:00:28of the penguin that is Chinstrap, Adelie
  8297. 5:00:30and Gentoo. And the data type is object
  8298. 5:00:32here. So, again we need to convert this
  8299. 5:00:34object into the numeric data type. So,
  8300. 5:00:36now what we are doing is we are using
  8301. 5:00:38the map function in Python and what we
  8302. 5:00:40do is we map Adelie to zero, Chinstrap
  8303. 5:00:42to one and Gentoo to two. So, this is
  8304. 5:00:45how we see then all the values are being
  8305. 5:00:48mapped to numeric. This is another way
  8306. 5:00:50to convert a categorical value into a
  8307. 5:00:52numeric value in Python. Now, what we do
  8308. 5:00:54is let us just drop the target value
  8309. 5:00:57species from our main data frame. So, we
  8310. 5:00:59just drop it and let's see our new data
  8311. 5:01:01frame. So, we see that we don't have any
  8312. 5:01:04target species here, right? Okay. So, in
  8313. 5:01:07X, let's store this new data and perform
  8314. 5:01:10the splitting of the data. So, what we
  8315. 5:01:13do is from sklearn.model_selection,
  8316. 5:01:15we will import our train_test_split
  8317. 5:01:18and we will split our training data into
  8318. 5:01:2070% and 30%. So, test data becomes 30%
  8319. 5:01:23and training data is some 70%. And this
  8320. 5:01:26random state is zero, which means that
  8321. 5:01:28I'm not fixing any random state. And
  8322. 5:01:31this is also useful for the code
  8323. 5:01:33reproducibility. Now, suppose if I again
  8324. 5:01:35run this code, I will get the same
  8325. 5:01:36result. It will not change. You can set
  8326. 5:01:39this random state to any of the random
  8327. 5:01:40number as per your choice and result
  8328. 5:01:43would differ. Okay. So, now let us print
  8329. 5:01:45the shape of X_train, Y_train, X_test
  8330. 5:01:48and Y_test. So, we see here that it has
  8331. 5:01:50been splitted into 70 and 30% and we get
  8332. 5:01:53X_train as 233 values here and seven
  8333. 5:01:56features. And X_test has 100 values and
  8334. 5:01:59seven features. Similarly, Y_train you
  8335. 5:02:01can see 233 values and Y_test has 100
  8336. 5:02:04values. That means the species. Okay.
  8337. 5:02:07So, that has been perfectly splitted
  8338. 5:02:09into 70 and 30%. Now, what we do is we
  8339. 5:02:12will train the random forest classifier
  8340. 5:02:14on the training set. How do we do it? We
  8341. 5:02:17will import the random forest classifier
  8342. 5:02:19from sklearn.ensemble.
  8343. 5:02:21So, we've already dealt with what is
  8344. 5:02:23ensemble. And then in classifier, we
  8345. 5:02:26will store this random forest and this
  8346. 5:02:28n_estimators is nothing but decision
  8347. 5:02:30tree. So, we are creating some five
  8348. 5:02:31decision trees here. And the criteria is
  8349. 5:02:33entropy. And again, random state is set
  8350. 5:02:36to zero. So, let's see. And then we will
  8351. 5:02:38fit this X_train and Y_train. So, this
  8352. 5:02:40has been fitted and the criteria is
  8353. 5:02:42entropy here. All right. So, now let's
  8354. 5:02:44make some predictions and let's create a
  8355. 5:02:46variable called Y_predict and we will
  8356. 5:02:49just predict it on X_test. And we've
  8357. 5:02:51also printed this Y prediction and now
  8358. 5:02:54let's bring the confusion matrix to
  8359. 5:02:56check the accuracy of random forest
  8360. 5:02:58algorithm. And what we do is from
  8361. 5:03:00matrices as sklearn matrices, we will
  8362. 5:03:02import classification report and
  8363. 5:03:04confusion matrix and also the accuracy
  8364. 5:03:06score. So, we will just import them and
  8365. 5:03:09then in CM variable, we will print the
  8366. 5:03:12confusion of Y test and Y predictions.
  8367. 5:03:15So, we will print it and we see here the
  8368. 5:03:17accuracy score also, which is 98%. So,
  8369. 5:03:21our random forest classifier is giving
  8370. 5:03:23us a very good accuracy of 98% and you
  8371. 5:03:26can see your confusion matrix that only
  8372. 5:03:27two cases have been misclassified. Rest
  8373. 5:03:30all the cases have been correctly
  8374. 5:03:32classified by random forest classifier.
  8375. 5:03:34Okay. So, now let's move on to printing
  8376. 5:03:36the classification report of Y test and
  8377. 5:03:38Y prediction. Let's see and we get the
  8378. 5:03:41precision as 96%. That means
  8379. 5:03:43the two predictions by the algorithm is
  8380. 5:03:4596%. The recall or the true prediction
  8381. 5:03:48rate is 100% which is very nice and Evan
  8382. 5:03:51score is also good which is 98%. So,
  8383. 5:03:54this is giving us a good result. But
  8384. 5:03:56what if if we change the criteria from
  8385. 5:03:58entropy to gini? So, let's just
  8386. 5:04:00experiment with that too. So, let's try
  8387. 5:04:03this with the different number of trees
  8388. 5:04:04and change the criteria to gini
  8389. 5:04:07coefficient. So, now again from
  8390. 5:04:09sklearn.ensemble, we will import random
  8391. 5:04:11forest classifier and fit it, okay? And
  8392. 5:04:14here what we are doing is just we are
  8393. 5:04:16using seven trees. Previously we used
  8394. 5:04:18five and now in the criteria, we will
  8395. 5:04:20use gini coefficient and random state is
  8396. 5:04:23zero. So, let's run this and see whether
  8397. 5:04:25there's a change in accuracy or not and
  8398. 5:04:27let's predict this and let's check the
  8399. 5:04:30accuracy score. What is the accuracy
  8400. 5:04:32score for this random forest classifier
  8401. 5:04:34with seven trees? So, we get 99%
  8402. 5:04:36accuracy with changing the criteria and
  8403. 5:04:39changing the number of trees. So, you
  8404. 5:04:40can just experiment with different
  8405. 5:04:42number of trees and different number of
  8406. 5:04:44decision trees. Let's just experiment
  8407. 5:04:45with, you know, 12 decision trees and
  8408. 5:04:48see what happens.
  8409. 5:04:50So, you can see the accuracy reduced to
  8410. 5:04:5198%. Okay? With seven, we were getting
  8411. 5:04:5599. So, let's just keep seven because it
  8412. 5:04:57is giving us really good accuracy. So,
  8413. 5:04:59this is about random forest classifier
  8414. 5:05:02and how it works with several trees and
  8415. 5:05:05different criteria to give us very good
  8416. 5:05:07accuracy on our training and test data.
  8417. 5:05:15>> [music]
  8418. 5:05:15>> The case that we are discussing is
  8419. 5:05:19basically the KNN algorithm, which is
  8420. 5:05:22an algorithm which we used for mostly
  8421. 5:05:25machine
  8422. 5:05:26learning, which is where you have
  8423. 5:05:28labeled data.
  8424. 5:05:30Okay, so
  8425. 5:05:33KNN algorithm. KNN algorithm is K
  8426. 5:05:36nearest neighbors algorithm.
  8427. 5:05:38And it is an example of supervised
  8428. 5:05:40learning algorithm where
  8429. 5:05:42basically, you try to classify a new
  8430. 5:05:45data point based on the neighbors of
  8431. 5:05:48that data point, which is basically
  8432. 5:05:50which data points are closer to it. For
  8433. 5:05:53example, here, as you can see,
  8434. 5:05:56you have on one side couple of cats and
  8435. 5:05:59on the other side you have couple of On
  8436. 5:06:02one side you have dogs
  8437. 5:06:04and on the other side you have cats.
  8438. 5:06:07Right? Now, if a new data's point is
  8439. 5:06:11given to us, there is a
  8440. 5:06:13picture of a new animal.
  8441. 5:06:15And
  8442. 5:06:17if it is lying somewhere here, right?
  8443. 5:06:20Then, we know that it is nearer to the
  8444. 5:06:23cats, right? And therefore, we will
  8445. 5:06:24classify it as cat.
  8446. 5:06:26Whereas,
  8447. 5:06:27if it is
  8448. 5:06:29sort of
  8449. 5:06:30uh
  8450. 5:06:31nearer to the dogs, then we classify it
  8451. 5:06:33as dog, right? So, that's the like, you
  8452. 5:06:36know, neighborhood for the dog. And
  8453. 5:06:38therefore, we sort of uh classify it
  8454. 5:06:42that new animal or the new picture as as
  8455. 5:06:45being
  8456. 5:06:46all that of a dog. And this is quite
  8457. 5:06:50you know, this is something which is
  8458. 5:06:52even seen our in our regular day-to-day
  8459. 5:06:54life, right? You know, we have had
  8460. 5:06:57examples where our parents keep telling
  8461. 5:06:59us, "Okay, don't play with those kinds
  8462. 5:07:02of you know, children or something
  8463. 5:07:04because they are not good in their
  8464. 5:07:06studies or probably they are not
  8465. 5:07:08so good in their behavior because you
  8466. 5:07:10would become like them, right?" So, it's
  8467. 5:07:12again an example from real life of
  8468. 5:07:14classifying a particular person based on
  8469. 5:07:16the company that they keep, right? Or
  8470. 5:07:18from the
  8471. 5:07:20with the kind of people that they are.
  8472. 5:07:22So, that's that's sort of
  8473. 5:07:26the example and now let me actually go
  8474. 5:07:28back.
  8475. 5:07:30Okay.
  8476. 5:07:31What are the features of of K nearest
  8477. 5:07:33neighbors algorithm?
  8478. 5:07:36Okay, so
  8479. 5:07:37let's talk about
  8480. 5:07:39the features of KNN. So, as as I said,
  8481. 5:07:42KNN is a supervised learning algorithm.
  8482. 5:07:45Typically, it is you know,
  8483. 5:07:47used for supervised learning kind of
  8484. 5:07:49problems.
  8485. 5:07:50It's very simple as we mentioned.
  8486. 5:07:52Intuitively, you can you know, it's
  8487. 5:07:55about you know, what kind of neighbors
  8488. 5:07:56do you have? So, your class is predicted
  8489. 5:07:59based on the your nearest neighbors as
  8490. 5:08:01the name suggests. And then it's a
  8491. 5:08:04non-parametric technique. So, I would
  8492. 5:08:06like to spend a couple of minutes here
  8493. 5:08:09to discuss about what we mean by by
  8494. 5:08:12non-parametric. So, typically, you know,
  8495. 5:08:15the supervised machine learning
  8496. 5:08:16algorithms
  8497. 5:08:18are of two kinds, right? One is the
  8498. 5:08:20parametric types and the second one is
  8499. 5:08:22the non-parametric type. When we say
  8500. 5:08:24parametric, what we mean is basically
  8501. 5:08:27that the machine or the algorithm
  8502. 5:08:29assumes
  8503. 5:08:31that there is an underlying function or
  8504. 5:08:34a distribution that is known of the
  8505. 5:08:36particular data set. Like for example,
  8506. 5:08:39the linear regression would assume that
  8507. 5:08:40the relationship between
  8508. 5:08:43two two things X and Y is is linear in
  8509. 5:08:46nature, right?
  8510. 5:08:47Um and and and similarly for a
  8511. 5:08:50distribution a Gaussian distribution it
  8512. 5:08:52would assume a normal distribution of
  8513. 5:08:54the data points and so on. So,
  8514. 5:08:57a lot of the supervised machine learning
  8515. 5:08:59algorithms they assume some kind of
  8516. 5:09:01function association between the
  8517. 5:09:04predictor which is basically the root
  8518. 5:09:06causes which help you in predicting and
  8519. 5:09:09the variable that you're trying to
  8520. 5:09:10predict.
  8521. 5:09:12Whereas there are certain algorithms
  8522. 5:09:14like the KNN
  8523. 5:09:15or the Parzen window or the linear
  8524. 5:09:18discriminant analysis which is of the
  8525. 5:09:20kind which is called non-parametric
  8526. 5:09:22because it does not assume
  8527. 5:09:24any particular kind of distribution or
  8528. 5:09:26any particular kind of functional
  8529. 5:09:28relationship of the data that you're
  8530. 5:09:30trying to you know predict or you're
  8531. 5:09:33trying to
  8532. 5:09:34learn the the pattern of.
  8533. 5:09:37So, KNN is one of the kind as I said
  8534. 5:09:39Parzen window is another one
  8535. 5:09:42uh which is basically where you uh in in
  8536. 5:09:46Parzen window essentially, you know, the
  8537. 5:09:49the volume of the data set uh or the
  8538. 5:09:53area that the data set covers is known
  8539. 5:09:55whereas you're trying to find the K
  8540. 5:09:57there where uh which is the number of
  8541. 5:09:59data points within that area or the
  8542. 5:10:02volume, right? Whereas in case of
  8543. 5:10:03non-parametric technique like KNN
  8544. 5:10:06it's the opposite where the K is known
  8545. 5:10:09which is basically you would like to
  8546. 5:10:12associate the class of of
  8547. 5:10:14the data point that you're trying to
  8548. 5:10:15predict based on the K number of
  8549. 5:10:19neighbors around it. Now, K can be
  8550. 5:10:21three, four, five, whatever, right? 10,
  8551. 5:10:2420, and so on.
  8552. 5:10:26Essentially, the difference between
  8553. 5:10:27Parzen window and KNN is that in KNN you
  8554. 5:10:31already know the K and then from the K
  8555. 5:10:33you try to find out the volume and
  8556. 5:10:35therefore then you try to find the the
  8557. 5:10:37probability of the density underlying
  8558. 5:10:39distribution.
  8559. 5:10:40And then there is the the discriminant
  8560. 5:10:43analysis which is off again two kinds
  8561. 5:10:45linear discriminant and multiple
  8562. 5:10:46discriminant analysis where basically
  8563. 5:10:48what you do is you you transform the
  8564. 5:10:52underlying data the features into a
  8565. 5:10:54higher dimension and in such a way that
  8566. 5:10:57in the new feature space after you have
  8567. 5:10:59transformed the data
  8568. 5:11:01you now try to apply a parametric
  8569. 5:11:03approach like for example you will try
  8570. 5:11:05to project the features onto a line or
  8571. 5:11:09if it is on a sub sub space which is
  8572. 5:11:13higher dimension than a line then
  8573. 5:11:15essentially it becomes multiple
  8574. 5:11:16discriminant analysis. So
  8575. 5:11:18basically
  8576. 5:11:19um those are the three kinds of
  8577. 5:11:21non-parametric techniques. So even if
  8578. 5:11:23you were not able to sort of get the
  8579. 5:11:25full hang of
  8580. 5:11:27what these three types are
  8581. 5:11:29what you need to keep in mind is that
  8582. 5:11:31non-parametric technique does not assume
  8583. 5:11:34any kind of distribution or any kind of
  8584. 5:11:36functional relationship of the
  8585. 5:11:38underlying data
  8586. 5:11:39and therefore it gives us a lot of
  8587. 5:11:40flexibility whereas the parametric
  8588. 5:11:43techniques they basically assume some
  8589. 5:11:45kind of functional relationship
  8590. 5:11:48between the data points
  8591. 5:11:50or they assume some kind of
  8592. 5:11:51distribution.
  8593. 5:11:53So where would you use a non-parametric
  8594. 5:11:55versus a parametric technique right?
  8595. 5:11:57So basically you would use
  8596. 5:11:59non-parametric technique you know where
  8597. 5:12:02you do not know about the functional
  8598. 5:12:03relationship that is one
  8599. 5:12:05second is that you know there is
  8600. 5:12:09maybe let's say large amount of data and
  8601. 5:12:12so on. And thirdly
  8602. 5:12:14non-parametric technique like KNN does
  8603. 5:12:16not work in very high dimensional data.
  8604. 5:12:20So you would also use parametric
  8605. 5:12:22techniques in that case.
  8606. 5:12:24Whereas if you have you know smaller
  8607. 5:12:26data sets, you would use, uh, you know,
  8608. 5:12:29typically non-parametric techniques. And
  8609. 5:12:32then also, if you know understand that
  8610. 5:12:34the relationship between the data points
  8611. 5:12:36might be, for example, linear or
  8612. 5:12:39something like that, you would use a
  8613. 5:12:40parametric technique. So, you're
  8614. 5:12:42assuming that there is some kind of
  8615. 5:12:43relationship, uh, like a linear
  8616. 5:12:45relationship or something like that in
  8617. 5:12:47between the data points.
  8618. 5:12:48So, that's where you will use the
  8619. 5:12:50parametric technique. So,
  8620. 5:12:51again, just to summarize, basically
  8621. 5:12:53non-parametric is an approach where you
  8622. 5:12:55do not assume any kind of distribution
  8623. 5:12:57or functional relationship, whereas
  8624. 5:12:59parametric assumes a functional
  8625. 5:13:00relationship or basically a distribution
  8626. 5:13:02between the data points.
  8627. 5:13:05The other feature of KNN is that it is a
  8628. 5:13:07lazy algorithm. So, what lazy algorithm?
  8629. 5:13:10Actually, in most of the supervised, uh,
  8630. 5:13:13learning algorithms,
  8631. 5:13:15uh, you basically train your model on
  8632. 5:13:17the training data set,
  8633. 5:13:19then you have your model,
  8634. 5:13:21and then you apply this particular model
  8635. 5:13:23to the test data set to then classify or
  8636. 5:13:26predict.
  8637. 5:13:27Uh, you know, for example, whether a new
  8638. 5:13:29image is that of a cat or a dog, right?
  8639. 5:13:33So, this can be, you know, some kind of
  8640. 5:13:35algorithm like support vector machine
  8641. 5:13:37or regression or logistic regression or
  8642. 5:13:39whatever, right? So,
  8643. 5:13:40you run this logistic regression or you
  8644. 5:13:42run this regression or support vector
  8645. 5:13:45machine on the training data set,
  8646. 5:13:47and it learns the features or it learns
  8647. 5:13:50the parameters of the model from that
  8648. 5:13:52data set, and then applies this learned
  8649. 5:13:55model on the test data set.
  8650. 5:13:58Whereas, in case of KNN, actually, there
  8651. 5:14:01is no training step at all. That's why
  8652. 5:14:04it's called a lazy algorithm.
  8653. 5:14:06Because what it does is,
  8654. 5:14:07at the time that you are actually now
  8655. 5:14:10want to predict,
  8656. 5:14:12at that point in time, actually, it will
  8657. 5:14:13go and it will do all the calculations
  8658. 5:14:15of the distance of the new data point,
  8659. 5:14:18like, for example, the new image of the
  8660. 5:14:20cat or dog, from all the other data
  8661. 5:14:23points that you have.
  8662. 5:14:25So, it will calculate all the distances
  8663. 5:14:27and then will check for the those data
  8664. 5:14:29points which are or the K data points
  8665. 5:14:31which are nearest to this.
  8666. 5:14:33So, that's why it's called the lazy
  8667. 5:14:35algorithm because nothing happens
  8668. 5:14:38till the point or no calculations happen
  8669. 5:14:40till the point you are actually trying
  8670. 5:14:41to predict something.
  8671. 5:14:43So, there is no training step involved.
  8672. 5:14:45Okay, and then
  8673. 5:14:47it's used for both classification and
  8674. 5:14:49regression as we just mentioned. So, it
  8675. 5:14:51can be used to predict the values as
  8676. 5:14:53well as be able to classify something
  8677. 5:14:56like, you know, okay, whether it is a
  8678. 5:14:57cat or dog or if you're trying to, let's
  8679. 5:15:00say, predict some value
  8680. 5:15:02some forecast or something that for that
  8681. 5:15:04also you can use it. And then it is
  8682. 5:15:06based on feature similarity, which is
  8683. 5:15:07basically what do we mean by feature
  8684. 5:15:10similarity? So, feature can be things
  8685. 5:15:11like if, for example, you know, you are
  8686. 5:15:14looking at classifying cats versus dog,
  8687. 5:15:16right? So, is the eyes like a dog? That
  8688. 5:15:20can be one of the features. Is the what
  8689. 5:15:22how do the ears look? That can be one of
  8690. 5:15:24the features.
  8691. 5:15:25What about the tongue, the face, and so
  8692. 5:15:27on? So, there can be multiple such
  8693. 5:15:28features.
  8694. 5:15:30And how similar it these features are
  8695. 5:15:33between two data points,
  8696. 5:15:35which is used basically by the KNN
  8697. 5:15:38algorithm.
  8698. 5:15:39And then
  8699. 5:15:40as I said, there is no training step
  8700. 5:15:41involved. So, these are the features of
  8701. 5:15:44KNN algorithm.
  8702. 5:15:45And therefore, now let's look at
  8703. 5:15:47actually just some simple examples of
  8704. 5:15:49how it works.
  8705. 5:15:52So, as you can see here in this slide,
  8706. 5:15:54we have
  8707. 5:15:56two classes of data. So, one is the all
  8708. 5:15:58these blue data points.
  8709. 5:16:00And there's another one which is the
  8710. 5:16:02orange data points.
  8711. 5:16:04Now, if you have a new data point, which
  8712. 5:16:06is this
  8713. 5:16:08pink one here, which class should it
  8714. 5:16:10belong to?
  8715. 5:16:12Should it belong to class A or should it
  8716. 5:16:14belong to class B? So, what you would do
  8717. 5:16:16is you would actually start calculating
  8718. 5:16:18the distance of this pink data point
  8719. 5:16:21from every square or blue triangle data
  8720. 5:16:24point. And then
  8721. 5:16:26you will decide you will have to assume
  8722. 5:16:28a particular K. Let's say
  8723. 5:16:30K is three.
  8724. 5:16:32Right? Which is I'm looking at the
  8725. 5:16:34nearest three data points. And in that
  8726. 5:16:38case, basically, as you can see if we
  8727. 5:16:41draw the circle, right?
  8728. 5:16:42Then we see that two of the nearest data
  8729. 5:16:46points within that circle is of the
  8730. 5:16:49square orange kind. So, basically, we
  8731. 5:16:51will predict that this particular new
  8732. 5:16:53data point belongs to class A.
  8733. 5:16:55Whereas, K value was seven, right? As in
  8734. 5:16:59this particular example now.
  8735. 5:17:01You would see that four out of the seven
  8736. 5:17:03is actually of the blue triangle kind.
  8737. 5:17:05And therefore, we will now classify it
  8738. 5:17:07as belonging to class B.
  8739. 5:17:09So, essentially, this prediction
  8740. 5:17:12changes, as you can see here, depending
  8741. 5:17:15on what is the K value.
  8742. 5:17:17So,
  8743. 5:17:18therefore, the question is what should
  8744. 5:17:20be the value of K, right? And typically,
  8745. 5:17:22what happens is you run a trial and
  8746. 5:17:24error, and basically, you will come up
  8747. 5:17:27with okay, what is the best K value. But
  8748. 5:17:29essentially, what one needs to
  8749. 5:17:31understand is that as the value of K
  8750. 5:17:35increases, basically, the the partition
  8751. 5:17:38line starts moving towards becoming more
  8752. 5:17:41and more linear. So, it starts becoming
  8753. 5:17:43less flexible, and it starts assuming
  8754. 5:17:45some kind of a linear dividing line or
  8755. 5:17:48something like that. So,
  8756. 5:17:50what happens is that in that as you
  8757. 5:17:53increase the K,
  8758. 5:17:54your bias
  8759. 5:17:56basically, increases, but your variation
  8760. 5:18:00reduces. So, we know that in
  8761. 5:18:03classification problems or in machine
  8762. 5:18:04learning problems, bias and variance are
  8763. 5:18:07two things that we are trying to manage,
  8764. 5:18:09right? Bias is basically how close you
  8765. 5:18:12are to the actual class or how or to the
  8766. 5:18:15actual
  8767. 5:18:16value. Whereas variation is how much
  8768. 5:18:19variability is there in your prediction.
  8769. 5:18:22So,
  8770. 5:18:22as the K increases, the bias
  8771. 5:18:25sort of increases, but the variance
  8772. 5:18:27reduces. And then it is vice versa. So,
  8773. 5:18:29if your K decreases, let's say
  8774. 5:18:31at K is equal to 1, where you are only
  8775. 5:18:34looking at just one nearest neighbor and
  8776. 5:18:36then
  8777. 5:18:37predicting based on that, actually the
  8778. 5:18:40bias is the least, which means it is the
  8779. 5:18:42most flexible.
  8780. 5:18:43K is equal to 1 is the will give you the
  8781. 5:18:45most flexible sort of demarcating line
  8782. 5:18:49or function.
  8783. 5:18:50Whereas the variability will be the
  8784. 5:18:52maximum.
  8785. 5:18:54So, that that's the sort of the
  8786. 5:18:55trade-off.
  8787. 5:18:57And that's how we actually determine K.
  8788. 5:19:00So, we have to get a K value in such a
  8789. 5:19:02way based on trial and error that
  8790. 5:19:04sort of maximizes our sort of or reduces
  8791. 5:19:07the bias as well as the variance. And
  8792. 5:19:10and that's the kind of optimization we
  8793. 5:19:11are trying to do.
  8794. 5:19:13Okay, so how do we calculate the
  8795. 5:19:14distance itself, right? And distance
  8796. 5:19:17typically can be of many kinds. So, you
  8797. 5:19:20know, the example here is of the
  8798. 5:19:21Euclidean distance, but you can have
  8799. 5:19:23other kinds of distances like Manhattan
  8800. 5:19:25distance or Mahalanobis distance and you
  8801. 5:19:28can look up references for other kinds
  8802. 5:19:31of distances.
  8803. 5:19:32Now, Euclidean distance is calculated
  8804. 5:19:35for the point P1 and P2 as given here.
  8805. 5:19:37Essentially, Euclidean distance is
  8806. 5:19:40nothing but the you know, square root of
  8807. 5:19:42sum of the X coordinates of these two
  8808. 5:19:45points P1 and P2 and then Y coordinate
  8809. 5:19:48square of
  8810. 5:19:49of of these two points P1 and P2.
  8811. 5:19:52And
  8812. 5:19:53this is just an example of one kind of
  8813. 5:19:55distance and
  8814. 5:19:56other kinds of distances like Manhattan
  8815. 5:19:59or Mahalanobis are also there.
  8816. 5:20:01And then this calculating this distance
  8817. 5:20:03becomes quite challenging,
  8818. 5:20:05especially in cases where
  8819. 5:20:07you know, you are trying to for example
  8820. 5:20:09calculate, let's say, how close two
  8821. 5:20:11LinkedIn profiles are, right? Or trying
  8822. 5:20:13to classify uh the category of
  8823. 5:20:16electrocardiogram and so on so forth.
  8824. 5:20:18So, there we have to bring in more
  8825. 5:20:20creativity to just decide what kind of
  8826. 5:20:22distance to use.
  8827. 5:20:24Okay. Now, let's move ahead. We will now
  8828. 5:20:27talk of some use cases where KNN can be
  8829. 5:20:29used and this is an example of how KNN
  8830. 5:20:32can be used for book recommendation. So,
  8831. 5:20:34if you have purchased books on Amazon or
  8832. 5:20:36whatever, right?
  8833. 5:20:38Some of these recommendations are based
  8834. 5:20:39on on KNN algorithm. And then, you know,
  8835. 5:20:43as we said, you know, KNN is like based
  8836. 5:20:45on features, right? So, maybe let's say
  8837. 5:20:48what will be the nearest neighbors of a
  8838. 5:20:50particular book. It can be based on who
  8839. 5:20:51is the author, what is the topic, and so
  8840. 5:20:54on so forth.
  8841. 5:20:55And then there are other use cases like
  8842. 5:20:58I mentioned. So, for classifying
  8843. 5:20:59satellite images, for classifying
  8844. 5:21:01handwritten digits,
  8845. 5:21:03uh on on image analytics or or for
  8846. 5:21:07classifying electrocardiograms,
  8847. 5:21:09um etc. Uh you know, typically KNN can
  8848. 5:21:12be used.
  8849. 5:21:13Okay. So, now actually we will get into
  8850. 5:21:15some hands-on.
  8851. 5:21:18Okay. So, to start the hands-on session,
  8852. 5:21:21I'll go to this Jupyter notebook that I
  8853. 5:21:25already have installed on my system.
  8854. 5:21:28And I have a certain
  8855. 5:21:31code written, which uh we will take two
  8856. 5:21:33examples.
  8857. 5:21:34Both the examples are based on data sets
  8858. 5:21:37which are available in the open source.
  8859. 5:21:39So, you can easily get access to that
  8860. 5:21:42data.
  8861. 5:21:43So, what we do is we start by importing
  8862. 5:21:47the necessary libraries.
  8863. 5:21:49So, we import pandas, seaborn, numpy,
  8864. 5:21:54and matplotlib. Basically, pandas and
  8865. 5:21:56numpy are there for doing the data
  8866. 5:21:59manipulation and also for storing data
  8867. 5:22:02as matrices or as arrays and and be able
  8868. 5:22:05to perform some mathematical procedures
  8869. 5:22:08on them.
  8870. 5:22:09And then seaborn is basically used for
  8871. 5:22:11plotting and matplotlib for plotting as
  8872. 5:22:13well.
  8873. 5:22:14And this line here get IPython just
  8874. 5:22:17helps us to run the images that we'll be
  8875. 5:22:20creating in line with Jupiter notebook
  8876. 5:22:22instead of opening up a new window.
  8877. 5:22:25So let's run this.
  8878. 5:22:27And what it will do is it will import
  8879. 5:22:28all these packages for us which we are
  8880. 5:22:30going to use.
  8881. 5:22:32And then we will first import the breast
  8882. 5:22:35cancer data that is available in your
  8883. 5:22:38scikit-learn datasets.
  8884. 5:22:40So we import that. And then let's just
  8885. 5:22:44initialize that data into a variable
  8886. 5:22:47here called cancer. So cancer here
  8887. 5:22:49represents all the load the breast
  8888. 5:22:51cancer data.
  8889. 5:22:53And now we will let's actually look at
  8890. 5:22:55what this data is.
  8891. 5:22:58It's a bunch of attributes in this
  8892. 5:23:01dictionary here. So you have data, the
  8893. 5:23:03target which is basically nothing but
  8894. 5:23:05whether it is a cancer or not. So
  8895. 5:23:08whether it is malignant or benign.
  8896. 5:23:09Malignant means it's a bad cancer and
  8897. 5:23:12benign means well it's just a tumor,
  8898. 5:23:13it's not cancerous.
  8899. 5:23:15Target name description, feature names
  8900. 5:23:18which is basically
  8901. 5:23:19the features that will tell us whether a
  8902. 5:23:21particular
  8903. 5:23:23case belongs to cancerous or or
  8904. 5:23:25non-cancer or malignant or benign.
  8905. 5:23:28And then
  8906. 5:23:29actually let's just print the
  8907. 5:23:30description of this particular data
  8908. 5:23:32here.
  8909. 5:23:34So as you can see we can use this
  8910. 5:23:36command to print the description here.
  8911. 5:23:39And then we see that there are 569 data
  8912. 5:23:42points with about 30 attributes.
  8913. 5:23:45And these attributes are radius,
  8914. 5:23:46texture, perimeter, etc.
  8915. 5:23:48And
  8916. 5:23:50the the max and min values of those are
  8917. 5:23:53given here.
  8918. 5:23:55And then now let's look at some of the
  8919. 5:23:57feature names.
  8920. 5:23:59So, these are the feature names, radius,
  8921. 5:24:03texture, and so on.
  8922. 5:24:06And now, let's actually set up a data
  8923. 5:24:08frame of this particular data here
  8924. 5:24:13using pandas, this function here. So,
  8925. 5:24:15there are 569
  8926. 5:24:18data points.
  8927. 5:24:19And all these are basically your
  8928. 5:24:21features, as we talked about.
  8929. 5:24:25And let's look at the target variable,
  8930. 5:24:28which is nothing but whether it is
  8931. 5:24:29telling us whether a particular data of
  8932. 5:24:31point belongs to malignant or benign.
  8933. 5:24:33So, zero is cancerous and one is
  8934. 5:24:35non-cancerous.
  8935. 5:24:37And then,
  8936. 5:24:38we convert the target into a data frame
  8937. 5:24:41as well.
  8938. 5:24:42And then, let's look at the couple of
  8939. 5:24:45examples of how the data points look
  8940. 5:24:48like. This is, you know, one row of the
  8941. 5:24:51data points which with all the several
  8942. 5:24:53feature values that we have.
  8943. 5:24:56So, basically, we use this um package
  8944. 5:25:00called standard scalar from scikit-learn
  8945. 5:25:02for pre-processing and for standardizing
  8946. 5:25:04the variables.
  8947. 5:25:06And we initialize this standard scalar
  8948. 5:25:09into a variable called scalar.
  8949. 5:25:11So, standardizing is nothing but, you
  8950. 5:25:13know, basically, bringing all the
  8951. 5:25:15samples to essentially the same range,
  8952. 5:25:18right? Because
  8953. 5:25:20uh what might happen is some of the data
  8954. 5:25:22point, like, for example, temperature
  8955. 5:25:23might be from zero to 100 and some price
  8956. 5:25:26might be from, let's say, 1,000 to
  8957. 5:25:30100,000 or whatever, right? So, the
  8958. 5:25:32absolute values can can lead to some
  8959. 5:25:34issues with respect to the prediction.
  8960. 5:25:36Therefore, we have to standardize it or
  8961. 5:25:37bring it between, let's say, minus one
  8962. 5:25:40and one. So, and then, a mean of zero,
  8963. 5:25:42right? So, we have to bring everything
  8964. 5:25:44to the same scale to be able to compare
  8965. 5:25:46the samples.
  8966. 5:25:47So, we first of all, we fit the this
  8967. 5:25:51standardization um or normalization on
  8968. 5:25:54the data set we have.
  8969. 5:25:56And that is we calculate the the
  8970. 5:25:59variance and and the means and then we
  8971. 5:26:02actually apply it on the data set to
  8972. 5:26:04transform it to the actual values.
  8973. 5:26:07And then
  8974. 5:26:09if we look at the scale values now, so
  8975. 5:26:11let's look at the scale values.
  8976. 5:26:14And this will give an example of the top
  8977. 5:26:17five rows here.
  8978. 5:26:18So we can see now the values are between
  8979. 5:26:20minus one and one.
  8980. 5:26:22Or rather it is standardized.
  8981. 5:26:24Essentially with a normal distribution.
  8982. 5:26:27And then we divide this data into
  8983. 5:26:30test and train. So basically we will
  8984. 5:26:34train the model and then we will test it
  8985. 5:26:36on a separate data set. If you use the
  8986. 5:26:39same random state, you should be able to
  8987. 5:26:41get the same result. Otherwise you may
  8988. 5:26:43get a different result here. And
  8989. 5:26:44essentially we are keeping the testing
  8990. 5:26:47size to 30 which means that we are
  8991. 5:26:48dividing the entire data set into two
  8992. 5:26:50parts. The train part which is having
  8993. 5:26:5270% of the data and the test part which
  8994. 5:26:55is having 30% of the data. And again we
  8995. 5:26:56are using this package called train test
  8996. 5:26:59split from the scikit-learn package.
  8997. 5:27:02So we get the X and the Ys which are
  8998. 5:27:05basically nothing but your train and the
  8999. 5:27:09X's are your predictors and Y is your
  9000. 5:27:11predicted variable whether it is
  9001. 5:27:13cancerous or not.
  9002. 5:27:15And then now let's import the K nearest
  9003. 5:27:17neighbors classifier. This is the actual
  9004. 5:27:19algorithm
  9005. 5:27:20which we are importing from scikit-learn
  9006. 5:27:23package.
  9007. 5:27:24And now we initialize this particular
  9008. 5:27:28algorithm.
  9009. 5:27:30And then we fit it on the
  9010. 5:27:33data.
  9011. 5:27:37And some of the parameters as you can
  9012. 5:27:40see
  9013. 5:27:41is basically what is the leaf size and
  9014. 5:27:42so on so forth.
  9015. 5:27:44The nearest N neighbors we are taking.
  9016. 5:27:46So we are taking K is equal to one here
  9017. 5:27:48basically as of now. We will see the
  9018. 5:27:50results based on that and then we will
  9019. 5:27:52change it and see how the results vary.
  9020. 5:27:56And we now
  9021. 5:27:59run it on the We now try to predict it.
  9022. 5:28:03And then we will now try to evaluate
  9023. 5:28:06what the results look like.
  9024. 5:28:08So, we have imported the classification
  9025. 5:28:10report and confusion matrix,
  9026. 5:28:12which is basically trying to see whether
  9027. 5:28:15we were able to correctly classify the
  9028. 5:28:18cancerous as cancerous and non-cancerous
  9029. 5:28:20as non-cancerous as or not.
  9030. 5:28:22So, we can see that this is the actual
  9031. 5:28:24and this is the predicted. So, basically
  9032. 5:28:26some data points here five and four are
  9033. 5:28:29classified wrongly, otherwise all the
  9034. 5:28:31others are classified well. So, if we
  9035. 5:28:33look at the accuracy calculated accuracy
  9036. 5:28:37actually, so
  9037. 5:28:39we see that the precision, which is true
  9038. 5:28:42alarm, right? Which is basically from
  9039. 5:28:44the cancerous
  9040. 5:28:45samples, how many were you able to
  9041. 5:28:47actually predict as cancerous?
  9042. 5:28:49If we see the accuracy is quite high,
  9043. 5:28:50almost 94 95%.
  9044. 5:28:54And then the recall, which is from all
  9045. 5:28:57of the cancerous samples, how many were
  9046. 5:28:59you able to actually predict accurately
  9047. 5:29:01is about again 94 95%. And F1 score is
  9048. 5:29:04nothing but a combination of both
  9049. 5:29:06precision as well as recall. And that's
  9050. 5:29:08quite good as well. So, with K is equal
  9051. 5:29:10to one, you're able to get some already
  9052. 5:29:12some good results. Now, let's try to see
  9053. 5:29:15how to choose the K value, right? So,
  9054. 5:29:17this is basically nothing but a
  9055. 5:29:20a bunch of code that actually runs the K
  9056. 5:29:23value from one to 40 and then tries to
  9057. 5:29:26check the accuracy.
  9058. 5:29:28And this is just like doing a trial and
  9059. 5:29:30error to see where we get the best
  9060. 5:29:32results so that we can then use the best
  9061. 5:29:34K value. So, if you can see this
  9062. 5:29:36particular plot here after we plot the
  9063. 5:29:38result from the running the trial and
  9064. 5:29:41error from one to 40, we see that the
  9065. 5:29:43error actually starts decreasing and
  9066. 5:29:45somewhere around this K is going to 21,
  9067. 5:29:48we get the minimum value of error. So,
  9068. 5:29:50for us, the best K value is
  9069. 5:29:5221. So, now if we compare the results
  9070. 5:29:55between K is equal to 1 and K is equal
  9071. 5:29:57to 21, we we should be able to see the
  9072. 5:29:59prediction results. So, as you can see,
  9073. 5:30:00this was the result with K is equal to
  9074. 5:30:021, which is
  9075. 5:30:03we get about 94 95% accuracy.
  9076. 5:30:06And then with K is equal to 21,
  9077. 5:30:09we will see whether the accuracy
  9078. 5:30:10improves, right? So, we see that yes,
  9079. 5:30:12the accuracy has now gone up to almost
  9080. 5:30:1599%, which is we earlier had nine
  9081. 5:30:19misclassified data points
  9082. 5:30:21out of all of the points. And then here
  9083. 5:30:24we have just two data points which are
  9084. 5:30:26misclassified from the test data set.
  9085. 5:30:28So, now this was one example of applying
  9086. 5:30:31KNN on the cancer data set, which is
  9087. 5:30:33available freely.
  9088. 5:30:35And now let's look at another example,
  9089. 5:30:37which is the Iris data set.
  9090. 5:30:39And
  9091. 5:30:40again, available freely as well.
  9092. 5:30:43So, Iris is a type of flower.
  9093. 5:30:45And we will see what flower is it. So,
  9094. 5:30:48just give me a moment here. So, we again
  9095. 5:30:49start by importing
  9096. 5:30:51the necessary libraries. And then
  9097. 5:30:54we'll look at what this Iris data set
  9098. 5:30:57is. So, the Iris data set comprises of
  9099. 5:31:0050 samples of three species of Iris
  9100. 5:31:03flower, which is Iris
  9101. 5:31:04setosa, Iris virginica, and Iris
  9102. 5:31:06versicolor. These are
  9103. 5:31:08the three types of Iris flowers. And if
  9104. 5:31:10you run this, basically we will find
  9105. 5:31:12that
  9106. 5:31:16Okay.
  9107. 5:31:17Sorry, we did not copy the entire the
  9108. 5:31:20code here. So,
  9109. 5:31:21it was giving an issue.
  9110. 5:31:23Let's just run it again.
  9111. 5:31:28Okay. So, we see that it is this
  9112. 5:31:29particular flower, which is Iris setosa.
  9113. 5:31:32So, the Iris setosa, you can see.
  9114. 5:31:35And
  9115. 5:31:36now let's look at the other two kinds of
  9116. 5:31:38flowers here. So, which is
  9117. 5:31:40Iris versicolor.
  9118. 5:31:44So, this is Iris versicolor. And then
  9119. 5:31:47you have the Iris virginica.
  9120. 5:31:55So, we see that this one is Iris
  9121. 5:31:57virginicas. So, essentially we now will
  9122. 5:32:01import the sort of data set. We have
  9123. 5:32:04already done that. And we will now use
  9124. 5:32:06the seaborn package to actually plot
  9125. 5:32:09some of this data and see
  9126. 5:32:12how it looks like.
  9127. 5:32:14Which is basically do some kind of
  9128. 5:32:15exploratory data analysis. So, if we
  9129. 5:32:18look at the data itself, so this is the
  9130. 5:32:20top five rows from the data, right? So,
  9131. 5:32:22essentially the data consists of
  9132. 5:32:25basically what is the sepal length,
  9133. 5:32:27sepal width,
  9134. 5:32:28petal length, and petal width.
  9135. 5:32:30And this is nothing but basically your
  9136. 5:32:33petal is your this colored part of the
  9137. 5:32:35flower and the sepal is basically your
  9138. 5:32:37green part here, right? So, it is
  9139. 5:32:39talking about what what is the sepal
  9140. 5:32:40length, sepal width, and petal length,
  9141. 5:32:42and petal width of each of the species,
  9142. 5:32:44whether it is setosa, virginica, or
  9143. 5:32:47versicolor. And we have we will see
  9144. 5:32:49whether we can use KNN to actually
  9145. 5:32:52classify these
  9146. 5:32:54these flowers into the data points into
  9147. 5:32:56these categories of flowers.
  9148. 5:32:58So, let's do some quick exploratory data
  9149. 5:33:01analysis
  9150. 5:33:02on this. So, we are running a pair plot
  9151. 5:33:05on the data set. And uh
  9152. 5:33:08in the meantime, I'll just copy another
  9153. 5:33:11part of the code here.
  9154. 5:33:13Okay, so now the plot has come up. So,
  9155. 5:33:16as we can see that the green is
  9156. 5:33:18basically your setosa flower. And
  9157. 5:33:22we can see that this pair plot actually
  9158. 5:33:24just plots
  9159. 5:33:25the sepal length, sepal width,
  9160. 5:33:28and petal length, petal width of each of
  9161. 5:33:30the samples of setosa, verse- color and
  9162. 5:33:33virginica. And we see that
  9163. 5:33:34the green dots which are the setosa
  9164. 5:33:36flower is actually quite separable from
  9165. 5:33:38the others. It's when you plot
  9166. 5:33:40let let's say for example sepal length
  9167. 5:33:42and
  9168. 5:33:43petal length, right? We see that this is
  9169. 5:33:45quite separate from the other data
  9170. 5:33:47points. So, let's see whether you know,
  9171. 5:33:49we can actually
  9172. 5:33:50be able to classify it using KNN
  9173. 5:33:53or not. And here we are running a kernel
  9174. 5:33:56density estimation function on the
  9175. 5:33:58setosa flower to check
  9176. 5:34:01what kind of distribution it has. So,
  9177. 5:34:03this is the kernel density estimation
  9178. 5:34:06plot using the SNS package. So, only for
  9179. 5:34:09the setosa flower. So, if we plot the
  9180. 5:34:12sepal length and sepal width, we get
  9181. 5:34:13something distribution like this. So,
  9182. 5:34:15essentially we see that the maximum
  9183. 5:34:17centered around here and then there is a
  9184. 5:34:19distribution as you can see here. So,
  9185. 5:34:21there's some kind of a linear
  9186. 5:34:22relationship here.
  9187. 5:34:24Okay, so now we will again do the same
  9188. 5:34:27standardization of the variables
  9189. 5:34:30that we had done in the cancer data set
  9190. 5:34:32case. So, we are importing the standard
  9191. 5:34:34scalar function
  9192. 5:34:36from the scikit-learn preprocessing. So,
  9193. 5:34:39we will initialize that.
  9194. 5:34:41So, we are again basically doing the
  9195. 5:34:42standardization or normalization of the
  9196. 5:34:45data.
  9197. 5:34:46And we will do the standardization on
  9198. 5:34:48everything except the species which is a
  9199. 5:34:50categorical value, right? So, it's
  9200. 5:34:52categorical whether it is which kind of
  9201. 5:34:54flower it is. So, we have removed that
  9202. 5:34:56and then we have
  9203. 5:34:58done the standardization or
  9204. 5:35:00normalization on rest of the data.
  9205. 5:35:02So, we now convert this into a data
  9206. 5:35:05frame, pandas data frame. And if we look
  9207. 5:35:08at the top five rows, now it's all
  9208. 5:35:09converted or transformed. So, the values
  9209. 5:35:12are now normally distributed basically.
  9210. 5:35:15Okay, so now we
  9211. 5:35:18divide the data again into train and
  9212. 5:35:20test.
  9213. 5:35:22And we again have training of about 70%
  9214. 5:35:27and test data set of about 30%. Um so we
  9215. 5:35:31are dividing that entire data set into
  9216. 5:35:32these two buckets.
  9217. 5:35:34And we will now use KNN
  9218. 5:35:38to see if we can use KNN to classify
  9219. 5:35:40them.
  9220. 5:35:42Again, the same and K is equal to 1.
  9221. 5:35:46And we will check the results and then
  9222. 5:35:47we will do a trial and error
  9223. 5:35:50to check what is the best value of K.
  9224. 5:35:52So here K is equal to 1.
  9225. 5:35:57And now we are going to predict on the
  9226. 5:36:00test data set
  9227. 5:36:02and look at the results.
  9228. 5:36:05So we are now importing the
  9229. 5:36:06classification report and the confusion
  9230. 5:36:08matrix.
  9231. 5:36:12So if you look at the confusion matrix,
  9232. 5:36:18we see that
  9233. 5:36:19kind of already we are getting quite
  9234. 5:36:21good
  9235. 5:36:22uh prediction. So just two misclassified
  9236. 5:36:24points.
  9237. 5:36:27And uh if we look at the accuracy,
  9238. 5:36:32we see that the accuracy is quite high,
  9239. 5:36:34around 96%
  9240. 5:36:37already.
  9241. 5:36:38Now we choose we have to see what is the
  9242. 5:36:41best value of K.
  9243. 5:36:43So
  9244. 5:36:45essentially we will
  9245. 5:36:49we will again run K is equal to 1 to 40
  9246. 5:36:52and check which is the best value.
  9247. 5:36:56So let's plot the errors when we vary
  9248. 5:36:59the K from 1 to 40. And we see that the
  9249. 5:37:02actually the error decreases and then
  9250. 5:37:04increases. So basically the error is
  9251. 5:37:07minimum with K is equal to let's say
  9252. 5:37:08three or even five or maybe 11. So let's
  9253. 5:37:12choose one of these values. So let's say
  9254. 5:37:14K is equal to three.
  9255. 5:37:16And let's see how the results look like.
  9256. 5:37:19Does it improve the
  9257. 5:37:21accuracy or not? So we now see that
  9258. 5:37:24even the two data points which are
  9259. 5:37:25misclassified earlier is now classified
  9260. 5:37:27properly. So, the accuracy improves to
  9261. 5:37:29100%.
  9262. 5:37:31So, that's the
  9263. 5:37:32example of how you can choose K.
  9264. 5:37:36>> [music]
  9265. 5:37:40>> What is Naive Bayes?
  9266. 5:37:42Let us understand Naive Bayes with an
  9267. 5:37:44example. Here, I just cannot seem to
  9268. 5:37:47figure out which are the best days to
  9269. 5:37:49play football with my friend. Can you
  9270. 5:37:52please help us out?
  9271. 5:37:54All possible conditions are given to us.
  9272. 5:37:56There is
  9273. 5:37:58summer, monsoon, and winter, which is
  9274. 5:38:00nothing but the outlook.
  9275. 5:38:03Am I correct in saying that?
  9276. 5:38:04Summer, monsoon, and winter is nothing
  9277. 5:38:06but the outlook. Then we have sunny or
  9278. 5:38:09not sunny. So, that basically is the
  9279. 5:38:12humidity, right? And then we have windy
  9280. 5:38:15or no windy. That speaks about
  9281. 5:38:18the winds. How are the winds?
  9282. 5:38:20Right?
  9283. 5:38:21So, if you look at these combinations,
  9284. 5:38:24okay, we will look at this using Naive
  9285. 5:38:25Bayes on how do we decide whether we can
  9286. 5:38:28play or not.
  9287. 5:38:29So, if I have noted down all the days it
  9288. 5:38:31was good, bad to play football, and the
  9289. 5:38:33combination of weather matrices on that
  9290. 5:38:35day, that will be perfect, right? That
  9291. 5:38:37is perfect, and we will be able to do
  9292. 5:38:39Naive Bayes classifiers using that. Now,
  9293. 5:38:42Naive Bayes classifier comes from the
  9294. 5:38:44Naive Bayes theorem, and Naive Bayes
  9295. 5:38:46theorem is purely and purely based on
  9296. 5:38:50the assumption of independence.
  9297. 5:38:53So, what does it mean? When I say
  9298. 5:38:55independence, what it means is that
  9299. 5:38:59this variable has no relationship, no
  9300. 5:39:03association with this variable. Now,
  9301. 5:39:06when I talk about this in a linear
  9302. 5:39:09context, obviously in summer, you will
  9303. 5:39:12see that we have more sunny days.
  9304. 5:39:16Yes? So, if you look at it from the
  9305. 5:39:18correlation area, from the linear
  9306. 5:39:21algebra concepts, linearly these two are
  9307. 5:39:25correlated to each other. Am I correct
  9308. 5:39:28in saying that?
  9309. 5:39:29Obviously, in monsoon we have less sunny
  9310. 5:39:32days. In winter we further have less
  9311. 5:39:34sunny days.
  9312. 5:39:36Yeah?
  9313. 5:39:36So,
  9314. 5:39:37although there is a relation,
  9315. 5:39:39Naive Bayes theorem
  9316. 5:39:42says that all these variables are
  9317. 5:39:45independent of each other.
  9318. 5:39:49What does the that mean? If this is
  9319. 5:39:51causing any kind of an effect, if this
  9320. 5:39:54is causing any kind of an impact,
  9321. 5:39:57this should not matter.
  9322. 5:40:00Okay? There's going to be no
  9323. 5:40:01relationship. Summer, monsoon, winter,
  9324. 5:40:04it has its own weightage. Okay? And it
  9325. 5:40:07has nothing to do with the other
  9326. 5:40:09conditions. Every condition is equally
  9327. 5:40:12significant.
  9328. 5:40:14All right? So, what happens in Naive
  9329. 5:40:15Bayes is we estimate the posterior
  9330. 5:40:18probability of every event happening.
  9331. 5:40:21Here, we calculate the posterior
  9332. 5:40:24probability of an event happening.
  9333. 5:40:27Okay? So, here if you see, if you look
  9334. 5:40:29at the sunny conditions, what we have
  9335. 5:40:31done, sunny conditions we have this
  9336. 5:40:33distribution. That there is no play
  9337. 5:40:35happening in summer, there is play
  9338. 5:40:36happening in monsoon, there is play
  9339. 5:40:38happening in winter.
  9340. 5:40:39Right? So, base of the season, we are
  9341. 5:40:42figuring that out. Similarly for windy
  9342. 5:40:44conditions, we are doing that.
  9343. 5:40:46Right? Again, then we do it for a
  9344. 5:40:49combination.
  9345. 5:40:50Right? Whether when windy conditions are
  9346. 5:40:53yes and no, what happens to play?
  9347. 5:40:56Okay? So, here
  9348. 5:40:57what at the end of the day, what gets
  9349. 5:41:00selected is the one which has a
  9350. 5:41:02posterior probability of greater than
  9351. 5:41:04five. Now, when I talk about posterior
  9352. 5:41:07probability,
  9353. 5:41:08what do I mean by posterior probability?
  9354. 5:41:10Let me have a blank slate. Here you go.
  9355. 5:41:13Actually, it's given. So, I need not
  9356. 5:41:15show you that. What is the simplistic
  9357. 5:41:17probabilistic classifier here? What is
  9358. 5:41:19the probability of an event A happening
  9359. 5:41:22given
  9360. 5:41:23B.
  9361. 5:41:24So, if you look at our problem context,
  9362. 5:41:27what is it that we are trying to figure
  9363. 5:41:29out?
  9364. 5:41:29We are trying to figure out what is the
  9365. 5:41:32probability of
  9366. 5:41:36play happening
  9367. 5:41:38given
  9368. 5:41:45the outlook is sunny,
  9369. 5:41:50{comma}
  9370. 5:41:53the
  9371. 5:41:54uh
  9372. 5:41:57winds
  9373. 5:42:03are normal,
  9374. 5:42:07and there is no rain.
  9375. 5:42:13On any given day,
  9376. 5:42:15on any given day when there is no rain,
  9377. 5:42:19there is no wind,
  9378. 5:42:21and the outlook is sunny,
  9379. 5:42:23whether play will happen or not. So,
  9380. 5:42:26what we end up doing is we calculate the
  9381. 5:42:28posterior probability of play happening
  9382. 5:42:30given these conditions. We also
  9383. 5:42:32calculate the posterior probability of
  9384. 5:42:35play not happening
  9385. 5:42:38given these conditions, and then we
  9386. 5:42:40normalize these probabilities.
  9387. 5:42:43Mathematically, we do all of these
  9388. 5:42:45calculations to figure out naive Bayes.
  9389. 5:42:47Okay? Now, here
  9390. 5:42:49see,
  9391. 5:42:50here we are talking about one event.
  9392. 5:42:53Here, we have three events. We have
  9393. 5:42:55outlook,
  9394. 5:42:57we have winds,
  9395. 5:42:59and we have rains.
  9396. 5:43:00So, what this becomes is
  9397. 5:43:04probability of three independent events.
  9398. 5:43:07So, what we will do, we figure out what
  9399. 5:43:10is the probability of play happening
  9400. 5:43:16given
  9401. 5:43:18outlook is sunny.
  9402. 5:43:22We also figure out
  9403. 5:43:26what is the probability of
  9404. 5:43:29play happening.
  9405. 5:43:34Multiply this with the probability of
  9406. 5:43:37wind as no.
  9407. 5:43:41Given no, then probability of play
  9408. 5:43:44happening
  9409. 5:43:49given
  9410. 5:43:51rain is no.
  9411. 5:43:53We figure out all the in three
  9412. 5:43:56independent probabilities multiplied.
  9413. 5:43:59All right? This is what we do
  9414. 5:44:01mathematically.
  9415. 5:44:03This is what is done mathematically in
  9416. 5:44:06Naive Bayes theorem.
  9417. 5:44:08Ultimately, the posterior probability
  9418. 5:44:12the posterior probability probability of
  9419. 5:44:14an event A happening given
  9420. 5:44:17B conditions is calculated by first
  9421. 5:44:21calculating the class probability.
  9422. 5:44:24What is the class probability here?
  9423. 5:44:25Probability of it raining given play was
  9424. 5:44:29happening in that day.
  9425. 5:44:31This is multiplied by the total
  9426. 5:44:33probability of play happening and
  9427. 5:44:35divided by the total probability of
  9428. 5:44:37sunny conditions.
  9429. 5:44:41Here, as you see, this is the posterior
  9430. 5:44:43probability.
  9431. 5:44:44This is the class probability.
  9432. 5:44:47Here is the predictors probability. This
  9433. 5:44:49is the outlook. X is the outlook. C is
  9434. 5:44:53what we are trying to predict.
  9435. 5:44:55Right?
  9436. 5:44:56And here is the likelihood.
  9437. 5:44:59So, how is this formulated into our
  9438. 5:45:01table? Now, if you correlate this to our
  9439. 5:45:03graph,
  9440. 5:45:04what will be the likelihood?
  9441. 5:45:07What will be the likelihood of sunny
  9442. 5:45:10conditions given play happens?
  9443. 5:45:13Sunny conditions play happens.
  9444. 5:45:152 / 3
  9445. 5:45:18out of
  9446. 5:45:19Sorry, how many sunny conditions do we
  9447. 5:45:21have? Six
  9448. 5:45:22conditions.
  9449. 5:45:24In six con-
  9450. 5:45:25ditions, how many days does play happen?
  9451. 5:45:27Two days.
  9452. 5:45:292 / 6 1 / 3. What is the class
  9453. 5:45:32probability? So, of all the events that
  9454. 5:45:34are given to us, how many days does play
  9455. 5:45:36happen?
  9456. 5:45:37Okay? This is how this is calculated.
  9457. 5:45:40So, here you see what we have done is
  9458. 5:45:43Let me go back to the previous slide.
  9459. 5:45:45Here you go.
  9460. 5:45:46Okay? Here, what we have done is we have
  9461. 5:45:48calculated this table. Now, it speaks
  9462. 5:45:51about a data set. Where can you get this
  9463. 5:45:54data set?
  9464. 5:45:55You can look for golf play days data set
  9465. 5:45:59online. You can look for golf play days
  9466. 5:46:02data set.
  9467. 5:46:03Okay?
  9468. 5:46:04In this data set, you will find all the
  9469. 5:46:07data
  9470. 5:46:08which is required for this particular
  9471. 5:46:11example to be done. So, what I will be
  9472. 5:46:13doing is I will be doing this example,
  9473. 5:46:15this Naive Bayes classification, with
  9474. 5:46:18you in Python. All right? So, whatever
  9475. 5:46:20calculations are being done here,
  9476. 5:46:23okay? I will do the same activity in
  9477. 5:46:26Python with pen and paper. And instead
  9478. 5:46:28of doing this
  9479. 5:46:32in a numeric way where I'm doing lot of
  9480. 5:46:34probabilistic calculations,
  9481. 5:46:37I will achieve this simply in
  9482. 5:46:40very limited lines of code.
  9483. 5:46:43Very limited lines of code with Python.
  9484. 5:46:47Once I'm I have done that, I will come
  9485. 5:46:49and explain this prob- probability table
  9486. 5:46:52to all of us.
  9487. 5:46:53I'm going to use some basic libraries.
  9488. 5:46:56All right. So, here, what I will be
  9489. 5:46:58doing is using some very basic libraries
  9490. 5:47:01for this activity. All right?
  9491. 5:47:57Done. Now, let me quickly go and uh
  9492. 5:48:01read the data set. So, for that what I
  9493. 5:48:04will do is quickly
  9494. 5:48:08change my working directory.
  9495. 5:48:43And now, let me quickly go and read my
  9496. 5:48:44data. So, my data frame is pd.
  9497. 5:49:09>> Here you go. This is my data set.
  9498. 5:49:12Right? So, if you look at this data set,
  9499. 5:49:14in this data set, you have 13 14 days.
  9500. 5:49:18In these 14 days, you have the outlook,
  9501. 5:49:21overcast,
  9502. 5:49:22rainy, and sunny. You have temperature,
  9503. 5:49:26temperature is hot, cool, and mild. You
  9504. 5:49:29have humidity, you have wind, and you
  9505. 5:49:32have play. Right? So, I will not be
  9506. 5:49:35using uh
  9507. 5:49:36Okay, let us use all four. In this
  9508. 5:49:39example, they're using only three
  9509. 5:49:40variables, but in our
  9510. 5:49:43hands-on, okay? In this hands-on, what I
  9511. 5:49:46will be doing is I will be uh
  9512. 5:49:49using all four variables. Let us do
  9513. 5:49:51that.
  9514. 5:49:52Okay?
  9515. 5:49:53So, before I do that, let me convert
  9516. 5:49:56everything into a category.
  9517. 5:49:58If you look at your data frame right
  9518. 5:49:59now, it's not everything is not into a
  9519. 5:50:02categorical variable.
  9520. 5:50:04Here you go, see.
  9521. 5:50:05Okay? So, let me quickly go and convert
  9522. 5:50:07everything into a category.
  9523. 5:50:26And once I have done this, uh
  9524. 5:50:29let me create a new data frame in which
  9525. 5:50:31I have everything as a category code.
  9526. 5:50:35So, that I have numbers. I'll show you
  9527. 5:50:37what What do I mean by this?
  9528. 5:50:52>> So, let me execute this. Here you go.
  9529. 5:50:55See, now I have two data frames. In the
  9530. 5:50:57first data frame I have all these
  9531. 5:50:59values. These are now categorical
  9532. 5:51:01variables, but in the second data frame
  9533. 5:51:03I have all ones and zeros. So, wherever
  9534. 5:51:06you see there is sunny conditions, now I
  9535. 5:51:08have a code two.
  9536. 5:51:09Rainy conditions, code one.
  9537. 5:51:11Similarly, when play happens I have a
  9538. 5:51:15one. When play does not happen I have a
  9539. 5:51:17zero.
  9540. 5:51:19This is what I have done.
  9541. 5:51:21This data frame is available online.
  9542. 5:51:24Okay? You can get this data frame
  9543. 5:51:27online.
  9544. 5:51:29All right? Now, my data frame is ready.
  9545. 5:51:32So, now what I'm going to do is I'm
  9546. 5:51:34going to divide my data frame
  9547. 5:51:37into training and testing. I have 14
  9548. 5:51:39records.
  9549. 5:51:41So, let's take 10 records for training.
  9550. 5:51:44I will give 10 records as an input.
  9551. 5:51:47And I will give four records, last four
  9552. 5:51:49records
  9553. 5:51:57as my test data frame. All right?
  9554. 5:52:00Uh
  9555. 5:52:00so, now I will need to create my X and
  9556. 5:52:03my Y.
  9557. 5:52:04So, how I will do that is
  9558. 5:52:06I'll say Y {underscore} train
  9559. 5:52:09is equal to from train
  9560. 5:52:12I don't want the play variable.
  9561. 5:52:14That play variable should be my Y. As
  9562. 5:52:16simple as that.
  9563. 5:52:18And I will say X {underscore} train is
  9564. 5:52:21equal to train.
  9565. 5:52:23Okay? And I will do the same thing for
  9566. 5:52:25my test data frame also.
  9567. 5:52:29Now, those who are new to Python will
  9568. 5:52:32find this a little bit strange. Please
  9569. 5:52:36bear with me.
  9570. 5:52:37But these are the only calculations,
  9571. 5:52:39only steps which need to be performed
  9572. 5:52:41every time.
  9573. 5:52:42You are trying to achieve maybe bias
  9574. 5:52:45algorithm or any kind of an algorithm.
  9575. 5:52:49All right. So, here now you see this is
  9576. 5:52:52my training data frame in which play
  9577. 5:52:54variable is not there.
  9578. 5:52:56Play variable is not there. This is my Y
  9579. 5:53:00in which only play variable is there.
  9580. 5:53:02This is my training data set. So, both
  9581. 5:53:04of them have 10 records with the
  9582. 5:53:06matching index.
  9583. 5:53:08Similarly, test data frame four records
  9584. 5:53:12four records with the matching index.
  9585. 5:53:15Right? So, that we know which data frame
  9586. 5:53:18is where.
  9587. 5:53:20Now, multinomial naive bias. Very
  9588. 5:53:23simple, three lines of code and my model
  9589. 5:53:26will be done.
  9590. 5:53:28Okay? First, I initialize my model.
  9591. 5:53:32Here you go. I have initialized my
  9592. 5:53:33model.
  9593. 5:53:34In this model, I fit my data.
  9594. 5:53:39In this model, I will fit my data. So,
  9595. 5:53:42to do that, what I say is fit
  9596. 5:53:45X underscore
  9597. 5:53:50train comma
  9598. 5:53:53comma Y underscore train.
  9599. 5:53:56Done.
  9600. 5:53:57Your model object is now ready.
  9601. 5:54:00And now you can simply get the
  9602. 5:54:03classification outcomes. So, we have in
  9603. 5:54:06our
  9604. 5:54:07test data frame, if you look at our test
  9605. 5:54:09data frame, this is our test data frame.
  9606. 5:54:11We have three four conditions. All four
  9607. 5:54:14are sunny,
  9608. 5:54:16high temperature, low humidity, and
  9609. 5:54:19windy. Right? And if you look at their
  9610. 5:54:22outcomes, these are their outcomes.
  9611. 5:54:24On the first two days, play is not
  9612. 5:54:26happening. On the next two days, play is
  9613. 5:54:28happening. Let us look at what is the
  9614. 5:54:30prediction of our model for this. So, to
  9615. 5:54:33do do that, what I simply do is
  9616. 5:54:37X out is equal to
  9617. 5:54:40model.predict
  9618. 5:54:46To this I give my X {underscore} test.
  9619. 5:54:50Here you go. Okay? Now you have your Y Y
  9620. 5:54:54out variable, so this is the prediction
  9621. 5:54:56for all the
  9622. 5:54:59four inputs that you give, and this is
  9623. 5:55:00the prediction.
  9624. 5:55:02First day
  9625. 5:55:03first day we say play does not happen.
  9626. 5:55:06Let us match it.
  9627. 5:55:08Let us try to match it with our
  9628. 5:55:10here.
  9629. 5:55:11See?
  9630. 5:55:12Out of four records, three records we
  9631. 5:55:14are predicting correctly.
  9632. 5:55:16Three records we are predicting
  9633. 5:55:18correctly. If you want to check the
  9634. 5:55:20accuracy, what is the accuracy of your
  9635. 5:55:23model? What you can simply do is print
  9636. 5:55:25Let us print the accuracy on both
  9637. 5:55:27training and testing.
  9638. 5:55:34Training accuracy. How do I get the
  9639. 5:55:36training accuracy? Very simple. model.
  9640. 5:55:39score
  9641. 5:55:43And here I give my X {underscore} train
  9642. 5:55:47{comma} Y {underscore} train.
  9643. 5:55:50And then we do the
  9644. 5:55:52testing accuracy also.
  9645. 5:56:06Here you go. So here you can see for our
  9646. 5:56:10model training we have 80% accuracy, and
  9647. 5:56:13for testing we have 75% accuracy. Okay?
  9648. 5:56:18So this is the advantage of doing this
  9649. 5:56:20activity in Python. But what is
  9650. 5:56:22happening in the back end?
  9651. 5:56:23Now let us go and also understand that
  9652. 5:56:25in terms of naive Bayes classifier.
  9653. 5:56:28We have successfully
  9654. 5:56:30we have successfully implemented the
  9655. 5:56:33Naive Bayes classifier in Python
  9656. 5:56:35programming language. But, here let us
  9657. 5:56:38try to understand Bayes theorem, what is
  9658. 5:56:41happening. So, from this data set all
  9659. 5:56:43the tabulated data frequency tables are
  9660. 5:56:45calculated.
  9661. 5:56:47Once the frequency tables are
  9662. 5:56:48calculated, they are substituted in our
  9663. 5:56:51formula to calculate the probabilistic
  9664. 5:56:53scores. So, what is the probability of
  9665. 5:56:56summer given it is playing conditions?
  9666. 5:57:00Total how many playing conditions are
  9667. 5:57:02there? Total there are nine playing
  9668. 5:57:03conditions. Nine days play happened.
  9669. 5:57:06That becomes our denominator. Out of
  9670. 5:57:08those days, how many days was summer is
  9671. 5:57:10our numerator. That is how for this we
  9672. 5:57:13get a probability of 0.33.
  9673. 5:57:16Then we calculate the class probability
  9674. 5:57:18where we look at how many days was it
  9675. 5:57:20summer? Out of total 14 days, five days
  9676. 5:57:23was summer, so that's the probability
  9677. 5:57:25and the class probability is 0.64. Put
  9678. 5:57:28everything into our equation.
  9679. 5:57:30Put everything into our equation and
  9680. 5:57:32this is what we get.
  9681. 5:57:34Okay? So, we do this for each and every
  9682. 5:57:37condition.
  9683. 5:57:38So, here we calculate it for winter.
  9684. 5:57:42All right? Once we have done it for all
  9685. 5:57:44three days,
  9686. 5:57:47winter, sunny and windy days,
  9687. 5:57:49we substitute those here
  9688. 5:57:51and that gives us the probability which
  9689. 5:57:53is
  9690. 5:57:54more than 0.5. Thus, now we can say that
  9691. 5:57:58if
  9692. 5:57:59it is winter, sunny and conditions are
  9693. 5:58:02sunny and there are winds. Conditions
  9694. 5:58:04are not sunny and there are winds. Play
  9695. 5:58:07can happen.
  9696. 5:58:09Look at another example. If a single
  9697. 5:58:11card is drawn from a standard deck of
  9698. 5:58:14playing cards, the probability that card
  9699. 5:58:16is a king is 4/52
  9700. 5:58:18since there are four kings in a standard
  9701. 5:58:21deck.
  9702. 5:58:22King is the event. This card is a king.
  9703. 5:58:25This is the event. The prior probability
  9704. 5:58:28of this is 1 by 13. If evidence is
  9705. 5:58:31provided, for instance, someone looks at
  9706. 5:58:33the card that the single card is a face
  9707. 5:58:35card, then the posterior probability can
  9708. 5:58:38be calculated using Bayes' theorem.
  9709. 5:58:47Okay? Since every king is also a face
  9710. 5:58:49card, the probability of face happening
  9711. 5:58:52given you getting a face card given it's
  9712. 5:58:54a king is one. Since there are three
  9713. 5:58:57face cards in each suit,
  9714. 5:58:59all right? It's actually four. Ace is
  9715. 5:59:01also a face card. So, it's jack, king,
  9716. 5:59:04queen, and uh ace. The probability of
  9717. 5:59:06the face card is 4 by 13.
  9718. 5:59:08Okay? So, if you combine these three
  9719. 5:59:10likelihoods, what you get is 13 by 4.
  9720. 5:59:13So, using Bayes' theorem, this is the
  9721. 5:59:15probability that you get.
  9722. 5:59:18>> [music]
  9723. 5:59:23>> What is support vector machine?
  9724. 5:59:25Support vector machine comes under
  9725. 5:59:27supervised machine learning.
  9726. 5:59:30And we use it specifically for
  9727. 5:59:32performing the task of classification.
  9728. 5:59:35So, support vector machine is a
  9729. 5:59:36discriminative classifier
  9730. 5:59:38that is formally designed by a separate
  9731. 5:59:41hyperplane.
  9732. 5:59:43Okay? It is a representation of examples
  9733. 5:59:45as points in a space that are mapped so
  9734. 5:59:48that the points of different categories
  9735. 5:59:50are separated by a gap as wide as
  9736. 5:59:53possible.
  9737. 5:59:55So, in this case of the support vector
  9738. 5:59:56machine,
  9739. 5:59:57let's say I have some data points. So,
  9740. 6:00:00there are some data points of X, and
  9741. 6:00:02there are data points of circle.
  9742. 6:00:05Now, this support vector machine
  9743. 6:00:07is a type of machine learning algorithm
  9744. 6:00:10where if I have the collection of
  9745. 6:00:12points, so here in this data points, I
  9746. 6:00:14have two classes. One is X, and the
  9747. 6:00:17another one is circle.
  9748. 6:00:19Now, given this kind of data points,
  9749. 6:00:22okay? Given this kind of binary
  9750. 6:00:23classification problem,
  9751. 6:00:26so the expectation is
  9752. 6:00:29in case of support vector machine, I'm
  9753. 6:00:31going to draw a hyperplane
  9754. 6:00:33which separates as much as possible.
  9755. 6:00:37Okay? So, I'm going to draw a hyperplane
  9756. 6:00:40which separates these two classes as
  9757. 6:00:43much as possible.
  9758. 6:00:46So, it says that the I'm going to draw
  9759. 6:00:49draw hyperplane
  9760. 6:00:51and it it will be separated by a gap as
  9761. 6:00:54wide as possible.
  9762. 6:00:56So, that is the intuition behind support
  9763. 6:00:58vector machine.
  9764. 6:01:00Okay. Now that you have an intuition
  9765. 6:01:02behind what is support vector machine,
  9766. 6:01:05let's understand as how does this SVM,
  9767. 6:01:08that is support vector machine, would
  9768. 6:01:09work.
  9769. 6:01:10So, in case of support vector machine,
  9770. 6:01:13so here there is one more example. I
  9771. 6:01:15have the set of
  9772. 6:01:17points which is green green color and I
  9773. 6:01:19have another set of points which are in
  9774. 6:01:21red color. So, these two
  9775. 6:01:24points are belonging to the different
  9776. 6:01:26different classes.
  9777. 6:01:28Now, what I'm going to do is I'm going
  9778. 6:01:30to draw a hyperplane which separates
  9779. 6:01:34these two classes data points as much as
  9780. 6:01:36possible. And when I'm drawing the
  9781. 6:01:38hyperplane, I'll make sure that this
  9782. 6:01:40hyperplane is
  9783. 6:01:42as this hyperplane is equidistant from
  9784. 6:01:46my support vectors.
  9785. 6:01:48Now, the support vectors is nothing but
  9786. 6:01:51the point which is closer to my
  9787. 6:01:52hyperplane.
  9788. 6:01:54Now, here in this example that you're
  9789. 6:01:55seeing,
  9790. 6:01:57the this data point and this data point
  9791. 6:02:01are called as support vectors because
  9792. 6:02:03these are the data points which are
  9793. 6:02:05nearest from my hyperplane that I've
  9794. 6:02:06just drawn.
  9795. 6:02:09In if I'm trying to make use of this SVM
  9796. 6:02:12model, it is going to draw this kind of
  9797. 6:02:14hyperplane
  9798. 6:02:16to make sure that it is separating two
  9799. 6:02:18classes. The two classes that we have
  9800. 6:02:20over here in this example is red and
  9801. 6:02:22green. It's going to separate these two
  9802. 6:02:24classes as much as possible and it will
  9803. 6:02:28be equidistant from my support vectors
  9804. 6:02:31and the support vectors are nothing but
  9805. 6:02:33the nearest point to my hyperplane.
  9806. 6:02:36And that is how I'm going to separate
  9807. 6:02:39between two classes when it comes to
  9808. 6:02:40support vector machines.
  9809. 6:02:44Now, here in this example that you're
  9810. 6:02:46currently seeing,
  9811. 6:02:47the hyperplane that I've just drawn, so
  9812. 6:02:49this is a simple linear hyperplane.
  9813. 6:02:53Just like a straight line that I'm
  9814. 6:02:54trying to draw if I want to separate two
  9815. 6:02:57classes of data points.
  9816. 6:02:59Now, apart from drawing this straight
  9817. 6:03:01line, we also have other kind of lines
  9818. 6:03:05as well which we can draw.
  9819. 6:03:07So, let's see how we can do that.
  9820. 6:03:11So, the types of line that we can draw
  9821. 6:03:13or the hyperplane that we can draw is
  9822. 6:03:16called as SVM kernels, that is support
  9823. 6:03:18vector machine kernels. The example that
  9824. 6:03:20we have seen, it's an example for linear
  9825. 6:03:23SVM kernels.
  9826. 6:03:25So, let's see what are the other types
  9827. 6:03:27of kernels that we have. So, when it
  9828. 6:03:29comes to SVM kernels, we have linear
  9829. 6:03:31kernels,
  9830. 6:03:33radial basis function kernel and along
  9831. 6:03:35with that, we also have polynomial
  9832. 6:03:38kernel.
  9833. 6:03:40Now, in case of linear kernel, I'm going
  9834. 6:03:42to draw a hyperplane which is like a
  9835. 6:03:44straight line.
  9836. 6:03:45In case of polynomial kernel, I can draw
  9837. 6:03:48my hyperplane on the basis of polynomial
  9838. 6:03:50function that I have created on the
  9839. 6:03:52basis of number of variables that I have
  9840. 6:03:54and the degree that I have over there in
  9841. 6:03:57case of polynomial. And in case of
  9842. 6:03:59radial basis function, so I'll make use
  9843. 6:04:01of radial basis to separate my data
  9844. 6:04:04points.
  9845. 6:04:07Okay, so these three are the important
  9846. 6:04:10kernels that we have in SVM and this is
  9847. 6:04:12one of the commonly asked interview
  9848. 6:04:13question when it comes to the topic of
  9849. 6:04:15support vector machines.
  9850. 6:04:18Now, let's look at some of the use cases
  9851. 6:04:21or the way we can do where we can use
  9852. 6:04:24this SVM to uh
  9853. 6:04:27work or let's look at some of the use
  9854. 6:04:29cases where we can use this SVM.
  9855. 6:04:32Okay.
  9856. 6:04:34So, we can use this SVM
  9857. 6:04:38on many of the use cases. So, to name a
  9858. 6:04:41few, we can use it in face detection.
  9859. 6:04:44We can use it in text and hypertext
  9860. 6:04:46categorization. We can use the SVM if
  9861. 6:04:49I'm trying to classify any images. I can
  9862. 6:04:52make use in bioinformatics.
  9863. 6:04:55And if I'm trying to detect something,
  9864. 6:04:57so I can
  9865. 6:04:59in the in an example here, remote
  9866. 6:05:01homology detection, handwriting
  9867. 6:05:03detection. Or in general, we can make
  9868. 6:05:06use of this generalized predictive
  9869. 6:05:08control. So, wherever we are dealing
  9870. 6:05:10with the task of classification, we can
  9871. 6:05:13use this SVM model. Okay. Now that we
  9872. 6:05:17have a theoretical understanding as what
  9873. 6:05:19is SVM and how it is actually going to
  9874. 6:05:22look like and how it will be,
  9875. 6:05:25let's have a quick walk through as how
  9876. 6:05:28we can implement this SVM. Now, to
  9877. 6:05:31implement this SVM,
  9878. 6:05:33these are the common steps that we are
  9879. 6:05:34going to follow.
  9880. 6:05:36We are going to load the data.
  9881. 6:05:39We'll explore the data.
  9882. 6:05:41And once we have explored the data, we
  9883. 6:05:43are going to split the data into two
  9884. 6:05:45parts. The reason is simple. One, I have
  9885. 6:05:47training, so I'll be making use of my
  9886. 6:05:50training data.
  9887. 6:05:51And once my training is complete, I'll
  9888. 6:05:54check how my model has been trained with
  9889. 6:05:56the help of my test data. So,
  9890. 6:05:58I'm going to split the data.
  9891. 6:06:00Now, once that is complete, we are going
  9892. 6:06:02to train this SVM model. And finally, we
  9893. 6:06:05can evaluate the model and observe as
  9894. 6:06:08how model is working.
  9895. 6:06:11So, this is the overview of the
  9896. 6:06:13implementation of support vector
  9897. 6:06:15machines.
  9898. 6:06:17So, let's do one thing. Let's
  9899. 6:06:21work it out and let's create the
  9900. 6:06:23notebook in Google Colab and let's see
  9901. 6:06:25it in action as how we can implement
  9902. 6:06:27this SVM.
  9903. 6:06:28I'll come back to my Google Colab.
  9904. 6:06:31So, this is the notebook that I have
  9905. 6:06:32already prepared and I'll give you a
  9906. 6:06:35walk-through as we proceed along.
  9907. 6:06:37Now, here in my first cell, I'm
  9908. 6:06:39importing my NumPy library, Pandas
  9909. 6:06:41library, and along with that, for
  9910. 6:06:43creation of plots, I'm importing my
  9911. 6:06:45Matplotlib library. Now, if you're
  9912. 6:06:48comfortable with Seaborn, you can use
  9913. 6:06:50the Seaborn library as well. So, in my
  9914. 6:06:52example, I'm just making use of
  9915. 6:06:54Matplotlib because we are not interested
  9916. 6:06:57in creation of visualization, but we
  9917. 6:06:59want to understand as how model is being
  9918. 6:07:02working.
  9919. 6:07:05Okay.
  9920. 6:07:06And I'm going to execute this cell.
  9921. 6:07:09So, this is going to take care of
  9922. 6:07:10necessary imports. I'm importing my
  9923. 6:07:12necessary libraries.
  9924. 6:07:14And once that is done,
  9925. 6:07:16here I'm importing this SVM. So, this
  9926. 6:07:20SVM model is available inside my
  9927. 6:07:22scikit-learn library. So, I've mentioned
  9928. 6:07:24as
  9929. 6:07:25scikit-learn .svm
  9930. 6:07:29and from scikit-learn.svm, I'm importing
  9931. 6:07:32my SVC. Okay? So, I'm importing my SVC.
  9932. 6:07:35So, I'll show you what is this SVC.
  9933. 6:07:38Um SVM SVC
  9934. 6:07:45So, it's C means support vector
  9935. 6:07:47classification.
  9936. 6:07:48Okay? Now, here when I'm instantiating
  9937. 6:07:52this SVC, I can mention what is the
  9938. 6:07:54kernel that I want to use. And if I'm
  9939. 6:07:57working with any polynomial kernel, then
  9940. 6:07:59I can also mention what is the degree of
  9941. 6:08:01polynomial that I want to use while
  9942. 6:08:03performing the fit for my data set.
  9943. 6:08:06So, I'm importing my SVC. And along with
  9944. 6:08:09that, I'm also importing the data sets.
  9945. 6:08:12So, in the scikit-learn library itself,
  9946. 6:08:14we have a data set. So, it the
  9947. 6:08:17scikit-learn host already like it it
  9948. 6:08:19actually scikit-learn has many toy data
  9949. 6:08:22set which will actually help us in our
  9950. 6:08:24learning journey. So, we are going to
  9951. 6:08:25use one of the data set, the famous Iris
  9952. 6:08:28data set. We use that for multi-class
  9953. 6:08:31classification.
  9954. 6:08:32So, I'm going to load that Iris data
  9955. 6:08:34set.
  9956. 6:08:35And I'm just going to extract only two
  9957. 6:08:37features. So, the two features that I'm
  9958. 6:08:39extracting is petal length and petal
  9959. 6:08:41width because I don't want to complicate
  9960. 6:08:43it. I just want to visualize the data.
  9961. 6:08:45So, in order to help in visualization, I
  9962. 6:08:48I'm just getting only two features of my
  9963. 6:08:50given data.
  9964. 6:08:52And I'm separating my Y as
  9965. 6:08:55Iris target. So, whatever the target
  9966. 6:08:56variable that I had, I'm assigning to my
  9967. 6:08:59variable of Y.
  9968. 6:09:01Then,
  9969. 6:09:02I'm going to uh
  9970. 6:09:04do this check whether it is setosa or
  9971. 6:09:07versicolor.
  9972. 6:09:08That means this default data set, which
  9973. 6:09:11is in multi-class classification, I'm
  9974. 6:09:13just going to convert it into a binary
  9975. 6:09:15classification task.
  9976. 6:09:17You'll get a better understanding once I
  9977. 6:09:19execute this next cell. So, this going
  9978. 6:09:21to prepare my data set and once the data
  9979. 6:09:24set is prepared, if I create a scatter
  9980. 6:09:26plot, so I'm just creating the scatter
  9981. 6:09:28plot to show us
  9982. 6:09:30what and how my data set looks like. So,
  9983. 6:09:33this is how my data set looks like.
  9984. 6:09:36On my X axis, I think I'm having petal
  9985. 6:09:38length. On my Y axis, I'm having petal
  9986. 6:09:40width.
  9987. 6:09:41And here,
  9988. 6:09:42the blue points refers to the class zero
  9989. 6:09:45and the orange points refers to the
  9990. 6:09:48class of one.
  9991. 6:09:50Okay? So, this is how my data set looks
  9992. 6:09:54like.
  9993. 6:09:55You can clearly see that I have one set
  9994. 6:09:57of points in one region and I have
  9995. 6:09:59another set of points in another region.
  9996. 6:10:01Now, this is a classic example to
  9997. 6:10:03understand about the SVM. How does it uh
  9998. 6:10:06draw a hyperplane?
  9999. 6:10:09So, we have the data set ready.
  10000. 6:10:11And as I mentioned already, in order to
  10001. 6:10:15fit this model, so when I say support
  10002. 6:10:17vector machine, I'm going to draw a
  10003. 6:10:18line.
  10004. 6:10:20This line that I have drawn, it will be
  10005. 6:10:23equidistant from my support vectors.
  10006. 6:10:25Now, here in this example, the support
  10007. 6:10:27vector is this because this is the only
  10008. 6:10:30point which is nearest to my line. And
  10009. 6:10:32here, I think this is the data point
  10010. 6:10:34which is nearest to my SVM SVM line,
  10011. 6:10:37that is this twisted line hyperplane
  10012. 6:10:38line.
  10013. 6:10:39So, I'll be placing this hyperplane such
  10014. 6:10:42that it is equidistant from the support
  10015. 6:10:45vectors.
  10016. 6:10:47That is how I'll be drawing this support
  10017. 6:10:49vector line.
  10018. 6:10:51So, we now have an intuition. Let's see
  10019. 6:10:53whether we get the same outcome as we
  10020. 6:10:55are expecting.
  10021. 6:10:57So, here I'm initializing my model. So,
  10022. 6:11:00for initialization, I'm saying it as
  10023. 6:11:02SVC. Use the kernel as linear because
  10024. 6:11:06I'm able to draw a line effectively. We
  10025. 6:11:08were We just seen. And I'm using the C
  10026. 6:11:11as infinity, that means it should be a
  10027. 6:11:13hard classifier. So, hard classifier
  10028. 6:11:15means I make I want the 100% result. I
  10029. 6:11:18mean, I don't want any loosens. I want
  10030. 6:11:20to draw a line which passes which
  10031. 6:11:23clearly separates two classes. So, I'm
  10032. 6:11:25saying it as C as infinity to mention
  10033. 6:11:27this as a hard classifier.
  10034. 6:11:30And once I initialize any model,
  10035. 6:11:32here in this scenario, SVM model, I'm
  10036. 6:11:35performing the fit on my data set. Now,
  10037. 6:11:36this is the common flow that we follow
  10038. 6:11:39whenever we are performing the fit. So,
  10039. 6:11:41we'll initialize the model and then we
  10040. 6:11:43perform the fit on a data set.
  10041. 6:11:46Now, since this SVM being a supervised
  10042. 6:11:49machine learning model, I have to
  10043. 6:11:51specify both my input X as well as my
  10044. 6:11:55output Y.
  10045. 6:11:56Hence,
  10046. 6:11:57SVM classifier.fit
  10047. 6:12:00X, Y.
  10048. 6:12:01So, this is going to perform the fit for
  10049. 6:12:03my data set.
  10050. 6:12:05I'll just execute this. So, this has
  10051. 6:12:08performed the fit and here it is giving
  10052. 6:12:10me the confirmation as this is the
  10053. 6:12:13parameter that are being used to perform
  10054. 6:12:15the fit.
  10055. 6:12:17Okay. Now, once I have drawn and once I
  10056. 6:12:20have found this fit,
  10057. 6:12:22next,
  10058. 6:12:23if I want to display the weight terms,
  10059. 6:12:26so, I can say it as SVM
  10060. 6:12:28classifier.coefficients.
  10061. 6:12:30So, these are the weight terms. And if I
  10062. 6:12:32want to display my bias term or the
  10063. 6:12:34intercept, it is minus 3.78.
  10064. 6:12:38Now, this means the line that I've just
  10065. 6:12:40drawn, so that line has the
  10066. 6:12:44uh
  10067. 6:12:45that line has the C term as or the W not
  10068. 6:12:48term as minus 3.78 and W1, W2 are 1.29
  10069. 6:12:53and 0.82, respectively.
  10070. 6:12:56So, that's how the data is distributed
  10071. 6:12:59for us. That's how the values has been
  10072. 6:13:02formed for our scenario.
  10073. 6:13:05Next, in order to get the better
  10074. 6:13:07visualization,
  10075. 6:13:08here I have created a function that is
  10076. 6:13:10called as plot SVC decision boundary and
  10077. 6:13:14this takes my SVM model,
  10078. 6:13:17the X min and the X max.
  10079. 6:13:21Now, W and B I'm extracting from the
  10080. 6:13:24coefficient and the intercept parameter
  10081. 6:13:26that we have over here.
  10082. 6:13:28So, we are extracting from this
  10083. 6:13:30uh
  10084. 6:13:31at from this attributes that we have
  10085. 6:13:33from this model.
  10086. 6:13:35And now, if I want to draw a decision
  10087. 6:13:37boundary,
  10088. 6:13:38so,
  10089. 6:13:39I need the set of points. So, in order
  10090. 6:13:42to get the points, I'm saying it as X
  10091. 6:13:43not is equal to np.linspace X max, X
  10092. 6:13:46min, X max, 200. And I'm specifying as
  10093. 6:13:50how does my decision boundary should
  10094. 6:13:52look like.
  10095. 6:13:53My decision boundary is given by w
  10096. 6:13:56naught into x naught plus w one into x
  10097. 6:13:58one plus b is equal to zero. So, this is
  10098. 6:14:00what my
  10099. 6:14:01decision boundary would look like. So, I
  10100. 6:14:03know what is x naught. I know w naught.
  10101. 6:14:06I also have w one and I also have b. So,
  10102. 6:14:09the only term that I do not have is my
  10103. 6:14:12x one.
  10104. 6:14:13Okay? So, the only term that I do not
  10105. 6:14:15have over here in this example is x one.
  10106. 6:14:18And the x one if I want it, so I just
  10107. 6:14:20have to substitute it. So, x one is
  10108. 6:14:22equal to minus w zero divided by w one
  10109. 6:14:25into x naught minus b divided by w one.
  10110. 6:14:30Now, I'm specifying the same equation
  10111. 6:14:33over here for my x two.
  10112. 6:14:35So, my x naught and the decision
  10113. 6:14:37boundary will give me the pair of input
  10114. 6:14:40and output. Okay?
  10115. 6:14:42Now, along with this
  10116. 6:14:44there is a property in SVM. Okay? So,
  10117. 6:14:47the property is given by whenever I have
  10118. 6:14:49a margin, so that margin is given by one
  10119. 6:14:53over w one.
  10120. 6:14:55Okay? So, the margin is nothing but the
  10121. 6:14:58distance between my hyper plane and the
  10122. 6:15:01support vector. So, that is given by ma
  10123. 6:15:04one by w one.
  10124. 6:15:06Hence, I have mentioned as gutter up and
  10125. 6:15:08down. Gutter up means one line or the
  10126. 6:15:11one line where the support vector lies.
  10127. 6:15:13So, that is given by decision boundary
  10128. 6:15:16plus margin.
  10129. 6:15:17And one line below my
  10130. 6:15:19one line below my hyper plane.
  10131. 6:15:22That is where another support vector
  10132. 6:15:24would lie. So, I mentioned as decision
  10133. 6:15:26boundary minus margin.
  10134. 6:15:29Okay?
  10135. 6:15:30Then
  10136. 6:15:31I'm defining where exactly my support
  10137. 6:15:34vectors are present.
  10138. 6:15:37My support vectors uh I can access the
  10139. 6:15:40support vectors coordinates by saying it
  10140. 6:15:42as
  10141. 6:15:43by accessing the attribute of my train
  10142. 6:15:45model support underscore vectors
  10143. 6:15:47underscore.
  10144. 6:15:49Now, I'm specifying where exactly those
  10145. 6:15:51support vectors are present with the
  10146. 6:15:53help of a simple scatter plot by
  10147. 6:15:55highlighting my support vectors.
  10148. 6:15:57And I'm specifying where exactly my
  10149. 6:15:59decision boundary is present. And I'm
  10150. 6:16:02also mentioning where is my line that is
  10151. 6:16:04gutter up and gutter down. So, let's do
  10152. 6:16:06one thing. I'll just execute this. This
  10153. 6:16:08is going to create me a function.
  10154. 6:16:11I'm going to call my function
  10155. 6:16:13support vector machine classification.
  10156. 6:16:16And I will specify my range of X and Y
  10157. 6:16:18as
  10158. 6:16:19here, yeah, X min and X max as 0 {comma}
  10159. 6:16:225.5.
  10160. 6:16:24I'll just execute this.
  10161. 6:16:28So,
  10162. 6:16:29what we have done just now is we have
  10163. 6:16:32created this hyperplane.
  10164. 6:16:36So, the middle one, the solid line that
  10165. 6:16:38you're seeing over here, so this solid
  10166. 6:16:40line is called as your hyperplane.
  10167. 6:16:43And these points that you're seeing over
  10168. 6:16:45here, so these two points which are
  10169. 6:16:47highlighted, these two points are
  10170. 6:16:49actually called as support vectors.
  10171. 6:16:54Okay? So, this dotted line that you're
  10172. 6:16:57seeing, so this dotted line refers to my
  10173. 6:17:00gutter up and gutter down which I've
  10174. 6:17:02found right here.
  10175. 6:17:05Let's do one thing. I'll add some label
  10176. 6:17:07so that you'll get some more
  10177. 6:17:09visualization in the plot itself. I'll
  10178. 6:17:11say label and I'll mention it as
  10179. 6:17:14hyperplane.
  10180. 6:17:19Okay.
  10181. 6:17:21And
  10182. 6:17:23there is one more, yeah.
  10183. 6:17:36These are support vectors.
  10184. 6:17:38And I'll say
  10185. 6:17:42plt.legend.
  10186. 6:18:00So, this clearly says which are all my
  10187. 6:18:04hyperplane and which are all my support
  10188. 6:18:06vectors.
  10189. 6:18:09So, this is the intuition behind support
  10190. 6:18:11vector machines.
  10191. 6:18:13So, we'll be drawing a hyperplane which
  10192. 6:18:16separates the points that we have.
  10193. 6:18:18Okay? And whichever the point which is
  10194. 6:18:20nearest to my hyperplane, we call that
  10195. 6:18:22point as a support vector. Now, to
  10196. 6:18:25access that support vector, we make use
  10197. 6:18:27of the attribute. So, let's do one
  10198. 6:18:29thing. Let's explore the same the
  10199. 6:18:31attributes.
  10200. 6:18:32svm.
  10201. 6:18:33support_vectors_.
  10202. 6:18:36So, this is going to tell me where
  10203. 6:18:37exactly my support vectors are present.
  10204. 6:18:40So, one point is given by 1.9.0.4.
  10205. 6:18:43I think this is the point that I'm
  10206. 6:18:45talking about.
  10207. 6:18:46And the another support vector that we
  10208. 6:18:48have is at the location 3, 1.1. So, 3
  10209. 6:18:51and 1.1. This is where we have another
  10210. 6:18:54support vector.
  10211. 6:18:56So, using all these attributes, we have
  10212. 6:18:58been able to create this visualization.
  10213. 6:19:03Okay.
  10214. 6:19:05Now,
  10215. 6:19:06whenever we are working with the support
  10216. 6:19:07vector machines,
  10217. 6:19:09it's very important that we scale the
  10218. 6:19:11data first. If I do not scale the data,
  10219. 6:19:14I'll not be able to get a better fit of
  10220. 6:19:17my SVM model.
  10221. 6:19:19So, here I've given one more example
  10222. 6:19:22where I have my X
  10223. 6:19:24uh is given as 1, 55, 23, 80. As you can
  10224. 6:19:28clearly see, it's it's not scaled. Okay?
  10225. 6:19:32So, I'm going to execute this cell. So,
  10226. 6:19:35this is going to tell me and give me a
  10227. 6:19:37visualization as how the
  10228. 6:19:39fit will be in case of scaled and
  10229. 6:19:42unscaled.
  10230. 6:19:44See, if it is unscaled
  10231. 6:19:47I'll If it is not scaled, okay? That
  10232. 6:19:49means if it is unscaled, we can clearly
  10233. 6:19:51see that the hyperplane that I'm drawing
  10234. 6:19:55and the distance from my hyperplane,
  10235. 6:19:57it's very close to each other.
  10236. 6:20:01And whenever I'm working, it's It's It
  10237. 6:20:03will be difficult for me to separate
  10238. 6:20:04those two data points.
  10239. 6:20:08But, if I scale them correctly
  10240. 6:20:11Now, here for scaling, I have made use
  10241. 6:20:12of a scalar standard scalar. Now, if I
  10242. 6:20:15scale it correctly, then in that
  10243. 6:20:17scenario, it will be easier for me and
  10244. 6:20:20it would actually work better when I
  10245. 6:20:22have scaled data.
  10246. 6:20:26Okay? So
  10247. 6:20:28this is about using the linear SVM model
  10248. 6:20:32to perform the fit on my given data set.
  10249. 6:20:36Now, if I go below, we have some more
  10250. 6:20:38examples about non-linear classifiers as
  10251. 6:20:41well.
  10252. 6:20:42Now, in order to test out the same
  10253. 6:20:45here
  10254. 6:20:46I'm creating an example data set and
  10255. 6:20:48that data set that I'm generating is
  10256. 6:20:50called as make moons data set and this
  10257. 6:20:53has been generated with the help of a
  10258. 6:20:55scalar data set generator.
  10259. 6:20:57Now, as you can clearly see, I cannot
  10260. 6:21:00make use of linear classifier. So,
  10261. 6:21:02linear classifier is nothing but a
  10262. 6:21:05classifier, okay? Which is an SVM model
  10263. 6:21:08where I'm drawing or where I'm using a
  10264. 6:21:10straight line to split my data points. I
  10265. 6:21:13can clearly see that I Wherever I I join
  10266. 6:21:16or wherever I try to draw a line over
  10267. 6:21:19here, I cannot split the data in an
  10268. 6:21:21effective manner.
  10269. 6:21:23Now, this brings us the challenge. Now,
  10270. 6:21:24if I have a data set in this way where I
  10271. 6:21:27cannot linearly separate it, how can we
  10272. 6:21:30go about and fit uh perform the fit on
  10273. 6:21:33our SVM model?
  10274. 6:21:34So, in order to save us, we have a model
  10275. 6:21:37that is called as uh SVM model, and from
  10276. 6:21:40that SVM model, we can actually create a
  10277. 6:21:43polynomial uh
  10278. 6:21:45polynomial kernel. So, we can make use
  10279. 6:21:46of polynomial kernel, and using that
  10280. 6:21:48polynomial kernel, I can actually uh
  10281. 6:21:52create it like this. I mean, using
  10282. 6:21:53polynomial kernel, I can perform
  10283. 6:21:55polynomial regression.
  10284. 6:21:57Or I can draw a line like this. Now, to
  10285. 6:21:59show you how it works,
  10286. 6:22:01I'm getting some data like this. So,
  10287. 6:22:03this is some uh random data.
  10288. 6:22:06And I'm making use of pipeline.
  10289. 6:22:08So, this pipeline is going to take care
  10290. 6:22:10of my stan- standard scaler as well as
  10291. 6:22:13kernel.
  10292. 6:22:14I'll do one thing, I'll just come below.
  10293. 6:22:16So, this is what we are currently
  10294. 6:22:17interested in.
  10295. 6:22:23So, here,
  10296. 6:22:26I'm importing the polynomial features,
  10297. 6:22:29and I'm generating the polynomial
  10298. 6:22:31features for my data.
  10299. 6:22:33I'm performing the fit and transform my
  10300. 6:22:35polynomial data. That means, I'm just
  10301. 6:22:37modifying my existing data, and I am
  10302. 6:22:39sending it
  10303. 6:22:41for my uh
  10304. 6:22:43X, okay? So, this is how my pair of
  10305. 6:22:45input X and Y looks like.
  10306. 6:22:47Now, I'll use my X. I'm going to
  10307. 6:22:49transform it with the help of my
  10308. 6:22:51polynomial features,
  10309. 6:22:53and then, I'm going to scale it with the
  10310. 6:22:56help of my standard scaler,
  10311. 6:22:58and I'm going to send it inside my
  10312. 6:23:01classifier, that is SVM classifier.
  10313. 6:23:08Okay? So, I'm going to
  10314. 6:23:11combine it together like this.
  10315. 6:23:14Now, observe what would happen.
  10316. 6:23:17Now, once that is complete,
  10317. 6:23:19see?
  10318. 6:23:20With the help of my polynomial uh
  10319. 6:23:24polynomial features that have applied on
  10320. 6:23:26my given linear data.
  10321. 6:23:28So, I have increased the degrees by
  10322. 6:23:32which my model can learn.
  10323. 6:23:36Now, instead of straight line, my model
  10324. 6:23:37is also having the ability to learn this
  10325. 6:23:40complex representation as well.
  10326. 6:23:43Because I have increased the model
  10327. 6:23:45complexity by adding my polynomial
  10328. 6:23:47features.
  10329. 6:23:50And while doing it, to make sure that we
  10330. 6:23:51follow a
  10331. 6:23:53clear path, so I have defined this is
  10332. 6:23:56scalar's pipeline. So, if you're new to
  10333. 6:23:58data science machine learning, I highly
  10334. 6:24:00recommend you to learn this concept of a
  10335. 6:24:02scalar pipeline. Now, this is scalar
  10336. 6:24:04pipeline helps us to combine multiple
  10337. 6:24:07operations in a single call.
  10338. 6:24:10So, here we have created a pipeline.
  10339. 6:24:12This pipeline is going to add some
  10340. 6:24:15polynomial features for my input data.
  10341. 6:24:18And on top of it, this is going to
  10342. 6:24:19perform scaling. And then I'm going to
  10343. 6:24:21perform this binomial classification
  10344. 6:24:24using this SVM.
  10345. 6:24:27And finally,
  10346. 6:24:29I'm performing the fit on my data set.
  10347. 6:24:31See, when I perform the fit, it takes my
  10348. 6:24:33input X and it's going to do all these
  10349. 6:24:36activities. It is going to chain all
  10350. 6:24:38these activities together, and then it
  10351. 6:24:41is going to perform the fit for my data
  10352. 6:24:43Y.
  10353. 6:24:45Once the fit has been complete, so we
  10354. 6:24:47can validate how my model is performing.
  10355. 6:24:50>> [music]
  10356. 6:24:55[music]
  10357. 6:24:56>> So, what is a clustering technique?
  10358. 6:24:57Clustering technique is something that
  10359. 6:24:59we will use it for grouping purpose.
  10360. 6:25:03So, especially there's a very easy way
  10361. 6:25:06to understand what is clustering
  10362. 6:25:07technique. You would have seen such a
  10363. 6:25:09while we are going through an
  10364. 6:25:11COVID-19 situation, the governments has
  10365. 6:25:13came up with creating some containment
  10366. 6:25:15zones.
  10367. 6:25:16As all of you must be knowing.
  10368. 6:25:19So, how on what criteria government has
  10369. 6:25:21taken that okay, which area supposed to
  10370. 6:25:23be a containment zone or which area
  10371. 6:25:25supposed to be applied with some some
  10372. 6:25:26restrictions and which areas can be can
  10373. 6:25:29be considered as normal? On what
  10374. 6:25:30criteria that they have created? So,
  10375. 6:25:32that's what using clustering technique.
  10376. 6:25:35Which means if the governments or when I
  10377. 6:25:37say government means that the people who
  10378. 6:25:38will be taking the final decision in
  10379. 6:25:40such criteria, either prime minister
  10380. 6:25:41either either the chief ministers of
  10381. 6:25:43that particular state will be be taking
  10382. 6:25:45decisions whether to go for lockdown
  10383. 6:25:47whether to not to go for lockdown or
  10384. 6:25:49which areas has to be considered as
  10385. 6:25:51containment zones or non-containment
  10386. 6:25:52zones.
  10387. 6:25:53So, those high-level decisions are
  10388. 6:25:55something which will be taken based on
  10389. 6:25:57the clustering technique output which is
  10390. 6:25:59generated by the these algorithms.
  10391. 6:26:02Based on a number of inputs, okay, what
  10392. 6:26:04is the population in a particular area?
  10393. 6:26:06How many number of people are affected?
  10394. 6:26:08How many number of hospitals which are
  10395. 6:26:10present?
  10396. 6:26:11How many number of
  10397. 6:26:13people who are been recovered? So,
  10398. 6:26:16likewise based on this these multiple
  10399. 6:26:18criteria, people will do some clustering
  10400. 6:26:20technique on top of the data and
  10401. 6:26:22according to that people will be
  10402. 6:26:24segregated or the areas will be
  10403. 6:26:26segregated.
  10404. 6:26:27So, that saying that okay, these are the
  10405. 6:26:28observations which belong to one
  10406. 6:26:29cluster, these are the observations
  10407. 6:26:31which belong to one cluster like that so
  10408. 6:26:32that people can cluster them which can
  10409. 6:26:34make organizations to take decisions on
  10410. 6:26:38a very high level.
  10411. 6:26:40Okay? That's what is all clustering
  10412. 6:26:41technique.
  10413. 6:26:42Which clustering technique output will
  10414. 6:26:45contain the different different groups.
  10415. 6:26:46It itself will group the different
  10416. 6:26:47different components
  10417. 6:26:49based on whatever the number of clusters
  10418. 6:26:51that you want to generate. That may not
  10419. 6:26:53give the direct output. On top of the
  10420. 6:26:55generated output, people will be taking
  10421. 6:26:56business related decisions. That's what
  10422. 6:26:58is all about clustering techniques.
  10423. 6:27:01Okay. So, now what we will do?
  10424. 6:27:03Let's take an example of within
  10425. 6:27:06clustering techniques, what are the
  10426. 6:27:07different types of clustering techniques
  10427. 6:27:08we have?
  10428. 6:27:09So, what What different types of
  10429. 6:27:10clustering that we have?
  10430. 6:27:13So, there are multiple types of there
  10431. 6:27:14are multiple types of ways based on the
  10432. 6:27:17type of output that we want to produce.
  10433. 6:27:18There are multiple different types of
  10434. 6:27:20clustering techniques we have, but out
  10435. 6:27:21of which the let's try to understand
  10436. 6:27:23about what are the very famous and most
  10437. 6:27:25widely used clustering technique
  10438. 6:27:26algorithm. Out of which we have
  10439. 6:27:28something called K-means clustering
  10440. 6:27:30algorithm is one of the very famous and
  10441. 6:27:32most widely used. More than 90% of the
  10442. 6:27:35people will end up with using K-means
  10443. 6:27:36clustering algorithm, which is very very
  10444. 6:27:38famous in clustering techniques.
  10445. 6:27:40Right? Which is very very famous in
  10446. 6:27:42clustering techniques.
  10447. 6:27:43So, what are these clustering
  10448. 6:27:44techniques? As I said, how this
  10449. 6:27:46clustering technique will work.
  10450. 6:27:48So, K-means clustering is nothing but
  10451. 6:27:49always remember one thing. If any one of
  10452. 6:27:51you were going to work in machine
  10453. 6:27:52learning or anywhere in anywhere,
  10454. 6:27:55wherever you see a notation called K,
  10455. 6:27:57by default K is nothing but you are you
  10456. 6:28:01are supposed to as a user, you are
  10457. 6:28:03supposed to provide what is the input of
  10458. 6:28:06K. Which means wherever you see there
  10459. 6:28:08are multiple techniques that we have in
  10460. 6:28:09machine learning like K-means clustering
  10461. 6:28:11technique, K nearest neighbor is one of
  10462. 6:28:13the algorithm, K-fold cross validation,
  10463. 6:28:15likewise. Wherever you see a notation
  10464. 6:28:17called K, what is what does a K means?
  10465. 6:28:19It's an input that you are supposed to
  10466. 6:28:21provide. Always remember this.
  10467. 6:28:24It's an input that you are supposed to
  10468. 6:28:25provide
  10469. 6:28:27to your algorithm. Your algorithm cannot
  10470. 6:28:29identify that K value. Of course,
  10471. 6:28:30everything else will be taken care by
  10472. 6:28:31your algorithm, but whenever you see K,
  10473. 6:28:33which means in K-means clustering, what
  10474. 6:28:35is the meaning of K-means clustering?
  10475. 6:28:37How many number of clusters that you
  10476. 6:28:38want to provide? That is something that
  10477. 6:28:40you have to input it to your algorithm.
  10478. 6:28:43That is something that you have to input
  10479. 6:28:44to your algorithm.
  10480. 6:28:45Right?
  10481. 6:28:47Here, the meaning of K is how many
  10482. 6:28:50clusters that you want to generate.
  10483. 6:28:52So, how many clusters that you want to
  10484. 6:28:54generate? Okay, when you have 1,000
  10485. 6:28:55observations which are present, when you
  10486. 6:28:57have 1,000 in input column input records
  10487. 6:28:59which are present in your historical
  10488. 6:29:01data, how many number of clusters that
  10489. 6:29:03you want to provide? Do you want to go
  10490. 6:29:04for one cluster? Obviously, one cluster
  10491. 6:29:06means that the entire data set will be
  10492. 6:29:07considered as is.
  10493. 6:29:09Do you want to create two clusters out
  10494. 6:29:11of the data?
  10495. 6:29:12Do you want to create three clusters out
  10496. 6:29:14of the data? Four clusters, five
  10497. 6:29:15clusters, or 10 clusters?
  10498. 6:29:17So, how this can be done?
  10499. 6:29:18There are multiple steps that are
  10500. 6:29:20involved in generating K-means
  10501. 6:29:22clustering algorithm. So, you can see,
  10502. 6:29:23choose the number of clusters. This is
  10503. 6:29:25what is nothing but your first step. It
  10504. 6:29:26means you need to decide what is your
  10505. 6:29:29K is nothing but number of clusters that
  10506. 6:29:31you want to produce.
  10507. 6:29:32And then, there is an initialization of
  10508. 6:29:35centroids will happen as a one-time
  10509. 6:29:37activity.
  10510. 6:29:38Right? So, there is an initialization of
  10511. 6:29:40centroids which will be which will be
  10512. 6:29:42declared that will that will be used as
  10513. 6:29:44your initial step for your machine. And
  10514. 6:29:45then, assign the clusters, move the
  10515. 6:29:47centroids, and optimization, and then
  10516. 6:29:49converge the the
  10517. 6:29:51all the clusters into one component.
  10518. 6:29:53Yes, I know it will be very difficult to
  10519. 6:29:54understand by looking at this thing. So,
  10520. 6:29:56let me show you a very simple example
  10521. 6:29:58how exactly it will be done. Maybe let
  10522. 6:29:59me take a simple diagram for you to show
  10523. 6:30:01how exactly that's going to work.
  10524. 6:30:05Okay? So, let's say for example, I'm
  10525. 6:30:07going to take some historical data just
  10526. 6:30:09to explain you on how exactly the
  10527. 6:30:11K-means clustering algorithm will work.
  10528. 6:30:13So, what is that it is written in the
  10529. 6:30:14first step?
  10530. 6:30:16What is the data is written in the first
  10531. 6:30:17step?
  10532. 6:30:18So, the choose the number of clusters.
  10533. 6:30:20Choose number of clusters. Now, it means
  10534. 6:30:22say for example, we need to take some
  10535. 6:30:24historical data. I'm considering some
  10536. 6:30:25historical data here. Let's assume this
  10537. 6:30:27is the historical data.
  10538. 6:30:29Let's assume this is the historical data
  10539. 6:30:31that we have.
  10540. 6:30:32So, as you can see, there are multiple
  10541. 6:30:35historical data points. As you can see,
  10542. 6:30:37the first step, choose number of
  10543. 6:30:38clusters, which means let's assume to
  10544. 6:30:40make this thing simple, I'm going to
  10545. 6:30:42choose that we want to have two clusters
  10546. 6:30:43generated. Okay? So, this is which means
  10547. 6:30:46we want to have two clusters created.
  10548. 6:30:47So, because my number of clusters that I
  10549. 6:30:49want to generate is two, I'm going to
  10550. 6:30:51consider that there are two centroids
  10551. 6:30:52which are present. This is my first
  10552. 6:30:54step. How this algorithm will do?
  10553. 6:30:56How algorithm will come to come with the
  10554. 6:30:58number of clusters? So, this is how it
  10555. 6:30:59will happen. So, the first step is to
  10556. 6:31:01choose the number of centroid and then
  10557. 6:31:03initialize your centroid. That's the
  10558. 6:31:04second step. So, now what is the third
  10559. 6:31:06step?
  10560. 6:31:07Let's assume this is observation number
  10561. 6:31:09one. This is our data point one.
  10562. 6:31:12So, now what what is the next step?
  10563. 6:31:14We will take the distance from every
  10564. 6:31:16observation to every centroid and every
  10565. 6:31:18observation to every centroid, which
  10566. 6:31:19means now tell me which observation
  10567. 6:31:23For this observation number one, which
  10568. 6:31:24centroid is more closer?
  10569. 6:31:26Is the green color centroid is more
  10570. 6:31:27closer or the red color centroid is more
  10571. 6:31:29closer for this?
  10572. 6:31:31Centroid means that data points, the one
  10573. 6:31:33which I have highlighted, that's what we
  10574. 6:31:34call a centroid, data points.
  10575. 6:31:37We have chosen two centroids only
  10576. 6:31:39because that is the as a user you're
  10577. 6:31:41supposed to input what's supposed to be
  10578. 6:31:42your K.
  10579. 6:31:43What's supposed to be your K value,
  10580. 6:31:45that's what I said. You need to know how
  10581. 6:31:47what is your K is nothing but how many
  10582. 6:31:48number of clusters that you want to
  10583. 6:31:49provide. Usually your business users are
  10584. 6:31:51going to provide that. In case if you're
  10585. 6:31:53going to work in this kind of
  10586. 6:31:54algorithms, they will provide that. Or
  10587. 6:31:56else there are some other methods
  10588. 6:31:58available like
  10589. 6:31:59elbow method available other things.
  10590. 6:32:01Yeah, we'll talk about that.
  10591. 6:32:03Now, this observation is more closer to
  10592. 6:32:05red color. So, now what happens? What
  10593. 6:32:07what is the next step? The algorithm
  10594. 6:32:08will assign this particular observation
  10595. 6:32:10one as a red color for now.
  10596. 6:32:12As considering that this is belong to
  10597. 6:32:13red color. Likewise for the second
  10598. 6:32:15observation, when the second observation
  10599. 6:32:16appears, what is the distance from the
  10600. 6:32:18second observation to both the
  10601. 6:32:20centroids? Now, which one is more
  10602. 6:32:21closer? I see green color is more
  10603. 6:32:23closer. So, now I'll mark this as green
  10604. 6:32:25color.
  10605. 6:32:26If both if what if the distance is same
  10606. 6:32:28equal? So, then the algorithm will force
  10607. 6:32:31any of the observation to get into any
  10608. 6:32:33of the centroid.
  10609. 6:32:35So, number of clusters number of
  10610. 6:32:37centroids that you will choose based on
  10611. 6:32:38the K value as I said.
  10612. 6:32:41Now, you're going to Likewise, you will
  10613. 6:32:43repeat the same process and whatever the
  10614. 6:32:45observation which is more closer to
  10615. 6:32:47whatever the centroid it is, you will
  10616. 6:32:48mark them as with their so-called mark
  10617. 6:32:51like this. Now, you're going to
  10618. 6:32:52initially mark them as observations into
  10619. 6:32:54either into green color either into red
  10620. 6:32:56color.
  10621. 6:32:56So, now, these observations are now
  10622. 6:32:59considered as green color observations,
  10623. 6:33:00and these observations are now
  10624. 6:33:02considered as red color observations.
  10625. 6:33:05This is step number one.
  10626. 6:33:07What is the step number two?
  10627. 6:33:09Step number two is segregate all these
  10628. 6:33:11red color observations and take the
  10629. 6:33:13average value of X and Y coordinates,
  10630. 6:33:15and repeat the same process for
  10631. 6:33:17calculating average coordinates of X and
  10632. 6:33:18Y coordinates for the green color
  10633. 6:33:20observations.
  10634. 6:33:21Take the average of all these
  10635. 6:33:22observations, calculate average, and
  10636. 6:33:24take the all these observations, take
  10637. 6:33:25the average. You end up with getting a
  10638. 6:33:27new centroid positions called XY. Which
  10639. 6:33:30means, now you ended up with getting a
  10640. 6:33:32new centroids in the initial step that
  10641. 6:33:35we have taken random centroid. Now, you
  10642. 6:33:37got the centroids that you can use based
  10643. 6:33:40on the previous iteration. Now, you
  10644. 6:33:41ended up with getting a new centroids.
  10645. 6:33:44You repeat the same process again.
  10646. 6:33:46Again, you repeat the same process. Take
  10647. 6:33:47the distance from every observation to
  10648. 6:33:49every centroid and assign the
  10649. 6:33:50observation based on the nearest nearest
  10650. 6:33:52to distance, and continue to mark every
  10651. 6:33:55observation either into red color or
  10652. 6:33:56either green color or whatever it is,
  10653. 6:33:58and you repeat the process until you
  10654. 6:34:00will be able to see there will be no
  10655. 6:34:02change applicable for your clusters.
  10656. 6:34:05You repeat the process. Which means, in
  10657. 6:34:07every step, you might end up with
  10658. 6:34:08changing your centroid. Every step will
  10659. 6:34:10continue to change your centroid. Now,
  10660. 6:34:12the centroid might become like this.
  10661. 6:34:13Then, later, your centroid will become
  10662. 6:34:15like this.
  10663. 6:34:17Likewise, your centroids will be keep on
  10664. 6:34:19moving. Initially, you have taken it
  10665. 6:34:20like this, but it it might continue to
  10666. 6:34:22move like this. Somewhere, it will be
  10667. 6:34:23fixed. And after that, there will be no
  10668. 6:34:25change that you will notice if you are
  10669. 6:34:26repeating the same process. At this
  10670. 6:34:28particular stage, whatever the
  10671. 6:34:30observations which are marked into which
  10672. 6:34:32are
  10673. 6:34:33grouped into green color, you say like
  10674. 6:34:35these are the green color observations.
  10675. 6:34:37Whatever the observations which are
  10676. 6:34:39marked into red color, you will you will
  10677. 6:34:40mark them as okay, these are the
  10678. 6:34:41observations which belong to red color.
  10679. 6:34:44These are the observations which belong
  10680. 6:34:45to red color. Likewise, you will
  10681. 6:34:47segregate all these observations either
  10682. 6:34:49into red color or green color.
  10683. 6:34:51Right? So, so, we can generate these
  10684. 6:34:54cluster techniques. That's how the
  10685. 6:34:55K-means clustering algorithm will
  10686. 6:34:56generate these algorithms
  10687. 6:34:58output of using this algorithm.
  10688. 6:35:00Okay? So, that's how K-means clustering
  10689. 6:35:02algorithm will work.
  10690. 6:35:04Okay?
  10691. 6:35:05So, now what likewise there are multiple
  10692. 6:35:06algorithms that we have. So, like
  10693. 6:35:09when we talk about machine learning, so
  10694. 6:35:10there are multiple types of algorithms
  10695. 6:35:12that we have. So, like how we have the
  10696. 6:35:14how does the K value K value will occur
  10697. 6:35:16as you can see on the PPTs which are
  10698. 6:35:17also mentioned. So, you're going to
  10699. 6:35:19choose some randomly generated K value
  10700. 6:35:21and you will be choosing the number of
  10701. 6:35:22case over here and you can see that it
  10702. 6:35:24will assign them based on the number of
  10703. 6:35:26these easy
  10704. 6:35:27most nearest distances and according to
  10705. 6:35:29that you will change your centroids.
  10706. 6:35:31Once you change your centroids, you will
  10707. 6:35:33repeat the same process until you're
  10708. 6:35:34able to change that your centroids don't
  10709. 6:35:37move further and then once it has been
  10710. 6:35:39finalized, you will say like this is the
  10711. 6:35:40final centroid.
  10712. 6:35:42Final cluster output that we can
  10713. 6:35:44generate out of this.
  10714. 6:35:46Right?
  10715. 6:35:47So, likewise we also have different
  10716. 6:35:48types of cluster technique and the
  10717. 6:35:49second type of cluster technique that we
  10718. 6:35:51have is a fuzzy or C-means clustering.
  10719. 6:35:53So, what is fuzzy or C-means clustering?
  10720. 6:35:54The output will remain same. So, C-means
  10721. 6:35:56clustering means that there are places
  10722. 6:35:58that one or two observations can belong
  10723. 6:36:00to one or two different clusters.
  10724. 6:36:03Like one or two different clusters,
  10725. 6:36:05which means usually the primary
  10726. 6:36:06difference between
  10727. 6:36:09the primary difference between your
  10728. 6:36:10C-means and K-means clustering technique
  10729. 6:36:12is in case if there are any observations
  10730. 6:36:15which are having equal amount of
  10731. 6:36:16distance, usually in K-means clustering
  10732. 6:36:18what we will do, we will force this
  10733. 6:36:20observation to be part of any of the
  10734. 6:36:22cluster. But in C-means clustering,
  10735. 6:36:25based on the distance that we see, there
  10736. 6:36:27are chances that an observation can go
  10737. 6:36:29to or can belong to one or two clusters.
  10738. 6:36:33So, it purely depends on the business
  10739. 6:36:34use case who is going to decide to
  10740. 6:36:36either to go for K-means clustering or
  10741. 6:36:38C-means clustering based on the the
  10742. 6:36:39business use case.
  10743. 6:36:41Based on our business outcome, so if
  10744. 6:36:43people are going to decide whether to go
  10745. 6:36:44for C-means clustering or K-means
  10746. 6:36:45clustering. So, there are n-number of
  10747. 6:36:48observations which might It's not
  10748. 6:36:50mandatory. Which might can belong to one
  10749. 6:36:52or more clusters. That can happen.
  10750. 6:36:55So, the third is agglomerative
  10751. 6:36:57clustering, so which which is the third
  10752. 6:36:58type of clustering technique that we
  10753. 6:37:00have. Okay? So, which is also one of one
  10754. 6:37:03of the clustering technique that we
  10755. 6:37:04have.
  10756. 6:37:06Okay?
  10757. 6:37:07So, now which can also be used here.
  10758. 6:37:09Okay? So, now what is this agglomerative
  10759. 6:37:12clustering? So, this is what we also
  10760. 6:37:14call it as hierarchical clustering. So,
  10761. 6:37:16in times you'll also call them as
  10762. 6:37:18hierarchical clustering. These
  10763. 6:37:19clustering techniques are built using
  10764. 6:37:20H-clustering. In short, we call it as
  10765. 6:37:22H-clustering.
  10766. 6:37:23And also people will also call it as
  10767. 6:37:24hierarchical clustering techniques. So,
  10768. 6:37:25what are this? So, based on the type of
  10769. 6:37:28algorithm
  10770. 6:37:29they will try to There is a There is a
  10771. 6:37:31mathematical expression which are
  10772. 6:37:32involved in it. But, then considering
  10773. 6:37:33that the limited time that we have, I'm
  10774. 6:37:35not going to take you through all the in
  10775. 6:37:36in detail depth of it. So, considering
  10776. 6:37:38the way how the data points are being
  10777. 6:37:39segregated, we'll it will build a kind
  10778. 6:37:41of dendrogram. So, on top of this
  10779. 6:37:43dendrogram, your observations can be
  10780. 6:37:45classified here like this, as you can
  10781. 6:37:46see on the screen. Is K-means clustering
  10782. 6:37:48sensitive to outlier?
  10783. 6:37:50You need to understand one thing.
  10784. 6:37:52When we are talking about unsupervised
  10785. 6:37:54learning algorithms, as I said, you may
  10786. 6:37:56or may not have clarity on the data.
  10787. 6:37:59Which means
  10788. 6:38:01your assumption is that at the least
  10789. 6:38:02level
  10790. 6:38:04you don't have clarity on the data.
  10791. 6:38:06Then if you don't When you don't have
  10792. 6:38:08clarity on the data, how can you say
  10793. 6:38:09that This is an outlier or this is not
  10794. 6:38:11an outlier?
  10795. 6:38:12When you have clarity on the data,
  10796. 6:38:13that's especially when you're working on
  10797. 6:38:14supervised learning algorithms, you can.
  10798. 6:38:16But, you don't have output column also.
  10799. 6:38:18How you will be able to evaluate how the
  10800. 6:38:20so-and-so-called output column can be
  10801. 6:38:22can be evaluated because this is not
  10802. 6:38:24being classified properly, this is not
  10803. 6:38:25being clustered properly based on the
  10804. 6:38:27historical data because there is no
  10805. 6:38:28output column.
  10806. 6:38:29So, those type of concepts are something
  10807. 6:38:31which you don't need to worry about when
  10808. 6:38:32you are working on supervised learning.
  10809. 6:38:34They will be primarily they'll be
  10810. 6:38:35constrained when you are working on
  10811. 6:38:37supervised learning algorithms.
  10812. 6:38:38And of course, in case if you see that
  10813. 6:38:40there are outliers which are present,
  10814. 6:38:41obviously it has to be It will be
  10815. 6:38:43considered It will be considered as one
  10816. 6:38:44of the cluster in of any any of any of
  10817. 6:38:46the cluster it belongs to the data. But
  10818. 6:38:48anyhow, that's the characteristic of the
  10819. 6:38:49data. You don't need to
  10820. 6:38:51You don't need to take care because you
  10821. 6:38:52might be killing the actual original
  10822. 6:38:54values which are present in the data.
  10823. 6:38:55But you cannot expect that outlier can
  10824. 6:38:57be recognized
  10825. 6:38:58for all the cases that you have in
  10826. 6:38:59unsupervised learning.
  10827. 6:39:01So, the next type of clustering
  10828. 6:39:02technique that we have is division
  10829. 6:39:04clustering. So, what is this division
  10830. 6:39:05clustering? As a division clustering is
  10831. 6:39:07also creates the data in a form of
  10832. 6:39:09dendrogram. But the difference is you
  10833. 6:39:11can see that the starts with all data
  10834. 6:39:12points in one cluster, splits the root
  10835. 6:39:14into child recursively based on the
  10836. 6:39:16dendrogram, and stops when there is a
  10837. 6:39:18single term clusters which are created.
  10838. 6:39:19Which means that for every cluster there
  10839. 6:39:21will be one observation which belong to.
  10840. 6:39:23So, likewise the clustering technique
  10841. 6:39:24will work.
  10842. 6:39:25Now, there is one more technique that we
  10843. 6:39:27have in terms of building a clustering
  10844. 6:39:28technique, that's what we call it the
  10845. 6:39:29mean shift clustering. So, what we will
  10846. 6:39:31do we'll take an average of every
  10847. 6:39:33cluster that we have. What we will do
  10848. 6:39:35we'll end up with reducing their means
  10849. 6:39:36into the into a single density of items,
  10850. 6:39:39and then we'll continue you to repeat
  10851. 6:39:40the process to see to that which
  10852. 6:39:42observations mean will belong to the
  10853. 6:39:43same cluster, and according to that
  10854. 6:39:45we'll continue to identify which
  10855. 6:39:47observation will belong to a cluster a
  10856. 6:39:48particular cluster.
  10857. 6:39:50So, likewise we can apply different
  10858. 6:39:52types of clustering technique that can
  10859. 6:39:53help us to cluster the data which is
  10860. 6:39:55part of unsupervised learning.
  10861. 6:39:57Right? So, now let's take a let's jump
  10862. 6:40:00into something called a small hands-on.
  10863. 6:40:01Okay, let's let's take a small
  10864. 6:40:03Python example, and we will see how do
  10865. 6:40:05we build that particular a clustering
  10866. 6:40:07technique on top of the data using one
  10867. 6:40:09of the simple data that we have.
  10868. 6:40:11Jupiter notebook, let me open uh how
  10869. 6:40:15we can use scikit-learn using one of the
  10870. 6:40:17example. So, meanwhile let me show you
  10871. 6:40:19some of the data set also.
  10872. 6:40:22Let me show you a data set that I'm
  10873. 6:40:23going to use as well.
  10874. 6:40:25So, I'm What I'm going to do is I'm
  10875. 6:40:27going to take this movie metadata
  10876. 6:40:28information. Let me open this data set.
  10877. 6:40:32I'm going to take this example of movie
  10878. 6:40:34metadata information where this this
  10879. 6:40:37data set has got a number of
  10880. 6:40:38observations which are present.
  10881. 6:40:41Okay, let me open this. Okay, you can
  10882. 6:40:43see that there are a number of movies
  10883. 6:40:44related information as you can see the
  10884. 6:40:45movie names. So, Avatar, Pirates of the
  10885. 6:40:48Caribbean, Spectre, The Dark Knight,
  10886. 6:40:50Star Wars, etc. etc. John Carter,
  10887. 6:40:52Spider-Man 3, Tangled, or etc. etc. We
  10888. 6:40:55have a lot of movies.
  10889. 6:40:56And about every movie we got a lot of
  10890. 6:40:58information which is present like who is
  10891. 6:40:59the director, who is the actor, what are
  10892. 6:41:01the director Facebook likes, what are
  10893. 6:41:03the actor Facebook likes, likewise we
  10894. 6:41:04got we got a lot of information which is
  10895. 6:41:06present as part of these particular
  10896. 6:41:08every observation.
  10897. 6:41:10So, now what we will do, we will try to
  10898. 6:41:12cluster these data points. You can see a
  10899. 6:41:13lot of observations which are given.
  10900. 6:41:15What is the gross of the movie? What is
  10901. 6:41:16the number of reviewers? What is the
  10902. 6:41:17IMDb rating? What is the so-and-so
  10903. 6:41:19called IMDb score? What is the movie
  10904. 6:41:21span? What is the gross? What is the
  10905. 6:41:22budget? And everything.
  10906. 6:41:24So, now what I'll be doing, I'll be
  10907. 6:41:26reading this data set using one of the
  10908. 6:41:27Pandas library that we have. Okay, so
  10909. 6:41:30I'll be reading this data set. Let me
  10910. 6:41:31open the Jupiter notebook.
  10911. 6:41:33I'll be using something called Pandas.
  10912. 6:41:35So, as you can see, read.pandas.csv.
  10913. 6:41:37I'll be reading this data set where you
  10914. 6:41:38can see that this is the data set that
  10915. 6:41:39I'm able to read. I got a number of
  10916. 6:41:41observations that are present. I can see
  10917. 6:41:43that director Facebook likes and actor
  10918. 6:41:45three Facebook likes.
  10919. 6:41:47So, now what I what is it I want to do
  10920. 6:41:48is instead of building this custom
  10921. 6:41:50technique on top of every observation,
  10922. 6:41:52so what I will do is I will take this
  10923. 6:41:54columns called
  10924. 6:41:55called number of Facebook likes on
  10925. 6:41:57director and number of actor Facebook
  10926. 6:41:58likes versus director Facebook likes.
  10927. 6:42:00So, where I can see that if I want to
  10928. 6:42:02select a director Facebook likes alone,
  10929. 6:42:03I'll be able to choose this director
  10930. 6:42:05Facebook likes alone. You can see for
  10931. 6:42:07every movie you got the number of
  10932. 6:42:08director Facebook likes that are
  10933. 6:42:09present. So, if there are more number of
  10934. 6:42:11Facebook likes, what does it mean?
  10935. 6:42:13The director is famous person.
  10936. 6:42:15Correct?
  10937. 6:42:17Or rather if there are more number of
  10938. 6:42:18famous Facebook likes that the actor has
  10939. 6:42:20got, which means that the person is or
  10940. 6:42:22the actor is very famous person. That's
  10941. 6:42:24what you can understand. So, now you can
  10942. 6:42:25see that I'll try to extract these
  10943. 6:42:27independent component that we talking
  10944. 6:42:29about. I will extract all the records in
  10945. 6:42:31all the director Facebook likes versus
  10946. 6:42:33actor Facebook likes where I'm just
  10947. 6:42:35going to form an object called new data
  10948. 6:42:37by applying some I location as a filter.
  10949. 6:42:39I location I LOC stands for index
  10950. 6:42:41location where I can filter out what are
  10951. 6:42:43the records that I wanted to what are
  10952. 6:42:45the column that I wanted to using this I
  10953. 6:42:47location function.
  10954. 6:42:49Now I got all the so and so called
  10955. 6:42:51number of director Facebook likes versus
  10956. 6:42:52actor Facebook likes which are present
  10957. 6:42:54as part of this. Okay, so this has been
  10958. 6:42:56loaded into an object called new data.
  10959. 6:42:59So now what is that I'll be doing? After
  10960. 6:43:00that I'm importing something called SK
  10961. 6:43:02learn K means cluster. So this algorithm
  10962. 6:43:04is available as part of this
  10963. 6:43:05scikit-learn algorithm. What are the
  10964. 6:43:07number of steps that we have discussed
  10965. 6:43:09you are not supposed to execute all
  10966. 6:43:10these steps manually and of course if
  10967. 6:43:12you want you can also write such a
  10968. 6:43:13program as well, but what is that we'll
  10969. 6:43:15be doing in scikit-learn
  10970. 6:43:17library there is a Python library called
  10971. 6:43:19scikit-learn which has got most of the
  10972. 6:43:20algorithm present and we'll be importing
  10973. 6:43:22this K means clustering algorithm. And
  10974. 6:43:25for this K means clustering I'm
  10975. 6:43:26providing the C is equal to number of
  10976. 6:43:28clusters is equal to five.
  10977. 6:43:29So here if you are if you are aware of
  10978. 6:43:31uh
  10979. 6:43:32object-oriented programming using Python
  10980. 6:43:34you'll be able to correlate. I'm
  10981. 6:43:35importing this K means clustering which
  10982. 6:43:37is implemented as a class here where I'm
  10983. 6:43:39creating an object called K means
  10984. 6:43:40clustering by providing an input called
  10985. 6:43:42N number of scores clusters is equal to
  10986. 6:43:44five which means that what is the
  10987. 6:43:45meaning of five? I want to build a five
  10988. 6:43:47clusters out of this. So where once you
  10989. 6:43:49create an object using K means I'm
  10990. 6:43:51calling this method called fit the
  10991. 6:43:53method by providing new data as my
  10992. 6:43:55independent variables.
  10993. 6:43:57I'm for calling this fit method which
  10994. 6:43:59means fit is a method that we're going
  10995. 6:44:00to invoke what supposed to be the
  10996. 6:44:01process that needs to be executed that's
  10997. 6:44:04going to build my clustering technique
  10998. 6:44:05algorithm by taking number of clusters
  10999. 6:44:07is equal to five.
  11000. 6:44:08So now I'm able to generate my algorithm
  11001. 6:44:10where by looking at my model I'll be
  11002. 6:44:12able to extract what are the centroids
  11003. 6:44:13that I got final centroids because
  11004. 6:44:15initially you'll be taking some random
  11005. 6:44:16centroids, but at the end you'll end up
  11006. 6:44:18with getting a final centroid position
  11007. 6:44:20somewhere fixed to it that is nothing
  11008. 6:44:22but a center point for every cluster
  11009. 6:44:24that we got. These are the final
  11010. 6:44:25centroids that we got.
  11011. 6:44:27Even if I want to print the what is the
  11012. 6:44:29labels which are generated labels is
  11013. 6:44:30nothing but it will extract the outcome.
  11014. 6:44:32You can see that these are the label
  11015. 6:44:33numbers which are added here. Out of
  11016. 6:44:355,000 movies we got the labels which are
  11017. 6:44:37added as an array. But we
  11018. 6:44:39[clears throat] won't be able to see it
  11019. 6:44:39like that. What is it I'm trying to do?
  11020. 6:44:41I'm trying to get all the unique values
  11021. 6:44:43present in this labels with respect to
  11022. 6:44:45two counts. Now you can see that my data
  11023. 6:44:47is now clustered into five different All
  11024. 6:44:49the movies are clustered into five
  11025. 6:44:50different clusters as you can see.
  11026. 6:44:52Cluster number zero has got 4,700 Most
  11027. 6:44:55of the observations are moved into
  11028. 6:44:56cluster number zero.
  11029. 6:44:57104 movies went into observation number
  11030. 6:45:00one. 11 movies are moved into cluster
  11031. 6:45:02cluster number two. And 87 movies are
  11032. 6:45:04moved into cluster number three and 67
  11033. 6:45:06into cluster number four.
  11034. 6:45:08That is how the data properties are
  11035. 6:45:09being distributed and that's how the
  11036. 6:45:11clustering technique has divided the
  11037. 6:45:12data into five clusters.
  11038. 6:45:14Now you can see what is it I'm trying to
  11039. 6:45:15do? I'm trying to put this into a new
  11040. 6:45:17data of cluster which means what are the
  11041. 6:45:19labels that are generated here.
  11042. 6:45:21This is my output column. I'm going to
  11043. 6:45:23create this into as a new column in my
  11044. 6:45:25new data. Where I'm using this Allen
  11045. 6:45:27plot that will print whatever the
  11046. 6:45:29columns that I have called director
  11047. 6:45:31Facebook likes versus actor Facebook
  11048. 6:45:32likes.
  11049. 6:45:33And I'm choosing this data is equal to
  11050. 6:45:35new data that will help me to that will
  11051. 6:45:37help me to identify what column can be
  11052. 6:45:39considered as few so that I'll be
  11053. 6:45:42printing it in a cool warm type chart.
  11054. 6:45:44You can see that palette type is equal
  11055. 6:45:46to cool warm type which will help me to
  11056. 6:45:49identify based on the cluster that you
  11057. 6:45:52have created.
  11058. 6:45:53Now if I print this you can see that
  11059. 6:45:54pretty much the every observation is now
  11060. 6:45:56categorized into individual cluster you
  11061. 6:45:58can see. These are the movies which are
  11062. 6:45:59now graphical representation. This is
  11063. 6:46:01the graphical representation that we are
  11064. 6:46:03using to see how these movies are being
  11065. 6:46:05segregated. You can see these movies are
  11066. 6:46:07nothing but cluster number zero.
  11067. 6:46:09You can see there are a few more movies
  11068. 6:46:11which are extracted over here. This is
  11069. 6:46:12the nothing but based on the color
  11070. 6:46:14indication this is cluster number three.
  11071. 6:46:16And these movies are created as a
  11072. 6:46:17cluster number one.
  11073. 6:46:19Cluster number two. These movies are
  11074. 6:46:21something which are created as cluster
  11075. 6:46:22number one.
  11076. 6:46:24And these movies are created as cluster
  11077. 6:46:25number four.
  11078. 6:46:27Zero.
  11079. 6:46:31One.
  11080. 6:46:32This is nothing but two. Cluster number
  11081. 6:46:34three. Cluster number
  11082. 6:46:35three and then cluster number four like
  11083. 6:46:37this.
  11084. 6:46:38Now you can clearly see that how is the
  11085. 6:46:39cluster came clustering has came up. The
  11086. 6:46:41movies which are made by new people with
  11087. 6:46:44the new directors, new actors are making
  11088. 6:46:46films with new directors.
  11089. 6:46:48Or new directors are making films with
  11090. 6:46:50new actors.
  11091. 6:46:51These are all You will see more number
  11092. 6:46:53of movies will fall into this category
  11093. 6:46:54because you'll end up with getting new
  11094. 6:46:56people into the into the into the film
  11095. 6:46:57industry most of the cases.
  11096. 6:47:00Lot of movies are being made with new
  11097. 6:47:02directors with new actors.
  11098. 6:47:06And you can see that these movies are
  11099. 6:47:07the clusters which are being segregated
  11100. 6:47:09very clearly. The
  11101. 6:47:11famous actors are making films with new
  11102. 6:47:13directors. Very famous actors are making
  11103. 6:47:15films with new directors.
  11104. 6:47:17And these movies are something where
  11105. 6:47:18very famous directors are less I mean
  11106. 6:47:21average paying directors are making
  11107. 6:47:23films with
  11108. 6:47:24some new actors. You can see these
  11109. 6:47:26movies are nothing but very famous
  11110. 6:47:27actors are making directors are making
  11111. 6:47:29films with very
  11112. 6:47:31new actors.
  11113. 6:47:33You can see very famous directors are
  11114. 6:47:35making films with very famous actors.
  11115. 6:47:37Like Liber be making a film with a Tom
  11116. 6:47:40Cruise or something like that.
  11117. 6:47:42So likewise you can clearly see that
  11118. 6:47:44where instead of if you do this activity
  11119. 6:47:46manually it might take a little longer
  11120. 6:47:48time for you to segregate each and every
  11121. 6:47:49component. But within five minutes we
  11122. 6:47:51are able to cluster this activity.
  11123. 6:47:53Right? So that's what the beauty of
  11124. 6:47:54algorithms. You don't need to manually
  11125. 6:47:56do this activity where it can provide
  11126. 6:47:58your data automatically based on the
  11127. 6:47:59properties of the data your algorithm
  11128. 6:48:01itself will will this clustering output
  11129. 6:48:02out of the data.
  11130. 6:48:04All right. So that's how we will be able
  11131. 6:48:05to build a clustering technique on top
  11132. 6:48:07of the given data. So that's what we
  11133. 6:48:09have as part of one of the hands on
  11134. 6:48:10example.
  11135. 6:48:15>> [music]
  11136. 6:48:17>> What is hierarchical clustering?
  11137. 6:48:20So, hierarchical clustering, it is also
  11138. 6:48:21known as HCA or hierarchical cluster
  11139. 6:48:24analysis, and this is a method of
  11140. 6:48:26cluster analysis, as we have seen. So,
  11141. 6:48:28what happens here is that this
  11142. 6:48:30clustering allows us to build the tree
  11143. 6:48:33structure from data similarities, like
  11144. 6:48:35we have built X and Y, and we have
  11145. 6:48:38created trees, and these trees are
  11146. 6:48:40actually called known as dendrogram. So,
  11147. 6:48:42the way you represent a hierarchical
  11148. 6:48:44cluster or a hierarchical clustering is
  11149. 6:48:47through dendrogram. So, we actually drew
  11150. 6:48:49a dendrogram, okay? So, this is how the
  11151. 6:48:52clustering is being formed, and this is
  11152. 6:48:54how the clusters are being made, and
  11153. 6:48:56this is how the relationship among the
  11154. 6:48:58clusters is being shown by a dendrogram.
  11155. 6:49:01So, based on these things, now we will
  11156. 6:49:03go on further to understanding what is
  11157. 6:49:05agglomerative clustering. So, when this
  11158. 6:49:08hierarchical clustering follows a
  11159. 6:49:09bottom-up approach, this is called
  11160. 6:49:12agglomerative clustering. And when this
  11161. 6:49:14is following up a top-down approach,
  11162. 6:49:16this is used in divisive clustering. So,
  11163. 6:49:19now let us understand what is
  11164. 6:49:20agglomerative clustering. So, the types
  11165. 6:49:22of hierarchical clustering are two, that
  11166. 6:49:24is agglomerative and divisive. So, now
  11167. 6:49:26moving on further, what is agglomerative
  11168. 6:49:28clustering? So, agglomerative
  11169. 6:49:30hierarchical clustering, this is also
  11170. 6:49:32known as AGNES, which means
  11171. 6:49:33agglomerative nesting hierarchical
  11172. 6:49:35clustering, and it follows a bottom-up
  11173. 6:49:37approach, which means that clustering or
  11174. 6:49:40clusters, they are formed from the
  11175. 6:49:41bottom and are again clustered till a
  11176. 6:49:44complete single cluster is formed. And
  11177. 6:49:47what happens then is that the clustering
  11178. 6:49:50continues until we obtain a single
  11179. 6:49:52cluster, and we will see how we obtain a
  11180. 6:49:54single cluster, and we also represent
  11181. 6:49:56it. So, individual data points, they are
  11182. 6:49:58clustered based on similarity, and we go
  11183. 6:50:01on clustering until there is only one
  11184. 6:50:03single cluster left. So, let us just
  11185. 6:50:05plot this agglomerative clustering and
  11186. 6:50:07make things really simple for us.
  11187. 6:50:10Now, suppose I have got these data
  11188. 6:50:12points scattered here A to G and these
  11189. 6:50:15data points have to be clustered. So
  11190. 6:50:17another important thing is that now we
  11191. 6:50:19will form clusters. So how clusters are
  11192. 6:50:21formed? So we can see that based on some
  11193. 6:50:23similarity like because of the distance
  11194. 6:50:26nearby distance A and B can be grouped
  11195. 6:50:28together in a single cluster. So I'm
  11196. 6:50:31just doing that.
  11197. 6:50:32C and D I form another cluster because
  11198. 6:50:34they are near. So I just club them and
  11199. 6:50:37again I would just club E and F based on
  11200. 6:50:39their distance and G is separate so I
  11201. 6:50:42will just form a separate cluster. Now
  11202. 6:50:44what happens in agglomerative clustering
  11203. 6:50:47is that I have to plot all these data
  11204. 6:50:49points like A B C right? D
  11205. 6:50:54E F and G. So these are separate
  11206. 6:50:57clusters. The clustering starts from the
  11207. 6:50:59bottom and each data point is treated as
  11208. 6:51:02a single cluster which we will also
  11209. 6:51:04understand with the help of an example
  11210. 6:51:06further but now for simplicity let's
  11211. 6:51:08take A B C D. So then what happens is
  11212. 6:51:11that since A and B are grouped as one
  11213. 6:51:14cluster so this is how I just group them
  11214. 6:51:17all right? And C and D is grouped as one
  11215. 6:51:19cluster this is how I group them. E and
  11216. 6:51:21F are grouped in one single cluster.
  11217. 6:51:24This is how now
  11218. 6:51:26which is not been grouped into any of
  11219. 6:51:28the cluster. So now clustering I said
  11220. 6:51:30that it is it continues until a single
  11221. 6:51:33cluster is left. So now I would have to
  11222. 6:51:36have another level of clustering. That
  11223. 6:51:38means that E F G since G is very close
  11224. 6:51:42to E F I will cluster it in one single
  11225. 6:51:44cluster okay? And since I can see that
  11226. 6:51:47both these A B and C D pairs these
  11227. 6:51:49clusters are again together. So I will
  11228. 6:51:52just cross this line and I will make one
  11229. 6:51:55single cluster of these four points
  11230. 6:51:58right? So what I do is since these two
  11231. 6:52:00are connected I connect them with the
  11232. 6:52:02help of this line figure that is tree
  11233. 6:52:04structured dendrogram and this G E and
  11234. 6:52:08F, they are connected somehow. I connect
  11235. 6:52:09them. All right? Now, what happens is
  11236. 6:52:12that I have got two big clusters, and
  11237. 6:52:14clustering continues until a single
  11238. 6:52:16cluster is obtained. So, in the end, I
  11239. 6:52:19will have to cluster everything into a
  11240. 6:52:22single cluster, and this is how I do
  11241. 6:52:25that. And to join it, I will again join
  11242. 6:52:27this entire graph. So, this is when it
  11243. 6:52:31follows a bottom-up to up approach. This
  11244. 6:52:34is called as agglomerative hierarchical
  11245. 6:52:37clustering.
  11246. 6:52:38Okay? So, now we will go and see an
  11247. 6:52:41example of this hierarchical
  11248. 6:52:43agglomerative clustering, right? Okay.
  11249. 6:52:46So, let us understand what is
  11250. 6:52:48agglomerative clustering with this
  11251. 6:52:50example. Now, we see that here the
  11252. 6:52:53clustering takes place from bottom to
  11253. 6:52:55up, and we have taken an example of
  11254. 6:52:57population, wherein we go on clustering
  11255. 6:52:59until we get population. So, here from
  11256. 6:53:02the bottom, the individual professions
  11257. 6:53:04are being plotted, and we see that let's
  11258. 6:53:07let's take for a convenience the
  11259. 6:53:09left-hand side. And on the left-hand
  11260. 6:53:11side, the red dots, as you see, this is
  11261. 6:53:14individual profession in public sector.
  11262. 6:53:16And on the another side, which we see in
  11263. 6:53:19the brown circles, is the private sector
  11264. 6:53:21employment. So, somehow there's
  11265. 6:53:23similarity between private sector, so uh
  11266. 6:53:25they are being clustered as one single
  11267. 6:53:27cluster, and the private sector as a
  11268. 6:53:29another cluster.
  11269. 6:53:31These again are being clustered into one
  11270. 6:53:33single cluster, and that is employment
  11271. 6:53:35cluster. Whereas on the another side, we
  11272. 6:53:37can see another cluster, which is
  11273. 6:53:39different, and that is unemployed
  11274. 6:53:41section of cluster of people. Now, they
  11275. 6:53:44again share one similarity, and that is
  11276. 6:53:46that they all belong to a single gender,
  11277. 6:53:49that is male. So, everything is being
  11278. 6:53:51clustered into one single cluster, that
  11279. 6:53:53is male. And then again, male and female
  11280. 6:53:56clusters are being clustered together to
  11281. 6:53:58form one single cluster, that is
  11282. 6:54:00population. Similarly, we also divide on
  11283. 6:54:02the right-hand side the individual
  11284. 6:54:05professions of women clubbed into
  11285. 6:54:07private and public sector. And then
  11286. 6:54:09again, we have separate clusters of
  11287. 6:54:11employed and unemployed women. And they
  11288. 6:54:13have been grouped into one single
  11289. 6:54:15cluster and that is women. Again, we
  11290. 6:54:17merge the two big clusters into one
  11291. 6:54:20single cluster that is population. So,
  11292. 6:54:23this is how the clustering is taking
  11293. 6:54:25place. The levels are being increasing
  11294. 6:54:27from bottom to up. So, when we are using
  11295. 6:54:30this bottom-up approach, this is
  11296. 6:54:31agglomerative clustering.
  11297. 6:54:35>> [music]
  11298. 6:54:39>> Now, many of us have visited retail
  11299. 6:54:41shops such as Walmart or Target for our
  11300. 6:54:43household needs. Or let's say that we
  11301. 6:54:45are planning to buy the new iPhone from
  11302. 6:54:47Target.
  11303. 6:54:48What we would typically do is search for
  11304. 6:54:50the model by visiting the mobile section
  11305. 6:54:52of the store and then select the product
  11306. 6:54:54and head towards the billing counter.
  11307. 6:54:57But in today's world, the goal of the
  11308. 6:54:59organization is to increase the revenue.
  11309. 6:55:01Can this be done by just pitching one
  11310. 6:55:03product at a time to the customer? Now,
  11311. 6:55:05the answer to this is clearly no.
  11312. 6:55:07Hence, organization began mining data
  11313. 6:55:10relating to frequently bought items.
  11314. 6:55:13So, market basket analysis is one of the
  11315. 6:55:15key techniques used by large retailers
  11316. 6:55:18to uncover associations between items.
  11317. 6:55:21Now, examples could be the customers who
  11318. 6:55:23purchase bread have a 60% likelihood to
  11319. 6:55:26also purchase jam.
  11320. 6:55:28Customers who purchase laptops are more
  11321. 6:55:30likely to purchase laptop bags as well.
  11322. 6:55:33They try to find out associations
  11323. 6:55:35between different items and products
  11324. 6:55:37that can be sold together, which gives
  11325. 6:55:40assisting in the right product
  11326. 6:55:41placement.
  11327. 6:55:43Typically, it figures out what products
  11328. 6:55:45are being bought together and
  11329. 6:55:47organizations can place products in a
  11330. 6:55:49similar manner.
  11331. 6:55:50For example, people who buy bread also
  11332. 6:55:53tend to buy butter, right?
  11333. 6:55:55And the marketing team at retail stores
  11334. 6:55:57should target customers who buy bread
  11335. 6:56:00and butter and provide an offer to them
  11336. 6:56:02so that they buy a third item, suppose
  11337. 6:56:05eggs.
  11338. 6:56:06So, if a customer buys bread and butter
  11339. 6:56:08and sees a discount offer on eggs, he
  11340. 6:56:10will be encouraged to spend more and buy
  11341. 6:56:12the eggs. And this is what market basket
  11342. 6:56:15analysis is all about. This is what we
  11343. 6:56:17are going to talk about in this session,
  11344. 6:56:19which is association rule mining and the
  11345. 6:56:21a priori algorithm.
  11346. 6:56:23Now, association rule can be thought of
  11347. 6:56:25as an if-then relationship. Just to
  11348. 6:56:28elaborate on that, we have come up with
  11349. 6:56:31a rule, suppose if an item A is being
  11350. 6:56:33bought by the customer, then the chances
  11351. 6:56:35of item B being picked by the customer
  11352. 6:56:38too under the same transaction ID is
  11353. 6:56:40found out. You need to understand here
  11354. 6:56:43that it's not a causality, rather it's a
  11355. 6:56:46co-occurrence pattern that comes to the
  11356. 6:56:48force.
  11357. 6:56:49Now, there are two elements to this
  11358. 6:56:50rule. First is the if and second is the
  11359. 6:56:53then.
  11360. 6:56:54Now, if is also known as antecedent.
  11361. 6:56:57This is an item or a group of items that
  11362. 6:57:00are typically found in the item set. And
  11363. 6:57:02the later one is called the consequent.
  11364. 6:57:06This comes along as an item with an
  11365. 6:57:08antecedent group or the group of
  11366. 6:57:10antecedents are purchased.
  11367. 6:57:13Now, if you look at the image here A
  11368. 6:57:14arrow B, it means that if a person buys
  11369. 6:57:17an item A, then he will also buy an item
  11370. 6:57:19B. Or he will most probably buy an item
  11371. 6:57:21B.
  11372. 6:57:22Now, the simple example that I gave you
  11373. 6:57:24about the bread and butter and the eggs
  11374. 6:57:27is just a small example. But what if you
  11375. 6:57:29have thousands and thousands of items?
  11376. 6:57:32If you go to any professional data
  11377. 6:57:34scientist with that data, you can just
  11378. 6:57:36imagine how much of profit you can make
  11379. 6:57:39if the data scientist provides you with
  11380. 6:57:41the right examples and the right
  11381. 6:57:42placement of the items which you can do.
  11382. 6:57:45And you can get a lot of insights. That
  11383. 6:57:47is why association rule mining is a very
  11384. 6:57:50good algorithm which helps the business
  11385. 6:57:52make profit. So, let's see how this
  11386. 6:57:55algorithm works.
  11387. 6:57:56So, association rule mining is all about
  11388. 6:57:58building the rules. And we have just
  11389. 6:58:01seen one rule that if you buy A, then
  11390. 6:58:05there's a slight possibility or there's
  11391. 6:58:07a chance that you might buy B also. This
  11392. 6:58:10type of relationship in which we can
  11393. 6:58:12find the relationship between these two
  11394. 6:58:14items is known as single cardinality.
  11395. 6:58:17But what if the customer who bought A
  11396. 6:58:20and B also wants to buy C?
  11397. 6:58:23Or if a customer who bought A, B, and C
  11398. 6:58:25also wants to buy D? Then in these
  11399. 6:58:28cases, the cardinality usually increases
  11400. 6:58:30and we can have a lot of combination
  11401. 6:58:33around these data.
  11402. 6:58:35And if you have around 10,000 or more
  11403. 6:58:37than 10,000 data or items, just imagine
  11404. 6:58:40how many rules you're going to create
  11405. 6:58:42for each product. That is why
  11406. 6:58:44association rule mining has such
  11407. 6:58:47measures so that we do not end up
  11408. 6:58:49creating tens of thousands of rules.
  11409. 6:58:52Now, that is where the a priori
  11410. 6:58:54algorithm comes in. But before we get
  11411. 6:58:56into the a priori algorithm, let's
  11412. 6:58:58understand what's the maths behind it.
  11413. 6:59:01Now, there are three types of matrices
  11414. 6:59:03which help to measure the association.
  11415. 6:59:06We have support, confidence, and lift.
  11416. 6:59:09So, support is the frequency of item A
  11417. 6:59:11or the combination of item A or B.
  11418. 6:59:14It's basically the frequency of the
  11419. 6:59:16items which we have bought and what are
  11420. 6:59:18the combination of the frequency of the
  11421. 6:59:20item we have bought. So, with this, what
  11422. 6:59:22we can do is filter out the items which
  11423. 6:59:25have been bought less frequently.
  11424. 6:59:28This is one of the measures which is
  11425. 6:59:29support.
  11426. 6:59:30Now, what confidence tells us?
  11427. 6:59:32So, confidence gives us how often the
  11428. 6:59:34items A and B occur together given the
  11429. 6:59:37number of times A occur.
  11430. 6:59:39Now, this also helps us solve a lot of
  11431. 6:59:41other problems because if somebody is
  11432. 6:59:43buying A and B together and not buying
  11433. 6:59:45C, we can just rule out C at that point
  11434. 6:59:47of time.
  11435. 6:59:48So, this solves another problem is that
  11436. 6:59:51we obviously do not need to analyze the
  11437. 6:59:53products which people just buy barely.
  11438. 6:59:56So, what we can do is according to the
  11439. 6:59:58sales, we can define our minimum support
  11440. 7:00:01and confidence. And when you have set
  11441. 7:00:03these values, we can put these values in
  11442. 7:00:05the algorithm and we can filter out the
  11443. 7:00:08data and we can create different rules.
  11444. 7:00:11And suppose even after filtering, you
  11445. 7:00:13have like 5,000 rules. And for every
  11446. 7:00:16item, we create these 5,000 rules. So,
  11447. 7:00:19that's practically impossible. So, for
  11448. 7:00:22that, we need the third calculation
  11449. 7:00:24which is the lift. So, lift is basically
  11450. 7:00:26the strength of any rule.
  11451. 7:00:28Now, let's have a look at the
  11452. 7:00:29denominator of the formula given here.
  11453. 7:00:32And if you see here, we have the
  11454. 7:00:34independent support values of A and B.
  11455. 7:00:37So, this gives us the independent
  11456. 7:00:39occurrence probability of A and B. And
  11457. 7:00:42obviously, there's a lot of difference
  11458. 7:00:44between this random occurrence and
  11459. 7:00:46association. And if the denominator of
  11460. 7:00:49the lift is more, what it means is that
  11461. 7:00:53the occurrence of randomness is more
  11462. 7:00:55rather than the occurrence because of
  11463. 7:00:58any association.
  11464. 7:00:59So, lift is the final verdict where we
  11465. 7:01:01know whether we have to spend time on
  11466. 7:01:03this particular rule what we have got
  11467. 7:01:06here or not. Now, let's have a look at a
  11468. 7:01:08simple example of association rule
  11469. 7:01:10mining.
  11470. 7:01:11So, suppose we have a set of items A, B,
  11471. 7:01:14C, D, and E and a set of transactions
  11472. 7:01:17T1, T2, T3, T4, and T5.
  11473. 7:01:20And as you can see here, we have the
  11474. 7:01:21transactions T1 in which we have ABC, T2
  11475. 7:01:24ACD, T3 BCD, T4 ADE, and T5 BCE.
  11476. 7:01:30Now, what we generally do is create some
  11477. 7:01:33rules or association rules such as A
  11478. 7:01:36gives D or C gives A, A gives C, B and C
  11479. 7:01:40gives A. What this basically means is
  11480. 7:01:43that if a person buys A, then he's most
  11481. 7:01:45likely to buy D. And if a person buys C,
  11482. 7:01:48then he's most likely to buy A. And if
  11483. 7:01:50you have a look at the last one, if a
  11484. 7:01:51person buys B and C, he's most likely to
  11485. 7:01:54buy the item A as well.
  11486. 7:01:56Now, if we calculate the support,
  11487. 7:01:58confidence, and lift using these rules,
  11488. 7:02:00as you can see here in the table, we
  11489. 7:02:02have the rule and the support,
  11490. 7:02:04confidence, and the lift values.
  11491. 7:02:06Now, let's discuss about a priori.
  11492. 7:02:09So, a priori algorithm uses the frequent
  11493. 7:02:12item sets to generate the association
  11494. 7:02:14rule.
  11495. 7:02:15And it is based on the concept that a
  11496. 7:02:17subset of a frequent item set must also
  11497. 7:02:20be a frequent item set itself.
  11498. 7:02:23Now, this raises the question, what
  11499. 7:02:24exactly is a frequent item set?
  11500. 7:02:26So, a frequent item set is an item set
  11501. 7:02:29whose support value is greater than the
  11502. 7:02:30threshold value. Now, just now we
  11503. 7:02:32discussed that the marketing team,
  11504. 7:02:34according to the sales, have a minimum
  11505. 7:02:36threshold value for the confidence as
  11506. 7:02:39well as the support.
  11507. 7:02:41So, frequent item set is that item set
  11508. 7:02:43whose support value is greater than the
  11509. 7:02:44threshold value already specified.
  11510. 7:02:47Now, example, if A and B is a frequent
  11511. 7:02:49item set, then A and B should also be
  11512. 7:02:52frequent item sets individually.
  11513. 7:02:55Now, let's consider the following
  11514. 7:02:56transaction to make the things a little
  11515. 7:02:59easier. Suppose we have transactions 1 2
  11516. 7:03:023 4 5 and these items are there.
  11517. 7:03:04So, T1 has 1 3 and 4, T2 has 2 3 and 5,
  11518. 7:03:08T3 has 1 2 3 5, T4 2 5, and T5 1 3 and
  11519. 7:03:125. Now, the first step is to build a
  11520. 7:03:15list of item sets of size one by using
  11521. 7:03:18this transactional data. And one thing
  11522. 7:03:20to note here is that the minimum support
  11523. 7:03:23count, which is given here, is two.
  11524. 7:03:25Let's suppose it's two.
  11525. 7:03:27So, the first step is to create item
  11526. 7:03:29sets of size one and calculate their
  11527. 7:03:31support values.
  11528. 7:03:32So, as you can see here, we have the
  11529. 7:03:33table C1 in which we have the item sets
  11530. 7:03:361 2 3 4 5, and the support values.
  11531. 7:03:39If you remember the formula of support,
  11532. 7:03:41it was frequency divided by the total
  11533. 7:03:43number of occurrence.
  11534. 7:03:45So, as you can see here, for the item
  11535. 7:03:46set 1, the support is three.
  11536. 7:03:49As you can see here, the item set 1
  11537. 7:03:51appears in T1, T3, and T5.
  11538. 7:03:54So, as you can see, its frequency is 1,
  11539. 7:03:562, and 3.
  11540. 7:03:57Now, as you can see here, the item set 4
  11541. 7:04:00has a support of 1 as it occurs only
  11542. 7:04:02once in transaction 1.
  11543. 7:04:04But, the minimum support value is two.
  11544. 7:04:06That's why it's going to be eliminated.
  11545. 7:04:09So, we have the final table, which is
  11546. 7:04:10the table F1, in which we have the item
  11547. 7:04:13sets 1, 2, 3, and 5, and we have the
  11548. 7:04:15support values 3, 3, 4, and 4.
  11549. 7:04:19Now, the next step is to create item
  11550. 7:04:20sets of size two and calculate the
  11551. 7:04:22support values.
  11552. 7:04:24Now, all the combination of the item
  11553. 7:04:25sets in the F1, which is the final table
  11554. 7:04:29in which you discarded the four, are
  11555. 7:04:31going to be used for this iteration. So,
  11556. 7:04:33we get the table C2. So, as you can see
  11557. 7:04:35here, we have 1, 2, 1, 3, 1, 5, 2, 3, 2,
  11558. 7:04:395, and 3, 5.
  11559. 7:04:40Now, if you calculate the support here
  11560. 7:04:42again, we can see that the item set 1, 2
  11561. 7:04:45has a support of 1, which is again less
  11562. 7:04:48than the specified threshold. So, we're
  11563. 7:04:51going to discard that.
  11564. 7:04:52So, if we have a look at the table F2,
  11565. 7:04:55we have 1, 3, 1, 5, 2, 3, 2, 5, and 3,
  11566. 7:04:595.
  11567. 7:05:00Again, we're going to move forward and
  11568. 7:05:02create the item set of size three and
  11569. 7:05:05calculate the support values.
  11570. 7:05:07Now, all the combinations are going to
  11571. 7:05:08be used from the item set F2 for this
  11572. 7:05:11particular iterations.
  11573. 7:05:13Now, before calculating support values,
  11574. 7:05:15let's perform pruning on the data set.
  11575. 7:05:18Now, what is pruning? Now, after the
  11576. 7:05:20combinations are being made, we divide
  11577. 7:05:21C3 item sets to check if there is
  11578. 7:05:23another subset whose support is less
  11579. 7:05:26than the minimum support value.
  11580. 7:05:28That is what frequent item set means.
  11581. 7:05:31So, if you have a look here, the item
  11582. 7:05:33sets we have is 1 2 3, 1 2, 1 3, 2 3 for
  11583. 7:05:38the first one. Because as you can see
  11584. 7:05:40here, if we have a look at the subsets
  11585. 7:05:42of 1 2 3, we have 1 {comma} 2 as well.
  11586. 7:05:46So, we are going to discard this whole
  11587. 7:05:48item set.
  11588. 7:05:49Same goes for the second one. We have 1
  11589. 7:05:512 5. We have 1 2 in that, which was
  11590. 7:05:53discarded in the previous set or the
  11591. 7:05:55previous step. That's why we're going to
  11592. 7:05:57discard that also.
  11593. 7:05:59Which leaves us with only two factors,
  11594. 7:06:01which is 1 3 5 item set and the 2 3 5.
  11595. 7:06:05And the support for this is two and two
  11596. 7:06:07as well.
  11597. 7:06:08Now, if we create the table C4 using
  11598. 7:06:12four elements, we're going to have only
  11599. 7:06:14one item set, which is 1 2 3 and 5. And
  11600. 7:06:18if we have a look at the table here, the
  11601. 7:06:20transaction table, 1 2 3 and 5 appears
  11602. 7:06:23only once. So, the support is one.
  11603. 7:06:25And since C4, the support of the whole
  11604. 7:06:28table C4 is less than two, so we're
  11605. 7:06:30going to stop here and return to the
  11606. 7:06:32previous item set that is three.
  11607. 7:06:34So, the frequent item sets are 1 3 5 and
  11608. 7:06:372 3 5.
  11609. 7:06:38Now, let's assume our minimum confidence
  11610. 7:06:40value is 60%.
  11611. 7:06:42For that, we're going to generate all
  11612. 7:06:43the non-empty subsets for each frequent
  11613. 7:06:46item sets.
  11614. 7:06:47Now, for I equals 1 {comma} 3 {comma} 5,
  11615. 7:06:50which is the item set, we get the subset
  11616. 7:06:521 3, 1 5, 3 5, 1, 3, and 5.
  11617. 7:06:58Similarly, for 2 3 5, we get 2 3, 2 5, 3
  11618. 7:07:015, 2, 3, and 5.
  11619. 7:07:04Now, this rule states that for every
  11620. 7:07:06subset S of I, the output of the rule
  11621. 7:07:09gives something like S gives I to S.
  11622. 7:07:13That implies S recommends I of S.
  11623. 7:07:16And this is only possible if the support
  11624. 7:07:18of I divided by the support of S is
  11625. 7:07:20greater than equal to the minimum
  11626. 7:07:22confidence value.
  11627. 7:07:24Now, applying these rules to the item
  11628. 7:07:26set of F3, we get rule one, which is 1,3
  11629. 7:07:30gives 1,3,5
  11630. 7:07:32and 1,3.
  11631. 7:07:33It means one and three gives five.
  11632. 7:07:36So, the confidence is equal to the
  11633. 7:07:39support of 1,3,5 / support of 1,3, that
  11634. 7:07:44equals 2 / 3, which is 66% and which is
  11635. 7:07:47greater than the 60%.
  11636. 7:07:49So, the rule one is selected.
  11637. 7:07:51Now, if we come to rule two, which is
  11638. 7:07:531,5, it gives 1,3,5
  11639. 7:07:56and 1,5. It means if we have one and
  11640. 7:07:59five, it implies we also going to have
  11641. 7:08:02three. Now, if we calculate the
  11642. 7:08:03confidence of this one, we're going to
  11643. 7:08:05have support 1,3,5 / support 1,5, which
  11644. 7:08:08gives us 100%, which means rule two is
  11645. 7:08:11selected as well. But again, if you have
  11646. 7:08:13a look at rule five and rule six over
  11647. 7:08:15here, similarly, if it select three
  11648. 7:08:18gives 1,3,5 and three, it means if you
  11649. 7:08:21have three, we also get one and five.
  11650. 7:08:23So, the confidence for this comes at
  11651. 7:08:2550%, which is less than the given 60%
  11652. 7:08:29target, so we're going to reject this
  11653. 7:08:31rule. And same goes for the rule number
  11654. 7:08:33six.
  11655. 7:08:34Now, one thing to keep in mind here is
  11656. 7:08:37that although the rule one and rule five
  11657. 7:08:39look a lot similar, they are not.
  11658. 7:08:42So, it really depends what's on the
  11659. 7:08:44left-hand side of the arrow and what's
  11660. 7:08:45on the right-hand side of the arrow.
  11661. 7:08:47It's the if-then possibility.
  11662. 7:08:49I'm sure you guys can understand what
  11663. 7:08:52exactly these rules are and how to
  11664. 7:08:54proceed with the rules.
  11665. 7:08:55So, let's see how we can implement the
  11666. 7:08:57same in Python, right?
  11667. 7:09:00So, for that, what I'm going to do is
  11668. 7:09:01create a new Python file and
  11669. 7:09:06I'm going to use the Jupyter Notebook.
  11670. 7:09:08You're free to use any sort of IDE.
  11671. 7:09:11I'm going to name it as Apriori.
  11672. 7:09:14So, the first thing what we're going to
  11673. 7:09:16do is we'll be using the online
  11674. 7:09:19transactional data of a retail store for
  11675. 7:09:21generating association rules. So,
  11676. 7:09:23firstly, what we need to do is get the
  11677. 7:09:25pandas and mlxtend libraries imported
  11678. 7:09:27and read the file.
  11679. 7:09:30So, as you can see here, we are using
  11680. 7:09:32the online retail.xlsx
  11681. 7:09:34format file.
  11682. 7:09:36And from mlxtend, we're going to import
  11683. 7:09:38a priori and association rules. It all
  11684. 7:09:40comes under mlxtend.
  11685. 7:09:43So, as you can see here, we have the
  11686. 7:09:45invoice, the stock code, the
  11687. 7:09:47description, the quantity, the invoice
  11688. 7:09:49data, unit price, customer ID, and the
  11689. 7:09:52country.
  11690. 7:09:53Now, next in this step, what we're going
  11691. 7:09:54to do is do data cleanup, which includes
  11692. 7:09:57removing the spaces from some of the
  11693. 7:09:59descriptions, and drop the rows that do
  11694. 7:10:01not have invoice numbers, and remove the
  11695. 7:10:03credit card transactions, because that
  11696. 7:10:06is of no use to us.
  11697. 7:10:13So, as you can see here, the output in
  11698. 7:10:15which we have like 532,000
  11699. 7:10:19rows with eight columns.
  11700. 7:10:21So, after the cleanup, we need to
  11701. 7:10:23consolidate the items into one
  11702. 7:10:25transaction per row with each product.
  11703. 7:10:27For the sake of keeping the data set
  11704. 7:10:29small, we are only looking at the sales
  11705. 7:10:31for France.
  11706. 7:10:36So, as you can see here, we have
  11707. 7:10:38excluded all the other sales. We are
  11708. 7:10:40just looking at the sales for France.
  11709. 7:10:42Now, there are a lot of zeros in the
  11710. 7:10:44data, but we also need to make sure any
  11711. 7:10:46positive values are converted to one,
  11712. 7:10:48and anything less than zero is set to
  11713. 7:10:49zero.
  11714. 7:10:52So, as you can see here, we have still
  11715. 7:10:53392 rows.
  11716. 7:10:56We're going to encode it and see.
  11717. 7:10:59Check again.
  11718. 7:11:00Now that you have structured the data
  11719. 7:11:02properly, in this step, what we're going
  11720. 7:11:03to do is generate frequent item sets
  11721. 7:11:05that have support at least 7%.
  11722. 7:11:08Now, this number is chosen so that you
  11723. 7:11:10can get close enough and generate the
  11724. 7:11:12rules with the corresponding support,
  11725. 7:11:13confidence, and lift.
  11726. 7:11:19So, guys, as you can see here, the
  11727. 7:11:20minimum support is 0.7. And what if we
  11728. 7:11:23add another constraint on the rules,
  11729. 7:11:25such as the lift is greater than six and
  11730. 7:11:28the confidence is greater than 0.8?
  11731. 7:11:31So as you can see here, we have the
  11732. 7:11:33left-hand side and the right-hand side
  11733. 7:11:34of the association rule, which is the
  11734. 7:11:36antecedent and the consequence.
  11735. 7:11:39We have the support, we have the
  11736. 7:11:40confidence, the lift, the leverage, and
  11737. 7:11:42the conviction.
  11738. 7:11:43So guys, that's it for this session.
  11739. 7:11:45That is how you create association rules
  11740. 7:11:48using the Apriori algorithm, which helps
  11741. 7:11:50a lot in the marketing business.
  11742. 7:11:53It runs on the principle of market
  11743. 7:11:55basket analysis, which is exactly what
  11744. 7:11:57big companies like Walmart, you have
  11745. 7:11:59Reliance, and Target too. Even IKEA does
  11746. 7:12:03it.
  11747. 7:12:04>> [music]
  11748. 7:12:10>> But the first thing that we need to
  11749. 7:12:11focus on is why artificial intelligence.
  11750. 7:12:13Why do we need artificial intelligence?
  11751. 7:12:16This with an example. So nowadays, if
  11752. 7:12:18you have noticed, if your car exceeds
  11753. 7:12:20the speed limit, so you'll get a letter
  11754. 7:12:22or basically a challan at your home. How
  11755. 7:12:24do you think that happens? Do you think
  11756. 7:12:26that there's a person who is sitting in
  11757. 7:12:27a chair and actually noting down all the
  11758. 7:12:29number plates that crosses the speed
  11759. 7:12:31limit? Well, that is not possible
  11760. 7:12:33because there might be millions of cars
  11761. 7:12:35that pass through that road. And at
  11762. 7:12:36once, there might be many cars that will
  11763. 7:12:38be passing through that road. So for a
  11764. 7:12:40human being to actually do this task is
  11765. 7:12:42next to impossible. Now, let us see
  11766. 7:12:44another approach to this particular
  11767. 7:12:46problem. So what we can do, we can
  11768. 7:12:48actually make use of cameras that will
  11769. 7:12:50click the picture of the car that
  11770. 7:12:52exceeds the speed limit. And then we
  11771. 7:12:54could convert that picture into a text.
  11772. 7:12:56For example, we have a UK plate.
  11773. 7:12:59So in this way, the human error, the
  11774. 7:13:01risk of human error has been reduced.
  11775. 7:13:03And at the same time, machines, they
  11776. 7:13:05never get tired. So because of that, you
  11777. 7:13:07can capture all the images of cars that
  11778. 7:13:09actually crosses the speed limit.
  11779. 7:13:11Similarly, you can think of uh many
  11780. 7:13:13other examples as well. It is used in
  11781. 7:13:15order to recognize a sign that is in
  11782. 7:13:16banks if you want to authenticate
  11783. 7:13:18whether that person is the bank customer
  11784. 7:13:20or not. Yeah, apart from that, it is
  11785. 7:13:22even used for self-driving cars as well.
  11786. 7:13:24So, in US around 30,000 people die every
  11787. 7:13:27year because of road accidents. So, that
  11788. 7:13:29can be completely removed if we use the
  11789. 7:13:31self-driving cars, which is based on the
  11790. 7:13:33concept of artificial intelligence. And
  11791. 7:13:35let me tell you guys, you might find it
  11792. 7:13:36very fascinating that people in MIT are
  11793. 7:13:39using artificial intelligence in order
  11794. 7:13:41to predict the future. So, you can
  11795. 7:13:43imagine why we need artificial
  11796. 7:13:44intelligence. If you have any questions,
  11797. 7:13:46any doubts, you can ask me. It is even
  11798. 7:13:48used in places where humans can't reach.
  11799. 7:13:50For example, uh deep oceans or
  11800. 7:13:52navigation in Mars. So, in those places
  11801. 7:13:55we need machines which are smart enough
  11802. 7:13:57to carry out tasks.
  11803. 7:13:58So, this is why we need artificial
  11804. 7:14:00intelligence. Let us move forward and
  11805. 7:14:01understand what exactly is artificial
  11806. 7:14:03intelligence.
  11807. 7:14:05Now, artificial intelligence, I know the
  11808. 7:14:06word sounds pretty complex and there are
  11809. 7:14:08a lot of Hollywood movies that are based
  11810. 7:14:10on artificial intelligence. If you have
  11811. 7:14:12seen Terminator or Matrix, all these
  11812. 7:14:14movies are based on artificial
  11813. 7:14:15intelligence. But, you don't need to
  11814. 7:14:17worry about it because till now we
  11815. 7:14:18haven't reached that level as they have
  11816. 7:14:20shown in movies like Terminator.
  11817. 7:14:22But, yeah, the concept is pretty
  11818. 7:14:23similar. So, basically, we want systems
  11819. 7:14:26and softwares in such a way that they
  11820. 7:14:28can imitate the human behavior.
  11821. 7:14:30Now, what happens in artificial
  11822. 7:14:31intelligence? Artificial intelligence is
  11823. 7:14:34accomplished by studying how human brain
  11824. 7:14:36thinks and how human brain learns,
  11825. 7:14:38decide, and work while trying to solve a
  11826. 7:14:41problem. And then we use outcome of this
  11827. 7:14:43study as the basis of development of
  11828. 7:14:45intelligent software and systems.
  11829. 7:14:48So, our major goal is to have systems or
  11830. 7:14:50softwares that can imitate the human
  11831. 7:14:53behavior. The way they think, the way
  11832. 7:14:55they decide, the way they solve a
  11833. 7:14:57problem. So, in that similar fashion, we
  11834. 7:14:59want our machines to do that. So, this
  11835. 7:15:02is basically artificial intelligence in
  11836. 7:15:03a nutshell.
  11837. 7:15:05Let us move forward and look at various
  11838. 7:15:06applications of artificial intelligence.
  11839. 7:15:09So, this slide basically talks about the
  11840. 7:15:11application of artificial intelligence.
  11841. 7:15:13Now, I've listed only three of them, but
  11842. 7:15:15there are millions of applications. For
  11843. 7:15:17example, it is used in speech
  11844. 7:15:18recognition. So, whenever you search
  11845. 7:15:20something on Google, so you can just
  11846. 7:15:21tell Google and it'll search it for you.
  11847. 7:15:23Similarly, it is used for understanding
  11848. 7:15:24natural language as well as for image
  11849. 7:15:26recognition as well. And there are many,
  11850. 7:15:28many other applications in which
  11851. 7:15:30artificial intelligence finds its use.
  11852. 7:15:32For example, it can be used in
  11853. 7:15:34self-driving cars. It can be used in
  11854. 7:15:36Siri for recommending some products. And
  11855. 7:15:38even when you go to websites like
  11856. 7:15:39YouTube or Pandora, YouTube knows which
  11857. 7:15:41video you want to next. Pandora knows
  11858. 7:15:43which song you want to listen to. How do
  11859. 7:15:45you think this happens? It happens all
  11860. 7:15:47because of artificial intelligence. So,
  11861. 7:15:49all of these are a few examples of
  11862. 7:15:51artificial intelligence, but nowadays,
  11863. 7:15:53it is used almost everywhere, guys.
  11864. 7:15:55Trust me on that. Now, let us move
  11865. 7:15:57forward and understand how to achieve
  11866. 7:15:59artificial intelligence.
  11867. 7:16:01Now, in order to achieve artificial
  11868. 7:16:02intelligence, there were a few
  11869. 7:16:04technologies that First came a machine
  11870. 7:16:06learning. Now, there were certain
  11871. 7:16:08limitations of machine learning. In
  11872. 7:16:10order to overcome those limitations,
  11873. 7:16:12came a deep learning. Now, let me tell
  11874. 7:16:14you guys, the concept of artificial
  11875. 7:16:16intelligence is not new. It was first
  11876. 7:16:18coined in 1956,
  11877. 7:16:20but it was just a theoretical concept.
  11878. 7:16:22Then in '80s and '90s, we were talking
  11879. 7:16:24about neural networks. But since we
  11880. 7:16:26didn't have enough computational power,
  11881. 7:16:29so we couldn't utilize it properly. But
  11882. 7:16:31in late '90s and 2000s, we started using
  11883. 7:16:33neural networks for machine learning.
  11884. 7:16:35Then in 2006, the term deep learning was
  11885. 7:16:38coined for the first time that overcame
  11886. 7:16:40the limitations of machine learning. And
  11887. 7:16:42from 2010, deep learning was used
  11888. 7:16:44commercially as well. So, this was just
  11889. 7:16:47a small history about artificial
  11890. 7:16:48intelligence, machine learning, and deep
  11891. 7:16:50learning. Now, in order to understand
  11892. 7:16:52this deep learning, we need to first
  11893. 7:16:54look at machine learning and what were
  11894. 7:16:55the biggest limitations of machine
  11895. 7:16:57learning that led to the evolution of
  11896. 7:16:58deep learning. So, we'll move forward
  11897. 7:17:00and understand what exactly is a machine
  11898. 7:17:03learning.
  11899. 7:17:04Now, what is machine learning? So,
  11900. 7:17:05machine learning is nothing but a type
  11901. 7:17:07of artificial intelligence or you can
  11902. 7:17:09say a subset of artificial intelligence.
  11903. 7:17:11And it provides computers with the
  11904. 7:17:13ability to learn without being
  11905. 7:17:15explicitly programmed. So, you don't
  11906. 7:17:16need to hardcode your machine for that.
  11907. 7:17:19Let us understand this with an example.
  11908. 7:17:21So, we have a problem statement in which
  11909. 7:17:23whenever you give a certain input, we
  11910. 7:17:25need to determine the species of the
  11911. 7:17:26flower. And what is that input? That
  11912. 7:17:28input will be sepal length, sepal width,
  11913. 7:17:31petal length, and petal width. So,
  11914. 7:17:33whenever we get these four parameters or
  11915. 7:17:35these four variables, our machine should
  11916. 7:17:37be able to predict what sort of a flower
  11917. 7:17:39it is. Now, how do you think that will
  11918. 7:17:41happen? First, what we need to do, we
  11919. 7:17:44need to train our machine on the basis
  11920. 7:17:46of the data that we have. So, in this
  11921. 7:17:49data, we have sepal length, sepal width,
  11922. 7:17:51petal length, and petal width, and we
  11923. 7:17:52have species. So, our machine will learn
  11924. 7:17:55from this data. It'll determine what
  11925. 7:17:57should be the length and width of the
  11926. 7:17:59sepal and petal in order to classify it
  11927. 7:18:01as setosa or versicolor or other species
  11928. 7:18:04of flowers as well. Now, what happens
  11929. 7:18:06next? So, you have trained your data.
  11930. 7:18:08So, you have trained your machine from
  11931. 7:18:09the data set. Then, what happens?
  11932. 7:18:11Whenever you give a new input to this
  11933. 7:18:13particular machine, it'll predict the
  11934. 7:18:15species of the flower.
  11935. 7:18:17So, this is how machine learning works.
  11936. 7:18:18It is nothing but machine learning in a
  11937. 7:18:19nutshell.
  11938. 7:18:21So, basically, I'll just revise it once
  11939. 7:18:23more.
  11940. 7:18:24So, you have a data set. So, you split
  11941. 7:18:26that data into training and testing
  11942. 7:18:28data. So, what happens with the help of
  11943. 7:18:30training data? You train your particular
  11944. 7:18:32machine. And after that, you test it in
  11945. 7:18:35order to determine the accuracy. And
  11946. 7:18:37once it is done, whenever you give the
  11947. 7:18:39new input, it'll predict the outcome or
  11948. 7:18:41the desired outcome. So, this is how
  11949. 7:18:43machine learning works, guys. Let us
  11950. 7:18:45move forward and understand various
  11951. 7:18:47types of machine learning.
  11952. 7:18:49So, the first type is called supervised
  11953. 7:18:51learning. Now, in supervised learning,
  11954. 7:18:54what happens? You have input variables X
  11955. 7:18:56and an output variable Y.
  11956. 7:18:58And you can use an algorithm to learn
  11957. 7:19:00mapping function from the input to the
  11958. 7:19:02output. Now, let me simplify it for you.
  11959. 7:19:04So, what happens in supervised learning,
  11960. 7:19:06the data that you have already contain
  11961. 7:19:09the classification. Now, let me talk
  11962. 7:19:11about the previous example itself. So,
  11963. 7:19:13from our data set, we knew that if we
  11964. 7:19:15have this width, this length of our
  11965. 7:19:17sepal and petal,
  11966. 7:19:19so that will be the species of flower.
  11967. 7:19:21So, the classifications are already
  11968. 7:19:23defined. So, that will be under
  11969. 7:19:25supervised learning. Now, let me tell
  11970. 7:19:27you how it actually works.
  11971. 7:19:28So, you have data. You divide that data
  11972. 7:19:30into training data as well as test data.
  11973. 7:19:33So, on the basis of this training data,
  11974. 7:19:35you train your machine. And after that,
  11975. 7:19:37you create a model. So, as you can see
  11976. 7:19:39that this phase is called training
  11977. 7:19:40phase. And after that, you create a
  11978. 7:19:42model. Now, in order to check this model
  11979. 7:19:45to get the accuracy, you have test data.
  11980. 7:19:47So, you'll pass that test data and
  11981. 7:19:49you'll see the accuracy. That is nothing
  11982. 7:19:51but the actual output minus the output
  11983. 7:19:54that is present in the test data. So,
  11984. 7:19:56with that, you can get the accuracy. So,
  11985. 7:19:58this is nothing but uh supervised
  11986. 7:20:00learning. And if you have any questions,
  11987. 7:20:02you can ask me right now.
  11988. 7:20:04So, we'll move forward and understand
  11989. 7:20:06unsupervised learning.
  11990. 7:20:07Now, in unsupervised learning, unlike
  11991. 7:20:09supervised learning, you don't have any
  11992. 7:20:11predefined classes. So, what happens,
  11993. 7:20:12you have data. So, on the basis of that
  11994. 7:20:15data, you try to create your own class.
  11995. 7:20:18You try to make sure that whatever class
  11996. 7:20:20you create has high intra-class
  11997. 7:20:21similarities and have a low inter-class
  11998. 7:20:24similarities. That means, if I've
  11999. 7:20:26created two class like this, class one
  12000. 7:20:28and class two, so the elements of this
  12001. 7:20:31particular class should have high
  12002. 7:20:33similarity, but at the same time, it
  12003. 7:20:35should have low similarity with the
  12004. 7:20:37elements of class two.
  12005. 7:20:39So, you can think of examples as well of
  12006. 7:20:40unsupervised learning. For example, if I
  12007. 7:20:43have a data about my customers. So, if I
  12008. 7:20:46have a website and there are millions of
  12009. 7:20:47visitors on my website, and I want to
  12010. 7:20:49make sure that I group people on various
  12011. 7:20:52criteria. For example, I can group
  12012. 7:20:54people on the basis of willingness to
  12013. 7:20:56purchase a product that is there on my
  12014. 7:20:57website or where they are coming from,
  12015. 7:21:00what is the source, all those things.
  12016. 7:21:01So, I want to group my customers and I
  12017. 7:21:04want to make sure that I have certain
  12018. 7:21:05high priority customers and I have low
  12019. 7:21:07priority customers and I have medium
  12020. 7:21:09priority customers. So, with the help of
  12021. 7:21:10unsupervised learning, I can actually do
  12022. 7:21:13that. I can make a certain classes of
  12023. 7:21:14people on whom I should focus more on as
  12024. 7:21:17compared to the other class. So, this
  12025. 7:21:18was just an example, guys. You can use
  12026. 7:21:20it in various other fields as well. So,
  12027. 7:21:22in marketing, this is how you can use
  12028. 7:21:24unsupervised learning.
  12029. 7:21:25So, this brings us to our next uh type
  12030. 7:21:27of machine learning, which is called a
  12031. 7:21:29reinforcement learning. Now, this is
  12032. 7:21:31reinforcement learning, guys. Now, what
  12033. 7:21:33happens in reinforcement learning, the
  12034. 7:21:34machine learns by interacting with space
  12035. 7:21:36or an environment. So, it learns with
  12036. 7:21:38its experience, with its past
  12037. 7:21:40experience, and also by new choice
  12038. 7:21:42exploration. Now, I'll take the analogy
  12039. 7:21:44of dogs. So, if you have any dog or a
  12040. 7:21:46pet at your home, so if you have trained
  12041. 7:21:48your dog in order to get the newspaper,
  12042. 7:21:50and if it gets it, then you reward it
  12043. 7:21:51with some chocolate or things that the
  12044. 7:21:54dog likes, right? So, the dog will know
  12045. 7:21:56whatever he has done, he's actually
  12046. 7:21:57rewarded for that. So, it'll continue
  12047. 7:21:59doing that. But, apart from that, if he
  12048. 7:22:01does something else, if instead of the
  12049. 7:22:02newspaper, he brings something else. So,
  12050. 7:22:04what do you do? You might even punish
  12051. 7:22:05it. So, because of that, the dog will
  12052. 7:22:07come to know that it has to get a
  12053. 7:22:09newspaper every morning.
  12054. 7:22:10Now, the same example is there in front
  12055. 7:22:12of your screen. So, you have this
  12056. 7:22:14machine. So, it has two choices, either
  12057. 7:22:16to touch the fire or touch the water.
  12058. 7:22:19Now, first what it does, it goes on and
  12059. 7:22:21touch the fire. So, because of that, it
  12060. 7:22:23gets some burning sensation. Now, it has
  12061. 7:22:24only other option, that is to touch the
  12062. 7:22:27water. So, when it touch the water, it
  12063. 7:22:29gets some reward. So, because of that,
  12064. 7:22:31it'll understand that it does not have
  12065. 7:22:32to touch fire ever again.
  12066. 7:22:35Now, there's a diagram that is there in
  12067. 7:22:36front of your screen. So, what happens
  12068. 7:22:38you have an agent, all right? That agent
  12069. 7:22:40performs some action. And on the basis
  12070. 7:22:42of that action, it'll be exposed to some
  12071. 7:22:44sort of an environment. Now, if that
  12072. 7:22:46action is correct, then it'll be
  12073. 7:22:48rewarded with that. But if it is not,
  12074. 7:22:50then it will change its choice, and it
  12075. 7:22:52will again perform some action. So, this
  12076. 7:22:54process will keep on repeating. So, this
  12077. 7:22:56is how reinforcement learning works.
  12078. 7:22:59So, let us move forward and understand
  12079. 7:23:01when we have machine learning, why do we
  12080. 7:23:03need deep learning? That is, we'll look
  12081. 7:23:05at various uh limitations of machine
  12082. 7:23:07learning.
  12083. 7:23:08Now, the first limitation is high
  12084. 7:23:09dimensionality of the data. Now, the
  12085. 7:23:11data that is now generated is huge in
  12086. 7:23:14size. So, we have a very large number of
  12087. 7:23:16inputs and outputs. So, due to that,
  12088. 7:23:18machine learning algorithms fail. So,
  12089. 7:23:20they cannot deal with high
  12090. 7:23:21dimensionality of data, or you can say
  12091. 7:23:22data with large number of inputs and
  12092. 7:23:25outputs.
  12093. 7:23:26Now, there's another problem as well, in
  12094. 7:23:27which it is unable to solve the crucial
  12095. 7:23:29AI problems, which can be natural
  12096. 7:23:31language processing, image recognition,
  12097. 7:23:32and uh things like that.
  12098. 7:23:34Now, one of the biggest challenges with
  12099. 7:23:35machine learning models is feature
  12100. 7:23:37extraction. Now, let me tell you what
  12101. 7:23:39are features. So, in statistics, we
  12102. 7:23:40consider features as variables, but when
  12103. 7:23:42we talk about artificial intelligence,
  12104. 7:23:44these variables are nothing but the
  12105. 7:23:45features.
  12106. 7:23:46Now, what happens because of that? The
  12107. 7:23:48complex problems such as object
  12108. 7:23:50recognition or handwriting recognition
  12109. 7:23:52becomes a huge challenge for machine
  12110. 7:23:53learning algorithms to solve. Now, let
  12111. 7:23:55me give you an example of this uh
  12112. 7:23:56feature extraction. Suppose, if you want
  12113. 7:23:58to predict that whether there'll be a
  12114. 7:24:00match today or not. So, it depends on a
  12115. 7:24:02various features. It depends on the
  12116. 7:24:03whether the weather is sunny, whether it
  12117. 7:24:05is windy, all those things. So, we have
  12118. 7:24:07provided all those features in our data
  12119. 7:24:09set. But, we have forgot one particular
  12120. 7:24:11feature that is humidity. And now, our
  12121. 7:24:13machine learning models are not that
  12122. 7:24:14efficient that they will automatically
  12123. 7:24:16generate that particular feature. So,
  12124. 7:24:18this is one huge problem, or you can say
  12125. 7:24:20limitation, with machine learning. Now,
  12126. 7:24:22obviously, we have limitation, and it
  12127. 7:24:23won't be fair that if I don't give you
  12128. 7:24:25the solution to this particular problem.
  12129. 7:24:27So, we'll move forward and understand
  12130. 7:24:28how deep learning solves these kind of
  12131. 7:24:29problems.
  12132. 7:24:31Now, as you can see that the first line
  12133. 7:24:32on your slide, which says that deep
  12134. 7:24:34learning models are capable to focus on
  12135. 7:24:36the right features by themselves,
  12136. 7:24:37requiring little guidance from the
  12137. 7:24:38programmer. So, with the help of little
  12138. 7:24:40guidance, what these deep learning
  12139. 7:24:42algorithms can do. They can generate the
  12140. 7:24:44features on which the outcome will
  12141. 7:24:46depend on. And at the same time, it also
  12142. 7:24:48solves the dimensionality problem as
  12143. 7:24:50well. If you have very large number of
  12144. 7:24:51inputs and outputs, you can make use of
  12145. 7:24:53a deep learning algorithm. Now, what
  12146. 7:24:55exactly is deep learning? Again, since
  12147. 7:24:57we know that it has been evolved by
  12148. 7:24:59machine learning and machine learning is
  12149. 7:25:01nothing but a subset of artificial
  12150. 7:25:02intelligence. And the idea behind
  12151. 7:25:03artificial intelligence is to imitate
  12152. 7:25:05the human behavior. The same idea is for
  12153. 7:25:07the deep learning as well is to build
  12154. 7:25:09learning algorithms that can mimic
  12155. 7:25:11brain.
  12156. 7:25:12Now, let us move forward and understand
  12157. 7:25:14deep learning what exactly it is.
  12158. 7:25:16Now, the deep learning is implemented
  12159. 7:25:18with the help of neural networks. And
  12160. 7:25:19the idea or the motivation behind neural
  12161. 7:25:21networks are nothing but neurons. What
  12162. 7:25:23are neurons? These are nothing but your
  12163. 7:25:24brain cells. Now, here's a diagram of
  12164. 7:25:26neuron. So, we have dendrites here,
  12165. 7:25:28which are used to provide input to a
  12166. 7:25:30neuron. As you can see, we have multiple
  12167. 7:25:32dendrites here. So, these many inputs
  12168. 7:25:34will be provided to a neuron. Now, this
  12169. 7:25:35is called cell body and inside the cell
  12170. 7:25:37body, we have a nucleus, which performs
  12171. 7:25:39some function. After that, that output
  12172. 7:25:41will travel through axon and it will go
  12173. 7:25:44towards the axon terminals. And then,
  12174. 7:25:46this neuron will fire this output
  12175. 7:25:48towards the next neuron. Now, the
  12176. 7:25:50studies tell us that the next neuron now
  12177. 7:25:52or you can say the two neurons are never
  12178. 7:25:54connected to each other. There's a gap
  12179. 7:25:55between them. So, that is called a
  12180. 7:25:57synapse. So, this is how basically a
  12181. 7:26:00neuron works like. And on the right hand
  12182. 7:26:02side of your slide, you can see an
  12183. 7:26:03artificial neuron. Now, let me explain
  12184. 7:26:05you that. So, over here, similar to
  12185. 7:26:07neurons, we have multiple inputs. Now,
  12186. 7:26:09these inputs will be provided to a
  12187. 7:26:11processing element like a cell body. And
  12188. 7:26:14over here in the processing element,
  12189. 7:26:15what will happen? Summation of your
  12190. 7:26:17inputs and weights. Now, when it moves
  12191. 7:26:20on, then what will happen? This input
  12192. 7:26:22will be multiplied with our weights. So,
  12193. 7:26:24in the beginning, what happens? These
  12194. 7:26:25weights are randomly assigned. So, what
  12195. 7:26:27will happen if I take the example of X1?
  12196. 7:26:29So, X1 multiplied by W1 will go towards
  12197. 7:26:32the processing element. Similarly, X2
  12198. 7:26:34and W2 will go towards the processing
  12199. 7:26:36element. And, similarly, the other
  12200. 7:26:38inputs as well. And, then summation will
  12201. 7:26:40happen which will generate a function of
  12202. 7:26:41S, that is f of S.
  12203. 7:26:43After that comes the concept of
  12204. 7:26:45activation function. Now, what is
  12205. 7:26:47activation function? It is nothing but
  12206. 7:26:49in order to provide a threshold. So, if
  12207. 7:26:51your output is above the threshold, then
  12208. 7:26:52only this neuron will fire, otherwise it
  12209. 7:26:54won't fire. So, you can use a step
  12210. 7:26:56function as an activation function, or
  12211. 7:26:57you can even use a sigmoid function as
  12212. 7:26:59your activation function. So, this is
  12213. 7:27:01how an artificial neuron it looks like.
  12214. 7:27:03So, a network will be multiple neurons
  12215. 7:27:05which are connected to each other will
  12216. 7:27:06form an artificial neural network. And,
  12217. 7:27:08this activation function can be a
  12218. 7:27:10sigmoid function or a step function,
  12219. 7:27:12that totally depends on your
  12220. 7:27:13requirement.
  12221. 7:27:14Now, once it exceeds the threshold, it
  12222. 7:27:16will fire. After that, what will happen?
  12223. 7:27:18It will check the output. Now, if this
  12224. 7:27:20output is not equal to the desired
  12225. 7:27:22output, so these are the actual outputs,
  12226. 7:27:24and we know the real output. So, we'll
  12227. 7:27:26compare both of that, and we'll find the
  12228. 7:27:28difference between the actual output and
  12229. 7:27:30the desired output. On the basis of that
  12230. 7:27:32difference, we are again going to update
  12231. 7:27:34our weights. And, this process will keep
  12232. 7:27:36on repeating until we get the desired
  12233. 7:27:39output as our actual output. Now, this
  12234. 7:27:41process of updating weight is nothing
  12235. 7:27:43but your back propagation method.
  12236. 7:27:45So, this is neural networks in a
  12237. 7:27:47nutshell. So, we'll move forward and
  12238. 7:27:49understand what are deep networks. So,
  12239. 7:27:51basically, deep learning is implemented
  12240. 7:27:53by the help of deep networks, and deep
  12241. 7:27:54networks are nothing but neural networks
  12242. 7:27:57with multiple hidden layers. Now, what
  12243. 7:27:59are hidden layers? Let me explain you
  12244. 7:28:00that. So, you have inputs that comes
  12245. 7:28:03here. So, this will be your input layer.
  12246. 7:28:05After that, some process happens, and
  12247. 7:28:07it'll go to the next node, or you can
  12248. 7:28:09say to the hidden layer nodes. So, this
  12249. 7:28:11is nothing but your hidden layer one.
  12250. 7:28:13So, every node is interconnected if you
  12251. 7:28:16can notice. After that, you have one
  12252. 7:28:18more hidden layer where some function
  12253. 7:28:19will happen. And, as you can see that
  12254. 7:28:22again these nodes are interconnected to
  12255. 7:28:23each other. After this hidden layer two
  12256. 7:28:26comes the output layer, and this output
  12257. 7:28:28layer again we are going to check the
  12258. 7:28:30output whether it is equal to the
  12259. 7:28:31desired output or not. If it is not, we
  12260. 7:28:33are again going to update the weights.
  12261. 7:28:35So, this is how a deep network looks
  12262. 7:28:37like. Now, there can be multiple hidden
  12263. 7:28:39layers. There can be hundreds of hidden
  12264. 7:28:41layers as well. But, when we talk about
  12265. 7:28:43machine learning, that was not the case.
  12266. 7:28:45We were not able to process multiple
  12267. 7:28:47hidden layers when we talk about machine
  12268. 7:28:49learning. So, because of deep learning,
  12269. 7:28:51we have multiple hidden layers at once.
  12270. 7:28:54Now, let us understand this with an
  12271. 7:28:55example. So, we'll take an image which
  12272. 7:28:57has four pixels. So, if you can notice,
  12273. 7:28:59we have four pixels here, among which
  12274. 7:29:01the top two pixels are bright, that is
  12275. 7:29:03they are black in color, whereas bottom
  12276. 7:29:04two pixels are white. Now, what happens?
  12277. 7:29:07We'll divide these pixels and we'll send
  12278. 7:29:08these pixels to each and every node. So,
  12279. 7:29:11for that, we need four nodes. So, this
  12280. 7:29:13particular pixel will go to this node,
  12281. 7:29:14it will go to this node, this pixel will
  12282. 7:29:16go to this node, and finally this pixel
  12283. 7:29:18will go to this particular node that I'm
  12284. 7:29:19highlighting with my cursor. Now, what
  12285. 7:29:21happens? We provide them random weights.
  12286. 7:29:24So, these white lines actually represent
  12287. 7:29:25the positive weights, and these black
  12288. 7:29:27lines represents the negative weights.
  12289. 7:29:29Now, this particular brightness, when we
  12290. 7:29:31display high brightness, we'll consider
  12291. 7:29:32it as negative. Now, what happens? When
  12292. 7:29:35you see the next output or the next
  12293. 7:29:36hidden layer, it'll be provided with the
  12294. 7:29:38input with this particular layer. So,
  12295. 7:29:40this will provide an input with positive
  12296. 7:29:42weight to this particular node, and the
  12297. 7:29:44second input will come from this
  12298. 7:29:45particular node. Since both of them are
  12299. 7:29:47positive, so we'll get this kind of a
  12300. 7:29:49node. Similarly, this node as well. Now,
  12301. 7:29:51when I talk about these two nodes, the
  12302. 7:29:52first node over here, so this is getting
  12303. 7:29:54input from this node as well as from
  12304. 7:29:56this node. Now, over here we have a
  12305. 7:29:58negative weight. So, because of that,
  12306. 7:30:00the value will be negative, and we have
  12307. 7:30:02represented that with black color.
  12308. 7:30:04Similarly, over here as well, we're
  12309. 7:30:06getting one input from here which has a
  12310. 7:30:07negative weight, and the another input
  12311. 7:30:09from here which has again has a negative
  12312. 7:30:10weight. So, accordingly, we get again a
  12313. 7:30:13negative value here. So, these two
  12314. 7:30:14becomes black in color. Now, if you
  12315. 7:30:17notice what'll happen next, we'll
  12316. 7:30:19provide one input here, which will be
  12317. 7:30:21negative and a positive weight, which
  12318. 7:30:23will be again negative, and this will be
  12319. 7:30:25also negative and a positive weight. So,
  12320. 7:30:27that will again come out to be negative.
  12321. 7:30:29So, that is why we have got this kind of
  12322. 7:30:31a structure. If you notice this this is
  12323. 7:30:33nothing but the inverse of this
  12324. 7:30:34particular image. When I talk about this
  12325. 7:30:36node over here, we are getting the
  12326. 7:30:38negative value with a positive weight,
  12327. 7:30:40which is negative, and a negative value
  12328. 7:30:41with a negative weight, which is
  12329. 7:30:42positive. So, we are getting something
  12330. 7:30:44which is positive here.
  12331. 7:30:46Now, obviously, I want this particular
  12332. 7:30:47image to get inverse. I want these black
  12333. 7:30:50strips to come up. So, what I'll do,
  12334. 7:30:52I'll actually calculate the inverse by
  12335. 7:30:54providing a negative weight like this.
  12336. 7:30:55Over here, I've provided a negative
  12337. 7:30:57weight, it'll come up. So, when I
  12338. 7:30:58provide a positive weight, so it'll stay
  12339. 7:31:00wherever it is. After that, it'll
  12340. 7:31:02detect, and the output you can see will
  12341. 7:31:04be a horizontal image, not a solid, not
  12342. 7:31:06a vertical, not a diagonal, but a
  12343. 7:31:08horizontal. And after that, we are going
  12344. 7:31:10to calculate the difference between the
  12345. 7:31:11actual output and the desired output,
  12346. 7:31:13and we are going to update the weights
  12347. 7:31:14accordingly. Now, this is just an
  12348. 7:31:16example, guys. So, guys, this is one
  12349. 7:31:18example of deep learning, where what
  12350. 7:31:19happens, we have images here. We provide
  12351. 7:31:22these raw data to the first layer to the
  12352. 7:31:24input layer.
  12353. 7:31:25Then, what happens, these input layers
  12354. 7:31:27will determine the patterns of local
  12355. 7:31:28contrast, or it'll fixate those patterns
  12356. 7:31:30of local contrast, which means that
  12357. 7:31:32it'll differentiate on the basis of
  12358. 7:31:34colors and luminosity and all those
  12359. 7:31:36things. So, it'll differentiate those
  12360. 7:31:37things. And after that, in the following
  12361. 7:31:39layer, what will happen, it'll determine
  12362. 7:31:41the face features, it'll fixate those
  12363. 7:31:43face features. So, it'll form nose,
  12364. 7:31:45eyes, ears, all those things. Then, what
  12365. 7:31:47will happen, it'll accumulate those
  12366. 7:31:49correct features for the correct face,
  12367. 7:31:51or you can say that and fixate those
  12368. 7:31:52features on the correct face template.
  12369. 7:31:55So, it'll actually determine the faces
  12370. 7:31:56here, as you can see it over here. And
  12371. 7:31:58then, it'll be sent to the output layer.
  12372. 7:32:00Now, basically, you can add more hidden
  12373. 7:32:02layers to solve more complex problem.
  12374. 7:32:04For example, if I want to find out a
  12375. 7:32:06particular kind of face, for example, a
  12376. 7:32:08face which has large eyes, or which has
  12377. 7:32:10light complexion. So, I can do that by
  12378. 7:32:12adding more hidden layers. And I can
  12379. 7:32:14increase the complexity also at the same
  12380. 7:32:16time, if I want to find which image
  12381. 7:32:18contains a dog. So, for for also, I can
  12382. 7:32:20have one more hidden layer. So, as and
  12383. 7:32:22when hidden layer increases, we are able
  12384. 7:32:23to solve more and more complex problem.
  12385. 7:32:25So, this is just a general overview of
  12386. 7:32:27how a deep network looks like. So, we
  12387. 7:32:29have first patterns of local contrast in
  12388. 7:32:31the first layer. Then what happens, we
  12389. 7:32:33fixate these patterns of local contrast
  12390. 7:32:35in order to form the face features such
  12391. 7:32:37as eyes, nose, ears, etc. And then we
  12392. 7:32:39accumulate these features for the
  12393. 7:32:41correct face and then we determine the
  12394. 7:32:43image. So, this is how uh deep learning
  12395. 7:32:46network or you can say deep network
  12396. 7:32:47looks like.
  12397. 7:32:48So, we'll move forward and I'll give you
  12398. 7:32:50some applications of deep learning. So,
  12399. 7:32:52here are few applications of deep
  12400. 7:32:53learning. It can be used in self-driving
  12401. 7:32:55cars. So, you must have heard about
  12402. 7:32:57self-driving cars. So, what happens,
  12403. 7:32:59it'll capture the images around it.
  12404. 7:33:00It'll process that huge amount of data
  12405. 7:33:02and then it'll decide what action should
  12406. 7:33:04it take. Should it take left, right?
  12407. 7:33:05Should it stop? So, accordingly it'll
  12408. 7:33:07decide what action should it take and
  12409. 7:33:09that will reduce the amount of accidents
  12410. 7:33:10that happens every year. Then when we
  12411. 7:33:12talk about voice control assistants, I'm
  12412. 7:33:14pretty sure you must have heard about
  12413. 7:33:15Siri. All the iPhone users know about
  12414. 7:33:17Siri, right? So, you can tell Siri
  12415. 7:33:19whatever you want to do. It'll search it
  12416. 7:33:20for you and display for you.
  12417. 7:33:22Then when we talk about automatic image
  12418. 7:33:24caption generation. So, what happens in
  12419. 7:33:26this, whatever image that you upload,
  12420. 7:33:27the algorithm is in such a way that
  12421. 7:33:29it'll generate the caption accordingly.
  12422. 7:33:31So, for example, if you have say blue
  12423. 7:33:33colored eyes, so it'll display a blue
  12424. 7:33:35colored eye caption uh at the bottom of
  12425. 7:33:37the image.
  12426. 7:33:38Now, when I talk about automatic machine
  12427. 7:33:39translation, so we can convert English
  12428. 7:33:42language into Spanish. Similarly,
  12429. 7:33:44Spanish to French. So, basically
  12430. 7:33:46automatic machine translation, you can
  12431. 7:33:47convert one language to another language
  12432. 7:33:49with the help of deep learning. And
  12433. 7:33:51these are just few examples, guys. There
  12434. 7:33:52are many, many other examples of deep
  12435. 7:33:54learning. It can be used in game
  12436. 7:33:56playing. It can be used in many other
  12437. 7:33:58things. And let me tell you one very
  12438. 7:33:59fascinating thing that I've told you in
  12439. 7:34:00the beginning as well. With the help of
  12440. 7:34:02deep learning, MIT is trying to predict
  12441. 7:34:04future. So, yeah, I know it is growing
  12442. 7:34:06exponentially right now, guys.
  12443. 7:34:14So, this is the problem statement, guys.
  12444. 7:34:15We need to figure out if the bank notes
  12445. 7:34:17are real or fake. And for that, we'll be
  12446. 7:34:19using artificial neural network. And
  12447. 7:34:21obviously, we need some sort of data in
  12448. 7:34:23order to train our network. So, let us
  12449. 7:34:25see how the data set looks like. So,
  12450. 7:34:27over here I've taken a screenshot of the
  12451. 7:34:29data set with few of the rows. In it,
  12452. 7:34:31data were extracted from images that
  12453. 7:34:33were taken from genuine and forged bank
  12454. 7:34:35note-like specimens.
  12455. 7:34:37After that, wavelet transform tools were
  12456. 7:34:39used to extract features from those
  12457. 7:34:41images. And these are few features that
  12458. 7:34:43I'm highlighting with my cursor. And the
  12459. 7:34:45final column or the last column actually
  12460. 7:34:47represents the label.
  12461. 7:34:48So, basically, label tells us to which
  12462. 7:34:50class that pattern represents, whether
  12463. 7:34:52that pattern represents a fake note or
  12464. 7:34:54it represents a real note. Let us
  12465. 7:34:56discuss these features and labels one by
  12466. 7:34:58one.
  12467. 7:34:59So, the first feature or the first
  12468. 7:35:00column is nothing but variance of a
  12469. 7:35:02wavelet transformed image. The second
  12470. 7:35:04column is about skewness. The third is
  12471. 7:35:06kurtosis of wavelet transformed image.
  12472. 7:35:08And finally, fourth one is entropy of
  12473. 7:35:10the image. After that, when I talk about
  12474. 7:35:12label, which is nothing but my last
  12475. 7:35:13column, over here if the value is one,
  12476. 7:35:15that means the pattern represents a real
  12477. 7:35:17note. Whereas, when value is zero, that
  12478. 7:35:19means it represents a fake note. So
  12479. 7:35:21guys, let's move forward and we'll see
  12480. 7:35:22what are the various steps involved in
  12481. 7:35:24order to implement this use case.
  12482. 7:35:26So, over here we'll first begin by
  12483. 7:35:28reading the data set that we have. We'll
  12484. 7:35:29define features and labels.
  12485. 7:35:32After that, we are going to encode the
  12486. 7:35:33dependent variable. And what is a
  12487. 7:35:35dependent variable? It is nothing but
  12488. 7:35:36your label.
  12489. 7:35:37Then, we are going to divide the data
  12490. 7:35:39set into two parts, one for training,
  12491. 7:35:41another for testing.
  12492. 7:35:42After that, we'll use TensorFlow data
  12493. 7:35:44structures for holding features, labels,
  12494. 7:35:46etc. And TensorFlow is nothing but a
  12495. 7:35:48Python library that is used in order to
  12496. 7:35:50implement deep learning models or you
  12497. 7:35:52can say neural networks.
  12498. 7:35:53Then, we'll write the code in order to
  12499. 7:35:55implement the model. And once this is
  12500. 7:35:57done, we will train our model on the
  12501. 7:35:58training data. We'll calculate the
  12502. 7:36:00error. The error is nothing but your
  12503. 7:36:02difference between the model output and
  12504. 7:36:04the actual output.
  12505. 7:36:05And we'll try to reduce this error. And
  12506. 7:36:07once this error becomes minimum, we'll
  12507. 7:36:09make prediction on the test data and
  12508. 7:36:11we'll calculate the final accuracy.
  12509. 7:36:13So guys, let me quickly open my PyCharm
  12510. 7:36:15and I'll show you how the output looks
  12511. 7:36:16like.
  12512. 7:36:18So this is my PyCharm, guys. Over here,
  12513. 7:36:19I've already written the code in order
  12514. 7:36:21to execute the use case. I'll go ahead
  12515. 7:36:23and run this and I'll show you the
  12516. 7:36:24output.
  12517. 7:36:29So over here, as you can see, with every
  12518. 7:36:30iteration, the accuracy is increasing.
  12519. 7:36:33So let me just stop it right here.
  12520. 7:36:35All right. Till now, any questions, any
  12521. 7:36:37doubts with respect to what is our use
  12522. 7:36:38case, what is the data set about? Any
  12523. 7:36:41questions, guys? You can go ahead and
  12524. 7:36:42ask me.
  12525. 7:36:44Okay, there's a question from Arpan.
  12526. 7:36:46He's asking, "Can you explain the code?"
  12527. 7:36:47Definitely, Arpan. I'll be doing that at
  12528. 7:36:49the end of this class when you are done
  12529. 7:36:51with all the fundamentals of neural
  12530. 7:36:52networks. I'll explain you the entire
  12531. 7:36:53code, how I've written that, and how
  12532. 7:36:55I've used TensorFlow in order to
  12533. 7:36:56implement a neural network.
  12534. 7:36:58I hope I you are satisfied with the
  12535. 7:37:00answer. Okay, he's fine with it. Any
  12536. 7:37:02other questions, any other doubts, guys?
  12537. 7:37:03Just go ahead and ask me. Over here, you
  12538. 7:37:05don't need to worry about code right
  12539. 7:37:06now, guys, because I'll explain this
  12540. 7:37:08later in the session. So what I'll do,
  12541. 7:37:10I'll open my slides once more and we'll
  12542. 7:37:12discuss the fundamentals of neural
  12543. 7:37:13networks that are required in order to
  12544. 7:37:15implement this use case.
  12545. 7:37:16So in order to understand why we need
  12546. 7:37:18neural networks, we are going to compare
  12547. 7:37:20the approach before and after neural
  12548. 7:37:21networks. And we'll see what were the
  12549. 7:37:23various problems that were there before
  12550. 7:37:25neural networks. So earlier,
  12551. 7:37:26conventional computers use an
  12552. 7:37:28algorithmic approach. That is, the
  12553. 7:37:30computer follows a set of instructions
  12554. 7:37:33in order to solve a problem. And unless
  12555. 7:37:35the specific steps that the computer
  12556. 7:37:37needs to follow are known, the computer
  12557. 7:37:39cannot solve the problem. So obviously,
  12558. 7:37:42we need a person who actually knows how
  12559. 7:37:44to solve that problem and then he or she
  12560. 7:37:45can provide the instructions to the
  12561. 7:37:47computer as to how to solve that
  12562. 7:37:48particular problem, right? So we first
  12563. 7:37:50should know the answer to that problem,
  12564. 7:37:52or we should know how to overcome that
  12565. 7:37:54challenge or problem which is there in
  12566. 7:37:55front of us. Then only we can provide
  12567. 7:37:57instructions to the computer.
  12568. 7:37:59So this restricts the problem-solving
  12569. 7:38:00capability of conventional computers to
  12570. 7:38:03problems that we already understand and
  12571. 7:38:05know how to solve. But what about those
  12572. 7:38:07problems whose answer we have no clue
  12573. 7:38:09of? So, that's where our traditional
  12574. 7:38:11approach was a failure. So, that's why
  12575. 7:38:14neural networks were introduced. Now,
  12576. 7:38:15let us see what was the scenario after
  12577. 7:38:17neural networks.
  12578. 7:38:18So, neural networks basically process
  12579. 7:38:20information in a similar way the human
  12580. 7:38:22brain does.
  12581. 7:38:23And these networks, they actually learn
  12582. 7:38:25from examples. You cannot program them
  12583. 7:38:27to perform a specific task. They will
  12584. 7:38:29learn from their examples, from their
  12585. 7:38:31experience. So, you don't need to
  12586. 7:38:33provide all the instructions to perform
  12587. 7:38:34a specific task, and your network will
  12588. 7:38:36learn on its own with its own
  12589. 7:38:38experience.
  12590. 7:38:39All right. So, this is what basically
  12591. 7:38:40neural network does.
  12592. 7:38:42So, even if you don't know how to solve
  12593. 7:38:43a problem, you can train your network in
  12594. 7:38:45such a way that with experience, it can
  12595. 7:38:47actually learn how to solve the problem.
  12596. 7:38:50So, that was a major reason why neural
  12597. 7:38:52networks came into existence.
  12598. 7:38:54We'll move forward and we'll understand
  12599. 7:38:56what is the motivation behind neural
  12600. 7:38:58networks.
  12601. 7:38:59So, these neural networks are basically
  12602. 7:39:01inspired by neurons, which are nothing
  12603. 7:39:02but your brain cells.
  12604. 7:39:04And the exact working of the human brain
  12605. 7:39:06is still a mystery, though.
  12606. 7:39:08So, as I've told you earlier as well
  12607. 7:39:09that neural networks work like human
  12608. 7:39:10brain and so the name.
  12609. 7:39:13And similar to a newborn human baby, as
  12610. 7:39:15he or she learns from his or her
  12611. 7:39:17experience, we want our network to do
  12612. 7:39:19that as well. But, we want it to do very
  12613. 7:39:21quickly.
  12614. 7:39:22So, here's a diagram of a neuron.
  12615. 7:39:24Basically, a biological neuron receives
  12616. 7:39:26input from other sources, combines them
  12617. 7:39:29in some way, perform a generally
  12618. 7:39:31non-linear operation on the result, and
  12619. 7:39:33then outputs the final result. So, here
  12620. 7:39:36if you notice these dendrites, these
  12621. 7:39:37dendrites will receive signals from the
  12622. 7:39:39other neurons. Then, what will happen?
  12623. 7:39:41It will transfer it to the cell body.
  12624. 7:39:43The cell body will perform some
  12625. 7:39:44function. It can be summation, it can be
  12626. 7:39:46multiplication. So, after performing
  12627. 7:39:48that summation on the set of inputs, via
  12628. 7:39:50axon it is transferred to the next
  12629. 7:39:52neuron.
  12630. 7:39:53Now, let's understand what exactly are
  12631. 7:39:55artificial neural networks.
  12632. 7:39:58It is basically a computing system that
  12633. 7:40:00is designed to simulate the way the
  12634. 7:40:02human brain analyzes and process the
  12635. 7:40:04information. Artificial neural networks
  12636. 7:40:06has self-learning capabilities that
  12637. 7:40:08enable it to produce better results as
  12638. 7:40:11more data becomes available. So, if you
  12639. 7:40:13train your network on more data, it will
  12640. 7:40:14be more accurate.
  12641. 7:40:16So, these neural networks, they actually
  12642. 7:40:17learn by example.
  12643. 7:40:19And you can configure your neural
  12644. 7:40:20network for specific applications. It
  12645. 7:40:22can be pattern recognition or it can be
  12646. 7:40:24data classification, anything like that,
  12647. 7:40:26all right?
  12648. 7:40:27So, because of neural networks, we see a
  12649. 7:40:29lot of new technology has evolved.
  12650. 7:40:31From translating web pages to other
  12651. 7:40:33languages to having a virtual assistant
  12652. 7:40:35to order groceries online to conversing
  12653. 7:40:37with chatbots. All of these things are
  12654. 7:40:40possible because of neural networks.
  12655. 7:40:43So, in a nutshell, if I need to tell
  12656. 7:40:44you, artificial neural network is
  12657. 7:40:46nothing but a network of various
  12658. 7:40:48artificial neurons.
  12659. 7:40:50All right? So, let me show you the
  12660. 7:40:51importance of neural network with two
  12661. 7:40:53scenarios, before and after neural
  12662. 7:40:55network.
  12663. 7:40:56So, over here we have a machine and we
  12664. 7:40:58have trained this machine on the four
  12665. 7:41:00types of dogs, as you can see where I'm
  12666. 7:41:02highlighting with my cursor.
  12667. 7:41:03And once the training is done, we
  12668. 7:41:05provide a random image to this
  12669. 7:41:06particular machine which has a dog. But
  12670. 7:41:09this dog is not like the other dogs on
  12671. 7:41:11which we have trained our system on.
  12672. 7:41:13So, without neural networks, our machine
  12673. 7:41:15cannot identify that dog in the picture,
  12674. 7:41:17as you can see it over here. Basically,
  12675. 7:41:19our machine will be confused. It cannot
  12676. 7:41:21figure out where the dog is. Now, when I
  12677. 7:41:23talk about neural networks, even if you
  12678. 7:41:25have not trained our machine on this
  12679. 7:41:26specific dog, but still it can identify
  12680. 7:41:29certain features of the dogs that we
  12681. 7:41:31have trained on and it can match those
  12682. 7:41:33features with the dog that is there in
  12683. 7:41:34this particular image and it can
  12684. 7:41:36identify that dog. So, this happens all
  12685. 7:41:39because of neural networks. So, this is
  12686. 7:41:41just an example to show you how
  12687. 7:41:42important are neural networks. Now, I
  12688. 7:41:44know you all must be thinking how neural
  12689. 7:41:47networks work.
  12690. 7:41:48So, for that, we'll move forward and
  12691. 7:41:50understand how it actually works.
  12692. 7:41:52So, over here I'll begin by first
  12693. 7:41:54explaining a single artificial neuron
  12694. 7:41:56that is called as perceptron.
  12695. 7:41:58So, this is an example of a perceptron.
  12696. 7:42:00Over here we have multiple inputs X1,
  12697. 7:42:02X2, {dash} {dash} {dash} till Xn. And we
  12698. 7:42:05have corresponding weights as well. W1
  12699. 7:42:07for X1, W2 for X2, similarly Wn for Xn.
  12700. 7:42:10Then what happens, we calculated the
  12701. 7:42:12weighted sum of these inputs. And after
  12702. 7:42:15doing that, we pass it through an
  12703. 7:42:16activation function. This activation
  12704. 7:42:19function is nothing but it provides a
  12705. 7:42:20threshold value. So, above that value my
  12706. 7:42:23neuron will fire, else it won't fire.
  12707. 7:42:26So, this is basically an artificial
  12708. 7:42:27neuron. So, when I talk about a neural
  12709. 7:42:29network, it involves a lot of these
  12710. 7:42:31artificial neurons with their own
  12711. 7:42:33activation function and their processing
  12712. 7:42:35element.
  12713. 7:42:36Now, we'll move forward and we'll
  12714. 7:42:38actually understand various modes of
  12715. 7:42:40this perceptron or single artificial
  12716. 7:42:42neuron. So, there are two modes in a
  12717. 7:42:44perceptron. One is training, another is
  12718. 7:42:45using mode. In training mode, the neuron
  12719. 7:42:48can be trained to fire for particular
  12720. 7:42:50input patterns, which means that we'll
  12721. 7:42:52actually train our neuron to fire on
  12722. 7:42:54certain set of inputs and to not fire on
  12723. 7:42:56the other set of inputs. That's what
  12724. 7:42:58basically training mode is. When I talk
  12725. 7:43:00about using mode, it means that when a
  12726. 7:43:01taught input pattern is detected at the
  12727. 7:43:03input, its associated output becomes the
  12728. 7:43:05current output, which means that once
  12729. 7:43:07the training is done and we provide an
  12730. 7:43:09input on which the neuron has been
  12731. 7:43:11trained on, so it will detect the input
  12732. 7:43:14and will provide the associated output.
  12733. 7:43:16So, that's what basically using mode is.
  12734. 7:43:18So, first you need to train it, then
  12735. 7:43:19only you can use your perceptron or your
  12736. 7:43:21network.
  12737. 7:43:23So, these were the two modes, guys. And
  12738. 7:43:24next up we'll understand what are the
  12739. 7:43:25various activation functions available.
  12740. 7:43:28So, these are the three activation
  12741. 7:43:29functions, although there are many more,
  12742. 7:43:30but I've listed down three. Step
  12743. 7:43:32function. So, over here the moment your
  12744. 7:43:34input is greater than this particular
  12745. 7:43:35value, your neuron will fire, else it
  12746. 7:43:37won't. Similarly for sigmoid and sign
  12747. 7:43:39function as well. So, these are three
  12748. 7:43:41activation functions. There are many
  12749. 7:43:42more that I've told you earlier as well.
  12750. 7:43:44So, yeah, these are the three majorly
  12751. 7:43:45used activation functions. Next up what
  12752. 7:43:48we are going to do, we are going to
  12753. 7:43:49understand how a neuron learns from its
  12754. 7:43:51experience. So, I'll give you a very
  12755. 7:43:53good analogy in order to understand
  12756. 7:43:55that. And later on when we talk about
  12757. 7:43:57the neural networks or you can say
  12758. 7:43:58multiple neurons in a network, I'll
  12759. 7:44:00explain you the math behind it. I'll
  12760. 7:44:02explain you the math behind learning how
  12761. 7:44:04it actually happens. So, right now I'll
  12762. 7:44:05explain you with an analogy. And guys,
  12763. 7:44:07trust me that analogy is pretty
  12764. 7:44:09interesting.
  12765. 7:44:10So, I know all of you must have guessed
  12766. 7:44:12it. So, these are two beer mugs and all
  12767. 7:44:14of you who love beer can actually relate
  12768. 7:44:15to this analogy a lot.
  12769. 7:44:17And I know most of you actually love
  12770. 7:44:19beer, so that's why I've chosen this
  12771. 7:44:20particular analogy so that all of you
  12772. 7:44:22can relate to it.
  12773. 7:44:24All right, jokes apart. So, fine guys,
  12774. 7:44:26so there's a beer festival happening
  12775. 7:44:27near your house.
  12776. 7:44:29And you want to badly go there. But your
  12777. 7:44:31decision actually depends on three
  12778. 7:44:32factors. First is how is the weather,
  12779. 7:44:34whether it is good or bad. Second is
  12780. 7:44:37your wife or husband is going with you
  12781. 7:44:38or not. And the third one is any public
  12782. 7:44:40transport is available. So, on these
  12783. 7:44:43three factors your decision will depend
  12784. 7:44:44whether you will go or not. So, we'll
  12785. 7:44:46consider these three factors as inputs
  12786. 7:44:49to our perceptron. And we'll consider
  12787. 7:44:51our decision of going or not going to
  12788. 7:44:53the beer festival as our output. So, let
  12789. 7:44:55us move forward with that. So, the first
  12790. 7:44:57input is how is the weather, we'll
  12791. 7:44:58consider it as X1. So, when weather is
  12792. 7:45:00good, it'll be one and when it is bad,
  12793. 7:45:02it'll be zero.
  12794. 7:45:03Similarly, your wife is going with you
  12795. 7:45:05or not, so that'd be your X2. If she is
  12796. 7:45:08going then it's one, if she's not going
  12797. 7:45:10then it's zero. Similarly for public
  12798. 7:45:11transport, if it is available then it is
  12799. 7:45:13one, else it is zero.
  12800. 7:45:14So, these are the three inputs that I'm
  12801. 7:45:15talking about. Let's see the output. So,
  12802. 7:45:18output will be one when you're going to
  12803. 7:45:19the beer festival and output will be
  12804. 7:45:21zero when you want to relax at home. You
  12805. 7:45:23want to have beer at home only, you
  12806. 7:45:25don't want to go outside. So, these are
  12807. 7:45:27the two outputs, whether you're going or
  12808. 7:45:28you're not going.
  12809. 7:45:30Now, what a human brain does. Over here,
  12810. 7:45:32okay, fine. I need to go to the beer
  12811. 7:45:34festival, but there are three things
  12812. 7:45:36that I need to consider. But will I give
  12813. 7:45:38importance to all these factors equally?
  12814. 7:45:41Definitely not. There'll be certain
  12815. 7:45:43factors which will be of high priority
  12816. 7:45:45for me. I'll focus on those factors
  12817. 7:45:47more. Whereas few factors won't affect
  12818. 7:45:50that much to me. All right. So, let's
  12819. 7:45:52prioritize our inputs or factors. So,
  12820. 7:45:54here our most important factor is
  12821. 7:45:56weather. So, if weather is good, I love
  12822. 7:45:58beer so much that I don't care even if
  12823. 7:45:59my wife is going with me or not or if
  12824. 7:46:01there is a public transport available.
  12825. 7:46:03So, I love beer that much that if
  12826. 7:46:05weather is good, then definitely I'm
  12827. 7:46:06going there. That means when X1 is high,
  12828. 7:46:08output will be definitely high.
  12829. 7:46:11So, how we do that? How we actually
  12830. 7:46:12prioritize our factors or how we
  12831. 7:46:14actually give importance more to a
  12832. 7:46:16particular input and less to another
  12833. 7:46:18input in a perceptron or in a neuron?
  12834. 7:46:21So, we do that by using weights. So, we
  12835. 7:46:23assign high weights to the more
  12836. 7:46:24important factors or more important
  12837. 7:46:27inputs and we assign low weights to
  12838. 7:46:28those particular inputs which are not
  12839. 7:46:30that important for us.
  12840. 7:46:31So, let's assign weights, guys. So,
  12841. 7:46:33weight W1 is associated with input X1,
  12842. 7:46:36W2 with X2, and similarly W3 with X3.
  12843. 7:46:39Now, as I've told you earlier as well
  12844. 7:46:41that weather is a very important factor,
  12845. 7:46:42so I'll assign a pretty high weight to
  12846. 7:46:43weather and I'll keep it at six.
  12847. 7:46:45Similarly, W2 and W3 are not that
  12848. 7:46:47important, so I'll keep it as two two.
  12849. 7:46:49After that, I've defined a threshold
  12850. 7:46:51value as five, which means that when the
  12851. 7:46:53weighted sum of my input is greater than
  12852. 7:46:55five, then only my neuron will fire or
  12853. 7:46:58you can say then only I'll be going to
  12854. 7:46:59the beer festival.
  12855. 7:47:00All right. So, I'll use my pen and we'll
  12856. 7:47:03see what happens when weather is good.
  12857. 7:47:06So, when weather is good, our X1 is one.
  12858. 7:47:09Our weight is six, we'll multiply it
  12859. 7:47:10with six.
  12860. 7:47:11Then,
  12861. 7:47:14if my wife decides that she is going to
  12862. 7:47:16stay at home and she will probably be
  12863. 7:47:18busy with cooking and she doesn't want
  12864. 7:47:20to drink beer with me. So, she's not
  12865. 7:47:22coming. So, that input becomes zero.
  12866. 7:47:24Zero into two will actually make no
  12867. 7:47:26difference because it'll be zero.
  12868. 7:47:29Then again, there's no public transport
  12869. 7:47:30available also. Then also this will be
  12870. 7:47:32zero into two.
  12871. 7:47:35So, what output I get here?
  12872. 7:47:37I get here as six.
  12873. 7:47:40I notice the threshold value that is
  12874. 7:47:42five. So, definitely six is greater than
  12875. 7:47:44five.
  12876. 7:47:46That means my output
  12877. 7:47:49will be one or you can say my neuron
  12878. 7:47:51will fire or I'll actually go to the
  12879. 7:47:53beer festival.
  12880. 7:47:54So, even if these two inputs are zero
  12881. 7:47:57for me, that means my wife is not
  12882. 7:47:58willing to go with me and there is no
  12883. 7:48:00public transport available, but weather
  12884. 7:48:02is good, which has very high weight
  12885. 7:48:03value and it actually matters a lot to
  12886. 7:48:05me.
  12887. 7:48:06So, if that is high, it doesn't really
  12888. 7:48:08matter whether the two inputs are high
  12889. 7:48:09or not. I'll go to the beer festival.
  12890. 7:48:11All right? Now, I'll explain you a
  12891. 7:48:13different scenario. So, over here our
  12892. 7:48:15threshold was five, but what if I change
  12893. 7:48:17this threshold to three? So, in that
  12894. 7:48:20scenario, even if my weather is not
  12895. 7:48:22good, uh I'll give it a zero. So, zero
  12896. 7:48:25into six.
  12897. 7:48:26But, my wife and public transport both
  12898. 7:48:30are available.
  12899. 7:48:31All right? So, one into two
  12900. 7:48:34plus
  12901. 7:48:35one into two.
  12902. 7:48:38Which is equal to four.
  12903. 7:48:41And it is definitely greater than three.
  12904. 7:48:45Then also my output will be one. That
  12905. 7:48:48means I will definitely go to the beer
  12906. 7:48:49festival even if weather is bad.
  12907. 7:48:52And my neuron will fire. So, these are
  12908. 7:48:54the two scenarios that I've discussed
  12909. 7:48:56with you. All right? So, there can be
  12910. 7:48:57many other ways in which you can
  12911. 7:48:59actually assign weight to your problem
  12912. 7:49:02or to your learning algorithm.
  12913. 7:49:04So, these are the two ways in which you
  12914. 7:49:05can assign weight and prioritize your
  12915. 7:49:07inputs or factors on which your output
  12916. 7:49:09will depend.
  12917. 7:49:10So, obviously on real life all the
  12918. 7:49:12inputs or all the factors are not as
  12919. 7:49:14important for you. So, you actually
  12920. 7:49:16prioritize them. And how you do that in
  12921. 7:49:17perceptron, you provide high weight to
  12922. 7:49:19it. This is just an analogy so that you
  12923. 7:49:22can relate to a perceptron to a real
  12924. 7:49:24life. We'll actually discuss the math
  12925. 7:49:26behind it later in the session as to how
  12926. 7:49:28a network or a neuron learns. All right?
  12927. 7:49:31So, how the weights are actually updated
  12928. 7:49:33and how the output is changing, that all
  12929. 7:49:36those things we'll be discussing later
  12930. 7:49:37in this session. But my aim is to make
  12931. 7:49:40you understand that you can actually
  12932. 7:49:42relate to a real life problem with that
  12933. 7:49:44of a perceptron. All right? And in real
  12934. 7:49:47life problems are not that easy. They
  12935. 7:49:49are very very complex problems that uh
  12936. 7:49:51we actually face. So in order to solve
  12937. 7:49:53those problems, a single neuron is
  12938. 7:49:55definitely not enough. So we need
  12939. 7:49:57networks of neuron. And that's where
  12940. 7:49:59artificial neural network, or you can
  12941. 7:50:01say multi-layer perceptron, comes into
  12942. 7:50:03the picture. Now let us discuss that.
  12943. 7:50:06Multi-layer perceptron or artificial
  12944. 7:50:07neural network.
  12945. 7:50:09So this is how an artificial neural
  12946. 7:50:10network actually looks like. So over
  12947. 7:50:12here we have multiple neurons in present
  12948. 7:50:14in different layers. The first layer is
  12949. 7:50:16always your input layer. This is where
  12950. 7:50:18you're actually feed in all of your
  12951. 7:50:19inputs. Then we have the first hidden
  12952. 7:50:21layer. Then we have second hidden layer,
  12953. 7:50:24and then we have the output layer.
  12954. 7:50:25Although the number of hidden layers
  12955. 7:50:26depend on your application, on what are
  12956. 7:50:28you working, what is your problem. So
  12957. 7:50:30that actually determines how many hidden
  12958. 7:50:32layers you'll have.
  12959. 7:50:33So let me explain you what is actually
  12960. 7:50:34happening here. So you provide in some
  12961. 7:50:36input to the first layer, which is
  12962. 7:50:38nothing but your input layer. You
  12963. 7:50:39provide inputs to these neurons. All
  12964. 7:50:41right? And after some function, the
  12965. 7:50:43output of these neurons will become the
  12966. 7:50:45input to the next layer, which is
  12967. 7:50:46nothing but your hidden layer one. Then
  12968. 7:50:48these hidden layers also have various
  12969. 7:50:50neurons. These neurons will have
  12970. 7:50:51different activation functions. So
  12971. 7:50:53they'll perform their own function on
  12972. 7:50:54the inputs that it receives from the
  12973. 7:50:56previous layer, and then the output of
  12974. 7:50:58this layer will be the input to the next
  12975. 7:51:00hidden layer, which is hidden layer two.
  12976. 7:51:02Similarly, the output of this hidden
  12977. 7:51:04layer will be the input to the output
  12978. 7:51:06layer. And finally, we get the output.
  12979. 7:51:09So this is how basically an artificial
  12980. 7:51:10neural network looks like. Now let me
  12981. 7:51:12explain you this with an example.
  12982. 7:51:14So over here I'll take an example of
  12983. 7:51:16image recognition using neural networks.
  12984. 7:51:19So over here what happens, we feed in a
  12985. 7:51:21lot of images to our input layer.
  12986. 7:51:23Now this input layer will actually
  12987. 7:51:25detect the patterns of local contrast.
  12988. 7:51:28And then we'll feed that to the next
  12989. 7:51:29layer, which is hidden layer one. So in
  12990. 7:51:31this hidden layer one, the face features
  12991. 7:51:34will be recognized. They'll recognize
  12992. 7:51:36eyes, nose, ears, things like that. And
  12993. 7:51:38then, that will be again fed as input to
  12994. 7:51:41the next hidden layer.
  12995. 7:51:42And in this hidden layer, we'll assemble
  12996. 7:51:44those features and we'll try to make a
  12997. 7:51:45face. And then, we'll get the output
  12998. 7:51:48that is the face will be recognized
  12999. 7:51:50properly. So, if you notice here, with
  13000. 7:51:52every layer, we're trying to get a more
  13001. 7:51:54abstract version or the generalized
  13002. 7:51:55version of the input. So, this is how
  13003. 7:51:58basically an artificial neural network
  13004. 7:52:00work, how it works. All right.
  13005. 7:52:02And there's a lot of training and
  13006. 7:52:03learning which is involved that I'll
  13007. 7:52:04show you now.
  13008. 7:52:06Training a neural network. So, how we
  13009. 7:52:07actually train our neural network? So,
  13010. 7:52:09basically, the most common algorithm for
  13011. 7:52:11training a network is called back
  13012. 7:52:12propagation.
  13013. 7:52:14So, what happens in back propagation?
  13014. 7:52:16After the weighted sum of inputs and
  13015. 7:52:17passing through an activation function
  13016. 7:52:19and getting the output, we compare that
  13017. 7:52:21output to the actual output that we
  13018. 7:52:22already know. We figure out how much is
  13019. 7:52:24the difference. We calculate the error.
  13020. 7:52:27And based on that error, what we do, we
  13021. 7:52:28propagate backwards. And we'll see what
  13022. 7:52:31happens when we change the weight. Will
  13023. 7:52:33the error decrease or will it increase?
  13024. 7:52:35And if it increases, when it increases
  13025. 7:52:37by increasing the value of the variables
  13026. 7:52:39or by decreasing the value of variables.
  13027. 7:52:41So, we kind of calculate all those
  13028. 7:52:43things and we update our variables in
  13029. 7:52:45such a way that our error becomes
  13030. 7:52:47minimum. And it takes a lot of
  13031. 7:52:49iterations. Trust me, guys. It takes a
  13032. 7:52:51lot of iterations. We get output a lot
  13033. 7:52:53of times and then we compare it with the
  13034. 7:52:54model with the actual output. Then
  13035. 7:52:56again, we propagate backwards. We change
  13036. 7:52:58the variables. Then again, we calculate
  13037. 7:52:59the output. We compare it again with the
  13038. 7:53:01desired output or the actual output.
  13039. 7:53:03Then again, we propagate backwards. So,
  13040. 7:53:05this process keeps on repeating until we
  13041. 7:53:06get the minimum value.
  13042. 7:53:08All right. So, there's an example that
  13043. 7:53:10is there in front of your screen. Don't
  13044. 7:53:11be scared of the terms that I used. I'll
  13045. 7:53:13actually explain you with an example.
  13046. 7:53:15So, this is the example over here. We
  13047. 7:53:16have zero, one, and two as inputs. And
  13048. 7:53:18our desired output or the output that we
  13049. 7:53:20already know is zero, one, and four. All
  13050. 7:53:22right. So, over here, we can actually
  13051. 7:53:23figure out that desired output is
  13052. 7:53:25nothing but twice of your input. But I'm
  13053. 7:53:27training a computer to do that, right?
  13054. 7:53:29The computer is not a human.
  13055. 7:53:31So, what happens? I actually initialize
  13056. 7:53:33my weight. I keep the value as three.
  13057. 7:53:35So, the model output will be 3 * 0 is 0,
  13058. 7:53:383 * 1 is 3, 3 * 2 is 6. Now, obviously
  13059. 7:53:42it is not equal to your desired output.
  13060. 7:53:43So, we check the error. Now, the error
  13061. 7:53:46that we have got here is 0, 1, and 2,
  13062. 7:53:48which is nothing but your difference.
  13063. 7:53:49So, 0 - 0 is 0, 3 - 2 is 1, 6 - 4 is 2.
  13064. 7:53:53Now, this is called an absolute error.
  13065. 7:53:55After squaring this error, we get square
  13066. 7:53:57error, which is nothing but 0, 1, and 4.
  13067. 7:54:00All right? So, now what we need to do,
  13068. 7:54:01we need to update the variables. We have
  13069. 7:54:03seen that the output that we got is
  13070. 7:54:05actually different from the desired
  13071. 7:54:07output. So, we need to update the value
  13072. 7:54:08of the weight. So, instead of three, our
  13073. 7:54:11computer makes it as four. After making
  13074. 7:54:13the value as four, we get the model
  13075. 7:54:15output as 0, 4, and 8.
  13076. 7:54:17And then we saw that the error has
  13077. 7:54:19actually increased. Instead of
  13078. 7:54:20decreasing, the error has increased. So,
  13079. 7:54:22after updating the variable, the error
  13080. 7:54:24has increased. So, you can see that
  13081. 7:54:26square error is now 0, 4, and 16, and
  13082. 7:54:28earlier it was 0, 1, and 4. That means
  13083. 7:54:30we cannot increase the weight value
  13084. 7:54:32right now. But if we decrease that, make
  13085. 7:54:34it as two, we get the output, which is
  13086. 7:54:37actually equal to desired output. But is
  13087. 7:54:39it always the case that we need to only
  13088. 7:54:41decrease the weight? Definitely not.
  13089. 7:54:44So, in this particular scenario,
  13090. 7:54:45whenever I'm increasing the weight,
  13091. 7:54:46error is increasing, and when I'm
  13092. 7:54:47decreasing the weight, error is
  13093. 7:54:49decreasing. But as I've told you earlier
  13094. 7:54:51as well, this is not the case every
  13095. 7:54:52time. Sometimes you need to increase the
  13096. 7:54:54weight as well. So, how we determine
  13097. 7:54:56that? All right. Fine, guys. This is how
  13098. 7:54:58basically a computer decide whether it
  13099. 7:54:59has to increase the weight or decrease
  13100. 7:55:01the weight. So, what happens here? This
  13101. 7:55:02is a graph of square error versus
  13102. 7:55:04weight.
  13103. 7:55:05So, over here, what happens? Suppose
  13104. 7:55:07your square error is somewhere here.
  13105. 7:55:09And your computer, it starts increasing
  13106. 7:55:11the weight in order to reduce the square
  13107. 7:55:13error. And it notices that whenever it
  13108. 7:55:14increases the weight, square error is
  13109. 7:55:16actually decreasing.
  13110. 7:55:17So, it'll keep on increasing until the
  13111. 7:55:19square error reaches a minimum value.
  13112. 7:55:22And after that, when it tries to still
  13113. 7:55:24increase the weight, the square error
  13114. 7:55:26will increase. So, at that time, our
  13115. 7:55:28network will recognize that whenever it
  13116. 7:55:30is increasing the weight after this
  13117. 7:55:31point, error is increasing. So,
  13118. 7:55:33therefore, it will stop right there, and
  13119. 7:55:34that will be our weight value.
  13120. 7:55:36Similarly, there can be one more
  13121. 7:55:38scenario. Suppose if we increase the
  13122. 7:55:40weight, but then also the square error
  13123. 7:55:41is increasing. So, at that time, we
  13124. 7:55:44cannot increase the weight. At that
  13125. 7:55:45time, computer will realize, "Okay,
  13126. 7:55:46fine. Whenever I'm increasing the
  13127. 7:55:47weight, the square error is increasing.
  13128. 7:55:49So, it'll go in the opposite direction."
  13129. 7:55:51So, it'll start decreasing the weight,
  13130. 7:55:52and it'll keep on doing that until the
  13131. 7:55:54square error becomes minimum. And the
  13132. 7:55:56moment it decreases more, the square
  13133. 7:55:58error is again increases. So, our
  13134. 7:56:00network will know that
  13135. 7:56:01whenever it decreases the weight value,
  13136. 7:56:04the square error is increasing. So, that
  13137. 7:56:05point will be our final weight value.
  13138. 7:56:08So, guys, this is what basically
  13139. 7:56:09backpropagation in a nutshell is. Fine.
  13140. 7:56:12So, we'll move forward, and now is the
  13141. 7:56:14correct time to understand how to
  13142. 7:56:15implement the use case that I was
  13143. 7:56:17talking about in the beginning. That is,
  13144. 7:56:19how to determine whether a node is fake
  13145. 7:56:20or real. So, for that, I'll open my
  13146. 7:56:22PyCharm.
  13147. 7:56:24This is my PyCharm again, guys. Let me
  13148. 7:56:26just close this. All right.
  13149. 7:56:28So, this is the code that I've written
  13150. 7:56:30in order to implement the use case. So,
  13151. 7:56:32over here, what we do, we import the
  13152. 7:56:33first important libraries which are
  13153. 7:56:35required. Matplotlib is used for
  13154. 7:56:36visualization. TensorFlow, we know, in
  13155. 7:56:38order to implement the neural network.
  13156. 7:56:40NumPy for arrays, Pandas for reading the
  13157. 7:56:42data set. Similarly, scikit-learn for
  13158. 7:56:44label encoding as well as for shopping,
  13159. 7:56:46and also to split the data set into
  13160. 7:56:48training and testing parts.
  13161. 7:56:49All right. Fine, guys. So, we'll begin
  13162. 7:56:51by first reading the data set, as I've
  13163. 7:56:52told you earlier as well when I was
  13164. 7:56:54explaining the steps. So, what I'll do,
  13165. 7:56:56I'll use Pandas in order to read the CSV
  13166. 7:56:58file, which has the data set.
  13167. 7:57:00After that, I'll define features and
  13168. 7:57:02labels. So, X will be my feature, and Y
  13169. 7:57:04will contain my label. So, basically, X
  13170. 7:57:06includes all the columns apart from the
  13171. 7:57:08last column, which is the fifth one. And
  13172. 7:57:10because the indexing starts from zero,
  13173. 7:57:12that's why we have written zero till
  13174. 7:57:14fourth. So, it won't include the fourth
  13175. 7:57:16column. All right? And so, our last
  13176. 7:57:18column will actually be our label.
  13177. 7:57:21Then, what we need to do, we need to
  13178. 7:57:22encode the dependent variable.
  13179. 7:57:24So, dependent variable, as I've told,
  13180. 7:57:26nothing but your label. So, I've
  13181. 7:57:28discussed encoding in TensorFlow
  13182. 7:57:29tutorial, you can go through it, and you
  13183. 7:57:31can actually get to know why and how we
  13184. 7:57:32do that. Then, what we have done, we
  13185. 7:57:35have uh read the data set. Then, what we
  13186. 7:57:37need to do is to split our data set into
  13187. 7:57:38training and testing. And uh these are
  13188. 7:57:41all optional steps. You can print the
  13189. 7:57:42shape of your training and test data. If
  13190. 7:57:44you don't want to do it, it's still
  13191. 7:57:45fine.
  13192. 7:57:46Then, we have defined learning rate. So,
  13193. 7:57:47learning rate is actually the steps in
  13194. 7:57:50which the weights will be updated, all
  13195. 7:57:52right? So, that is what basically
  13196. 7:57:53learning rate is. Then, when we talk
  13197. 7:57:55about epochs means iterations.
  13198. 7:57:58Then, we have defined cost history, that
  13199. 7:57:59will be an empty NumPy array, and its
  13200. 7:58:02shape will be one, and it will include
  13201. 7:58:03the float type object. Then, we have
  13202. 7:58:05defined N dim, which is nothing but your
  13203. 7:58:07X shape of axis one, which means your
  13204. 7:58:09column. Then, we'll print that. After
  13205. 7:58:12that, we have defined the number of
  13206. 7:58:13classes. So, there can be only two
  13207. 7:58:14class, whether the note can be fake or
  13208. 7:58:16it can be real. And this model path I've
  13209. 7:58:19given in order to save my model. So,
  13210. 7:58:21I've just given a path where I need to
  13211. 7:58:23save it. So, I'll just save it here
  13212. 7:58:24only, in the current working directory.
  13213. 7:58:26Now is the time to actually define our
  13214. 7:58:29neural network. So, we'll first make
  13215. 7:58:31sure that we have defined the important
  13216. 7:58:33parameters like hidden layers, number of
  13217. 7:58:35neurons in hidden layers. So, I'll take
  13218. 7:58:3610 neurons in every hidden layer, and
  13219. 7:58:37I'm taking four layers like that. Then,
  13220. 7:58:40X will be my placeholder, and the shape
  13221. 7:58:42of this particular placeholder is none,
  13222. 7:58:43{comma} N {underscore} dim. N
  13223. 7:58:45{underscore} dim value I'll get it from
  13224. 7:58:47here, and none can be at any value. I'll
  13225. 7:58:49define one variable W, and I'll
  13226. 7:58:51initialize it with zeros, and this will
  13227. 7:58:53be the shape of my weight. Similarly,
  13228. 7:58:56for bias as well, this will be the
  13229. 7:58:57particular shape. And there will be one
  13230. 7:58:59more placeholder Y dash, which will
  13231. 7:59:01actually be used in order to provide us
  13232. 7:59:03with the actual output of the model.
  13233. 7:59:05There'll be one model output, and
  13234. 7:59:07there'll be one actual output, which we
  13235. 7:59:08use in order to calculate the
  13236. 7:59:09difference, right? So, we'll feed in the
  13237. 7:59:11actual values of the labels in this
  13238. 7:59:13particular placeholder Y dash.
  13239. 7:59:16And now we'll define the model. So, over
  13240. 7:59:18here we have named the function as
  13241. 7:59:20multilayer perceptron, and in it we'll
  13242. 7:59:22first define the first layer. So, the
  13243. 7:59:24first hidden layer, and we are going to
  13244. 7:59:26name it as layer underscore one, which
  13245. 7:59:28will be nothing but the a matrix
  13246. 7:59:30multiplication of X and weights of H1,
  13247. 7:59:33that is the hidden layer one.
  13248. 7:59:35And that'll be added to your biases B1.
  13249. 7:59:37After that, we'll pass it through a
  13250. 7:59:39sigmoid activation function. Similarly,
  13251. 7:59:40in layer two as well, matrix
  13252. 7:59:42multiplication of layer one and weights
  13253. 7:59:45of H2. So, if you can notice, layer one
  13254. 7:59:47was the network layer just before the
  13255. 7:59:50layer two, right? So, the output of this
  13256. 7:59:52layer one will become input to the layer
  13257. 7:59:53two. And that's why we have written
  13258. 7:59:55layer underscore one. It'll be
  13259. 7:59:56multiplied by weights H2, and then we'll
  13260. 7:59:58add it with the bias.
  13261. 8:00:00Similarly, for this particular hidden
  13262. 8:00:01layer as well, and this particular layer
  13263. 8:00:03as well. But, over here we are going to
  13264. 8:00:05use a ReLU activation function instead
  13265. 8:00:07of sigmoid. Then, we are going to define
  13266. 8:00:09the weights and biases. So, this is how
  13267. 8:00:12we basically define weights. This is how
  13268. 8:00:13we basically define weights. So, weights
  13269. 8:00:15H1 will be a variable which will be a
  13270. 8:00:18truncated normal with the shape of N
  13271. 8:00:20underscore dim and N underscore hidden
  13272. 8:00:22underscore one. So, these are nothing
  13273. 8:00:24but your shapes. All right.
  13274. 8:00:26And after that, what we have done, we
  13275. 8:00:27have defined biases as well. Then, we
  13276. 8:00:29need to initialize all the variables.
  13277. 8:00:31So, all the guys, in brief, let's talk
  13278. 8:00:34about TensorFlow.
  13279. 8:00:36Since in TensorFlow, we need to
  13280. 8:00:37initialize a variable before we use it.
  13281. 8:00:42That's how we
  13282. 8:00:43do it. We first initialize it.
  13283. 8:00:46And then we need to run it. That's when
  13284. 8:00:48your variables will be initialized.
  13285. 8:00:51After that, we are going to
  13286. 8:00:52create a stateful object, and then
  13287. 8:00:55finally, I'm going to call my model.
  13288. 8:00:58And then comes it
  13289. 8:01:00part where the training happens. Cost
  13290. 8:01:02function. Cost function
  13291. 8:01:04is nothing but you can say an error that
  13292. 8:01:07will be calculated between the actual
  13293. 8:01:09output
  13294. 8:01:10and the model output.
  13295. 8:01:13All right, so Y is nothing but our model
  13296. 8:01:15output and
  13297. 8:01:17that is nothing but actual output or the
  13298. 8:01:19output that we already know.
  13299. 8:01:21All right, and then we are going to use
  13300. 8:01:23a gradient
  13301. 8:01:24descent optimizer to reduce the error.
  13302. 8:01:27Then
  13303. 8:01:28we are going to create a session object
  13304. 8:01:30and uh finally we are going to run the
  13305. 8:01:32session.
  13306. 8:01:33So,
  13307. 8:01:34this is how we basically for every
  13308. 8:01:35calculated change as
  13309. 8:01:38well as the accuracy that comes after
  13310. 8:01:40every
  13311. 8:01:41the epoch on the training data.
  13312. 8:01:44After we have calculated the accuracy on
  13313. 8:01:46the training data, we are going to plot
  13314. 8:01:47it for every
  13315. 8:01:49accuracy is.
  13316. 8:01:51And after
  13317. 8:01:52after plotting that we have accuracy
  13318. 8:01:53with our tell using the same prediction
  13319. 8:01:55on the test and after the print the and
  13320. 8:01:57the mean score. So, let's do this, guys.
  13321. 8:02:00All right, so training and what
  13322. 8:02:01See, accuracy epochs see has 99%. So,
  13323. 8:02:05with every epoch it is actually
  13324. 8:02:06increasing apart from a couple of
  13325. 8:02:08instances, it is actually keep on
  13326. 8:02:09increasing. So, the more data you train
  13327. 8:02:11your model on, it will be more accurate.
  13328. 8:02:14Let me just close it. So, now the model
  13329. 8:02:16has also been saved where I wanted it to
  13330. 8:02:18be. This is my final test accuracy and
  13331. 8:02:21this is the mean squared error. All
  13332. 8:02:23right, so these are the files that will
  13333. 8:02:24appear once you save your model.
  13334. 8:02:26These are the four files that I've
  13335. 8:02:27highlighted. Now, what we need to do is
  13336. 8:02:29restore this particular model and I've
  13337. 8:02:32explained this in detail how how to re-
  13338. 8:02:35restore a model that you have already
  13339. 8:02:37saved. So, over here what I'll take I've
  13340. 8:02:39taken before to 768.
  13341. 8:02:41So, all the values in the row of 754 and
  13342. 8:02:45768 will be fed to our model and our
  13343. 8:02:48model will make prediction on that. So,
  13344. 8:02:50let us go ahead and run this.
  13345. 8:02:53So, when I'm restoring my model, it
  13346. 8:02:55seems that my model is 100% I'll use a
  13347. 8:02:57value
  13348. 8:02:58fed in. So, whatever values that I have
  13349. 8:03:00actually given as input to my model, it
  13350. 8:03:02has correctly identified its class,
  13351. 8:03:04whether it's a
  13352. 8:03:06fake note or a real note, because fake
  13353. 8:03:09note and one stands for real note, okay?
  13354. 8:03:11So, original class is nothing but a set,
  13355. 8:03:13so it is zero already. And what
  13356. 8:03:15prediction my model has made is zero,
  13357. 8:03:17that means it is fake. percent.
  13358. 8:03:19Similarly, for other values as well.
  13359. 8:03:23Fine, guys. So, this is how we basically
  13360. 8:03:24implement the use case that we saw in
  13361. 8:03:26the beginning.
  13362. 8:03:27So, in this slide you can notice that
  13363. 8:03:29I've listed down only two applications,
  13364. 8:03:30although there are many more.
  13365. 8:03:32So, neural networks in medicine.
  13366. 8:03:34Artificial neural networks are currently
  13367. 8:03:36a very hot research area in medicine,
  13368. 8:03:38and it is believed that they will
  13369. 8:03:39receive extensive application to
  13370. 8:03:42biomedical systems in the next few
  13371. 8:03:43years. And currently, the research is
  13372. 8:03:46mostly on modeling parts of human body
  13373. 8:03:48and uh recognizing diseases from various
  13374. 8:03:50scans. For example, it can be
  13375. 8:03:51cardiograms, CAT scans, ultrasonic
  13376. 8:03:53scans, etc.
  13377. 8:03:55And uh currently, the research is going
  13378. 8:03:57uh mostly on uh two major areas. First
  13379. 8:03:59is modeling and diagnosing the
  13380. 8:04:00cardiovascular system.
  13381. 8:04:02So, neural networks are used
  13382. 8:04:03experimentally to model the human
  13383. 8:04:05cardiovascular system.
  13384. 8:04:07Diagnosis can be achieved by building a
  13385. 8:04:08model of the cardiovascular system of an
  13386. 8:04:10individual and comparing it with the
  13387. 8:04:12real-time physiological measurements
  13388. 8:04:14taken from the patient. And trust me,
  13389. 8:04:16guys, if this routine is carried out
  13390. 8:04:18regularly, potential harmful medical
  13391. 8:04:21conditions can be detected at an early
  13392. 8:04:23stage and thus, make the process of
  13393. 8:04:25combating disease much easier.
  13394. 8:04:27Apart from that, it is currently being
  13395. 8:04:29used in electronic noses as well.
  13396. 8:04:31Electronic noses have several potential
  13397. 8:04:33applications in telemedicine. Now, let
  13398. 8:04:36me just give you an introduction to
  13399. 8:04:37telemedicine. Telemedicine is a practice
  13400. 8:04:39of medicine over long distance via a
  13401. 8:04:41communication link. So, what the
  13402. 8:04:43electronic noses will do, they would
  13403. 8:04:45identify odors in the remote surgical
  13404. 8:04:47environment. These identified odors
  13405. 8:04:50would then be electronically transmitted
  13406. 8:04:52to another site, wherein odor generation
  13407. 8:04:54system would recreate them.
  13408. 8:04:57Because the sense of the smell can be an
  13409. 8:04:58important sense to the surgeon,
  13410. 8:05:00tele-smell would enhance tele-present
  13411. 8:05:02surgery.
  13412. 8:05:04So, these are the two ways in which you
  13413. 8:05:05can use it in medicine. You can use it
  13414. 8:05:08in business as well, guys. So, business
  13415. 8:05:10is basically a diverted field with
  13416. 8:05:12several general areas of specialization
  13417. 8:05:14such as accounting or financial
  13418. 8:05:16analysis. Almost any neural network
  13419. 8:05:18application would fit into one business
  13420. 8:05:20area or financial analysis.
  13421. 8:05:22Now, there is some potential for using
  13422. 8:05:24neural networks for business purposes
  13423. 8:05:25including resource allocation and
  13424. 8:05:27scheduling. I've listed down two major
  13425. 8:05:29areas where it can be used. One is
  13426. 8:05:31marketing.
  13427. 8:05:32So, there is a marketing application
  13428. 8:05:33which has been integrated with a neural
  13429. 8:05:35network system.
  13430. 8:05:37The airline marketing tactician is a
  13431. 8:05:39computer system made of various
  13432. 8:05:41intelligent technologies including
  13433. 8:05:43expert systems. A feedforward neural
  13434. 8:05:45network is integrated with the AMT,
  13435. 8:05:47which is nothing but airline marketing
  13436. 8:05:49tactician, and was trained using
  13437. 8:05:51backpropagation to assist the marketing
  13438. 8:05:53control of airline seat allocation.
  13439. 8:05:56So, it has wide applications in
  13440. 8:05:58marketing as well.
  13441. 8:06:00Now, the second area is credit
  13442. 8:06:01evaluation. Now, I'll give you an
  13443. 8:06:02example here. The HNC company has
  13444. 8:06:05developed several neural network
  13445. 8:06:06applications, and one of them is a
  13446. 8:06:08credit scoring system which increases
  13447. 8:06:10the profitability of existing model up
  13448. 8:06:12to
  13449. 8:06:13So, these are few applications that I'm
  13450. 8:06:15telling you guys. Neural network is
  13451. 8:06:17actually the future.
  13452. 8:06:19People are talking about neural networks
  13453. 8:06:21everywhere, and especially after the
  13454. 8:06:23intro-
  13455. 8:06:24duction of GPUs and the amount of data
  13456. 8:06:26that we have now, neural network is
  13457. 8:06:28actually spreading like plague right
  13458. 8:06:29now.
  13459. 8:06:31>> [music]
  13460. 8:06:37[music]
  13461. 8:06:40>> So, this is an
  13462. 8:06:41image of New York's this picture. So,
  13463. 8:06:43when a human will see this image, he'll
  13464. 8:06:45a lot of buildings in different colors
  13465. 8:06:47and stuff like that. But, how are this
  13466. 8:06:49image? So, So there'll be three
  13467. 8:06:51channels. Red,
  13468. 8:06:52another will be green and finally we
  13469. 8:06:53have blue channel which is popularly
  13470. 8:06:55known as RGB. So all each of these
  13471. 8:06:58channels will they have their own
  13472. 8:06:59respective pixel values as you can see
  13473. 8:07:01it over here. So when I say size is B
  13474. 8:07:04cross A cross 3, it means that there are
  13475. 8:07:07B
  13476. 8:07:08rows, A columns and three channels. All
  13477. 8:07:10right? So So if somebody tells you that
  13478. 8:07:13the size of an image is 28 cross 28
  13479. 8:07:15cross three pixels, it means that it has
  13480. 8:07:1728 rows, 28 columns and three channels.
  13481. 8:07:20So this is how
  13482. 8:07:21this is for colored images for we have
  13483. 8:07:23only two channels. So let's move forward
  13484. 8:07:25and we'll see why can't we use for image
  13485. 8:07:27classification.
  13486. 8:07:28So consider an image which has 28 three
  13487. 8:07:30pixels.
  13488. 8:07:32So when I feed in this image to a fully
  13489. 8:07:33con-
  13490. 8:07:34nected network like this, then the total
  13491. 8:07:36number of weights required in the fully
  13492. 8:07:38connected 2,352.
  13493. 8:07:40You can just go ahead and multiply it
  13494. 8:07:41your-
  13495. 8:07:42self. All right?
  13496. 8:07:44But in real life the images are not that
  13497. 8:07:46small. All right? So whatever images
  13498. 8:07:47that we have, they are definitely above
  13499. 8:07:49200 cross 200 cross three pixels.
  13500. 8:07:51So if I take an image which has 200
  13501. 8:07:53cross 200 cross three pixels and I feed
  13502. 8:07:55it to a fully connected network, that
  13503. 8:07:57time the number of weights required
  13504. 8:07:58itself will be 120,000 guys. So we need
  13505. 8:08:00to deal with such huge amount of
  13506. 8:08:02parameters and obviously we require more
  13507. 8:08:04number of neurons. So that can
  13508. 8:08:05eventually lead to overfitting. So
  13509. 8:08:07that's why we can't use network for
  13510. 8:08:09image classification. Let's see why we
  13511. 8:08:10need convolutional neural networks.
  13512. 8:08:12Basically in convolutional neural
  13513. 8:08:14network in the layer will only be
  13514. 8:08:15connected to a small region of the layer
  13515. 8:08:17before it. So if you consider this
  13516. 8:08:18particular neuron which I'm highlighting
  13517. 8:08:20right now is only connected to three
  13518. 8:08:21other neurons. Unlike the fully
  13519. 8:08:23connected network where this particular
  13520. 8:08:24neuron will be connected to all these
  13521. 8:08:25five neurons. Because of this we need to
  13522. 8:08:27handle less amount of weights and in
  13523. 8:08:29turn we need less number of neurons as
  13524. 8:08:31well. So let us understand what exactly
  13525. 8:08:33is convolutional neural network. So
  13526. 8:08:35convolutional neural networks are
  13527. 8:08:36special type
  13528. 8:08:38of feedforward artificial neural
  13529. 8:08:40networks which is inspired from visual
  13530. 8:08:43cortex. So visual cortex
  13531. 8:08:45This but a small region in our brain
  13532. 8:08:47brain which is present somewhere here
  13533. 8:08:49where you can see the bulb and basically
  13534. 8:08:52what happened
  13535. 8:08:53was an experiment conducted and people
  13536. 8:08:55got to know that visual cortex is small
  13537. 8:08:57regions of cells that are sensitive to
  13538. 8:08:58specific regions of visual field.
  13539. 8:09:01So what I'm
  13540. 8:09:02example some neurons in the visual
  13541. 8:09:03cortex exposed to vertical edges. Some
  13542. 8:09:05will fire when exposed to horizontal
  13543. 8:09:07edges. Some will fire when exposed to
  13544. 8:09:08diagonal edges and that is nothing but
  13545. 8:09:10the motivation behind convolutional
  13546. 8:09:12neural network. So now let us
  13547. 8:09:14convolutional neural network work force.
  13548. 8:09:15So generally a collect work has three
  13549. 8:09:17layers convolution labeling layer and
  13550. 8:09:18fully connected layer. We'll understand
  13551. 8:09:20each of these layers one by one. We'll
  13552. 8:09:22take an example of a classifier that can
  13553. 8:09:24classify an image of an X as well as an
  13554. 8:09:26O. So with this example we'll be
  13555. 8:09:28understanding all these four layers. So
  13556. 8:09:30let's begin guys. Now there are certain
  13557. 8:09:32trickier cases. So what I mean by that
  13558. 8:09:33is X can be represented in these four
  13559. 8:09:36forms as well, right? So these are
  13560. 8:09:38nothing but the deformed images of X.
  13561. 8:09:39Similarly for O as well. So these are
  13562. 8:09:42deformed images. So even I want to
  13563. 8:09:43classify these images either X or O. All
  13564. 8:09:46right, because even this is X, this is
  13565. 8:09:47X, this is X, this is X. But all these
  13566. 8:09:49are deformed images. But they are in
  13567. 8:09:52turn X, right? So I want my classifier
  13568. 8:09:53to classify them as X. So basically
  13569. 8:09:56that's what I want. So if you can notice
  13570. 8:09:58here this is a proper image of an X and
  13571. 8:10:00which is actually equal to this
  13572. 8:10:02particular X which is a deformed image.
  13573. 8:10:03Same goes for this O as well. So now
  13574. 8:10:05what we are going to do is we know that
  13575. 8:10:07a computer understands an image using
  13576. 8:10:08numbers at each pixels. So what we'll
  13577. 8:10:10do, whatever the white pixels that we
  13578. 8:10:12have we are going to assign a value
  13579. 8:10:13minus one to it and whatever the black
  13580. 8:10:15pixels we have we are going to assign a
  13581. 8:10:16value one to it. When we use normal
  13582. 8:10:18techniques to compare these two images,
  13583. 8:10:20one is a proper image of X and another
  13584. 8:10:21is a deformed image of X, we got to know
  13585. 8:10:23that a computer is not able to classify
  13586. 8:10:25the deformed image of X correctly. Why?
  13587. 8:10:27Because it is comparing it with the
  13588. 8:10:29proper image of X, right? So when you go
  13589. 8:10:31ahead and add the pixel values of both
  13590. 8:10:33of these images you get something like
  13591. 8:10:35this. So basically our computer is not
  13592. 8:10:37able to recognize whether it is an X or
  13593. 8:10:39not. Now what we do with the help of CNN
  13594. 8:10:41we take small patches of our image. So,
  13595. 8:10:43these patches or these pieces are known
  13596. 8:10:46as nothing but features or filters. So,
  13597. 8:10:48what we do, by finding rough feature
  13598. 8:10:50matches in roughly the same positions in
  13599. 8:10:52two images, CNN gets a lot better at
  13600. 8:10:54seeing the similarity between the whole
  13601. 8:10:56image matching schemes. What I mean by
  13602. 8:10:57that is, we have these filters, right?
  13603. 8:10:59We have these filters that you can see.
  13604. 8:11:01So, consider this first filter. This is
  13605. 8:11:03exactly equal to the feature or the part
  13606. 8:11:05of the image in the deformed image as
  13607. 8:11:07well. So, this is our proper image and
  13608. 8:11:08this is our deformed image, all right?
  13609. 8:11:10Right? So, this particular feature or
  13610. 8:11:11this particular part of the image is
  13611. 8:11:13actually equal to this particular part
  13612. 8:11:14of the image. Same goes for this
  13613. 8:11:16particular feature or filter as well.
  13614. 8:11:18And similarly, we have this filter as
  13615. 8:11:19well, which is actually equal to this
  13616. 8:11:21particular part of the deformed image,
  13617. 8:11:24all right? So, let's move forward and
  13618. 8:11:25we'll see we're taking in our example.
  13619. 8:11:27So, we'll be considering these three
  13620. 8:11:29features or filters. This is a diagonal
  13621. 8:11:31filter, this is again a diagonal filter
  13622. 8:11:32and this is nothing but a small x. So,
  13623. 8:11:34we'll take these three filters and we'll
  13624. 8:11:36move forward. So, what we are going to
  13625. 8:11:37do is we are going to compare these
  13626. 8:11:39features, the small pieces of the bigger
  13627. 8:11:41image, we are going to put it on the
  13628. 8:11:43input image and if it matches, then the
  13629. 8:11:45image will be classified correctly. Now,
  13630. 8:11:47we'll begin, guys. The first layer is
  13631. 8:11:48convolution layer. So, these are the
  13632. 8:11:50beginning two steps of this particular
  13633. 8:11:51layer. First, we need to line up the
  13634. 8:11:53feature in the image and then multiply
  13635. 8:11:54image by the corresponding feature
  13636. 8:11:56pixel. Now, let me explain you with an
  13637. 8:11:57example. So, this is our first diagonal
  13638. 8:11:59feature that we'll take. We are going to
  13639. 8:12:01put this particular feature on our image
  13640. 8:12:04of x, all right? And we're going to
  13641. 8:12:05multiply the corresponding pixel value.
  13642. 8:12:07So, one will be multiplied with one,
  13643. 8:12:09we'll get one and we'll put it in
  13644. 8:12:10another matrix. Similarly, we are going
  13645. 8:12:13to move forward and we're going to
  13646. 8:12:14multiply minus one with minus one. We're
  13647. 8:12:16going to multiply minus one with minus
  13648. 8:12:18one, as you can see. Similarly, we
  13649. 8:12:19multiply this result, minus one into
  13650. 8:12:21minus one, then again minus one into
  13651. 8:12:22minus one. So, we are going to complete
  13652. 8:12:24this whole process and we're going to
  13653. 8:12:25finish up this matrix, all right? And
  13654. 8:12:27once we are done finishing up the
  13655. 8:12:28multiplication of all the corresponding
  13656. 8:12:30pixels in the feature as well as in the
  13657. 8:12:32image, we need to follow two more steps.
  13658. 8:12:34We need to add them up and divide by the
  13659. 8:12:36total number of the pixels in the
  13660. 8:12:37feature. So, what I mean by that is
  13661. 8:12:39after the multiplication of the
  13662. 8:12:41corresponding pixel values, what we do,
  13663. 8:12:43we add all these values, we divide by
  13664. 8:12:45the total number of pixels, and we get
  13665. 8:12:47some value, right? And then now our next
  13666. 8:12:49step is to create a map and put the
  13667. 8:12:51value of the filter at that particular
  13668. 8:12:53place. We saw that after multiplying the
  13669. 8:12:54pixel value of a feature with the
  13670. 8:12:56corresponding pixel value of with that
  13671. 8:12:58of our image, we get the output which is
  13672. 8:13:00one. So, we place one here. Similarly,
  13673. 8:13:03we are going to move this filter
  13674. 8:13:05throughout the image. Next up, we are
  13675. 8:13:06going to move this filter here. After
  13676. 8:13:08that, we're going to move it here, here,
  13677. 8:13:09here, everywhere on the image we are
  13678. 8:13:11going to move it and we're going to
  13679. 8:13:12follow the same process. All right, so
  13680. 8:13:14yeah, this is one more example where
  13681. 8:13:15I've moved my filter in between and
  13682. 8:13:17after doing that, I've got the output
  13683. 8:13:19something like this, 1 1 -1 and all. So,
  13684. 8:13:22over here if you notice, I've got couple
  13685. 8:13:24of times -1 as well, due to which my
  13686. 8:13:26output that comes is 0.55, right? So,
  13687. 8:13:29I'm going to place 0.55 here. Similarly,
  13688. 8:13:31after moving the pixel after moving the
  13689. 8:13:33filter throughout the image, I got this
  13690. 8:13:35particular matrix. All right? And this
  13691. 8:13:37is for one particular feature. After
  13692. 8:13:39performing the same process for the
  13693. 8:13:41other two filters as well, I've got
  13694. 8:13:43these two values. So, we have these
  13695. 8:13:45three values after passing through the
  13696. 8:13:46convolution layer. Let me give you a
  13697. 8:13:48quick recap of what happens in
  13698. 8:13:49convolution layer. So, basically we have
  13699. 8:13:51taken three features, all right? And one
  13700. 8:13:53by one we'll take one feature, move it
  13701. 8:13:55through the entire image, and when we
  13702. 8:13:56are moving it, at that time we are
  13703. 8:13:57multiplying the pixel value of the image
  13704. 8:13:59with that of the corresponding pixel
  13705. 8:14:00value of the filter, adding them up,
  13706. 8:14:02dividing by the total number of pixels
  13707. 8:14:04to get the output. So, when we do that
  13708. 8:14:07for all the filters, we get we got these
  13709. 8:14:09three outputs, all right? So, let's move
  13710. 8:14:10forward and we'll see what happens in
  13711. 8:14:12ReLU layer. So, this is ReLU layer,
  13712. 8:14:14guys, and people who have gone through
  13713. 8:14:15the previous tutorial actually know what
  13714. 8:14:17it is. So, let me just give you a quick
  13715. 8:14:18introduction of ReLU layer. So, ReLU is
  13716. 8:14:20nothing but a activation function. All
  13717. 8:14:22right? So, what I mean by that is it
  13718. 8:14:23will only activate a node if the input
  13719. 8:14:26is above a certain quantity. While the
  13720. 8:14:28input is below zero, the output is also
  13721. 8:14:30zero, all right? And when the input
  13722. 8:14:32rises above the certain threshold, it
  13723. 8:14:34has a linear relationship with the
  13724. 8:14:36dependent variable. Now, I'll explain
  13725. 8:14:38you with an example. We have a graph of
  13726. 8:14:40ReLU function here. So, my function says
  13727. 8:14:42that when f of x is equal to zero if x
  13728. 8:14:45is less than zero, and it is equal to x
  13729. 8:14:47when x is greater than zero. All right?
  13730. 8:14:49So, whatever values that I have which
  13731. 8:14:51are below zero will actually in turn
  13732. 8:14:53become zero, and whatever values that
  13733. 8:14:54are above zero, our function value will
  13734. 8:14:57also be equal to that particular value.
  13735. 8:14:59So, f of x will be equal to x if it is
  13736. 8:15:01greater than or equal to zero, and it
  13737. 8:15:03will be zero if it is less than zero.
  13738. 8:15:05So, if I have x value as minus three, so
  13739. 8:15:07definitely it is less than zero, so f of
  13740. 8:15:09x becomes zero. Similarly, if I have
  13741. 8:15:11minus five x value, then that again it
  13742. 8:15:12is less than zero, so my f of x value
  13743. 8:15:14becomes zero. But, when I consider three
  13744. 8:15:16as my x value, then my f of x becomes
  13745. 8:15:19equal to x, which is nothing but three.
  13746. 8:15:20So, over here I'll have three. Again, if
  13747. 8:15:22I take my x value as five, then
  13748. 8:15:24obviously it is greater than or equal to
  13749. 8:15:26zero, then my f of x becomes equal to x,
  13750. 8:15:30so my f of x value becomes five. So,
  13751. 8:15:32this is how our ReLU function works. So,
  13752. 8:15:34why are we using ReLU function here is
  13753. 8:15:35we want to remove all the negative
  13754. 8:15:37values from our output that we got
  13755. 8:15:39through the convolution layer. So, we'll
  13756. 8:15:41only take the first output that we got
  13757. 8:15:43by moving one feature throughout the
  13758. 8:15:45image. So, this is the output that we
  13759. 8:15:47have got for only one filter. All right?
  13760. 8:15:49So, over here I'm going to remove all
  13761. 8:15:50negative values. So, over here you can
  13762. 8:15:52see that it it was minus point one one
  13763. 8:15:54before, and I've converted that to zero.
  13764. 8:15:56Similarly, I'm going to repeat the whole
  13765. 8:15:57process for the entire matrix. And once
  13766. 8:16:00I'm done with that, I get this
  13767. 8:16:01particular value. Now, remember this is
  13768. 8:16:03only for the output that we got through
  13769. 8:16:05one feature. All right? So, when we were
  13770. 8:16:07doing convolution at that time, we were
  13771. 8:16:08you
  13772. 8:16:09So, this is the output only for one
  13773. 8:16:10filter. After doing it for the output of
  13774. 8:16:12the other two filters as well, we have
  13775. 8:16:14got these two values more. So, totally
  13776. 8:16:16we have these three values after passing
  13777. 8:16:18through ReLU activation function. Next
  13778. 8:16:20up, we'll see what exactly is pooling
  13779. 8:16:21layer. So, in pooling layer what we do,
  13780. 8:16:23we take a window size of two, and we
  13781. 8:16:25move it across the entire matrix that we
  13782. 8:16:27have got after passing through ReLU
  13783. 8:16:28layer. And we take only the maximum
  13784. 8:16:31value from there so that we can shrink
  13785. 8:16:33the image. So what we are actually doing
  13786. 8:16:35is we are reducing the size of our
  13787. 8:16:37image. So let me explain you with an
  13788. 8:16:38example. So this is basically one output
  13789. 8:16:41that we have got after passing through
  13790. 8:16:42ReLU layer. And over here we have taken
  13791. 8:16:44a window size of two cross two. So when
  13792. 8:16:46we keep this window at this particular
  13793. 8:16:48position, we see that one is the highest
  13794. 8:16:50value. So we're going to keep one here.
  13795. 8:16:52And we are going to repeat the same
  13796. 8:16:54process for this particular window as
  13797. 8:16:56well. So over here the maximum value is
  13798. 8:16:570.33 so 0.33 will come. So if you notice
  13799. 8:17:00here, earlier we had
  13800. 8:17:02seven cross seven matrix and now we have
  13801. 8:17:04reduced that to four cross four matrix.
  13802. 8:17:06So after doing that for the entire
  13803. 8:17:08image, we have got this as our output.
  13804. 8:17:11This output we have got after moving our
  13805. 8:17:13window throughout the image that we have
  13806. 8:17:16got after passing through ReLU layer,
  13807. 8:17:17right? And when we repeat this process
  13808. 8:17:19for all the three outputs that we have
  13809. 8:17:21got after the ReLU layer, then we get
  13810. 8:17:23this particular output after pooling
  13811. 8:17:25layer. Right? So basically we have
  13812. 8:17:27shrinked our image to a four cross four
  13813. 8:17:29matrix. Now comes the tricky part. So
  13814. 8:17:31what we are going to do now is stack up
  13815. 8:17:33all these layers. So we have discussed
  13816. 8:17:34convolution layer, ReLU layer and
  13817. 8:17:36pooling layer. So I'll just give you a
  13818. 8:17:38brief recap of what all things we have
  13819. 8:17:39discussed. In convolution layer what we
  13820. 8:17:41did, we took three features and then
  13821. 8:17:43after that one by one we moved each
  13822. 8:17:45filter throughout the image. And when we
  13823. 8:17:47were moving it, we were continuously
  13824. 8:17:49multiplying the image pixel value with
  13825. 8:17:51that of the corresponding filter pixel
  13826. 8:17:53value and then we were dividing it by
  13827. 8:17:55the total number of pixels. All right?
  13828. 8:17:57With that we got three output after
  13829. 8:17:58passing through the convolution layer.
  13830. 8:18:00Then those three output we passed
  13831. 8:18:02through a ReLU layer where we have
  13832. 8:18:03removed the negative value. All right?
  13833. 8:18:05And after removing negative value again
  13834. 8:18:07we have got the three outputs.
  13835. 8:18:09Then those three outputs we passed
  13836. 8:18:10through pooling layer. So basically
  13837. 8:18:11we're trying to shrink our image. And
  13838. 8:18:13what we did, we took a window size of
  13839. 8:18:15two cross two, moved it through all the
  13840. 8:18:17three outputs that we have got through
  13841. 8:18:19ReLU layer. And after doing that, we
  13842. 8:18:21were only taking the maximum value pixel
  13843. 8:18:23value in that particular window and then
  13844. 8:18:26we were putting it in a different matrix
  13845. 8:18:27so that we get a shrinked image. And
  13846. 8:18:29after passing it through pooling layer,
  13847. 8:18:30we have got a four cross four matrix.
  13848. 8:18:32Since we took three features in the
  13849. 8:18:34beginning, so therefore we have got the
  13850. 8:18:36three outputs after passing through
  13851. 8:18:37pooling layer. All right. Next up, we
  13852. 8:18:39are going to stack up all the layers,
  13853. 8:18:41all right? So, let's do that. So, after
  13854. 8:18:43passing through convolution, relu and
  13855. 8:18:44pooling, we have got this four cross
  13856. 8:18:46four matrix. This was our input image.
  13857. 8:18:48Now, when we add one more layer of
  13858. 8:18:50convolution, relu and pooling, we have
  13859. 8:18:52shrinked our image from four cross four
  13860. 8:18:53to two cross two as you can notice here.
  13861. 8:18:55Now, we are going to use fully connected
  13862. 8:18:57layer. Now, what happens in fully
  13863. 8:18:58connected layer? The actual
  13864. 8:18:59classification happens here, guys, okay?
  13865. 8:19:01So, what we are doing here is we are
  13866. 8:19:03going to take the shrinked images and
  13867. 8:19:05put it into a single list. So, basically
  13868. 8:19:08this is what we have got after passing
  13869. 8:19:10through two layers of convolution, relu
  13870. 8:19:12and pooling and this is what we have
  13871. 8:19:13got. So, basically we're converting into
  13872. 8:19:15a single list or a vector. How we do
  13873. 8:19:17that? We take the first value one, then
  13874. 8:19:18we take 0.55, then we take 0.55, then we
  13875. 8:19:21take one again. Then we take one, then
  13876. 8:19:22we take 0.55, 0.55, 0.55. Then we again
  13877. 8:19:25take 0.55, one, one and 0.55. So, this
  13878. 8:19:29is nothing but a vector or you can say a
  13879. 8:19:31list. If you notice here that there are
  13880. 8:19:33certain values in my list which are high
  13881. 8:19:35for X and similarly if I repeat the
  13882. 8:19:37entire process that we have discussed
  13883. 8:19:39for O, there'll be certain will be high.
  13884. 8:19:41So, for X we have first, fourth, fifth,
  13885. 8:19:4410th and 11th element vector values are
  13886. 8:19:47high. For O we have second, third, ninth
  13887. 8:19:51and 12th element vector which are high.
  13888. 8:19:53So, basically we know now if if we have
  13889. 8:19:55an input image
  13890. 8:19:57which has a first, fourth, 10th and 11th
  13891. 8:20:00element vector values high, we know that
  13892. 8:20:03we can classify it as X. Similarly, if
  13893. 8:20:05our input image has a list which has the
  13894. 8:20:07second, third,
  13895. 8:20:09ninth and 12th element vector values
  13896. 8:20:12high, then we can classify it as zero.
  13897. 8:20:13Now, let me explain you with an example.
  13898. 8:20:15So, after the training is done, after
  13899. 8:20:17the after doing the entire process for
  13900. 8:20:19both X and O, you know that our model is
  13901. 8:20:21trained now, okay? So, we have given one
  13902. 8:20:23a new input image and that input image
  13903. 8:20:25passes through all the layers. And once
  13904. 8:20:26it has passed through all the layers, we
  13905. 8:20:28have got this 12-element vector. Now, it
  13906. 8:20:30has 0.9, 0.65, all these values, right?
  13907. 8:20:33Now, how do we classify it whether it is
  13908. 8:20:35an X or O? So, what we do, we'll compare
  13909. 8:20:37this with a list of X and O, right? So,
  13910. 8:20:39we have got the list in the previous uh
  13911. 8:20:41slide, if you notice. We have got two
  13912. 8:20:43different lists for X and O. We are
  13913. 8:20:45going to compare this new input image
  13914. 8:20:47list that we have got with that of X and
  13915. 8:20:49O, right? So, first let us compare that
  13916. 8:20:51with X. Now, as I've told you earlier as
  13917. 8:20:54well, for X there are certain values
  13918. 8:20:55which will be higher, which is nothing
  13919. 8:20:57but first, fourth, fifth, 10th, and 11th
  13920. 8:20:59value, right? So, I'm going to sum
  13921. 8:21:01first, fourth, fifth, 10th, and 11th
  13922. 8:21:03value and I've got five. 1 + 1 + 1 + 1
  13923. 8:21:06and + 1. So, five times one, I've got
  13924. 8:21:08five. And now, I'm going to sum the
  13925. 8:21:10corresponding values of my input image
  13926. 8:21:11vector as well. So, the first value is
  13927. 8:21:130.9. Then, the fourth value is 0.87,
  13928. 8:21:16fifth value is 0.96, 10th value is 0.89,
  13929. 8:21:19and the 11th value is 0.94. So, after
  13930. 8:21:21this doing the sum of these values, I've
  13931. 8:21:23got 4.56. When I divide this by five, I
  13932. 8:21:25got 0.91, right? Now, this is for X.
  13933. 8:21:28Now, when I do the same process for O,
  13934. 8:21:30so in O, if you notice, I have second,
  13935. 8:21:32third, ninth, and 12th element vector
  13936. 8:21:35values is high. So, when I sum these
  13937. 8:21:37values, I get four.
  13938. 8:21:38And when I do the sum of the
  13939. 8:21:40corresponding values in my input image,
  13940. 8:21:41I've got 2.07. When I divide that by
  13941. 8:21:44four, I got 4 for 0.51.
  13942. 8:21:46So, now we notice that 0.91 is a higher
  13943. 8:21:49value compared to 0.51. So, we have when
  13944. 8:21:51we have compared our input image with
  13945. 8:21:52the values of X, we got a higher value
  13946. 8:21:55than the value that we have got after
  13947. 8:21:56comparing the input image with the
  13948. 8:21:58values of O. So, the input image is
  13949. 8:22:00classified as X. All right, so now let
  13950. 8:22:02us move towards our use case. So, this
  13951. 8:22:04is our use case, guys. So, over here,
  13952. 8:22:06what we are going to do is we are going
  13953. 8:22:08to train our model on different types of
  13954. 8:22:11dogs and cats images and then we are
  13955. 8:22:14going to provide an input and it will
  13956. 8:22:16classify whether the input is of a dog
  13957. 8:22:19or a cat. Now, let me tell you the steps
  13958. 8:22:21involved in it. So, what we are going to
  13959. 8:22:23do in the beginning is obviously first
  13960. 8:22:24we need to download the data set. After
  13961. 8:22:26that we are going to write a function to
  13962. 8:22:28encode the labels. Labels are nothing
  13963. 8:22:30but the dependent variable that we are
  13964. 8:22:32trying to predict. So, in our training
  13965. 8:22:34data and testing data, obviously we know
  13966. 8:22:36the labels, right? So, on that basis
  13967. 8:22:38only we can train our model. So, we are
  13968. 8:22:39going to encode those labels. After that
  13969. 8:22:41we'll resize the image to 50 cross 50
  13970. 8:22:43pixel and we are going to read it as a
  13971. 8:22:45grayscale image. Then we are going to
  13972. 8:22:47split the data, 24,000 images for
  13973. 8:22:50training and 50 for testing. Once this
  13974. 8:22:52is done, we are going to reshape the
  13975. 8:22:54data appropriately for TensorFlow. Now,
  13976. 8:22:56TensorFlow, I think everyone knows about
  13977. 8:22:58TensorFlow. It's nothing but a Python
  13978. 8:23:00library for implementing deep learning
  13979. 8:23:01models. Then we are going to build the
  13980. 8:23:03model, calculate the loss. It is nothing
  13981. 8:23:05but categorical cross entropy. Then we
  13982. 8:23:07are going to reduce the loss by using
  13983. 8:23:10Adam optimizer with a learning rate set
  13984. 8:23:11up set to point double zero one. Then we
  13985. 8:23:14are going to train the train the deep
  13986. 8:23:15neural network for 10 epochs and finally
  13987. 8:23:18we are going to make predictions. All
  13988. 8:23:19right, so I'll just quickly open my
  13989. 8:23:21PyCharm and I'll show you the code how
  13990. 8:23:22it looks like.
  13991. 8:23:24So, this is the code that I've written
  13992. 8:23:25in order to implement the use case. In
  13993. 8:23:26the beginning I need to import the
  13994. 8:23:28libraries that I require.
  13995. 8:23:31And once it is done, what I mean my
  13996. 8:23:32training data and the testing data. So,
  13997. 8:23:36train and test one contains as well as
  13998. 8:23:39testing data respectively. Then I've
  13999. 8:23:41taken my image size as 50 and learning
  14000. 8:23:44rate I've defined here and I've given a
  14001. 8:23:46name to my model. You can give whatever
  14002. 8:23:47name you want. All right, the first
  14003. 8:23:49thing that we saw we need to encode the
  14004. 8:23:51dependent variable. That's what we are
  14005. 8:23:52doing here. We are encoding our
  14006. 8:23:54dependent variable. So, whenever the
  14007. 8:23:57label is cat, then it will be converted
  14008. 8:23:59to an array of one comma zero and when
  14009. 8:24:01it is dog it will be converted to an
  14010. 8:24:02array of zero comma one. So, why why we
  14011. 8:24:04are actually encoding the label? Because
  14012. 8:24:06our code cannot understand the
  14013. 8:24:07categorical variable. So, we need to
  14014. 8:24:09encode it. Right? Next, what I'm doing
  14015. 8:24:11is I'm resizing my image to 50 cross 50
  14016. 8:24:14and I am converting it to a grayscale
  14017. 8:24:15image. Right? And once this is done, I'm
  14018. 8:24:17going to split my data set into training
  14019. 8:24:20and testing parts.
  14020. 8:24:21So, yeah, we are basically splitting the
  14021. 8:24:22data set into two parts for training and
  14022. 8:24:25testing.
  14023. 8:24:26And here
  14024. 8:24:27we are defining a model. So, you can
  14025. 8:24:29just I can just go ahead and throw in a
  14026. 8:24:32comment here.
  14027. 8:24:35Building the model. Yeah. So, so this is
  14028. 8:24:38where we are building the model. So,
  14029. 8:24:40basically what we have done here is we
  14030. 8:24:41have resized our image to 50 cross 50
  14031. 8:24:44cross one matrix and that is the size of
  14032. 8:24:46the input that we are using, right?
  14033. 8:24:48Talking about. Then, what we have done
  14034. 8:24:50here we have defined two filters and a
  14035. 8:24:53stride of five with an activation
  14036. 8:24:55function. After that
  14037. 8:24:57that we have added a pooling layer, max
  14038. 8:24:59pool layer. Okay? What we have done, we
  14039. 8:25:01have repeated the same process, but over
  14040. 8:25:04here we are taking 64 filters and five
  14041. 8:25:07passing it through a real activation
  14042. 8:25:08function. And after that we have
  14043. 8:25:10have a
  14044. 8:25:11repeated the 128 filters. After that we
  14045. 8:25:12have repeated for 64 filters, then for
  14046. 8:25:1432 filters. Then after that we are using
  14047. 8:25:17a fully connected layer with 1024
  14048. 8:25:19neurons. And finally we are using the
  14049. 8:25:21dropout layer with key probability of
  14050. 8:25:240.8 to finish our models. This is where
  14051. 8:25:26our model is actually finished. And then
  14052. 8:25:28what we are doing is we are using the
  14053. 8:25:30Adam optimizer to optimize our model.
  14054. 8:25:33So, basically whatever the loss that we
  14055. 8:25:35have, we are trying to reduce it. And
  14056. 8:25:37this is basically for your TensorBoard.
  14057. 8:25:39So, we are creating some log files and
  14058. 8:25:41then with that log file TensorBoard will
  14059. 8:25:43create a pretty fancy graphs for us that
  14060. 8:25:46helps us to visualize the entire model.
  14061. 8:25:48And then what we are doing is we are
  14062. 8:25:50trying to fit the model. And we have
  14063. 8:25:51defined epochs as 10, that is the number
  14064. 8:25:53of iterations that will happen will be
  14065. 8:25:5610. And yeah, so this is pretty much it.
  14066. 8:25:58Model name we have given. Then input is
  14067. 8:26:00X score to check the accuracy. Similarly
  14068. 8:26:04uh the target will be Y test labels
  14069. 8:26:07associated with that test data will be a
  14070. 8:26:09Y test and which we have encoded
  14071. 8:26:11basically. So this is how we are going
  14072. 8:26:13to actually calculate the accuracy and
  14073. 8:26:16we'll try to reduce the loss as much as
  14074. 8:26:17possible in 10 epochs. So till now our
  14075. 8:26:20model is complete. We are done with it.
  14076. 8:26:22Next what I'm doing is I'm feeding in
  14077. 8:26:24some random input from the test data and
  14078. 8:26:26I'm validating whether my model is
  14079. 8:26:28predicting it correct or not. All right.
  14080. 8:26:30So I've already trained the model
  14081. 8:26:32because it takes a lot of time and yeah
  14082. 8:26:34I cannot do it here. So I've already
  14083. 8:26:36trained the model and you can see that
  14084. 8:26:38the loss that came after the 10th epoch
  14085. 8:26:41is 0.2973
  14086. 8:26:43and the accuracy is somewhere around 88%
  14087. 8:26:45which is pretty good guys and yeah and
  14088. 8:26:48I've done the prediction on the test
  14089. 8:26:49data as well. So let me just show it to
  14090. 8:26:51you that. So this is the prediction that
  14091. 8:26:53it has done on few of the images in the
  14092. 8:26:55test data. So yeah it is a cat predicted
  14093. 8:26:57as cat cat predicted as a cat cat cat
  14094. 8:26:59cat and dogs as well. There are certain
  14095. 8:27:01dogs as well.
  14096. 8:27:04>> [music]
  14097. 8:27:08>> Why can't we use feed forward networks?
  14098. 8:27:11Now let us take an example of a feed
  14099. 8:27:13forward network that is used for image
  14100. 8:27:14classification. So we have trained this
  14101. 8:27:17particular network for classifying
  14102. 8:27:18various images of animals. Now if you
  14103. 8:27:20feed in an image of a dog it'll identify
  14104. 8:27:23that image and will provide a relevant
  14105. 8:27:25label to that particular image.
  14106. 8:27:26Similarly if you feed in an image of an
  14107. 8:27:28elephant it'll provide relevant label to
  14108. 8:27:31that particular image as well. Now if
  14109. 8:27:32you notice the new output that we have
  14110. 8:27:34got that is classifying an elephant has
  14111. 8:27:37no relation with the previous output
  14112. 8:27:39that is of a dog. Or you can say that
  14113. 8:27:41the output at time T is independent of
  14114. 8:27:44output at time T minus one. As we can
  14115. 8:27:46see that there is no relation between
  14116. 8:27:48the new output and the previous output.
  14117. 8:27:50So we can say that in feed forward
  14118. 8:27:52networks outputs are independent to each
  14119. 8:27:54other. Now But are few scenarios where
  14120. 8:27:56we actually need the previous output to
  14121. 8:27:57get the new output. Let us discuss one
  14122. 8:28:00such scenario.
  14123. 8:28:01Now what happens when you read a book?
  14124. 8:28:03You'll understand that book only on the
  14125. 8:28:05understanding of your previous words.
  14126. 8:28:07All right, so if I use a feedforward
  14127. 8:28:08network and try to predict the next word
  14128. 8:28:10in a sentence, I can't do that. Why
  14129. 8:28:12can't I do that? Because my output will
  14130. 8:28:15actually depend on the previous outputs.
  14131. 8:28:17But in the feedforward network, my new
  14132. 8:28:20output is independent of the previous
  14133. 8:28:21outputs. That is, output at t plus one
  14134. 8:28:24has no relation with output at t minus
  14135. 8:28:26two, t minus one and at t. So basically,
  14136. 8:28:28we cannot use feedforward networks for
  14137. 8:28:30predicting the next word in a sentence.
  14138. 8:28:32Similarly, you can think of many other
  14139. 8:28:34examples where we need the previous
  14140. 8:28:36output, some information from the
  14141. 8:28:37previous output, so as to infer the new
  14142. 8:28:39output. This is just one small example.
  14143. 8:28:41There are many other examples that you
  14144. 8:28:43can think of. So we'll move forward and
  14145. 8:28:45understand how we can solve this
  14146. 8:28:46particular problem. So over here, what
  14147. 8:28:48we have done, we have input at t minus
  14148. 8:28:50one. We'll feed it to our network, then
  14149. 8:28:53we'll get the output at t minus one.
  14150. 8:28:55Then at the next time stamp, that is at
  14151. 8:28:57time t, we have input at time t. That
  14152. 8:28:59will be given to our network along with
  14153. 8:29:01the information from the previous time
  14154. 8:29:03stamp, that is t minus one, and that
  14155. 8:29:06will help us to get the output at t.
  14156. 8:29:08Similarly, at output for t plus one, we
  14157. 8:29:11have two inputs. One is new input that
  14158. 8:29:13we give. Another is the information
  14159. 8:29:15coming from the previous time stamp,
  14160. 8:29:16that is t, in order to get the output at
  14161. 8:29:19time t plus one. Similarly, it can go
  14162. 8:29:21on. So over here, I have just written a
  14163. 8:29:23generalized way to represent it. There's
  14164. 8:29:25a loop where the information from the
  14165. 8:29:27previous time stamp is flowing. This is
  14166. 8:29:29how we can solve this particular
  14167. 8:29:30challenge. Now let us understand what
  14168. 8:29:32exactly are recurrent neural networks.
  14169. 8:29:35So for understanding recurrent neural
  14170. 8:29:36network, I'll take an analogy. Suppose
  14171. 8:29:38your gym trainer has made a schedule for
  14172. 8:29:40you. The exercises are repeated after
  14173. 8:29:42every third day.
  14174. 8:29:44Now this is the order of your exercises.
  14175. 8:29:46First day you'll be doing shoulder,
  14176. 8:29:47second day you'll be doing biceps, third
  14177. 8:29:49day you'll be doing cardio. And all
  14178. 8:29:50these exercises are repeated in a proper
  14179. 8:29:52order. Now, what happens when we use a
  14180. 8:29:54feedforward network for predicting the
  14181. 8:29:56exercise today? So, we'll provide in the
  14182. 8:29:58input such as day of the week, month of
  14183. 8:30:00the year, and health status. All right,
  14184. 8:30:02and we need to train our model or our
  14185. 8:30:04network on the exercises that we have
  14186. 8:30:06done in the past. After that, there'll
  14187. 8:30:07be a complex voting procedure involved
  14188. 8:30:09that will predict the exercise for us.
  14189. 8:30:11And that procedure won't be that
  14190. 8:30:13accurate. So, whatever output we'll get
  14191. 8:30:15won't be as accurate as we want it to
  14192. 8:30:17be. Now, what if I change my inputs and
  14193. 8:30:20I make my inputs as what exercise I've
  14194. 8:30:22done yesterday? So, if I've done
  14195. 8:30:23shoulder, then definitely today I'll be
  14196. 8:30:24doing biceps. Similarly, if I've done
  14197. 8:30:26biceps yesterday, today I'll be doing
  14198. 8:30:28cardio. Similarly, if I've done cardio
  14199. 8:30:30yesterday, today I'll be doing shoulder.
  14200. 8:30:32Now, there can be one scenario where you
  14201. 8:30:34are unable to go to gym for 1 day. Due
  14202. 8:30:36to some personal reasons, you could not
  14203. 8:30:38go to the gym. Now, what will happen at
  14204. 8:30:40that time?
  14205. 8:30:41We'll go one time step back and we'll
  14206. 8:30:43feed in what exercise that happened day
  14207. 8:30:45before yesterday. So, if the exercise
  14208. 8:30:47that happened day before yesterday was
  14209. 8:30:49shoulder, then yesterday there were
  14210. 8:30:50biceps exercises. All right, similarly,
  14211. 8:30:52biceps happened day before yesterday,
  14212. 8:30:54then yesterday would have been cardio
  14213. 8:30:56exercises. Similarly, if cardio would
  14214. 8:30:57have happened day before yesterday,
  14215. 8:30:59yesterday would have been shoulder
  14216. 8:31:00exercises. All right, and this
  14217. 8:31:02prediction, the prediction for the
  14218. 8:31:04exercise that happened yesterday, will
  14219. 8:31:06be fed back to our network and these
  14220. 8:31:08predictions will be used as inputs in
  14221. 8:31:10order to predict what exercise will
  14222. 8:31:12happen today. Similarly, if you have
  14223. 8:31:14missed your gym, say for 2 days, 3 days,
  14224. 8:31:16or 1 week, so you need to roll back. You
  14225. 8:31:19need to go to the last day when you went
  14226. 8:31:21to the gym. You need to figure out what
  14227. 8:31:23exercise you did on that day, feed that
  14228. 8:31:25as an input, and then only you'll be
  14229. 8:31:26getting the relevant output as to what
  14230. 8:31:28exercise will happen today.
  14231. 8:31:30Now, what I'll do, I'll convert these
  14232. 8:31:31things into a vector. Now, what is a
  14233. 8:31:33vector? Vector is nothing but a list of
  14234. 8:31:35numbers. All right, so this is the new
  14235. 8:31:37information, guys, along with the
  14236. 8:31:39information from the prediction at the
  14237. 8:31:40previous time step. So, we need both of
  14238. 8:31:43these in order to get the prediction at
  14239. 8:31:44time t. Imagine if I've done shoulder
  14240. 8:31:47exercises yesterday, so this will be
  14241. 8:31:49one, this will be zero, this will be
  14242. 8:31:50zero. Now, the prediction that will
  14243. 8:31:52happen will be biceps exercise because
  14244. 8:31:53if I have done shoulder yesterday, it's
  14245. 8:31:55related to biceps. So, my output will be
  14246. 8:31:57zero, one, and zero. And this is how
  14247. 8:31:59vectors work, guys. So, I hope you have
  14248. 8:32:01understood this, guys. Now, this is how
  14249. 8:32:03a neural network looks like, guys. We
  14250. 8:32:05have new information along with the
  14251. 8:32:08information from the previous time step.
  14252. 8:32:10The output that we have got in the
  14253. 8:32:11previous time step will certain
  14254. 8:32:13information from that. We'll feed into
  14255. 8:32:15our network as inputs, and then that
  14256. 8:32:17will help us to get the new output.
  14257. 8:32:19Similarly, this new output that we have
  14258. 8:32:21got will take some information from
  14259. 8:32:23that, feed in as an input to our network
  14260. 8:32:25along with the new information to get
  14261. 8:32:26the new prediction, and this process
  14262. 8:32:28keeps on repeating.
  14263. 8:32:29Now, let me show you the math behind the
  14264. 8:32:31recurrent neural networks.
  14265. 8:32:33So, this is the structure of a recurrent
  14266. 8:32:34neural network, guys. Let me explain you
  14267. 8:32:36what happens here. Now, consider at time
  14268. 8:32:38t equals to zero, we have input x
  14269. 8:32:40naught, and we need to figure out what
  14270. 8:32:41is x naught. So, according to this
  14271. 8:32:43equation, h of zero is equal to w i,
  14272. 8:32:47weight matrix, multiplied by our input x
  14273. 8:32:49of zero plus w r into h of zero minus
  14274. 8:32:54one, which is h of minus one, and time
  14275. 8:32:56can never be negative, so we this
  14276. 8:32:58particular equation cannot be applied
  14277. 8:33:00here, plus a bias. So, w i into x of
  14278. 8:33:03zero plus b h passes through a function
  14279. 8:33:05g of h to get h of zero over here. After
  14280. 8:33:08that, I want to calculate y naught. So,
  14281. 8:33:10for y naught, I'll multiply h of zero
  14282. 8:33:12with the weight matrix w i, and I'll add
  14283. 8:33:14a bias to it and pass it through a
  14284. 8:33:16function g of i to get y naught. Now, in
  14285. 8:33:18the next time step, that is at time t
  14286. 8:33:20equals to one, things become a bit
  14287. 8:33:22tricky. Now, let me explain you what
  14288. 8:33:24happens here. So, at time t equals to
  14289. 8:33:26one, I have input x one, I need to
  14290. 8:33:27figure out what is x one. So, for that,
  14291. 8:33:29I'll use this equation. So, I'll
  14292. 8:33:31multiply w i, that is the weight matrix,
  14293. 8:33:34by the input x one plus w r into h of
  14294. 8:33:38one minus one, which is zero. H of zero,
  14295. 8:33:40we know what we got from here. So, WR
  14296. 8:33:42into H of zero plus the bias, pass it
  14297. 8:33:45through a function G of H to get the
  14298. 8:33:47output as H1. Now, this H1 will use to
  14299. 8:33:50get Y1. We'll multiply H1 with WY plus a
  14300. 8:33:53bias and we'll pass it through a
  14301. 8:33:55function G of Y to get Y1.
  14302. 8:33:57Similarly, the next time stamp, that is
  14303. 8:33:59at time T equals to two, we have input
  14304. 8:34:01X2. We need to figure out what will be
  14305. 8:34:03H2. So, we'll multiply the weight matrix
  14306. 8:34:05WI with X of two plus WR into H of one
  14307. 8:34:08that we have got here plus B of H and
  14308. 8:34:11pass it through a function G of H to get
  14309. 8:34:13H of two. From H of two, we'll calculate
  14310. 8:34:15Y of two. WY into H of two plus BY, that
  14311. 8:34:18is the bias, pass it through a function
  14312. 8:34:20G of Y to get Y2. And this is how
  14313. 8:34:22recurrent neural network works, guys.
  14314. 8:34:24Now, you must be thinking how to train a
  14315. 8:34:26recurrent neural network.
  14316. 8:34:28So, a recurrent neural network uses back
  14317. 8:34:29propagation algorithm for training. But
  14318. 8:34:31back propagation happens for every time
  14319. 8:34:34stamp. That is why it is commonly called
  14320. 8:34:36as back propagation through time.
  14321. 8:34:38Over here, I won't be discussing back
  14322. 8:34:39propagation in detail. I'll just give
  14323. 8:34:41you a brief introduction of what it is.
  14324. 8:34:44Now, with back propagation, there are
  14325. 8:34:45certain issues, namely vanishing and
  14326. 8:34:47exploding gradients. Let us see those
  14327. 8:34:49one by one.
  14328. 8:34:50So, in vanishing gradient, what happens?
  14329. 8:34:52When you use back propagation, you tend
  14330. 8:34:54to calculate the error, which is nothing
  14331. 8:34:56but the actual output that you already
  14332. 8:34:58know minus the model output, output that
  14333. 8:35:01you got through your model, and the
  14334. 8:35:02square of that.
  14335. 8:35:03So, you figure out the error. With that
  14336. 8:35:05error, what do you do? You tend to find
  14337. 8:35:08out the change in error with respect to
  14338. 8:35:10change in weight or any variable. So,
  14339. 8:35:13we'll call it weight here. So, change of
  14340. 8:35:15error with respect to weight multiplied
  14341. 8:35:17by learning rate will give you the
  14342. 8:35:18change in weight. Then you need to add
  14343. 8:35:20that change in weight to the old weight
  14344. 8:35:22to get the new weight. All right? So,
  14345. 8:35:25obviously, what we are trying to do, we
  14346. 8:35:26are trying to reduce the error. So, for
  14347. 8:35:28that, we need to figure out what will be
  14348. 8:35:30the change in error if my variables are
  14349. 8:35:32changed, right? So, that way we can get
  14350. 8:35:34the change in in variable and add it to
  14351. 8:35:36our old variable to get the new
  14352. 8:35:37variable. Now, over here, what can
  14353. 8:35:39happen if the value dE by dW, that is a
  14354. 8:35:42gradient, or you can say the rate of
  14355. 8:35:44change of error with respect to our
  14356. 8:35:45variable weight, becomes very small than
  14357. 8:35:47one, like it is 0.00 something. So, if
  14358. 8:35:50you multiply that with the a learning
  14359. 8:35:52rate, which is definitely smaller than
  14360. 8:35:54one, then you get the change of weight,
  14361. 8:35:56which is negligible. All right? So,
  14362. 8:35:59there might be certain examples where,
  14363. 8:36:00you know, you are trying to predict,
  14364. 8:36:02say, a next word in a sentence, and that
  14365. 8:36:03sentence is pretty long. For example, if
  14366. 8:36:05I say, "I went to France {dash} {dash}
  14367. 8:36:08{dash} I went to France." Then there are
  14368. 8:36:10certain words. Then I say, "Few of them
  14369. 8:36:13speak {dash}." Now, I need to predict
  14370. 8:36:15speak, what will come after speak. So,
  14371. 8:36:18for that, I need to go back in time and
  14372. 8:36:20check what was the context, which will
  14373. 8:36:22be very complex. And due to that,
  14374. 8:36:24there'll be a lot of iterations. And
  14375. 8:36:26because of that, this error, this change
  14376. 8:36:28in weight, will become very small, very
  14377. 8:36:31small. So, the new weight that we'll get
  14378. 8:36:32will be actually almost equal to your
  14379. 8:36:35old weight. So, there won't be any
  14380. 8:36:37updation of weight that will be
  14381. 8:36:38happening. And that is nothing but your
  14382. 8:36:40vanishing gradient. All right, I'll
  14383. 8:36:42repeat it once more. So, what happens in
  14384. 8:36:44back propagation, you first calculate
  14385. 8:36:46the error. This error is nothing but the
  14386. 8:36:48difference between the actual output and
  14387. 8:36:49the model output and the square of that.
  14388. 8:36:52With that error, we figure out what will
  14389. 8:36:53be the change in error when we change a
  14390. 8:36:55particular variable, say, weight. So, dE
  14391. 8:36:58by dW, multiply it with learning rate to
  14392. 8:37:00get the change in the variable or change
  14393. 8:37:02in the weight. Now, we'll add that
  14394. 8:37:03change in the weight to our old weight
  14395. 8:37:05to get the new weight. This is back
  14396. 8:37:07propagation, is guys, all right? I'm
  14397. 8:37:08just giving you a small introduction to
  14398. 8:37:10back propagation. Now, consider a
  14399. 8:37:12scenario where you need to predict the
  14400. 8:37:14next word in a sentence. And your
  14401. 8:37:15sentence is something like this. "I have
  14402. 8:37:18been to France." Then there are a lot of
  14403. 8:37:20words. After that, few people speak. And
  14404. 8:37:24then you need to predict what comes
  14405. 8:37:25after speak. Now, if I need to do that,
  14406. 8:37:27I need to go back and understand the
  14407. 8:37:29context, what is it talking about?
  14408. 8:37:32And that is nothing but your long-term
  14409. 8:37:34dependencies. So, what happens during
  14410. 8:37:35long-term dependencies if this DE by DW
  14411. 8:37:38becomes very small? Then, when you
  14412. 8:37:40multiply it with N, which is again
  14413. 8:37:41smaller than one, you get delta W, which
  14414. 8:37:44will be very, very small. That will be
  14415. 8:37:46negligible. So, the new weight that
  14416. 8:37:48you'll get here will be almost equal to
  14417. 8:37:50your old weight. So, I hope you're
  14418. 8:37:52getting my point. So, this new weight
  14419. 8:37:54So, there will be no updation of
  14420. 8:37:56weights, guys. This new weight will
  14421. 8:37:58definitely be will always be almost
  14422. 8:38:00equal to our old weight. There won't be
  14423. 8:38:02any learning here. So, that is nothing
  14424. 8:38:04but your vanishing gradient problem.
  14425. 8:38:06Similarly, when I talk about exploding
  14426. 8:38:08gradient, it is just the opposite of
  14427. 8:38:09vanishing gradient. So, what happens
  14428. 8:38:11when your gradient or DE by DW becomes
  14429. 8:38:13very uh large, becomes greater than
  14430. 8:38:15greater than one? All right? And you
  14431. 8:38:17have some long-term dependencies. So, at
  14432. 8:38:19that time, your DE by DW will keep on
  14433. 8:38:22increasing. Delta W will become large.
  14434. 8:38:24And because of that, your weights, the
  14435. 8:38:26new weight with that will come will be
  14436. 8:38:28very different from your old weight. So,
  14437. 8:38:30these two are the problems with back
  14438. 8:38:31propagation. Now, let us see how to
  14439. 8:38:33solve these problems.
  14440. 8:38:35Now, exploding gradients can be solved
  14441. 8:38:37with the help of truncated BPTT, back
  14442. 8:38:38propagation through time. So, instead of
  14443. 8:38:40starting back propagation at the last
  14444. 8:38:42time stamp, we can choose a smaller time
  14445. 8:38:44stamp like 10. Or we can clip the
  14446. 8:38:47gradients at a threshold. So, there can
  14447. 8:38:48be a threshold value where we can, you
  14448. 8:38:50know, clip the gradients. And we can
  14449. 8:38:52adjust the learning rate as well. Now,
  14450. 8:38:53for vanishing gradient, we can use a
  14451. 8:38:55ReLU activation function. We have
  14452. 8:38:56discussed ReLU activation function in
  14453. 8:38:58artificial neural network tutorial,
  14454. 8:38:59guys. Similarly, we can also use LSTM
  14455. 8:39:02and GRUs. In this tutorial, we'll be
  14456. 8:39:04discussing LSTMs that are long
  14457. 8:39:06short-term memory units. Now, let us
  14458. 8:39:09understand what exactly are LSTMs.
  14459. 8:39:12So, guys, we saw what are the two
  14460. 8:39:13limitations with the recurrent neural
  14461. 8:39:15networks. Now, we'll understand how we
  14462. 8:39:17can solve that with the help of LSTMs.
  14463. 8:39:19Now, what are LSTMs? Long short-term
  14464. 8:39:21memory networks, usually called as
  14465. 8:39:23LSTMs, are nothing but a special kind of
  14466. 8:39:25recurrent neural network. And these
  14467. 8:39:27recurrent neural networks are capable of
  14468. 8:39:29learning long-term dependencies. Now,
  14469. 8:39:31what are long-term dependencies? I've
  14470. 8:39:33discussed on the previous slide, but
  14471. 8:39:35I'll just explain it to you here as
  14472. 8:39:36well. Now, what happens sometimes we
  14473. 8:39:38only need to look at the recent
  14474. 8:39:39information to perform the present task.
  14475. 8:39:42Now, let me give you an example.
  14476. 8:39:43Consider a language model trying to
  14477. 8:39:45predict the next word based on the
  14478. 8:39:47previous ones. If we are trying to
  14479. 8:39:49predict the last word in the sentence,
  14480. 8:39:51say, "The clouds are in the sky." So, we
  14481. 8:39:54don't need any further context. It's
  14482. 8:39:55pretty obvious that the next word is
  14483. 8:39:57going to be sky. Now, in such cases
  14484. 8:39:59where the gap between the relevant
  14485. 8:40:01information and the place that it's
  14486. 8:40:03needed is small, RNNs can learn to use
  14487. 8:40:06the past information. And at that time,
  14488. 8:40:08there won't be such problems like
  14489. 8:40:09vanishing and exploding gradient. But,
  14490. 8:40:11there are few cases where we need more
  14491. 8:40:14context. Consider trying to predict the
  14492. 8:40:16last word in the text, "I grew up in
  14493. 8:40:19France." Then, there are some words.
  14494. 8:40:20After that comes, "I speak fluent
  14495. 8:40:23French." Now, recent information
  14496. 8:40:25suggests that word is probably the name
  14497. 8:40:27of a language. But, if we want to narrow
  14498. 8:40:29down which language, we need the context
  14499. 8:40:32of France from further back. And it's
  14500. 8:40:35entirely possible for the gap between
  14501. 8:40:37the relevant information and the point
  14502. 8:40:39where it is needed to become very large.
  14503. 8:40:41And this is nothing but long-term
  14504. 8:40:42dependencies. And the LSTMs are capable
  14505. 8:40:45of handling such long-term dependencies.
  14506. 8:40:47Now, LSTMs also have a chain-like
  14507. 8:40:50structure like recurrent neural
  14508. 8:40:51networks. Now, all the recurrent neural
  14509. 8:40:53networks have the form of a chain of
  14510. 8:40:54repeating modules of neural networks.
  14511. 8:40:56Now, in standard RNNs, the repeating
  14512. 8:40:58module will have a very simple structure
  14513. 8:41:00such as a single tan h layer that you
  14514. 8:41:01can see. Now, this tan h layer is
  14515. 8:41:03nothing but a squashing function. Now,
  14516. 8:41:05what I mean by squashing function is to
  14517. 8:41:07convert my values between minus one and
  14518. 8:41:10one. All right, that's why we use tan h.
  14519. 8:41:12And this is an example of an RNN. Now,
  14520. 8:41:15we'll understand what exactly are LSTMs.
  14521. 8:41:17Now, this is a structure of an LSTM. All
  14522. 8:41:20If you notice, LSTM also have a chain
  14523. 8:41:22like structure. But the ripple has
  14524. 8:41:24different structures. Instead of having
  14525. 8:41:26single neural network here, there are
  14526. 8:41:28four interacting in a very special way.
  14527. 8:41:30Now, the key to LSTM is the cell state.
  14528. 8:41:32Now, this particular line that I'm
  14529. 8:41:34highlighting, this is what what is
  14530. 8:41:36called the cell state. The horizontal
  14531. 8:41:38line running through the top of the
  14532. 8:41:39diagram. So, this is nothing but your
  14533. 8:41:40cell state. Now, you can consider the
  14534. 8:41:42cell state as a kind of a conveyor belt.
  14535. 8:41:45It runs straight down the entire chain
  14536. 8:41:47with only some minor linear
  14537. 8:41:48interactions. Now, what I'll do, I'll
  14538. 8:41:50give you a walk through of LSTM step by
  14539. 8:41:52step, all right? So, we'll start with
  14540. 8:41:54the first step.
  14541. 8:41:55All right, guys. So, the first step in
  14542. 8:41:57our LSTM is to decide what information
  14543. 8:41:59we are going to throw away from the cell
  14544. 8:42:01state. And you know what is the cell
  14545. 8:42:03state, right? I've discussed in the
  14546. 8:42:04previous slide. Now, this decision is
  14547. 8:42:06made by the sigmoid layer. So, the layer
  14548. 8:42:09that I'm highlighting with my cursor, it
  14549. 8:42:10is the sigmoid layer. Called the forget
  14550. 8:42:12gate layer. It looks at HT minus one,
  14551. 8:42:15that is the information from the
  14552. 8:42:16previous time step, and XT, which is the
  14553. 8:42:19new input, and outputs a number between
  14554. 8:42:21zeros and ones for each number in the
  14555. 8:42:23cell state, CT minus one, which is
  14556. 8:42:25coming from the previous time step. A
  14557. 8:42:27one represents completely keep this,
  14558. 8:42:29while a zero represents completely get
  14559. 8:42:31rid of this. Now, if we go back to our
  14560. 8:42:33example of a language model trying to
  14561. 8:42:35predict the next word based on all the
  14562. 8:42:37previous ones, in such a problem, the
  14563. 8:42:39cell state might include the gender of
  14564. 8:42:41the present subject so that the correct
  14565. 8:42:43pronouns can be used. When we see a new
  14566. 8:42:45subject, we want to forget the gender of
  14567. 8:42:47the old subject, right? We want to use
  14568. 8:42:50the gender of the new subject. So, we'll
  14569. 8:42:52forget the gender of the previous
  14570. 8:42:53subject here. This is just an example to
  14571. 8:42:56explain you what is happening here.
  14572. 8:42:58Uh now, let me explain you the equations
  14573. 8:42:59which I've written here. So, FT will be
  14574. 8:43:02uh combining with the cell state later
  14575. 8:43:04on, that I'll tell you. So, currently,
  14576. 8:43:06FT will be nothing but the weight matrix
  14577. 8:43:09multiplied by HT minus one and XT, and
  14578. 8:43:13uh plus the bias, and this equation is
  14579. 8:43:15passed through a sigmoid layer. All
  14580. 8:43:17right? And we get an output that is zero
  14581. 8:43:19and one. Zero means completely get rid
  14582. 8:43:21of this and one means completely keep
  14583. 8:43:23this. All right, so this is what
  14584. 8:43:24basically is happening in the first
  14585. 8:43:26step. Now, let us see what happens in
  14586. 8:43:28the next step. So, the next step is to
  14587. 8:43:30decide what information we are going to
  14588. 8:43:32store. In the previous step, we decided
  14589. 8:43:34what information we are going to keep,
  14590. 8:43:35but here we are going to decide what
  14591. 8:43:37information we are going to store here.
  14592. 8:43:39All right, what new information we are
  14593. 8:43:41going to store in the cell state. Now,
  14594. 8:43:42this has two parts. First, a sigmoid
  14595. 8:43:45layer, this is called a sigmoid layer
  14596. 8:43:46and which is also known as an input gate
  14597. 8:43:48layer, decide which values will update.
  14598. 8:43:51All right, so what values we need to
  14599. 8:43:52update. Then there's also a tan h layer
  14600. 8:43:54that creates a vector of the candidate
  14601. 8:43:56values c bar of t minus one that will be
  14602. 8:44:00added to the state later on. All right,
  14603. 8:44:02so let me explain it to you in a simpler
  14604. 8:44:03terms. So, whatever input that we are
  14605. 8:44:05getting from the previous time stamp and
  14606. 8:44:07the new input, it will be passed through
  14607. 8:44:09a sigmoid function, which will give us i
  14608. 8:44:11of t. All right, and this i of t will be
  14609. 8:44:14multiplied by c t but coming from the
  14610. 8:44:17previous time stamp and the new input
  14611. 8:44:19with that is passed through a tan h that
  14612. 8:44:21will result in c t. And this will be
  14613. 8:44:23later added on to our cell state. In the
  14614. 8:44:25next step, we'll combine these two to
  14615. 8:44:27update the states. Now, let me explain
  14616. 8:44:28the equations. So, i of t will be what?
  14617. 8:44:31Weight matrix and then we have h t minus
  14618. 8:44:33one comma x t multiplied by the weight
  14619. 8:44:35matrix plus the bias pass it through a
  14620. 8:44:37sigmoid function, we get i of t. c bar
  14621. 8:44:39of t will get by passing a weight matrix
  14622. 8:44:41h t minus one x t plus bias through a
  14623. 8:44:44tan h square function and we'll get c
  14624. 8:44:46bar of t. All right, so as I've told you
  14625. 8:44:48earlier as well in the next step, we'll
  14626. 8:44:49combine these two to update the state.
  14627. 8:44:51Let us see how we do that. So, now is
  14628. 8:44:54the time to update the old cell state c
  14629. 8:44:56t minus one with the new cell state c t.
  14630. 8:44:59All right, in the previous steps, we
  14631. 8:45:00have already decided what to do. We just
  14632. 8:45:02need to actually do it. So, what we'll
  14633. 8:45:04do, we'll multiply the old cell state c
  14634. 8:45:06t minus one with f t that we got in the
  14635. 8:45:08first step for getting the things that
  14636. 8:45:10we decided to forget earlier in the
  14637. 8:45:12first step if you can recall. Then what
  14638. 8:45:14we do, we add it to IT and CT. Then we
  14639. 8:45:18add it by the term that will come after
  14640. 8:45:20multiplication of IT and C bar T. And
  14641. 8:45:22this new candidate value scaled by how
  14642. 8:45:24much we decided to update each state
  14643. 8:45:26value. All right? So, in the case of the
  14644. 8:45:29language model that we are discussing,
  14645. 8:45:30this is where we would actually drop the
  14646. 8:45:32information about the old subject gender
  14647. 8:45:35and add the new information as we
  14648. 8:45:36decided in the previous steps. So, I
  14649. 8:45:38hope you are able to follow me guys. All
  14650. 8:45:40right? So, let us move forward and we'll
  14651. 8:45:42see what is the next step. Now, our last
  14652. 8:45:44step is to decide what we are going to
  14653. 8:45:46output. And this output will depend on
  14654. 8:45:48our cell state, but [snorts] it will be
  14655. 8:45:50a filtered version. Now, finally what we
  14656. 8:45:52need to do is we need to decide what we
  14657. 8:45:53are going to output. And this output
  14658. 8:45:55will be based on our cell state. First,
  14659. 8:45:57we need to pass HT minus one and XT
  14660. 8:45:59through a sigmoid activation function so
  14661. 8:46:02that we get output that is OT. All
  14662. 8:46:04right? And this OT will be in turn
  14663. 8:46:06multiplied by the cell state after
  14664. 8:46:08passing it through an NH squashing
  14665. 8:46:10function or an activation function. And
  14666. 8:46:12why we do that? Just to push the values
  14667. 8:46:14between minus one and one. So, after
  14668. 8:46:17multiplying OT, that is this value, and
  14669. 8:46:20a tan at CT, we'll get the output H2,
  14670. 8:46:23which will be our new output. And that
  14671. 8:46:25will only output the part that we
  14672. 8:46:27decided to. Whatever we have decided in
  14673. 8:46:29the previous steps, it will only output
  14674. 8:46:30that value. All right? Now, I'll take
  14675. 8:46:32the example of that language model
  14676. 8:46:34again. Since it just saw a subject, it
  14677. 8:46:36might want to output information
  14678. 8:46:38relevant to a verb and in case that's
  14679. 8:46:40what is coming next.
  14680. 8:46:42For example, it might output whether the
  14681. 8:46:44subject is singular or plural. So, that
  14682. 8:46:46we know what form of a verb should be
  14683. 8:46:48conjugated into. All right? And uh you
  14684. 8:46:51can see from the uh you can see the
  14685. 8:46:52equations as well. Again, we have a
  14686. 8:46:54sigmoid function. Then that uh whatever
  14687. 8:46:57output we get from there, we multiply it
  14688. 8:46:59with tan at CT to get the new output.
  14689. 8:47:01All right, guys? So, this is basically
  14690. 8:47:03uh LSTMs in a nutshell. So, in the first
  14691. 8:47:06step, we decided what we need to forget.
  14692. 8:47:08In the next step, we decided what are we
  14693. 8:47:10going
  14694. 8:47:11to our cell state, what new information
  14695. 8:47:13going to add to cell state, and we were
  14696. 8:47:15taking example of the gender throughout
  14697. 8:47:17this whole process. All right? And in
  14698. 8:47:18the third step, what we do, we actually
  14699. 8:47:20combined it to get the new cell state.
  14700. 8:47:22Now, in the fourth step, what we did, we
  14701. 8:47:24finally got the output that we want. And
  14702. 8:47:27how we did that? Just by passing HT - 1
  14703. 8:47:29and HT through a sigmoid function,
  14704. 8:47:31multiplying it with the tan H CT, the
  14705. 8:47:33tan H new cell state, and we get the new
  14706. 8:47:36output. Fine, guys? So, this is what
  14707. 8:47:38basically LSTM is, guys. Now, we'll look
  14708. 8:47:40at a use case where we'll be using LSTM
  14709. 8:47:43to predict the next word in a sentence.
  14710. 8:47:45All right? Let me show you how we are
  14711. 8:47:46going to do that.
  14712. 8:47:48So, this is what we are trying to do in
  14713. 8:47:49our use case, guys. We'll feed LSTM with
  14714. 8:47:52correct sequences from the text of three
  14715. 8:47:54symbols. For example, had a general and
  14716. 8:47:57a label that is counsel in this
  14717. 8:47:59particular example. Eventually, our
  14718. 8:48:01network will learn to predict the next
  14719. 8:48:03symbol correctly. So, obviously, we need
  14720. 8:48:05to train it on something. Let us see
  14721. 8:48:06what we are going to train it on.
  14722. 8:48:08So, we'll be training LSTM to predict
  14723. 8:48:10the next word using a sample short story
  14724. 8:48:12that you can see over here.
  14725. 8:48:14All right? So, it has basically 112
  14726. 8:48:16unique symbols. So, even comma and full
  14727. 8:48:18stop are considered as symbols. All
  14728. 8:48:20right? So, this is what we are going to
  14729. 8:48:22train it on.
  14730. 8:48:23So, technically, we know that LSTMs can
  14731. 8:48:25only understand real numbers. All right?
  14732. 8:48:27So, what we need to do is we need to
  14733. 8:48:29convert these unique symbols into a
  14734. 8:48:31unique integer value based on the
  14735. 8:48:33frequency of occurrence. And like that,
  14736. 8:48:35we'll create a dictionary. For example,
  14737. 8:48:37we have had here that will have value
  14738. 8:48:3920. A will have value six. General will
  14739. 8:48:42have value 33. All right? And then, what
  14740. 8:48:45happens, our LSTM will create a 112
  14741. 8:48:48element vector that will contain the
  14742. 8:48:50probability of each of these words or
  14743. 8:48:53each of these unique integer values. All
  14744. 8:48:55right? So, since 0.6 has the highest
  14745. 8:48:57probability in this particular vector,
  14746. 8:48:59it'll pick the index value of 0.6. Then,
  14747. 8:49:02it will see it what symbol is attached
  14748. 8:49:04to that particular integer value. So, 37
  14749. 8:49:06is attached to counsel. So, this will be
  14750. 8:49:08our prediction, which is absolutely
  14751. 8:49:09correct as the label is also counsel
  14752. 8:49:11according to our training data. All
  14753. 8:49:13right. So, this is what we are going to
  14754. 8:49:15do in our use case. So, guys, this is
  14755. 8:49:17what we'll be doing in our today's use
  14756. 8:49:18case. Now, I'll quickly open my PyCharm
  14757. 8:49:21and I'll show you how you can implement
  14758. 8:49:22it using Python. We'll be using
  14759. 8:49:24TensorFlow, which is a popular Python
  14760. 8:49:26library for implementing deep neural
  14761. 8:49:28networks or neural networks in general.
  14762. 8:49:30All right. So, I'll quickly open my
  14763. 8:49:31PyCharm now. So, guys, this is my
  14764. 8:49:33PyCharm and I over here I've already
  14765. 8:49:35written the code in order to execute the
  14766. 8:49:36use case that we have. So, first we need
  14767. 8:49:38to do is import the libraries, NumPy for
  14768. 8:49:41arrays, TensorFlow we know,
  14769. 8:49:42tensorflow.contrib from that we need to
  14770. 8:49:44import RNN and random collections and
  14771. 8:49:46time. All right. So, this particular
  14772. 8:49:48block of code is used to evaluate the
  14773. 8:49:51time taken for the training. After that,
  14774. 8:49:53we have log_path and this log_path is
  14775. 8:49:56basically telling us the path where the
  14776. 8:49:58graph will be stored. All right. So,
  14777. 8:49:59there will be a graph that will be
  14778. 8:50:00created and then that graph will be
  14779. 8:50:02launched. Then only our RNN model will
  14780. 8:50:04be executed. Then that's how TensorFlow
  14781. 8:50:06works, guys.
  14782. 8:50:07So, that graph will be created in this
  14783. 8:50:09particular path. All right. And we are
  14784. 8:50:11using summary writer. So, that will
  14785. 8:50:13actually create the log file that will
  14786. 8:50:14be used in order to display the graph
  14787. 8:50:17using TensorBoard. All right. So, then
  14788. 8:50:19we have defined training_file, which
  14789. 8:50:21will have our story on which we'll train
  14790. 8:50:23our model on. Then what we need to do is
  14791. 8:50:25read this file. So, how are we going to
  14792. 8:50:27do that? First is read line by line
  14793. 8:50:29whatever content that we have in our
  14794. 8:50:31file. Then we are going to strip it.
  14795. 8:50:33That means we are going to remove the
  14796. 8:50:35first and the last white space. Then
  14797. 8:50:37again, we are splitting it just to
  14798. 8:50:40remove all the white spaces that are
  14799. 8:50:41there. After that, we're creating an
  14800. 8:50:43array and then we're reshaping it. Now,
  14801. 8:50:45in during the reshape, if you notice
  14802. 8:50:46this minus one value tells us the
  14803. 8:50:48compatibility. All right. So, when
  14804. 8:50:50you're reshaping it, you need to make
  14805. 8:50:51sure that
  14806. 8:50:53you know, we are providing in the
  14807. 8:50:54correct parameters to reshape it. So,
  14808. 8:50:56you can convert a three cross two matrix
  14809. 8:50:58to a two cross three matrix, like right?
  14810. 8:51:01So, just to make sure that that it is
  14811. 8:51:02compatible enough, we add this minus one
  14812. 8:51:04and it'll be done automatically. All
  14813. 8:51:06right? Then, return content. After that,
  14814. 8:51:09what we are doing, we are feeding in the
  14815. 8:51:11training data that we have, training
  14816. 8:51:12{underscore} file. We are feeding in our
  14817. 8:51:14story and calling the function read
  14818. 8:51:16{underscore} data. Then, what we are
  14819. 8:51:17doing, we are creating a dictionary.
  14820. 8:51:19What is a dictionary? We all know, key
  14821. 8:51:20value pairs based on the frequency of
  14822. 8:51:22occurrences of each symbol. All right?
  14823. 8:51:24So, from here, collections.counter
  14824. 8:51:26words.most_common. So, most common words
  14825. 8:51:29with their frequency of occurrence,
  14826. 8:51:30there'll be a dictionary created. And
  14827. 8:51:32after that, uh we'll call this dict
  14828. 8:51:34function and this dict function will
  14829. 8:51:36feed in word and which is equal to
  14830. 8:51:38length of dictionary. That means
  14831. 8:51:40whatever the length of that particular
  14832. 8:51:42dictionary, how many time it is
  14833. 8:51:43repeated. So, we'll have the frequency
  14834. 8:51:45as well as a symbol. That'll be our key
  14835. 8:51:47value pair and we're reversing it as
  14836. 8:51:49well.
  14837. 8:51:50Then, what we are doing, we are calling
  14838. 8:51:51it build {underscore} data set and we're
  14839. 8:51:54feeding in our training data there. This
  14840. 8:51:56is our vocabulary size, which is nothing
  14841. 8:51:57but the length of your dictionary. Then,
  14842. 8:51:59we have defined various parameters such
  14843. 8:52:01as learning rate, uh iterations or
  14844. 8:52:03epochs. Then, we have display step and
  14845. 8:52:05{underscore} input. Now, learning rate,
  14846. 8:52:07we all know what it is, uh the steps in
  14847. 8:52:09which our variables are updated.
  14848. 8:52:11Training {underscore} iterations is
  14849. 8:52:12nothing but your epochs, the total
  14850. 8:52:14number of iterations. So, we have given
  14851. 8:52:1550,000 iterations here. Then, we have
  14852. 8:52:17display {underscore} step, that is
  14853. 8:52:191,000, which is basically your batch
  14854. 8:52:20size. So, batch size is what? After
  14855. 8:52:23every 1,000 epochs, you'll see the
  14856. 8:52:24output. All right? So, it'll be
  14857. 8:52:25processing it in batches of 1,000
  14858. 8:52:27iterations. Then, we have n {underscore}
  14859. 8:52:29input as three. Now, the number of units
  14860. 8:52:31in the RNN cell, we'll keep it as 512.
  14861. 8:52:34Then, we need to define X and Y. So, X
  14862. 8:52:37will be our placeholder that will have
  14863. 8:52:38the input values and Y will have all the
  14864. 8:52:41labels. All right? vocab size.
  14865. 8:52:44So, X is a placeholder where we'll be
  14866. 8:52:46feeding in our input dictionary.
  14867. 8:52:47Similarly, Y is also one more
  14868. 8:52:49placeholder and it'll have a shape of
  14869. 8:52:51none {comma} vocab size. Vocab size we
  14870. 8:52:53have defined earlier.
  14871. 8:52:55As you can see, which is nothing but the
  14872. 8:52:56length of your dictionary. Then we're
  14873. 8:52:57defining weights as well as biases.
  14874. 8:53:00After that, we have defined our model.
  14875. 8:53:02All right. So, this is how we are going
  14876. 8:53:03to define it. We'll
  14877. 8:53:05create a function RNN when we'll have X
  14878. 8:53:08weights and biases. And after that, we
  14879. 8:53:10are calling in RNN.multi_rnn_cell
  14880. 8:53:12function. And this is basically to
  14881. 8:53:14create a two-layer LSTM. And each layer
  14882. 8:53:17has n_hidden_units.
  14883. 8:53:19After that, what we are doing, we are
  14884. 8:53:20generating the predictions. But once we
  14885. 8:53:22have generated the prediction, there are
  14886. 8:53:24n_input_outputs,
  14887. 8:53:25but we only want the last output. For
  14888. 8:53:28that, we have written this particular
  14889. 8:53:29line. And then finally, we are making a
  14890. 8:53:30prediction. We are calling this RNN and
  14891. 8:53:32function feeding in X weights and
  14892. 8:53:34biases. After that, we are calculating
  14893. 8:53:36the loss as and then we are optimizing
  14894. 8:53:38it. For calculating the loss, we are
  14895. 8:53:41using reduce_mean softmax_cross_entropy.
  14896. 8:53:44And this will give us basically the
  14897. 8:53:46probability of each symbol. And then we
  14898. 8:53:48are optimizing it using RMS uh prop
  14899. 8:53:50optimizer. All right. And this gives
  14900. 8:53:52actually a better accuracy than Adam
  14901. 8:53:54optimizer. And that's the reason why we
  14902. 8:53:56are using it. Then we are going to
  14903. 8:53:57calculate the accuracy. And after that,
  14904. 8:54:00we are going to initialize the variables
  14905. 8:54:01that we have used. As we have seen in
  14906. 8:54:03TensorFlow, that we need to initialize
  14907. 8:54:04all the variables, unlike constants and
  14908. 8:54:06placeholders in TensorFlow. All right.
  14909. 8:54:08And once we are done with that, we are
  14910. 8:54:10feeding in our values, then calculating
  14911. 8:54:12the accuracy, how accurate it is. And
  14912. 8:54:14then when optimization is done, we are
  14913. 8:54:16calculating the elapsed time as well.
  14914. 8:54:18So, that will give us how much time it
  14915. 8:54:20took in order to train our model. Then
  14916. 8:54:22this is just to run the TensorBoard on
  14917. 8:54:24our local host 6006. And yeah, and this
  14918. 8:54:28particular block of code is is used in
  14919. 8:54:30order to handle the exceptions. So,
  14920. 8:54:32exceptions can be like whatever word
  14921. 8:54:34that we are putting in might not be
  14922. 8:54:36there in our dictionary or might not be
  14923. 8:54:37there in our training data. So, those
  14924. 8:54:39exceptions will be handled here. And if
  14925. 8:54:41it is not there in our dictionary, then
  14926. 8:54:42it will print word not in our
  14927. 8:54:44dictionary. All right. So, fine guys.
  14928. 8:54:46Let's
  14929. 8:54:47input some values and we'll have some
  14930. 8:54:49fun with this model. All right? So, the
  14931. 8:54:51first thing that I'm going to feed in is
  14932. 8:54:53had general. So, whenever I feed in
  14933. 8:54:56these three values, had a general,
  14934. 8:54:58there'll be a story that will be
  14935. 8:54:59generated by feeding back the predicted
  14936. 8:55:01output as the next symbol in the inputs.
  14937. 8:55:04All right? So, when I feed in had a
  14938. 8:55:05general, so it'll predict the correct
  14939. 8:55:07output as counsel. And this counsel will
  14940. 8:55:10be fed back as a part of the new input
  14941. 8:55:12and our new input will be a general
  14942. 8:55:14counsel. So, it'll be a general counsel.
  14943. 8:55:16All right? So, these three words will
  14944. 8:55:18become our new input to predict the new
  14945. 8:55:19output, which is two. All right? And so
  14946. 8:55:21on. So, surprisingly, LSTM actually
  14947. 8:55:24creates a story that, you know, somehow
  14948. 8:55:26makes sense. So, let's just read it. Had
  14949. 8:55:28a general counsel to consider what
  14950. 8:55:30measures they could take to outwit their
  14951. 8:55:32common enemy, the cat. By this means, we
  14952. 8:55:35should always know when she was about
  14953. 8:55:37and could easily. All right? So, somehow
  14954. 8:55:39it actually makes sense when you feed in
  14955. 8:55:40that. So, what'll happen when you feed
  14956. 8:55:42in these three inputs, it'll predict the
  14957. 8:55:44next word, that is counsel. After that,
  14958. 8:55:46it'll take counsel and it'll feed back
  14959. 8:55:48as an input along with a general. So, a
  14960. 8:55:50general counsel will be your next input
  14961. 8:55:53to predict two. Similarly, in the next
  14962. 8:55:55iteration, it'll take general counsel
  14963. 8:55:57two and predict counsel for us. And this
  14964. 8:55:59will keep on repeating.
  14965. 8:56:02>> [music]
  14966. 8:56:09>> processing and why do we even need it?
  14967. 8:56:12You see, natural language
  14968. 8:56:13analysis in both audible data as well as
  14969. 8:56:16the text document. NLP system can
  14970. 8:56:19capture meaning from an input such as
  14971. 8:56:21sentences, paragraphs, pages, and give
  14972. 8:56:23out a desired output based on our
  14973. 8:56:25application. So, why do we need NLP? You
  14974. 8:56:28see, natural language processing helps
  14975. 8:56:29computer communicate with humans in
  14976. 8:56:31their own language and scale other
  14977. 8:56:33language-related task. For example, NLP
  14978. 8:56:36makes it possible for computers to read
  14979. 8:56:38text, hear speeches, interpret it,
  14980. 8:56:41measure the sentiment, and then
  14981. 8:56:42determine which part of it are
  14982. 8:56:44important. Today's machines can analyze
  14983. 8:56:46more language-based data than humans.
  14984. 8:56:48That too with consistency, accuracy, and
  14985. 8:56:51in an unbiased manner. More or less, we
  14986. 8:56:53all know that there is a staggering
  14987. 8:56:55amount of unstructured data that is
  14988. 8:56:57generated every day. Be it from a
  14989. 8:56:59medical record or to a social media.
  14990. 8:57:01Automating NLP task will critically be
  14991. 8:57:04helpful in future for analyzing text and
  14992. 8:57:06speech data efficiently. So, moving
  14993. 8:57:09ahead, let us now see the ways we can
  14994. 8:57:10process our textual data. We can process
  14995. 8:57:13our textual data in one of two ways. One
  14996. 8:57:15is a machine learning way, and other one
  14997. 8:57:17is a deep learning method. In machine
  14998. 8:57:19learning, we can make use of algorithms
  14999. 8:57:21such as bag of words, TF-IDF to classify
  15000. 8:57:24and predict the desired output. But, the
  15001. 8:57:26drawback of this is that these machine
  15002. 8:57:28learning algorithms do not consider the
  15003. 8:57:30context of the word or a sequence. Here,
  15004. 8:57:32the way it works is on the base on
  15005. 8:57:34number of times a word is repeating, a
  15006. 8:57:36probability is derived out of it, and
  15007. 8:57:38then performs a classification task.
  15008. 8:57:41This is the reason why we have deep
  15009. 8:57:42learning model for NLP tasks. Speaking
  15010. 8:57:45about deep learning model, we have
  15011. 8:57:46something like recurrent neural network,
  15012. 8:57:48LSTM, transformer network, Google's BERT
  15013. 8:57:51algorithm, and many more. This deep
  15014. 8:57:53learning model learns the pattern of a
  15015. 8:57:55word or a sequence, and then tries to
  15016. 8:57:57predict the desired outcome of the task.
  15017. 8:57:59We can perform NLP using deep learning
  15018. 8:58:01in one of two ways. One by pre-trained
  15019. 8:58:03models such as Google's Word2vec or
  15020. 8:58:06global vector models. And the other way
  15021. 8:58:08is to train our own model. If you're
  15022. 8:58:10trying to train our own model, it would
  15023. 8:58:12require a very huge amount of data and
  15024. 8:58:14also a compute power to support it. In
  15025. 8:58:16most of the cases, we'll be using
  15026. 8:58:18pre-trained models. Moving ahead, let us
  15027. 8:58:20now discuss recurrent neural networks.
  15028. 8:58:23As I mentioned earlier, the bag of word
  15029. 8:58:25or TF-IDF model for processing our text
  15030. 8:58:28is very inefficient. As I mentioned
  15031. 8:58:30earlier, the bag of word or TF-IDF model
  15032. 8:58:32that was used in machine learning to
  15033. 8:58:34process our textual data is very
  15034. 8:58:36inefficient as it takes one word at a
  15035. 8:58:38time and also the context of the word in
  15036. 8:58:41which it is being spoken about is
  15037. 8:58:42totally ignored. Although this would
  15038. 8:58:44give us some prediction, but we can
  15039. 8:58:45expect lot of loss. This is why we use
  15040. 8:58:48RNN model or recurrent neural network.
  15041. 8:58:50Here it requires a sequential data.
  15042. 8:58:53Now you might be wondering what does
  15043. 8:58:54this sequential data mean, right? In
  15044. 8:58:56simple words, sequential data is
  15045. 8:58:58dependent on the past value. What I'm
  15046. 8:59:00trying to say here is that we can read
  15047. 8:59:02our document, right? We can read our
  15048. 8:59:04document or English documents only from
  15049. 8:59:06left to right side. Or take example of
  15050. 8:59:08stock price prediction. We cannot
  15051. 8:59:10randomly place the dates, right? If you
  15052. 8:59:12have to make a prediction for next 2-3
  15053. 8:59:14months, we'll obviously refer to the
  15054. 8:59:16data that was previously recorded. So
  15055. 8:59:18this is what a sequential data means. So
  15056. 8:59:21what makes RNN capable of handling
  15057. 8:59:23sequential data? Well, you see RNN model
  15058. 8:59:26makes use of something called a state.
  15059. 8:59:28This is nothing but a temporary memory
  15060. 8:59:30that stores the previous data. As you
  15061. 8:59:32can see here in an image, this is the
  15062. 8:59:34general architecture of a recurrent
  15063. 8:59:36neural network. So let me now move to my
  15064. 8:59:38canvas and show you how recurrent neural
  15065. 8:59:40network works and what are its internal
  15066. 8:59:42workings.
  15067. 8:59:44All right. So as we have seen in an
  15068. 8:59:45image, right? Back in our slide, you saw
  15069. 8:59:47that, you know, recurrent neural network
  15070. 8:59:49have something called as states and then
  15071. 8:59:51we also had some boxes, right? So what
  15072. 8:59:54does this boxes represents? So let me
  15073. 8:59:56quickly draw over here and show you what
  15074. 8:59:57does this box represents. So if you
  15075. 9:00:00remember, right? It goes something like
  15076. 9:00:02we have a box here. Okay? And then we
  15077. 9:00:04also had another box.
  15078. 9:00:07And then let's consider like two more
  15079. 9:00:09boxes that would be sufficient. Okay?
  15080. 9:00:11And then the way we provide input for
  15081. 9:00:13our recurrent neural network is over
  15082. 9:00:15here.
  15083. 9:00:16Okay? Now based on our application, we
  15084. 9:00:18can demand our recurrent neural network
  15085. 9:00:20to provide output in either one of three
  15086. 9:00:22ways. It can either be like we can have
  15087. 9:00:24multiple inputs, as you can see here, or
  15088. 9:00:26we can have single input and multiple
  15089. 9:00:28outputs, or the other way around is we
  15090. 9:00:30can have single input and single output.
  15091. 9:00:32So here what we'll consider is we are
  15092. 9:00:34using multiple inputs and also we have
  15093. 9:00:36multiple outputs.
  15094. 9:00:38Okay. So now let us see as this is a
  15095. 9:00:40supervised learning model, right?
  15096. 9:00:42Obviously it will have some kind of
  15097. 9:00:43input and then it will also have labels
  15098. 9:00:45to clarify the data. So let's take
  15099. 9:00:48something like if this is our X data,
  15100. 9:00:49right? So if this is the X data, so
  15101. 9:00:52let's take this as a list and then over
  15102. 9:00:53here we'll have data which would
  15103. 9:00:55represent something like X1, then we'll
  15104. 9:00:57have X2,
  15105. 9:00:59then we'll have X3,
  15106. 9:01:00and then we'll have something like Xn.
  15107. 9:01:04Okay? So what does this X over here
  15108. 9:01:07represents? See, here let's take an
  15109. 9:01:09example that Uri, which was a movie, is
  15110. 9:01:12a good movie. We can have couple more
  15111. 9:01:13X's and we'll just put it here as good
  15112. 9:01:16movie.
  15113. 9:01:27Okay. So this is nothing but the inputs.
  15114. 9:01:30Each of these words. Now what our model
  15115. 9:01:32over here does, let's take an example
  15116. 9:01:34that we are trying to find name entity
  15117. 9:01:36prediction, which stands for NER, right?
  15118. 9:01:38So what does name entity prediction does
  15119. 9:01:40is, you know, trains the model in such a
  15120. 9:01:42manner that it finds a pattern. So as
  15121. 9:01:44you can see here, we this is Uri, right?
  15122. 9:01:47Uri, it's it's something like a name,
  15123. 9:01:48right? So this this goes as one.
  15124. 9:01:50And is is not a name. So this would be
  15125. 9:01:52zero. Good is this is also not a name
  15126. 9:01:55and movie is also not a name. So the
  15127. 9:01:57output over here would be something like
  15128. 9:01:581000. All right? So now let's see what
  15129. 9:02:01are the inputs and how they would look
  15130. 9:02:03like. So first off we have inputs. So if
  15131. 9:02:06this is our X data, so we'll have input
  15132. 9:02:08something like Xi. This should be in
  15133. 9:02:10lower case.
  15134. 9:02:11Okay? So this would be X1,
  15135. 9:02:13then we'll have X2,
  15136. 9:02:15we'll have X3, and then we'll have X4.
  15137. 9:02:19Similarly, the Y the output over here
  15138. 9:02:21would be something like we'll give it Y
  15139. 9:02:23hat of 1, [snorts]
  15140. 9:02:25Y hat of 2,
  15141. 9:02:27Y hat off three, at the same time we'll
  15142. 9:02:29also have why hat off four. Now if
  15143. 9:02:31you're wondering what does this why hat
  15144. 9:02:33represents, why hat over here is nothing
  15145. 9:02:34but
  15146. 9:02:35predicted values.
  15147. 9:02:39And why is nothing but you know the
  15148. 9:02:41labeled value or trained values.
  15149. 9:02:45And now as we all know that the main
  15150. 9:02:48thing or the main feature behind
  15151. 9:02:50recurrent neural network is nothing but
  15152. 9:02:52the states. And the way the states are
  15153. 9:02:54represented is by using the state vector
  15154. 9:02:56or it can also be called as context
  15155. 9:02:57vector. So the state vector over here
  15156. 9:02:59starts with A. Let me give it as a
  15157. 9:03:01different color here. Let me give it as
  15158. 9:03:02blue. So here we'll have A and this
  15159. 9:03:05should be zero. Okay? And now this A
  15160. 9:03:08would be passed down to this. Okay, so
  15161. 9:03:10this would be A of one and over here
  15162. 9:03:12we'll have A of two,
  15163. 9:03:14A of three and then finally A of four.
  15164. 9:03:18I'm pretty sure you might be wondering
  15165. 9:03:19what does this A contains, right? So
  15166. 9:03:21over here as I have mentioned, let me
  15167. 9:03:23take this very example. So here Uri is
  15168. 9:03:26good movie, okay? And so what does A1
  15169. 9:03:29will contain? A1 will contain Uri. See,
  15170. 9:03:33if we take the normal algorithm, what's
  15171. 9:03:34going to happen is it will just consider
  15172. 9:03:36only this one particular block. Okay? If
  15173. 9:03:39this was just a normal algorithm or
  15174. 9:03:41something which is used in olden days,
  15175. 9:03:42it will just consider this particular
  15176. 9:03:44block and it won't be considering the
  15177. 9:03:45previous values. So this previous value
  15178. 9:03:48is being stored over here in the form of
  15179. 9:03:50a memory. So now A1 will contain Uri and
  15180. 9:03:53what A2 will contain over here? So A2
  15181. 9:03:57will be having something like Uri is.
  15182. 9:04:00And similarly as it goes through here,
  15183. 9:04:02it will collect each and every word and
  15184. 9:04:04finally A4 will be complete context
  15185. 9:04:06vector. So it will have the entire
  15186. 9:04:07sentence Uri is a good movie.
  15187. 9:04:12Okay, I hope you understood what does
  15188. 9:04:14this A signifies over here. Let me
  15189. 9:04:16quickly erase all of these. Okay, so the
  15190. 9:04:19another important part which goes over
  15191. 9:04:20here in any machine learning or deep
  15192. 9:04:22learning model is nothing but the
  15193. 9:04:24weights.
  15194. 9:04:25So, let's consider that we have weights
  15195. 9:04:27which is nothing but W or let's take it
  15196. 9:04:29as U. Let me give another color here, U,
  15197. 9:04:32V, and W. In this RNN model, right? What
  15198. 9:04:35we're going to do is we're going to have
  15199. 9:04:37a single or we'll have just only one
  15200. 9:04:39weights, right? So, the U weight is same
  15201. 9:04:42for all of these inputs here. V weight
  15202. 9:04:45over here is same for all the outputs.
  15203. 9:04:47The weight W is same for all our state
  15204. 9:04:50metrics. Okay? So, this is how the
  15205. 9:04:52weights are determined.
  15206. 9:04:54So, now to get a better understanding of
  15207. 9:04:56what's happening over here and to derive
  15208. 9:04:58a mathematical equation for feedforward
  15209. 9:05:00network, let's take a single block over
  15210. 9:05:02here. And let's see how this internal
  15211. 9:05:04working of this is working, okay? So,
  15212. 9:05:06over here we'll have a box.
  15213. 9:05:09Okay? And this will be our input. So,
  15214. 9:05:11let's give our input as X of T because
  15215. 9:05:14this is a generalized model, right? And
  15216. 9:05:16now what you're going to do over here is
  15217. 9:05:17this would be A
  15218. 9:05:19of T minus 1. That's because A of T is
  15219. 9:05:22will be present over here. And now if
  15220. 9:05:24you have an output which is present over
  15221. 9:05:26here and this would be nothing but Y hat
  15222. 9:05:29of T.
  15223. 9:05:30All right? So, as I've mentioned earlier
  15224. 9:05:31that we will be having weights. So, the
  15225. 9:05:33weights over here is nothing but U, V,
  15226. 9:05:35and W. So, what's happening over here?
  15227. 9:05:38Okay. So, first off, there will be a
  15228. 9:05:41matrix multiplication between these two
  15229. 9:05:42values.
  15230. 9:05:44Okay? And then there will be a matrix
  15231. 9:05:45multiplication between these two values.
  15232. 9:05:47And once they are done, we'll add these
  15233. 9:05:49two values and then give a activation
  15234. 9:05:51function to it. Okay? And the activation
  15235. 9:05:54function that will be used over here is
  15236. 9:05:55tan H. So, if I have to put it in an
  15237. 9:05:58equation form over here, so the product
  15238. 9:06:00of these two, X of T
  15239. 9:06:02times or it would be a dot product
  15240. 9:06:05of U, okay? And the sum of these two, so
  15241. 9:06:09it will be T minus 1 times W. And now
  15242. 9:06:12what we're going to do is we're going to
  15243. 9:06:13pass an activation function which is
  15244. 9:06:14nothing but tan H over here.
  15245. 9:06:16So, let's just give a small dotted
  15246. 9:06:18notation over here.
  15247. 9:06:20So, whatever the output which comes
  15248. 9:06:22it'll be performing an addition of these
  15249. 9:06:25two and also give an activation function
  15250. 9:06:27which will represent here by f.
  15251. 9:06:29All right? So, this is how we get an
  15252. 9:06:31activation function over here. So, what
  15253. 9:06:33does this value signifies this? This is
  15254. 9:06:35nothing but this output over here, a of
  15255. 9:06:37t. So, I hope you understand what is a
  15256. 9:06:39of t. A of t is this value. Okay, let me
  15257. 9:06:42quickly highlight that for you. So, a of
  15258. 9:06:44t is this value. Okay? So, a of t is
  15259. 9:06:47represented over here. And the way we
  15260. 9:06:49get this is by this particular equation.
  15261. 9:06:52Now, what about the value for y hat of t
  15262. 9:06:54or the prediction of y of t? So, for y
  15263. 9:06:57of t, it's going to be something like
  15264. 9:06:58this. So, y hat of t, this would be
  15265. 9:07:01nothing but we have to multiply this
  15266. 9:07:03particular a of t with this matrix over
  15267. 9:07:06here. Okay? Or the weights I can say.
  15268. 9:07:08So, it's going to be nothing but a of t
  15269. 9:07:11times the weight matrix that is v. And
  15270. 9:07:14then we have to pass activation
  15271. 9:07:16function. Usually, the activation
  15272. 9:07:17function that is going to be used over
  15273. 9:07:19here is softmax or sigmoid. It totally
  15274. 9:07:21depends upon the what kind of output
  15275. 9:07:24you're expecting. All right? And this is
  15276. 9:07:26how it works. And now, an important
  15277. 9:07:28thing that I would like to mention over
  15278. 9:07:30here is that we also have to add
  15279. 9:07:32something called as bias. So, it would
  15280. 9:07:34be b over here. I'll just represented
  15281. 9:07:36this by a different color.
  15282. 9:07:38So, this is an equation for our
  15283. 9:07:40recurrent neural network in a
  15284. 9:07:41mathematical form.
  15285. 9:07:43So, now what's going to happen is in
  15286. 9:07:44order to find the loss, right? We have
  15287. 9:07:46to subtract whatever value we have
  15288. 9:07:47predicted with the given value. So, the
  15289. 9:07:50predicted value over here, let's say
  15290. 9:07:51this is as capital y. Okay? And we also
  15291. 9:07:54have been given the train value. So,
  15292. 9:07:55this would be looking something like
  15293. 9:07:57this.
  15294. 9:07:58So, we'll obviously have the train
  15295. 9:07:59values.
  15296. 9:08:01Let's just give a random output, but
  15297. 9:08:03there'll be four. So, it'll be 0 but the
  15298. 9:08:05train values. And now we'll have the
  15299. 9:08:07other, this is nothing but the given
  15300. 9:08:08label data. So, this would be correct
  15301. 9:08:10values. So, 1 0 0 1. As be correct
  15302. 9:08:12values. So, 1 0 0 1. As this is a
  15303. 9:08:15supervised learning, so we obviously
  15304. 9:08:17will be given this label data over here.
  15305. 9:08:19So, now what this will do is this will
  15306. 9:08:20subtract each of these values and
  15307. 9:08:23calculate a loss. So, how do you
  15308. 9:08:25calculate a loss, right?
  15309. 9:08:27Okay. So, as you can see here, if I give
  15310. 9:08:30this L, like let me change the color
  15311. 9:08:32here. So, if I say L is loss, so this
  15312. 9:08:35would be nothing but Y hat of first
  15313. 9:08:37value minus the actual value. Okay, so
  15314. 9:08:40this is nothing but the predicted value
  15315. 9:08:42and this is nothing but the given value.
  15316. 9:08:44So, if I try to find out the loss of
  15317. 9:08:46this, then what this would look like is
  15318. 9:08:48this is just for the one value, right?
  15319. 9:08:49So, if I have to do it for all the
  15320. 9:08:50values, then it would be the summation.
  15321. 9:08:52So, it would be nothing but loss is
  15322. 9:08:54equal to summation I which ranges from 1
  15323. 9:08:57to n, right? And then we'll have loss
  15324. 9:09:01and then theta I. Okay, so this LI over
  15325. 9:09:04here represent these values. Okay, so
  15326. 9:09:06this is just for one. So, if you want me
  15327. 9:09:07to explain you this in detail, so over
  15328. 9:09:09here we have Y1, Y2, Y3 and Y4. How this
  15329. 9:09:12would look like is something So, this
  15330. 9:09:14would be for one plus Y hat of two minus
  15331. 9:09:18actual value of Y of two, then summation
  15332. 9:09:21predicted value of Y3 minus the actual
  15333. 9:09:24value of Y3 and then it'll be predicted
  15334. 9:09:28value of Y4 minus the actual value of
  15335. 9:09:30Y4. And when I perform addition over
  15336. 9:09:33here or when I perform summation, I can
  15337. 9:09:35generalize this equation into this form.
  15338. 9:09:37But we're not yet done over here. As you
  15339. 9:09:39can see, we have something called as
  15340. 9:09:41theta values.
  15341. 9:09:42So, what does this theta values
  15342. 9:09:43represents?
  15343. 9:09:44You see, the main agenda behind finding
  15344. 9:09:47a loss is to increase our accuracy,
  15345. 9:09:48right? So, if I say this is my gradient
  15346. 9:09:51descent and this is the lowest global
  15347. 9:09:54minima, right? So, my agenda over here
  15348. 9:09:56is to reach this global minima. So, now
  15349. 9:09:59in order for me to do this, in order to
  15350. 9:10:01increase my accuracy, I have to
  15351. 9:10:02obviously change my values. So, how do I
  15352. 9:10:05do that? How do I increase an accuracy?
  15353. 9:10:07It's obviously by these weights. Okay?
  15354. 9:10:09It's by this W, U, and V. So, W, U, and
  15355. 9:10:13V are nothing but the weights. So, what
  15356. 9:10:15this theta over here represents, let me
  15357. 9:10:17quickly erase this.
  15358. 9:10:19Okay, so what this theta over here
  15359. 9:10:20represents is nothing but the values of
  15360. 9:10:22W, U, and V. So, now what we're going to
  15361. 9:10:25do is we're going to have a partial
  15362. 9:10:27derivative of the loss with respect to
  15363. 9:10:31U, and then we'll have a partial
  15364. 9:10:33derivative of loss with respect to V,
  15365. 9:10:35and then we'll also have partial
  15366. 9:10:37derivative of loss with respect to W.
  15367. 9:10:40These are nothing but weights, and we
  15368. 9:10:41are trying to train the weights.
  15369. 9:10:44This is done in order to increase our
  15370. 9:10:45accuracy.
  15371. 9:10:46Okay? And the way these models get
  15372. 9:10:48trained is with the help of back
  15373. 9:10:50propagation. And the way the back
  15374. 9:10:51propagation works over here is by
  15375. 9:10:53partial derivative. So, what I'm trying
  15376. 9:10:55to say here is as if I get my weights
  15377. 9:10:57over here, so let me just take another
  15378. 9:10:59color. So, this V we know that it is
  15379. 9:11:01totally dependent upon this value over
  15380. 9:11:03here. This won't be V, this would be W.
  15381. 9:11:05Okay? So, this W value will be totally
  15382. 9:11:07dependent on the previous one. And this
  15383. 9:11:09value will be dependent upon this one,
  15384. 9:11:10and this value will be dependent upon
  15385. 9:11:12this one, and this one would be finally
  15386. 9:11:13dependent on this. So, this is in order
  15387. 9:11:16to move from here to here to here and
  15388. 9:11:18then to here, we'll use something called
  15389. 9:11:20as partial derivatives, right? So, this
  15390. 9:11:23how we update our old weights.
  15391. 9:11:25All right. Another important concept
  15392. 9:11:27that make neural network or recurrent
  15393. 9:11:29neural network very important is nothing
  15394. 9:11:31but embedding layer. So, let me quickly
  15395. 9:11:33draw a boundary over here.
  15396. 9:11:35Okay. So, embedding layer.
  15397. 9:11:40All right. So, what does this embedding
  15398. 9:11:42layer signifies? Okay, so if I take a
  15399. 9:11:44convention way, right? So, like let's
  15400. 9:11:46say that, you know, our X or, you know,
  15401. 9:11:49our input over here, this is nothing but
  15402. 9:11:51X over here, right? So, the what is the
  15403. 9:11:52dimension or the shape of this X? The
  15404. 9:11:54shape of this X is nothing but 1 {comma}
  15405. 9:11:574, right? 1 {comma} 4. Similarly, it's 1
  15406. 9:11:59{comma} 4 here, 1 {comma} 4, and 1
  15407. 9:12:02{comma} 4. But, this is not usually the
  15408. 9:12:04case when you're working with the real
  15409. 9:12:06world examples. When you're working with
  15410. 9:12:08real world examples, you won't be having
  15411. 9:12:09four different words, right? We'll
  15412. 9:12:11obviously have lots of values. So, for
  15413. 9:12:13example, let's say we have X and the
  15414. 9:12:16shape of this X is something like we
  15415. 9:12:18have 10,000 values. Okay? So, now if I
  15416. 9:12:21try to feed this 10,000 values into my
  15417. 9:12:25into my network over here, obviously I
  15418. 9:12:26would be using batch propagation. So, it
  15419. 9:12:29would take a lot of time, right? Because
  15420. 9:12:31it's 10,000 values after all. So, in
  15421. 9:12:33order to overcome this, what we're going
  15422. 9:12:34to do is we're going to reduce the size
  15423. 9:12:36of this. Okay? So, basically what word
  15424. 9:12:38embedding layer does is think that we
  15425. 9:12:40have a matrix. This is our input layer,
  15426. 9:12:43right? Our input word. So, that we call
  15427. 9:12:45this as sparse matrix.
  15428. 9:12:47This is nothing but, you know, an
  15429. 9:12:49individual value. That's XI or X1, X2,
  15430. 9:12:51X3. Let's for generalization we'll give
  15431. 9:12:53you here as XI. And the shape of this is
  15432. 9:12:55nothing but 1,V. V here represents the
  15433. 9:12:58end number of dimensions. And one is
  15434. 9:13:00because it's just going to be one word,
  15435. 9:13:02right? So, it's going to be one. So, if
  15436. 9:13:03I consider with respect to this, this is
  15437. 9:13:05one and this is V. Okay?
  15438. 9:13:07So, now what will happen is if V is is
  15439. 9:13:1010,000 or pretty great,
  15440. 9:13:12it obviously we won't have much
  15441. 9:13:14we're going to lose a lot of time on
  15442. 9:13:15computation and also take huge compute
  15443. 9:13:18power. So, what we'll do is I'll
  15444. 9:13:19multiply this with an embedding layer.
  15445. 9:13:24And what this embedding matrix does is
  15446. 9:13:26it is basically a set of features. So,
  15447. 9:13:28this is something you know, this is a
  15448. 9:13:30black box model. And we'll all this
  15449. 9:13:32would do is this would attract couple of
  15450. 9:13:34features. If you want me to give you a
  15451. 9:13:36better analogy of what this is, and if
  15452. 9:13:38you want me to compare this with respect
  15453. 9:13:40to a CNN, which is nothing but another
  15454. 9:13:41great algorithm for image processing
  15455. 9:13:43using deep learning, right? Over there
  15456. 9:13:45we are going to use something called as
  15457. 9:13:46filters or kernels. So, each of those
  15458. 9:13:48filters or kernels is responsible for
  15459. 9:13:50extracting one specific feature, right?
  15460. 9:13:52So, this is what embedding layer does.
  15461. 9:13:54And what this would do, for example, now
  15462. 9:13:57let's say that the size of embedding
  15463. 9:13:58layer is V {comma} K. K is something
  15464. 9:14:00that we provide an input over here. So,
  15465. 9:14:02now what this would happen is this would
  15466. 9:14:05give us a new matrix or embedded matrix
  15467. 9:14:08whose size would be 1 {comma} K. I'm
  15468. 9:14:10pretty sure you didn't understand this
  15469. 9:14:12because over here I'm using, you know,
  15470. 9:14:13these these letters. So, in order to
  15471. 9:14:15make you better understand this, what
  15472. 9:14:16I'm going to do is let me take a matrix
  15473. 9:14:18over here. Okay? Let me take something
  15474. 9:14:20like, you know, because this would be in
  15475. 9:14:22an embedded form, right? So, this would
  15476. 9:14:23be a sparse matrix. So, it would be like
  15477. 9:14:250 0 0 0 1 then we'll have 0 0 and so on.
  15478. 9:14:30Okay? So, now what this would do is
  15479. 9:14:33we'll also create a matrix over here.
  15480. 9:14:35The size of this would be something
  15481. 9:14:36similar to that of
  15482. 9:14:38V.
  15483. 9:14:39This is nothing but 1 {comma} V and over
  15484. 9:14:42here it would be K. So, the matrix shape
  15485. 9:14:44over here would be V {comma} K, right?
  15486. 9:14:47To give you a better analogy, let me
  15487. 9:14:48also draw a couple of boxes here.
  15488. 9:14:51So, this is the matrix and this is the
  15489. 9:14:52matrix over here again. So, now what
  15490. 9:14:54will happen over here is when I try to
  15491. 9:14:56perform this matrix multiplication,
  15492. 9:14:58right? What this would do, you know, as
  15493. 9:15:00everything is zero and only one value is
  15494. 9:15:02true, so let's say this is the one
  15495. 9:15:04value, right? And this would be going
  15496. 9:15:06across, you know, from left to right and
  15497. 9:15:08this would be from top to bottom. Only
  15498. 9:15:10one part over here would be marked and
  15499. 9:15:12rest everything would be zero.
  15500. 9:15:13Therefore, reducing the dimension. Let
  15501. 9:15:16me give some random values like 0.5,
  15502. 9:15:181.8, 0.5, just some random values. So,
  15503. 9:15:22now this would obviously reduce the size
  15504. 9:15:24of 1 {comma} K. So, what I'm trying to
  15505. 9:15:26say here is now, for example, say that I
  15506. 9:15:29have a size over here as 1 {comma}
  15507. 9:15:3110,000. Okay, which is a very huge
  15508. 9:15:33matrix and the shape of this, let's say
  15509. 9:15:35that it's 10,000 {comma} 200.
  15510. 9:15:39When I perform this embedding, right? Or
  15511. 9:15:40embedding, the shape of the new matrix
  15512. 9:15:43would be nothing but 1 {comma} 200. If I
  15513. 9:15:45compare this part over here to the
  15514. 9:15:47embedding whatever we have received over
  15515. 9:15:49here. So, let me just give a quick
  15516. 9:15:51brief. So, this is our embedded layer.
  15517. 9:15:53So, this is really 1,200. So, you will
  15518. 9:15:56see that we have decreased the
  15519. 9:15:57dimensions by a drastic amount. Okay, so
  15520. 9:16:00this is 10,000 and this is only 200. And
  15521. 9:16:02this would be very efficient when we are
  15522. 9:16:04trying to feed this to our recurrent
  15523. 9:16:06neural network. And let me quickly show
  15524. 9:16:08you how this would go. So, first let me
  15525. 9:16:10draw our architecture. Let's take this
  15526. 9:16:14blocks like this.
  15527. 9:16:17And then we'll have an output over here.
  15528. 9:16:19Okay. So, this would be our inputs,
  15529. 9:16:21right? So, let me give something like
  15530. 9:16:23this.
  15531. 9:16:24So, initially, we used to provide X
  15532. 9:16:26values over here, right? Now, we won't
  15533. 9:16:28be doing that. We won't be providing any
  15534. 9:16:30X values directly. Instead of that, what
  15535. 9:16:32I'm going to do is I'll have an
  15536. 9:16:33embedding layer over here.
  15537. 9:16:38And this will have the X values. So, let
  15538. 9:16:41me give here as X of 1, so X of 2, X of
  15539. 9:16:453, and then we'll have X of 4, and then
  15540. 9:16:49similarly, let's take this model to be
  15541. 9:16:51multiple input and single output. So,
  15542. 9:16:53here we'll be have Y hat of T. And then
  15543. 9:16:56we'll have weights, obviously. So, this
  15544. 9:16:58would be U, V, and W. And this is
  15545. 9:17:02nothing but our matrix over here. So,
  15546. 9:17:04this would be A of 0, A of 1, A of 2, A
  15547. 9:17:09of 3,
  15548. 9:17:10and finally A of 4.
  15549. 9:17:12Okay, so this is our context matrix. And
  15550. 9:17:14obviously, we'll be performing an
  15551. 9:17:15activation function here. So, I'll just
  15552. 9:17:17give it a F. You can put F, you can put
  15553. 9:17:19G, it's totally up to you. So, this is
  15554. 9:17:21how our recurrent neural network would
  15555. 9:17:22actually work. All right?
  15556. 9:17:25So, now that we know how RNN works, let
  15557. 9:17:27us now understand what is LSTM. Or we
  15558. 9:17:30can also say it as long short-term
  15559. 9:17:31memory. You see, traditional RNNs are
  15560. 9:17:34not good at capturing long-range
  15561. 9:17:36dependencies. What I mean to say here is
  15562. 9:17:38that when we tend to work with a very
  15563. 9:17:39huge data set and multiple RNN layer, we
  15564. 9:17:42are at the risk of vanishing gradient
  15565. 9:17:44problem. Now, you might be wondering
  15566. 9:17:46what is this vanishing gradient, right?
  15567. 9:17:48Well, you see when training a very deep
  15568. 9:17:50neural network, gradient or the
  15569. 9:17:52derivatives decrease exponentially as it
  15570. 9:17:54propagates down the layer. This is known
  15571. 9:17:56as vanishing gradient problem. These
  15572. 9:17:58gradients are actually used to update
  15573. 9:18:00the weights of a neural network. But
  15574. 9:18:02when the gradients vanish, these weights
  15575. 9:18:04will not get updated. In the worst case
  15576. 9:18:06scenario, it will completely stop the
  15577. 9:18:08neural network from training. This
  15578. 9:18:10vanishing gradient problem is a common
  15579. 9:18:12issue in very deep neural networks. So
  15580. 9:18:15to overcome this vanishing gradient
  15581. 9:18:16problem in RNNs, long short-term memory
  15582. 9:18:19was introduced. You see LSTM or long
  15583. 9:18:22short memory is a modification to RNNs
  15584. 9:18:24hidden layer. LSTM is capable of
  15585. 9:18:26remembering RNNs weights and their
  15586. 9:18:28inputs over a very long period of time.
  15587. 9:18:31In LSTM, in addition to the hidden
  15588. 9:18:32state, cell state is passed down to the
  15589. 9:18:34next block. The way LSTM works is that
  15590. 9:18:37it can capture long-range dependencies,
  15591. 9:18:40that is old weights. It can have memory
  15592. 9:18:42of previous inputs for a very extended
  15593. 9:18:44time duration. The way LSTM cell does
  15594. 9:18:46this is by using three main gates. First
  15595. 9:18:49one is a forget gate. Forget gate
  15596. 9:18:51removes the information that is no
  15597. 9:18:52longer useful in the cell state. Then we
  15598. 9:18:55have input gate. Additional information
  15599. 9:18:57to the cell state is added by input
  15600. 9:18:59gate. And finally, we have something
  15601. 9:19:01called as output gate. Additional useful
  15602. 9:19:03information to the cell state is also
  15603. 9:19:05added by an output gate. This gating
  15604. 9:19:07mechanism of LSTM has allowed network to
  15605. 9:19:10learn the conditions for when to forget,
  15606. 9:19:12ignore, or keep information in the
  15607. 9:19:14memory cell.
  15608. 9:19:15So let me now quickly move to my Jupiter
  15609. 9:19:17notebook and show you how I can
  15610. 9:19:19implement LSTM on name entity
  15611. 9:19:21prediction. All right, so let me quickly
  15612. 9:19:23move there. All right, so over here
  15613. 9:19:25first off, I'll be opening my Google
  15614. 9:19:28Colab.
  15615. 9:19:32Okay, so let us give a name for our
  15616. 9:19:34Google Colab over here.
  15617. 9:19:36Let's give a short term, right? Name
  15618. 9:19:38entity prediction. And let's connect our
  15619. 9:19:40Google Colab to our server.
  15620. 9:19:42Okay, meanwhile that's connecting. So,
  15621. 9:19:44now you might be wondering from where am
  15622. 9:19:46I going to use my data set? So, for me
  15623. 9:19:48to use my data set, I'll just go for
  15624. 9:19:49Kaggle, k a g g l e
  15625. 9:19:52baby names. So, let me just quickly show
  15626. 9:19:55you how this data set would look like.
  15627. 9:19:57So, this is a CSV file over here. All
  15628. 9:19:59right, so as you can see here, we have
  15629. 9:20:01over 93,889
  15630. 9:20:03unique values. Okay, so this is a very
  15631. 9:20:06huge data set. And let's try downloading
  15632. 9:20:09this. To download this is pretty simple.
  15633. 9:20:11All you need to do is click this and it
  15634. 9:20:13will get downloaded. As I've already
  15635. 9:20:15downloaded this file, let me quickly
  15636. 9:20:17upload this on my Jupyter notebook. So,
  15637. 9:20:19let me go here and upload it from here.
  15638. 9:20:23Okay, so let me go to this upload file.
  15639. 9:20:26And yeah, so I have my CSV file here and
  15640. 9:20:29let me open this. As this is a pretty
  15641. 9:20:31huge data set, it will take some time.
  15642. 9:20:32Meanwhile that's loading, let's see what
  15643. 9:20:34we can do.
  15644. 9:20:36So, first off let's import couple of
  15645. 9:20:37libraries. So, we'll have import pandas
  15646. 9:20:42as pd.
  15647. 9:20:44And then we're going to import
  15648. 9:20:46NumPy as np. And then we also need to
  15649. 9:20:50have matplotlib. So, from sklearn
  15650. 9:20:53All right, and we also need something
  15651. 9:20:55like label encoder, but I'll show you a
  15652. 9:20:57shortcut way to you know bypass label
  15653. 9:20:59encoding. Okay, so let's try to load our
  15654. 9:21:02cell here. And in order for us to read
  15655. 9:21:04this data, so it's pretty simple. All
  15656. 9:21:06we're going to do is let's give this as
  15657. 9:21:08a data. This would be nothing but
  15658. 9:21:10pandas.read_csv
  15659. 9:21:12and then we're going to pass our file
  15660. 9:21:15name. Let me change this to our root
  15661. 9:21:16directory
  15662. 9:21:18by putting a dot over here. Okay, so I
  15663. 9:21:20won't be executing this as of now
  15664. 9:21:21because it's trying to load our file.
  15665. 9:21:26All right, so now that we have
  15666. 9:21:27successfully loaded our data so, let's
  15667. 9:21:30try running this cell over here. Okay,
  15668. 9:21:32so let me close this and let me zoom in
  15669. 9:21:35over here.
  15670. 9:21:36So, now what we're going to do is let's
  15671. 9:21:38see the shape of our data. So, let's see
  15672. 9:21:40what's the data shape. data.shape
  15673. 9:21:43and now let's see what it would be like.
  15674. 9:21:45Okay, so as you can see here, we have
  15675. 9:21:47five columns. But, the number of rows
  15676. 9:21:50that we have is 1.8 million. That is
  15677. 9:21:52approximately 18 lakhs, right? So, this
  15678. 9:21:55is a pretty huge value. So, now what
  15679. 9:21:57we're going to do is we'll just see how
  15680. 9:21:59our data is looking like. So, we'll see
  15681. 9:22:01data.head.
  15682. 9:22:03And let's see what we need. So, as you
  15683. 9:22:05can see here, we have ID, which is of no
  15684. 9:22:08use for us. Then we have name. Okay,
  15685. 9:22:10then this year, I don't think it's of
  15686. 9:22:12any use for us. Then we have gender and
  15687. 9:22:14count. Count here represents, you know,
  15688. 9:22:17how many people have the name Mary, how
  15689. 9:22:19many people have the name Anna, how many
  15690. 9:22:21people have the name Emma, Elizabeth,
  15691. 9:22:23and Minnie. This is over here, out of
  15692. 9:22:25this if you see, right? There are a
  15693. 9:22:27couple of things that we don't need. We
  15694. 9:22:28can drop them out. You know, all we need
  15695. 9:22:30is a name. Okay, and then we also need
  15696. 9:22:32the gender. Because this is going to be
  15697. 9:22:35our prediction. We're going to predict a
  15698. 9:22:36we'll give our own custom name and then
  15699. 9:22:38we'll see whether the name that is
  15700. 9:22:40you're giving is male or a female. Okay?
  15701. 9:22:44So, now what we're going to do is let's
  15702. 9:22:45see how many unique values we have. So,
  15703. 9:22:47let me quickly erase this first. Okay,
  15704. 9:22:50so what I'm going to do is data.names.
  15705. 9:22:53So, this should give us here name. And
  15706. 9:22:56then we'll type here as unique.
  15707. 9:22:58Okay, so this should give me unique
  15708. 9:23:00values. Okay, so over here I have 93,889
  15709. 9:23:04unique names. Okay, so now what we're
  15710. 9:23:07going to do is we want to label encode
  15711. 9:23:09this, right? So, we want our female, uh
  15712. 9:23:11which is nothing but F, we want female
  15713. 9:23:13to be zero and then male to be one or
  15714. 9:23:15vice versa. So, in order to do that,
  15715. 9:23:17either we can use label encoder or
  15716. 9:23:20there's a shortcut method to this. Let
  15717. 9:23:21me quickly show you how that works. So,
  15718. 9:23:23first of all, we'll take our data frame,
  15719. 9:23:25so it's data. And which column do you
  15720. 9:23:27want to do this for? We want to do this
  15721. 9:23:29for our gender column, right? So, let me
  15722. 9:23:32pass this and give gender. And now what
  15723. 9:23:35we're going to do is
  15724. 9:23:37Okay?
  15725. 9:23:38We'll take this as as type.
  15726. 9:23:41Okay, this would be obviously in the
  15727. 9:23:42form of category.
  15728. 9:23:44And now what we'll do is this is cat
  15729. 9:23:47dot codes.
  15730. 9:23:50Okay, so this is nothing but panda
  15731. 9:23:51shortcut, you know, to label encoding.
  15732. 9:23:53Let's try to execute this and see what
  15733. 9:23:55it would look like. So, as you can see
  15734. 9:23:57here, we have couple of zeros and, you
  15735. 9:23:59know, ones. This is nothing but it's
  15736. 9:24:00representing females with one and males
  15737. 9:24:03with zeros. Okay? So, now what we'll do
  15738. 9:24:05is we have to update this column.
  15739. 9:24:09So, we'll paste this. And this should be
  15740. 9:24:11something like this over here. And let
  15741. 9:24:13me execute this. Okay, so if you want to
  15742. 9:24:16see how our data would look like now,
  15743. 9:24:18let me just quickly run this once again.
  15744. 9:24:20So, you'll see here now the values has
  15745. 9:24:22been label encoded. Okay? So, now what
  15746. 9:24:25we're going to do is we obviously need
  15747. 9:24:27to take the unique names, right? And
  15748. 9:24:30then we'll obviously group it by, right?
  15749. 9:24:31So, what we'll do for this is we'll take
  15750. 9:24:33something like data. We'll group this by
  15751. 9:24:37the names. So, group by
  15752. 9:24:39names.
  15753. 9:24:40All right? And now what we'll do is
  15754. 9:24:42we'll calculate the mean
  15755. 9:24:44of the genders.
  15756. 9:24:46We'll reset the index. The reason why we
  15757. 9:24:47want to reset the index is because, you
  15758. 9:24:49know, if you don't give the index then
  15759. 9:24:50our name over here will become the
  15760. 9:24:52index, right? So, we'll give reset
  15761. 9:24:54{underscore} index.
  15762. 9:24:56All right? So, let's give this to a new
  15763. 9:24:59data frame and we'll call this as DF.
  15764. 9:25:02Okay, let me execute this now.
  15765. 9:25:04And let's see how this DF would look
  15766. 9:25:05like. Okay, let me execute this right
  15767. 9:25:08after this.
  15768. 9:25:09So, as you can see here, it has grouped
  15769. 9:25:11by by names, all everything in an
  15770. 9:25:13ascending order. So, if this is all in
  15771. 9:25:15an alphabetical manner. And yeah.
  15772. 9:25:19And now only thing that I want to work
  15773. 9:25:21on is this gender.
  15774. 9:25:22Okay? The reason is because over here
  15775. 9:25:24I'm getting a floating point value. I
  15776. 9:25:26don't want this floating point value. I
  15777. 9:25:28want to change this to integer value,
  15778. 9:25:30right? So, what I'll do is
  15779. 9:25:32DF gender
  15780. 9:25:34This would be nothing but
  15781. 9:25:36DF gender. Then I'll all I'm going to do
  15782. 9:25:39is as type.
  15783. 9:25:40I'll just put here as int. So, let's now
  15784. 9:25:42see what this value would look like.
  15785. 9:25:45Fantastic. We over here have now, you
  15786. 9:25:47know, ones and zeros, which is nothing
  15787. 9:25:48but an integer value. Okay? So, if you
  15788. 9:25:51want to see this, so I either I can
  15789. 9:25:53write DF or I can also put as head.
  15790. 9:25:56Okay, so these are the first five
  15791. 9:25:57values.
  15792. 9:25:58Okay. So, now the way our neural network
  15793. 9:26:01is going to work or the recurrent neural
  15794. 9:26:02network is going to work is that, you
  15795. 9:26:04know, I hope you remember these boxes,
  15796. 9:26:06right? So, when I was talking or when I
  15797. 9:26:09was explaining this RNN, I was saying
  15798. 9:26:11that I would be passing around the
  15799. 9:26:13words. But here in this project or in
  15800. 9:26:16this program, we won't be passing words
  15801. 9:26:18over here. You know, we won't be passing
  15802. 9:26:20like Abba or Abida or Adam. We won't be
  15803. 9:26:23passing these words. Instead of that,
  15804. 9:26:25we'll be passing letters.
  15805. 9:26:27So, over here it's going to be like
  15806. 9:26:29alphabets. So, A, B. It can be any
  15807. 9:26:32alphabet. It can be Z here. So,
  15808. 9:26:34basically it depends upon whatever the
  15809. 9:26:36value is coming here. So, in order to do
  15810. 9:26:38that, we have to find number of unique
  15811. 9:26:40alphabets. So, we know how many unique
  15812. 9:26:41alphabets we have, right? So, we it's
  15813. 9:26:4326. So, in order to get these alphabets,
  15814. 9:26:46what we'll do is let me first quickly
  15815. 9:26:47erase this.
  15816. 9:26:49Erase all drawing.
  15817. 9:26:50So, now we have 26 alphabets. We have to
  15818. 9:26:53create our own vocabulary. So, what I'm
  15819. 9:26:54going to do is I'm going to import
  15820. 9:26:56string. So, now I need letters, right?
  15821. 9:26:59So, l e t t e r s. This would be nothing
  15822. 9:27:01but list of string
  15823. 9:27:05.ascii.
  15824. 9:27:06Okay? And if you want to see what this
  15825. 9:27:07would give me, this would be nothing but
  15826. 9:27:10the list of alphabets, which are in
  15827. 9:27:11lower cases.
  15828. 9:27:13Okay?
  15829. 9:27:14And now what we'll do is we'll try to
  15830. 9:27:15create a label encoding or we have to
  15831. 9:27:18create a vocabulary, right? So, we'll
  15832. 9:27:19have something like vocab. This would be
  15833. 9:27:21nothing but I'll be using dictionary.
  15834. 9:27:24And now what I want is zip. The way I
  15835. 9:27:26want over here is, you know, for every
  15836. 9:27:28individual values of this A B C D, I
  15837. 9:27:31want to label encode this to 0 1 and
  15838. 9:27:35whatever the value it is, right? So, it
  15839. 9:27:36would be from 1 to 27.
  15840. 9:27:38A unique numbers, right? So, this would
  15841. 9:27:40be nothing but letters. And then uh
  15842. 9:27:42we'll be need something like uh range
  15843. 9:27:451 {comma} 27. So, this would give me the
  15844. 9:27:48matrix from 1 to 26, right? And let's
  15845. 9:27:50now see what this would look like. So,
  15846. 9:27:52we have vocab.
  15847. 9:27:54And let me execute this. This should
  15848. 9:27:56give me a dictionary, okay? So, here
  15849. 9:27:58we'll convert A to 1.
  15850. 9:28:00Okay? And then B would be 2, C would be
  15851. 9:28:033, and so on, Z would be 26.
  15852. 9:28:06And now what we're going to do is uh
  15853. 9:28:08we'll just try to create the reverse
  15854. 9:28:10vocabulary. And the reason is we
  15855. 9:28:12obviously won't be needing this, but uh
  15856. 9:28:14you know, just in case you want to use
  15857. 9:28:16it would be something very similar to
  15858. 9:28:17this. Let me just copy the exact same
  15859. 9:28:19thing.
  15860. 9:28:20And paste it over here.
  15861. 9:28:22So, we'll just do it as reverse, right?
  15862. 9:28:24So, it will be R {underscore}
  15863. 9:28:26R {underscore} And here, instead of
  15864. 9:28:28numbers being second,
  15865. 9:28:30we'll just cut this letters and we'll
  15866. 9:28:32pass letters over here.
  15867. 9:28:34And now you'll see if you're trying to
  15868. 9:28:36decode whatever we have predicted, you
  15869. 9:28:37know, we can just pass it down like
  15870. 9:28:39this.
  15871. 9:28:40Okay. So, now what what will happen is
  15872. 9:28:42we need to do something like, you know,
  15873. 9:28:44all our data, whatever is there, we have
  15874. 9:28:45to convert them into a lowercase.
  15875. 9:28:48So, and then once we convert them into a
  15876. 9:28:50lowercase, we have to encode them into a
  15877. 9:28:52numbers.
  15878. 9:28:54So, whatever I'm saying is this A A B A
  15879. 9:28:57N, right? A ban. So, we this A A
  15880. 9:28:59obviously first of we have to convert
  15881. 9:29:00all of these into a lowercase,
  15882. 9:29:02you know, this value. And then whatever
  15883. 9:29:04the equivalent value of A, the numerical
  15884. 9:29:07value of A, so it's obviously going to
  15885. 9:29:09be one. We'll substitute that with this.
  15886. 9:29:11And it's going to be a list, right? So,
  15887. 9:29:12how do I do that? So, for that I'll
  15888. 9:29:14write a function.
  15889. 9:29:16So, we'll have DEF word to number,
  15890. 9:29:18right? Word to
  15891. 9:29:21So, now what I'm going to do is I'm
  15892. 9:29:22going to have for loop for I in range.
  15893. 9:29:26So, this would be nothing but
  15894. 9:29:29we have to go through the entire shape,
  15895. 9:29:31right? So, d f dot shape.
  15896. 9:29:33This should give me a list and I just
  15897. 9:29:35need the first index.
  15898. 9:29:37Okay? So, now what I'm going to do is
  15899. 9:29:39I'll create one new list sequence.
  15900. 9:29:42This would be nothing but for letters in
  15901. 9:29:46d f. Obviously, we want the names part.
  15902. 9:29:50And in this we're going to pass the
  15903. 9:29:51index value. It's going to be I.
  15904. 9:29:53Let me just give some space here just so
  15905. 9:29:55that you better understand this.
  15906. 9:29:57And now what I'm going to do is, you
  15907. 9:29:59know, I'll have this vocabulary.
  15908. 9:30:01vocab See, every time I pass a letter
  15909. 9:30:03it'll convert it into, you know, this
  15910. 9:30:05individual letter it'll convert it into
  15911. 9:30:07a list all the equivalent, you know,
  15912. 9:30:09numerical representation. It'll be
  15913. 9:30:11letters and obviously it has to be in
  15914. 9:30:13lower so it'll be lower.
  15915. 9:30:15And then we'll just close this bracket
  15916. 9:30:16here.
  15917. 9:30:17So, now what we're going to do is before
  15918. 9:30:19we execute this function, we'll have to
  15919. 9:30:22append this so it'll be d f.
  15920. 9:30:24And this is going to be names
  15921. 9:30:27dot I.
  15922. 9:30:28We'll replace the name in that index
  15923. 9:30:30with this particular sequence.
  15924. 9:30:32Okay? So, now all we need to do is run
  15925. 9:30:34this function over here.
  15926. 9:30:36And yeah.
  15927. 9:30:38This will take some time. The reason is
  15928. 9:30:39because we have almost around 18 lakh
  15929. 9:30:42values. So, yeah, this should take some
  15930. 9:30:44time. Meanwhile, let me just comment
  15931. 9:30:46this.
  15932. 9:30:53Okay? So, in the next stage what we're
  15933. 9:30:54going to do is let's see how our this
  15934. 9:30:57value over here would look like. So, let
  15935. 9:31:00us now first execute this.
  15936. 9:31:03All right. So, let us now see how our
  15937. 9:31:04data frame will look like. So, let me
  15938. 9:31:06execute this block now.
  15939. 9:31:09So, as you can see here, our names have
  15940. 9:31:11been completely changed or converted
  15941. 9:31:13into list of numbers. But now, only
  15942. 9:31:15issue that we are trying to have is the
  15943. 9:31:18imbalance in the size of the list.
  15944. 9:31:20Because when we are trying to have the
  15945. 9:31:22number of boxes, right? We won't be
  15946. 9:31:24having variable number of boxes. Okay?
  15947. 9:31:27So, what we're going to do is either we
  15948. 9:31:28set a value like something like take an
  15949. 9:31:30average number like 10, 20, or you can
  15950. 9:31:33take something like, you know, something
  15951. 9:31:35like you take you either depend on
  15952. 9:31:36maximum number or the minimum number of
  15953. 9:31:38list. But the thing is, if you take the
  15954. 9:31:41maximum number, then we have to pad a
  15955. 9:31:42lot of zeros, and this would lead to a
  15956. 9:31:44loss.
  15957. 9:31:45So, if I reduce the size, this would
  15958. 9:31:46also decrease the accuracy, right? So,
  15959. 9:31:48what we're going to do is we'll plot
  15960. 9:31:49this name and gender in the form of a
  15961. 9:31:52histogram. So, let's take here X. This
  15962. 9:31:55would be DF names.
  15963. 9:31:58And we'll give here as dot values.
  15964. 9:32:00And then same thing we'll do it for Y,
  15965. 9:32:02DF gender.
  15966. 9:32:04And this should be dot values.
  15967. 9:32:06So, what we'll do is we also need a
  15968. 9:32:09list, okay? So, now as we are going to
  15969. 9:32:10plot this on a histogram, and what we're
  15970. 9:32:12going to see in the histogram is just to
  15971. 9:32:15analyze, you know, this this is a graph.
  15972. 9:32:17We want to analyze, you know, where does
  15973. 9:32:19the highest number of sequence, or if
  15974. 9:32:22suppose this is a size, if this is size
  15975. 9:32:23eight, and this is like 8,000 words or
  15976. 9:32:268,000 names have the size eight, then
  15977. 9:32:28you know, we can keep our average
  15978. 9:32:30somewhere near, and then we can also
  15979. 9:32:32decide, you know, if if the number after
  15980. 9:32:3410, if not many names have a longer
  15981. 9:32:36number or the longer length of that
  15982. 9:32:38name, you know, so we can keep our
  15983. 9:32:40average somewhere around nine or 10,
  15984. 9:32:42okay? So, let's now quickly see how we
  15985. 9:32:44can do that.
  15986. 9:32:46To get the length of our names, so
  15987. 9:32:47length X or name length.
  15988. 9:32:51This would be like list comprehension
  15989. 9:32:53for I in range 0, DF.shape
  15990. 9:32:58of 0.
  15991. 9:33:00And now what we're going to do is we
  15992. 9:33:01need to find the length. So, this would
  15993. 9:33:03be length X of I.
  15994. 9:33:06Or here, you can either give BF or you
  15995. 9:33:08can also give this X, right? So, it
  15996. 9:33:10would be length of X.
  15997. 9:33:11Okay, so let me quickly execute this.
  15998. 9:33:14Okay, so here we're getting an error.
  15999. 9:33:15Oh, yeah. It's not O, it's going to be
  16000. 9:33:17zero, right? So, let me execute this
  16001. 9:33:19now.
  16002. 9:33:19Okay, so let me show you how this would
  16003. 9:33:21look like. Name length, and let me print
  16004. 9:33:24this off.
  16005. 9:33:25So, as you can see, this is giving me
  16006. 9:33:26list of names.
  16007. 9:33:27So, there are huge amount of names. So,
  16008. 9:33:29as you can see, first we had five, five,
  16009. 9:33:31and then nine. So, let's now plot this
  16010. 9:33:34and see how it would look like. Import
  16011. 9:33:37Matplotlib as plt.
  16012. 9:33:39All right, this is perfect. So, now what
  16013. 9:33:40I'm going to do is I have to plot this,
  16014. 9:33:42right? So, all I'm going to do is
  16015. 9:33:44plt.hist.
  16016. 9:33:45All right, and now I'm just going to
  16017. 9:33:47give name length.
  16018. 9:33:49And number of bins, this would be like
  16019. 9:33:51let's give 20, okay? And then plt.show.
  16020. 9:33:54Okay, so what do we find from this graph
  16021. 9:33:57over here? You see, this is nothing but
  16022. 9:33:58the length of the names. So, two, four,
  16023. 9:34:01six, eight, 12, all these are length of
  16024. 9:34:03the names. And this is nothing but zero,
  16025. 9:34:055,000, 10,000, this is nothing but
  16026. 9:34:07number of names that have a length four
  16027. 9:34:09or number of names that have the length
  16028. 9:34:11six. So, as you can see, right? The
  16029. 9:34:13there are around almost 25,000 names
  16030. 9:34:15whose length is six. And then as I cross
  16031. 9:34:18like 10 or as I cross 12, not many names
  16032. 9:34:21are there whose length is greater than,
  16033. 9:34:23you know, 12. So, what I'm going to do
  16034. 9:34:25here now is now we have to pad, right?
  16035. 9:34:27Now, we have to pad number of zeros. So,
  16036. 9:34:29in order to pad zeros, we have something
  16037. 9:34:31called as built-in function from Keras.
  16038. 9:34:33So, from Keras or you can also set as
  16039. 9:34:35Keras.preprocessing
  16040. 9:34:38.sequence import pad_sequence.
  16041. 9:34:42So, let me quickly execute this now.
  16042. 9:34:44And now what I'm going to do is I'm
  16043. 9:34:45going to create a new list.
  16044. 9:34:47So, let this be X. This is in lowercase.
  16045. 9:34:50So, pad_sequence. And the things that it
  16046. 9:34:52this is going to take is obviously the
  16047. 9:34:54sequence. We have to give a list of
  16048. 9:34:55sequence.
  16049. 9:34:57Okay, and then we'll give something like
  16050. 9:34:59df.names.
  16051. 9:35:01And then this is going to be values.
  16052. 9:35:04And now we want to define the max
  16053. 9:35:06length. So, this is going to be 10. We
  16054. 9:35:08also have an option of providing where
  16055. 9:35:10do you want to do the padding? So, we
  16056. 9:35:11can also do it as pre or post. We'll
  16057. 9:35:14obviously be doing pre. So, let's see
  16058. 9:35:17how do we do that. This would be nothing
  16059. 9:35:19but, you know, if you can see this
  16060. 9:35:20sequence over here. So, we have padding
  16061. 9:35:22is equal to pre. It's so it's by
  16062. 9:35:24default, right? So, pre. So, let me
  16063. 9:35:26execute this now. And let's see how this
  16064. 9:35:28X would look like.
  16065. 9:35:30Okay. So, as you can see here, X is a
  16066. 9:35:31matrix whose length is 10. Okay? So,
  16067. 9:35:34each of these like this is this is
  16068. 9:35:36nothing but, you know, 19 million cross
  16069. 9:35:3810. So, there are 10 columns throughout
  16070. 9:35:40all.
  16071. 9:35:41So, now what we're going to do is we're
  16072. 9:35:42going to create our own model. So, for
  16073. 9:35:46that we'll do from keras.layers
  16074. 9:35:49import
  16075. 9:35:51input layer.
  16076. 9:35:52And then we have to have embedding layer
  16077. 9:35:54cuz if you don't have embedding layer,
  16078. 9:35:56then you know, it it would be like each
  16079. 9:35:57input would be something like 1 comma or
  16080. 9:36:001.9 million. That is 18 lakhs. So, it's
  16081. 9:36:02a pretty huge value to compute. So, we
  16082. 9:36:04don't want that. So, that's why we'll
  16083. 9:36:05use embedding layer. Then we have dense
  16084. 9:36:06layer and then we have LSTM.
  16085. 9:36:09We also, you know, rather than taking
  16086. 9:36:11this as a sequential model, we'll take
  16087. 9:36:12it as, you know, feed forward. So, what
  16088. 9:36:15we'll do is from keras.models
  16089. 9:36:19import model.
  16090. 9:36:21So, now what we're going to do is we'll
  16091. 9:36:23have to create our input layer. So, this
  16092. 9:36:25would be input is equal to input.
  16093. 9:36:28And now the shape that we're going to
  16094. 9:36:29pass over here for for this shape
  16095. 9:36:33So, how many columns do we have? We
  16096. 9:36:34obviously have 10 columns, right? So,
  16097. 9:36:36it's going to be 10.
  16098. 9:36:37So, now what we're going to do is next
  16099. 9:36:38we're going to have embedding layer. So,
  16100. 9:36:40let's say this is EMB and this would be
  16101. 9:36:43embedding.
  16102. 9:36:44So, input over here
  16103. 9:36:47or the input dimension over here is
  16104. 9:36:48nothing but vocab size.
  16105. 9:36:50We haven't defined this vocab size, so
  16106. 9:36:52let's quickly do that. So, vocab
  16107. 9:36:55size this would be nothing but length
  16108. 9:36:59of vocabularies
  16109. 9:37:01plus one. The reason why I'm doing plus
  16110. 9:37:03one is because we also have zeros over
  16111. 9:37:05here, right?
  16112. 9:37:06And this would be like vocab size if you
  16113. 9:37:08want to see.
  16114. 9:37:09And let me execute this.
  16115. 9:37:11So, we have 27, right? So, 26 are the
  16116. 9:37:13number of alphabets and one is because
  16117. 9:37:15we have number of zeros. So, we'll pass
  16118. 9:37:17this as vocab size.
  16119. 9:37:19Okay? And now we're also going to pass
  16120. 9:37:21output dimension. So, output dimension,
  16121. 9:37:23this is nothing but, you know, how many
  16122. 9:37:24dimensions we want. So, now this is
  16123. 9:37:26going to be five. All right? And the
  16124. 9:37:28input for this embedded layer is going
  16125. 9:37:29to be from INT.
  16126. 9:37:31Okay? And now we are going to have our
  16127. 9:37:33first LSTM layer. So, it's going to be
  16128. 9:37:34LSTM
  16129. 9:37:36one. So, this would be LSTM layer.
  16130. 9:37:40Number of units we have to define here.
  16131. 9:37:42So, units, this is going to be like 32.
  16132. 9:37:45The units over here does not represent
  16133. 9:37:46the number of boxes.
  16134. 9:37:48The units over here represent the A
  16135. 9:37:49values, right? So, now we have to do
  16136. 9:37:52return sequence and this is going to be
  16137. 9:37:54true.
  16138. 9:37:55And the input for this is going to be
  16139. 9:37:56from embedded layer.
  16140. 9:37:57Then we have LSTM second layer.
  16141. 9:38:00And this is be LSTM units we're going to
  16142. 9:38:03pass. So, number of units that we're
  16143. 9:38:05going to pass now is 64.
  16144. 9:38:07And now the input for this is going to
  16145. 9:38:09be LSTM one. Finally, we have an output
  16146. 9:38:12layer.
  16147. 9:38:13So, at the end, right? We're going to
  16148. 9:38:14have a dense layer, right? So, dense.
  16149. 9:38:17So, the number of units or number of
  16150. 9:38:18neurons at the end we are going to have
  16151. 9:38:19one.
  16152. 9:38:20And kind of activation function that I'm
  16153. 9:38:22going to have here is going to be
  16154. 9:38:23sigmoid because we have to predict
  16155. 9:38:25either it's a male or a female. Okay?
  16156. 9:38:27So, sigmoid.
  16157. 9:38:29And the input for this is going to be
  16158. 9:38:31LSTM two.
  16159. 9:38:33So, finally, we have to add this to our
  16160. 9:38:35model. So, I'll do it as my model.
  16161. 9:38:38This would be model.
  16162. 9:38:40So, now I have to define the inputs. So,
  16163. 9:38:42I N P U T S, this is going to be inputs
  16164. 9:38:44I N P
  16165. 9:38:46and outputs.
  16166. 9:38:48This is going to be out.
  16167. 9:38:49Okay? So, let me quickly execute this
  16168. 9:38:51now.
  16169. 9:38:52Okay, so we have this error.
  16170. 9:38:54Please provide either a shape. Okay.
  16171. 9:38:57Let's see what's Oh, yeah. Oh, here I've
  16172. 9:38:59given it as pass, right? So, it's not
  16173. 9:39:00going to be this pass. It's going to be
  16174. 9:39:01shape.
  16175. 9:39:02So, let me quickly execute this once
  16176. 9:39:04again.
  16177. 9:39:05Okay, as you can see, we have
  16178. 9:39:06successfully executed this. And now
  16179. 9:39:08let's see the model. summary.
  16180. 9:39:10So, as you can see here, first off, we
  16181. 9:39:12have input layer.
  16182. 9:39:13Okay. So, we can have a number of
  16183. 9:39:15values, but there will be only 10
  16184. 9:39:17features. Okay, that's 10 columns.
  16185. 9:39:19And then we're going to have once you go
  16186. 9:39:21through this embedding layer, then you
  16187. 9:39:23know, instead of having 10
  16188. 9:39:25you know, instead of having that 1
  16189. 9:39:27million or whatever it is, 1.8 million,
  16190. 9:39:29we'll have just 135 parameters. Okay.
  16191. 9:39:32Similarly, over here and finally at the
  16192. 9:39:33dense, we have 65 parameters and then we
  16193. 9:39:35have one.
  16194. 9:39:36The reason why we have 65 here, just for
  16195. 9:39:39if you don't know, is because 64 + 1
  16196. 9:39:41bias. And this will give us 65.
  16197. 9:39:43Okay. So, now finally, we're going to
  16198. 9:39:45train our model. But before that, we
  16199. 9:39:47have to compile it. It's going to be my
  16200. 9:39:48model.
  16201. 9:39:49model.compile
  16202. 9:39:51Okay. So, we have to find an optimizer.
  16203. 9:39:53So, optimizer, best one that I feel is
  16204. 9:39:56Adam.
  16205. 9:39:57Then we have to find the loss.
  16206. 9:39:59So, as we're going to use just two
  16207. 9:40:01predictions, right? It's either it's
  16208. 9:40:02male or female, we'll use binary
  16209. 9:40:04cross-entropy.
  16210. 9:40:06And then finally, the matrix that we
  16211. 9:40:07want to use here is
  16212. 9:40:10This will be accuracy.
  16213. 9:40:12So, let me execute this now.
  16214. 9:40:14And finally, we are going to compile
  16215. 9:40:15this. So, we have history.
  16216. 9:40:17So, this will be model.fit.
  16217. 9:40:20Okay. And now we're going to pass our
  16218. 9:40:22values. We're going to give X. We're
  16219. 9:40:23going to give Y.
  16220. 9:40:25As you know, X is nothing but a matrix
  16221. 9:40:26which has a padding. And Y is nothing
  16222. 9:40:28but you know, the the classes. They're
  16223. 9:40:29telling us either it's ones or zeros.
  16224. 9:40:31And number of epochs
  16225. 9:40:34is going to be 10.
  16226. 9:40:35Batch size, as this is a pretty used
  16227. 9:40:37data set, we are going to keep a pretty
  16228. 9:40:38high batch size.
  16229. 9:40:40So, I'll give a batch size here as 256.
  16230. 9:40:43And then finally, we need validation
  16231. 9:40:45split.
  16232. 9:40:46Okay, it's not going to be my model,
  16233. 9:40:47it's going to be my model, right? So, my
  16234. 9:40:49underscore model. So, finally we're
  16235. 9:40:52going to have validation split here.
  16236. 9:40:54So, let's give it as 20%. So, it's going
  16237. 9:40:56to be 0.2.
  16238. 9:40:58All right. So, finally it's the moment
  16239. 9:41:00of truth. Let us now execute our code.
  16240. 9:41:03This will take some time to execute.
  16241. 9:41:06So, let's see how this would look like.
  16242. 9:41:11Okay, so if you can analyze this data
  16243. 9:41:13over here,
  16244. 9:41:15so as you can see, right? This
  16245. 9:41:16validation accuracy has to increase.
  16246. 9:41:19And we cannot see see every time, you
  16247. 9:41:21know, if our model is over fitting, the
  16248. 9:41:23accuracy over here will keep on
  16249. 9:41:24increasing. All right. So, the more
  16250. 9:41:27reliable source over here to see is
  16251. 9:41:29nothing but validation accuracy. So, if
  16252. 9:41:31validation accuracy is increasing, that
  16253. 9:41:33means our model is neither over fitting
  16254. 9:41:34or under fitting. And you can also see
  16255. 9:41:36that we have our validation loss, which
  16256. 9:41:38is kind of decreasing. And over here as
  16257. 9:41:40well, we can see the validation loss
  16258. 9:41:41over here. We can see the loss of our
  16259. 9:41:43model is decreasing from 60 then 40 then
  16260. 9:41:4639 39 and 80 38. So, let's now wait for
  16261. 9:41:49a few more epochs. So, we have four more
  16262. 9:41:52to go.
  16263. 9:41:55Okay, so as you can see here, you know,
  16264. 9:41:57our validation accuracy has been
  16265. 9:41:59increasing. So, this is a very healthy
  16266. 9:42:01growth.
  16267. 9:42:02And even over here, our accuracy of our
  16268. 9:42:04model is also increasing. Fine? And the
  16269. 9:42:06loss is decreasing. It's decreased from
  16270. 9:42:0860% to 37% and over here, our validation
  16271. 9:42:11loss decreased from 42% to 36%.
  16272. 9:42:14So, let us now map this like whatever
  16273. 9:42:16values we have received. So, in order to
  16274. 9:42:18map this, we have H.
  16275. 9:42:20You know, this model over here retrieves
  16276. 9:42:21us the history function, right? So,
  16277. 9:42:23we'll give hist
  16278. 9:42:24dot history.
  16279. 9:42:25And let me execute this.
  16280. 9:42:27Okay. So, now what we're going to do is
  16281. 9:42:29this is nothing but key value pairs. So,
  16282. 9:42:31if I put H,
  16283. 9:42:32you know, if I give something like
  16284. 9:42:33accuracy, okay, we let's plot this.
  16285. 9:42:36So, model dot plot, right? So, plt dot
  16286. 9:42:40plot.
  16287. 9:42:41Okay.
  16288. 9:42:42We'll have accuracy. We'll compare this
  16289. 9:42:44with respect to accuracy and then
  16290. 9:42:46plt.plot
  16291. 9:42:49then we'll have validation accuracy. And
  16292. 9:42:51we want to show this, right? So,
  16293. 9:42:52plt.show.
  16294. 9:42:53So, as you can see here, okay, so just
  16295. 9:42:55to give a better analogy, let's let's
  16296. 9:42:57execute one of these first.
  16297. 9:42:59Okay, so the blue line over here
  16298. 9:43:00represents the accuracy of our model and
  16299. 9:43:02then this is nothing but the accuracy of
  16300. 9:43:04our training data, right? Or testing
  16301. 9:43:06data. So, as you can see, our model
  16302. 9:43:07accuracy isn't decreasing. So, it's a
  16303. 9:43:09very good model and it has trained very
  16304. 9:43:11well. So, now coming down to the moment
  16305. 9:43:13of truth, so let's now, you know, take a
  16306. 9:43:15random name and see whether it can
  16307. 9:43:17predict whether the name is true or
  16308. 9:43:18false. Okay, so we'll give here as
  16309. 9:43:20test_name.
  16310. 9:43:21So, let's give something like, you know,
  16311. 9:43:24we'll this will be like name, right? So,
  16312. 9:43:25we'll give name.lower.
  16313. 9:43:28And now we're going to pass the name
  16314. 9:43:29over here.
  16315. 9:43:31So, this would be, let's say, Tom.
  16316. 9:43:34Okay, so we have to convert this into
  16317. 9:43:35letters, right? So, it'll be vocab of I
  16318. 9:43:39for
  16319. 9:43:40I in test name.
  16320. 9:43:42And now we'll give the name here as
  16321. 9:43:44X_test.
  16322. 9:43:46So, this would be nothing but we have to
  16323. 9:43:47pad the sequence, pad sequence. We'll
  16324. 9:43:50pass this in a form of a tuple, then
  16325. 9:43:51this would be seq.
  16326. 9:43:54And then we know that we have to pad 10,
  16327. 9:43:55right? And anyways, we don't have to say
  16328. 9:43:57whether it's pre or post because by
  16329. 9:43:59default it's going to be pre.
  16330. 9:44:01Fine. So, let us now see how our text
  16331. 9:44:03data would look like. So, X_test.
  16332. 9:44:06String attribute has Okay.
  16333. 9:44:08Oh, yeah. So, I have done a typo over
  16334. 9:44:10here. It's going to be l o w e r.
  16335. 9:44:13Let me execute this once again.
  16336. 9:44:15So, as you can see here, we have a
  16337. 9:44:17matrix which has a size of 10. And let's
  16338. 9:44:19now see what this would predict. So,
  16339. 9:44:22it's the moment of truth. So, y.predict.
  16340. 9:44:25This would be model.predict.
  16341. 9:44:28And I'm going to pass here as X_test.
  16342. 9:44:30And let's see what does this predict.
  16343. 9:44:32So, we'll give here as y_pred. And yeah,
  16344. 9:44:35let us execute this.
  16345. 9:44:37Okay, so we are getting this in a form
  16346. 9:44:38of a array. So, what this tells us, you
  16347. 9:44:41know, this tells us that, you know, this
  16348. 9:44:43is like 70% chances that this name is
  16349. 9:44:46Tom. In order to make this, you know,
  16350. 9:44:48layman's stuff, so what we're going to
  16351. 9:44:49do is we'll have if y pred Let me
  16352. 9:44:52execute this first.
  16353. 9:44:54So, let me go down to another block.
  16354. 9:44:56So, this would be something like if
  16355. 9:44:59y pred is less than 0.5, then we'll say
  16356. 9:45:03the name is female, okay?
  16357. 9:45:08Else print name is masculine or name is
  16358. 9:45:13Always let let be male.
  16359. 9:45:15Okay. So, same thing over here, we'll
  16360. 9:45:16just give it as
  16361. 9:45:18male.
  16362. 9:45:19So, let's now see what this thing
  16363. 9:45:20predicts. Okay, so this thing predicts
  16364. 9:45:22male.
  16365. 9:45:23So, let's take another common name.
  16366. 9:45:26Let's take something like Let's go to
  16367. 9:45:28Google and see what name can we take.
  16368. 9:45:31Yeah, we can take up something like Brad
  16369. 9:45:32Pitt.
  16370. 9:45:33Okay, so let's execute this and let's
  16371. 9:45:36execute this again.
  16372. 9:45:37And then this.
  16373. 9:45:39So, as you can see here, this is giving
  16374. 9:45:40me a male name. And let's give a female
  16375. 9:45:42name over here. Let's give as Mary.
  16376. 9:45:45And let's see whether this would predict
  16377. 9:45:47it as male or female. So, as you can
  16378. 9:45:49see, it's female.
  16379. 9:45:50Now, as if you have seen the data set,
  16380. 9:45:52right? It says this is the name from the
  16381. 9:45:54US kids, right? What about What will
  16382. 9:45:56happen if I give a Indian-based name?
  16383. 9:45:59So, let me give a Indian-based name like
  16384. 9:46:02Priyanka.
  16385. 9:46:03And let me execute this.
  16386. 9:46:05It's giving me a female name, right? So,
  16387. 9:46:07this is something pretty astonishing,
  16388. 9:46:09right? So, why do you think it gave me a
  16389. 9:46:10female name? Well, the reason is because
  16390. 9:46:12when we are training this model like RNN
  16391. 9:46:15using RNN, right? It's not looking at
  16392. 9:46:17the name. It doesn't know whether the
  16393. 9:46:18name is female or not. But, as a matter
  16394. 9:46:21of fact, it is looking at the pattern.
  16395. 9:46:23Okay? So, it might be looking at, you
  16396. 9:46:25know, if the name ends with so and so,
  16397. 9:46:26it is a female. If the name starts with
  16398. 9:46:29this or if a name have something like
  16399. 9:46:31this, it means, you know, it's a female
  16400. 9:46:33or a male. So, to give you a better
  16401. 9:46:35analogy, let's give something like
  16402. 9:46:37Julia.
  16403. 9:46:39So, let me execute this.
  16404. 9:46:41This will give me a female, right? So,
  16405. 9:46:43what if I give something like Juneid?
  16406. 9:46:46So, it give it as male. So, all it's
  16407. 9:46:48trying to do is it's trying to see, you
  16408. 9:46:50know, uh recognize a pattern. That's why
  16409. 9:46:52it's taking individual words at the same
  16410. 9:46:54time. So, this is what makes NLP using
  16411. 9:46:57LSTM very effective.
  16412. 9:47:00All right. So, moving ahead, let's see
  16413. 9:47:02some of the LSTMs use cases.
  16414. 9:47:04You see, LSTM is a very popular deep
  16415. 9:47:06learning algorithm for sequential
  16416. 9:47:08models.
  16417. 9:47:09Apple Siri and Google's voice search are
  16418. 9:47:11some of the real-world examples that
  16419. 9:47:13have used LSTM. And you won't believe
  16420. 9:47:15it, LSTM is a success story for those
  16421. 9:47:17algorithm.
  16422. 9:47:18So, let us now have a look and see how
  16423. 9:47:20LSTM changed that technology.
  16424. 9:47:22Okay. So, starting off with Apple, in
  16425. 9:47:242003, Apple was the first major tech
  16426. 9:47:26company to integrate a smart assistant,
  16427. 9:47:28that is Siri, into their operating
  16428. 9:47:30system. And the Siri was actually a
  16429. 9:47:32byproduct of some other company. So,
  16430. 9:47:34Siri was a company's adoption of a
  16431. 9:47:36standalone app that has been purchased
  16432. 9:47:38along with the creators who made it. It
  16433. 9:47:39was somewhere in 2010. The initial
  16434. 9:47:42reviews about Siri was that it was
  16435. 9:47:43intense. But, over the next few months
  16436. 9:47:45or and years, the users became more
  16437. 9:47:47impatient with the shortcomings. And all
  16438. 9:47:49too often, it wrongly interpreted
  16439. 9:47:51commands. And then, you know, no matter
  16440. 9:47:54what you do, there was no fix for it.
  16441. 9:47:56So, this is when Apple moved Siri's
  16442. 9:47:57voice recognition to a neural-based
  16443. 9:47:59system.
  16444. 9:48:00Some of the previous technique remained
  16445. 9:48:02operational, something like, you know,
  16446. 9:48:03applying hidden Markov models. But, most
  16447. 9:48:05of the time, you know, DNN or deep
  16448. 9:48:07neural network using LSTM was used.
  16449. 9:48:10Although people did not find any changes
  16450. 9:48:12on the outside, but from within, it was
  16451. 9:48:14a supercharged deep learning model.
  16452. 9:48:16Speaking about Google's implementation,
  16453. 9:48:19Google implemented Google voice search
  16454. 9:48:20somewhere around 2009.
  16455. 9:48:22Google voice transcription had initially
  16456. 9:48:24used something called as Gaussian
  16457. 9:48:25mixture model.
  16458. 9:48:27This was This was nothing but an
  16459. 9:48:28acoustic model and this was something
  16460. 9:48:29considered to be a state-of-the-art
  16461. 9:48:31speech recognition for almost 30 plus
  16462. 9:48:33years.
  16463. 9:48:34But it was in 2012, there was a boom in
  16464. 9:48:36deep neural network.
  16465. 9:48:37And when Google implemented deep neural
  16466. 9:48:39network, that too using multiple layer
  16467. 9:48:41networks, there was a huge performance
  16468. 9:48:43gap.
  16469. 9:48:44But things really improved when the
  16470. 9:48:46recurrent neural network, especially
  16471. 9:48:47with LSTM RNN, first launched on an
  16472. 9:48:49Android speech recognition in May 2012.
  16473. 9:48:52Compared to deep neural network, LSTM
  16474. 9:48:54RNNs have additional recurrent
  16475. 9:48:56connections and memory cell that allows
  16476. 9:48:58them to remember the previous data.
  16477. 9:49:00All right. So, moving ahead to our last
  16478. 9:49:01topic of our session, let us now see
  16479. 9:49:03some of the real-world applications of
  16480. 9:49:04LSTM RNN networks. First off, we can
  16481. 9:49:07perform named entity recognition. This
  16482. 9:49:09is something which we did in our
  16483. 9:49:10previous demo, right? So, what is this
  16484. 9:49:12named entity recognition? You see, named
  16485. 9:49:14entity recognition is a subtask of
  16486. 9:49:16information extraction that seeks to
  16487. 9:49:18locate and classify named entity
  16488. 9:49:20mentioned in an unstructured data. Okay?
  16489. 9:49:23Next, we have something called as
  16490. 9:49:24sentiment analysis.
  16491. 9:49:26Sentiment analysis is a predictive
  16492. 9:49:28modeling task where model is trained to
  16493. 9:49:30predict the polarity of a textual data
  16494. 9:49:32or sentiments like positive, neutral, or
  16495. 9:49:34negative.
  16496. 9:49:35Sentiment analysis is performed by
  16497. 9:49:37various businesses to understand their
  16498. 9:49:39consumers' behavior towards the product.
  16499. 9:49:42Then we have machine translation. The
  16500. 9:49:44task of machine translation consists of
  16501. 9:49:45reading text in one language and
  16502. 9:49:47generating text in another language.
  16503. 9:49:49When neural networks are used for this
  16504. 9:49:51task, we talk about neural machine
  16505. 9:49:53translation.
  16506. 9:49:55Within neural machine translation, an
  16507. 9:49:56encoder-decoder structure is quite a
  16508. 9:49:58popular LSTM RNN architecture.
  16509. 9:50:02>> [music]
  16510. 9:50:06>> Why should we choose deep learning for,
  16511. 9:50:08you know, various tasks?
  16512. 9:50:10So, the big advantage of using deep
  16513. 9:50:12learning is that we can extract more
  16514. 9:50:14number of features. And when we have
  16515. 9:50:16more number of features and when we can
  16516. 9:50:17work at the same time with huge amount
  16517. 9:50:19of data, we can perceive an object like
  16518. 9:50:21a human being does. What I'm trying to
  16519. 9:50:23say over here is like if you want to
  16520. 9:50:25perform a classification task between
  16521. 9:50:27pen and a pencil, you'll obviously know
  16522. 9:50:29as a human being you'll know the
  16523. 9:50:30difference because you have look at a
  16524. 9:50:32pen and a pencil continuous number of
  16525. 9:50:34times. And now when you're trying to
  16526. 9:50:36actually classify it, you can do it with
  16527. 9:50:38ease.
  16528. 9:50:39Okay? And the reason for this is because
  16529. 9:50:40you know the features of a pen and you
  16530. 9:50:42know the features of a pencil. Okay?
  16531. 9:50:44Similarly, this is how deep learning
  16532. 9:50:46works. More the data you feed, more the
  16533. 9:50:48dimensions it can analyze. More the
  16534. 9:50:50dimensions it can learn. All right? So,
  16535. 9:50:52as I've already mentioned, one of the
  16536. 9:50:54most popular application of deep
  16537. 9:50:56learning is image classification. And
  16538. 9:50:58when it comes to image classification,
  16539. 9:51:00it can be something as simple as
  16540. 9:51:01classifying between two different
  16541. 9:51:02animals to something as complicated as,
  16542. 9:51:05you know, hiding data or trying to run
  16543. 9:51:08automated cars using classification
  16544. 9:51:10task. Okay? All right. So, next type of
  16545. 9:51:13application using deep learning is using
  16546. 9:51:15on sequential data. Sequential data
  16547. 9:51:18basically refers to something like time
  16548. 9:51:20series data or having to understand
  16549. 9:51:22natural language.
  16550. 9:51:23So, the reason why we call it sequential
  16551. 9:51:25data is because here the previous word
  16552. 9:51:28or the previous feature is dependent
  16553. 9:51:30upon the next feature. Okay? So, as you
  16554. 9:51:32can see over here, we have what time is
  16555. 9:51:34it, right? So, if I just say it is like
  16556. 9:51:37over here what time is and it are
  16557. 9:51:39basically features, right? And in order
  16558. 9:51:40for you to make an analogy or to
  16559. 9:51:42understand, obviously have to know what
  16560. 9:51:44has happened in the past. So, in order
  16561. 9:51:45to do this, we use something called as
  16562. 9:51:47RNNs. Okay? And there are various
  16563. 9:51:49versions of RNN that go around in order
  16564. 9:51:51to overcome the disadvantages which
  16565. 9:51:53we'll look in in sometime. All right.
  16566. 9:51:56So, moving on to the next application
  16567. 9:51:57that is GANs. GANs, which stands for
  16568. 9:52:00generative adversarial network, is an
  16569. 9:52:02unsupervised part of a deep learning
  16570. 9:52:04application. Some of the common
  16571. 9:52:06application which you can see in recent
  16572. 9:52:07days is nothing but deep fakes and many
  16573. 9:52:09more. Finally, coming down to performing
  16574. 9:52:11classification and regression task using
  16575. 9:52:13multi-layer perceptron. If you remember
  16576. 9:52:15or if you're well versed with machine
  16577. 9:52:17learning, in order to perform
  16578. 9:52:18classification in machine learning, we
  16579. 9:52:20had algorithms like decision tree,
  16580. 9:52:21random forest, or something very simple
  16581. 9:52:24as linear regression or logistic
  16582. 9:52:26regression. But, let me tell you what.
  16583. 9:52:28When we try to perform classification
  16584. 9:52:30using MLP, or multi-layer perceptron, we
  16585. 9:52:32get a very high accuracy even compared
  16586. 9:52:34to SVM and decision trees. All right.
  16587. 9:52:37So, now that we know what exactly is
  16588. 9:52:38deep learning and why we use it, let's
  16589. 9:52:41now stream down to understand how can we
  16590. 9:52:43process natural language data using
  16591. 9:52:45RNNs.
  16592. 9:52:46So, what are RNNs, right? Well, RNN
  16593. 9:52:48basically stands for recurrent neural
  16594. 9:52:50network. And we usually use this in
  16595. 9:52:52order to deal with a sequential data.
  16596. 9:52:54Sequential data can be something like a
  16597. 9:52:56time series data, or a textual data of
  16598. 9:52:58any format. So, why should one use RNN,
  16599. 9:53:00right? Well, this is because there's a
  16600. 9:53:02concept of internal memory here. RNN can
  16601. 9:53:04remember important things about the
  16602. 9:53:06input it has received. Which allows them
  16603. 9:53:09to be very precise in predicting what
  16604. 9:53:11can be the next outcome. So, this is the
  16605. 9:53:13reason why they are performed or
  16606. 9:53:14preferred on a sequential data
  16607. 9:53:16algorithm, okay? And some of the
  16608. 9:53:17examples of sequential data can be
  16609. 9:53:19something like time series, speech,
  16610. 9:53:21text, financial data, audio, video,
  16611. 9:53:23weather, and many more. Although RNN
  16612. 9:53:26were the state-of-the-art algorithm for
  16613. 9:53:27dealing with sequential data, they come
  16614. 9:53:29up with their own drawbacks. And some of
  16615. 9:53:31the popular drawbacks over here can be
  16616. 9:53:33like, due to the complication or the
  16617. 9:53:35complexity of the algorithm, the neural
  16618. 9:53:37network is pretty slow to train. And as
  16619. 9:53:39there are a huge amount of dimensions
  16620. 9:53:40here, the training is very long and
  16621. 9:53:43difficult to do, okay? Apart from that,
  16622. 9:53:45the most decisive feature for RNN, or
  16623. 9:53:47for the improvement in RNN, is that of a
  16624. 9:53:50vanishing gradient. What this vanishing
  16625. 9:53:52gradient is is that, you know, when we
  16626. 9:53:54go deeper and deeper into our neural
  16627. 9:53:56network, the previous data is lost. This
  16628. 9:53:59is because of a concept called as
  16629. 9:54:01vanishing gradient. And due to this, we
  16630. 9:54:03cannot work on a large or a longer
  16631. 9:54:05sequence of data. Okay? To overcome
  16632. 9:54:08this, we came up with some new or
  16633. 9:54:10upgrades to the current recurrent neural
  16634. 9:54:12networks or RNNs.
  16635. 9:54:14Starting off with bidirectional
  16636. 9:54:15recurrent neural network. You see,
  16637. 9:54:17bidirectional recurrent neural network
  16638. 9:54:19connect two hidden layers of opposite
  16639. 9:54:21direction into the same output. With
  16640. 9:54:23this form of generative deep learning,
  16641. 9:54:25the output layer can get information
  16642. 9:54:27from past future states simultaneously.
  16643. 9:54:30So, as you can see here, we have two
  16644. 9:54:32layers over here, and as they are
  16645. 9:54:34bidirectional, what happens is when the
  16646. 9:54:36algorithm feels that it is kind of
  16647. 9:54:37losing its gradients or the previous
  16648. 9:54:39data, it can go back and get the data
  16649. 9:54:41from the past. So, why do we need
  16650. 9:54:44bidirectional recurrent neural network?
  16651. 9:54:45Well, bidirectional recurrent neural
  16652. 9:54:47network duplicates RNN processing chain
  16653. 9:54:50so that the input process both forward
  16654. 9:54:52and reverse time order,
  16655. 9:54:53thus allowing bidirectional recurrent
  16656. 9:54:55neural network to look into future
  16657. 9:54:57context as well. The next one is long
  16658. 9:54:59short-term memory. Long short-term
  16659. 9:55:01memory or also sometime referred to as
  16660. 9:55:03LSTM is a artificial recurrent neural
  16661. 9:55:05network architecture used in the field
  16662. 9:55:07of deep learning. Unlike standard
  16663. 9:55:09feedforward neural network, LSTM has a
  16664. 9:55:11feedback connections. It can not only
  16665. 9:55:13process single data point, but also the
  16666. 9:55:15entire sequence of data. So, as you can
  16667. 9:55:17see here, from what I'm trying to say is
  16668. 9:55:19with LSTM or long short-term memory, it
  16669. 9:55:22has something like, you know, we can
  16670. 9:55:23feed a longer sequence compared to what
  16671. 9:55:25it was with bidirectional RNN or RNNs.
  16672. 9:55:29So, why is LSTM better than RNN? We can
  16673. 9:55:31say that when we move from RNN to LSTM,
  16674. 9:55:34we are introducing more and more control
  16675. 9:55:36over the sequence of the data that we
  16676. 9:55:38can provide. The LSTM gives us more
  16677. 9:55:40control ability and does better results.
  16678. 9:55:43All right. So, the next type of
  16679. 9:55:44recurrent neural network is the gated
  16680. 9:55:46recurrent neural network or also
  16681. 9:55:48referred to as GRUs. You see, GRU is a
  16682. 9:55:50type of recurrent neural network that
  16683. 9:55:52is, in certain cases, is advantageous
  16684. 9:55:55over long short-term memory. GRU makes
  16685. 9:55:57use of less memory and also is faster
  16686. 9:55:59than LSTM. But thing is, LSTMs are more
  16687. 9:56:02accurate while using longer data sets.
  16688. 9:56:05I'm sure by now you might have got a
  16689. 9:56:07hint about the trend that has led to the
  16690. 9:56:09improvement, right? So, the trend over
  16691. 9:56:11here is, you know, the model should be
  16692. 9:56:13capable of remembering and taking in on
  16693. 9:56:16a longer input sequence.
  16694. 9:56:18The game-changer part for the sequential
  16695. 9:56:20data was developed when we came up with
  16696. 9:56:22something called as transformers. And
  16697. 9:56:24this paper was something which is based
  16698. 9:56:26on a concept called as attention is
  16699. 9:56:29everything.
  16700. 9:56:30All right. So, let's take a look at
  16701. 9:56:31this.
  16702. 9:56:32The paper attention is all you need
  16703. 9:56:35introduces a novel architecture called
  16704. 9:56:37as transformers. Like LSTM, transformers
  16705. 9:56:40is an architecture for transforming one
  16706. 9:56:42sequence into another while helping
  16707. 9:56:44adapt to parts, that is encoders and
  16708. 9:56:46decoders. But it differs from previously
  16709. 9:56:48described sequence to sequence model
  16710. 9:56:50because it does not work like GRUs,
  16711. 9:56:52okay? So, it does not implements uh
  16712. 9:56:55recurrent neural networks.
  16713. 9:56:57Recurrent neural network until now were
  16714. 9:56:59one of the best ways to capture the
  16715. 9:57:00timely dependence on a sequence.
  16716. 9:57:03However, the team presenting this paper,
  16717. 9:57:05that is attention is all you need,
  16718. 9:57:06proved that an architecture with only
  16719. 9:57:08attention mechanism does not use RNN can
  16720. 9:57:11improve its result in translation task
  16721. 9:57:14and other NLP task. One of the best
  16722. 9:57:16examples for transformers is Google's
  16723. 9:57:18BERT. So, what exactly is this
  16724. 9:57:20transformer, right? You see, here we
  16725. 9:57:22have encoder on the top and decoder on
  16726. 9:57:24the bottom. Both encoder and decoder are
  16727. 9:57:26comprised of modules that can stick onto
  16728. 9:57:29the top of each other multiple times.
  16729. 9:57:31So, what happens here is the inputs and
  16730. 9:57:33outputs are first embedded into
  16731. 9:57:35N-dimension space since we cannot use
  16732. 9:57:37this directly. So, we obviously have to
  16733. 9:57:39encode our inputs, whatever we are
  16734. 9:57:41providing here. One slight but important
  16735. 9:57:43part of this model is the positional
  16736. 9:57:45encoding of different words. Since we
  16737. 9:57:47have no recurrent neural network that
  16738. 9:57:49can remember how sequence are fed into
  16739. 9:57:51the model, we need to somehow give every
  16740. 9:57:53word or part of our sequence a relative
  16741. 9:57:55position since the sequence depends on
  16742. 9:57:58the order of the elements, okay? These
  16743. 9:58:00positions are added to the embedded
  16744. 9:58:02representation of each words. All right.
  16745. 9:58:04So, this was the brief about
  16746. 9:58:06transformers. So, let us now move ahead
  16747. 9:58:08and see some of the popular language
  16748. 9:58:10models that are available in the market.
  16749. 9:58:12All right. So, let us now start off by
  16750. 9:58:14understanding OpenAI's GPT-3. The
  16751. 9:58:16successor to GPT and GPT-2 is the GPT-3
  16752. 9:58:20and is one of the most controversial
  16753. 9:58:22pre-trained models by OpenAI. The
  16754. 9:58:24large-scale transformer-based language
  16755. 9:58:26model has been trained on 175 billion
  16756. 9:58:29parameters, which is 10 times more than
  16757. 9:58:31any previous non-sparse language model.
  16758. 9:58:34The model has been trained to achieve
  16759. 9:58:36strong performance on many NLP data set,
  16760. 9:58:38including tasks like translation,
  16761. 9:58:41answering questions, as well as several
  16762. 9:58:42other tasks. Then we have Google's BERT.
  16763. 9:58:45BERT stands for bidirectional encoder
  16764. 9:58:47representations from transformers. It is
  16765. 9:58:50a pre-trained NLP model, which is
  16766. 9:58:52developed by Google in 2018. With this,
  16767. 9:58:54anyone in the world can train either
  16768. 9:58:56their own question answering module with
  16769. 9:58:58up to 30 minutes on a single cloud TPU
  16770. 9:59:01or few hours using single GPU. The
  16771. 9:59:04company then released this showcasing
  16772. 9:59:06the performance of 11 NLP tasks,
  16773. 9:59:08including very competitive Stanford
  16774. 9:59:10dataset questions.
  16775. 9:59:12Unlike other language model, BERT has
  16776. 9:59:14only been pre-trained on 250 million
  16777. 9:59:16words of Wikipedia and 800 million words
  16778. 9:59:18of book corpus and has been successfully
  16779. 9:59:21used as a pre-trained model in deep
  16780. 9:59:23neural network. According to
  16781. 9:59:24researchers, BERT has achieved 93%
  16782. 9:59:27accuracy, which has surpassed any
  16783. 9:59:28previous language models.
  16784. 9:59:30Next, we have ELMo. ELMo, also known as
  16785. 9:59:33embedding for language model, is a deep
  16786. 9:59:35contextualized word representation that
  16787. 9:59:38models syntax and semantic words, as
  16788. 9:59:40well as their logistic context. The
  16789. 9:59:42model developed by Allen NLP has been
  16790. 9:59:44pre-trained on a huge text corpus and
  16791. 9:59:47learned functions from bidirectional
  16792. 9:59:49models, that is BiLM. ELMo can easily be
  16793. 9:59:52added to their existing models, which
  16794. 9:59:54drastically improves the features of
  16795. 9:59:56functions across vast NLP problem,
  16796. 9:59:59including answering questions, textual
  16797. 10:00:01entailment, and sentiment analysis.
  16798. 10:00:06>> [music]
  16799. 10:00:09>> What are GANs?
  16800. 10:00:11So, we're going to start with generative
  16801. 10:00:13models.
  16802. 10:00:14So, generative models are nothing but
  16803. 10:00:16those models that use an unsupervised
  16804. 10:00:19learning approach.
  16805. 10:00:20In a generative model, there are samples
  16806. 10:00:23in the data that is input variables X,
  16807. 10:00:26but it lacks a output variable Y. And we
  16808. 10:00:29use the only input variables to train
  16809. 10:00:31the generative model, and it recognizes
  16810. 10:00:34patterns from the input variables to
  16811. 10:00:36generate an output that is unknown and
  16812. 10:00:39based on the training data only.
  16813. 10:00:41In supervised learning, we are more
  16814. 10:00:43aligned towards creating predictive
  16815. 10:00:45models from the input variables.
  16816. 10:00:48And this type of modeling is also known
  16817. 10:00:50as discriminative modeling.
  16818. 10:00:52And in a classification problem, the
  16819. 10:00:54model has to discriminate as to which
  16820. 10:00:57class the example belongs to. And on the
  16821. 10:00:59other hand, unsupervised models are used
  16822. 10:01:01to create or generate new examples in
  16823. 10:01:04the input distribution.
  16824. 10:01:06To define a generative model in layman
  16825. 10:01:09terms, we can say generative models are
  16826. 10:01:12able to generate new examples from the
  16827. 10:01:15sample that are not only similar to the
  16828. 10:01:17examples, but are indistinguishable as
  16829. 10:01:20well.
  16830. 10:01:21And the most common example of a
  16831. 10:01:22generative model is a naive Bayes
  16832. 10:01:24classifier, which is more often used as
  16833. 10:01:27a discriminative model.
  16834. 10:01:29Other examples of generative models
  16835. 10:01:31include Gaussian mixture model and a
  16836. 10:01:33rather modern example, that is
  16837. 10:01:35generative adversarial networks.
  16838. 10:01:38So, let us try to understand what
  16839. 10:01:39exactly are GANs, or generative
  16840. 10:01:42adversarial networks.
  16841. 10:01:44Generative adversarial networks, or
  16842. 10:01:46GANs, are a deep learning-based
  16843. 10:01:48generative model that is used for
  16844. 10:01:50unsupervised learning.
  16845. 10:01:52It is basically a system where two
  16846. 10:01:54competing neural networks compete with
  16847. 10:01:56each other to create or generate
  16848. 10:01:58variations in the data.
  16849. 10:02:00It was first described in a paper in
  16850. 10:02:022014 by Ian Goodfellow and a
  16851. 10:02:05standardized and much stable model
  16852. 10:02:07theory was proposed by Alec Radford in
  16853. 10:02:102016, which is also known as DCGAN.
  16854. 10:02:15Also known as DCGAN or we can call it as
  16855. 10:02:18deep convolutional generative
  16856. 10:02:20adversarial networks.
  16857. 10:02:22And most of the GANs today use deep
  16858. 10:02:24convolutional generative adversarial
  16859. 10:02:25networks.
  16860. 10:02:27The GANs architecture consists of two
  16861. 10:02:29sub models known as the generator model
  16862. 10:02:32and the discriminator model.
  16863. 10:02:34So, a generator network takes a sample
  16864. 10:02:36and generates sample of data.
  16865. 10:02:38A discriminator network decides whether
  16866. 10:02:40the data is generated or taken from the
  16867. 10:02:42real sample using a binary
  16868. 10:02:44classification problem with the help of
  16869. 10:02:46a sigmoid function that gives the output
  16870. 10:02:49in the form or the range zero and one.
  16871. 10:02:52So, let us go ahead and take a look at
  16872. 10:02:54how GANs actually work.
  16873. 10:02:57To understand how GANs work, let's break
  16874. 10:02:59it down.
  16875. 10:03:00So, generative means that the model
  16876. 10:03:02follows the unsupervised learning
  16877. 10:03:04approach and is a generative model.
  16878. 10:03:07When we talk about adversarial, the
  16879. 10:03:09model is trained in an adversarial
  16880. 10:03:11setting.
  16881. 10:03:12And network simply means for the
  16882. 10:03:14training of the model, we use the neural
  16883. 10:03:16networks as artificial intelligence
  16884. 10:03:18algorithms.
  16885. 10:03:20In GANs, there is a generator network
  16886. 10:03:22that takes a sample and generates a
  16887. 10:03:24sample of data.
  16888. 10:03:26And after this, the discriminator
  16889. 10:03:27network decides whether the data is
  16890. 10:03:29generated or taken from the real sample
  16891. 10:03:31using a binary classification problem
  16892. 10:03:33with the help of a sigmoid function that
  16893. 10:03:36gives the output in the range zero to
  16894. 10:03:37one.
  16895. 10:03:38The generative model analyzes the
  16896. 10:03:40distribution of the data in such a way
  16897. 10:03:42that after the training phase, the
  16898. 10:03:44probability of the discriminator making
  16899. 10:03:46a mistake maximizes. and the
  16900. 10:03:48discriminator on the other hand is based
  16901. 10:03:50on a model that will estimate the
  16902. 10:03:52probability that the sample is coming
  16903. 10:03:54from the real data or not the generator.
  16904. 10:03:57The whole process can be formalized in a
  16905. 10:03:59mathematical formula.
  16906. 10:04:01So G over here is generator, D is equal
  16907. 10:04:04to discriminator, P data X is the
  16908. 10:04:07distribution of real data, P data Z is
  16909. 10:04:10the distributor of generator, X is the
  16910. 10:04:13sample from the real data, and Z is the
  16911. 10:04:16sample from generator. Where DX is the
  16912. 10:04:18discriminator network and GZ is a
  16913. 10:04:21generator network.
  16914. 10:04:23So let's take a look at the flowchart
  16915. 10:04:24once again, guys.
  16916. 10:04:25So we have the training data, which is
  16917. 10:04:27going to give the real sample. And the
  16918. 10:04:29generator network is going to generate
  16919. 10:04:31the sample from the random noise or the
  16920. 10:04:33examples.
  16921. 10:04:34And then it will go to the discriminator
  16922. 10:04:36network, where it's going to check if
  16923. 10:04:38the sample that is coming is real or
  16924. 10:04:41fake.
  16925. 10:04:42So that is how a GAN actually work.
  16926. 10:04:44Now let's take a look at the training
  16927. 10:04:46phase, like how a generative adversarial
  16928. 10:04:48network is actually trained.
  16929. 10:04:50So it happens in two phases, guys.
  16930. 10:04:53So the first phase is where we train the
  16931. 10:04:55discriminator and we actually freeze the
  16932. 10:04:57generator, which means that the training
  16933. 10:05:00set for the generator is done false and
  16934. 10:05:02the network will only do the forward
  16935. 10:05:04pass and there will not be any back
  16936. 10:05:06propagation.
  16937. 10:05:08Basically, the discriminator is trained
  16938. 10:05:10with real data and checks if it can
  16939. 10:05:12predict them correctly.
  16940. 10:05:14And the same with the fake data to
  16941. 10:05:15identify them as fake.
  16942. 10:05:18After this, there's the second part
  16943. 10:05:20where we train the generator and freeze
  16944. 10:05:22the discriminator.
  16945. 10:05:24So we get the result from the first
  16946. 10:05:25phase and we use them to make better
  16947. 10:05:28from the previous state to try and fool
  16948. 10:05:30the discriminator better.
  16949. 10:05:32So to understand this in the layman's
  16950. 10:05:33term, I'm going to tell you a few steps
  16951. 10:05:35for training, like how you should start.
  16952. 10:05:38So the first step is you have to define
  16953. 10:05:40in problem.
  16954. 10:05:41You've got to define the problem and
  16955. 10:05:43collect the data.
  16956. 10:05:44After this, the second step is you have
  16957. 10:05:47to choose the architecture of GAN.
  16958. 10:05:49So, in this step, depending on your
  16959. 10:05:50problem, you have to choose how your GAN
  16960. 10:05:52should look like.
  16961. 10:05:54The third step is training the
  16962. 10:05:55discriminator on real data. So, we train
  16963. 10:05:58the discriminator with real data to
  16964. 10:06:00predict them as real for n number of
  16965. 10:06:02times, so we call it a epochs as well.
  16966. 10:06:04And then we generate the fake inputs
  16967. 10:06:06from the generator.
  16968. 10:06:08So, in this step, we are going to
  16969. 10:06:09generate the fake samples from the
  16970. 10:06:11generator.
  16971. 10:06:12And the next step is we train the
  16972. 10:06:14discriminator on fake data.
  16973. 10:06:17So, whatever samples are generated from
  16974. 10:06:18the generator network, you're going to
  16975. 10:06:20train the discriminator to predict the
  16976. 10:06:21generated data as fake.
  16977. 10:06:24So, that's how we know that
  16978. 10:06:25discriminator is actually predicting the
  16979. 10:06:27values as correctly.
  16980. 10:06:28And the last step is we train the
  16981. 10:06:30generator with the output of
  16982. 10:06:31discriminator. So, after getting the
  16983. 10:06:33discriminator predictions, we train the
  16984. 10:06:36generator to fool the discriminator.
  16985. 10:06:39So, that's how we train the GAN to
  16986. 10:06:41actually get our solution from the
  16987. 10:06:43problem. Which is like defining the
  16988. 10:06:45problem.
  16989. 10:06:46So, you'll understand this when I'm
  16990. 10:06:47talking about the applications, guys. No
  16991. 10:06:49worry.
  16992. 10:06:50Now, let's go ahead and take a look at a
  16993. 10:06:51few challenges of generative adversarial
  16994. 10:06:54networks.
  16995. 10:06:55So, the concept of GANs is rather
  16996. 10:06:57fascinating, but there are a lot of
  16997. 10:07:00setbacks that can cause a lot of
  16998. 10:07:01hindrance in its path.
  16999. 10:07:03Some of the major challenges faced by
  17000. 10:07:05GANs are
  17001. 10:07:06The first one is the stability.
  17002. 10:07:08So, there has to be a stability that is
  17003. 10:07:09required between discriminator and the
  17004. 10:07:11generator network, otherwise the whole
  17005. 10:07:13network would just fall.
  17006. 10:07:15For example, in case, let's say if the
  17007. 10:07:17discriminator is too powerful, the
  17008. 10:07:20generator will fail to train altogether.
  17009. 10:07:22Won't be able to push fake samples to
  17010. 10:07:24that discriminator, and it will always
  17011. 10:07:27identify them as fake.
  17012. 10:07:29And let's say if the network is too
  17013. 10:07:30lenient, the discriminator network is
  17014. 10:07:32too lenient,
  17015. 10:07:34so any image that would be generated by
  17016. 10:07:36the generator network would make the
  17017. 10:07:38network useless.
  17018. 10:07:39The next challenge that is faced by GANs
  17019. 10:07:42is GANs fail miserably in determining
  17020. 10:07:44the positioning of the objects in terms
  17021. 10:07:46of how many times the objects should
  17022. 10:07:48occur at that location. Suppose we have
  17023. 10:07:50a image in which we have, let's say,
  17024. 10:07:52three dogs with two eyes and sometimes a
  17025. 10:07:56GAN will fail to, you know, determine
  17026. 10:07:58the positioning of the objects in terms
  17027. 10:07:59of it will generate an image with like
  17028. 10:08:01one dog and six eyes. So, that's kind of
  17029. 10:08:04a problem that we face while working on
  17030. 10:08:06GANs.
  17031. 10:08:07And the next challenge is 3D perspective
  17032. 10:08:10troubles GANs as it is not able to
  17033. 10:08:12understand the perspective also.
  17034. 10:08:15So, it will often give a flat image for
  17035. 10:08:16a 3D object. So, that's one challenge
  17036. 10:08:19that we face with GANs as well.
  17037. 10:08:21And GANs have a problem of understanding
  17038. 10:08:24the global objects and it cannot
  17039. 10:08:26differentiate or understand a holistic
  17040. 10:08:28structure.
  17041. 10:08:29Like, if you're talking about trees or
  17042. 10:08:31if you're talking about flowers, that's
  17043. 10:08:33a problem that GANs will follow.
  17044. 10:08:35And last but not least, newer types of
  17045. 10:08:37GANs are more advanced that are brought
  17046. 10:08:39about that is deep convolutional
  17047. 10:08:41generative adversarial networks and are
  17048. 10:08:44expected to overcome these shortcomings
  17049. 10:08:46altogether. So, that we don't have to
  17050. 10:08:47worry about these. These are the
  17051. 10:08:49shortcomings that we face with normal
  17052. 10:08:51GANs, uh initial generative adversarial
  17053. 10:08:53networks. Now that they have become more
  17054. 10:08:55advanced, they actually overcome these
  17055. 10:08:58shortcomings, so you don't have to
  17056. 10:08:59worry, guys.
  17057. 10:09:00So, last but not the least, I want to
  17058. 10:09:02talk about a few applications of
  17059. 10:09:04generative adversarial networks.
  17060. 10:09:06So, the first one is prediction of next
  17061. 10:09:08frame in a video.
  17062. 10:09:10So, let's say the prediction of future
  17063. 10:09:11events in a video frame is made possible
  17064. 10:09:13with the help of GANs and DVD GAN or we
  17065. 10:09:17can call it as dual video discriminator
  17066. 10:09:19GAN can generate a 256 by 256 videos of
  17067. 10:09:23notable fidelity up to 48 frames in
  17068. 10:09:26length.
  17069. 10:09:27And this can be used for various
  17070. 10:09:28purposes including surveillance in which
  17071. 10:09:31we can determine the activities in a
  17072. 10:09:32frame that gets distorted due to other
  17073. 10:09:35factors like rain, dust, smoke, etc.
  17074. 10:09:38So, the possibilities are immense with
  17075. 10:09:40this if you're able to predict the next
  17076. 10:09:42frame in a video. That actually helps in
  17077. 10:09:44a lot of things like surveillance,
  17078. 10:09:45security, and we can predict outcomes
  17079. 10:09:48based on these frames that we generate
  17080. 10:09:50from a video.
  17081. 10:09:51After this comes the text to image
  17082. 10:09:53generation.
  17083. 10:09:54So, basically object-driven attentive
  17084. 10:09:56GAN, which is also known as object GAN,
  17085. 10:09:58performs the text to image synthesis in
  17086. 10:10:01two steps. So, the first step is
  17087. 10:10:03generating the semantic layout and then
  17088. 10:10:06generating the image by synthesizing the
  17089. 10:10:07image by using a deconvolutional image
  17090. 10:10:10generator is the final step.
  17091. 10:10:13So, this could be used intensively to
  17092. 10:10:14generate images by understanding the
  17093. 10:10:16captions, the layouts, and refine
  17094. 10:10:18details by synthesizing the words.
  17095. 10:10:21And there is another study about the
  17096. 10:10:23story GANs that can synthesize the whole
  17097. 10:10:25storyboards from mere paragraphs.
  17098. 10:10:28So, that's actually very good idea if
  17099. 10:10:30you're talking about GANs. So, you can
  17100. 10:10:31just give a few layouts and captions.
  17101. 10:10:35Based on that, it will generate image
  17102. 10:10:36for us.
  17103. 10:10:38Talking about the next application, we
  17104. 10:10:39have image to image translation.
  17105. 10:10:42So, Pix2Pix is a model which is designed
  17106. 10:10:44for general purpose image to image
  17107. 10:10:46translation.
  17108. 10:10:48So, let's say we have three images.
  17109. 10:10:50We have a real image.
  17110. 10:10:51Then we'll be having a generated image,
  17111. 10:10:54which is basically a fake, and then it
  17112. 10:10:56will be reconstructed to the previous
  17113. 10:10:58image which was real.
  17114. 10:11:00So, this is how image to image
  17115. 10:11:01translation work, guys.
  17116. 10:11:03And after this, we have enhancing the
  17117. 10:11:04resolution of an image.
  17118. 10:11:06So, super-resolution generative
  17119. 10:11:08adversarial network, or also known as
  17120. 10:11:10SRGAN, is a GAN which can generate the
  17121. 10:11:14super-resolution images from
  17122. 10:11:16low-resolution images with finer details
  17123. 10:11:18and better quality.
  17124. 10:11:20So, this is actually a very good
  17125. 10:11:22application of GANs, guys. The
  17126. 10:11:24applications can be immense.
  17127. 10:11:26So, you imagine a higher quality image
  17128. 10:11:29with finer details generated from a low
  17129. 10:11:31resolution image. The amount of help it
  17130. 10:11:34would produce to identify details in
  17131. 10:11:36lower resolution images can be used for
  17132. 10:11:38wider purposes including surveillance.
  17133. 10:11:41We can use it for documentation
  17134. 10:11:43security. We can use it for detecting
  17135. 10:11:45patterns, etc.
  17136. 10:11:47And last but not least, we have
  17137. 10:11:48interactive image generation.
  17138. 10:11:51So, GANs can be used to generate
  17139. 10:11:53interactive images as well.
  17140. 10:11:55And computer science and artificial
  17141. 10:11:56intelligence laboratory also known as
  17142. 10:11:58CSAIL
  17143. 10:11:59has developed a GAN that can generate 3D
  17144. 10:12:02models with realistic lighting and
  17145. 10:12:04reflections enabled by the shape and
  17146. 10:12:06texture editing.
  17147. 10:12:08And more recently, researchers have come
  17148. 10:12:10up with a model that can synthesize a
  17149. 10:12:13re-enacted face animated by a person's
  17150. 10:12:16movement while preserving the appearance
  17151. 10:12:19of the face at the same time.
  17152. 10:12:21There are a lot more applications we can
  17153. 10:12:23work on.
  17154. 10:12:25>> [music]
  17155. 10:12:30>> Evolution of AI. So, AI as we know today
  17156. 10:12:33is entirely different from where it
  17157. 10:12:34started. Back in the 18th century, it
  17158. 10:12:37was entirely based on myths,
  17159. 10:12:39speculations, and fiction. But, it did
  17160. 10:12:41start taking shape and the real
  17161. 10:12:43initiation in its truest essence took
  17162. 10:12:45place in 1956.
  17163. 10:12:47The AI search began with six major
  17164. 10:12:49design goals. The first one was teach
  17165. 10:12:51the machines to reason in accordance to
  17166. 10:12:53perform sophisticated mental tasks like
  17167. 10:12:55playing chess, providing mathematical
  17168. 10:12:57theorems, and others. The second one is
  17169. 10:13:00knowledge representation for machines to
  17170. 10:13:02interact with the real world as humans
  17171. 10:13:04do. Like machines needed to be able to
  17172. 10:13:06identify objects, people, and languages.
  17173. 10:13:09Programming language Lisp was developed
  17174. 10:13:11for this very purpose.
  17175. 10:13:13The third one is teach the machines to
  17176. 10:13:15plan and navigate around the world we
  17177. 10:13:17live in. With this, machines could
  17178. 10:13:18autonomously move around by navigating
  17179. 10:13:21themselves.
  17180. 10:13:22The fourth one is enable the machines to
  17181. 10:13:24process natural language so that they
  17182. 10:13:26can understand the language,
  17183. 10:13:28conversations, and the context of
  17184. 10:13:29speech.
  17185. 10:13:31The fifth one is train the machines to
  17186. 10:13:33perceive the way humans do, like touch,
  17187. 10:13:36feel, sight, hearing, and taste.
  17188. 10:13:38And general intelligence that included
  17189. 10:13:40emotional intelligence, intuition, and
  17190. 10:13:42creativity was the sixth point. Talking
  17191. 10:13:44about machine learning,
  17192. 10:13:46machine learning as we know can be
  17193. 10:13:47remotely explained with the evolution of
  17194. 10:13:49robots in the past years. Although,
  17195. 10:13:51machine learning isn't just a machine
  17196. 10:13:53that is going to learn stuff. It has a
  17197. 10:13:55lot more to it. Basically, we have data
  17198. 10:13:57at our bay. We train and test the model,
  17199. 10:14:00which in this case can be a robot, and
  17200. 10:14:02then make it to do task relevant to the
  17201. 10:14:04learning. And then again, learning can
  17202. 10:14:05be of different types, which is
  17203. 10:14:07supervised, unsupervised, reinforcement,
  17204. 10:14:09etc. To know more about machine learning
  17205. 10:14:11in detail, refer to our machine learning
  17206. 10:14:13full course tutorial to get on speed.
  17207. 10:14:15Now, let us go ahead and take a look at
  17208. 10:14:17what exactly is AI and machine learning.
  17209. 10:14:20So, what exactly is AI? According to the
  17210. 10:14:22Merriam-Webster dictionary, artificial
  17211. 10:14:24intelligence is a branch of computer
  17212. 10:14:25science dealing with the simulation of
  17213. 10:14:27intelligent behavior in computers.
  17214. 10:14:30AI is a technique that enables machines
  17215. 10:14:32to mimic human behavior. Artificial
  17216. 10:14:35intelligence is the theory and
  17217. 10:14:36development of computer systems able to
  17218. 10:14:38perform task normally requiring human
  17219. 10:14:40intelligence, such as visual perception,
  17220. 10:14:43speech recognition, decision-making, and
  17221. 10:14:45translation between languages.
  17222. 10:14:47If you ask me, AI is the simulation of
  17223. 10:14:49human intelligence done by machines
  17224. 10:14:51programmed by us. The machines need to
  17225. 10:14:54learn how to reason and do some
  17226. 10:14:56self-correction as needed along the way.
  17227. 10:14:58And artificial intelligence is
  17228. 10:14:59accomplished by studying how human brain
  17229. 10:15:02thinks, learns, and decide to work while
  17230. 10:15:04trying to solve a problem. And then
  17231. 10:15:06using the outcomes of this study as a
  17232. 10:15:07basis of developing intelligent software
  17233. 10:15:09and systems.
  17234. 10:15:11Now, let us go ahead and take a look at
  17235. 10:15:12what exactly is machine learning. So,
  17236. 10:15:14machine learning is a concept which
  17237. 10:15:16allows the machines to learn from
  17238. 10:15:18examples and experiences.
  17239. 10:15:20And that too without being explicitly
  17240. 10:15:22programmed. So, instead of you writing
  17241. 10:15:24the code, what you do is you feed the
  17242. 10:15:26data to the generic algorithm and the
  17243. 10:15:29algorithm or the machine builds the
  17244. 10:15:30logic based on the given data.
  17245. 10:15:32Machine learning algorithms are an
  17246. 10:15:34evolution of normal algorithms and they
  17247. 10:15:36make your program smarter by allowing
  17248. 10:15:38them to automatically learn from the
  17249. 10:15:40data that you provide. Now that we know
  17250. 10:15:42what AI and ML actually is, let us go
  17251. 10:15:44ahead and take a look at a few
  17252. 10:15:45applications of AI and ML today.
  17253. 10:15:47So, AI and ML can be widely used in so
  17254. 10:15:49many applications and I have listed down
  17255. 10:15:52a few applications to give you a wider
  17256. 10:15:53perspective. So, first of all, I'm going
  17257. 10:15:55to talk about the usage of AI and ML in
  17258. 10:15:57healthcare.
  17259. 10:15:58AI and ML in healthcare is an angel in
  17260. 10:16:01disguise. To understand this, imagine
  17261. 10:16:03you have the past data of millions of
  17262. 10:16:05patients with diseases in the past. Now,
  17263. 10:16:07all this data can be put to an effective
  17264. 10:16:10use in a sense that we would be able to
  17265. 10:16:12detect a disease in early stages with
  17266. 10:16:15the help of machine learning and
  17267. 10:16:16artificial intelligence algorithms.
  17268. 10:16:18And since the numbers never lie, we can
  17269. 10:16:21be pretty sure about the accuracy of the
  17270. 10:16:22results. Even so, if we have any doubts,
  17271. 10:16:25we can always check the accuracy in
  17272. 10:16:27almost all the cases.
  17273. 10:16:29Now, let's talk about the usage of AI
  17274. 10:16:31and ML in finance. So, artificial
  17275. 10:16:33intelligence in finance is transforming
  17276. 10:16:35the way we interact with money.
  17277. 10:16:37AI is helping the financial industry to
  17278. 10:16:39streamline and optimize processes
  17279. 10:16:42ranging from credit decisions to
  17280. 10:16:44quantitative trading and financial risk
  17281. 10:16:46management. And if we have that figured
  17282. 10:16:48out, it saves us from a lot of bad and
  17283. 10:16:51risky decisions.
  17284. 10:16:53Now, let us go ahead and take a look at
  17285. 10:16:54object detection in which we use AI and
  17286. 10:16:56ML. So, object detection today is
  17287. 10:16:58playing an important part in the IT
  17288. 10:17:00industry. For example, surveillance has
  17289. 10:17:02never looked more tech-savvy. Google
  17290. 10:17:04Lens is one example that uses image
  17291. 10:17:06recognition to identify the images on
  17292. 10:17:08the camera in real time.
  17293. 10:17:09Now, let us go ahead and take a look at
  17294. 10:17:11AI and ML in risk detection and
  17295. 10:17:13predictive analysis. So, predictive
  17296. 10:17:15analysis has proven its metal in the
  17297. 10:17:16industry already with almost every
  17298. 10:17:18organization using it to derive
  17299. 10:17:19conclusions based on previous data. So,
  17300. 10:17:22one example is how sports franchises,
  17301. 10:17:25such as a cricket team, would take
  17302. 10:17:26account of the performance of a player,
  17303. 10:17:28let's say a batsman, facing the
  17304. 10:17:30deliveries of bouncer and short ball
  17305. 10:17:32deliveries. So, they will be able to
  17306. 10:17:34figure out the best possible outcomes
  17307. 10:17:36based on the short selection and
  17308. 10:17:38strategies using the artificial
  17309. 10:17:40intelligence and machine learning
  17310. 10:17:41algorithms.
  17311. 10:17:42So, this is one way we can use
  17312. 10:17:43predictive analysis or engine it's just
  17313. 10:17:45one example. We can use it for many
  17314. 10:17:47purposes. Like, we can use it to predict
  17315. 10:17:49stock prices that we can do in finance
  17316. 10:17:51and we can use it to predict the weather
  17317. 10:17:53based on the hundreds and hundreds of
  17318. 10:17:55years of data that we already have.
  17319. 10:17:57Now, let's talk about AI and ML in
  17320. 10:17:58marketing and advertising. So, marketing
  17321. 10:18:01and advertising industry is the most
  17322. 10:18:03benefited with the evolution of AI and
  17323. 10:18:05ML. They are able to recognize the
  17324. 10:18:07browsing patterns of users through data
  17325. 10:18:09and target users with specific content
  17326. 10:18:11on the internet. For example, to reach
  17327. 10:18:13out in a subtle way, you only see those
  17328. 10:18:15ads which interest you. And you must
  17329. 10:18:17have felt sometimes like you keep
  17330. 10:18:19getting ads or you know the content on
  17331. 10:18:21the internet that you talk about or you
  17332. 10:18:22were talking about or you were thinking
  17333. 10:18:24about. It's basically nothing but AI and
  17334. 10:18:26machine learning that is learning a
  17335. 10:18:28pattern through the data and the your
  17336. 10:18:30browsing history or your browsing
  17337. 10:18:31pattern and reaches out to you in the
  17338. 10:18:33form of targeted marketing.
  17339. 10:18:35And now that we have talked about the
  17340. 10:18:36current applications, let us talk about
  17341. 10:18:38how AI and ML will shape up in the
  17342. 10:18:40future and how it would look like 10
  17343. 10:18:42years, maybe 20 years from now.
  17344. 10:18:44So, we cannot be sure about if AI and ML
  17345. 10:18:45would shape in the future like we have
  17346. 10:18:47seen in the movies, but for now, we will
  17347. 10:18:49stick to the realistic possibilities.
  17348. 10:18:51Although when I say realistic
  17349. 10:18:52possibilities, we are not really sure
  17350. 10:18:54how it would look like, but we can take
  17351. 10:18:56a guess.
  17352. 10:18:57So, when we talk about future of health
  17353. 10:18:59care, health care would seem pretty
  17354. 10:19:01reachable and advanced. Even now,
  17355. 10:19:03researchers are working on detecting
  17356. 10:19:05diseases in the early stages based on
  17357. 10:19:07the lifestyle and other relevant data.
  17358. 10:19:09And in the future, we can expect more
  17359. 10:19:10advancements in the psychological part
  17360. 10:19:12as well, where we will be able to
  17361. 10:19:14identify traits and warnings in the
  17362. 10:19:16early stages and work on it before it
  17363. 10:19:18gets better off us.
  17364. 10:19:20Now imagine being able to cure a disease
  17365. 10:19:22before even getting the hint of it. That
  17366. 10:19:24is what researchers are aiming for and
  17367. 10:19:26it looks pretty promising, guys. Let me
  17368. 10:19:28tell you. Now let's talk about the
  17369. 10:19:29future of self-driving cars. So
  17370. 10:19:31self-driving cars looks like a dream
  17371. 10:19:33come true today. But in the coming
  17372. 10:19:35decades, we're going to see a lot of
  17373. 10:19:37developments in the self-driving cars.
  17374. 10:19:39We would have overcome all the
  17375. 10:19:40challenges that we face today and I'm
  17376. 10:19:42not saying we will have levitating cars
  17377. 10:19:44driving you to your destinations, but it
  17378. 10:19:47won't be less than a fascinating
  17379. 10:19:48experience that may look like a dream
  17380. 10:19:49today or you might have seen in the
  17381. 10:19:50movies.
  17382. 10:19:52Now let's talk about the AI ML in
  17383. 10:19:53manufacturing that would take the
  17384. 10:19:55future.
  17385. 10:19:56So robots in manufacturing is one thing
  17386. 10:19:58that is going to change the future for
  17387. 10:20:00us.
  17388. 10:20:01Manufacturing industries will have the
  17389. 10:20:02best ever workforce and I'm not talking
  17390. 10:20:05about the human aspect of it. The
  17391. 10:20:06manufacturing would be so much easier
  17392. 10:20:08with the robots and since they don't get
  17393. 10:20:10tired and they won't even ask for leaves
  17394. 10:20:13or they might. We We never know. And AI
  17395. 10:20:15and robotics is a risky slope, although,
  17396. 10:20:18but that is not entirely true.
  17397. 10:20:19Researchers and experts are working day
  17398. 10:20:22and night to make it as safe as
  17399. 10:20:24possible.
  17400. 10:20:25And then again, let me talk about future
  17401. 10:20:26of finance with AI and ML. So managing
  17402. 10:20:29finance and risk detection would become
  17403. 10:20:31a piece of cake and to understand this
  17404. 10:20:33in layman terms, you will be able to do
  17405. 10:20:35your taxes without even lifting a pen.
  17406. 10:20:37And the fraud detection and trading
  17407. 10:20:38would become a lot easier and
  17408. 10:20:40accessible.
  17409. 10:20:41Financial advisory would take a much
  17410. 10:20:43advanced shape as we are already seeing
  17411. 10:20:45it with a lot of trading applications in
  17412. 10:20:47the market.
  17413. 10:20:48Then again, we have computer vision
  17414. 10:20:49which is going to change the future for
  17415. 10:20:50us. And I'm sure most of you are aware
  17416. 10:20:52of the concept of God's eye that we have
  17417. 10:20:54already seen in the movies. Although it
  17418. 10:20:56is fictional, but not sure for the wrong
  17419. 10:20:57reasons, but for the greater good, this
  17420. 10:20:59might be the possibility in the coming
  17421. 10:21:01years as computer vision has started to
  17422. 10:21:03overcome a lot of challenges in the
  17423. 10:21:04real-time image recognition.
  17424. 10:21:06And then we have the future of NLP
  17425. 10:21:08natural language processing and it's
  17426. 10:21:10going to you know it would open a lot of
  17427. 10:21:12linguistic barriers in the
  17428. 10:21:13conversational AI. The conversational AI
  17429. 10:21:15that we see today is limited to certain
  17430. 10:21:17tasks, but in coming years it could be
  17431. 10:21:19like a personal assistant or even a life
  17432. 10:21:21guide as well.
  17433. 10:21:23And with the recent advancements we are
  17434. 10:21:24aiming for a very flexible interface
  17435. 10:21:27that is going to work for everyone with
  17436. 10:21:29any linguistic experience or any
  17437. 10:21:31linguistic expectations.
  17438. 10:21:34And then we have a rather fictional
  17439. 10:21:36concept that I'm going to talk about
  17440. 10:21:37which is immortality through AI. So
  17441. 10:21:39there are scientists and researchers who
  17442. 10:21:41are trying to figure out a way to map
  17443. 10:21:43the brain simulation on a computer. So
  17444. 10:21:45immortality isn't just living until the
  17445. 10:21:47very eternity, it is in my opinion
  17446. 10:21:49leaving a legacy and but in hindsight
  17447. 10:21:52this task is pretty impossible, but we
  17448. 10:21:53never know in the future this might be a
  17449. 10:21:55possibility and we will be able to live
  17450. 10:21:57through a computer where people would
  17451. 10:21:59have figured out a way to shift all of
  17452. 10:22:01our brain simulations onto a computer
  17453. 10:22:04and we'll be able to think on its own
  17454. 10:22:05like our own very image. So that is one
  17455. 10:22:08possibility with the future in AI and ML
  17456. 10:22:11that many researchers are actually
  17457. 10:22:13aiming for and we might as well get
  17458. 10:22:15through with it. So hang in there guys
  17459. 10:22:18and now let me just talk about a key
  17460. 10:22:20skills of AI and ML specialist. There
  17461. 10:22:22was a lot of applications that I just
  17462. 10:22:24told you about. Now let's take a look at
  17463. 10:22:25what are the skill sets that are
  17464. 10:22:27required to become a AI ML specialist in
  17465. 10:22:29today's world.
  17466. 10:22:30So first of all you have to be familiar
  17467. 10:22:31with programming in Python
  17468. 10:22:33and you must be very well aware of maths
  17469. 10:22:35and statistics as well because it needs
  17470. 10:22:36a lot of logic to build algorithms and
  17471. 10:22:38understand them and there's a lot of
  17472. 10:22:40applied mathematics behind it as well.
  17473. 10:22:42So you have to be familiar with maths
  17474. 10:22:43and statistics and programming language
  17475. 10:22:45which is Python and the versioning tools
  17476. 10:22:47like TensorFlow, Keras, PyTorch, etc.
  17477. 10:22:50And then you must have an expertise in
  17478. 10:22:51one of the following machine learning
  17479. 10:22:53domains which is image processing,
  17480. 10:22:55computer vision, language processing,
  17481. 10:22:56speech signal processing, etc. And then
  17482. 10:22:59you must have an advanced knowledge in
  17483. 10:23:00deep learning algorithms as well because
  17484. 10:23:02it is a very important aspect of AIML.
  17485. 10:23:04And there has to be an effective
  17486. 10:23:06communication skills and a you have to
  17487. 10:23:07be a problem solver because in
  17488. 10:23:09hindsight, if you get a problem, the
  17489. 10:23:11only thing that an employer seeks from
  17490. 10:23:13you is the solution of the problem. So,
  17491. 10:23:15anyways, you have to be a problem solver
  17492. 10:23:17in that skill set. And you must be
  17493. 10:23:20experienced in data visualization tools
  17494. 10:23:21and methods like Tableau, Matplotlib,
  17495. 10:23:23Power BI, etc. And there has to be a
  17496. 10:23:25knowledge in ensemble and online
  17497. 10:23:27learning. And you must have an
  17498. 10:23:29experience with SQL and other database
  17499. 10:23:31related languages. And you have to be
  17500. 10:23:33familiar with parallel computing using
  17501. 10:23:35GPUs and unique big systems. So, these
  17502. 10:23:37are the key skills that are required to
  17503. 10:23:39become an AIML specialist, guys.
  17504. 10:23:41Now, let me just walk you through the
  17505. 10:23:42market trends that looks pretty
  17506. 10:23:44promising for now. So, first of all, I'm
  17507. 10:23:46going to talk about the market trends
  17508. 10:23:47and it looks pretty solid, guys, with
  17509. 10:23:49the amount of data flowing in each year
  17510. 10:23:51that it is pretty obvious that it will
  17511. 10:23:53be opening a lot of doors for skilled
  17512. 10:23:54professionals. And the correct approach
  17513. 10:23:56is to get skilled since it is still in
  17514. 10:23:58the evolution phase and we have to scale
  17515. 10:24:01a lot more possibilities in these
  17516. 10:24:02domains. So, if you have the skill set
  17517. 10:24:04for it, it's going to be a very bright
  17518. 10:24:06future for you guys. And to be specific,
  17519. 10:24:08there are a lot of opportunities in
  17520. 10:24:09health care, finance, conversational AI
  17521. 10:24:12or we can call it chatbots, and object
  17522. 10:24:14detection, etc. Almost every industry
  17523. 10:24:16would move to automating their
  17524. 10:24:17processes. And what else than machine
  17525. 10:24:19learning and AI to do your job? So, it
  17526. 10:24:21is the best time to learn AI and ML if
  17527. 10:24:23you are looking for a bright career
  17528. 10:24:25right now. And let us go ahead and take
  17529. 10:24:27a look at the salary trends as well so
  17530. 10:24:28you get the perspective of how much
  17531. 10:24:29you're going to get paid. So, I have
  17532. 10:24:31categorized the salary trends in a few
  17533. 10:24:33job profiles in AI and ML. For a machine
  17534. 10:24:35learning engineer, the takeaway fruits
  17535. 10:24:37of your labor would look around $114,000
  17536. 10:24:40a year. And for a machine learning
  17537. 10:24:41scientist, the average salary looks
  17538. 10:24:43around $120,000 a year and can go as
  17539. 10:24:46high as $150,000 a year.
  17540. 10:24:48And for an AI engineer, the average
  17541. 10:24:50salary is around a $90,000, but it can
  17542. 10:24:53go as high as $140,000 a year as well.
  17543. 10:24:56And for an AI researcher, it goes from a
  17544. 10:24:58$125,000 average to as high as $150,000
  17545. 10:25:03a year. And now let us go ahead and take
  17546. 10:25:04a look at a few companies that are
  17547. 10:25:06hiring right now for AI and ML
  17548. 10:25:07specialists. And although there are a
  17549. 10:25:09lot more companies that are hiring for
  17550. 10:25:11AI and ML specialists right now, I've
  17551. 10:25:12just listed down a few over here. So, we
  17552. 10:25:14have Ford Motors, we have Capgemini,
  17553. 10:25:16Accenture, Dell, Deloitte, Google,
  17554. 10:25:18Amazon. And there are a lot of startups
  17555. 10:25:20as well which are actually artificial
  17556. 10:25:22intelligence and machine learning based.
  17557. 10:25:24So, there's a lot of opportunity for you
  17558. 10:25:26guys. And let me just tell you the best
  17559. 10:25:28approach to actually get a job in AI and
  17560. 10:25:30ML industry. So, the best approach to
  17561. 10:25:33find a job as an AI and ML specialist,
  17562. 10:25:35even if you are a beginner or an
  17563. 10:25:37experienced professional, this works for
  17564. 10:25:39everyone. So, first of all, you have to
  17565. 10:25:41start with a programming language,
  17566. 10:25:42preferably Python, because it works best
  17567. 10:25:44with AI and ML algorithms. And after you
  17568. 10:25:47are done mastering these basics, start
  17569. 10:25:49with machine learning and artificial
  17570. 10:25:50intelligence algorithms. And before
  17571. 10:25:52that, make sure you are sophisticated
  17572. 10:25:53enough to work with data. I mean, you
  17573. 10:25:55can analyze the data, clean it, prepare
  17574. 10:25:57it for model building, and etc. And try
  17575. 10:25:59to learn all of them with the
  17576. 10:26:00implementations on unique data instead
  17577. 10:26:03of the generic data that you find on the
  17578. 10:26:04internet. The next step would be to
  17579. 10:26:06learn the advanced concepts in AI and ML
  17580. 10:26:08like TensorFlow for object detection,
  17581. 10:26:10speech recognition, image processing,
  17582. 10:26:11etc. And after you have mastered the
  17583. 10:26:14versioning tools, you must make sure
  17584. 10:26:16that you have a credibility in order to
  17585. 10:26:17get a job. Because as an employer,
  17586. 10:26:20anyone would look for credible person
  17587. 10:26:22proficient enough to do the job. And
  17588. 10:26:24where will you get that? If you have a
  17589. 10:26:26relevant master's degree, it is well and
  17590. 10:26:28good. But if you don't have a degree,
  17591. 10:26:30you can always go for a certification,
  17592. 10:26:32which will give you the credibility, and
  17593. 10:26:34it is going to be the best option to
  17594. 10:26:35prove your mettle. And then there's one
  17595. 10:26:38way to look at it. I mean, you can take
  17596. 10:26:39up the Edureka's post graduate program
  17597. 10:26:41in artificial intelligence and machine
  17598. 10:26:43learning, which is going to be a a good
  17599. 10:26:45deal for you because it is an
  17600. 10:26:46affiliation with a top college and then
  17601. 10:26:49you are good to go. You can apply for
  17602. 10:26:51jobs and you'll get the job easily if
  17603. 10:26:52you have all the skill sets and you have
  17604. 10:26:54the experience in making relevant, you
  17605. 10:26:56know, you are working on real-time
  17606. 10:26:58projects as well. So, these are going to
  17607. 10:26:59be very useful for you.
  17608. 10:27:02>> [music]
  17609. 10:27:07>> Let's get started with our basic level
  17610. 10:27:09questions.
  17611. 10:27:10So, first we have what is the difference
  17612. 10:27:12between AI, machine learning, and deep
  17613. 10:27:14learning? I'm sure all of you have this
  17614. 10:27:16question at the top of your mind because
  17615. 10:27:18there's a huge confusion between AI,
  17616. 10:27:19machine learning, and deep learning. So,
  17617. 10:27:21let's try to understand how they are
  17618. 10:27:23different. Now, first of all, AI came
  17619. 10:27:25into existence at around 1950s. All
  17620. 10:27:28right, this was followed by machine
  17621. 10:27:29learning and then deep learning was
  17622. 10:27:31introduced. Now, AI basically represents
  17623. 10:27:34simulated intelligence in machines,
  17624. 10:27:36which means that it represents any robot
  17625. 10:27:39or any machine that can mimic the
  17626. 10:27:41behavior of a human being. Machine
  17627. 10:27:43learning on the other hand is a practice
  17628. 10:27:46of getting machines to make decisions
  17629. 10:27:48without being explicitly programmed to
  17630. 10:27:50do so. Now, if you don't program a
  17631. 10:27:52machine, how are you going to let it
  17632. 10:27:54make decisions? Now, the way machines
  17633. 10:27:57learn is through data. So, the most
  17634. 10:27:59important thing in machine learning is
  17635. 10:28:01the data. All right, you're going to
  17636. 10:28:02train machines using data so that they
  17637. 10:28:04can make their own decisions. Next, we
  17638. 10:28:06have deep learning. Now, deep learning
  17639. 10:28:08is basically the process of using
  17640. 10:28:10artificial neural networks to solve
  17641. 10:28:12complex problems. So, basically you can
  17642. 10:28:15think of deep learning as a field that
  17643. 10:28:17tries to mimic our brain. Okay, so how
  17644. 10:28:20we have neural networks in our brain,
  17645. 10:28:22that's exactly how deep learning uses
  17646. 10:28:24the concepts of artificial neural
  17647. 10:28:26networks in order to solve problems.
  17648. 10:28:28Now, AI is a subset of data science. So
  17649. 10:28:31guys, first of all, data science is the
  17650. 10:28:32process of deriving useful insights from
  17651. 10:28:35data. All right, it's a process of
  17652. 10:28:37extracting information from data that
  17653. 10:28:39will help you solve problems. So, AI is
  17654. 10:28:42a subset of data science. Now, on the
  17655. 10:28:44other hand, machine learning is a subset
  17656. 10:28:46of AI and data science because machine
  17657. 10:28:48learning comes after AI. So, basically
  17658. 10:28:51in AI, you're going to make use of
  17659. 10:28:53techniques and concepts of machine
  17660. 10:28:55learning in order to solve problems.
  17661. 10:28:57Then, we have deep learning. So, it's
  17662. 10:28:59sort of a hierarchy. First, we have data
  17663. 10:29:01science, then we have AI, then we have
  17664. 10:29:03machine learning, and then we have deep
  17665. 10:29:04learning. Deep learning is a subset of
  17666. 10:29:06machine learning, AI, and data science.
  17667. 10:29:09Okay, I hope this is clear. Now, the
  17668. 10:29:11main aim of artificial intelligence is
  17669. 10:29:13to build machines in such a way that
  17670. 10:29:15they're capable of thinking like human
  17671. 10:29:17beings. All right, so basically, they
  17672. 10:29:19must be able to mimic the behavior of a
  17673. 10:29:22human being. Now, the aim of machine
  17674. 10:29:24learning on the other hand is to make
  17675. 10:29:26machines learn by providing them a lot
  17676. 10:29:28of data. Okay, once you make a machine
  17677. 10:29:30learn through data, it's going to be
  17678. 10:29:32able to solve complex problems and find
  17679. 10:29:34solutions. Now, the aim of deep learning
  17680. 10:29:37is to build neural networks that are
  17681. 10:29:40able to solve more advanced and complex
  17682. 10:29:42problems. Okay, now like I mentioned,
  17683. 10:29:44deep learning is like an artificial
  17684. 10:29:47brain. All right, you're basically
  17685. 10:29:48building an artificial brain that is
  17686. 10:29:50able to think exactly like how we do.
  17687. 10:29:53Okay, that's what deep learning is. It's
  17688. 10:29:55a little more advanced than machine
  17689. 10:29:57learning. Now guys, in short, AI,
  17690. 10:29:59machine learning, and deep learning are
  17691. 10:30:01used to solve problems through data. So,
  17692. 10:30:03basically, AI makes use of techniques
  17693. 10:30:06and methods of machine learning and deep
  17694. 10:30:08learning to solve problems or to draw
  17695. 10:30:10useful insights from data. So, this is
  17696. 10:30:13the difference between AI, machine
  17697. 10:30:14learning, and deep learning. I hope all
  17698. 10:30:16of you are clear with this. Now, let's
  17699. 10:30:18look at our question number two. The
  17700. 10:30:20question is, "What is artificial
  17701. 10:30:22intelligence? Give an example of where
  17702. 10:30:25AI is used on a daily basis." So, there
  17703. 10:30:27are a lot of definitions of AI on the
  17704. 10:30:29internet. A few of them are, "Artificial
  17705. 10:30:32intelligence is an area of computer
  17706. 10:30:34science that emphasizes on the creation
  17707. 10:30:37of intelligent machines that work and
  17708. 10:30:39react like humans. So, like I said,
  17709. 10:30:41basically, a machine that is able to
  17710. 10:30:43mimic the behavior of a human being is
  17711. 10:30:46known as artificial intelligence.
  17712. 10:30:48Another such definition is the
  17713. 10:30:49capability of a machine to imitate the
  17714. 10:30:51intelligent human behavior. All right?
  17715. 10:30:54So, artificial intelligence, in short,
  17716. 10:30:56is basically a machine that we created
  17717. 10:30:58who can act and think like a human
  17718. 10:31:00being. Now, where do you think AI is
  17719. 10:31:02used on a daily basis? There are tons of
  17720. 10:31:05applications that make use of AI, but
  17721. 10:31:07one of the most popular applications of
  17722. 10:31:09AI is a Google search engine. Now, if
  17723. 10:31:12you just open up Google search and you
  17724. 10:31:13start typing anything, immediately you
  17725. 10:31:15get recommendations. These
  17726. 10:31:17recommendations you derive by using
  17727. 10:31:19machine learning algorithms, by using
  17728. 10:31:21deep neural networks, and so on. So, on
  17729. 10:31:24the top of my head, the most general
  17730. 10:31:25example of AI is a Google search engine.
  17731. 10:31:28All of us use Google search engine, and
  17732. 10:31:30we know how quick it is with its results
  17733. 10:31:32and how relevant searches it gives us.
  17734. 10:31:35All this is because of AI. All right?
  17735. 10:31:38Now, let's look at our next question,
  17736. 10:31:39which states, "What are the different
  17737. 10:31:41types of AI?" Now, a lot of people might
  17738. 10:31:43not be aware of this because there are a
  17739. 10:31:46couple of types of AI or a couple of
  17740. 10:31:48types of machines which are
  17741. 10:31:50hypothetical. Okay, we haven't actually
  17742. 10:31:52implemented these machines in the real
  17743. 10:31:54world. We just have a theoretical
  17744. 10:31:56definition of these. Okay, let's look at
  17745. 10:31:58what I'm talking about. So, first of
  17746. 10:32:00all, we have reactive machines AI. Now,
  17747. 10:32:03these machines are all based on the
  17748. 10:32:05present actions. Okay, they have no
  17749. 10:32:07memory or they have no concept of
  17750. 10:32:09storing memory so that they can learn
  17751. 10:32:11from that experience. They just react at
  17752. 10:32:14the moment. Okay, so they're based on
  17753. 10:32:15present actions, and they cannot use
  17754. 10:32:18previous experiences to form current
  17755. 10:32:20decisions and update their memory. Then,
  17756. 10:32:22we have limited memory AI. Now, this
  17757. 10:32:25type of AI has some temporary storage of
  17758. 10:32:27memory in it. Now, if we have some
  17759. 10:32:29memory stored in a machine, we know that
  17760. 10:32:31it can look back into the memory and it
  17761. 10:32:34can try to make decisions based on
  17762. 10:32:36previous or past experiences.
  17763. 10:32:38So, limited memory AI makes use of that
  17764. 10:32:40concept. We have temporary memory here.
  17765. 10:32:42We do not have permanent memory, but one
  17766. 10:32:45of the top applications of limited
  17767. 10:32:47memory AI is the self-driving cars. I'm
  17768. 10:32:50sure all of you have heard of
  17769. 10:32:51self-driving cars. They make use of
  17770. 10:32:53limited memory AI in order to run. Then
  17771. 10:32:56we have theory of mind AI. Now, like I
  17772. 10:32:59mentioned earlier, there are a couple of
  17773. 10:33:01types of artificial intelligent machines
  17774. 10:33:03which are not actually implemented in
  17775. 10:33:05the real world. An example of that is
  17776. 10:33:07theory of mind AI. Okay, this is
  17777. 10:33:09basically an advanced machine which will
  17778. 10:33:12have the ability to understand emotions,
  17779. 10:33:15people, and other things in the real
  17780. 10:33:16world. We might have come close to this
  17781. 10:33:19type of AI, but we haven't actually
  17782. 10:33:20developed something that can understand
  17783. 10:33:22emotions. Next, we have self-aware AI.
  17784. 10:33:25Now, this is another such example of a
  17785. 10:33:27machine that is not built in the real
  17786. 10:33:29world. This basically includes any
  17787. 10:33:31machine that has consciousness or that
  17788. 10:33:34can react just like a human being. Okay,
  17789. 10:33:36so basically a machine that can take own
  17790. 10:33:38decisions, that can form own
  17791. 10:33:40conclusions, and these are machines that
  17792. 10:33:42have the capability of making their own
  17793. 10:33:44decisions without any human
  17794. 10:33:45intervention. Now, this kind of AI is
  17795. 10:33:47not developed, like I mentioned, because
  17796. 10:33:49it's going to take up a lot of resources
  17797. 10:33:51and we still haven't reached that peak
  17798. 10:33:54of evolution yet. Then we have
  17799. 10:33:56artificial narrow intelligence. Now,
  17800. 10:33:58these are the general purpose AI that we
  17801. 10:34:00see on a daily basis. I'm sure all of
  17802. 10:34:03you have used Google Assistant, you've
  17803. 10:34:05used Siri. All of that comes under
  17804. 10:34:07artificial narrow intelligence. After
  17805. 10:34:09that, we have artificial general
  17806. 10:34:11intelligence. Now, these are a little
  17807. 10:34:13more advanced than the artificial narrow
  17808. 10:34:15intelligence.
  17809. 10:34:16Then we have artificial superhuman
  17810. 10:34:18intelligence. Now, these are one of the
  17811. 10:34:21most advanced type of AIs that are
  17812. 10:34:23there. Now, like I mentioned earlier,
  17813. 10:34:25there are a couple of types of
  17814. 10:34:27artificial intelligent machines which
  17815. 10:34:29are not actually implemented in the real
  17816. 10:34:31world. An example of that is artificial
  17817. 10:34:34super human intelligence. So guys, these
  17818. 10:34:36were the different types of AI. Now,
  17819. 10:34:39let's look at the next question which
  17820. 10:34:41says explain the different domains of
  17821. 10:34:43artificial intelligence. Now, AI covers
  17822. 10:34:45a lot of different domains starting with
  17823. 10:34:48machine learning. Okay, so machine
  17824. 10:34:50learning like I mentioned earlier is the
  17825. 10:34:52science of getting computers to act by
  17826. 10:34:54feeding them data and by letting them
  17827. 10:34:56learn a few tricks on their own without
  17828. 10:34:59being programmed to do so. Okay, so
  17829. 10:35:01you're not explicitly programming the
  17830. 10:35:03machine, instead you're feeding it a lot
  17831. 10:35:04of data so that it understands the data
  17832. 10:35:07and it makes its own decisions. Then we
  17833. 10:35:09have neural networks. Now, neural
  17834. 10:35:11networks are basically a set of
  17835. 10:35:13algorithms or you can say a set of
  17836. 10:35:14techniques which are modeled in
  17837. 10:35:17accordance with a human brain. Okay,
  17838. 10:35:19like I mentioned earlier, deep learning
  17839. 10:35:21or neural networks is almost the same
  17840. 10:35:23thing. Deep learning makes use of neural
  17841. 10:35:25networks in order to solve complex
  17842. 10:35:27problems. Now, we have robotics. Now,
  17843. 10:35:29robotics is a subset of AI which
  17844. 10:35:32includes different branches and
  17845. 10:35:33applications of robots. These robots are
  17846. 10:35:36basically artificial agents which act in
  17847. 10:35:39a real world environment. Okay, so an AI
  17848. 10:35:41robot works by manipulating the objects
  17849. 10:35:43in its surrounding by perceiving,
  17850. 10:35:45moving, and taking relevant actions.
  17851. 10:35:48Then we have expert systems. Now, an
  17852. 10:35:50expert system is basically a computer
  17853. 10:35:52system that mimics the decision-making
  17854. 10:35:54ability of a human being. Now, I know
  17855. 10:35:56all of these domains sound very similar,
  17856. 10:35:59but they have a very different approach
  17857. 10:36:01with which they solve the problem. All
  17858. 10:36:03right, that's the main difference
  17859. 10:36:04between these domains. Next, we have
  17860. 10:36:06fuzzy logic systems. Now, traditional
  17861. 10:36:09systems usually give out output in the
  17862. 10:36:11form of binary. So, usually if you feed
  17863. 10:36:14something to a machine, it's always in
  17864. 10:36:15the binary form. The output is also
  17865. 10:36:17usually in the form of yes, no, true,
  17866. 10:36:19false, and so on. But when it comes to
  17867. 10:36:21fuzzy logic, it tries to give an output
  17868. 10:36:23in the form of degrees of truth. Okay,
  17869. 10:36:26so it's very different when compared to
  17870. 10:36:28the traditional computer systems or the
  17871. 10:36:29traditional programs. Next, we have
  17872. 10:36:32natural language processing. Now, this
  17873. 10:36:34is a field of AI that analyzes natural
  17874. 10:36:37human language to derive useful insights
  17875. 10:36:40so that it can solve problems. Now, NLP
  17876. 10:36:42is used majorly in social media
  17877. 10:36:44platforms. So, Twitter sentimental
  17878. 10:36:46analysis is done via NLP. Even Facebook
  17879. 10:36:49uses NLP in a lot of things. All right,
  17880. 10:36:52so NLP, fuzzy logic, expert systems,
  17881. 10:36:54machine learning, neural networks, and
  17882. 10:36:56robotics are the different domains of
  17883. 10:36:58AI. I hope all of you are clear with the
  17884. 10:37:00domains. Now, let's look at our next
  17885. 10:37:03question. Okay, so how is machine
  17886. 10:37:05learning related to artificial
  17887. 10:37:07intelligence? There is a huge confusion
  17888. 10:37:10between machine learning and AI. A lot
  17889. 10:37:12of people tend to believe that AI and
  17890. 10:37:13machine learning is one in the same
  17891. 10:37:15thing. All right, I would say that you
  17892. 10:37:17cannot compare AI and machine learning
  17893. 10:37:19because machine learning is a subset of
  17894. 10:37:21AI. So, basically, AI makes use of
  17895. 10:37:24machine learning algorithms and machine
  17896. 10:37:26learning concepts to solve problems.
  17897. 10:37:28That's the basic difference or that is
  17898. 10:37:30where the confusion ends. Machine
  17899. 10:37:32learning is a technique which is
  17900. 10:37:34implemented in artificial intelligence
  17901. 10:37:36in order to solve problems. I hope this
  17902. 10:37:39is clear. Now, let's look at what are
  17903. 10:37:41the different types of machine learning.
  17904. 10:37:43So, there are three types of machine
  17905. 10:37:45learning. We have supervised,
  17906. 10:37:46unsupervised, and reinforcement
  17907. 10:37:48learning. Now, supervised learning is
  17908. 10:37:51the type of learning in which the
  17909. 10:37:52machine learns by using labeled data.
  17910. 10:37:55Now, to make you understand, let's look
  17911. 10:37:56at an example. Okay, let's say that
  17912. 10:37:58you've input images of apples and
  17913. 10:38:01oranges to your machine and you've
  17914. 10:38:03labeled them. You've told the machine
  17915. 10:38:05like, "Listen, this is the apple, this
  17916. 10:38:07is an orange, and the output should also
  17917. 10:38:09look like this." Okay, so you're
  17918. 10:38:11labeling the input as apple and an
  17919. 10:38:13orange, and then you're asking the
  17920. 10:38:15machine to output an apple and an
  17921. 10:38:17orange. But, when it comes to
  17922. 10:38:18unsupervised learning, you're not going
  17923. 10:38:20to label them. You're just going to give
  17924. 10:38:22them images of apple and oranges, and it
  17925. 10:38:24has to figure out on its own. It has to
  17926. 10:38:26try and understand the difference
  17927. 10:38:28between apple and oranges, try and
  17928. 10:38:30understand how they look different, or
  17929. 10:38:32how they have a different color. So,
  17930. 10:38:34basically in unsupervised learning, you
  17931. 10:38:35don't have a labeled data set. Okay,
  17932. 10:38:37you're going to give it an unlabeled
  17933. 10:38:38data set, and you're going to ask it to
  17934. 10:38:40find out and classify which is an apple
  17935. 10:38:43and which is an orange. Okay, that's the
  17936. 10:38:44difference between supervised and
  17937. 10:38:46unsupervised. Now, reinforcement
  17938. 10:38:47learning is comparatively different.
  17939. 10:38:50Let's imagine that you were put off in
  17940. 10:38:52an island. Okay, let's say that you were
  17941. 10:38:53left in an isolated island. What would
  17942. 10:38:56you do? Now, initially, we'll all panic,
  17943. 10:38:58and we won't know what to do. But, after
  17944. 10:39:00a point, you'll start exploring the
  17945. 10:39:02island. You'll start adapting to the
  17946. 10:39:04change in the climate conditions, you'll
  17947. 10:39:06start looking for food, and then you'll
  17948. 10:39:08try and understand which food is right
  17949. 10:39:09for you and which food is wrong for you.
  17950. 10:39:11You know, you'll learn from your
  17951. 10:39:12experience. So, in reinforcement
  17952. 10:39:15learning, basically, an agent interacts
  17953. 10:39:17with its environment by producing
  17954. 10:39:19actions and discovers errors or rewards.
  17955. 10:39:22Now, the type of problems that
  17956. 10:39:23supervised learning is used to solve is
  17957. 10:39:25regression and classification. When it
  17958. 10:39:27comes to unsupervised, it is association
  17959. 10:39:29and clustering. And in reinforcement
  17960. 10:39:31learning, it's all the reward-based
  17961. 10:39:33problems. The type of data for
  17962. 10:39:35supervised learning is labeled data. For
  17963. 10:39:37unsupervised, it is unlabeled. And for
  17964. 10:39:39reinforcement, it is no predefined data.
  17965. 10:39:42Now, when I say no predefined data, I
  17966. 10:39:43mean that the reinforcement learning
  17967. 10:39:45agent has to start collecting the data.
  17968. 10:39:48So, basically, in reinforcement
  17969. 10:39:50learning, from data collection to model
  17970. 10:39:52evaluation, it does everything. In terms
  17971. 10:39:54of training, supervised learning
  17972. 10:39:56provides external supervision in the
  17973. 10:39:58form of labeled data set. In
  17974. 10:40:00unsupervised learning, there's no
  17975. 10:40:02supervision. That's why it's called
  17976. 10:40:03unsupervised learning. Again, in
  17977. 10:40:05reinforcement learning, there's no
  17978. 10:40:06supervision at all. The agent has to
  17979. 10:40:08figure everything out. Now, how
  17980. 10:40:10supervised learning works is uh you map
  17981. 10:40:13the labeled input to the known output.
  17982. 10:40:15So, basically you teach the machine like
  17983. 10:40:17you tell it that this is the input and
  17984. 10:40:19this has to be the output. When it comes
  17985. 10:40:21to unsupervised learning, you just
  17986. 10:40:22provide data to the machine and it has
  17987. 10:40:24to understand patterns and it has to
  17988. 10:40:26discover the output. Now, in
  17989. 10:40:28reinforcement learning, it has to follow
  17990. 10:40:30the trial and error method. Okay,
  17991. 10:40:32there's no particular way in which the
  17992. 10:40:35agent learns. It just has to explore the
  17993. 10:40:37environment, try out a few things, and
  17994. 10:40:39learn from that experience. Popular
  17995. 10:40:41supervised learning algorithms include
  17996. 10:40:43linear regression, logistic regression.
  17997. 10:40:46For unsupervised, we have K-means. And
  17998. 10:40:47for reinforcement learning, we have
  17999. 10:40:49Q-learning. So, guys, these were the
  18000. 10:40:51different types of machine learning and
  18001. 10:40:53I also discussed the difference between
  18002. 10:40:55the three. Now, let's move on and look
  18003. 10:40:57at our next question, which is what is
  18004. 10:40:59Q-learning?
  18005. 10:41:00In the previous slide itself, I told you
  18006. 10:41:02that a type of reinforcement learning
  18007. 10:41:03algorithm is Q-learning. So, basically,
  18008. 10:41:06here what happens is an agent tries to
  18009. 10:41:08learn the optimal policy from its past
  18010. 10:41:11experience with the environment. The
  18011. 10:41:13past experience of an agent are a
  18012. 10:41:15sequence of action, state, and rewards.
  18013. 10:41:18So, what happens is, first of all, you
  18014. 10:41:20take an agent and you put it in state
  18015. 10:41:22zero. Okay, let's say there's some state
  18016. 10:41:24known as state zero. Now, this agent is
  18017. 10:41:27going to perform some action A0. On
  18018. 10:41:30performing this action, it is going to
  18019. 10:41:32get a reward R1. And if it gets a reward
  18020. 10:41:35R1, then it's going to move to state S1.
  18021. 10:41:38But in case the action is wrong, then
  18022. 10:41:40it's going to get a negative reward, as
  18023. 10:41:43in some points are going to be reduced.
  18024. 10:41:45So, guys, think of Q-learning as a game.
  18025. 10:41:47You're in state zero, and then you do
  18026. 10:41:49some action, and either you get a reward
  18027. 10:41:51and go to the next state, or else you
  18028. 10:41:53lose and you go back to the same state.
  18029. 10:41:56So, until you learn, you're going to be
  18030. 10:41:57in the same state. But if you keep
  18031. 10:41:59learning and if you keep receiving
  18032. 10:42:01positive rewards, then you're going to
  18033. 10:42:02move on to state one, and similarly, you
  18034. 10:42:04move on to state two, three, and so on.
  18035. 10:42:06This is what Q learning is about.
  18036. 10:42:08Now, the next question is what is deep
  18037. 10:42:10learning? Now, deep learning, like I
  18038. 10:42:12mentioned earlier, basically mimics the
  18039. 10:42:15way our brain works. Okay, it learns
  18040. 10:42:17from experience. Now, the main concept
  18041. 10:42:19behind deep learning is neural networks.
  18042. 10:42:22In our brain also, we have neural
  18043. 10:42:23networks. [clears throat] So, what deep
  18044. 10:42:24learning tries to do is it tries to use
  18045. 10:42:26the concept of neural networks in order
  18046. 10:42:29to solve complex problems. So,
  18047. 10:42:30basically, we're trying to mimic our
  18048. 10:42:32brain. Any deep neural network will have
  18049. 10:42:35three types of layers. The first is the
  18050. 10:42:37input layer. Now, this layer will
  18051. 10:42:39basically receive all the input, and it
  18052. 10:42:41will forward them to the hidden layer.
  18053. 10:42:43Now, in the hidden layer, all the
  18054. 10:42:45analysis and the computation takes
  18055. 10:42:47place. All right, once the computation
  18056. 10:42:49is done, the result is transferred to
  18057. 10:42:51the output layer. Now, there can be n
  18058. 10:42:53number of hidden layers depending on the
  18059. 10:42:55type of problem you're trying to solve.
  18060. 10:42:57Then, we have the output layer. So,
  18061. 10:42:59basically, this layer is responsible for
  18062. 10:43:01transferring the information from the
  18063. 10:43:03neural network to the outside world. So,
  18064. 10:43:05it's as simple as that. It's pretty
  18065. 10:43:07obvious. Input layer will take in the
  18066. 10:43:09input, hidden layer will perform the
  18067. 10:43:11computations, and the output layer will
  18068. 10:43:13give out the output. This is a small
  18069. 10:43:15explanation of what deep learning is.
  18070. 10:43:17Now, of course, this is much more
  18071. 10:43:18complex than this, but in short, this is
  18072. 10:43:21exactly what deep learning is. Now,
  18073. 10:43:23let's look at our next question, which
  18074. 10:43:25is explain how deep learning works. So,
  18075. 10:43:28basically, deep learning is a concept
  18076. 10:43:30based on something known as neuron.
  18077. 10:43:33Okay, neuron is a basic unit of the
  18078. 10:43:35brain. Inspired from this neuron, they
  18079. 10:43:37came up with something known as
  18080. 10:43:39perceptrons or artificial neurons. Now,
  18081. 10:43:42in this image on the left-hand side, you
  18082. 10:43:44can see that there is something known as
  18083. 10:43:46dendrite. These are modules which
  18084. 10:43:48receive the input. It basically receives
  18085. 10:43:50all the signals that we send to our
  18086. 10:43:52brain. Okay, similar to the dendrites
  18087. 10:43:55are the input layer in our artificial
  18088. 10:43:57neural networks. Now, in the previous
  18089. 10:43:59slide, we discussed that the input layer
  18090. 10:44:00takes in all the input from the outside.
  18091. 10:44:03That's exactly what a dendrite does. So,
  18092. 10:44:05basically a perceptron receives multiple
  18093. 10:44:08inputs. It applies various
  18094. 10:44:10transformations and functions, and then
  18095. 10:44:12it provides an output.
  18096. 10:44:14So, basically guys, just like how our
  18097. 10:44:16brain contains multiple connected
  18098. 10:44:18neurons called neural networks, we also
  18099. 10:44:21have a network of artificial neurons
  18100. 10:44:23called perceptrons to form a deep neural
  18101. 10:44:25network. So, basically an artificial
  18102. 10:44:27neuron or a perceptron, it models a
  18103. 10:44:30neuron which has a set of inputs, each
  18104. 10:44:33of which is assigned some specific
  18105. 10:44:35weight. Okay, all of these inputs will
  18106. 10:44:37have a specific weight, and the neuron
  18107. 10:44:39will compute some function on these
  18108. 10:44:41weighted inputs and give you the output.
  18109. 10:44:44So, the neuron will basically perform
  18110. 10:44:45analysis and all of that on these
  18111. 10:44:47weighted inputs to give you some output.
  18112. 10:44:50This is a basic concept of deep
  18113. 10:44:52learning. So, there are inputs which
  18114. 10:44:54have some weight on it, and these inputs
  18115. 10:44:56are then formulated and analyzed in
  18116. 10:44:59order to give you an output. Now, let's
  18117. 10:45:01look at our next question, which is
  18118. 10:45:03explain the commonly used artificial
  18119. 10:45:05neural networks. Now, this is a very
  18120. 10:45:07theoretical question because in order to
  18121. 10:45:09make you understand how each of them
  18122. 10:45:11work will take a lot of time. Okay, so
  18123. 10:45:13I'm just going to briefly tell you what
  18124. 10:45:15each of these networks are and what they
  18125. 10:45:17do. Now, feedforward neural network is
  18126. 10:45:19the most basic kind of artificial neural
  18127. 10:45:21network. So, basically the feedforward
  18128. 10:45:23neural network is unidirectional. The
  18129. 10:45:26data passes through the input nodes and
  18130. 10:45:28leaves through the output nodes. In
  18131. 10:45:30feedforward neural network, usually the
  18132. 10:45:32number of hidden layers depends on the
  18133. 10:45:34complexity of the problem. Coming to
  18134. 10:45:36convolutional neural networks, here
  18135. 10:45:39basically the input features are taken
  18136. 10:45:42in small sets. Okay, or they're taken in
  18137. 10:45:44batches. This will help the network
  18138. 10:45:47remember better because you're feeding
  18139. 10:45:49batches of images or you're feeding
  18140. 10:45:51batches of input to the neural network.
  18141. 10:45:54Now, this type of neural network is
  18142. 10:45:56mainly used for signal and image
  18143. 10:45:57processing. Next, we have recurrent
  18144. 10:46:00neural networks. These are also known as
  18145. 10:46:03long short-term memory networks. So,
  18146. 10:46:05this basically works on the principle of
  18147. 10:46:07feeding the output of a layer back into
  18148. 10:46:10the input layer in order to predict the
  18149. 10:46:12outcomes. Okay, this way it's more
  18150. 10:46:14precise and it is a little more complex
  18151. 10:46:16when compared to convolutional networks.
  18152. 10:46:18Now, one main important point of
  18153. 10:46:20recurrent neural networks is that they
  18154. 10:46:22have something known as memory. So,
  18155. 10:46:24basically each neuron will have some
  18156. 10:46:26information or some memory stored in
  18157. 10:46:28them so that if they can use this memory
  18158. 10:46:31in order to take actions in the future.
  18159. 10:46:33So, they have some experiences stored in
  18160. 10:46:36the form of memory so that they can make
  18161. 10:46:37their decisions based on previous
  18162. 10:46:39actions. Now, finally, we have
  18163. 10:46:41autoencoders. Now, autoencoders are
  18164. 10:46:44mainly used in dimensionality reduction
  18165. 10:46:46for learning generative models. Okay,
  18166. 10:46:49and one more important thing about
  18167. 10:46:50autoencoders is that the number of units
  18168. 10:46:53in the output layer and the input layer
  18169. 10:46:55is the same. This is because the output
  18170. 10:46:57layer has to reconstruct its own inputs.
  18171. 10:46:59So, these were the different types of
  18172. 10:47:02artificial neural networks. Now, let's
  18173. 10:47:04look at our next question, which is what
  18174. 10:47:06are Bayesian networks? Okay, so a
  18175. 10:47:08Bayesian network is a statistical model
  18176. 10:47:11that represents a set of variables and
  18177. 10:47:13the conditional dependencies in the form
  18178. 10:47:16of a directed acyclic graph. Now,
  18179. 10:47:18basically, on the occurrence of any
  18180. 10:47:20event, a Bayesian network can be used to
  18181. 10:47:22predict the likelihood that any one of
  18182. 10:47:25several possible known causes was a
  18183. 10:47:27contributing factor. An example of this
  18184. 10:47:30is a Bayesian network could be used to
  18185. 10:47:32study the relationship between diseases
  18186. 10:47:34and symptoms. So, given a set of
  18187. 10:47:36symptoms, the Bayesian network can be
  18188. 10:47:38used to find out the probability of the
  18189. 10:47:41presence of any diseases. All right, so
  18190. 10:47:43the next question is explain the
  18191. 10:47:45assessment that is used to test the
  18192. 10:47:47intelligence of a machine. Now, guys,
  18193. 10:47:49this is a a common question and it is
  18194. 10:47:52sort of a general knowledge-based
  18195. 10:47:53question. All right, I'm hoping that
  18196. 10:47:55most of you know the answer to this. So,
  18197. 10:47:57let's look at what the answer is. I'm
  18198. 10:48:00not sure how many of you have heard of
  18199. 10:48:01Alan Turing. So, Alan Turing was the one
  18200. 10:48:04who came up with the Turing test. Now,
  18201. 10:48:06this test is basically to determine
  18202. 10:48:08whether or not a computer is capable of
  18203. 10:48:11thinking like a human being.
  18204. 10:48:13So, if a machine or if a computer passes
  18205. 10:48:15this exam, it means that that machine is
  18206. 10:48:18capable of thinking like a human being.
  18207. 10:48:20It means that it is successfully an
  18208. 10:48:22artificial intelligent machine, meaning
  18209. 10:48:24that it can make its own decisions and
  18210. 10:48:26interpret data and form their own
  18211. 10:48:28formulations or form their own
  18212. 10:48:30conclusions about the data. Now, sadly,
  18213. 10:48:32I don't think there are a lot of
  18214. 10:48:33machines that have passed the Turing
  18215. 10:48:35test. In fact, I'm not sure if there is
  18216. 10:48:37any machine that's passed the Turing
  18217. 10:48:39test as of now, but in the near future,
  18218. 10:48:41I'm sure that we'll see machines who are
  18219. 10:48:44more smarter than human beings and who
  18220. 10:48:46have passed this test. Now, for a
  18221. 10:48:48machine, it might be very easy to do
  18222. 10:48:50computations, but it might be very hard
  18223. 10:48:52for a machine to just get up and walk
  18224. 10:48:54around. All right, the simple things
  18225. 10:48:56that us humans can do is very
  18226. 10:48:58complicated for a machine. They can do
  18227. 10:49:00computations which we can do in probably
  18228. 10:49:02a year, they can do those computations
  18229. 10:49:04in maybe a week or less than a week.
  18230. 10:49:07But, doing simple things such as walking
  18231. 10:49:09up to the fridge or walking up to the
  18232. 10:49:11kitchen is very hard for the machines.
  18233. 10:49:14So, to achieve that level of
  18234. 10:49:15intelligence, we're going to take a
  18235. 10:49:16while, but in the near future, I'm sure
  18236. 10:49:18we'll see machines which are way more
  18237. 10:49:20capable than human beings. Now, let's
  18238. 10:49:22move on to our next level. Now, here
  18239. 10:49:25I'll basically be discussing
  18240. 10:49:27intermediate level artificial
  18241. 10:49:28intelligence questions. So, let's look
  18242. 10:49:30at the first question. All right, the
  18243. 10:49:32first question is how does reinforcement
  18244. 10:49:35learning work? Explain with an example.
  18245. 10:49:38Okay, so first of all, reinforcement
  18246. 10:49:40learning is a type of machine learning.
  18247. 10:49:42We discussed about reinforcement
  18248. 10:49:44learning earlier. Reinforcement learning
  18249. 10:49:46is a type of machine learning wherein
  18250. 10:49:48there's an agent and you put this agent
  18251. 10:49:50in an unknown environment. All right,
  18252. 10:49:52now the agent has to figure out actions,
  18253. 10:49:54what sort of actions it must take, and
  18254. 10:49:56how it's going to get rewards so that it
  18255. 10:49:58can move from state zero to state one.
  18256. 10:50:01It's sort of like a video game. If
  18257. 10:50:02you're in a video game, let's say if
  18258. 10:50:04you're playing Counter-Strike, you're in
  18259. 10:50:06level zero or state zero. Now, if you
  18260. 10:50:09perform some action and if you get some
  18261. 10:50:11rewards, you're going to move to state
  18262. 10:50:12one. That's exactly how reinforcement
  18263. 10:50:14learning works. If you perform the
  18264. 10:50:16relevant actions and the correct
  18265. 10:50:17actions, you're going to get a reward
  18266. 10:50:19and you'll move on to the next state.
  18267. 10:50:21But in case you perform a wrong action,
  18268. 10:50:23you'll get negative rewards and you'll
  18269. 10:50:25stay in the same state unless and until
  18270. 10:50:27you don't learn. All right, so if you
  18271. 10:50:29learn and achieve, then you'll move to
  18272. 10:50:31the next state. So, basically a
  18273. 10:50:33reinforcement learning system will have
  18274. 10:50:35two main components. It'll have an agent
  18275. 10:50:38and an environment. Now, the agent I've
  18276. 10:50:40been repetitively saying an agent An
  18277. 10:50:42agent is basically the reinforcement
  18278. 10:50:44learning algorithm. It is the model. The
  18279. 10:50:47model has to learn everything on its
  18280. 10:50:48own. It has to collect data on its own.
  18281. 10:50:51It has to draw useful insights on its
  18282. 10:50:53own. Okay, you're not going to feed any
  18283. 10:50:55predefined data to this reinforcement
  18284. 10:50:57learning agent. All right, he has to
  18285. 10:50:59figure out everything on its own. So,
  18286. 10:51:01let's look at an example of
  18287. 10:51:02Counter-Strike. Okay, I'm not sure how
  18288. 10:51:04many of you play the game, but yeah.
  18289. 10:51:06What happens here is the reinforcement
  18290. 10:51:08learning agent or the player one
  18291. 10:51:10collects a state S0 from the
  18292. 10:51:12environment. Okay, so let's suppose that
  18293. 10:51:14you're playing Counter-Strike and you're
  18294. 10:51:16in state zero. Now, you'll perform some
  18295. 10:51:19action A0. All right, initially it's
  18296. 10:51:21going to be a random action. So,
  18297. 10:51:23obviously if you're put in an unknown
  18298. 10:51:24environment, your first action is going
  18299. 10:51:26to be random, correct? Because you don't
  18300. 10:51:28know what's right, you don't know what's
  18301. 10:51:29wrong. So, in your state zero, you'll
  18302. 10:51:31take an action A0. This will result in a
  18303. 10:51:34new state S1. And on achieving state S1,
  18304. 10:51:38the agent will get a reward R1, okay,
  18305. 10:51:40from the environment. Now, in the case
  18306. 10:51:43of Counter-Strike games, if you've
  18307. 10:51:45observed, whenever you win a state or
  18308. 10:51:47you pass a level, you're going to get
  18309. 10:51:49some rewards. Maybe you'll get more
  18310. 10:51:51weapons or you'll get more points. Okay,
  18311. 10:51:53just like that, in reinforcement
  18312. 10:51:55learning problem, you'll get some reward
  18313. 10:51:57R1. Okay, it's basically a plus point.
  18314. 10:52:00You might get a negative reward or a
  18315. 10:52:02positive reward based on the action that
  18316. 10:52:04you take. Now, this loop will go on
  18317. 10:52:07until the agent is dead or it reaches
  18318. 10:52:09the destination. So, in Counter-Strike,
  18319. 10:52:11until you have failed the level, you
  18320. 10:52:14will keep playing the game, right?
  18321. 10:52:15You'll keep moving from state one, state
  18322. 10:52:17two, state three, and so on. Or, if
  18323. 10:52:19you've reached the destination, then
  18324. 10:52:20it's the end game. That's exactly how it
  18325. 10:52:23works in reinforcement learning. If the
  18326. 10:52:25agent has explored the entire
  18327. 10:52:26environment and reached the end state,
  18328. 10:52:29that's when the loop will end. All
  18329. 10:52:31right, that's exactly how reinforcement
  18330. 10:52:33learning works. It is very similar to
  18331. 10:52:35the games that we play. All right, it's
  18332. 10:52:37very understandable. Now, let's move on
  18333. 10:52:39and discuss the next question. So, the
  18334. 10:52:41next question is explain Markov decision
  18335. 10:52:44process with an example. Now, the
  18336. 10:52:46solution for a reinforcement learning
  18337. 10:52:48problem is achieved through the Markov
  18338. 10:52:50decision process. It's basically a
  18339. 10:52:53mathematical approach that maps the
  18340. 10:52:55solution in reinforcement learning.
  18341. 10:52:57Okay, so now to understand this, there
  18342. 10:52:59are a couple of parameters in a Markov
  18343. 10:53:02decision process. They're going to be a
  18344. 10:53:04set of actions called A. Okay, you can
  18345. 10:53:06name them A. A set of states, there's
  18346. 10:53:09going to be reward, there's going to be
  18347. 10:53:11policy, and there's going to be value.
  18348. 10:53:13To sum it up, what exactly happens in a
  18349. 10:53:15Markov decision process is that the
  18350. 10:53:18agent takes an action A to transition
  18351. 10:53:21from the start state to the end state.
  18352. 10:53:23Now, while doing so, the agent receives
  18353. 10:53:26some reward R for each action that he
  18354. 10:53:28takes. The series of actions taken by
  18355. 10:53:30the agent will define a policy or an
  18356. 10:53:33approach. And the rewards collected will
  18357. 10:53:36define the value. So, the main goal in a
  18358. 10:53:39Markov decision process is to maximize
  18359. 10:53:41the rewards by choosing the most optimum
  18360. 10:53:44policy. Meaning that you're going to
  18361. 10:53:46choose the best path or the best
  18362. 10:53:47solution in order to get the most number
  18363. 10:53:50of rewards. Now, in order to make you
  18364. 10:53:52all understand this better, let's solve
  18365. 10:53:54the shortest path problem by using
  18366. 10:53:56Markov decision process. I'm sure all of
  18367. 10:53:59you have heard of shortest path problem.
  18368. 10:54:01This was I think taught to us when we
  18369. 10:54:03were in 11th or 12th, I'm not sure. Look
  18370. 10:54:06at the diagram that is over here. This
  18371. 10:54:08is basically a representation of our
  18372. 10:54:11problem. Given this representation, our
  18373. 10:54:13goal here is to find the shortest path
  18374. 10:54:16between the node A and node D. All
  18375. 10:54:18right, you can see nodes A, B, C, and D.
  18376. 10:54:21We have to find the shortest path
  18377. 10:54:23between node A and node D. Now, the link
  18378. 10:54:25between these two nodes has a number on
  18379. 10:54:28it. Okay, for example, between A and C
  18380. 10:54:30you can see there's a number 15. Okay,
  18381. 10:54:32this basically denotes the cost to
  18382. 10:54:34traverse that edge. So, if you want to
  18383. 10:54:36go from A to C, you'll spend around 15
  18384. 10:54:39points. So, our end goal here is to
  18385. 10:54:41travel between node A and node D with
  18386. 10:54:44minimal possible cost. We should travel
  18387. 10:54:47between A to D in such a way that our
  18388. 10:54:49cost is minimal. Now, in this problem if
  18389. 10:54:51you notice that we have a set of states.
  18390. 10:54:54Okay, these are denoted by the nodes A,
  18391. 10:54:56B, C, D. Now, like I mentioned earlier,
  18392. 10:54:59a Markov decision process has a set of
  18393. 10:55:01states. Similarly, in this problem the
  18394. 10:55:03set of states are A, B, C, D. The action
  18395. 10:55:06is to traverse from one node to the
  18396. 10:55:08other. So, going from A to B is
  18397. 10:55:10basically an action. Going from A to C
  18398. 10:55:12is another action. Going from A to D is
  18399. 10:55:15another action, and so on. Now, reward
  18400. 10:55:17is represented by the cost on each of
  18401. 10:55:20these links. And the policy is the path
  18402. 10:55:22which is taken to reach the destination.
  18403. 10:55:25So, our aim here is to choose a policy
  18404. 10:55:28that gets us to node D in the minimum
  18405. 10:55:30cost possible. So, how do you think you
  18406. 10:55:33can solve this problem? All right, you
  18407. 10:55:34can start off at node A and you can take
  18408. 10:55:37baby steps to your destination. Now,
  18409. 10:55:39initially only the next possible node is
  18410. 10:55:41visible to you. Like I mentioned
  18411. 10:55:43earlier, the initial action taken in a
  18412. 10:55:46reinforcement learning problem is always
  18413. 10:55:48random. So, at random you'll choose any
  18414. 10:55:50node. Let's say you take A to B. Now, if
  18415. 10:55:53you go from A to B, you can go B to D
  18416. 10:55:55and you'll reach the destination. So,
  18417. 10:55:57policy is the path which is taken to
  18418. 10:56:00reach the destination. All right, so it
  18419. 10:56:02can go from A to B to D or you can go
  18420. 10:56:04from A to C to D or you can go A C B D.
  18421. 10:56:08All right, now it's up to you to figure
  18422. 10:56:09out which is the shortest path. All
  18423. 10:56:11right, you have to choose a path in such
  18424. 10:56:13a way that the cost between A to D is
  18425. 10:56:16minimized. So guys, this was a simple
  18426. 10:56:18problem of how Markov decision process
  18427. 10:56:21is used to solve the shortest path
  18428. 10:56:22problem. Now, let's move on and look at
  18429. 10:56:25our next question. All right, now the
  18430. 10:56:27next question is explain reward maximiza
  18431. 10:56:30tion in reinforcement learning. So,
  18432. 10:56:32basically a reinforcement learning agent
  18433. 10:56:35works based on the theory of reward
  18434. 10:56:37maximization. Okay, in the previous
  18435. 10:56:39question itself I told you that the main
  18436. 10:56:41aim of reinforcement learning is to
  18437. 10:56:43maximize the reward. So, that's why a
  18438. 10:56:46reinforcement learning agent must be
  18439. 10:56:48trained in such a way that he takes the
  18440. 10:56:50best action so that the reward is
  18441. 10:56:52maximum. Okay, this is exactly what
  18442. 10:56:54reward maximization means. He has to
  18443. 10:56:57choose the best policy in such a way
  18444. 10:56:59that the reward is maximum. Now, let me
  18445. 10:57:02explain this with a small game. So, in
  18446. 10:57:04the figure you can see a fox, you can
  18447. 10:57:06see some meat and you can see a tiger.
  18448. 10:57:08Now, our reinforcement learning agent is
  18449. 10:57:10the fox. His end goal is to eat the
  18450. 10:57:13maximum amount of meat before being
  18451. 10:57:15eaten by the tiger. Okay, so he has to
  18452. 10:57:18explore around, eat the maximum number
  18453. 10:57:20of meat that he can eat before the tiger
  18454. 10:57:22kills him. Since the fox is a clever
  18455. 10:57:25fellow, he eats the meat that is closer
  18456. 10:57:27to him. Okay, so rather than eating the
  18457. 10:57:30meat which is close to the tiger, he
  18458. 10:57:32eats the meat which is only close to
  18459. 10:57:33him. This is because the closer he gets
  18460. 10:57:35to the tiger, the higher are his chances
  18461. 10:57:38of getting killed. So, as a result of
  18462. 10:57:40this, the rewards near the tiger, even
  18463. 10:57:43if they are bigger meat chunks, will be
  18464. 10:57:45discounted. So, because the fox is not
  18465. 10:57:48going closer to the tiger and eating the
  18466. 10:57:50meat chunks closer to the tiger, this
  18467. 10:57:52reward will get discounted. Now, I know
  18468. 10:57:55you're all wondering what discounted is.
  18469. 10:57:57Now, this is done because of the
  18470. 10:57:59uncertainty factor that the tiger might
  18471. 10:58:01kill the fox. So, what is discounting of
  18472. 10:58:04reward? Okay, how does it work? To
  18473. 10:58:06understand this, we define a discount
  18474. 10:58:09rate called gamma. Okay, this is a
  18475. 10:58:10parameter and the value of gamma always
  18476. 10:58:14ranges between zero and one. So, the
  18477. 10:58:16smaller the gamma, the larger the
  18478. 10:58:18discount and so on. So, guys, this was
  18479. 10:58:20reward maximization. So, here basically
  18480. 10:58:23the fox will try to get as much as meat
  18481. 10:58:25chunks as he can and he'll also try to
  18482. 10:58:28avoid getting killed because that will
  18483. 10:58:30end the reinforcement learning loop. We
  18484. 10:58:32also discussed the discounted factor.
  18485. 10:58:34All right, now that is not needed to
  18486. 10:58:36understand reward maximization, but I
  18487. 10:58:38just thought I'll add on some extra
  18488. 10:58:39info.
  18489. 10:58:40Now, let's look at the next question
  18490. 10:58:42which is what is exploitation and
  18491. 10:58:44exploration trade-off?
  18492. 10:58:46So, basically exploration is, like the
  18493. 10:58:49name suggests, it is about exploring and
  18494. 10:58:52capturing more information about an
  18495. 10:58:53environment. Now, on the other hand,
  18496. 10:58:56exploitation is about using the already
  18497. 10:58:58known exploited information to heighten
  18498. 10:59:01the rewards. So, consider the same
  18499. 10:59:03example that we discussed in the
  18500. 10:59:04previous question. Here, the fox only
  18501. 10:59:07eats the meat chunks which are close to
  18502. 10:59:09him. Okay, he does not eat the bigger
  18503. 10:59:11chunks because even though the bigger
  18504. 10:59:13chunks would give him more rewards, it
  18505. 10:59:15would get him killed. Okay, he does not
  18506. 10:59:17go towards the tiger itself. Now, if the
  18507. 10:59:19fox only focuses on the closest reward,
  18508. 10:59:23he will never reach the big chunks of
  18509. 10:59:24meat. Okay, this is what exploitation
  18510. 10:59:27is. He's sticking only to the
  18511. 10:59:29information that he knows and he's
  18512. 10:59:31trying to get the most number of rewards
  18513. 10:59:32from it. But if the fox decides to
  18514. 10:59:35explore a bit, it can find the bigger
  18515. 10:59:37rewards. Okay, the bigger rewards are
  18516. 10:59:39basically the big chunks of meat which
  18517. 10:59:41are near the tiger. And this is exactly
  18518. 10:59:43what exploration is. Okay, exploitation
  18519. 10:59:46is about using the already known
  18520. 10:59:48information to heighten your rewards.
  18521. 10:59:51Exploration on the other hand is about
  18522. 10:59:53exploring and capturing more information
  18523. 10:59:55about an environment.
  18524. 10:59:57All right, so that was about
  18525. 10:59:58exploitation and exploration. Now, let's
  18526. 11:00:01move on to our next question. So, this
  18527. 11:00:04is a difference question which ask the
  18528. 11:00:07difference between parametric and
  18529. 11:00:09non-parametric models. So, a parametric
  18530. 11:00:12model basically uses a fixed number of
  18531. 11:00:14parameters to build the model. Now,
  18532. 11:00:16first of all guys, what are parameters?
  18533. 11:00:18Now, parameters are basically predictor
  18534. 11:00:21variables that are used to build a
  18535. 11:00:23machine learning model or build any
  18536. 11:00:25predictive analytics model. Now that we
  18537. 11:00:27know what parameters are, let's try to
  18538. 11:00:29understand the difference between a
  18539. 11:00:31parametric and a non-parametric model.
  18540. 11:00:34Now, parametric model basically uses a
  18541. 11:00:35fixed number of parameters to build the
  18542. 11:00:37model. A non-parametric model uses
  18543. 11:00:40flexible number of parameters to build
  18544. 11:00:42the model. When it comes to a parametric
  18545. 11:00:44model, the assumptions about the data
  18546. 11:00:46are very strong. In a non-parametric
  18547. 11:00:49model, there are fewer assumptions about
  18548. 11:00:51the data. A parametric model has the
  18549. 11:00:53fixed number of parameters. Everything
  18550. 11:00:55is defined over here. So, the
  18551. 11:00:57computation is very fast. Okay, you know
  18552. 11:00:59what sort of variables you'll need to
  18553. 11:01:01predict the outcome. Okay, you have a
  18554. 11:01:03defined set of variables or a defined
  18555. 11:01:05set of predictor variables that will
  18556. 11:01:07compute your outcome. So, that's why the
  18557. 11:01:09computation is a bit faster. When you
  18558. 11:01:12compare to non-parametric models, there
  18559. 11:01:14are a lot of parameters taken into
  18560. 11:01:16account. Now, when it comes to a
  18561. 11:01:18non-parametric model, you do not have a
  18562. 11:01:20fixed number of parameters. All right,
  18563. 11:01:22you do not have a fixed number of
  18564. 11:01:24predictor variables that will help you
  18565. 11:01:26get to the outcome. So, the computation
  18566. 11:01:28is a bit slower. Now, parametric models
  18567. 11:01:31require lesser data and non-parametric
  18568. 11:01:34require more data. Example of parametric
  18569. 11:01:37models include logistic regression and
  18570. 11:01:39naive bias. And for non-parametric
  18571. 11:01:41models, we have KNN and decision tree
  18572. 11:01:43models. Now, logistic and naive bias
  18573. 11:01:46models are very strong models because
  18574. 11:01:49they have a fixed number of parameters
  18575. 11:01:51or a fixed number of predictor variables
  18576. 11:01:53and they will give you an immediate
  18577. 11:01:55output. Okay, when it comes to
  18578. 11:01:57non-parametric models like decision tree
  18579. 11:01:58models and KNN, you might even observe a
  18580. 11:02:01little bit of overfitting. Okay, this
  18581. 11:02:03happens because you have a fewer number
  18582. 11:02:06of assumptions about the data and also
  18583. 11:02:08because your parameters are not fixed.
  18584. 11:02:10Now, that's not the reason for
  18585. 11:02:12overfitting, but it's seen that in some
  18586. 11:02:14of the non-parametric models,
  18587. 11:02:16overfitting occurs more often. Now,
  18588. 11:02:18let's discuss the next question, which
  18589. 11:02:21is what is the difference between
  18590. 11:02:22hyperparameters and model parameters?
  18591. 11:02:25Now, model parameters are the predictor
  18592. 11:02:27variables that I was speaking about
  18593. 11:02:29earlier. Hyperparameters, let's discuss
  18594. 11:02:32what they are. Okay, model parameters
  18595. 11:02:34are the features of training data that
  18596. 11:02:36will learn on its own during training.
  18597. 11:02:38Whereas model hyperparameters are the
  18598. 11:02:40parameters that determine the training
  18599. 11:02:42process.
  18600. 11:02:43Now, let's see that you want to
  18601. 11:02:45determine the height of an individual
  18602. 11:02:47depending on his weight. The height and
  18603. 11:02:50weight will become your model
  18604. 11:02:52parameters. But, your hyperparameter is
  18605. 11:02:55basically the learning rate. It's the
  18606. 11:02:57rate at which your model is going to
  18607. 11:02:59learn this correlation between the
  18608. 11:03:01height and the weight. So, this is the
  18609. 11:03:03difference between model parameters and
  18610. 11:03:05hyperparameters. Model parameters are
  18611. 11:03:07the ones that you find in your data.
  18612. 11:03:10These are all the variables that you use
  18613. 11:03:12to predict your outcomes.
  18614. 11:03:13Hyperparameters will define your
  18615. 11:03:15training process. There is a huge
  18616. 11:03:17difference between model and
  18617. 11:03:18hyperparameters. Another difference is
  18618. 11:03:21that they are internal to the model and
  18619. 11:03:23their value can be estimated from the
  18620. 11:03:24data. Hyperparameters are external to
  18621. 11:03:27the model and their value cannot be
  18622. 11:03:29estimated from data. Now, like I said,
  18623. 11:03:31model parameters are derived from your
  18624. 11:03:33data itself. Okay, these are the
  18625. 11:03:35parameters that are there in your data.
  18626. 11:03:37Hyperparameters are the ones that you
  18627. 11:03:39define in order to train your entire
  18628. 11:03:42data. So, that is the difference between
  18629. 11:03:44hyperparameters and model parameters.
  18630. 11:03:47So, next question is what are
  18631. 11:03:48hyperparameters in deep neural networks?
  18632. 11:03:51So, guys, like I mentioned in the
  18633. 11:03:53previous example, hyperparameters are
  18634. 11:03:56variables such as the learning rate.
  18635. 11:03:58This will define how your entire data
  18636. 11:04:00training process goes. For those of you
  18637. 11:04:02don't know, in order to build a model,
  18638. 11:04:04you first need to train the model and
  18639. 11:04:06then you need to test it. Okay, now
  18640. 11:04:08while training the model, you're going
  18641. 11:04:09to make the model learn a lot of things.
  18642. 11:04:12You're going to give it a lot of data.
  18643. 11:04:14It has to figure out relations between
  18644. 11:04:16various variables and how these
  18645. 11:04:18variables are affecting the output. All
  18646. 11:04:21of this training will depend on a few
  18647. 11:04:23variables such as the learning rate.
  18648. 11:04:25Okay, these are basically called
  18649. 11:04:26hyperparameters in deep neural networks.
  18650. 11:04:29So, these parameters will define the
  18651. 11:04:31number of hidden layers that are present
  18652. 11:04:33in the network. Okay, and more the
  18653. 11:04:35number of hidden layers, the more
  18654. 11:04:36accurate your network is going to be.
  18655. 11:04:39Whereas, if you have less a number of
  18656. 11:04:40units, you may cause underfitting in
  18657. 11:04:42your data. Underfitting will also result
  18658. 11:04:45in inaccurate predictions. So, that's
  18659. 11:04:47why you need to make sure that the
  18660. 11:04:49number of hidden units in your hidden
  18661. 11:04:50layers are perfect, are ideal. Okay, and
  18662. 11:04:53this is determined by your learning rate
  18663. 11:04:56or by your hyperparameters. Not by your
  18664. 11:04:58learning rate specifically, but by the
  18665. 11:05:00number of hyperparameters you have.
  18666. 11:05:02These number of hidden layers are
  18667. 11:05:04determined by the hyperparameters. Okay,
  18668. 11:05:07that's why hyper parameters are very
  18669. 11:05:08important in deep neural networks. I
  18670. 11:05:11hope you all are clear with this. Now,
  18671. 11:05:13let's look at our next question, which
  18672. 11:05:14is explain the different algorithms used
  18673. 11:05:17for hyperparameter optimization. We'll
  18674. 11:05:20discuss the three methods, which are
  18675. 11:05:21grid search, random search, and Bayesian
  18676. 11:05:24optimization. Okay, now grid search
  18677. 11:05:26basically will train the network only on
  18678. 11:05:29the two sets of hyper parameters, which
  18679. 11:05:31are learning rate and the number of
  18680. 11:05:32layers. Okay, so it's going to use every
  18681. 11:05:34combination of these two sets in order
  18682. 11:05:37to train the network. Okay, after that
  18683. 11:05:39it'll evaluate the efficiency of the
  18684. 11:05:41model by using the cross-validation
  18685. 11:05:43techniques. Cross-validation is the best
  18686. 11:05:46improvement method. Okay, it's the best
  18687. 11:05:48way to check if your model is optimal or
  18688. 11:05:51not. Then we have random search. Now,
  18689. 11:05:54this will randomly select samples, and
  18690. 11:05:56it will evaluate sets for a particular
  18691. 11:05:58probability distribution. Now, in random
  18692. 11:06:01search there is no fixed number of hyper
  18693. 11:06:03parameters that it's going to evaluate.
  18694. 11:06:05So, it'll randomly select a set of hyper
  18695. 11:06:08parameters. Okay, for example, now
  18696. 11:06:10instead of checking your entire sample
  18697. 11:06:13or your entire Let's say that you have
  18698. 11:06:1510,000 samples. Instead of checking all
  18699. 11:06:17of these samples, it'll randomly select
  18700. 11:06:19100 parameters that can be checked.
  18701. 11:06:21Okay, and then it'll use this to build
  18702. 11:06:23the model.
  18703. 11:06:24After that, we have Bayesian
  18704. 11:06:25optimization.
  18705. 11:06:27Okay, now Bayesian optimization
  18706. 11:06:29basically uses something known as a
  18707. 11:06:31Gaussian process. Basically, the
  18708. 11:06:33Gaussian process will help in model
  18709. 11:06:35tuning. Okay, model tuning or you can
  18710. 11:06:37also say parameter tuning. So, parameter
  18711. 11:06:39tuning will help you tweak the
  18712. 11:06:41parameters a little bit in order to
  18713. 11:06:43improve the efficiency of the model.
  18714. 11:06:45Okay, so Bayesian optimization basically
  18715. 11:06:47makes use of the Gaussian process, which
  18716. 11:06:50will provide model tuning to your
  18717. 11:06:51algorithm and thus improve the
  18718. 11:06:53efficiency. Now guys, the one of the
  18719. 11:06:55most important ways to improve the
  18720. 11:06:57efficiency of a model is by
  18721. 11:06:59hyperparameter optimization. Okay, if
  18722. 11:07:01you're tuning your hyper parameters and
  18723. 11:07:03if you're trying to check in which way
  18724. 11:07:05these hyper parameters will give you the
  18725. 11:07:07most accurate outcome, that's when your
  18726. 11:07:10result will be very good. Okay, so
  18727. 11:07:12that's the best way to improve the
  18728. 11:07:13efficiency of the model. All right, now
  18729. 11:07:15let's look at our next question. The
  18730. 11:07:18next question is how does data
  18731. 11:07:20overfitting occur and how can it be
  18732. 11:07:22fixed? Now guys, this is a very common
  18733. 11:07:24question in a machine learning or in an
  18734. 11:07:27artificial intelligence interview. Okay,
  18735. 11:07:29people expect you to understand what
  18736. 11:07:31data overfitting is and how you can fix
  18737. 11:07:34these problems. Okay, because data
  18738. 11:07:36overfitting occurs pretty often,
  18739. 11:07:37especially if you're using decision
  18740. 11:07:40trees or if you're using random forest.
  18741. 11:07:43Okay, random forest actually reduces
  18742. 11:07:44overfitting, but sometimes with these
  18743. 11:07:46complex models you can get data
  18744. 11:07:48overfitting. Now, to answer this
  18745. 11:07:49question, first of all, let's understand
  18746. 11:07:51what overfitting really is. So,
  18747. 11:07:54overfitting occurs when a machine
  18748. 11:07:56learning algorithm captures the noise of
  18749. 11:07:58the data. Okay, this causes an algorithm
  18750. 11:08:01to show low bias, but high variance in
  18751. 11:08:03the outcome. Now, what overfitting
  18752. 11:08:05really means is you have trained your
  18753. 11:08:08model way too many times on the training
  18754. 11:08:10data. Okay, so basically the model has
  18755. 11:08:13memorized the training data. It has
  18756. 11:08:15memorized the noise in the training
  18757. 11:08:17data. Okay, so if you feed new data to
  18758. 11:08:20the model during the testing stage, it
  18759. 11:08:22will not be able to recognize the noise
  18760. 11:08:24or it will not be able to recognize any
  18761. 11:08:26sort of correlation in that data. Okay,
  18762. 11:08:28that's why it won't be able to get a
  18763. 11:08:30proper outcome. Okay, that's when
  18764. 11:08:32overfitting happens. You have trained
  18765. 11:08:34the model way too much with the training
  18766. 11:08:36data and this has resulted in inaccurate
  18767. 11:08:39outcome during the testing phase. Okay,
  18768. 11:08:42that's what overfitting is about. Now,
  18769. 11:08:44how do you avoid overfitting? First of
  18770. 11:08:46all, is cross-validation. Now, before
  18771. 11:08:49this also I mentioned that
  18772. 11:08:50cross-validation is the best way to
  18773. 11:08:52obtain a more optimal solution. Now, the
  18774. 11:08:55general idea behind cross-validation is
  18775. 11:08:58to split the training data in order to
  18776. 11:09:00generate multiple mini train test
  18777. 11:09:03splits. Okay, these splits can be used
  18778. 11:09:05to tune your model. Okay, so you're
  18779. 11:09:07basically splitting the training data in
  18780. 11:09:09such a way that, you know, the model
  18781. 11:09:11does not just use the entire training
  18782. 11:09:13data and memorize it. Instead, it's
  18783. 11:09:15going to check the different sets in the
  18784. 11:09:17training data and the different sets in
  18785. 11:09:18the testing data and learn from it.
  18786. 11:09:20Okay, so cross validation is one of the
  18787. 11:09:22best ways to prevent overfitting.
  18788. 11:09:24Another method to prevent overfitting is
  18789. 11:09:27by training the model with more data.
  18790. 11:09:30So, feeding more data to the machine
  18791. 11:09:31learning model will help in better
  18792. 11:09:33analysis and classification. However,
  18793. 11:09:35this method is not always going to work,
  18794. 11:09:38but yeah, this is also one of the ways
  18795. 11:09:40to prevent overfitting. Okay, next we
  18796. 11:09:42have removing features. Now, many times
  18797. 11:09:45the data set contains irrelevant
  18798. 11:09:47features or predictor variables, which
  18799. 11:09:49are not needed for analysis. Such
  18800. 11:09:51features will only increase the
  18801. 11:09:53complexity of the model. Therefore, it
  18802. 11:09:55lead to possibilities of data
  18803. 11:09:57overfitting. Okay, so if you have
  18804. 11:09:59irrelevant data, like for example, if
  18805. 11:10:02you're trying to understand the weight
  18806. 11:10:04of a person depending on its height, and
  18807. 11:10:06you have another variable, let's say,
  18808. 11:10:08you have a variable like the name of the
  18809. 11:10:10person. Okay, now the name of the person
  18810. 11:10:12is not relevant in understanding the
  18811. 11:10:14height of an individual. So, if you have
  18812. 11:10:17irrelevant predictor variables, then it
  18813. 11:10:20will just increase the complexity of the
  18814. 11:10:21model because you have an extra
  18815. 11:10:23irrelevant variable. All right, this
  18816. 11:10:25will only increase the complexity of the
  18817. 11:10:27model. It will not help the model in any
  18818. 11:10:29way. So, make sure you remove irrelevant
  18819. 11:10:31features or you remove redundant
  18820. 11:10:33features. Okay, the next method is early
  18821. 11:10:36stopping. Now, a machine learning model
  18822. 11:10:38is trained iteratively. This will allow
  18823. 11:10:41us to check how well each iteration of
  18824. 11:10:43the model performs. But, after a certain
  18825. 11:10:46number of iterations, the model's
  18826. 11:10:48performance starts to saturate. Further
  18827. 11:10:51training will only result in
  18828. 11:10:52overfitting. Okay, so like I mentioned,
  18829. 11:10:55if you train the model with the same
  18830. 11:10:57data and you make the model memorize the
  18831. 11:11:00data, then it'll just saturate. It won't
  18832. 11:11:02be able to predict any outcomes after a
  18833. 11:11:04point. What you have to do is you have
  18834. 11:11:06to understand where you need to stop
  18835. 11:11:09training the model. So, this can be
  18836. 11:11:11achieved by using a mechanism known as
  18837. 11:11:13early stopping. So, at this point you
  18838. 11:11:15know that you have to stop training the
  18839. 11:11:16model because this might result in
  18840. 11:11:18overfitting. Now, regularization is one
  18841. 11:11:21of the most common ways to prevent
  18842. 11:11:22overfitting. Regularization can be done
  18843. 11:11:25in n number of ways. Okay, the method
  18844. 11:11:27will always depend on the type of
  18845. 11:11:29learner you're implementing. For
  18846. 11:11:30example, pruning is performed on
  18847. 11:11:32decision trees. Now, pruning is a type
  18848. 11:11:35of regularization. Similarly, the
  18849. 11:11:37dropout technique can be used on neural
  18850. 11:11:39networks. And also, there are other
  18851. 11:11:40methods like parameter tuning which can
  18852. 11:11:42help to solve overfitting.
  18853. 11:11:44The next way to prevent overfitting is
  18854. 11:11:46by using ensemble models. Now, ensemble
  18855. 11:11:49learning is a technique that is used to
  18856. 11:11:51create multiple machine learning models
  18857. 11:11:53which are then combined to produce more
  18858. 11:11:56accurate results. So, basically if you
  18859. 11:11:58have one problem statement in machine
  18860. 11:12:00learning, you're going to use like five
  18861. 11:12:03to 10 different models and then you're
  18862. 11:12:05going to calculate the accuracy
  18863. 11:12:07depending on the average of the result
  18864. 11:12:09from each of these models. By this way,
  18865. 11:12:11you will reduce overfitting. Now,
  18866. 11:12:13ensemble models is one of the best ways
  18867. 11:12:16to prevent overfitting. An example is
  18868. 11:12:18the random forest. Random forest uses
  18869. 11:12:21ensemble of decision trees to make more
  18870. 11:12:23accurate predictions and to avoid
  18871. 11:12:25overfitting. So, basically random forest
  18872. 11:12:28is a set of decision trees. So, here
  18873. 11:12:30you're going to train the model by using
  18874. 11:12:32a set of decision trees and this way
  18875. 11:12:34you'll have different data sets and on
  18876. 11:12:36each of these data sets you'll have a
  18877. 11:12:38different decision tree model. Okay,
  18878. 11:12:40this will reduce overfitting to a very
  18879. 11:12:42large extent. That's why in most of the
  18880. 11:12:44cases when you see a decision tree
  18881. 11:12:46having overfitting issues, you'll be
  18882. 11:12:48asked to use random forest. So guys,
  18883. 11:12:51those were the different ways to prevent
  18884. 11:12:52overfitting. Now the next question is
  18885. 11:12:55mention a technique that helps to avoid
  18886. 11:12:57overfitting in a neural network. Now the
  18887. 11:13:00most famous method to prevent
  18888. 11:13:02overfitting in neural networks is
  18889. 11:13:04dropout technique. Okay, now dropout is
  18890. 11:13:06a type of regularization technique which
  18891. 11:13:08is used to avoid overfitting in a neural
  18892. 11:13:10network. So here what you do is you
  18893. 11:13:12randomly select neurons and you drop
  18894. 11:13:15them during the training phase. Right?
  18895. 11:13:17So the dropout value also has to be
  18896. 11:13:19chosen very carefully because a higher
  18897. 11:13:21dropout value will result in under
  18898. 11:13:23learning by the network. So if you're
  18899. 11:13:25dropping out too many predictor
  18900. 11:13:26variables or if you're dropping out too
  18901. 11:13:28many neurons in a neural network, then
  18902. 11:13:30the model will not learn enough. Okay,
  18903. 11:13:33because there's not enough predictor
  18904. 11:13:34variables or not enough neurons. But if
  18905. 11:13:37you have too much of a low rate for a
  18906. 11:13:39dropout value, then this might have a
  18907. 11:13:41very minimal effect. So make sure your
  18908. 11:13:43dropout value is very optimal depending
  18909. 11:13:45on the problem you're trying to solve.
  18910. 11:13:47Okay, so dropout is the technique which
  18911. 11:13:49is used to avoid overfitting in a neural
  18912. 11:13:51network. Next question is what is the
  18913. 11:13:54purpose of deep learning framework such
  18914. 11:13:56as Keras, TensorFlow, and PyTorch? So
  18915. 11:13:59Keras is basically an open-source neural
  18916. 11:14:01network library which is written in
  18917. 11:14:03Python. So basically it is designed to
  18918. 11:14:06enable fast experimentation with deep
  18919. 11:14:08neural networks. Now TensorFlow is
  18920. 11:14:10another open-source software library for
  18921. 11:14:12data flow programming. TensorFlow is
  18922. 11:14:14mainly used in machine learning
  18923. 11:14:16applications. Similarly, PyTorch is
  18924. 11:14:18again an open-source machine learning
  18925. 11:14:20library for Python. Its applications are
  18926. 11:14:23mainly in the field of natural language
  18927. 11:14:24processing. Now I'd say that these three
  18928. 11:14:27deep learning frameworks are the most
  18929. 11:14:29important when it comes to machine
  18930. 11:14:30learning and deep learning because they
  18931. 11:14:32have a varied set of functions in them
  18932. 11:14:35which help in building a better machine
  18933. 11:14:36learning model or a better deep learning
  18934. 11:14:39network. Now let's look at question
  18935. 11:14:41number 24, which is differentiate
  18936. 11:14:43between NLP and text mining. So guys,
  18937. 11:14:46NLP stands for natural language
  18938. 11:14:47processing for those of you who don't
  18939. 11:14:49know. Now first of all, let me clear out
  18940. 11:14:51a confusion between text mining and
  18941. 11:14:53natural language processing. A lot of
  18942. 11:14:55people tend to think that text mining
  18943. 11:14:57and NLP are the same thing, but text
  18944. 11:14:59mining is the broader field and NLP is
  18945. 11:15:02basically an application of text mining
  18946. 11:15:04or it's basically a technique used in
  18947. 11:15:06text mining. So the aim of text mining
  18948. 11:15:09is to extract useful insights from
  18949. 11:15:10structured and unstructured text.
  18950. 11:15:12Whereas the aim of NLP is to understand
  18951. 11:15:15what is conveyed in these texts. Now
  18952. 11:15:17text mining can be done using text
  18953. 11:15:19processing languages like Perl and NLP
  18954. 11:15:22can be achieved using advanced machine
  18955. 11:15:24learning models such as deep neural
  18956. 11:15:25networks. Now the outcome for text
  18957. 11:15:28mining is you'll calculate the frequency
  18958. 11:15:30of words, you'll understand the patterns
  18959. 11:15:32between different words, you'll
  18960. 11:15:33understand the correlations between two
  18961. 11:15:35different words and you'll see how these
  18962. 11:15:37two words occur together more frequently
  18963. 11:15:39and why they occur together more
  18964. 11:15:41frequently. So text mining basically
  18965. 11:15:43will give you a more understanding about
  18966. 11:15:45the words that are used in a document.
  18967. 11:15:48Whereas in NLP, you'll understand the
  18968. 11:15:50grammar behind the text. You'll
  18969. 11:15:52understand in more depth about the
  18970. 11:15:54language that is used in the document or
  18971. 11:15:57in whatever you're trying to analyze. So
  18972. 11:15:59that is the difference between NLP and
  18973. 11:16:01text mining. NLP is a little more
  18974. 11:16:03advanced field because you use deep
  18975. 11:16:05neural networks to perform this. Text
  18976. 11:16:07mining on the other hand makes use of
  18977. 11:16:09NLP. Next question is what are the
  18978. 11:16:12different components of NLP? Now there
  18979. 11:16:14are two components of natural language
  18980. 11:16:16processing, which is natural language
  18981. 11:16:18understanding and natural language
  18982. 11:16:20generation. In natural language
  18983. 11:16:22understanding, you'll basically map your
  18984. 11:16:24input to some useful representation.
  18985. 11:16:26This means that you'll try to understand
  18986. 11:16:28the correlations in your language and
  18987. 11:16:30it'll also include analyzing different
  18988. 11:16:32aspects of the language. All right, so
  18989. 11:16:34this is majorly about understanding your
  18990. 11:16:36text. When it comes to natural language
  18991. 11:16:38generation, here you'll understand how
  18992. 11:16:40to generate text by having a brief plan
  18993. 11:16:43about the text. You'll have sentence
  18994. 11:16:45planning and you'll have text
  18995. 11:16:46realization. Now, natural language
  18996. 11:16:48generation will basically break down
  18997. 11:16:50sentences or will break down text in
  18998. 11:16:52order to understand it better. Okay,
  18999. 11:16:54that's what natural language generation
  19000. 11:16:56is. Natural language understanding is
  19001. 11:16:58more about analyzing your language or
  19002. 11:17:00analyzing the text that you have at hand
  19003. 11:17:02and predicting some useful outcome out
  19004. 11:17:04of it. Generation is more focused on the
  19005. 11:17:07planning aspect of your text. So, these
  19006. 11:17:10are the different components of natural
  19007. 11:17:11language processing. Now, let's look at
  19008. 11:17:13what is stemming and lemmatization in
  19009. 11:17:16natural language processing. Now, what
  19010. 11:17:18is stemming? It is an algorithm which
  19011. 11:17:20works by cutting off the end or the
  19012. 11:17:22beginning of the word and only taking
  19013. 11:17:25into account a list of common prefixes
  19014. 11:17:27and suffixes that can be found in
  19015. 11:17:29inflicted words. Now, for example, on
  19016. 11:17:32the screen you can see that there is a
  19017. 11:17:34detections, detected, detection, and
  19018. 11:17:36detecting.
  19019. 11:17:38Now, if you apply stemming on these four
  19020. 11:17:40words, it will lead to detect. Okay,
  19021. 11:17:43because at the end of the day,
  19022. 11:17:44detections, detected, detection, and
  19023. 11:17:46detecting is the same thing as detect.
  19024. 11:17:48So, stemming will help you remove all of
  19025. 11:17:50these unwanted prefixes and suffixes.
  19026. 11:17:53This way you can analyze the importance
  19027. 11:17:55of the word. All right, you don't have
  19028. 11:17:57to have extra suffix or prefix before
  19029. 11:17:59the word. Now, sometimes during
  19030. 11:18:01stemming, cutting off the ends of the
  19031. 11:18:03words will form an inaccurate result.
  19032. 11:18:05Okay, that's why we have lemmatization.
  19033. 11:18:08In lemmatization, the most important
  19034. 11:18:10thing is the morphological analysis of
  19035. 11:18:12the word. Okay, so here, in order to
  19036. 11:18:14perform lemmatization, you have to have
  19037. 11:18:17a detailed dictionaries which the
  19038. 11:18:18algorithm can look through and it can
  19039. 11:18:21form back to its lemma.
  19040. 11:18:22So, the main difference between stemming
  19041. 11:18:24and lemmatization is that stemming will
  19042. 11:18:26just crop the prefix and the suffix,
  19043. 11:18:28whereas lemmatization will try to
  19044. 11:18:30understand the word in a grammatical way
  19045. 11:18:33and give you an actual word as the
  19046. 11:18:34output. Next is to explain the fuzzy
  19047. 11:18:37logic architecture. All right, so the
  19048. 11:18:39fuzzy logic architecture looks like what
  19049. 11:18:42is shown on the screen. Okay, so
  19050. 11:18:44basically the input is fed into
  19051. 11:18:46something known as the fuzzifier. Okay,
  19052. 11:18:48the fuzzifier or the fuzzification
  19053. 11:18:50module will transform the system's input
  19054. 11:18:53into a number of fuzzy sets. Okay, after
  19055. 11:18:55that it's fed to the controller. Now,
  19056. 11:18:57the controller will have knowledge base
  19057. 11:18:59and the inference engine. Knowledge base
  19058. 11:19:01is basically a set of rules or you can
  19059. 11:19:03say it's an algorithm which is provided
  19060. 11:19:06by experts. Inference engine, like the
  19061. 11:19:08name suggests, will basically infer
  19062. 11:19:10meaning out of these rules. Okay, so
  19063. 11:19:12once you've applied the rules to your
  19064. 11:19:14input, you'll have to draw some useful
  19065. 11:19:16insights or you'll have to infer these
  19066. 11:19:18inputs. Okay, for that you use the
  19067. 11:19:20inference engine. After that, whatever
  19068. 11:19:23inferences and analysis you've formed
  19069. 11:19:25from your inference engine is passed on
  19070. 11:19:26to the defuzzification module. Now, the
  19071. 11:19:29defuzzification will just give you a
  19072. 11:19:31crisp output. All right, it'll give you
  19073. 11:19:33a clear and cut output. That is the
  19074. 11:19:36whole fuzzy logic architecture.
  19075. 11:19:38Now, let's understand the components of
  19076. 11:19:40an expert system. Now, there are three
  19077. 11:19:42important components in an expert
  19078. 11:19:44system, which is knowledge base,
  19079. 11:19:46inference engine, and user interface.
  19080. 11:19:48Now, like I mentioned in fuzzy logic,
  19081. 11:19:50the knowledge base and inference engine
  19082. 11:19:52will play the same part. The user
  19083. 11:19:54interface is basically to provide
  19084. 11:19:56interaction between the users of the
  19085. 11:19:58expert system and the expert system.
  19086. 11:20:01Okay, the expert system is basically a
  19087. 11:20:03program that helps in decision-making
  19088. 11:20:05process. Okay, so here the knowledge
  19089. 11:20:07base will contain some high-quality
  19090. 11:20:09knowledge or it contain rules and
  19091. 11:20:11algorithms. The inference engine will
  19092. 11:20:13acquire all the knowledge that is needed
  19093. 11:20:16to solve the problem. And the user
  19094. 11:20:18interface is just for the users to
  19095. 11:20:19interact with the expert system. Okay,
  19096. 11:20:21this is the whole expert system
  19097. 11:20:23component. Now, obviously this This a
  19098. 11:20:25little more complex than this, but uh
  19099. 11:20:27stick to how this works. All right, I'm
  19100. 11:20:29just going to tell you the working of
  19101. 11:20:30expert systems and fuzzy logic. If I
  19102. 11:20:33start to explain each and everything,
  19103. 11:20:35it's going to take a lot of time. All
  19104. 11:20:36right, so let's move on to our next
  19105. 11:20:38question, which is how is computer
  19106. 11:20:40vision and AI related? Now, computer
  19107. 11:20:43vision is a field of artificial
  19108. 11:20:45intelligence that is used to obtain
  19109. 11:20:47information from images or
  19110. 11:20:49multi-dimensional data. Now, computer
  19111. 11:20:51vision is basically the concept behind
  19112. 11:20:53the self-driving cars that you see these
  19113. 11:20:55days. All right, computer vision
  19114. 11:20:57involves a lot of image processing. So,
  19115. 11:20:59machine learning algorithms like K-means
  19116. 11:21:01can be used in image segmentation.
  19117. 11:21:03Support vector machines can be used for
  19118. 11:21:05image classification. Okay, that's how
  19119. 11:21:07computer vision and AI are related.
  19120. 11:21:09Because most of the things that happen
  19121. 11:21:11in computer vision like image processing
  19122. 11:21:13and segmentation make use of machine
  19123. 11:21:15learning algorithms like K-means and
  19124. 11:21:17support vector machines. So, to sum it
  19125. 11:21:19up, computer vision makes use of
  19126. 11:21:21artificial intelligence technologies to
  19127. 11:21:24solve complex problems such as object
  19128. 11:21:26detection, image processing, and so on.
  19129. 11:21:28That is the relationship between
  19130. 11:21:30computer vision and AI. Now, question
  19131. 11:21:32number 30 is which is better for image
  19132. 11:21:35classification? Is it supervised or
  19133. 11:21:38unsupervised classification? So, guys,
  19134. 11:21:40earlier in the session we discussed what
  19135. 11:21:42supervised learning is and what
  19136. 11:21:43unsupervised learning is. In supervised
  19137. 11:21:46learning, the images are interpreted
  19138. 11:21:48manually by the machine learning expert
  19139. 11:21:50to create feature classes. Now, what
  19140. 11:21:52this means is you're manually going to
  19141. 11:21:54feed a labeled set of data to the
  19142. 11:21:56supervised learning model. All right,
  19143. 11:21:58that's how supervised learning works.
  19144. 11:22:00You're manually going to feed a set of
  19145. 11:22:02images which are labeled to the
  19146. 11:22:04classifier. In unsupervised learning,
  19147. 11:22:06the machine learning software creates
  19148. 11:22:08feature classes based on image pixel
  19149. 11:22:10values. So, basically in unsupervised
  19150. 11:22:12classification, the model itself has to
  19151. 11:22:15figure out what to do and what not to
  19152. 11:22:17do. Okay, so it'll create a own feature
  19153. 11:22:19class based on some values such as image
  19154. 11:22:22pixels or it can also use the image
  19155. 11:22:24color or it can use intensity factors in
  19156. 11:22:27order to classify. So, if you ask me it
  19157. 11:22:29is better to opt for supervised
  19158. 11:22:31classification because you're manually
  19159. 11:22:33inputting images with a lot more
  19160. 11:22:35information. Okay, whereas in
  19161. 11:22:37unsupervised learning you're totally
  19162. 11:22:38letting the model perform everything.
  19163. 11:22:40Okay, so in image classification, I
  19164. 11:22:42think it's better to go for supervised
  19165. 11:22:44learning. Now, let's look at question
  19166. 11:22:46number 31. The next question is finite
  19167. 11:22:50difference filters in image processing
  19168. 11:22:51are very susceptible to noise. To cope
  19169. 11:22:54up with this, which method can you use
  19170. 11:22:56so that there would be minimal
  19171. 11:22:58distortions by noise? Now, the noise in
  19172. 11:23:01an image can be due to high intensity or
  19173. 11:23:03high contrast. Okay, so if you increase
  19174. 11:23:06the contrast and increase the intensity
  19175. 11:23:08of an image, you won't be able to
  19176. 11:23:10understand each pixel. Okay, so each
  19177. 11:23:12pixel will have a value associated to it
  19178. 11:23:15and if the intensity and the contrast of
  19179. 11:23:17that pixel is a little too much, it'll
  19180. 11:23:19be hard for us to understand the image
  19181. 11:23:21properly. It'll be hard to perform image
  19182. 11:23:24analysis because we don't have a clear
  19183. 11:23:26image. Contrast and intensity will just
  19184. 11:23:28cause noise in an image. So, the best
  19185. 11:23:31method to remove this is image
  19186. 11:23:32smoothing. Okay, it is used for reducing
  19187. 11:23:35noise by forcing pixels to be more like
  19188. 11:23:38their neighbors. Okay, this way you'll
  19189. 11:23:40have a faded image or you'll have a more
  19190. 11:23:42equalized image. Now, the next question
  19191. 11:23:45is how is game theory and AI related? So
  19192. 11:23:48guys, AI is actually applied in a vast
  19193. 11:23:51number of fields. Okay, so a lot of
  19194. 11:23:53fields from computer vision to game
  19195. 11:23:55theory to machine learning, AI is always
  19196. 11:23:58a concept behind these fields. Most of
  19197. 11:24:01the game examples that we see make use
  19198. 11:24:03of reinforcement learning or deep neural
  19199. 11:24:05networks. Now, deep neural networks and
  19200. 11:24:07reinforcement learning are very closely
  19201. 11:24:09related to AI because they are branches
  19202. 11:24:11of machine learning. So, machine
  19203. 11:24:13learning is majorly involved in game
  19204. 11:24:15theory. An example of this is in Dota 2
  19205. 11:24:18also they make use of machine learning.
  19206. 11:24:20So, game theory is just a very logical
  19207. 11:24:23approach to solving a problem. And
  19208. 11:24:25machine learning is the best way to
  19209. 11:24:27implement game theory. Now, question
  19210. 11:24:29number three is what is the minimax
  19211. 11:24:31algorithm? Explain the terminologies
  19212. 11:24:33involved in the problem.
  19213. 11:24:35Now guys, minimax is one of the main
  19214. 11:24:37algorithms which is used in game theory.
  19215. 11:24:39All right, it is used to choose an
  19216. 11:24:41optimal move for a player assuming that
  19217. 11:24:43the other player is also playing
  19218. 11:24:45optimally. Meaning that both of these
  19219. 11:24:47players are playing in order to win and
  19220. 11:24:50you're going to use the minimax
  19221. 11:24:51algorithm on one of these players so
  19222. 11:24:53that they choose the optimal move. In
  19223. 11:24:56order to understand the minimax
  19224. 11:24:57algorithm, you need to know what are the
  19225. 11:24:59components in a game. Okay, there's
  19226. 11:25:01something known as game tree. It is
  19227. 11:25:03basically a tree structure which
  19228. 11:25:05contains all the possible moves in a
  19229. 11:25:07game. If it's up, down, right, left, any
  19230. 11:25:09strategy, everything is mentioned in the
  19231. 11:25:11game tree. Now, initial state is
  19232. 11:25:13obviously the initial position of the
  19233. 11:25:15player on the board. All right, the
  19234. 11:25:17successor function it defines all the
  19235. 11:25:19possible moves that a player can make.
  19236. 11:25:22We'll understand this in the next
  19237. 11:25:23question itself, so don't worry if you
  19238. 11:25:25haven't understood this properly.
  19239. 11:25:27Terminal state is obviously the end of
  19240. 11:25:29the game. It's basically the state which
  19241. 11:25:31will lead to the end game or it will
  19242. 11:25:33lead to your destination. Utility
  19243. 11:25:35function is a numerical value for the
  19244. 11:25:37output of the game. So guys, these were
  19245. 11:25:39the terminologies and this is what the
  19246. 11:25:41minimax algorithm is. It is basically a
  19247. 11:25:44game theory algorithm which helps a
  19248. 11:25:46player choose the best optimal policy in
  19249. 11:25:49order to win a game. I'll explain this
  19250. 11:25:51in more depth in the upcoming slides.
  19251. 11:25:54So, let's move on. Now, the next couple
  19252. 11:25:56of questions are going to be
  19253. 11:25:57scenario-based questions. Now, such
  19254. 11:25:59questions are very important in an
  19255. 11:26:01interview because this is where the
  19256. 11:26:03interviewer will understand how well you
  19257. 11:26:05know the concepts. So, the first
  19258. 11:26:07question is show the working of the
  19259. 11:26:09minimax algorithm using the tic-tac-toe
  19260. 11:26:11game. Now, one of the major applications
  19261. 11:26:14of the minimax algorithm is the
  19262. 11:26:16tic-tac-toe game. Okay, you can
  19263. 11:26:18understand and analyze all the possible
  19264. 11:26:20outcomes of the tic-tac-toe game by
  19265. 11:26:22using the minimax algorithm. Let's see
  19266. 11:26:24how this happens. Now, first of all, in
  19267. 11:26:27a minimax algorithm or in a game, there
  19268. 11:26:29are two players involved. Okay, the max
  19269. 11:26:32is the player that tries to get the
  19270. 11:26:33highest possible score, and min is the
  19271. 11:26:36player that tries to get the lowest
  19272. 11:26:37possible score. So, this algorithm is
  19273. 11:26:40designed in such a way that assuming
  19274. 11:26:42that there going to be two players, and
  19275. 11:26:44obviously one player is going to win the
  19276. 11:26:45game, and that is the max player, and
  19277. 11:26:48min is the player which loses the game
  19278. 11:26:50and has the lowest possible score. Now,
  19279. 11:26:52the first step in the minimax algorithm
  19280. 11:26:54is to generate the entire game tree.
  19281. 11:26:57Okay, the game tree is all the possible
  19282. 11:26:59outcomes that can happen in tic-tac-toe.
  19283. 11:27:01Okay, in the figure you can see that
  19284. 11:27:02first X is aligned in the first box,
  19285. 11:27:04then in the second box, third box, and
  19286. 11:27:06so on. All the possible actions that you
  19287. 11:27:09can take in a tic-tac-toe game are put
  19288. 11:27:11in this game tree. And then, step number
  19289. 11:27:14two is to apply the utility function to
  19290. 11:27:16get the utility values from all the
  19291. 11:27:18terminal states. Getting utility value
  19292. 11:27:21is important because this is how you'll
  19293. 11:27:23understand your outcome. Okay, you'll
  19294. 11:27:24understand if you're going to win or
  19295. 11:27:25lose. Now, in the terminal states,
  19296. 11:27:28whatever numbers you see over here,
  19297. 11:27:30these are the utility values. Now, step
  19298. 11:27:32three is determine the utilities of the
  19299. 11:27:34higher nodes with the help of utilities
  19300. 11:27:36of the terminal nodes. Now, in this
  19301. 11:27:39diagram, you can see that in the
  19302. 11:27:40terminal nodes, we have the utility
  19303. 11:27:42values. The step three is to get utility
  19304. 11:27:45values in the higher stages, which is
  19305. 11:27:48the min stage. All right, these two
  19306. 11:27:50circles, you need to fill in the utility
  19307. 11:27:51values by using the utility values which
  19308. 11:27:54are in the terminal state.
  19309. 11:27:55Now, how do you calculate the utility
  19310. 11:27:57value? Let's start by calculating the
  19311. 11:28:00utility value of the left node. Okay,
  19312. 11:28:02this red color node, we'll start by
  19313. 11:28:05calculating this.
  19314. 11:28:06Now, you calculate that by finding the
  19315. 11:28:08minimum of the three nodes that it's
  19316. 11:28:11leading to. Now, this red node is
  19317. 11:28:13leading to three, five, and 10. And the
  19318. 11:28:15minimum out of three, five, 10 is three.
  19319. 11:28:17So, the utility value for this red node
  19320. 11:28:20is going to be three. Okay, similarly
  19321. 11:28:22for this green node, it's going to be
  19322. 11:28:23two because the minimum value between
  19323. 11:28:25two and two is still two. Now, step four
  19324. 11:28:27is to fill in these utility values that
  19325. 11:28:29you've calculated. So, now we have a
  19326. 11:28:32minimax algorithm which has all the
  19327. 11:28:34utility values filled in. Now, the only
  19328. 11:28:36utility value which isn't filled is the
  19329. 11:28:38one with max. Okay, the one on the root
  19330. 11:28:41node. Here, we haven't filled the
  19331. 11:28:43utility value. Again, to fill this
  19332. 11:28:45value, you're going to check the nodes
  19333. 11:28:46which are directly connected to it,
  19334. 11:28:48which is three and two. You'll find the
  19335. 11:28:50maximum between these two because this
  19336. 11:28:52is the max function. All right, so here
  19337. 11:28:54you'll get a value of three. So, that's
  19338. 11:28:57why the best opening move for max is the
  19339. 11:28:59left node. Okay, you can make use of the
  19340. 11:29:02left node in order to win the game. This
  19341. 11:29:04is the first step that the max player
  19342. 11:29:06has to take in order to get to the path
  19343. 11:29:08of winning the game. So guys, by doing
  19344. 11:29:11this for each and every step, you can
  19345. 11:29:13win the game. Okay, so you'll have to
  19346. 11:29:15calculate the utility value at the
  19347. 11:29:17terminal nodes. You'll have to move up
  19348. 11:29:19to the other hierarchical nodes above
  19349. 11:29:21it, calculate the utility values there
  19350. 11:29:23until you reach the root node. Okay,
  19351. 11:29:25once you reach the root node, you'll get
  19352. 11:29:26a utility value and that utility value
  19353. 11:29:29will be connected to some move or some
  19354. 11:29:31node. You'll have to take that node or
  19355. 11:29:34you'll have to take that move in the
  19356. 11:29:36game in order to win the game. So, this
  19357. 11:29:38way you'll have to calculate the utility
  19358. 11:29:40value for each and every move that the
  19359. 11:29:42player makes so that the player will win
  19360. 11:29:44the game.
  19361. 11:29:45So guys, minimax algorithm is quite easy
  19362. 11:29:47and it's very understandable. All you
  19363. 11:29:49need to know is a little bit of math in
  19364. 11:29:51order to solve this problem. Question
  19365. 11:29:52number 35 is which method is used for
  19366. 11:29:56optimizing a minimax based game? Now,
  19367. 11:29:58this is not a scenario-based question,
  19368. 11:30:00but this question is usually asked if an
  19369. 11:30:02interviewer asks you about a minimax
  19370. 11:30:05game.
  19371. 11:30:05Now, the best way to optimize a minimax
  19372. 11:30:08game is by using something known as
  19373. 11:30:10alpha-beta pruning. Now, the main thing
  19374. 11:30:12about alpha-beta pruning is that it'll
  19375. 11:30:14remove all the nodes that are not
  19376. 11:30:16affecting the final decision. It's just
  19377. 11:30:18a faster way to reach your outcome.
  19378. 11:30:21That's what alpha-beta pruning is all
  19379. 11:30:23about. So, let's look at an example to
  19380. 11:30:26understand this. Okay, let's say there
  19381. 11:30:27was another node over here. Okay, here
  19382. 11:30:29you can see that this is going down to a
  19383. 11:30:31terminal state with utility value two.
  19384. 11:30:34Okay, now you don't know the value of
  19385. 11:30:36the other two nodes, but if you use
  19386. 11:30:38minimax to calculate the utility of the
  19387. 11:30:41other two nodes, you'll get a value of
  19388. 11:30:43three. So, in this example again, we'll
  19389. 11:30:45start at the terminal nodes. So, three,
  19390. 11:30:47five, 10 are the utility values here.
  19391. 11:30:50So, this will give us a value of three
  19392. 11:30:52because we're calculating the minimum
  19393. 11:30:53over here. Now, here you have two and
  19394. 11:30:56you have two unknown values. You have A
  19395. 11:30:58or B. Okay, I've named them as A and B.
  19396. 11:31:01Okay, let's leave this for now. Let's go
  19397. 11:31:02to the next node. Okay, here the
  19398. 11:31:04possibilities are two, seven, and three.
  19399. 11:31:07So, the minimum between two, seven,
  19400. 11:31:08three is two.
  19401. 11:31:10Okay, so here there's going to be three.
  19402. 11:31:11There's going to be a value, let's say
  19403. 11:31:13C, and here there's going to be a value,
  19404. 11:31:15let's say two. Now, we know that the
  19405. 11:31:18maximum between three, C, and something
  19406. 11:31:20else will be three. Okay, that's because
  19407. 11:31:23two is the minimum value over here, and
  19408. 11:31:25the maximum will obviously be three. So,
  19409. 11:31:27the hint here is in the two AB node. We
  19410. 11:31:30know that the value or the utility value
  19411. 11:31:32will obviously be equal to two or it'll
  19412. 11:31:35be less than two because you're
  19413. 11:31:36calculating the minimum in this step.
  19414. 11:31:38Now, if you calculate the max out of
  19415. 11:31:40these three values, we'll obviously get
  19416. 11:31:42the answer as three.
  19417. 11:31:44So, this way this entire node itself is
  19418. 11:31:46removed because you don't need it to get
  19419. 11:31:48to the final answer. Okay, that's what
  19420. 11:31:50alpha-beta pruning is all about. It'll
  19421. 11:31:52identify the nodes which are not going
  19422. 11:31:54to affect the final decisions, and it'll
  19423. 11:31:56just remove those nodes. So guys, this
  19424. 11:31:58is how the optimization for a minimax
  19425. 11:32:01game is done. It's done using the
  19426. 11:32:03alpha-beta pruning.
  19427. 11:32:04The next question is which algorithm
  19428. 11:32:06does Facebook use for face verification?
  19429. 11:32:09Now guys, even though this might seem
  19430. 11:32:11like a general knowledge question, this
  19431. 11:32:14is actually a very important sort of
  19432. 11:32:16question in artificial intelligence.
  19433. 11:32:18Okay, even if you don't know the answer
  19434. 11:32:20to this, you should have an idea of how
  19435. 11:32:22the algorithm might work. Okay, that's
  19436. 11:32:24exactly what the interviewer wants to
  19437. 11:32:26know. He wants to know whether you know
  19438. 11:32:28how the algorithm works step-by-step.
  19439. 11:32:30You might not know the final answer or
  19440. 11:32:32you might not know the exact algorithm
  19441. 11:32:34which Facebook uses because obviously
  19442. 11:32:35Facebook uses more than one algorithm to
  19443. 11:32:38achieve this, but you must know the
  19444. 11:32:40steps in which the face verification
  19445. 11:32:42works. Okay, that's the main goal behind
  19446. 11:32:44this question. Now anyway, the algorithm
  19447. 11:32:47used by Facebook is the deep face. Okay,
  19448. 11:32:49deep face makes use of a lot of neural
  19449. 11:32:51networks and a lot of algorithms. Okay,
  19450. 11:32:53so it works on artificial intelligence
  19451. 11:32:56techniques, like I mentioned earlier.
  19452. 11:32:58Now how would a face verification work?
  19453. 11:33:00How do you think it works? Now it starts
  19454. 11:33:02by an input. So the idea here is you
  19455. 11:33:05have to scan a huge number of photos and
  19456. 11:33:07you'll have to feed it to the algorithm.
  19457. 11:33:10Okay, now these photos can have a lot of
  19458. 11:33:12disturbance, a lot of distortions and it
  19459. 11:33:14can have different angles or anything
  19460. 11:33:16like that. Okay, you have to feed any
  19461. 11:33:18sort of photos that are possible. Okay,
  19462. 11:33:20even they are complex to understand, but
  19463. 11:33:22you have to still feed the model with
  19464. 11:33:24all the possible photos that you can
  19465. 11:33:25get. Now the next step is the main
  19466. 11:33:27process. Here there are a few important
  19467. 11:33:30things which is detect, align, represent
  19468. 11:33:32and classify. Detect is basically you'll
  19469. 11:33:34detect facial features. All right,
  19470. 11:33:37you'll try to understand the distance
  19471. 11:33:38between the eyes and the nose of a
  19472. 11:33:40person, the way the lips is aligned or
  19473. 11:33:43anything like that. That's what aligning
  19474. 11:33:45is about. You'll align and compare the
  19475. 11:33:46various features in the face in order to
  19476. 11:33:49understand the facial features. You'll
  19477. 11:33:51represent the key patterns by using some
  19478. 11:33:533D graphs or 3D models. Okay, it's very
  19479. 11:33:55important to visualize whatever you get
  19480. 11:33:58because visualization will help you
  19481. 11:34:00understand the correlation. It'll help
  19482. 11:34:02you understand that okay, the eyes are
  19483. 11:34:03at this distance, the nose is at this
  19484. 11:34:05distance, and so on. Finally, you'll
  19485. 11:34:07classify the images based on the
  19486. 11:34:09similarity. All right, that's how the
  19487. 11:34:11output comes out. And basically, the
  19488. 11:34:13output is you need to detect whether two
  19489. 11:34:15images represent the same person or not.
  19490. 11:34:18Okay, so by studying the facial features
  19491. 11:34:20and by using image processing and by
  19492. 11:34:22using computer vision, Facebook's
  19493. 11:34:24achieves face verification. You need to
  19494. 11:34:26know the basic concept behind face
  19495. 11:34:28verification. You need to know that it
  19496. 11:34:30starts with image collection or data
  19497. 11:34:32acquisition. After that, you're going to
  19498. 11:34:34perform image processing or
  19499. 11:34:36pre-processing. All right, and this
  19500. 11:34:38might involve performing conversions
  19501. 11:34:40from RGB to any other state like YCbCr.
  19502. 11:34:44Okay, I'm not going to go in depth of
  19503. 11:34:45this because the video will get to about
  19504. 11:34:472-3 hours. So, there are a lot of ways
  19505. 11:34:49in which you can convert an image and
  19506. 11:34:51you know, you can understand the image
  19507. 11:34:52more properly. Also, an important thing
  19508. 11:34:54in image processing is it's not done
  19509. 11:34:57just based on the image. All right.
  19510. 11:34:59You're going to take the image, you're
  19511. 11:35:00going to form a matrix, and you're going
  19512. 11:35:02to have pixel values in these matrix.
  19513. 11:35:04So, it is a very in-depth approach. All
  19514. 11:35:06right, it's not a very simple approach.
  19515. 11:35:08When I'm speaking about it, it might
  19516. 11:35:10seem simple, but image analysis is very
  19517. 11:35:13in-depth. After image analysis, you can
  19518. 11:35:15perform image segmentation. All right,
  19519. 11:35:17image segmentation is basically dividing
  19520. 11:35:19the image into different segments and
  19521. 11:35:21studying each image segment separately.
  19522. 11:35:24Then after that, you can do feature
  19523. 11:35:25extraction. Here, you'll try to
  19524. 11:35:27understand the features and how they are
  19525. 11:35:29related to each other. Finally, you'll
  19526. 11:35:31classify the images and see whether two
  19527. 11:35:34images represent the same person or not.
  19528. 11:35:37So, the main idea behind the Facebook
  19529. 11:35:39algorithm is image processing, neural
  19530. 11:35:41networks, machine learning, and computer
  19531. 11:35:43vision. All right, and all of this comes
  19532. 11:35:45down to artificial intelligence. Next,
  19533. 11:35:48we have explain the logic behind
  19534. 11:35:50targeted marketing and how can machine
  19535. 11:35:52learning help with this? Now, target
  19536. 11:35:54marketing is something that we see very
  19537. 11:35:56often. All right, let's say that you
  19538. 11:35:58were looking for some shoe on Amazon. In
  19539. 11:36:01a day or two, you just open up YouTube
  19540. 11:36:03and Facebook. You'll see that you'll get
  19541. 11:36:05ads of shoes from Amazon. Okay, this is
  19542. 11:36:08targeted marketing. So, basically Amazon
  19543. 11:36:10knows that we've been looking for a
  19544. 11:36:12particular type of shoe, so it's going
  19545. 11:36:14to target you with that particular ad.
  19546. 11:36:16Okay, this is what targeted marketing is
  19547. 11:36:18in short. Targeted marketing can be done
  19548. 11:36:20in different ways. For example, it can
  19549. 11:36:23be done depending on your geography or
  19550. 11:36:25it can be done depending on your social
  19551. 11:36:27economic profile. Okay, let's say that
  19552. 11:36:30Amazon has details about your age, it
  19553. 11:36:33has details about what sport you like to
  19554. 11:36:35play. Let's say that you've been
  19555. 11:36:36browsing through a lot of sports. Okay,
  19556. 11:36:39you've been browsing through a lot of
  19557. 11:36:40sport equipments or something like that.
  19558. 11:36:43Amazon will know that you're interested
  19559. 11:36:44in this by using machine learning, of
  19560. 11:36:46course, and it will send you ads based
  19561. 11:36:49on what you're interested in. Okay, this
  19562. 11:36:51is what target marketing really is.
  19563. 11:36:53Now, how does machine learning come into
  19564. 11:36:55target marketing? Okay, so there's
  19565. 11:36:57something known as text analytics
  19566. 11:36:59systems. Now, the applications for text
  19567. 11:37:01analytics ranges from search
  19568. 11:37:03applications, text classification, named
  19569. 11:37:06entity recognition, or pattern search.
  19570. 11:37:09Okay, so it's basically a way to
  19571. 11:37:10understand what you're looking for.
  19572. 11:37:12Okay, they'll try to understand your
  19573. 11:37:14search history and they'll try to target
  19574. 11:37:16you by using your interests. Clustering
  19575. 11:37:19is another way of targeted marketing.
  19576. 11:37:21All right, you'll cluster customers who
  19577. 11:37:23have similar interests and you'll send
  19578. 11:37:25them similar ads or you'll send them
  19579. 11:37:26similar offers. Classification is
  19580. 11:37:28another method used for targeted
  19581. 11:37:30marketing. Now, here you'll make use of
  19582. 11:37:33algorithms like decision trees and
  19583. 11:37:34neural networks. Now, recommender
  19584. 11:37:36systems is what I spoke about earlier.
  19585. 11:37:38When it comes to Amazon, they recommend
  19586. 11:37:40items to you based on your interest or
  19587. 11:37:43based on people who have similar
  19588. 11:37:44interests like you. Market basket
  19589. 11:37:47analysis is another method that machine
  19590. 11:37:49learning uses for marketing. Okay, here
  19591. 11:37:51basically you'll understand a
  19592. 11:37:53combination of products that are
  19593. 11:37:54frequently bought. Okay, by
  19594. 11:37:56understanding what two products are
  19595. 11:37:58frequently bought, you can give some
  19596. 11:37:59offers or you can give some discounts on
  19597. 11:38:01those products so that people buy more
  19598. 11:38:03and more. Okay, that's how market basket
  19599. 11:38:05analysis also works. So guys, this is
  19600. 11:38:07what targeted marketing is.
  19601. 11:38:10Now next is how can AI be used to detect
  19602. 11:38:13fraud? AI is used in a lot of ways in
  19603. 11:38:16credit card fraud detection. It's used
  19604. 11:38:18in detecting anomalies and all of that.
  19605. 11:38:20It basically makes use of machine
  19606. 11:38:22learning algorithms to do this. Now
  19607. 11:38:24let's try to understand how this process
  19608. 11:38:26works. Okay, first it begins with data
  19609. 11:38:28extraction or data collection. So at
  19610. 11:38:30this stage data is either collected
  19611. 11:38:32through a survey or through web
  19612. 11:38:34scraping. Okay, if you're trying to
  19613. 11:38:35detect credit card fraud then
  19614. 11:38:37information about the customer's
  19615. 11:38:39collected. All right, this includes any
  19616. 11:38:41transactional or any shopping and
  19617. 11:38:43personal details. Next is data cleaning.
  19618. 11:38:45So at this stage the redundant data must
  19619. 11:38:48be removed. Any inconsistencies or any
  19620. 11:38:50missing values that you have in your
  19621. 11:38:52data, it has to be removed because they
  19622. 11:38:54lead to wrongful prediction. Okay, so
  19623. 11:38:56therefore you have to get rid of any
  19624. 11:38:58inconsistencies in this stage. Next we
  19625. 11:39:01have data exploration and analysis. Now
  19626. 11:39:03this is the most important step in AI.
  19627. 11:39:06Okay, here you study the relationship
  19628. 11:39:08between various predictable variables.
  19629. 11:39:10For example, let's say that a person has
  19630. 11:39:12spent an unusual sum of money on a
  19631. 11:39:15particular day. Now the chances for a
  19632. 11:39:17fraudulent occurrence is very high
  19633. 11:39:19because usually the person is not used
  19634. 11:39:21to spending this much money. So such
  19635. 11:39:23patterns have to be detected and
  19636. 11:39:24understood in data exploration and
  19637. 11:39:26analysis. This is followed by building a
  19638. 11:39:29machine learning model. Now here there
  19639. 11:39:31are any machine learning algorithms that
  19640. 11:39:33can be used for fraud detection or
  19641. 11:39:35anomaly detection. One such example is
  19642. 11:39:37logistic regression, okay, which is a
  19643. 11:39:39classification algorithm and it can be
  19644. 11:39:42used to classify events into two
  19645. 11:39:44classes. Okay, you can use them to
  19646. 11:39:46classify a person or classify an event
  19647. 11:39:49as either fraudulent and non-fraudulent.
  19648. 11:39:52Then comes model evaluation. Here you'll
  19649. 11:39:54basically test the efficiency of the
  19650. 11:39:56machine learning model. Okay, so if
  19651. 11:39:58there's any room for improvement, then
  19652. 11:40:00you can perform parameter tuning and you
  19653. 11:40:02can improve the model. This will just
  19654. 11:40:04improve the accuracy of the model. So
  19655. 11:40:06guys, all of these complex problems like
  19656. 11:40:09fraud detection or object detection, all
  19657. 11:40:12of this is done through a process. Okay,
  19658. 11:40:14and in general the process is data
  19659. 11:40:16collection, data cleaning, exploration
  19660. 11:40:18and analysis, building a model and model
  19661. 11:40:21evaluation. Most of these complex
  19662. 11:40:23problems can be solved by using this
  19663. 11:40:25approach. Now let's look at our next
  19664. 11:40:27question.
  19665. 11:40:28Okay, a bank manager is given a data set
  19666. 11:40:31containing records of thousands of
  19667. 11:40:33applicants who have applied for loan.
  19668. 11:40:35How can AI help the manager understand
  19669. 11:40:37which loans he can approve?
  19670. 11:40:39To be more specific, this problem
  19671. 11:40:41statement can easily be solved by using
  19672. 11:40:43the KNN algorithm. Okay, KNN is
  19673. 11:40:46basically stands for K nearest neighbor.
  19674. 11:40:49All right, this is a classification and
  19675. 11:40:51a regression algorithm.
  19676. 11:40:53So if you use a KNN algorithm, it'll
  19677. 11:40:55form two classes. One is the loan is
  19678. 11:40:57approved and the other is applicants
  19679. 11:40:59whose loan is not been approved. So like
  19680. 11:41:02I said, K nearest neighbor is a
  19681. 11:41:03supervised learning algorithm that
  19682. 11:41:05classifies a new data point into the
  19683. 11:41:08target class depending on the features
  19684. 11:41:10of its neighboring data points. All
  19685. 11:41:12right, so KNN basically focuses on the
  19686. 11:41:14neighbors and it understands that if a
  19687. 11:41:16new data point is similar to one of its
  19688. 11:41:18neighbors, then it has to classify that
  19689. 11:41:20new data point into that neighbor's
  19690. 11:41:22class. Now again, the methodology for
  19691. 11:41:24solving this problem is same. You start
  19692. 11:41:27by data collection, data cleaning,
  19693. 11:41:29exploration and analysis, building a
  19694. 11:41:31model and model evaluation. All right.
  19695. 11:41:34So, while data collection, you can
  19696. 11:41:35collect data like account balance,
  19697. 11:41:37credit amount, age, occupation, loan
  19698. 11:41:40records, and all of that. So, by using
  19699. 11:41:42this data, you can predict whether or
  19700. 11:41:43not to approve the loan of an applicant.
  19701. 11:41:46Data cleaning, again, you have to remove
  19702. 11:41:48any variables which will not help the
  19703. 11:41:50model. Okay, any variables which will
  19704. 11:41:52just increase the complexity of the
  19705. 11:41:54model. Okay, so you'll remove such
  19706. 11:41:55variables at this stage. In data
  19707. 11:41:58exploration and analysis, you will
  19708. 11:42:00understand the patterns in your data.
  19709. 11:42:02Okay, let's see that a person has a
  19710. 11:42:04history of unpaid loans. Okay, if any
  19711. 11:42:07person or any applicant has a history of
  19712. 11:42:09unpaid loans, then the chances are that
  19713. 11:42:11he might not get approval on his loan
  19714. 11:42:13application. Okay, this is obvious
  19715. 11:42:15because the manager is going to see that
  19716. 11:42:17his previous loans are still due. So,
  19717. 11:42:20that's why he won't be able to approve
  19718. 11:42:21the application. So, these are the kind
  19719. 11:42:23of patterns that are detected in
  19720. 11:42:25exploration and analysis. Now, building
  19721. 11:42:28a machine learning model, you can use n
  19722. 11:42:30number of models when it comes to
  19723. 11:42:31predicting whether an applicant loan
  19724. 11:42:34request is approved or not. Now, like I
  19725. 11:42:36mentioned, one of the easy algorithms
  19726. 11:42:38that you can implement is the K nearest
  19727. 11:42:39algorithm. Okay, it can be used for both
  19728. 11:42:42classification and regression. It
  19729. 11:42:44classify the applicant's loan request
  19730. 11:42:46into approved or disapproved based on
  19731. 11:42:49the socio-economic profile of the
  19732. 11:42:50applicant. Okay, based on variables like
  19733. 11:42:53loan based on variables like the salary,
  19734. 11:42:55the occupation of the applicant. Now,
  19735. 11:42:58model evaluation, again, is the same
  19736. 11:43:00thing. You'll basically evaluate the
  19737. 11:43:02efficiency of the model. You'll try to
  19738. 11:43:03improve the accuracy of the model by
  19739. 11:43:05using parameter tuning or cross
  19740. 11:43:07validation. So, guys, this is how a bank
  19741. 11:43:10manager can understand whether a loan
  19742. 11:43:12can be approved or not. Again, in this
  19743. 11:43:14question, they are just trying to test
  19744. 11:43:16if you know how the flow of the problem
  19745. 11:43:18will go. If you know how this problem
  19746. 11:43:20can be solved. You don't have to know
  19747. 11:43:22the exact details, but you have to know
  19748. 11:43:24how you can approach the problem.
  19749. 11:43:26Okay, now let's move on and look at
  19750. 11:43:28question number 40. Now the question
  19751. 11:43:30here is place an agent in any one of the
  19752. 11:43:32rooms and the goal is to reach outside
  19753. 11:43:35the building. Can this be achieved
  19754. 11:43:37through AI? If yes, explain how it can
  19755. 11:43:39be done. Now in this question there is a
  19756. 11:43:42diagram along with a small explanation.
  19757. 11:43:44Okay, so basically there are four rooms
  19758. 11:43:47in this diagram. Basically 0 1 2 3 and 4
  19759. 11:43:50represent rooms and this 5 represents
  19760. 11:43:54outside the building. Okay, now the goal
  19761. 11:43:56is to place an agent in any one of these
  19762. 11:43:59rooms in such a way that he has to reach
  19763. 11:44:01room number five or he has to reach
  19764. 11:44:03outside the building. Now they've also
  19765. 11:44:05mentioned that a room number one and
  19766. 11:44:07room number four directly lead outside
  19767. 11:44:10the building. That's correct because
  19768. 11:44:11room number one is directly connected to
  19769. 11:44:13five and four is also directly connected
  19770. 11:44:16to five. Right? Four leads outside the
  19771. 11:44:18building. Now if you look at room number
  19772. 11:44:21zero, if you want to go from zero to
  19773. 11:44:23five, first from zero you'll have to go
  19774. 11:44:25to four and then only you can go to
  19775. 11:44:26five. Similarly, if you look at room
  19776. 11:44:28number three, if you want to go from
  19777. 11:44:30three to five, you'll have to take
  19778. 11:44:32three, then you'll have to go to one and
  19779. 11:44:34then only you'll have to go to five.
  19780. 11:44:36These are not directly connected to
  19781. 11:44:38outside the building, whereas room
  19782. 11:44:39number one and four are directly
  19783. 11:44:41connected to outside the building. All
  19784. 11:44:43right, I hope the question is clear. Now
  19785. 11:44:45as soon as you read the question, you
  19786. 11:44:47must know that this is a reinforcement
  19787. 11:44:49learning question. All right, it's
  19788. 11:44:51pretty clear because they have mentioned
  19789. 11:44:53that there is an agent which is going to
  19790. 11:44:55be placed in any one of the rooms and he
  19791. 11:44:57has to basically explore the environment
  19792. 11:45:00and reach five, which is basically
  19793. 11:45:02outside the building. So as soon as you
  19794. 11:45:03read the question, the first thing that
  19795. 11:45:05should come into your head is that this
  19796. 11:45:06is a reinforcement learning problem. Now
  19797. 11:45:09this problem can be solved by using the
  19798. 11:45:11Q-learning algorithm. Now if you
  19799. 11:45:13remember earlier in the session, I
  19800. 11:45:15discussed what exactly Q-learning is and
  19801. 11:45:17how it works. So Q-learning is basically
  19802. 11:45:20a reinforcement learning algorithm which
  19803. 11:45:22is used to solve reward-based problems.
  19804. 11:45:24So, now let's look at how we'll solve
  19805. 11:45:26the problem. First step would be to
  19806. 11:45:29represent the rooms on a graph. All
  19807. 11:45:31right, so each room over here you'll
  19808. 11:45:33represent it as a node and each door
  19809. 11:45:35will represent a link. So, if you look
  19810. 11:45:37at this figure over here, this is our
  19811. 11:45:39original figure that was given in the
  19812. 11:45:41question and now this is the graph that
  19813. 11:45:43we draw from this figure. So, we have
  19814. 11:45:46node one. Let's look at how node one is
  19815. 11:45:49connected to node three. So, basically
  19816. 11:45:51you can go from node one to node three
  19817. 11:45:53and you can go from node three to node
  19818. 11:45:55one. If you look at the diagram, there
  19819. 11:45:57is a direct connection from one to three
  19820. 11:45:59and three to one. Okay, let's look at
  19821. 11:46:01one and two. Now, there's no link
  19822. 11:46:03between one and two because if you look
  19823. 11:46:05over here, if you want to go from room
  19824. 11:46:07number one to room number two, you
  19825. 11:46:08cannot directly go. You'll have to go
  19826. 11:46:11from room number one to room number
  19827. 11:46:12three and only then you'll reach room
  19828. 11:46:14number two. All right, that's why
  19829. 11:46:15there's no link between room number one
  19830. 11:46:18and node number two.
  19831. 11:46:19Similarly, if you look at node one and
  19832. 11:46:22four, they are directly connected to
  19833. 11:46:24five. This is because room number one
  19834. 11:46:26and four directly lead to this goal. Our
  19835. 11:46:29goal is to reach room number five. So,
  19836. 11:46:32one and four are directly connected to
  19837. 11:46:34five, whereas the others are connected
  19838. 11:46:36just like how they're shown in this
  19839. 11:46:38figure. Okay, it's pretty
  19840. 11:46:39understandable, guys. This is just
  19841. 11:46:41logic. Now, let's look at the next step.
  19842. 11:46:43Now, the next step is to associate a
  19843. 11:46:46reward value to each door. What we're
  19844. 11:46:48going to do here is we're going to build
  19845. 11:46:50a reward matrix. Now, I'll tell you what
  19846. 11:46:52that exactly means. So, for the doors
  19847. 11:46:55that lead directly to the end goal,
  19848. 11:46:57which is room number five, you'll assign
  19849. 11:46:59a reward of 100 to those doors. So, if
  19850. 11:47:01you're traversing from node number one
  19851. 11:47:03to node number five, you'll get a reward
  19852. 11:47:05of 100. Similarly, if you're traversing
  19853. 11:47:08from four to five, you'll get a reward
  19854. 11:47:09of 100. Five to five also you'll get a
  19855. 11:47:12reward of 100 because your end goal is
  19856. 11:47:14five, right? So, basically any link that
  19857. 11:47:16leads directly to our end goal, for that
  19858. 11:47:19link, we're going to assign a reward of
  19859. 11:47:20100. Now, doors that are not directly
  19860. 11:47:23connected to the target room will have a
  19861. 11:47:25reward zero. This is because if you take
  19862. 11:47:28room number two or if you take room
  19863. 11:47:29number three, you won't reach room
  19864. 11:47:31number five directly. And our goal here
  19865. 11:47:34is to reach room number five. That's why
  19866. 11:47:36for the other doors, we've given a
  19867. 11:47:38reward of zero. For door number one and
  19868. 11:47:40door number four, however, we have
  19869. 11:47:42rewards of 100. Similarly, for door
  19870. 11:47:44number five or room number five, also we
  19871. 11:47:46have a reward of 100. So, basically,
  19872. 11:47:49each action or each link will represent
  19873. 11:47:52a reward. So, let's say that you're
  19874. 11:47:54traversing from room number one to room
  19875. 11:47:56number four. If you go from room number
  19876. 11:47:59one to room number three and then three
  19877. 11:48:01to four, your reward is going to remain
  19878. 11:48:03zero because you're not reaching the end
  19879. 11:48:05goal here. Only if you traverse from one
  19880. 11:48:07to five or if you traverse from four to
  19881. 11:48:09five or five to five, you'll get a
  19882. 11:48:11reward of 100. All right, it's as simple
  19883. 11:48:14as that. Now, let's see how the
  19884. 11:48:15Q-learning algorithm works in this
  19885. 11:48:18particular problem statement. Now, there
  19886. 11:48:20are two main components in this
  19887. 11:48:22algorithm. All right, there is state and
  19888. 11:48:24there is action. Now, basically, all
  19889. 11:48:26these rooms will represent the state and
  19890. 11:48:28the agent's movement from one room to
  19891. 11:48:30the other will represent an action. So,
  19892. 11:48:33basically, 0 1 2 3 4 and 5 represent the
  19893. 11:48:36state and let's say you're traversing
  19894. 11:48:38from two to three. This two to three
  19895. 11:48:40will basically represent an action. In
  19896. 11:48:42order to make you understand how this
  19897. 11:48:44works, let's say that you're traversing
  19898. 11:48:46from room number two to room number
  19899. 11:48:47five. Your initial state is going to be
  19900. 11:48:50room number two. Your next state is
  19901. 11:48:52going to be room number three. All
  19902. 11:48:53right, so you're moving from two to
  19903. 11:48:55three and you're getting a reward of
  19904. 11:48:56zero. Remember that. Now, from state
  19905. 11:48:58three, you'll either be going to state
  19906. 11:49:00two or you can go to state one or state
  19907. 11:49:03four. All right, if you choose four, you
  19908. 11:49:06can uh directly go to five and here
  19909. 11:49:08you'll get a reward of 100. If you
  19910. 11:49:10choose one, again, you'll go directly to
  19911. 11:49:12five. You'll get a reward of 100. But if
  19912. 11:49:15you go back to two, you'll get a reward
  19913. 11:49:16zero. So guys, let me tell you that the
  19914. 11:49:18agent is going to explore over here,
  19915. 11:49:20okay? He has no idea about the
  19916. 11:49:22environment, so he's not going to go
  19917. 11:49:23from two, three, one, five directly, all
  19918. 11:49:26right? He's not going to know that this
  19919. 11:49:28will lead to the output. He has to
  19920. 11:49:30explore, he has to make mistakes, he has
  19921. 11:49:32to learn, and he has to find out the
  19922. 11:49:34best path to reach room number five. Now
  19923. 11:49:36next is our reward matrix. So guys,
  19924. 11:49:39there are two main matrices in
  19925. 11:49:40Q-learning algorithm. One is a reward
  19926. 11:49:42matrix, and the other is going to be the
  19927. 11:49:44Q matrix or the memory matrix. All
  19928. 11:49:47right, I'll be discussing the memory
  19929. 11:49:48matrix in a while, but for now let's
  19930. 11:49:50look at the reward matrix. So what I'm
  19931. 11:49:53doing here is I'm basically putting all
  19932. 11:49:55the reward values for traversing from
  19933. 11:49:57one node to the other node in a matrix
  19934. 11:49:59known as the reward matrix. Now the
  19935. 11:50:01minus one will basically represent the
  19936. 11:50:03null values. What I'm trying to say is
  19937. 11:50:06there is no connection from zero to
  19938. 11:50:07zero, all right? That's why I'm giving a
  19939. 11:50:09value of minus one. But you might say
  19940. 11:50:11that why is the reward from five to five
  19941. 11:50:14hundred? Now this is because if you go
  19942. 11:50:16from room number five to room number
  19943. 11:50:18five, you're still reaching the end
  19944. 11:50:20goal, all right? That's why there's an a
  19945. 11:50:22reward of 100 over here. Let's look at
  19946. 11:50:24zero {comma} four, all right? There is a
  19947. 11:50:26link from zero to four, but the reward
  19948. 11:50:28is going to be zero because four is not
  19949. 11:50:31the end goal, all right? Four is not
  19950. 11:50:33your goal room or anything, that's why
  19951. 11:50:35your reward is going to be zero. Now the
  19952. 11:50:37reward is 100 only for one {comma} five
  19953. 11:50:39that is if you traverse from one to
  19954. 11:50:41five, for four {comma} five, which is if
  19955. 11:50:44you traverse from four to five, and five
  19956. 11:50:46{comma} five. Now this is because
  19957. 11:50:48through all these three actions you're
  19958. 11:50:49going back to room number five, all
  19959. 11:50:51right? Which is your end goal, that's
  19960. 11:50:53why we have a reward of 100 for these
  19961. 11:50:55three actions, all right? So like I
  19962. 11:50:57mentioned earlier, we're going to have
  19963. 11:50:59another matrix known as a Q matrix,
  19964. 11:51:01which will basically represent the
  19965. 11:51:03memory of what the agent has learned
  19966. 11:51:05through experience. Okay, that's the
  19967. 11:51:07only way the agent will actually learn
  19968. 11:51:09further. If the agent forgets everything
  19969. 11:51:11that he's learned, then there's no point
  19970. 11:51:13because he'll have to redo everything
  19971. 11:51:14from scratch and again he'll forget
  19972. 11:51:16everything. That's why we have a matrix
  19973. 11:51:18known as a Q matrix, which represents
  19974. 11:51:20the memory of the agent. Okay, so if a
  19975. 11:51:22agent has traveled from room number two
  19976. 11:51:24to three, he's going to remember the
  19977. 11:51:26reward. Okay, that reward is going to be
  19978. 11:51:28stored in the Q matrix. And the rows of
  19979. 11:51:31the Q matrix will represent the current
  19980. 11:51:32state and the columns will represent the
  19981. 11:51:35possible actions which lead to the next
  19982. 11:51:37state. Now, the formula to calculate the
  19983. 11:51:39Q matrix is the following. You have Q
  19984. 11:51:42state {comma} action is equals to R
  19985. 11:51:44state {comma} action. Okay, let's say
  19986. 11:51:46the state is your S state number one and
  19987. 11:51:48you're moving to state number two. Okay,
  19988. 11:51:50so Q 1 {comma} 2 is going to be so on. R
  19989. 11:51:53state {comma} action will represent the
  19990. 11:51:55reward of 1 {comma} 2. Okay, let's try
  19991. 11:51:57to understand the reward. So, the reward
  19992. 11:52:00of 1 {comma} 2 is minus 1. Okay, so here
  19993. 11:52:03you'll get a value of minus 1. And then
  19994. 11:52:05you'll have plus the gamma parameter
  19995. 11:52:08into the maximum of the next state and
  19996. 11:52:10the next possible action that you can
  19997. 11:52:12take. Okay, now what is a gamma
  19998. 11:52:14parameter? So, the gamma parameter has a
  19999. 11:52:17range of 0 to 1. Okay, so it can be
  20000. 11:52:19between the value of 0 and 1. Now, if
  20001. 11:52:22gamma is closer to 0, then the agent
  20002. 11:52:24will tend to consider only immediate
  20003. 11:52:26rewards. But if the gamma parameter is
  20004. 11:52:29close to 1, the agent will consider
  20005. 11:52:31future rewards with greater weight. Now,
  20006. 11:52:33I don't know if this reminds you of
  20007. 11:52:35something, but earlier in the session, I
  20008. 11:52:37discussed two important concepts of
  20009. 11:52:39reinforcement learning, which was
  20010. 11:52:41exploration and exploitation trade-off.
  20011. 11:52:44Now, if you're exploring, then the gamma
  20012. 11:52:46parameter is going to be closer to 1,
  20013. 11:52:48but if you're exploiting, then the gamma
  20014. 11:52:50parameter is going to be closer to 0.
  20015. 11:52:53It's better if the gamma parameter is
  20016. 11:52:54closer to 1 because it means that you're
  20017. 11:52:56exploring the entire environment and
  20018. 11:52:59you're trying to get future rewards.
  20019. 11:53:00You're going to get greater weightage
  20020. 11:53:02and more rewards.
  20021. 11:53:03So guys, to sum up the entire thing,
  20022. 11:53:05let's look at how the Q-learning
  20023. 11:53:07algorithm will solve the problem. So you
  20024. 11:53:09begin by setting the gamma parameter and
  20025. 11:53:12the environment rewards in reward matrix
  20026. 11:53:14R. We already did that. After that, you
  20027. 11:53:17set the matrix Q to zero because
  20028. 11:53:19initially the agent will start with no
  20029. 11:53:21knowledge of the environment. As the
  20030. 11:53:24agent explores the environment, the Q
  20031. 11:53:26matrix will start filling up. After
  20032. 11:53:28that, the next step will be select a
  20033. 11:53:30random initial state. Like I mentioned
  20034. 11:53:33earlier, initially you'll randomly
  20035. 11:53:34select any state because the agent has
  20036. 11:53:37no idea about the environment. So you'll
  20037. 11:53:39randomly select a state at step number
  20038. 11:53:41three. Then step number four is you set
  20039. 11:53:44the initial state as current state. Step
  20040. 11:53:47number five, select one among all
  20041. 11:53:49possible actions for the current state.
  20042. 11:53:51Okay, so if you've chosen the current
  20043. 11:53:52state as one, let's say that all the
  20044. 11:53:55possible states that you can traverse to
  20045. 11:53:57from one is two, three, and so on. So
  20046. 11:54:00these are all the possible actions for
  20047. 11:54:02the current state. Now use the possible
  20048. 11:54:04actions and consider going to the next
  20049. 11:54:06state. Once you know what are the
  20050. 11:54:08possible actions from the current state,
  20051. 11:54:10you're going to go to one of those
  20052. 11:54:12possible actions and then you move to
  20053. 11:54:14the next state. After that, get maximum
  20054. 11:54:16Q value for this next state based on all
  20055. 11:54:19possible actions. If you get the maximum
  20056. 11:54:21Q value, it means that you've chosen an
  20057. 11:54:24optimal policy in order to reach your
  20058. 11:54:26end state. Finally, you'll compute the Q
  20059. 11:54:29value by using the formula that we
  20060. 11:54:30discussed earlier. And the last step is
  20061. 11:54:33you have to repeat all of these states
  20062. 11:54:35until your current state is equal to
  20063. 11:54:37your goal state. And our goal state is
  20064. 11:54:39room number five. Now guys, this is a
  20065. 11:54:41very logical solution because you can
  20066. 11:54:43easily understand what is happening over
  20067. 11:54:45here. You have an agent, he has to
  20068. 11:54:47explore through all the states in such a
  20069. 11:54:49way that he reaches the end goal using
  20070. 11:54:51the optimum policy. Okay, that's the end
  20071. 11:54:54goal of Q-learning algorithm and that's
  20072. 11:54:56exactly how you're going to solve this
  20073. 11:54:58problem.
  20074. 11:54:59And

About this transcript

This page contains the full transcript of AI & ML Full Course 2026 | Complete Artificial Intelligence and Machine Learning Tutorial | Edureka by edureka!, generated from the public captions YouTube serves with the video. The transcript has 130,062 words across 20,074 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.