AI & ML Full Course 2026 | Complete Artificial Intelligence and Machine Learning Tutorial | Edureka — Transcript
Full transcript
- 0:09Hello everyone and welcome to the AI
- 0:11[music] and machine learning full
- 0:12course. Artificial intelligence and
- 0:15machine learning are rapidly changing
- 0:17the way organization [music] analyze
- 0:18data, automate processes, and build
- 0:21intelligent applications. [music]
- 0:23This course provides a complete
- 0:25introduction to a key ideas, techniques,
- 0:27and tools
- 0:28>> [music]
- 0:28>> used in modern AI and machine learning.
- 0:31Throughout the course, you will explore
- 0:33how machines [music] learn from data,
- 0:35understand the role of algorithms and
- 0:37models, and see how AI systems are
- 0:40applied to solve real-world problems.
- 0:41[music]
- 0:42The course is designed to build your
- 0:44understanding step-by-step, making
- 0:46complex concepts [music] easier to
- 0:48understand. And by the end of this
- 0:50course, you will have a strong
- 0:51foundation in AI and machine learning
- 0:54and a clear view [music] of how these
- 0:56technologies are driving innovation
- 0:57across industries. So, before we begin,
- 1:00please like, share, and subscribe to
- 1:01Edureka's YouTube channel and hit the
- 1:03bell icon to stay updated on the latest
- 1:05tech content [music] from Edureka. Also,
- 1:08check out Edureka's postgraduate program
- 1:10in generative [music] AI and machine
- 1:12learning in collaboration with Illinois
- 1:14Tech. It offers a unique opportunity to
- 1:16explore [music] the cutting-edge world
- 1:18of generative AI and develop advanced
- 1:20AI-powered solutions. This program
- 1:23[music] covers in-demand topics
- 1:24including machine learning, deep
- 1:26learning, natural language processing,
- 1:28[music] prompt engineering, generative
- 1:30AI, LLMs, RAG, agentic AI, and much
- 1:33more. Learn from industry [music]
- 1:35experts through a curriculum built
- 1:37around real-world hands-on use cases
- 1:39designed to equip you with a practical
- 1:41and job-ready skills. So, check out the
- 1:43course link given in the description box
- 1:45[music] below. Now, let us get started
- 1:48by understanding what artificial
- 1:49intelligence is.
- 1:51AI or artificial intelligence is a
- 1:54branch of computer science where focused
- 1:56on creating systems that were performed
- 1:59task and that would be normally required
- 2:01by human intelligence. These tasks can
- 2:04range understanding of natural language.
- 2:07Secondly, recognizing patterns, then
- 2:10making decisions, and lastly, learning
- 2:12from the experiences.
- 2:14AI is a collaboration of ideas, methods,
- 2:18and knowledge where from the multiple
- 2:20academic disciplines work on a different
- 2:22problem-solving and share their
- 2:24knowledge to a better understanding and
- 2:26come up with a good solution.
- 2:28But, it can also be rule-based and
- 2:31operate under a set of rules and
- 2:33conditions only.
- 2:34So, I would like to tell you all guys a
- 2:37small real-time experience or an
- 2:39experiment done by Alan Turing to
- 2:42propagate and to establish the
- 2:44artificial intelligence.
- 2:46Alan Turing was the first person to
- 2:48conduct the sustainable research in the
- 2:50field that he called machine
- 2:52intelligence.
- 2:53The Turing test was conducted to explore
- 2:55whether machines could exhibit
- 2:57human-like intelligence.
- 2:59Proposed by Alan Turing in 1950, it
- 3:02involves a human evaluator communicating
- 3:05with both a human and a machine through
- 3:08a text interface.
- 3:09If a evaluator cannot distinguish
- 3:11between the two based on their
- 3:13responses, the machine is said to have
- 3:15passed the test. It serves as a
- 3:17benchmark of the assessing the progress
- 3:20of AI and discussions of the nature and
- 3:23intelligence and consciousness.
- 3:25So, this is to evolve and to establish
- 3:28that even machines can work as humans.
- 3:32And that is how it is made to bring up
- 3:35the machines' knowledge and the humans'
- 3:37knowledge into machines.
- 3:39There are two types of artificial
- 3:40intelligence.
- 3:42First, let's talk about weak AI. Before
- 3:44getting into what is weak AI, I would
- 3:46like to tell you with an example that is
- 3:49performed in a real world.
- 3:51I know you're all guys will be knowing
- 3:53about Alexa, Apple Siri, and also
- 3:55self-driving vehicles.
- 3:57So, all of these considered as in weak
- 4:00AI. It is also known as a narrow AI or
- 4:04an AI narrow intelligent that is trained
- 4:07by the AI and focused to perform a
- 4:09specific task. Weak AI drives most of
- 4:12the AI that surrounds by us today.
- 4:14Narrow might be a more apt descriptor
- 4:17for this type of AI as it is anything
- 4:19but weak.
- 4:20Next, let us learn about strong AI.
- 4:23Strong AI is made up of artificial
- 4:25general intelligence or artificial super
- 4:28intelligence.
- 4:30Simply, it could be told as where a
- 4:31machine would have a intelligence equal
- 4:34to humans.
- 4:35Where humans will track what AI needs to
- 4:37be done.
- 4:39It would be able to self-aware with the
- 4:41consciousness that would be this ability
- 4:43to solve problems, learn, and also plan
- 4:45for the future.
- 4:46Have you ever thought what does AI do at
- 4:49its core?
- 4:50I'm here to tell you what. AI is
- 4:52essential. It works by analyzing a lot
- 4:55of data
- 4:57to find patterns and a useful
- 4:58information.
- 4:59It even learns from this data to get
- 5:02better at the task over time. With this
- 5:04learning, AI can make decisions, predict
- 5:06future events, and do tasks
- 5:08automatically that would normally need
- 5:10human intelligence.
- 5:11This help business and other
- 5:13organizations work more efficiently.
- 5:16So, AI is about making computers smarter
- 5:19and more helpful in everyday life. So,
- 5:22let me tell you some use cases that is
- 5:24happening in daily uses or daily real
- 5:27life based.
- 5:28Firstly, I have taken is about
- 5:30cybersecurity. As it is a very important
- 5:33and a vital role for many platforms here
- 5:35after.
- 5:36Cybersecurity is a critical concern for
- 5:38individual businesses and also
- 5:40governments as cyber threats continue to
- 5:43evolve in a complexity and
- 5:44sophistication. It would play a very
- 5:47role for augmenting cybersecurity
- 5:49defenses.
- 5:50It is all based on the false positive
- 5:53effects that is made by the false
- 5:55information given by any criteria.
- 5:58Leading the alert of fatigue and reduced
- 6:01operational efficiencies.
- 6:03Cyber attacks can exploit
- 6:05vulnerabilities in AI models
- 6:08by invading detection and compromising
- 6:10security defenses.
- 6:12They have some private data where it is
- 6:15trained on the basis of incomplete data
- 6:18sets may produce the outcomes of private
- 6:21data. This will raise an ethical concern
- 6:24and also regulatory compliances issues.
- 6:27Having a very complicated problems, too.
- 6:29So, the next one will be your
- 6:31entertainment. As people know,
- 6:33entertainment is taking a very huge part
- 6:36in everyone's life.
- 6:37For example, it would be your social
- 6:39media, too.
- 6:40So, AI is transforming the entertainment
- 6:42industry by revolutionizing content
- 6:45creation, personalization, and audience
- 6:47engagement. So, this can be such as
- 6:50television, gaming, music, and also
- 6:52digital media platforms.
- 6:54One significant use of AI in
- 6:56entertainment is personalized content
- 6:59recommendation.
- 7:01So, personalized entertainment
- 7:02experiences are enhanced through AI
- 7:04driven. It is recommended through
- 7:06systems.
- 7:07They are even having some platforms like
- 7:10Netflix,
- 7:11Prime Amazon, and also Spotify
- 7:13that is making people engaged in a very
- 7:15hype as of now.
- 7:17These recommendations improve over time.
- 7:20So, I would like to conclude by telling
- 7:22artificial intelligence can do amazing
- 7:24things like analyzing data and making
- 7:27task easier.
- 7:28But, it also brings up huge question
- 7:31about fairness, jobs, and who controls
- 7:33it.
- 7:34We need to be careful in how we develop
- 7:36and use the AI.
- 7:37Making sure it helps everyone and
- 7:39doesn't cause harm at all.
- 7:41So, this makes easier for people to get
- 7:43into creativity and make rules and
- 7:45guidelines.
- 7:49>> [music]
- 7:52>> So, now let's get started with the first
- 7:53topic, which is history of artificial
- 7:56intelligence.
- 7:57The concept of AI goes back to the
- 7:59classical ages. Under Greek mythology,
- 8:02the concept of machines and mechanical
- 8:04men were well thought of. An example is
- 8:07Talos. Talos was supposedly a giant
- 8:10animated bronze warrior who was
- 8:12programmed to guard the island of Crete.
- 8:15Now, let's get back to the 19th century.
- 8:17In 1950, Alan Turing proposed the Turing
- 8:20test. The Turing test basically
- 8:22determines whether or not a computer can
- 8:25intelligently think like a human being.
- 8:27The Turing test was the first serious
- 8:29proposal in the philosophy of artificial
- 8:31intelligence.
- 8:331951 marked the era for game artificial
- 8:36intelligence. This period was called
- 8:38game AI because here a lot of computer
- 8:40scientists developed programs for
- 8:42checkers and for chess. However, these
- 8:45programs were later rewritten and redone
- 8:47in a better way.
- 8:491956 marked the most important year for
- 8:52artificial intelligence.
- 8:54During this year, John McCarthy first
- 8:56coined the term artificial intelligence.
- 8:58This was followed by the first AI
- 9:00laboratory, which was set up in 1959.
- 9:03MIT AI Lab was the first setup, which
- 9:06was basically dedicated to the research
- 9:08of AI.
- 9:09In 1960, the first robot was introduced
- 9:12to the General Motors assembly line. In
- 9:151961, the first AI chatbot called Eliza
- 9:19was introduced. In 1997, IBM's Deep Blue
- 9:23beats the world champion Garry Kasparov
- 9:25in the game of chess.
- 9:272005 marks for the year when an
- 9:29autonomous robotic car called Stanley
- 9:32won the DARPA Grand Challenge.
- 9:34In 2011, IBM's question-answering
- 9:37machine Watson defeated the two greatest
- 9:40Jeopardy champions Brad Rutter and Ken
- 9:42Jennings. So, that was a brief history
- 9:45of AI. Now guys, since the emergence of
- 9:47artificial intelligence in 1950s, we
- 9:50have seen an exponential growth in its
- 9:52potential. AI covers domains such as
- 9:55machine learning, deep learning, neural
- 9:57networks, natural language processing,
- 9:59knowledge-based expert systems, and so
- 10:02on.
- 10:02Now that you know a brief history of
- 10:04artificial intelligence, let's move on
- 10:06and understand what exactly artificial
- 10:08intelligence is. So, the term artificial
- 10:11intelligence was first coined by John
- 10:13McCarthy, like I mentioned earlier. He
- 10:15defined AI as a science and engineering
- 10:18of making intelligent machines. In other
- 10:20words, artificial intelligence can also
- 10:23be defined as a development of computer
- 10:25systems that are capable of performing
- 10:28tasks that require human intelligence
- 10:30such as decision-making, object
- 10:32detection, solving complex problems, and
- 10:34so on. So, like I mentioned, artificial
- 10:37intelligence helps in decision-making,
- 10:39solving complex problems, it performs
- 10:42high-level computations, and also
- 10:44increases the accuracy of your
- 10:46predictions. Right? These are the main
- 10:48features of AI.
- 10:49So, now let's understand the different
- 10:51stages of artificial intelligence. So,
- 10:54basically, when I was doing my research,
- 10:55I found a lot of videos and a lot of
- 10:58articles that stated that artificial
- 11:00general intelligence, artificial narrow
- 11:03intelligence, and artificial super
- 11:04intelligence are the different types of
- 11:07AI. If I have to be more precise with
- 11:09you, then artificial intelligence has
- 11:11three different stages. Right? The types
- 11:13of AI are completely different from the
- 11:15stages of AI. So, under the stages of
- 11:17artificial intelligence, we have
- 11:19artificial narrow intelligence,
- 11:21artificial general intelligence, and
- 11:23artificial super intelligence.
- 11:25So, what is artificial narrow
- 11:26intelligence? Artificial narrow
- 11:28intelligence, also known as weak AI, is
- 11:31a stage of artificial intelligence that
- 11:33involves machines that can perform only
- 11:36a narrowly defined set of specific
- 11:38tasks. Right? At this stage, the
- 11:40machines don't possess any thinking
- 11:42ability. They just perform a set of
- 11:44predefined functions. Examples of weak
- 11:47AI include Siri, Alexa, AlphaGo, Sophia,
- 11:50the self-driving cars, and so on. Almost
- 11:53all the AI-based systems that are built
- 11:55till this date fall under the category
- 11:57of weak AI or artificial narrow
- 11:59intelligence.
- 12:01Next, we have something known as
- 12:02artificial general intelligence.
- 12:04Artificial general intelligence is also
- 12:06known as strong AI. This stage is the
- 12:09evolution of artificial intelligence,
- 12:11wherein machines will possess the
- 12:13ability to think and make decisions just
- 12:16like human beings. There are currently
- 12:18no existing examples of strong AI, but
- 12:21it's believed that we will soon be able
- 12:23to create machines that are as smart as
- 12:25human beings. Strong AI is actually
- 12:28considered a threat to human existence
- 12:30by many scientists. This includes
- 12:32Stephen Hawking. Stephen Hawking quoted
- 12:35that the development of full artificial
- 12:37intelligence could spell the end of
- 12:40human race. Moving on to our last stage,
- 12:42which is artificial super intelligence.
- 12:45Artificial super intelligence is that
- 12:47stage of AI when the capability of
- 12:49computers will surpass human beings.
- 12:52Artificial super intelligence is
- 12:54currently seen as a hypothetical
- 12:56situation as depicted in movies and
- 12:58science fiction books. You see a lot of
- 13:00movies which show that machines are
- 13:02taking over the world. All of that is
- 13:04artificial super intelligence. Now, I
- 13:06believe that machines are not very far
- 13:08from reaching the stage taking into
- 13:10consideration our current pace.
- 13:12However, such systems don't currently
- 13:14exist, right? We don't have any machine
- 13:16that is capable of thinking better than
- 13:19a human being or reasoning in a better
- 13:21way than a human. Artificial super
- 13:23intelligence, basically any robot that
- 13:24is much smarter than humans. Now, moving
- 13:27on to the different types of artificial
- 13:29intelligence. Based on the functionality
- 13:32of AI-based systems, artificial
- 13:34intelligence can be categorized into
- 13:36four types. The first type is reactive
- 13:38machines AI. This type of AI includes
- 13:41machines that operate solely based on
- 13:44the present data and take into
- 13:46consideration only the current
- 13:48situation.
- 13:49Reactive AI machines cannot form
- 13:51inferences from the data to evaluate any
- 13:54future actions. They can perform a
- 13:56narrowed range of predefined tasks.
- 13:59An example of reactive AI is the famous
- 14:02IBM chess program that beat the world
- 14:04champion Garry Kasparov.
- 14:06This is one of the most impressive AI
- 14:08machines built so far.
- 14:10Next, we have limited memory AI. Now,
- 14:13like the name suggests, limited memory
- 14:15AI can make informed and improved
- 14:17decisions by studying the past data from
- 14:20its memory. So, such an AI has a
- 14:22short-lived or you can say a temporary
- 14:25memory that can be used to store past
- 14:27experiences and hence evaluate your
- 14:29future actions.
- 14:31Self-driving cars are limited memory AI
- 14:33that use the data collected in the
- 14:35recent past to make immediate decisions.
- 14:38For example, self-driving cars use
- 14:40sensors to identify civilians that are
- 14:43crossing the road. They identify any
- 14:45steep roads or traffic signals and they
- 14:48use this to make better driving
- 14:49decisions. This also helps in preventing
- 14:52any future accidents. Next, we have
- 14:54something known as theory of mind
- 14:56artificial intelligence. The theory of
- 14:58mind AI is a more advanced type of
- 15:00artificial intelligence. This category
- 15:03is speculated to play a very important
- 15:06role in psychology. This type of AI will
- 15:08mainly focus on emotional intelligence
- 15:11so that human beliefs and thoughts can
- 15:13be better comprehended. The theory of
- 15:15mind AI has not been fully developed
- 15:18yet, but rigorous research is happening
- 15:20in this area.
- 15:21Moving on to our last type of artificial
- 15:23intelligence is the self-aware
- 15:25artificial intelligence.
- 15:27So guys, let us fold hands and pray that
- 15:29we don't reach the state of AI where
- 15:32machines have their own consciousness
- 15:34and become self-aware. This type of AI
- 15:37is a little far-fetched, but in the
- 15:38future achieving a stage of
- 15:40superintelligence might be possible.
- 15:43Geniuses like Elon Musk and Stephen
- 15:45Hawking have constantly warned us about
- 15:47evolution of AI.
- 15:49So guys, let me know your thoughts in
- 15:50the comment section. Do you ever think
- 15:52we'll reach the stage of artificial
- 15:54superintelligence?
- 15:56Moving on to the last topic of today's
- 15:58session is the different domains or the
- 16:00different branches of artificial
- 16:02intelligence. So artificial intelligence
- 16:04can be used to solve real-world problems
- 16:06by implementing machine learning, deep
- 16:08learning, natural language processing,
- 16:10robotics, expert systems, and fuzzy
- 16:13logic. Now guys, these are the different
- 16:15domains or you can say the different
- 16:16branches that AI uses in order to solve
- 16:19any problem. Recently, AI has also been
- 16:22used as an application in computer
- 16:24vision and image processing. Right, for
- 16:26now let me tell you briefly about each
- 16:28of these domains. Machine learning is
- 16:30basically the science of getting
- 16:32machines to interpret, process, and
- 16:34analyze data in order to solve
- 16:36real-world problems. Right, under
- 16:38machine learning there's supervised,
- 16:40unsupervised, and reinforcement
- 16:41learning. If any of you are interested
- 16:43in learning about these technologies,
- 16:45I'll leave a link in the description
- 16:46box. You all can go through that
- 16:47content. Next, we have deep learning or
- 16:50neural networks. So deep learning is a
- 16:52process of implementing neural networks
- 16:54on high-dimensional data to gain
- 16:57insights and form solutions. It is
- 16:59basically the logic behind the face
- 17:01verification algorithm on Facebook. It
- 17:04is the logic behind the self-driving
- 17:06cars, virtual assistants like Siri and
- 17:08Alexa. Then we have natural language
- 17:10processing. Natural language processing
- 17:12refers to the science of drawing
- 17:14insights from natural human language in
- 17:16order to communicate with machines and
- 17:19grow businesses. So, an example of NLP
- 17:22is Twitter and Amazon. Twitter uses NLP
- 17:25to filter out terroristic language in
- 17:27their tweets. Amazon uses NLP to
- 17:30understand customer reviews and improve
- 17:32user experience. Then we have robotics.
- 17:35Robotics is a branch of artificial
- 17:37intelligence which focuses on the
- 17:39different branches and applications of
- 17:41robots. AI robots are artificial agents
- 17:44which act in the real world environment
- 17:47to produce results by taking some
- 17:49accountable actions.
- 17:51So, I'm sure all of you have heard of
- 17:52Sophia. Sophia the humanoid is a very
- 17:55good example of AI in robotics. Then we
- 17:58have fuzzy logic. So, fuzzy logic is a
- 18:00computing approach that is based on the
- 18:02principle of degree of truth instead of
- 18:05the usual modern logic that we use which
- 18:08is basically the Boolean logic. Fuzzy
- 18:10logic is used in medical fields to solve
- 18:12complex problems which involve decision
- 18:15making. It is also used in automating
- 18:18gear systems in your cars and all of
- 18:20that. Then we have expert systems. An
- 18:22expert system is an AI-based computer
- 18:24system that learns and reciprocates the
- 18:27decision-making ability of a human
- 18:29expert. Expert systems use if-then logic
- 18:32notions in order to solve any complex
- 18:34problem. They do not rely on
- 18:37conventional procedural programming.
- 18:39Expert systems are mainly used in
- 18:41information management. They're seen to
- 18:43be used in fraud detection, virus
- 18:46detection, also in managing medical and
- 18:48hospital records, and so on. So, guys,
- 18:50to sum it up, these were the different
- 18:52branches of artificial intelligence.
- 18:55>> [music]
- 19:00>> These are the term which have confused a
- 19:02lot of people. And if you too are one
- 19:04among them, let me resolve it for you.
- 19:07Well, artificial intelligence is a
- 19:09broader umbrella under which machine
- 19:11learning and deep learning come. You can
- 19:13also see in the diagram that even deep
- 19:15learning is a subset of machine
- 19:17learning. So, you can say that all three
- 19:19of them, the AI, the machine learning,
- 19:21and deep learning, are just the subset
- 19:24of each other. So, let's move on and
- 19:26understand how exactly they differ from
- 19:28each other. So, let's start with
- 19:30artificial intelligence. The term
- 19:32artificial intelligence was first coined
- 19:35in the year 1956.
- 19:37The concept is pretty old, but it has
- 19:39gained its popularity recently. But why?
- 19:42Well, the reason is earlier we had very
- 19:45small amount of data. The data we had
- 19:48was not enough to predict the accurate
- 19:50result. But now, there's a tremendous
- 19:52increase in the amount of data.
- 19:54Statistics suggest that by 2020, the
- 19:57accumulated volume of data will increase
- 20:00from 4.4 zettabytes to roughly around 44
- 20:03zettabytes, or 44 trillion GBs of data.
- 20:06Along with such enormous amount of data,
- 20:09now we have more advanced algorithm and
- 20:12high-end computing power and storage
- 20:14that can deal with such large amount of
- 20:15data.
- 20:16As a result, it is expected that 70% of
- 20:19enterprise will implement AI over the
- 20:21next 12 months, which is up from 40% in
- 20:242016 and 51% in 2017.
- 20:28Just for your understanding, what is AI?
- 20:31Well, it's nothing but a technique that
- 20:33enables the machine to act like humans
- 20:35by replicating the behavior and nature.
- 20:38With AI, it is possible for machine to
- 20:40learn from the experience. The machines
- 20:43adjust their responses based on new
- 20:45input, thereby performing human-like
- 20:47tasks.
- 20:48Artificial intelligence can be trained
- 20:50to accomplish specific tasks by
- 20:51processing large amount of data and
- 20:53recognizing pattern in them.
- 20:55You can consider that building an
- 20:57artificial intelligence is like building
- 20:59a church. The first church took
- 21:01generations to finish. So, most of the
- 21:04workers who were working on it never saw
- 21:06the final outcome. Those working on it
- 21:08took pride in their crafts, building
- 21:10bricks and chiseling stone that was
- 21:12going to be placed into the great
- 21:13structure. So, as AI researchers, we
- 21:16should think of ourselves as humble
- 21:18brick makers whose job is to study how
- 21:21to build components, example parsers,
- 21:23planners, or learning algorithm, or
- 21:25etc., anything that someday someone and
- 21:27somewhere will integrate into the
- 21:29intelligent systems. Some of the
- 21:31examples of artificial intelligence from
- 21:33our day-to-day life are Apple series,
- 21:36chess-playing computer, Tesla's
- 21:38self-driving car, and many more. These
- 21:40examples are based on deep learning and
- 21:42natural language processing.
- 21:44Well, this was about what is AI and how
- 21:46it gained its height. So, moving on
- 21:48ahead, let's discuss about machine
- 21:50learning and see what it is and why it
- 21:53was even introduced. Well, machine
- 21:55learning came into existence in the late
- 21:57'80s and the early '90s. But, what were
- 21:59the issues with the people which made
- 22:01the machine learning come into
- 22:02existence? Let us discuss them one by
- 22:05one.
- 22:05In the field of statistics, the problem
- 22:08was how to efficiently train large
- 22:10complex model. In the field of computer
- 22:12science and artificial intelligence, the
- 22:14problem was how to train more robust
- 22:16version of AI system. While in the case
- 22:18of neuroscience, problem faced by the
- 22:20researchers was how to design
- 22:22operational model of the brain.
- 22:24So, these were some of the issues which
- 22:26had the largest influence and led to the
- 22:28existence of the machine learning.
- 22:30Now, this machine learning shifted its
- 22:32focus from the symbolic approaches it
- 22:33had inherited from the AI and moved
- 22:36towards the methods and model it had
- 22:38bought from statistics and probability
- 22:40theory.
- 22:41So, let's proceed and see what exactly
- 22:43is machine learning. Well, machine
- 22:45learning is a subset of AI which enables
- 22:48the computer to act and make data-driven
- 22:50decisions to carry out a certain task.
- 22:52These programs or algorithms are
- 22:54designed in a way that they can learn
- 22:56and improve over time when exposed to
- 22:58new data. Let's see an example of
- 23:00machine learning. Let's say you want to
- 23:02create a system which tells the expected
- 23:04weight of a person based on its height.
- 23:07The first thing you do is you collect
- 23:08the data. Let's see, this how your data
- 23:10looks like. Now, each point on the graph
- 23:13represent one data point. To start with,
- 23:16we can draw a simple line to predict the
- 23:18weight based on the height. For example,
- 23:20a simple line W equal H minus 100, where
- 23:23W is weight in kg and H is height in cm.
- 23:27This line can help us to make the
- 23:28prediction. Our main goal is to reduce
- 23:31the difference between the estimated
- 23:32value and the actual value. So, in order
- 23:35to achieve it, we try to draw a straight
- 23:37line that fits through all these
- 23:39different points and minimize the error.
- 23:41So, our main goal is to minimize the
- 23:43error and make them as small as
- 23:45possible. Decreasing the error or the
- 23:47difference between the actual value and
- 23:49estimated value increases the
- 23:50performance of the model. Further on,
- 23:53the more data points we collect, the
- 23:55better our model will become. We can
- 23:56also improve our model by adding more
- 23:58variables and creating different
- 24:00prediction lines for them. Once the line
- 24:02is created, so from the next time if we
- 24:04feed a new data, for example, height of
- 24:06a person to the model, it would easily
- 24:08predict the data for you and it will
- 24:10tell you what its predicted weight could
- 24:12be. I hope you got a clear understanding
- 24:14of machine learning. So, moving on
- 24:16ahead, let's learn about deep learning.
- 24:18Now, what is deep learning? You can
- 24:20consider deep learning model as a rocket
- 24:22engine and its fuel is its huge amount
- 24:25of data that we feed to these
- 24:26algorithms.
- 24:27The concept of deep learning is not new.
- 24:30But recently, it's hype has increased
- 24:32and deep learning is getting more
- 24:33attention.
- 24:34This field is a particular kind of
- 24:36machine learning that is inspired by the
- 24:38functionality of our brain cells called
- 24:39neuron, which led to the concept of
- 24:42artificial neural network.
- 24:44It simply takes the data connection
- 24:45between all the artificial neurons and
- 24:47adjust them according to the data
- 24:49pattern. More neurons are added if the
- 24:51size of the data is large. It
- 24:53automatically features learning at
- 24:55multiple levels of abstraction, thereby
- 24:57allowing a system to learn complex
- 24:59function mapping without depending on
- 25:01any specific algorithm. You know what?
- 25:04No one actually knows what happens
- 25:06inside a neural network and why it works
- 25:08so well. So, currently you can call it
- 25:10as a black box. Let [snorts] us discuss
- 25:12some of the example of deep learning and
- 25:14understand it in a better way. Let me
- 25:16start with a simple example and explain
- 25:18you how things happen at a conceptual
- 25:21level. Let us try and understand how you
- 25:23recognize a square from other shapes.
- 25:26The first thing you do is you check
- 25:28whether there are four lines associated
- 25:30with the figure or not. Simple concept,
- 25:32right? If yes, we further check if they
- 25:35are connected and closed. Again, if yes,
- 25:37we finally check whether it is
- 25:39perpendicular and all its sides are
- 25:41equal. Correct? If everything fulfills,
- 25:44yes, it is a square.
- 25:46Well, it is nothing but a nested
- 25:47hierarchy of concepts.
- 25:50What we did here, we took a complex task
- 25:52of identifying a square in this case and
- 25:54broke it into simpler task. Now, this
- 25:56deep learning also does the same thing
- 25:58but at a larger scale. Let's take an
- 26:01example of machine which recognizes the
- 26:03animal. The task of the machine is to
- 26:05recognize whether the given image is of
- 26:07a cat or of a dog.
- 26:09What if we were asked to resolve the
- 26:10same issue using the concept of machine
- 26:12learning? What we would do? First, we
- 26:15would define the features such as check
- 26:17whether the animal has whiskers or not
- 26:19or check if the animal has pointed ears
- 26:21or not or whether its tail is straight
- 26:23or curved. In short, we will define the
- 26:25facial features and let the system
- 26:27identify which features are more
- 26:29important in classifying a particular
- 26:31animal. Now, when it comes to deep
- 26:34learning, it takes this to one step
- 26:35ahead. Deep learning automatically finds
- 26:38out the feature which are most important
- 26:40for classification compared to machine
- 26:42learning where we had to manually give
- 26:44out that features.
- 26:46By now, I guess you have understood that
- 26:48AI is a bigger picture and machine
- 26:50learning and deep learning are its
- 26:51subpart. So, let's move on and focus our
- 26:53discussion on machine learning and deep
- 26:55learning.
- 26:56The easiest way to understand the
- 26:58difference between the machine learning
- 26:59and deep learning is to know that deep
- 27:01learning is machine learning. More
- 27:03specifically, it is the next evolution
- 27:05of machine learning. Let's take few
- 27:07important parameter and compare machine
- 27:09learning with deep learning. So,
- 27:11starting with data dependencies. The
- 27:13most important difference between deep
- 27:15learning and machine learning is its
- 27:17performance as the volume of the data
- 27:19gets increased. From the below graph,
- 27:21you can see that when the size of the
- 27:23data is small, deep learning algorithm
- 27:25doesn't perform that well. But, why?
- 27:28Well, this is because deep learning
- 27:30algorithm needs a large amount of data
- 27:32to understand it perfectly.
- 27:34On the other hand, the machine learning
- 27:36algorithm can easily work with smaller
- 27:38data set. Fine?
- 27:40Next comes the hardware dependencies.
- 27:42Deep learning algorithms are heavily
- 27:44dependent on high-end machines, while
- 27:46the machine learning algorithm can work
- 27:48on low-end machines as well.
- 27:50This is because the requirement of deep
- 27:52learning algorithm include GPUs, which
- 27:55is an integral part of its working.
- 27:57The deep learning algorithm require GPUs
- 27:59as they do a large amount of matrix
- 28:01multiplication operations, and these
- 28:03operations can only be efficiently
- 28:06optimized using a GPU as it is built for
- 28:09this purpose only.
- 28:10Our third parameter will be feature
- 28:12engineering. Well, feature engineering
- 28:15is a process of putting the domain
- 28:17knowledge to reduce the complexity of
- 28:19the data and make patterns more visible
- 28:21to learning algorithms.
- 28:23This process is difficult and expensive
- 28:25in terms of time and expertise.
- 28:28In case of machine learning, most of the
- 28:29features are needed to be identified by
- 28:31an expert and then hand coded as per the
- 28:34domain and the data type. For example,
- 28:37the features can be a pixel value,
- 28:38shapes, texture, position, orientation,
- 28:41or anything. Fine? The performance of
- 28:44most of the machine learning algorithm
- 28:46depends on how accurately the features
- 28:48are identified and extracted.
- 28:50Whereas in case of deep learning
- 28:52algorithms, it try to learn high-level
- 28:54features from the data. This is a very
- 28:56distinctive part of deep learning, which
- 28:57makes it way ahead of traditional
- 28:59machine learning.
- 29:01Deep learning reduces the task of
- 29:03developing new feature extractor for
- 29:04every problem. Like in the case of CNN
- 29:07algorithm, it first try to learn the
- 29:09low-level features of the image, such as
- 29:11edges and lines, and then it proceeds to
- 29:13the parts of faces of people, and then
- 29:16finally to the high-level representation
- 29:17of the face. I hope the things are
- 29:19getting clear to you.
- 29:21So, let's move on ahead and see the next
- 29:23parameter. So, our next parameter is
- 29:25problem-solving approach.
- 29:27When we are solving a problem using
- 29:29traditional machine learning algorithm,
- 29:30it is generally recommended that we
- 29:33first break down the problem into
- 29:34different sub parts, solve them
- 29:36individually, and then finally combine
- 29:38them to get the desired result. This is
- 29:41how the machine learning algorithm
- 29:42handles the problem. On the other hand,
- 29:45the deep learning algorithm solves the
- 29:46problem from end to end.
- 29:48Let's take an example to understand
- 29:50this.
- 29:51Suppose you have a task of multiple
- 29:52object detection, and your task is to
- 29:54identify what is the object and where it
- 29:57is present in the image. So, let's see
- 29:59and compare how will you tackle this
- 30:01issue using the concept of machine
- 30:03learning and deep learning.
- 30:04Starting with machine learning, in a
- 30:06typical machine learning approach, you
- 30:08would first divide the problem into two
- 30:10step. First, object detection and then
- 30:13object recognition. First of all, you'd
- 30:16use a bounding box detection algorithm
- 30:18like GrabCut for example,
- 30:20to scan through the image and find out
- 30:22all the possible objects. Now, once the
- 30:25objects are recognized, you'd use object
- 30:27recognition algorithm like SVM with HOG,
- 30:31to recognize relevant objects.
- 30:33Now, finally when you combine the
- 30:35result, you would be able to identify
- 30:37what is the object and where it is
- 30:38present in the image.
- 30:40On the other hand, in deep learning
- 30:42approach, you would do the process from
- 30:44end to end. For example, in a YOLO net,
- 30:46which is a type of deep learning
- 30:48algorithm, you would pass an image and
- 30:50it would give out the location along
- 30:52with the name of the object. Now, let's
- 30:54move on to our fifth comparison
- 30:56parameter.
- 30:57It's execution time.
- 30:59Usually, a deep learning algorithm takes
- 31:01a long time to train. This is because
- 31:03there are so many parameter in a deep
- 31:05learning algorithm that makes the
- 31:06training longer than usual. The training
- 31:09might even last for 2 weeks or more than
- 31:11that if you're training completely from
- 31:13the scratch. Whereas in the case of
- 31:15machine learning, it relatively takes
- 31:17much less time to train, ranging from a
- 31:19few weeks to few hours.
- 31:21Now, the execution time is completely
- 31:23reversed when it comes to the testing of
- 31:25data. During testing, the deep learning
- 31:28algorithm takes much less time to run.
- 31:30Whereas if you compare it with a KNN
- 31:32algorithm, which is a type of machine
- 31:33learning algorithm, the test time
- 31:35increases as the size of the data
- 31:36increase.
- 31:38Last but not the least, we have
- 31:39interpretability as a factor for
- 31:41comparison of machine learning and deep
- 31:43learning. This factor is the main reason
- 31:46why deep learning is still thought 10
- 31:48times before anyone uses it in the
- 31:50industry. Let's take an example. Suppose
- 31:53we use deep learning to give automated
- 31:56scoring to essays. The performance it
- 31:58gives in scoring is quite excellent and
- 32:00is near to the human performance. But
- 32:02there's an issue with it. It does not
- 32:04reveal why it has given that score.
- 32:06Indeed, mathematically it is possible to
- 32:09find out that which node of a deep
- 32:11neural network were activated, but we
- 32:13don't know what the neurons are supposed
- 32:15to model and what these layers of neuron
- 32:17were doing collectively.
- 32:19So, we failed to interpret the result.
- 32:21On the other hand, machine learning
- 32:22algorithm like decision tree gives us a
- 32:25crisp rule for why it chose and what it
- 32:27chose. So, it is particularly easy to
- 32:30interpret the reasoning behind it.
- 32:32Therefore, the algorithms like decision
- 32:33tree and linear or logistic regression
- 32:36are primarily used in industry for
- 32:38interpretability.
- 32:45So, why exactly are we using Python for
- 32:47artificial intelligence? Why aren't we
- 32:49using any other language? Right? Now,
- 32:52there are a couple of reasons as to why
- 32:54Python is so popular when it comes to
- 32:56AI, machine learning, and deep learning.
- 32:58The first reason is less coding is
- 33:00required. Now, artificial intelligence
- 33:02has a lot of algorithms. If you have to
- 33:04implement AI in any code or in any
- 33:07problem, then there are going to be tons
- 33:09and tons of machine learning algorithms
- 33:11involved, deep learning algorithms
- 33:13involved, right? Now, testing all of
- 33:15these can become a very tiresome task.
- 33:18That's where Python usually comes in
- 33:20handy. Now, the language has something
- 33:23known as check as you code methodology,
- 33:26which eases the process of testing,
- 33:28right? You can check your program as you
- 33:30code it. Basically, as you're typing
- 33:32each sentence, your errors or your any
- 33:34sort of mistakes in your code will be
- 33:36given to you. Right? So, testing becomes
- 33:38much easier when it comes to Python.
- 33:40The next important reason why we're
- 33:42choosing Python is it has support for
- 33:44pre-built libraries. Right? Python is
- 33:47very convenient for AI developers
- 33:50because all of the algorithms, machine
- 33:52learning algorithms, and deep learning
- 33:53algorithms are already predefined in
- 33:56libraries, right? So, you don't have to
- 33:57actually sit down and code each and
- 33:59every algorithm. That would take a lot
- 34:01of time. And that's a very
- 34:03time-consuming task. And thanks to
- 34:05Python, you don't have to do that
- 34:06because they have libraries and packages
- 34:09that have all the algorithms built in
- 34:11them, right? So, if you want to run any
- 34:14algorithm, all you have to do is you
- 34:15have to call the function and load the
- 34:17library. That's all. It's as simple as
- 34:19that. Now, the next reason is ease of
- 34:21learning. So, guys, Python is actually
- 34:24the most simplest programming language,
- 34:26right? If you ask me, I think it is is
- 34:28the most easiest programming language.
- 34:30It's very similar to English language,
- 34:32right? If you read a couple of lines in
- 34:34Python, you'll understand what exactly
- 34:36the code is doing. It has a very simple
- 34:38syntax, and this simple syntax can be
- 34:41implemented to solve simple problems
- 34:43like addition of two strings, and it can
- 34:45also be used to solve complex problems
- 34:48like building machine learning models
- 34:50and deep learning models. So, ease of
- 34:52learning is a major factor when it comes
- 34:54to why Python is chosen for artificial
- 34:57intelligence, right? Next, we have
- 34:59platform independent.
- 35:01So, good thing about Python is that you
- 35:02can get your project running on
- 35:04different operating systems, right? And
- 35:07what happens when you transfer your code
- 35:09from one operating system to another
- 35:11operating system is we find a lot of
- 35:13dependency issues. To solve that, Python
- 35:16has a couple of packages such as there
- 35:18is a package known as PyInstaller,
- 35:20right? This PyInstaller will take care
- 35:22of all the dependency issues when you're
- 35:24transferring your code from one platform
- 35:26to the other platform. So, all of this
- 35:29support is provided by Python. The last
- 35:31reason is massive community support.
- 35:33This is a very important point because
- 35:36it is important that you have a large
- 35:38community that will help you out with
- 35:40any errors or with any sort of problems
- 35:43in your code, right? So, Python has
- 35:46several communities and several forums
- 35:48and groups on Facebook. So, if you have
- 35:50any doubts regarding any error, you can
- 35:53just post those errors in these groups,
- 35:55and you'll have like a bunch of people
- 35:56helping you out. Right? So, guys, these
- 35:59are a couple of reasons as to why Python
- 36:01is chosen for artificial intelligence.
- 36:03It's actually considered the most
- 36:04popular and the most used language for
- 36:07data science, AI, machine learning, and
- 36:09deep learning. To prove that to you,
- 36:11here is a stat from Stack Overflow.
- 36:13Stack Overflow recently stated that
- 36:16Python is the fastest growing
- 36:17programming language. If you look at the
- 36:19graph, you can see that it has taken
- 36:21over JavaScript and Java and C#, C++,
- 36:25and PHP, right? So, Python is actually
- 36:28growing at an exponential rate,
- 36:30especially when it comes to data science
- 36:32and artificial intelligence. A lot of
- 36:34developers are very comfortable with the
- 36:35Python language because, you know, it's
- 36:37a general-purpose language, first of
- 36:38all. So, most of the developers are
- 36:40already aware of Python. And then, using
- 36:43the same language in order to solve
- 36:45complex problems like artificial
- 36:47intelligence, machine learning, and deep
- 36:49learning is something every developer
- 36:51wants, right? They want a simple
- 36:52language in order to code all the
- 36:54complex algorithms or the complex
- 36:56models. Right? So, that's why Python is
- 36:59the best choice for artificial
- 37:00intelligence.
- 37:02For those of you who are not aware of
- 37:03Python programming and don't know much
- 37:05about Python, I'm going to leave a
- 37:07couple of links in the description box.
- 37:09Right? You can go through those links
- 37:11and study a little bit more about how
- 37:12Python works or how the coding part
- 37:15works. Right? I'm going to be focusing
- 37:17mainly on artificial intelligence, and
- 37:19I'll be showing you a lot of demos. So,
- 37:21those of you are not aware of Python,
- 37:23make sure you check the description box.
- 37:25Right?
- 37:26Next, I'm going to discuss the different
- 37:28Python packages for artificial
- 37:30intelligence. Now, these are the
- 37:31packages that are specifically for
- 37:33machine learning, deep learning, natural
- 37:35language processing, and so on. So,
- 37:37let's take a look at all these packages.
- 37:39So, first, we have TensorFlow. If you
- 37:42are currently working on a machine
- 37:44learning project in Python, then you
- 37:46must have heard of this popular
- 37:48open-source library known as TensorFlow.
- 37:51Right? This library was developed by
- 37:52Google in collaboration with Brain team.
- 37:55TensorFlow is used in almost every
- 37:57Google application for machine learning.
- 37:59Now, let me just discuss a few features
- 38:01of TensorFlow. It has a responsive
- 38:03construct, meaning that with TensorFlow,
- 38:06we can easily visualize each and every
- 38:08part of the graph, which is not an
- 38:10option when you're using other packages
- 38:12such as NumPy or scikit. Right? Another
- 38:15feature is that it's very flexible. Now,
- 38:17one of the most important TensorFlow
- 38:19features is that it is flexible in
- 38:22operability. Meaning that it has
- 38:24modularity and the parts of which you
- 38:27want to make standalone, it offers you
- 38:29that option. Right? It's very flexible
- 38:31in that way. It'll give you exactly what
- 38:33you want. Now, good feature about
- 38:35TensorFlow is that you can train it on
- 38:37both CPU and GPU. Right? So, for
- 38:39distributed computing, you can have both
- 38:42these options. Also, it supports
- 38:44parallel neural network training. So,
- 38:46TensorFlow offers pipelining in the
- 38:49sense that you can train multiple neural
- 38:51networks and multiple GPUs, which makes
- 38:54the models very efficient on any
- 38:56large-scale system. Right? So, parallel
- 38:58neural network training is supported by
- 39:00TensorFlow. Right? This is one of the
- 39:03most important features of TensorFlow.
- 39:05Apart from this, it has a very large
- 39:07community. And needless to say, if it
- 39:09has been developed by Google, then
- 39:11there's already a large team of software
- 39:13engineers who work on stability,
- 39:16improvements, and all of that. Right?
- 39:19The next library I'm going to talk about
- 39:20is scikit-learn.
- 39:22Now, scikit-learn is a Python library
- 39:24that is associated with NumPy and SciPy.
- 39:26Right? That's why it has the name
- 39:28scikit-learn. Now, this is considered to
- 39:30be one of the best uh libraries for
- 39:32working with complex data. And there are
- 39:34a lot of changes that are being made in
- 39:36this library. And one modification is
- 39:39the cross-validation feature, which
- 39:41provides the ability to use more than
- 39:43one metric. Right? Cross-validation is
- 39:46one of the most important and one of the
- 39:47most easiest methods for checking the
- 39:50accuracy of a model. Right? So,
- 39:51cross-validation is being implemented in
- 39:53scikit-learn. And apart from that,
- 39:56again, there are a large spread of
- 39:57algorithms that you can implement by
- 39:59using scikit-learn. Right? These include
- 40:01unsupervised learning algorithms,
- 40:03starting from clustering, factor
- 40:05analysis, principal component analysis,
- 40:08to all the unsupervised neural networks.
- 40:11Scikit-learn is also very essential uh
- 40:13for feature extracting in images and
- 40:16text.
- 40:17So, mainly scikit-learn is used for
- 40:19implementing all the standard machine
- 40:21learning and data mining tasks like
- 40:24reducing dimensionality, classification,
- 40:26regression, clustering, and model
- 40:28selection. Next up, we have NumPy. Now,
- 40:30NumPy is considered as one of the most
- 40:33popular machine learning libraries in
- 40:35Python. Now, let me tell you that
- 40:36TensorFlow and other libraries, they
- 40:38make use of NumPy internally for
- 40:41performing multiple operations on
- 40:43tensors. The most important feature of
- 40:47NumPy is the array interface. It
- 40:49supports multi-dimensional arrays.
- 40:51Right? That's one of the most important
- 40:53features of NumPy. Another feature is uh
- 40:55it makes complex mathematical
- 40:57implementations very simple. Right? It's
- 41:00mainly known for computing mathematical
- 41:03data. So, NumPy is a package that you
- 41:05should be using for any sort of
- 41:07statistical analysis or data analysis
- 41:10that involves a lot of math. Apart from
- 41:12that, it makes coding very easy and
- 41:15grasping the concept is extremely easy
- 41:16with NumPy. Now, NumPy is mainly used
- 41:20for expressing images, sound waves, and
- 41:22other mathematical computations.
- 41:24All right? Moving on to our next
- 41:26library, we have Theano. Theano is a
- 41:29computational framework which is used
- 41:32for computing multi-dimensional arrays.
- 41:34Right? Theano actually works very
- 41:36similar to TensorFlow, but the only
- 41:38drawback is that you can't fit Theano
- 41:41into production environments. But apart
- 41:43from that, Theano allows you to define,
- 41:46optimize, and evaluate mathematical
- 41:48expressions that involve
- 41:49multi-dimensional arrays. Right? This is
- 41:52another library that lets you implement
- 41:54multi-dimensional arrays. Features of
- 41:56Theano include tight integration with
- 41:58NumPy. An advantage of Theano is that
- 42:01you can easily implement NumPy arrays in
- 42:03Theano. Right? That's why there's a
- 42:05connection between Theano and NumPy
- 42:07because both of them effectively use
- 42:09multi-dimensional arrays. Transparent
- 42:11use of GPU. Now, performing data
- 42:14intensive computations are much faster
- 42:16when it comes to uh Theano because of
- 42:18its use of GPU, right? Theano also lets
- 42:21you detect and diagnose multiple types
- 42:24of errors and any sort of ambiguity in
- 42:27the model. So, guys, Theano was actually
- 42:29designed to handle the types of
- 42:31computations required for large neural
- 42:34network algorithms, right? It was mainly
- 42:36built for deep learning and neural
- 42:38networks. It was one of the first
- 42:41libraries of its kind and it is
- 42:43considered as an industry standard for
- 42:46deep learning research and development.
- 42:48Theano is being used in multiple neural
- 42:50networks projects and the popularity of
- 42:52Theano is only going to grow with time,
- 42:54right? A lot of people actually haven't
- 42:56heard of Theano, but let me tell you
- 42:58that this is one of the best ways to
- 42:59implement deep learning and neural
- 43:01network models.
- 43:03Moving on, uh we have Keras. Now, Keras
- 43:05is considered to be the most popular
- 43:08Python package. It provides some of the
- 43:10best functionalities for compiling
- 43:12models, processing your data sets, and
- 43:15visualizing graphs. It is also popular
- 43:17in the implementation of neural
- 43:19networks, right? It is considered to be
- 43:21the simplest package uh with which you
- 43:23can implement neural networks. In fact,
- 43:25in our today's demo for deep learning,
- 43:27we'll be implementing Keras in order to
- 43:29understand how neural networks work. Few
- 43:32of the features of Keras include that it
- 43:34runs very smoothly on both CPU and GPU.
- 43:37It supports almost all the models of the
- 43:40neural network, right? From fully
- 43:42connected, convolutional, pooling,
- 43:44recurrent, embedding, all of these
- 43:46models are supported by Keras.
- 43:48And not only that, you can combine these
- 43:50models to build more complex models.
- 43:53Keras is completely Python-based, which
- 43:55makes it very easy to debug and explore,
- 43:58right? Since Python has a huge community
- 44:00of followers, it's very simple in order
- 44:03to debug any sort of error that you find
- 44:05while implementing Keras. So, the
- 44:07libraries that I discussed so far were
- 44:09dedicated to machine learning and deep
- 44:11learning. For natural language
- 44:13processing, we have the most famous
- 44:15library known as the Natural Language
- 44:17Toolkit, which is an open-source Python
- 44:19library, mainly used for natural
- 44:21language processing, text analysis, and
- 44:24text mining. The main features include
- 44:26that it studies and analyzes natural
- 44:28language text in order to draw useful
- 44:31information from all this natural
- 44:32language text. It performs text analysis
- 44:36and sentimental analysis by performing
- 44:38tasks such as stemming, lemmatization,
- 44:41tokenization, and so on. Now, don't
- 44:44worry if you don't know what any of
- 44:45those terms mean. I'll be discussing all
- 44:47of those terms with you by the end of
- 44:49today's session. So guys, these were a
- 44:51couple of Python-based libraries, which
- 44:53are very essential for implementing
- 44:55machine learning and deep learning and
- 44:58artificial intelligence when you're
- 45:00using Python, right? These libraries are
- 45:02perfect for implementing AI. So guys, if
- 45:05any of you have any doubts regarding the
- 45:07libraries or if you want to learn more
- 45:09about the libraries, I will leave a
- 45:11couple of links in the description box.
- 45:13You can go through those videos as well.
- 45:15So now, let's move on to the main topic
- 45:17of discussion, which is artificial
- 45:18intelligence.
- 45:20Now, before we get started with the
- 45:22demand of artificial intelligence, let
- 45:25me tell you that AI was invented long
- 45:27ago. AI goes back to the 19th century.
- 45:29It was not something that was recently
- 45:31invented, even though AI has recently
- 45:33gained a lot of popularity. We can say
- 45:36that in the past decade, AI has gained
- 45:39the maximum popularity. But, it was
- 45:41actually invented in the 19th century.
- 45:44Now, especially in the year 1950, there
- 45:46was somebody known as Alan Turing. I'm
- 45:48sure a lot of you have heard about the
- 45:50Turing test. The Turing test is
- 45:52basically used to determine whether or
- 45:54not a machine is artificially
- 45:57intelligent, meaning that whether a
- 45:59machine can think intelligently like a
- 46:01human being.
- 46:02Right? This was the first proposition
- 46:04and this was one of the most important
- 46:07breakthroughs in artificial
- 46:08intelligence. Right? Somebody known as
- 46:10Alan Turing, he published a landmark
- 46:13paper in which he speculated about the
- 46:15possibility of creating machines that
- 46:17think. Right? So, the Turing test was
- 46:20the first serious proposal in the
- 46:22philosophy of artificial intelligence.
- 46:24This was done in 1950. Right? After
- 46:27this, we had eras of AI. We had the game
- 46:30AI which was in 1951. Now, since the
- 46:33emergence of AI in 1950s, we have seen
- 46:37an exponential growth in its potential.
- 46:39Right? AI covers domains like machine
- 46:41learning, deep learning, neural
- 46:43networks, natural language processing,
- 46:45knowledge base, and so on. It's also
- 46:47made its way into computer vision and
- 46:49image processing. But, the question is
- 46:51if AI has been here for over half a
- 46:54century,
- 46:55why has it suddenly gained so much
- 46:57importance? Right? Why are we talking
- 47:00about artificial intelligence now? The
- 47:02main reasons for the vast popularity of
- 47:04AI are the following. Right? The first
- 47:06reason is more computational power. Now,
- 47:09AI requires a lot of computing power.
- 47:12Recently, many advances have been made
- 47:14and complex deep learning models can be
- 47:16deployed. And one of the greatest
- 47:18technology that made this possible are
- 47:20GPUs. Since the invention of GPUs, we
- 47:23can compute much more with our
- 47:25computers. Initially, we could barely
- 47:27process 1 GB of data. Right? We only had
- 47:30hard disk to store additional memory and
- 47:32all of that. Now, our computers can
- 47:34process tons and tons of data. So, now
- 47:37we have more computational power, which
- 47:39is one of the main reasons behind why AI
- 47:41became so popular. So, by having more
- 47:44computational power, it becomes much
- 47:46easier to implement artificial
- 47:47intelligence. Next reason is more data.
- 47:50Now, big data is one of the most
- 47:52important reasons behind the development
- 47:55of artificial intelligence. Now, AI and
- 47:58data science and machine learning, deep
- 48:00learning, all of these processes are
- 48:03here only because we have a lot of data
- 48:05at present. Now, the main idea behind
- 48:08all these technologies is to draw useful
- 48:10insights from data. Now, since we start
- 48:12generating a lot of data, we need to
- 48:14find a method that can process this much
- 48:17data and draw useful insights from data
- 48:20such that it benefits an organization or
- 48:23it grows a business. That's why
- 48:25artificial intelligence and machine
- 48:27learning comes into the picture. Right?
- 48:28So, more data led to the demand of
- 48:31artificial intelligence. Apart from
- 48:33this, we also have better algorithms
- 48:35now, right? We have state-of-the-art
- 48:37algorithms. Most of them are based on
- 48:40the idea of neural networks and these
- 48:42are constantly getting better. Neural
- 48:43networks are actually one of the most
- 48:45significant discoveries in artificial
- 48:48intelligence because with neural
- 48:50networks, you can take in thousand
- 48:52layers of input data. Right? You can
- 48:54take in a lot of input data to perform
- 48:56computations. So, through neural
- 48:58networks, we are actually able to solve
- 49:00a lot of problems including healthcare
- 49:02problems, fraud detection problems, and
- 49:04so on. Another reason is broad
- 49:06investment. So, our universities and
- 49:09governments and startups and any tech
- 49:12giants like Google, Amazon, and
- 49:14Facebook, they are all investing heavily
- 49:16in artificial intelligence, which also
- 49:18led to the demand of AI. So, AI is
- 49:21rapidly growing both as a field of study
- 49:23and also as an economy. Right? It's
- 49:26adding a lot to the economy and I think
- 49:29this is the perfect time for you to get
- 49:30into the field of artificial
- 49:31intelligence because right now AI is in
- 49:34a really high demand. AI, machine
- 49:36learning, data science, all of this are
- 49:38of really high demand at present. All
- 49:41right. So, this is the perfect time for
- 49:42you to get started with artificial
- 49:43intelligence. Now, let me tell you that
- 49:45the term artificial intelligence was
- 49:47first coined in the year 1956
- 49:50by a scientist known as John McCarthy.
- 49:53Now, John McCarthy defined artificial
- 49:56intelligence as the science and
- 49:57engineering of making intelligent
- 50:00machines. So, now let's move on and talk
- 50:02about how artificial intelligence is
- 50:05different from machine learning and deep
- 50:06learning. A lot of people uh tend to
- 50:08assume that artificial intelligence,
- 50:10machine learning, and deep learning are
- 50:12the same because they have common
- 50:14applications, right? For example, Siri
- 50:17is an application of AI, machine
- 50:19learning, and deep learning. So, how are
- 50:22these technologies uh related, right? Or
- 50:24how are they different from each other?
- 50:26Now, artificial intelligence is the
- 50:28science of getting machines to mimic the
- 50:30behavior of human beings. Machine
- 50:33learning is the subset of artificial
- 50:36intelligence that focuses on getting
- 50:38machines to make decisions by feeding
- 50:40them data. Deep learning, on the other
- 50:43hand, is a subset of machine learning
- 50:45that uses the concept of neural networks
- 50:48to solve complex problems.
- 50:50So, to sum it up to you, artificial
- 50:52intelligence, machine learning, and deep
- 50:54learning are heavily interconnected
- 50:56fields, right? Machine learning and deep
- 50:58learning aids artificial intelligence by
- 51:00providing a set of algorithms and neural
- 51:03networks to solve data-driven problems.
- 51:05However, AI is not restricted to only
- 51:08machine learning and deep learning,
- 51:09right? It covers a vast domain of fields
- 51:12which include natural language
- 51:13processing, object detection, computer
- 51:15vision, robotics, expert systems, and so
- 51:18on, right? So, AI is a very vast field.
- 51:20Guys, I hope I cleared the difference
- 51:22between AI, machine learning, and deep
- 51:24learning. Also, a lot of you might be
- 51:26confused about data science. Data
- 51:28science is now an umbrella term, right?
- 51:31Data science basically means to derive
- 51:33useful insights from data. So, data
- 51:35science actually uh uses AI, machine
- 51:38learning, and deep learning, right? So,
- 51:40it implements all of these three
- 51:42technologies in order to derive useful
- 51:44insights from data, right? Now, let's
- 51:47move on to the most interesting topic in
- 51:49artificial intelligence, which is
- 51:51machine learning. Now guys, the term
- 51:53machine learning was first coined by a
- 51:55scientist known as Arthur Samuel in the
- 51:58year 1959.
- 51:59Looking back, that year was probably the
- 52:02most significant in terms of
- 52:03technological advancements.
- 52:05In order to define machine learning, if
- 52:08you browse the internet for what is
- 52:10machine learning, you'll get at least
- 52:11100 different definitions.
- 52:13In simple terms, machine learning is a
- 52:16subset of artificial intelligence, which
- 52:18provides machines the ability to learn
- 52:21automatically and improve from
- 52:23experience without being explicitly
- 52:26programmed to do so. In a sense, it is
- 52:29the practice of getting machines to
- 52:31solve problems by gaining the ability to
- 52:33think. Now the question here is, can a
- 52:36machine think or can a machine make
- 52:38decisions? Well, if you feed a machine a
- 52:41good amount of data, it will learn how
- 52:43to interpret, process, and analyze this
- 52:46data by using something known as machine
- 52:48learning algorithms. To give you a basic
- 52:51idea of how the machine learning process
- 52:53works, look at the figure on this slide.
- 52:56A machine learning process always begins
- 52:58by feeding the machine lots and lots of
- 53:00data. Now by using this data, the
- 53:03machine is trained to detect any hidden
- 53:05insights and trends in the data. These
- 53:08insights are then used to build a
- 53:10machine learning model by using a
- 53:12machine learning algorithm in order to
- 53:14solve a problem. The basic aim of
- 53:17machine learning is to solve a problem
- 53:19or find a solution by using data. Now
- 53:22moving ahead, I'll be discussing the
- 53:23machine learning process in depth,
- 53:25right? So don't worry if you haven't got
- 53:27the exact idea of what machine learning
- 53:29is.
- 53:30Now the machine learning process
- 53:32involves building a predictive model
- 53:34that can be used to find a solution for
- 53:37a particular problem. A well-defined
- 53:39machine learning process will have
- 53:41around seven steps. It always begins
- 53:44with defining the objective followed by
- 53:46data gathering or data collection. Then
- 53:49we have something known as preparing
- 53:51data, which is also called data
- 53:53pre-processing. Then we have data
- 53:55exploration or exploratory data
- 53:57analysis. This is followed by building a
- 54:00machine learning model.
- 54:02Then we have model evaluation and
- 54:04finally predictions. This is how the
- 54:06process of machine learning works. To
- 54:08understand the machine learning process,
- 54:10let's assume that you've been given a
- 54:12problem that needs to be solved by using
- 54:14machine learning. Let's say that the
- 54:16problem is to predict the occurrence of
- 54:18rain in your local area by using machine
- 54:21learning. Now, the first step is to
- 54:23define the objective of the problem.
- 54:25Right? At this step we must understand
- 54:27what exactly needs to be predicted. In
- 54:29our case, the objective is to predict
- 54:31the possibility of rain by studying the
- 54:34weather conditions. So, at this stage it
- 54:36is essential to take mental notes on
- 54:39what kind of data can be used to solve
- 54:41this problem or the type of approach
- 54:43that you must follow to get to the
- 54:45solution.
- 54:46The questions you should be asking
- 54:48yourself is what are we trying to
- 54:50predict? Right? Here we're trying to
- 54:51predict whether it'll rain or not.
- 54:54Right? You need to understand what are
- 54:56the target features. Target features are
- 54:59basically the variable that you need to
- 55:01predict. Here we need to predict a
- 55:03variable that'll show us whether it's
- 55:04going to rain tomorrow or not. Then you
- 55:07must also understand what kind of data
- 55:09you'll need to solve this problem. Apart
- 55:11from that, you need to know what kind of
- 55:13problem you're facing. Is it a binary
- 55:15classification problem or is it a
- 55:17clustering problem? Now, if you don't
- 55:19know what classification and clustering
- 55:21is, don't worry. I'll be talking about
- 55:23all of these things in the upcoming
- 55:24slides. So, your first step is to define
- 55:27the objective of your problem. You need
- 55:30to understand what exactly needs to be
- 55:32done here. Right? How can you solve this
- 55:34problem?
- 55:35Moving on, your next step is to gather
- 55:37the data that you need. At this stage,
- 55:39you must be asking questions such as
- 55:42what kind of data is needed to solve
- 55:44this problem. Is the data available to
- 55:46me? And if it's not available, how can I
- 55:48get the data? Right? Once you know the
- 55:51type of data that is required, you must
- 55:53understand how you can derive this data.
- 55:56Data collection can be either done
- 55:58manually or it can be done by web
- 56:00scraping. But don't worry if you're a
- 56:02beginner and you're just looking to
- 56:04learn machine learning, you don't have
- 56:05to worry about getting the data.
- 56:08There are thousands of data resources on
- 56:10the web. You can just download the data
- 56:12set and you can get going.
- 56:14Coming back to the problem at hand, the
- 56:16data needed for weather forecasting
- 56:18includes measures such as humidity
- 56:20level, your temperature, the pressure,
- 56:23the locality, whether or not you live in
- 56:26a hill station, and so on.
- 56:28Such data must be collected and it has
- 56:30to be stored for analysis. This is where
- 56:33you collect all the data. Now, moving on
- 56:35to step number three is data
- 56:37preparation. The data that you collected
- 56:40is almost never in the right format. All
- 56:43right, even if you collect it from a
- 56:45internet resource, if you download it
- 56:47from some website, even then your data
- 56:50is not going to be clean. Right? It's
- 56:52not going to be in the correct format.
- 56:53There's always going to be some sort of
- 56:55inconsistencies in your data.
- 56:58Inconsistencies include any missing
- 57:00values or any redundant variables,
- 57:03duplicate values. All of these are
- 57:05inconsistencies.
- 57:07Removing all of this is very essential
- 57:09because they might lead to any wrongful
- 57:11computation. Therefore, at this stage,
- 57:13you can scan the entire data set for any
- 57:16missing values and you have to fix them
- 57:18here itself.
- 57:20Now, actually, this is one of the most
- 57:21time-consuming steps in a machine
- 57:23learning process. If you ask a data
- 57:26scientist which step he hates the most
- 57:28or which step is, you know, the most
- 57:30time-consuming, they're probably going
- 57:32to tell you data processing and data
- 57:33cleaning. Right? It's one of the most
- 57:36tiresome task because you need to look
- 57:38at all the values that are there. You
- 57:39need to find any missing values, any
- 57:41data that is not relevant to you. Right?
- 57:43All of this has to be removed so that
- 57:45you can analyze the data in a better
- 57:47way.
- 57:48Now, step number four is exploratory
- 57:50data analysis.
- 57:52So, guys, this stage is all about
- 57:54getting deep into your data and finding
- 57:57all the hidden data mysteries.
- 58:00EDA or exploratory data analysis is like
- 58:03the brainstorming stage of machine
- 58:05learning.
- 58:06Data exploration involves understanding
- 58:08the patterns and the trends in your
- 58:09data. So, at this stage all the useful
- 58:12insights are drawn and any correlations
- 58:15between the variables are understood.
- 58:17For example, in the case of predicting
- 58:19rainfall, we know that there is a strong
- 58:22possibility of rain if the temperature
- 58:24has fallen low. Such correlations have
- 58:27to be understood and mapped at this
- 58:29stage.
- 58:30EDA is actually the most important step
- 58:32in a machine learning process because
- 58:34here is where you understand your data.
- 58:37You understand how your data is going to
- 58:39help you predict the outcome.
- 58:41Moving on to step number five, we have
- 58:44building a machine learning model. So,
- 58:46all the insights and all the patterns
- 58:49that you got from your data exploration
- 58:51stage, those insights are used to build
- 58:53the machine learning model. So, this
- 58:55stage always begins by splitting the
- 58:57data set into two parts, that is
- 59:00training and testing data. Now, remember
- 59:02that the training data will be used to
- 59:05build and analyze the model.
- 59:07The model is basically the machine
- 59:09learning algorithm that predicts the
- 59:11output by using the data that you feed
- 59:13to it. An example of machine learning
- 59:15algorithm is logistic regression and
- 59:18linear regression. All of these are
- 59:20machine learning algorithms.
- 59:22Now, don't worry about choosing the
- 59:23right algorithm. Right? First, we'll
- 59:25focus on what the machine learning
- 59:26process is.
- 59:28But anyway, choosing the right algorithm
- 59:30will depend on several factors, right?
- 59:32It depends on the type of problem you're
- 59:34trying to solve, the data set, and the
- 59:36level of complexity of the problem.
- 59:39In the upcoming sections, we'll discuss
- 59:40all the different types of problems that
- 59:42can be solved by using machine learning.
- 59:44Moving on to step number six, we have
- 59:46model evaluation and optimization.
- 59:49Now, after you build a model by using
- 59:51the training data set, it is finally
- 59:53time to put the model to a test. The
- 59:56testing data set is used to check the
- 59:58efficiency of the model and how
- 1:00:00accurately it can predict the outcome.
- 1:00:03Now, once the accuracy is calculated and
- 1:00:06any further improvements in the model,
- 1:00:08they have to be implemented at this
- 1:00:10stage. Methods like parameter tuning and
- 1:00:13cross-validation can be used to improve
- 1:00:15the performance of the model.
- 1:00:17Before I move any further, I don't know
- 1:00:19if all of you know what training and
- 1:00:21testing data set means. In machine
- 1:00:23learning, the input data is always
- 1:00:25divided into two sets. We have something
- 1:00:28known as the training data set, and we
- 1:00:29have something known as the testing data
- 1:00:31set.
- 1:00:32So, in machine learning, you always
- 1:00:34split the data into two parts, right?
- 1:00:36This process is known as a data
- 1:00:38splicing. Now, the training data set
- 1:00:40will be used to build the machine
- 1:00:42learning model, and the testing data set
- 1:00:44will be used to test the efficiency of
- 1:00:47the model that you built. This is what
- 1:00:49training and testing data set is.
- 1:00:51They're not any different data that you
- 1:00:53derive. They're the same as the input
- 1:00:55data set. The only thing is you are
- 1:00:57splitting the data set so that you can
- 1:00:59train the model on one data and test the
- 1:01:01model on another data.
- 1:01:03Now, remember that the training data set
- 1:01:05is always larger in size when compared
- 1:01:08to the testing data set. Because
- 1:01:10obviously, you are training and building
- 1:01:12the model by using the training data
- 1:01:13set. The testing data set is just for
- 1:01:16evaluating the performance of your
- 1:01:18model.
- 1:01:19Now, let's move on and understand step
- 1:01:21number seven, which is predictions. Now,
- 1:01:24once a model is evaluated and you've
- 1:01:26improved the model, it is finally used
- 1:01:29to make predictions. The final output
- 1:01:31can be a categorical variable or it can
- 1:01:34be a continuous quantity. Right? All of
- 1:01:36this depends on the type of problem
- 1:01:38you're trying to solve. Don't worry,
- 1:01:39I'll be discussing the type of problems
- 1:01:41that can be solved using machine
- 1:01:43learning in the upcoming slides. In our
- 1:01:45case for predicting the occurrence of
- 1:01:47rainfall, the output will be a
- 1:01:49categorical variable.
- 1:01:51Categorical variable is anything that
- 1:01:53has some categorical value. For example,
- 1:01:56gender is a categorical variable. Gender
- 1:01:59has either male, female, or other.
- 1:02:02It has a defined set of values. That is
- 1:02:04a categorical variable. So guys, that
- 1:02:06was the entire machine learning process.
- 1:02:09Now, as we continue with this tutorial,
- 1:02:11in the upcoming sections, I will be
- 1:02:14running a demo in Python, in which we
- 1:02:16will be performing weather forecasting.
- 1:02:18So, make sure you remember all these
- 1:02:20steps that I spoke about because I'll be
- 1:02:21going through all these steps by using
- 1:02:24Python. We'll be coding all of this that
- 1:02:26we just spoke about.
- 1:02:28Now, the next topic we're going to
- 1:02:29discuss is the types of machine
- 1:02:31learning.
- 1:02:32A machine can learn to solve a problem
- 1:02:34by following any one of the three
- 1:02:37approaches.
- 1:02:38You can say that there are three ways in
- 1:02:40which a machine learns. The three ways
- 1:02:43are supervised learning, unsupervised
- 1:02:45learning, and reinforcement learning.
- 1:02:48These are the three methods in which you
- 1:02:50can train a machine to learn.
- 1:02:52So first, let's discuss supervised
- 1:02:54learning.
- 1:02:55So, what is supervised learning?
- 1:02:57Supervised learning is a technique in
- 1:02:59which we teach or train the machine by
- 1:03:01using data which is labeled. To
- 1:03:04understand this better, let's consider
- 1:03:06an analogy.
- 1:03:07As kids, we all needed guidance to solve
- 1:03:10math problems. At least I had a really
- 1:03:12tough time solving math problems. Yeah,
- 1:03:15so our teachers always helped us
- 1:03:17understand what addition is and how it
- 1:03:19is done. Similarly, you can think of
- 1:03:21supervised learning as a type of machine
- 1:03:24learning that involves a guide. The
- 1:03:26label data set is a teacher that will
- 1:03:28train the machine to understand the
- 1:03:30patterns in the data. The label data set
- 1:03:33is nothing but the training data set.
- 1:03:35So, to better understand this, consider
- 1:03:37the figure. Right here, we're feeding
- 1:03:40the machine images of Tom and Jerry, and
- 1:03:42the goal is for the machine to identify
- 1:03:45and classify the images into two
- 1:03:47separate groups. Basically, one group
- 1:03:49will contain Tom images, and the other
- 1:03:51group will contain images of Jerry. Now,
- 1:03:54pay attention to the training data set.
- 1:03:56The training data set that is fed to a
- 1:03:58model is labeled. As in, we're telling
- 1:04:01the machine, "Listen, this is how Tom
- 1:04:03looks, and this is how Jerry looks."
- 1:04:05But, basically, labeling each data point
- 1:04:08that we're feeding to the machine.
- 1:04:10Right? If the image is of Tom's, we've
- 1:04:12labeled it as Tom, and if the image is a
- 1:04:15Jerry image, then we're going to label
- 1:04:17it as Jerry. By doing this, you're
- 1:04:19training the machine by using labeled
- 1:04:22data. So, to sum it up, in supervised
- 1:04:24learning, there is a well-defined
- 1:04:26training phase done with the help of
- 1:04:28labeled data. Right? The rest of the
- 1:04:30process is the same. After you feed the
- 1:04:33machine labeled data, you're going to
- 1:04:34perform data cleaning, then exploratory
- 1:04:37data analysis, followed by building the
- 1:04:39machine learning model, and then model
- 1:04:41evaluation, and finally, your
- 1:04:43predictions. Also, one more point to
- 1:04:45remember is that the output that you're
- 1:04:47going to get in a supervised learning
- 1:04:49algorithm is a labeled output. This
- 1:04:52Jerry will be labeled as Jerry, and this
- 1:04:54Tom will be labeled as Tom. Basically,
- 1:04:56you'll get a labeled output. Now, let's
- 1:04:59understand what is unsupervised
- 1:05:01learning.
- 1:05:02Unsupervised learning involves training
- 1:05:05by using unlabeled data and allowing the
- 1:05:07model to act on that information without
- 1:05:10any guidance. So, think of unsupervised
- 1:05:13learning as a smart kid that learns
- 1:05:15without any guidance. In this type of
- 1:05:17machine learning, the model is not fed
- 1:05:20with any label data. As in, the model
- 1:05:22has no clue that this image is Tom and
- 1:05:25this image is Jerry. It figures out
- 1:05:28patterns and the differences between Tom
- 1:05:30and Jerry on its own by taking in tons
- 1:05:32of data. For example, it identifies
- 1:05:35prominent features of Tom such as pointy
- 1:05:38ears, bigger in size, and so on to
- 1:05:40understand that this image is of type
- 1:05:42one.
- 1:05:43Similarly, it finds such features in
- 1:05:45Jerry and knows that this is another
- 1:05:47type of image, [clears throat] maybe
- 1:05:48type two. Right? Therefore, it
- 1:05:50classifies the images into two different
- 1:05:52clusters without knowing who is Tom and
- 1:05:55who is Jerry. Now, the main idea behind
- 1:05:57unsupervised learning is to understand
- 1:05:59the patterns in your data set and form
- 1:06:01clusters based on feature similarity.
- 1:06:04Basically, it'll feature similar images
- 1:06:06or similar data points into one cluster,
- 1:06:09and it'll form another cluster which is
- 1:06:11totally different from the first
- 1:06:13cluster. So, look at the output over
- 1:06:15here. The unlabeled output is basically
- 1:06:17clusters or groups of two different
- 1:06:20data.
- 1:06:21Next, we have something known as the
- 1:06:22reinforcement learning. Now,
- 1:06:24reinforcement learning is comparatively
- 1:06:26different, right? It's pretty different
- 1:06:28from supervised and unsupervised.
- 1:06:31It is basically a part of machine
- 1:06:32learning where you put an agent in an
- 1:06:35environment, and this agent learns to
- 1:06:37behave in the environment by performing
- 1:06:40certain actions and observing the
- 1:06:42rewards which it gets from these
- 1:06:44actions. To understand reinforcement
- 1:06:46learning, imagine that you were dropped
- 1:06:49off at an isolated island. What would
- 1:06:51you do? Initially, we'd all panic,
- 1:06:54right? But, as time passes by, you will
- 1:06:56learn how to live on the island. You
- 1:06:59will explore the environment. You will
- 1:07:01understand the climate conditions.
- 1:07:03You'll understand the type of food that
- 1:07:05grows there. You'll know what is
- 1:07:07dangerous to you and what is not. You'll
- 1:07:09understand which food is good for you
- 1:07:11and which is not. This is exactly how
- 1:07:13reinforcement learning works. It
- 1:07:15involves an agent, which is basically
- 1:07:17you stuck on the island, that is put in
- 1:07:20an unknown environment, which is the
- 1:07:22island, where the agent must learn by
- 1:07:24observing and performing actions that
- 1:07:27result in rewards.
- 1:07:28Reinforcement learning is mainly used in
- 1:07:30advanced machine learning areas such as
- 1:07:33self-driving cars, AlphaGo, and so on.
- 1:07:35So, guys, that sums up the types of
- 1:07:37machine learning. Before we go any
- 1:07:39further, I'd like to discuss the
- 1:07:40difference between supervised,
- 1:07:42unsupervised, and reinforcement
- 1:07:43learning. Now, first of all, we have the
- 1:07:45definition. Supervised learning is all
- 1:07:48about teaching a machine by using
- 1:07:50labeled data. Unsupervised learning,
- 1:07:53like the name suggests, there is no
- 1:07:55supervision over here. The machine is
- 1:07:57trained on unlabeled data without any
- 1:08:00guidance.
- 1:08:01Reinforcement learning is totally
- 1:08:02different. Here, you have an agent who
- 1:08:04interact with the environment by
- 1:08:06producing actions and discover some
- 1:08:09errors and rewards.
- 1:08:10Now, the type of problem that is solved
- 1:08:12using supervised learning is regression
- 1:08:14and classification problems. We'll
- 1:08:16discuss what regression, classification,
- 1:08:18and clustering is in the upcoming slide,
- 1:08:21right? So, don't worry if you don't know
- 1:08:22what it is. Unsupervised learning is
- 1:08:24mainly to solve association and
- 1:08:26clustering problems.
- 1:08:28Reinforcement learning is for
- 1:08:29reward-based problems.
- 1:08:31Now, what is the type of data in
- 1:08:33supervised learning? It is labeled data.
- 1:08:35That is the main difference between
- 1:08:37supervised and any other type of machine
- 1:08:39learning.
- 1:08:40In supervised, you have labeled data. In
- 1:08:42unsupervised, we have unlabeled data.
- 1:08:44Whereas, in reinforcement learning, we
- 1:08:46have no predefined data at all. The
- 1:08:49machine has to perform everything from
- 1:08:51scratch. It has to collect data,
- 1:08:53analyze, do everything on its own. Now,
- 1:08:55the training in supervised learning is
- 1:08:57external supervision, meaning that we
- 1:09:00have external supervision in the form of
- 1:09:02the labeled training data set. In
- 1:09:05unsupervised, there is obviously no
- 1:09:07supervision. There is an unlabeled data
- 1:09:09set, therefore there's no supervision.
- 1:09:11In reinforcement learning, there is no
- 1:09:13supervision at all. Now, the approach to
- 1:09:15solving supervised learning problem is
- 1:09:17basically you're going to map your
- 1:09:19labeled input to your known output. In
- 1:09:22unsupervised learning, the machine is
- 1:09:23going to understand the patterns and
- 1:09:25discover the output on its own.
- 1:09:27Reinforcement learning, here the agent
- 1:09:30will follow something known as a trial
- 1:09:32and error method. Right? It's totally
- 1:09:34based on the concept of trial and error.
- 1:09:37Popular algorithms under supervised
- 1:09:39learning are linear regression, logistic
- 1:09:42regression, support vector machines, and
- 1:09:44so on. Under unsupervised learning, we
- 1:09:46have the famous K-means clustering
- 1:09:48algorithm. Under reinforcement learning,
- 1:09:50we have the Q-learning algorithm, which
- 1:09:52is one of the most important algorithms.
- 1:09:55It is basically the logic behind the
- 1:09:57famous AlphaGo game. I'm sure all of you
- 1:09:59have heard of that. So, guys, these were
- 1:10:01the differences between supervised,
- 1:10:03unsupervised, and reinforcement
- 1:10:04learning. Now, let's move on and discuss
- 1:10:07the type of problems that you can solve
- 1:10:09by using machine learning.
- 1:10:11Now, there are three types of problems
- 1:10:13in machine learning. Now, any problem
- 1:10:15that needs to be solved in machine
- 1:10:17learning can fall into one of these
- 1:10:19three categories. Now, what is a
- 1:10:22regression? In this type of problem, the
- 1:10:24output is a continuous quantity. For
- 1:10:27example, if you want to predict the
- 1:10:29speed of a car given the distance. That
- 1:10:32means it is a regression problem.
- 1:10:34First of all, what is a continuous
- 1:10:36quantity? A continuous quantity is any
- 1:10:38variable that can hold a continuous
- 1:10:40value. A continuous variable is any
- 1:10:43variable that can have infinite number
- 1:10:45of values.
- 1:10:47For example, the height of a person or
- 1:10:49the weight of a person is a continuous
- 1:10:51quantity. Right? I can have a weight of
- 1:10:5350.1 kgs or 50.12 or 50.112 kg.
- 1:10:59This is a continuous quantity.
- 1:11:01Regression problems can be solved by
- 1:11:03using supervised learning algorithms.
- 1:11:06Another type of problem is a
- 1:11:07classification problem. Here, the output
- 1:11:10is always a categorical value.
- 1:11:13Classifying emails into two classes, for
- 1:11:15example, classifying your email as spam
- 1:11:17and non-spam is a classification
- 1:11:19problem.
- 1:11:20Here again, you'll be using supervised
- 1:11:22learning classification algorithms such
- 1:11:24as support vector machines, naive bias,
- 1:11:27logistic regression, and so on.
- 1:11:29Then we have clustering problem. And
- 1:11:31this type of problem involves assigning
- 1:11:34the input into two or more clusters
- 1:11:36based on feature similarity. For
- 1:11:39example, clustering the viewers into
- 1:11:41similar groups based on their interest
- 1:11:44or based on their age or geography can
- 1:11:46be done by using unsupervised learning
- 1:11:48algorithms like K-means clustering.
- 1:11:51One thing you need to understand is
- 1:11:53under supervised learning, you can solve
- 1:11:55regression and classification problems.
- 1:11:58Under unsupervised learning, you can
- 1:11:59solve clustering problems. Reinforcement
- 1:12:02learning is something else altogether,
- 1:12:05right? You can solve reward-based
- 1:12:06problems and more complex and deep
- 1:12:08problems. So, now let's move on and
- 1:12:12understand the different machine
- 1:12:13learning algorithms. Now, I will not be
- 1:12:16going into depth for machine learning
- 1:12:18algorithms because there are a lot of
- 1:12:19algorithms to cover, but we have content
- 1:12:22around almost every machine learning
- 1:12:24algorithm out there. So, I'll be leaving
- 1:12:26a couple of links in the description
- 1:12:28box, right? You can check out all these
- 1:12:30links and understand how each of these
- 1:12:32machine learning algorithms work in
- 1:12:34depth. So, I'm just going to show you a
- 1:12:36hierarchical diagram of how the
- 1:12:38algorithms are structured. So, under
- 1:12:40machine learning, we have three types of
- 1:12:42learning. We have supervised,
- 1:12:43unsupervised, and reinforcement. Under
- 1:12:45supervised learning, we have regression
- 1:12:47and classification problems. And under
- 1:12:50unsupervised learning, we have
- 1:12:51clustering problems. Reinforcement
- 1:12:54learning is completely different. I'll
- 1:12:55be leaving a link in the description
- 1:12:57specifically for reinforcement learning.
- 1:12:59You can check out the entire content of
- 1:13:01reinforcement learning there.
- 1:13:03Now, regression problems can be solved
- 1:13:05by using linear regression algorithm
- 1:13:07such as linear regression, decision
- 1:13:10trees, and random forest can also be
- 1:13:11used in regression problems. But,
- 1:13:14usually decision trees and random
- 1:13:16forest, all of these are used to solve
- 1:13:18classification problems. Famous
- 1:13:20classification algorithms include
- 1:13:22K-nearest neighbor, which is basically
- 1:13:24KNN, decision trees and random forest,
- 1:13:27logistic regression, naive bias, support
- 1:13:30vector machines. All of these are
- 1:13:32classification algorithms. Coming to
- 1:13:34unsupervised learning, we have
- 1:13:36clustering and association analysis. And
- 1:13:39clustering problems can be solved by
- 1:13:40using K-means. And association analysis
- 1:13:43can be solved by using a priori
- 1:13:45algorithm. A priori algorithm is mainly
- 1:13:48used in market basket analysis. Right?
- 1:13:51For this algorithm as well, I'll be
- 1:13:52leaving a link in the description. We've
- 1:13:55performed a very excellent demo where in
- 1:13:57we've shown how market basket analysis
- 1:13:59can be done by using a priori algorithm.
- 1:14:02Markov model is also explained in one of
- 1:14:05the videos. I'll be leaving that link in
- 1:14:07the description box.
- 1:14:08Now, to sum up machine learning to you,
- 1:14:10I'll be running a small demonstration in
- 1:14:13Python. Right? Like I promised earlier,
- 1:14:16I'll be using Python to understand the
- 1:14:18whole machine learning process. All
- 1:14:20right. So, let's get started with that
- 1:14:21demo.
- 1:14:22So guys, for those of you who don't know
- 1:14:24Python, I will leave a couple of links
- 1:14:26in the description box so that you
- 1:14:27understand Python. But, apart from that,
- 1:14:30Python is pretty understandable. If you
- 1:14:31just look at the code, you'll know what
- 1:14:33exactly I'm talking about. Right? So,
- 1:14:35don't worry. And also, I'll be
- 1:14:36explaining everything in the code.
- 1:14:39So, I'm using PyCharm in order to run
- 1:14:42the demo.
- 1:14:43Right? So guys, like I said, if you
- 1:14:44don't know Python, I'll leave a couple
- 1:14:46of links in the description box. You can
- 1:14:48go through those videos as well. The
- 1:14:51main aim of our demo is to build a
- 1:14:53machine learning model that will predict
- 1:14:55whether or not it will rain tomorrow by
- 1:14:58studying the past data set. Now, this
- 1:15:00data set contains around 145,000
- 1:15:03observations on the daily weather
- 1:15:05conditions as observed in Australia.
- 1:15:08Right, the data set has around 24
- 1:15:10features and we will be using 23
- 1:15:13features out of that to predict the
- 1:15:15target variable which is rain tomorrow.
- 1:15:18So, this data set I collected from
- 1:15:19Kaggle. Right, for those of you don't
- 1:15:21know Kaggle is a online platform where
- 1:15:23you can find hundreds of data sets and
- 1:15:26you know, there are a lot of
- 1:15:26competitions held by machine learning
- 1:15:28engineers and all of that. It's an
- 1:15:31interesting website.
- 1:15:33Now, the problem statement itself is to
- 1:15:35build a machine learning model that will
- 1:15:37predict whether or not it will rain
- 1:15:39tomorrow.
- 1:15:40This is clearly a classification
- 1:15:42problem. The machine learning model has
- 1:15:44to classify the output into two classes,
- 1:15:47that is either yes or no. Yes will stand
- 1:15:50for it will rain tomorrow and no will
- 1:15:52basically denote that it will not rain
- 1:15:54tomorrow. Right, this is a
- 1:15:55classification problem.
- 1:15:57So, I hope the objective is clear.
- 1:15:59Right, so we'll begin the demonstration
- 1:16:01by importing the required libraries. So,
- 1:16:04first of all, for mathematical
- 1:16:06computations, we'll be importing the
- 1:16:07NumPy library. We'll also be importing
- 1:16:10the Pandas library for data processing.
- 1:16:13Next, we will load the CSV file.
- 1:16:15Basically, my data is stored in a CSV
- 1:16:18format in this file. weatherAUS.csv is
- 1:16:21my data set.
- 1:16:22So, basically I've saved this file in
- 1:16:24this path. Right, so that's what I'm
- 1:16:26doing here. I'm loading my data set and
- 1:16:28I'm storing it in a variable known as
- 1:16:31DF.
- 1:16:32Next, what we'll do is we'll see the
- 1:16:33size of our data frame.
- 1:16:36Let's print the size of the data frame.
- 1:16:38We'll also display the first five
- 1:16:40observations in our data frame.
- 1:16:42Let's look at the output.
- 1:16:45Basically, around 145,000
- 1:16:47observations and 24 features. Now, 24
- 1:16:51features are basically the variables
- 1:16:53that are there in my data set. You know,
- 1:16:55for example, date is a variable,
- 1:16:57location is a variable, minimum
- 1:16:59temperature till rain tomorrow. All of
- 1:17:01these are variables. So, I have around
- 1:17:0224 features in my data set, right? Now,
- 1:17:05the variable that I have to predict is
- 1:17:07rain tomorrow. Okay? If the value of
- 1:17:10rain tomorrow is no, it denotes that it
- 1:17:12will not rain tomorrow. But, if the
- 1:17:14value is yes, then it will denote that
- 1:17:16it will rain tomorrow.
- 1:17:18So, rain tomorrow is basically my target
- 1:17:20variable.
- 1:17:21Right? I'll be finding out whether it's
- 1:17:23going to rain tomorrow or not. So, this
- 1:17:25is my target variable, also known as
- 1:17:27your output variable.
- 1:17:28My input variables will be the other 23
- 1:17:31variables. Date, location, minimum
- 1:17:33temperature, rain today, risk, all of
- 1:17:36this will be my input variables. Now,
- 1:17:38these variables are also known as
- 1:17:40predictor variables. Basically, they're
- 1:17:42used to predict your outcome. So, these
- 1:17:44are also known as predictor variables.
- 1:17:47Now, the next step is checking for null
- 1:17:49values.
- 1:17:50This is basically data pre-processing.
- 1:17:53Let me just comment it for you.
- 1:17:55This is data preparation or data
- 1:18:01Right. So, this stage is data
- 1:18:03pre-processing. Right here, we start
- 1:18:05checking for any null values or any
- 1:18:07missing values. This is exactly what I'm
- 1:18:09doing over here. I am checking for any
- 1:18:11missing or null values in my data set.
- 1:18:14If you notice the output, it shows that
- 1:18:16the first four columns have more than
- 1:18:1840% null values. Right? So, it's always
- 1:18:21best for us to remove features or such
- 1:18:23variables because they will not help us
- 1:18:25in our prediction.
- 1:18:26Now, during data pre-processing, it is
- 1:18:28always necessary to remove the variables
- 1:18:31that are not significant. Unnecessary
- 1:18:33data will just increase our
- 1:18:35computations. That's why it's always
- 1:18:37best if you remove the unwanted or
- 1:18:39unnecessary variables.
- 1:18:41Now, apart from removing these four
- 1:18:42variables, we'll also remove the
- 1:18:44location variable, and we will remove
- 1:18:47the date variable. Right, I'll come to
- 1:18:49this variable in a minute. We'll also be
- 1:18:51removing location and date variable
- 1:18:54because both of these variables are not
- 1:18:56needed in order to predict whether it'll
- 1:18:57rain tomorrow. Right, we do not need to
- 1:19:00know the location and the date.
- 1:19:02Now, we'll also be removing this a risk
- 1:19:04mm variable. Risk mm variable basically
- 1:19:07tells us the amount of rain that might
- 1:19:09occur the next day.
- 1:19:11Right, now this is a very informative
- 1:19:13variable, and it might actually leak
- 1:19:15some information to our model.
- 1:19:17By using this variable, we'll easily be
- 1:19:19able to predict, and there's no point of
- 1:19:21doing that. So, this variable will give
- 1:19:23us too much information. And so, that's
- 1:19:25why we're going to remove this variable
- 1:19:26as well. It'll leak a lot of
- 1:19:28information. So, after that, if you
- 1:19:30print the shape of your data frame,
- 1:19:33we have only 17 variables and so many
- 1:19:37observations.
- 1:19:38Now, after this, we'll just uh look at
- 1:19:40any null values, and we'll remove them.
- 1:19:43This drop.any function will just remove
- 1:19:45all the null values. Right, then if you
- 1:19:47print the shape of your data frame,
- 1:19:49we'll have around 112,000 rows with 17
- 1:19:53variables. This is the shape of the data
- 1:19:55set after removing all the null values
- 1:19:57and all the redundant or unnecessary
- 1:20:00variables.
- 1:20:01Now, it's time to remove the outliers in
- 1:20:03the data. So, after you remove any null
- 1:20:05values, we should also check our data
- 1:20:07set for any outliers. An outlier is a
- 1:20:11data point that is very different from
- 1:20:13your other observations.
- 1:20:15Outliers usually occur because of
- 1:20:17miscalculations while collecting the
- 1:20:19data.
- 1:20:20These are some sort of errors in your
- 1:20:22data set.
- 1:20:23So, in this whole code snippet, we're
- 1:20:25just getting rid of outliers.
- 1:20:29This is the output that we get. All our
- 1:20:31outliers.
- 1:20:32Next, what we'll be doing is we will be
- 1:20:35assigning zeros and ones in the place of
- 1:20:38yes and no.
- 1:20:39The only thing is we're going to change
- 1:20:40the categorical variables from yes and
- 1:20:42no to zero and one. Right, that's
- 1:20:44exactly what we're doing over here.
- 1:20:46Now, if there are any unique values such
- 1:20:48as any character values which are not
- 1:20:50supposed to be there, we'll be changing
- 1:20:52them into integer values.
- 1:20:54That's all we're doing over here.
- 1:20:56After this, we'll be normalizing our
- 1:20:58data set.
- 1:20:59This is a very important step because in
- 1:21:02order to avoid any biasness in your
- 1:21:04output, you have to normalize your input
- 1:21:06variables.
- 1:21:08Right, to do this, we can make use of
- 1:21:09the min-max scalar function which Python
- 1:21:11provides in a package known as
- 1:21:13scikit-learn. You can use that package
- 1:21:16in order to normalize your data set.
- 1:21:18So, after normalizing our data set, this
- 1:21:20is what our data set looks like.
- 1:21:23This is before normalization. You can
- 1:21:25see that these are in two digits,
- 1:21:27whereas these values are in single
- 1:21:29digits. Right? This causes a lot of
- 1:21:30biasness. But once we normalize the
- 1:21:33values, we know that all of the values
- 1:21:35are in a similar range. We have
- 1:21:37everything in decimals. Right, so
- 1:21:39normalization is something that has to
- 1:21:41be performed because if you have a data
- 1:21:43set like this, your output is not going
- 1:21:44to be correct. And that's why we perform
- 1:21:46normalization.
- 1:21:48So, now that we are done with
- 1:21:50pre-processing, what we're going to do
- 1:21:51is it's time for exploratory data
- 1:21:54analysis.
- 1:21:55Let me just comment it for you. This is
- 1:21:58exploratory data analysis.
- 1:22:01So, basically here, what we're going to
- 1:22:02do is we're going to analyze and
- 1:22:04identify the significant variables that
- 1:22:07will help us predict the outcome. To do
- 1:22:09this, we'll be using the select key best
- 1:22:12function which is present in the
- 1:22:13scikit-learn library.
- 1:22:15There's a predefined function in Python
- 1:22:17called select key best, which will
- 1:22:19basically select the most significant
- 1:22:21predictor variables in our data set.
- 1:22:24When we run that line of code,
- 1:22:26we get these three variables to be the
- 1:22:29most significant variables in our data
- 1:22:31set. Right, the main aim of this demo is
- 1:22:33to make you understand how machine
- 1:22:35learning works. That's why to simplify
- 1:22:37the competition, we'll assign only one
- 1:22:39of these significant variables as the
- 1:22:42input. Instead of taking all three
- 1:22:44variables as input, we'll select one
- 1:22:46variable and we'll take that as the
- 1:22:48input and the output is the rain
- 1:22:50tomorrow variable.
- 1:22:52So, basically, we are creating a data
- 1:22:54frame of all the significant variables.
- 1:22:56Basically, we're choosing this variable
- 1:22:58in order to predict our outcome.
- 1:23:00Obviously, our outcome is rain tomorrow
- 1:23:02variable.
- 1:23:03So, our input is humidity level and our
- 1:23:06output is to detect whether it'll rain
- 1:23:08tomorrow.
- 1:23:09The next step is data modeling. All of
- 1:23:11you are aware of what data modeling is.
- 1:23:13To solve this, we'll be using
- 1:23:15classification algorithms over here.
- 1:23:17We'll use logistic regression. We will
- 1:23:20use random forest classifier, which is
- 1:23:23another machine learning algorithm.
- 1:23:25We'll also use the decision tree
- 1:23:26classifier and support vector machine.
- 1:23:30Right, we'll be using all of these
- 1:23:31algorithms in order to predict the
- 1:23:33outcome. We'll also check which
- 1:23:35algorithm gives us the best accuracy.
- 1:23:38So guys, we're just using multiple
- 1:23:40algorithms or multiple classification
- 1:23:42algorithms on the same data set. We're
- 1:23:44not doing anything very complex over
- 1:23:46here.
- 1:23:47So, we start by importing all the
- 1:23:48necessary libraries for the logistic
- 1:23:50regression algorithm.
- 1:23:52We're also going to import time because
- 1:23:54we'll be calculating the accuracy and
- 1:23:56the time taken by the algorithm to get
- 1:23:58the output.
- 1:23:59So, the first step is data splicing.
- 1:24:01I've already mentioned data splicing is
- 1:24:03splitting your data set into your
- 1:24:05testing data set and into your training
- 1:24:07data set. That's exactly what we're
- 1:24:09doing over here. So, 25% of your data is
- 1:24:12assigned for the testing data and the
- 1:24:14remaining 75% is your training data.
- 1:24:17Here, you're creating the instance of
- 1:24:19the logistic regression algorithm. This
- 1:24:21is an instance that you created. Then
- 1:24:23you'll fit the model by using your
- 1:24:25training data set. So basically, to
- 1:24:27build your machine learning algorithm,
- 1:24:29you'll be fitting your training data
- 1:24:30set. So X_train and Y_train variables
- 1:24:33have your training data set.
- 1:24:35After that, you will be evaluating the
- 1:24:37model by using your testing data set.
- 1:24:40Then you'll calculate the accuracy
- 1:24:42score. Right? I'll also be printing the
- 1:24:44accuracy using logistic regression and
- 1:24:47the time taken using logistic
- 1:24:48regression. Let's look at the accuracy.
- 1:24:51Don't worry about these warnings. They
- 1:24:53are not important. So accuracy using
- 1:24:55logistic regression is around 0.83%,
- 1:24:59which is 83% accuracy, approximately
- 1:25:0184%. And this is the time taken. So the
- 1:25:05accuracy is actually pretty good, right?
- 1:25:0684% is a good number. Then we have
- 1:25:09random forest classifier. Here again,
- 1:25:11we'll import the libraries that are
- 1:25:13needed to run random forest classifier.
- 1:25:16Then we're again calculating the
- 1:25:17accuracy and the time taken by the
- 1:25:19classifier.
- 1:25:20Data splicing, like I mentioned,
- 1:25:22splitting the data into testing and
- 1:25:24training data set. Then you're just
- 1:25:26building the model by using the training
- 1:25:28data set. After that, you'll evaluate
- 1:25:31the model by using the testing data set
- 1:25:33and you'll finally calculate the
- 1:25:35accuracy. The accuracy using random
- 1:25:37forest is again approximately 84%, which
- 1:25:40is a really good number. Then we have
- 1:25:42decision tree classifier. Here again,
- 1:25:45we'll be importing the libraries needed
- 1:25:47for this classifier. We'll be
- 1:25:49calculating the accuracy and the time
- 1:25:51taken by this classifier. Data splicing
- 1:25:54followed by building the model by using
- 1:25:56the training data set, evaluating the
- 1:25:58model by using the testing data set, and
- 1:26:00finally calculating the accuracy and
- 1:26:02printing the accuracy.
- 1:26:04So let's see the accuracy using decision
- 1:26:06tree classifier. Again, we have an
- 1:26:09accuracy of around 83 to 84%.
- 1:26:12This is a pretty good number. And last,
- 1:26:14we're going to do this by using another
- 1:26:16classification algorithm known as
- 1:26:18support vector machine.
- 1:26:20Here again, we're importing the needed
- 1:26:22libraries. Then we're calculating the
- 1:26:24accuracy and the time, performing data
- 1:26:26splicing.
- 1:26:27Then we're building the model by using
- 1:26:29the training data set, testing the model
- 1:26:31using the testing data set, and finally
- 1:26:33printing the accuracy.
- 1:26:35So guys, all the classification models
- 1:26:38gave us an accuracy score of
- 1:26:39approximately 84% to 83%.
- 1:26:43So this is exactly how a machine
- 1:26:44learning process works. Right? You begin
- 1:26:47by importing all your data, then you
- 1:26:49perform data pre-processing or data
- 1:26:51cleaning. After that, you perform
- 1:26:53exploratory data analysis, where you
- 1:26:55understand the important patterns or the
- 1:26:58important variables in your data set.
- 1:27:00After that, you build a model, then you
- 1:27:03will evaluate the model by using the
- 1:27:05testing data set, and finally calculate
- 1:27:07the accuracy. I showed you all the steps
- 1:27:09in the machine learning process by using
- 1:27:11a practical demonstration in Python.
- 1:27:14So guys, give yourself a pat on the back
- 1:27:16because we just understood the whole
- 1:27:17machine learning process with a small
- 1:27:19implementation in Python.
- 1:27:22Now let's move on to our next topic,
- 1:27:24which is limitations of machine
- 1:27:26learning.
- 1:27:27Before we understand what deep learning
- 1:27:29is, it's important to know the
- 1:27:30limitations of machine learning, and why
- 1:27:33these limitations gave rise to the
- 1:27:35concept of deep learning. One major
- 1:27:37problem in machine learning is machine
- 1:27:40learning algorithms and models are not
- 1:27:42capable of handling high-dimensional
- 1:27:45data.
- 1:27:46Right? We can take in data with 20 to 30
- 1:27:49feature variables, but when it comes to
- 1:27:51data sets which have thousands of
- 1:27:53variables, machine learning does not
- 1:27:55work. Machine learning is not capable
- 1:27:58enough to process that much data.
- 1:28:01So high-dimensional data cannot be
- 1:28:03analyzed, processed, and modeled by
- 1:28:05using machine learning.
- 1:28:07Another limitation is that it cannot be
- 1:28:09used in image recognition and object
- 1:28:12detection because these applications
- 1:28:14require the implementation of
- 1:28:16high-dimensional data. Another major
- 1:28:19challenge in machine learning is to tell
- 1:28:21the machine what are the important
- 1:28:23features it should look for in order to
- 1:28:25precisely predict the outcome. So,
- 1:28:28basically you're selecting the important
- 1:28:29features for the machine learning model
- 1:28:31and you're telling them like these are
- 1:28:32the important features and this is what
- 1:28:35you should use in order to build the
- 1:28:36model.
- 1:28:37This process is known as feature
- 1:28:39extraction.
- 1:28:40Now, in machine learning this is a
- 1:28:41manual process. You're going to manually
- 1:28:43input as a programmer, you're going to
- 1:28:45tell that these are the important
- 1:28:47predictor variables. But, what happens
- 1:28:49when your data set has hundreds of
- 1:28:51variables?
- 1:28:52How are you going to sit and choose
- 1:28:54every variable and perform analysis on
- 1:28:56each variable to understand which is a
- 1:28:58really significant variable? That's
- 1:29:01going to become a very tedious task,
- 1:29:03right? It's not possible for you to
- 1:29:04manually sit down with 100 variables,
- 1:29:06check the correlation with each variable
- 1:29:08and understand which variable is
- 1:29:09significant in predicting the output.
- 1:29:12So, performing feature extraction
- 1:29:14manually is very tedious and that is one
- 1:29:16of the major limitations of machine
- 1:29:17learning. Now, deep learning comes to
- 1:29:20the rescue to all of these problems.
- 1:29:22So, let's understand what deep learning
- 1:29:25is and why we have deep learning in the
- 1:29:27first place.
- 1:29:28So, deep learning is actually one of the
- 1:29:30only methods by which we can overcome
- 1:29:33the challenge of feature extraction.
- 1:29:35This is because deep learning models are
- 1:29:37capable of learning to focus on the
- 1:29:39right features by themselves requiring
- 1:29:42minimal human intervention. Meaning that
- 1:29:45feature extraction will be performed by
- 1:29:47the deep learning model itself. You
- 1:29:49don't have to manually tell that this
- 1:29:50feature is important, that feature is
- 1:29:52important, choose this feature for
- 1:29:54predicting the output. All of this is
- 1:29:56not needed in deep learning. The model
- 1:29:58itself will learn which features are
- 1:30:00most significant in predicting the
- 1:30:02output.
- 1:30:03Also, deep learning is mainly used to
- 1:30:05deal with high-dimensional data, right?
- 1:30:08It is based on the concept of neural
- 1:30:10networks and is often used in object
- 1:30:13detection and image processing. This is
- 1:30:15exactly why we need deep learning. It
- 1:30:17solves the problem of processing
- 1:30:19high-dimensional data and manual feature
- 1:30:22extraction.
- 1:30:23Now, how exactly does deep learning
- 1:30:25work? Now, deep learning mimics the
- 1:30:27basic component of the human brain
- 1:30:29called the brain cell. The brain cell is
- 1:30:32also known as a neuron.
- 1:30:34So, inspired from a neuron, an
- 1:30:37artificial neuron was developed. Deep
- 1:30:39learning is based on the functionality
- 1:30:41of a biological neuron. So, let's
- 1:30:44understand how we mimic this
- 1:30:46functionality in an artificial neuron.
- 1:30:49Now guys, an artificial neuron is also
- 1:30:51known as a perceptron.
- 1:30:53Let's understand what this biological
- 1:30:55neuron does and how deep learning is
- 1:30:57based on this concept.
- 1:30:59In a biological neuron, you can see
- 1:31:01these dendrites, right? In this image,
- 1:31:03you see something known as dendrites.
- 1:31:06These dendrites are used to receive any
- 1:31:08input. These inputs are summed in the
- 1:31:11cell body and through the axon, it is
- 1:31:13passed on to the next neuron.
- 1:31:16So, similar to the biological neuron, a
- 1:31:18perceptron or a artificial neuron
- 1:31:20receives multiple inputs, applies
- 1:31:23various transformations and functions,
- 1:31:25and provides an output.
- 1:31:27Right? So, that's how artificial neural
- 1:31:29networks or that's how deep learning
- 1:31:31works.
- 1:31:32Now guys, the human brain consists of
- 1:31:34multiple connected neurons called a
- 1:31:36neural network. Similarly, by combining
- 1:31:39multiple perceptrons, we've developed
- 1:31:41what is known as deep neural networks.
- 1:31:44The main idea behind deep learning is
- 1:31:46neural networks and that's what we're
- 1:31:48going to learn about. So now, let's
- 1:31:50understand what exactly deep learning
- 1:31:52is.
- 1:31:53Deep learning is a collection of
- 1:31:55statistical machine learning techniques
- 1:31:57used to learn feature hierarchies based
- 1:32:00on the concept of artificial neural
- 1:32:02networks. So, the main idea behind deep
- 1:32:05learning is to use the concept of neural
- 1:32:08networks.
- 1:32:09A deep neural network will have three
- 1:32:11layers. Okay, there's something known as
- 1:32:13the input layer followed by the hidden
- 1:32:15layers and then we have the output
- 1:32:17layer. The input layer is basically the
- 1:32:19first layer and it receives all the
- 1:32:21inputs. So, all the inputs are fed into
- 1:32:24this input layer.
- 1:32:25The last layer is obviously the output
- 1:32:27layer. This layer will provide your
- 1:32:30desired output. Now, all the layers
- 1:32:32between the input and your output layer
- 1:32:34are known as the hidden layers.
- 1:32:36Now, the number of hidden layers in a
- 1:32:39deep learning network will depend on the
- 1:32:41type of problem you're trying to solve
- 1:32:42and the data that you have.
- 1:32:44We'll get into depth of what exactly a
- 1:32:46hidden layer does, but for now this is
- 1:32:48how a neural network is structured in
- 1:32:51deep learning. So, guys uh deep learning
- 1:32:53is used in highly computational use
- 1:32:56cases such as face verification,
- 1:32:58self-driving cars, and so on. Right? So,
- 1:33:00let's understand the importance of deep
- 1:33:02learning by looking at a real-world use
- 1:33:05case.
- 1:33:06So, I'm sure all of you have heard of
- 1:33:07the company PayPal. Now, PayPal makes
- 1:33:10use of deep learning to identify any
- 1:33:13possible fraudulent activities.
- 1:33:16So, the company makes use of deep
- 1:33:18learning for fraud detection. Now,
- 1:33:20PayPal recently processed over 235
- 1:33:24billion dollars in payments from 4
- 1:33:27billion transactions by its more than
- 1:33:30170 million customers. So, basically it
- 1:33:33processed this much data by using deep
- 1:33:35learning. PayPal uses machine learning
- 1:33:38and deep learning algorithms to mine
- 1:33:40data from the customers purchasing
- 1:33:42history in addition to reviewing
- 1:33:44patterns of any sort of fraud stored in
- 1:33:47the database and it will do this to
- 1:33:49predict whether a particular transaction
- 1:33:51is fraudulent or not. Now, the company
- 1:33:53has been relying on deep learning and
- 1:33:55machine learning technology for around
- 1:33:5710 years.
- 1:33:59Initially, the fraud monitoring team
- 1:34:01used simple linear models, right? They
- 1:34:03used machine learning, but over the
- 1:34:05years the company switched to more
- 1:34:07advanced machine learning technology
- 1:34:09called deep learning. This shows how
- 1:34:11deep learning is used in more advanced
- 1:34:14and more complicated use cases.
- 1:34:17The fraud risk manager and the data
- 1:34:19scientist at PayPal, he quoted that what
- 1:34:23we enjoy from more modern advanced
- 1:34:25machine learning is its ability to
- 1:34:27consume a lot more data, handle layers
- 1:34:30and layers of abstraction, and be able
- 1:34:32to see things that a simpler technology
- 1:34:35would not be able to see. Even human
- 1:34:37beings might not able to see. This is
- 1:34:40exactly what he quoted. He said that a
- 1:34:42simple linear model is capable of
- 1:34:44consuming around 20 variables, but with
- 1:34:48deep learning technology, you can run
- 1:34:50thousands of data points.
- 1:34:52He also quoted that there is a magnitude
- 1:34:55of difference. You'll be able to analyze
- 1:34:57a lot more information and identify
- 1:35:00patterns that are a lot more
- 1:35:02sophisticated.
- 1:35:03So, by implementing deep learning
- 1:35:05technology, PayPal can finally analyze
- 1:35:07millions of transactions to identify any
- 1:35:10fraudulent activity.
- 1:35:12This is how PayPal makes use of deep
- 1:35:14learning.
- 1:35:15Not only PayPal, we also have Facebook,
- 1:35:17right? Facebook makes use of deep
- 1:35:19learning technology for face
- 1:35:21verification.
- 1:35:22You've all seen the tagging feature at
- 1:35:24Facebook where we tag our friends in
- 1:35:26photos. All of that is based on deep
- 1:35:29learning and machine learning.
- 1:35:31So guys, that was a real-world use case
- 1:35:33to make you understand how important
- 1:35:35deep learning is.
- 1:35:36Now, let's move on and look at what
- 1:35:39exactly a perceptron is, right? We'll be
- 1:35:41going in depth about deep learning.
- 1:35:44A perceptron is basically a single-layer
- 1:35:47neural network that is used to classify
- 1:35:49linear data.
- 1:35:51It is the most basic component of a
- 1:35:53neural network. Now, a perceptron has
- 1:35:55four important components. It has
- 1:35:58something known as inputs, weights, and
- 1:36:00bias, summation functions, activation
- 1:36:03and transformation functions.
- 1:36:05These are four important parts of a
- 1:36:07perceptron.
- 1:36:09Now, before I discuss this diagram with
- 1:36:11you, let me tell you the basic logic
- 1:36:13behind a perceptron.
- 1:36:15There is something known as inputs,
- 1:36:17right? The input X here, you can see X1,
- 1:36:19X2 till Xn.
- 1:36:21So, let me explain the structure of a
- 1:36:22perceptron. What you're going to do is
- 1:36:24you're going to input variables into the
- 1:36:26perceptron, right? This X1, X2 till Xn
- 1:36:29basically stands for input. W1, W2 till
- 1:36:33Wn stands for the weight assigned to
- 1:36:36each of these inputs.
- 1:36:38Right? There is a specific weight
- 1:36:40that'll be randomly initialized in the
- 1:36:42beginning for each of your input.
- 1:36:45Next, you have something known as the
- 1:36:46summation element. Here, what you do is
- 1:36:48you multiply the respective input with
- 1:36:51the respective weight, and you add all
- 1:36:54these products. Right? That is basically
- 1:36:57your summation function. After this is
- 1:36:59what is your transfer function, also
- 1:37:01known as activation function.
- 1:37:03Right? The activation function map your
- 1:37:05input to your desired output. So, your
- 1:37:08input will go through these processes.
- 1:37:11It'll go through summation and
- 1:37:12activation function in order to get to
- 1:37:14the output. So, guys, remember that the
- 1:37:17neural networks work the same way as a
- 1:37:19perceptron. So, if you want to
- 1:37:21understand how deep neural networks
- 1:37:23work, you need to understand what a
- 1:37:24perceptron does.
- 1:37:26A deep neural network is nothing but
- 1:37:27multiple perceptrons.
- 1:37:29So, let me tell you how the entire works
- 1:37:31once again.
- 1:37:32So, basically, all your inputs are
- 1:37:34multiplied with their respective
- 1:37:36weights. Now, you add all the multiplied
- 1:37:39values, and you call them as a weight
- 1:37:41sum. You use the summation function to
- 1:37:43add all of this. After that, you apply
- 1:37:46the weighted sum to the correct
- 1:37:48activation transfer function. Activation
- 1:37:50function is very similar to a function
- 1:37:53in our brain. The neurons become active
- 1:37:56in our brain after a certain potential
- 1:37:59is reached. That threshold is known as
- 1:38:01the activation potential.
- 1:38:03So, mathematically, there are a few
- 1:38:05functions which represent the activation
- 1:38:07function. Basically, the signum, the
- 1:38:09sigmoid, the tan h, all of these are
- 1:38:11activation functions. You can think of
- 1:38:13activation function as a function that
- 1:38:15maps the input to the respective output.
- 1:38:19Then, I spoke about something known as
- 1:38:20weights and biases. All right, now you
- 1:38:23must be wondering, why do we have to
- 1:38:25assign weights to each of our input?
- 1:38:28Weights basically show the strength of a
- 1:38:30particular input or how important a
- 1:38:33particular input is for predicting the
- 1:38:35output. In simple words, the weightage
- 1:38:38denotes the importance of an input. Bias
- 1:38:41is basically a value which allows you to
- 1:38:44shift the activation function curve in
- 1:38:47order to get a precise output.
- 1:38:50All right, so that's exactly what
- 1:38:51weights are.
- 1:38:52I hope all of you are clear with inputs,
- 1:38:54weights, summation, and activation
- 1:38:57function. Also, one important thing I
- 1:38:59forgot to mention in a perceptron is a
- 1:39:01single layer perceptron will have no
- 1:39:04hidden layers. All right, there'll only
- 1:39:05be an input layer, an output layer, and
- 1:39:07a couple of transformation function in
- 1:39:09between. That's all will be there in a
- 1:39:11perceptron. Now, perceptron, like I
- 1:39:13mentioned, is used to solve only linear
- 1:39:15problems. If you look at this data
- 1:39:17distribution, how do you think we can
- 1:39:19solve this? This data is not linearly
- 1:39:22separable. So, you cannot use a single
- 1:39:24layer perceptron to separate this data.
- 1:39:27All right, that's why we need something
- 1:39:28known as a multi-layer perceptron with
- 1:39:31backpropagation.
- 1:39:32I'll be explaining this in the next
- 1:39:34slide. So, complex problems that involve
- 1:39:38a lot of parameters and high-dimensional
- 1:39:41data can be solved by using multiple
- 1:39:43layer perceptron. Now, a multi-layer
- 1:39:46perceptron is the same as a single layer
- 1:39:48perceptron. The only difference is that
- 1:39:50a multi-layer perceptron will have
- 1:39:52hidden layers.
- 1:39:53So, the number of hidden layers in a
- 1:39:56model depends upon various factors. I
- 1:39:58told you it depends on the complexity of
- 1:40:00the problem you're trying to solve. It
- 1:40:01depends on the number of inputs in your
- 1:40:03data and so on. So, it works in the same
- 1:40:05way. All your inputs are multiplied with
- 1:40:08your weights and then you do the
- 1:40:09summation and then there is a
- 1:40:11transformation function or a activation
- 1:40:14function.
- 1:40:15While designing a neural network, in the
- 1:40:17beginning itself I told you we
- 1:40:19initialize weights with some random
- 1:40:21values. We do not have some specific
- 1:40:23[clears throat] value for each
- 1:40:24weightage. Initially, we've selected
- 1:40:27random values.
- 1:40:28It is always important that whatever
- 1:40:31weight values we have selected will be
- 1:40:33correct.
- 1:40:34Now, whatever weight values we've
- 1:40:36assigned to each input, it denotes the
- 1:40:38importance of that input variable. So,
- 1:40:41we need to assign the weights in such a
- 1:40:43way or we need to update the weights in
- 1:40:45such a way that it denotes the
- 1:40:47significance of that particular input.
- 1:40:50So, initially we're selecting some
- 1:40:52random value for weight. And let's say
- 1:40:54that we use this weight value to get our
- 1:40:57output.
- 1:40:58Now, what happens is the output is
- 1:41:00actually very different or it is not
- 1:41:02precise when compared to our actual
- 1:41:04output. Basically, the error value is
- 1:41:07very huge.
- 1:41:08So, how will you reduce the error?
- 1:41:10The main thing in a neural network is
- 1:41:13the weightage that you give to a input
- 1:41:15variable, right? Depending on the
- 1:41:16weightage that you give to a input
- 1:41:18variable, you're telling the neural
- 1:41:19network how important that variable is.
- 1:41:22Now, what if you randomly give some
- 1:41:24weightage and your output is wrong? The
- 1:41:28first thing that comes into your mind is
- 1:41:29that you need to change the weight
- 1:41:31because the weight signifies the
- 1:41:33importance of a variable.
- 1:41:35So, basically what we need to do is we
- 1:41:37need to somehow explain to the model to
- 1:41:39change the weight in such a way that the
- 1:41:41error becomes minimum. Let's put it in
- 1:41:44another way. So, basically, we need to
- 1:41:46train a model. One way to train a model
- 1:41:48is called as backpropagation.
- 1:41:51So, in backpropagation, what happens is
- 1:41:54once you've initialized a weight to each
- 1:41:56of the input, you calculate the output.
- 1:41:59Right? You get an output, and let's say
- 1:42:01you have a very high error value in that
- 1:42:03output. What you do is you'll
- 1:42:05backpropagate as in you'll go back to
- 1:42:08the weight, and you'll keep updating the
- 1:42:10weight in such a way that your error
- 1:42:12becomes minimum.
- 1:42:14This is exactly what backpropagation is.
- 1:42:16You'll be going back to the first layer.
- 1:42:18You'll be updating each of the weights
- 1:42:20in such a way that your output is more
- 1:42:23precise. So, guys, basically, the weight
- 1:42:26and the error in a neural network is
- 1:42:29highly related. By updating the weight
- 1:42:32in a particular way, your error will
- 1:42:33decrease. So, you need to figure out how
- 1:42:36you need to update the weight. Do you
- 1:42:38have to increase the weight or decrease
- 1:42:40the weight? Once you figure out whether
- 1:42:41you have to increase or decrease the
- 1:42:43weight, you have to just follow that
- 1:42:45direction in such a way that your error
- 1:42:47is minimized. And that's exactly what
- 1:42:50backpropagation is. So, the final output
- 1:42:53of backpropagation is you're going to
- 1:42:55select the weight that minimizes the
- 1:42:57error function. And then you're going to
- 1:42:59use that weight to solve the whole
- 1:43:01problem.
- 1:43:02Right? This is what backpropagation is
- 1:43:04about.
- 1:43:05Now, in order to make you understand
- 1:43:07deep neural networks, let's look at a
- 1:43:09practical implementation.
- 1:43:12So, again, guys, I'll be using Python to
- 1:43:14run the demo. If you don't have a good
- 1:43:16idea about Python, check the
- 1:43:17description. I'll leave a couple of
- 1:43:19links about Python programming. Now, in
- 1:43:21this demo, I'll be walking you through
- 1:43:23one of the most important applications
- 1:43:25of deep learning. I will demonstrate how
- 1:43:28you can construct a high-performance
- 1:43:30model to detect credit card fraud.
- 1:43:32Right? We'll be using deep learning
- 1:43:34models to do this. Now, before that, let
- 1:43:36me just tell you something about our
- 1:43:37data set. Right? The data set contains
- 1:43:40transactions made by credit cards in the
- 1:43:43year September 2013 by European card
- 1:43:46holders. This data set presents
- 1:43:48transactions that occurred in 2 days,
- 1:43:51where we have 492 frauds out of 285,000
- 1:43:56transactions. Approximately 285,000
- 1:44:00transactions. Out of these transactions,
- 1:44:02492 were frauds, and the data set is
- 1:44:05quite unbalanced. Right? The positive
- 1:44:07class accounts for 0.172%.
- 1:44:11So, the positive class, basically the
- 1:44:13fraudulent class. So, again, we're going
- 1:44:15to start by importing the required
- 1:44:17packages. We're going to import Keras,
- 1:44:19Matplotlib library, Seaborn library, and
- 1:44:22scikit-learn for preprocessing. Right?
- 1:44:24Again, min-max scalar, which is for
- 1:44:26normalization. We're going to import our
- 1:44:29data set and store it in this variable.
- 1:44:31Right? This is the path to my data set.
- 1:44:34My data set is in the CSV format, or
- 1:44:36also known as comma-separated version.
- 1:44:39Now, we're going to print out the first
- 1:44:41five rows of our data set. Right? Let's
- 1:44:44take a look at the output.
- 1:44:47So, here is the time of the transaction.
- 1:44:49V1, V2, V3, etc. These are all the
- 1:44:52features of our data set. I'm not going
- 1:44:55to go into depth of what these features
- 1:44:57stand for, because this demo is all
- 1:44:59about understanding deep learning. Now,
- 1:45:01these V1, V2, V3, these are all
- 1:45:03predictor variables, which will help us
- 1:45:05predict our class.
- 1:45:07So, guys, don't worry about what these
- 1:45:09features are. These features are just
- 1:45:11information and details about your
- 1:45:13transaction, such as the amount you
- 1:45:15spend, or the time of transaction, and
- 1:45:18so on.
- 1:45:19So, here we have the amount variable,
- 1:45:20which denotes the amount spent. After
- 1:45:23that, we have the class variable. Now,
- 1:45:25this class variable is your output
- 1:45:27variable or your target variable. So,
- 1:45:30your class is basically your output
- 1:45:32variable. Value zero denotes that there
- 1:45:34has been no fraudulent activity, but if
- 1:45:36you get a class of one, it means that
- 1:45:39this transaction is a fraudulent
- 1:45:41transaction. For example, this
- 1:45:43transaction is not fraudulent, and
- 1:45:45that's why we have a value of zero over
- 1:45:47here. All right, so this is our data
- 1:45:49set.
- 1:45:51Next what we're doing is we're counting
- 1:45:53the number of samples for each class.
- 1:45:55Right, we have class zero and class one,
- 1:45:57where in class zero denotes the normal
- 1:45:59transaction, which is non-fraudulent
- 1:46:02transaction, and class one will denote
- 1:46:04the fraudulent transactions. Right, so
- 1:46:06we have around 492 fraudulent
- 1:46:08transactions and around 284,315
- 1:46:13non-fraudulent transactions.
- 1:46:15So, when you see this, you know that our
- 1:46:17data set is highly unbalanced. Highly
- 1:46:19unbalanced means that one class has a
- 1:46:22really small number when compared to the
- 1:46:24other class. Right, there's no balance
- 1:46:25between the two classes.
- 1:46:27So, here what we're doing is we are
- 1:46:29starting the data set by class for
- 1:46:31stratified sampling. Stratified sampling
- 1:46:34is a statistical technique for sampling
- 1:46:37your data set. Now, this type of
- 1:46:39sampling is always good if you have an
- 1:46:41unbalanced data set.
- 1:46:43Next what we're going to do is we're
- 1:46:44going to perform data pre-processing.
- 1:46:47Data pre-processing in deep learning
- 1:46:49mainly has a method known as dropout
- 1:46:52method.
- 1:46:53Next what we're going to do is we're
- 1:46:54going to uh drop out the entire time
- 1:46:57column. We do not need the time of the
- 1:47:00transaction in order to understand if
- 1:47:02the transaction was fraudulent or not.
- 1:47:05Right, so that's why we're getting rid
- 1:47:06of unnecessary variables. Right, so
- 1:47:08we're dropping out that variable.
- 1:47:12So, after dropping out the time
- 1:47:14variable, we're going to assign the
- 1:47:16first 3,000 samples to our new data
- 1:47:18frame. Right, this DF sample will have
- 1:47:20our first 3,000 samples, and we're going
- 1:47:24to use those 3,000 samples.
- 1:47:26So, here we're just counting the number
- 1:47:27of class for each of these samples.
- 1:47:30After that, we're just counting the
- 1:47:31number of samples for each of the class.
- 1:47:33Like we're doing the same thing again
- 1:47:34and here we get class zero has 2,508
- 1:47:38samples and class one has 492 samples.
- 1:47:41Now, this makes the data set quite
- 1:47:43balanced, right? It's very balanced when
- 1:47:45compared to our old data set.
- 1:47:48Next, we'll just randomly shuffle our
- 1:47:49data set, right? In order to remove any
- 1:47:52sort of biasness in the data.
- 1:47:54After that, we'll split our data set
- 1:47:56into two parts. One is for training and
- 1:47:59your other data set is for testing,
- 1:48:01right? This is also known as data
- 1:48:03splicing.
- 1:48:05Then, we'll be splitting each data frame
- 1:48:06into feature and label, meaning that
- 1:48:09your input and your output.
- 1:48:12We'll be doing this for your training
- 1:48:13data and for your testing data, right?
- 1:48:15All you're doing is you're separating
- 1:48:17your input from your output.
- 1:48:19Next, we're looking at our training data
- 1:48:21set, right? We're printing the shape of
- 1:48:23our training data set. The training data
- 1:48:25set has around 2,400
- 1:48:27observations and 29 variables or 29
- 1:48:31features.
- 1:48:32Similarly, we'll be printing out the
- 1:48:34size of our test data frame, right?
- 1:48:36That's exactly what we're doing over
- 1:48:37here.
- 1:48:38After that, we'll perform normalization,
- 1:48:41right? For this, we'll be using the
- 1:48:42min-max scaler.
- 1:48:44So, in normalization, we'll basically be
- 1:48:46scaling all our predictor variables
- 1:48:48around the same range so that there is
- 1:48:50no biasness in our prediction.
- 1:48:53After this, we'll be plotting a function
- 1:48:55for each of the learning curves. For
- 1:48:57your training phase and for your testing
- 1:48:59phase, you'll be plotting a learning
- 1:49:01curve. Now, I'll show you the output of
- 1:49:03this in a couple of minutes. For now,
- 1:49:05let's move on to the main part, which is
- 1:49:07model creation, right? In this demo,
- 1:49:10we'll use three fully connected layers.
- 1:49:13We'll also use dropout technique. Now,
- 1:49:15dropout is a type of regularization
- 1:49:18technique that is used to avoid any sort
- 1:49:20of overfitting in a neural network. It
- 1:49:23is a technique where you select neurons
- 1:49:25and you drop them during the training
- 1:49:26phase.
- 1:49:27We'll be using the ReLU as the
- 1:49:29activation function, which is a type of
- 1:49:31activation function just like sigmoid
- 1:49:33and tanh.
- 1:49:35So, the type of model that we'll be
- 1:49:36using is the sequential model. Right,
- 1:49:39sequential is the easiest way to build a
- 1:49:41model in Keras. Right, we're using the
- 1:49:43Keras library over here. If you
- 1:49:45remember, I imported that in the
- 1:49:47beginning. Right, it allows you to build
- 1:49:49a model layer by layer. So, each layer
- 1:49:52has weights that correspond to the layer
- 1:49:54that follows it. After this, you'll use
- 1:49:57the add function to add the dense
- 1:50:00layers. Basically, your hidden layers
- 1:50:01you're going to add over here. So, in
- 1:50:03our model, we'll be adding two dense
- 1:50:05layers or hidden layers, you can say.
- 1:50:08So, here what we're doing is we're
- 1:50:10adding the first dense layer. Now guys,
- 1:50:12a dense layer is standard layer type
- 1:50:15that works for most cases. Right, in a
- 1:50:18dense layer, all the nodes in the
- 1:50:19previous layer connect to the nodes in
- 1:50:21the current layer.
- 1:50:23So guys, don't get too involved into
- 1:50:25what exactly is happening here. All I'm
- 1:50:27doing is I'm creating a sequential
- 1:50:29model, and what is happening is I'm just
- 1:50:32assigning the number of inputs for each
- 1:50:34of the dense layer or for each of the
- 1:50:36hidden layer. I'm also assigning dropout
- 1:50:39value. Dropout is basically to prevent
- 1:50:41overfitting. Overfitting might occur
- 1:50:43when your model memorizes the training
- 1:50:45data set. Overfitting basically reduces
- 1:50:48the accuracy of a model. That's why
- 1:50:50we're using the dropout method to
- 1:50:52prevent overfitting. So, in the first
- 1:50:54hidden layer, we have around 200 units.
- 1:50:56Right, we have the activation function
- 1:50:58ReLU. Then, we're adding the second
- 1:51:01dense layer with again 200 neurons and
- 1:51:03the ReLU activation function. Kernel
- 1:51:06initializer is uniform, meaning that
- 1:51:08it's just sequential and normal. Then,
- 1:51:10we're again adding a dropout layer of
- 1:51:120.5. The dropout value of a network has
- 1:51:16to be chosen very wisely. Okay, a value
- 1:51:19that is too low will result in a minimal
- 1:51:21effect and a value that is too high will
- 1:51:23result in under learning by the network.
- 1:51:26So, 0.5 is a standard dropout value.
- 1:51:30Now, this last layer is our output
- 1:51:32layer. In the output layer, we'll
- 1:51:33obviously have only one neuron. We'll
- 1:51:35have one neuron that will show us the
- 1:51:37output class, either zero or one. Zero
- 1:51:41will show us non-fraudulent transactions
- 1:51:43and one will denote fraudulent
- 1:51:45transaction. Right, that's why we have
- 1:51:46only one neuron over here. And the
- 1:51:49activation function here is sigmoid.
- 1:51:51Right, since the number of neurons is
- 1:51:52only one.
- 1:51:54After that, we're printing the model
- 1:51:55summary. Now, I'll show you the summary
- 1:51:58and everything. Before that, let us
- 1:51:59understand what exactly optimization
- 1:52:02functions are. We'll understand what
- 1:52:04this optimization function does. Now, an
- 1:52:06optimizer takes care of the necessary
- 1:52:09computations that are used to change the
- 1:52:12network's weights and bias. So,
- 1:52:14basically, your optimizers will take
- 1:52:16care of all your computations such as
- 1:52:19changing the weight or updating the
- 1:52:21weight. If you all remember, I spoke
- 1:52:22about backpropagation, right? Where
- 1:52:24you'll update the weight and all of
- 1:52:25that. That is done by using optimizers.
- 1:52:28Here, we're selecting an optimizer known
- 1:52:30as the Adam optimizer.
- 1:52:33So, Adam optimizer is one of the current
- 1:52:35default optimizers in deep learning.
- 1:52:38Right, it stands for adaptive moment
- 1:52:40estimation. We don't have to get into
- 1:52:42the depth of all of this. Right, all of
- 1:52:44these are predefined optimizers in our
- 1:52:46Keras package itself. After this, we're
- 1:52:48going to fit our model by using the
- 1:52:50training features. We're also setting
- 1:52:53200 epochs and also there's something
- 1:52:56known as epochs and batch size. Right,
- 1:52:58we're setting epochs as 200 and batch
- 1:53:00size as 500. I'll tell you what exactly
- 1:53:03this means. Now, batch sizes are
- 1:53:05basically used so that we don't overfit
- 1:53:08our model, right? We're going to
- 1:53:09basically split our data set into 500
- 1:53:12batches. So, our input will be going in
- 1:53:14the form of batches, right? And our
- 1:53:17batch size is 500 inputs per batch, and
- 1:53:20we'll be going through 200 epochs.
- 1:53:23Meaning that our training will iterate
- 1:53:25200 times. This is basically the number
- 1:53:27of times that training our model. All
- 1:53:29right, that's what epoch and batch size
- 1:53:31is. After that, we're just showing our
- 1:53:34training history. I mean, just printing
- 1:53:35the accuracy curve for our training
- 1:53:37phase. We're also going to print our
- 1:53:39loss curves for our training phase,
- 1:53:41basically the error curves. And then
- 1:53:43finally, we have the evaluation. Here,
- 1:53:45we'll be testing our model by using our
- 1:53:47testing data set.
- 1:53:49Then we're finally printing the accuracy
- 1:53:51on our testing data set. After that,
- 1:53:54we're just going to plot a heat map,
- 1:53:56which I'll be showing y'all. Let me just
- 1:53:58show you the output.
- 1:54:00So guys, in this entire line of code,
- 1:54:01all we're doing is we're printing an
- 1:54:03accuracy plot. All right, basically
- 1:54:05we're printing a heat map. I'll show you
- 1:54:07what the heat map looks like.
- 1:54:10This is just to check the accuracy.
- 1:54:12We're comparing all the correctly
- 1:54:14predicted values to our incorrectly
- 1:54:16predicted values. So, this is our
- 1:54:19training history. Here, blue stands for
- 1:54:21our training phase, and this is our
- 1:54:22validation or our prediction stage.
- 1:54:26That was our training curve, and this is
- 1:54:28our loss curve. Now, when you compare it
- 1:54:30to the actual validation stage, it's
- 1:54:32quite similar, right? Meaning that our
- 1:54:34model is doing pretty well. So guys,
- 1:54:37this is the heat map that I was talking
- 1:54:38about. This is basically going to give
- 1:54:40us the class for each of our
- 1:54:41predictions. All right, it basically
- 1:54:43plots the classes that we correctly
- 1:54:46predicted, right? Basically, for each
- 1:54:48data point, it's just going to tell us
- 1:54:49whether we predicted it correctly or
- 1:54:51not. It's sort of a confusion matrix in
- 1:54:53the form of a heat map. So guys, these
- 1:54:56are all our epochs, basically the 200
- 1:54:59iterations that we went through, right?
- 1:55:01This is the 50th iteration is showing us
- 1:55:03our loss, it's showing us our accuracy
- 1:55:05as well. And here we have 88%, 90%, 92%.
- 1:55:10Now, if you carefully look at the epoch
- 1:55:12accuracy values, you see that as we
- 1:55:15train our model even more, our accuracy
- 1:55:17keeps increasing. Initially, our
- 1:55:19accuracy was around 83, right? At epoch
- 1:55:22number 15, our accuracy was around 83%.
- 1:55:26But as we kept training our model a
- 1:55:29little bit more, our accuracy kept
- 1:55:31increasing. We have 90, we have 91, 94,
- 1:55:3595, 96, and so on. All right, so
- 1:55:38basically, the more you train your
- 1:55:39model, the better it's going to be. So
- 1:55:42guys, this was our entire demo. Now, in
- 1:55:44the end, I'm printing out the false
- 1:55:46positive rate and the false negative
- 1:55:48rate. All of this basically denotes how
- 1:55:50many of the data points was I correctly
- 1:55:53able to predict as fraudulent and how
- 1:55:55many did I predict wrongly. That's all
- 1:55:58the false negative and the false
- 1:56:00positive rate denotes. So guys, this was
- 1:56:02the entire demo on deep learning. Now,
- 1:56:06if you have any doubts regarding the
- 1:56:07deep learning demo, please mention them
- 1:56:09in the comment section and I will solve
- 1:56:11your queries. All right, now let's look
- 1:56:13at our last topic for the day, which is
- 1:56:15natural language processing. Now, before
- 1:56:18we understand what is natural language
- 1:56:20processing, let's understand the need
- 1:56:21for natural language processing and a
- 1:56:24process known as text mining. Text
- 1:56:26mining and natural language processing
- 1:56:28are heavily correlated. All right, I'll
- 1:56:29talk about both of these in the upcoming
- 1:56:32slides. For now, let me tell you why we
- 1:56:34need natural language processing or text
- 1:56:36mining. So guys, the amount of data that
- 1:56:38we're generating these days is
- 1:56:40unbelievable. It is a known fact that
- 1:56:42we're creating 2.5 quintillion bytes of
- 1:56:46data every day, and this number is only
- 1:56:48going to grow. With the evolution of
- 1:56:50communication through social media, we
- 1:56:52generate tons and tons of data. All
- 1:56:55right, the numbers are on your screen.
- 1:56:57So, basically, we post around 1.7
- 1:56:59million pictures on Instagram per
- 1:57:01minute. Right, I'm talking about post
- 1:57:04per minute. All of these numbers are per
- 1:57:06minute values. These are the amount of
- 1:57:08tweets, 347,000
- 1:57:10tweets per minute. Right, this is a lot
- 1:57:13of data. We're generating data while
- 1:57:15we're watching YouTube videos, when
- 1:57:17we're sending emails, when we are
- 1:57:19chatting, and all of that. Right, even
- 1:57:21the IoT devices at our house, right, we
- 1:57:24have Alexa all of This is generating a
- 1:57:25lot of data. A single click on your
- 1:57:27phone is generating a lot of data. Now,
- 1:57:30not only that, out of all the data that
- 1:57:32we generate, only 21% of the data is
- 1:57:35structured and well formatted. Right,
- 1:57:38the remaining of the data is
- 1:57:39unstructured. And the major sources of
- 1:57:42unstructured data include text messages
- 1:57:44from WhatsApp, Facebook likes, comments
- 1:57:46on Instagram, the bulk emails, and all
- 1:57:49of this. Right, all of this accounts for
- 1:57:51the unstructured data that we have
- 1:57:53today. Now, the data we generate is used
- 1:57:56to grow a business. So, by analyzing and
- 1:57:58mining the data, we can add more value
- 1:58:01to a business. This is exactly what
- 1:58:04natural language processing and text
- 1:58:06mining is all about. Text mining and NLP
- 1:58:09is a subset of artificial intelligence,
- 1:58:11wherein we try and understand the
- 1:58:14natural language text that we get from
- 1:58:16text messages and so on, in order to
- 1:58:19derive useful insights and grow
- 1:58:21businesses by using these insights. So,
- 1:58:23what exactly is text mining? Text mining
- 1:58:26is a process of deriving meaningful
- 1:58:28insights or information from natural
- 1:58:31language text. So, all the data that we
- 1:58:34generate through text messages, emails,
- 1:58:36and documents are written in natural
- 1:58:38language text. Right, and we're going to
- 1:58:40use text mining and natural language
- 1:58:42processing to draw useful insights or
- 1:58:44patterns from such data in order to grow
- 1:58:47a business.
- 1:58:48Now, let's understand where exactly do
- 1:58:50we make use of natural language
- 1:58:51processing and text mining? Now, have
- 1:58:53you ever noticed that if you start
- 1:58:55typing a word on Google, you immediately
- 1:58:58get suggestions, right? This feature is
- 1:59:00known as auto complete. It will
- 1:59:02basically suggest the rest of the word
- 1:59:04to you. We also have something known as
- 1:59:05spam detection, right? Here's an example
- 1:59:08of how Google recognizes this
- 1:59:10misspelling Netflix and shows results
- 1:59:13for the keyword that matches your
- 1:59:14misspelling. Let me show you a couple of
- 1:59:17more examples. We also have predictive
- 1:59:20typing and spell checkers and features
- 1:59:22like auto correct, email classification.
- 1:59:25So, predictive typing and spell
- 1:59:27checkers, all of these are applications
- 1:59:30of natural language processing. All of
- 1:59:32this basically involves processing the
- 1:59:34natural language that we use and
- 1:59:36deriving some useful information from
- 1:59:38it, right? Or running businesses from
- 1:59:40it. Netflix uses natural language
- 1:59:42processing in a really good fashioned
- 1:59:45way, right? It basically studies the
- 1:59:47reviews that customer gives for a
- 1:59:49particular movie and it tries to figure
- 1:59:51out if that movie is good or bad
- 1:59:53depending on the review. So, Netflix
- 1:59:55actually uses NLP in a very interesting
- 1:59:58manner. It tries to understand the type
- 2:00:00of movies that a person likes by the way
- 2:00:03a person has rated the movie or by the
- 2:00:05way the person has reviewed a movie. So,
- 2:00:08by understanding what type of review a
- 2:00:10person is giving to a movie, Netflix
- 2:00:12will recommend more movies that you
- 2:00:15like. That's how important NLP has
- 2:00:17become. Now, let's look at what exactly
- 2:00:19NLP is. NLP, which also stands for
- 2:00:22natural language processing, is a part
- 2:00:24of computer science and artificial
- 2:00:26intelligence, which deals with human
- 2:00:28language. Right? It's basically the
- 2:00:30process of processing natural language
- 2:00:32in order to derive some useful
- 2:00:34information from it. For those of you
- 2:00:36who have studied natural language
- 2:00:38processing or have heard of natural
- 2:00:40language processing, there is a huge
- 2:00:42confusion between text mining and
- 2:00:44natural language processing. So, text
- 2:00:46mining is the process of deriving high
- 2:00:48quality information from text. But, the
- 2:00:51overall goal is to turn the text into
- 2:00:53data for analysis by using natural
- 2:00:56language processing. So, basically text
- 2:00:58mining is implemented by using natural
- 2:01:01language processing techniques. Right?
- 2:01:03There are various techniques in natural
- 2:01:04language processing that can help us
- 2:01:06perform text mining. That's how text
- 2:01:08mining and natural language processing
- 2:01:10are related. Natural language processing
- 2:01:12is the techniques that are used to solve
- 2:01:15the problem of text mining, text
- 2:01:17analysis, and all of that. Let's look at
- 2:01:19a couple more applications. Sentimental
- 2:01:22analysis is one of the major
- 2:01:23applications of natural language
- 2:01:25processing. You see Twitter performs
- 2:01:27sentimental analysis, Facebook, Google,
- 2:01:30all of these perform sentimental
- 2:01:31analysis. Sentimental analysis mainly
- 2:01:34used to analyze social media content
- 2:01:36that can help us determine the public
- 2:01:38opinion on a certain topic. Then we have
- 2:01:41chatbots. Now, chatbots use natural
- 2:01:43language processing to convert human
- 2:01:45language into desirable actions. We also
- 2:01:48have machine translation. NLP is used in
- 2:01:51machine translation by studying the
- 2:01:52morphological analysis of each word and
- 2:01:55translating it to another language.
- 2:01:58Advertisement matching is also done
- 2:01:59using NLP in order to recommend ads
- 2:02:02based on your history. Right? These are
- 2:02:04few of the applications of NLP. Now, let
- 2:02:07me tell you the basic terminologies
- 2:02:08under natural language processing. So,
- 2:02:10tokenization is the most basic step in
- 2:02:13natural language processing.
- 2:02:14Tokenization means breaking down the
- 2:02:17data into smaller chunks or tokens so
- 2:02:20that they can be easily analyzed. So,
- 2:02:22the first step is you'll break a complex
- 2:02:24sentence into words, then you'll
- 2:02:26understand the importance of each of the
- 2:02:28word with respect to that sentence in
- 2:02:31order to produce a structural
- 2:02:33description on an input sentence. So,
- 2:02:35for example, take this sentence. How
- 2:02:37would I perform tokenizations on the
- 2:02:39sentence?
- 2:02:40Let's say that tokens are simple is a
- 2:02:43sentence and I want to perform
- 2:02:44tokenization on the sentence. This is
- 2:02:47what I'm going to do. I'm going to split
- 2:02:48the sentence into different words. I'm
- 2:02:50going to understand each word with
- 2:02:52respect to that sentence. Right? This is
- 2:02:55done to simplify operations in natural
- 2:02:57language processing. Right? It's always
- 2:02:59simpler to analyze a single token
- 2:03:02instead of analyzing an entire sentence.
- 2:03:04Then we have something known as
- 2:03:05stemming. Now look at this example.
- 2:03:08Right here we have words such as
- 2:03:10detection, detecting, detected, and
- 2:03:12detections. We all know that the root
- 2:03:15word for all of these words is detect.
- 2:03:17So stemming algorithm basically does
- 2:03:20that. It works by cutting off the end or
- 2:03:23the beginning of the word and taking
- 2:03:25into account a list of common prefixes
- 2:03:28and suffixes that can be found in an
- 2:03:30inflicted word. Stemming basically helps
- 2:03:33us in analyzing a lot of words. We know
- 2:03:36that detections, detected, and detection
- 2:03:38basically mean the same thing. So all
- 2:03:40we're doing is we're going to ease our
- 2:03:42analysis by removing prefixes and
- 2:03:44suffixes which not make sense. Right? We
- 2:03:47just need to understand the
- 2:03:48morphological analysis of the word.
- 2:03:50Right? So that's why we're randomly
- 2:03:51cutting the prefixes and suffixes in
- 2:03:53such a way that we only get the
- 2:03:55important part of the word. This is
- 2:03:57called stemming. Now this cutting of
- 2:03:59words can be successful in some
- 2:04:02occasions, but not always. That is why
- 2:04:04we say that stemming approach has a few
- 2:04:08limitations. In order to get over these
- 2:04:11limitations, we have a process known as
- 2:04:13lemmatization. Right? Lemmatization on
- 2:04:16the other hand takes into consideration
- 2:04:18the morphological analysis of the words.
- 2:04:21It does not randomly cut the word in the
- 2:04:23beginning and the ending. It understands
- 2:04:25what the word means and only then it
- 2:04:27cuts the word. For example, let's
- 2:04:29consider the word recap. If we perform
- 2:04:32stemming on the word recap, we'll get
- 2:04:35cap. Right? The output will be cap. But,
- 2:04:38cap and recap do not have the same
- 2:04:40meaning, do they? They have absolutely
- 2:04:41different meanings. That's why stemming
- 2:04:43is sometimes not considered to be the
- 2:04:45right thing to do. But, when it comes to
- 2:04:47lemmatization, it's going to understand
- 2:04:49the meaning of recap. Only then will it
- 2:04:52perform any sort of change in the word,
- 2:04:54or it'll cut down the word. So,
- 2:04:56basically, it groups together different
- 2:04:58inflected forms of a word called lemma.
- 2:05:01Lemmatization is similar to stemming
- 2:05:03because it maps several words into one
- 2:05:06common root. But, the output of a
- 2:05:09lemmatization process is always a proper
- 2:05:11word. An example of lemmatization is to
- 2:05:15map gone, going, and went into go. Gone,
- 2:05:18going, went, all of them mean go. So,
- 2:05:21basically, by lemmatization, you can
- 2:05:22just output the words as go. That is
- 2:05:25what lemmatization is. Next, we have
- 2:05:27something known as stop words, right?
- 2:05:29Stop words are basically a set of
- 2:05:31commonly used words in any language,
- 2:05:34right? Not just English, any language.
- 2:05:36The reason why stop words are critical
- 2:05:38to many applications is that if we
- 2:05:41remove the words that are very commonly
- 2:05:44used in a given language, we can finally
- 2:05:46focus on the important words. For
- 2:05:48example, in the context of Let's say you
- 2:05:51open up Google and you look for
- 2:05:53strawberry milkshake recipe. Instead of
- 2:05:55typing strawberry milkshake recipe,
- 2:05:57let's say you type how to make
- 2:05:59strawberry milkshake. Now, here, what
- 2:06:02Google will do is it'll find results for
- 2:06:04how, to, and make. Instead, if you just
- 2:06:07type strawberry milkshake recipe, you'll
- 2:06:10get the most desired output. That's why
- 2:06:13it's always considered a good practice
- 2:06:15in natural language processing to get
- 2:06:17rid of stop words, right? Stop words
- 2:06:19will just increase our computation, and
- 2:06:21it'll just add additional work to us.
- 2:06:23They are not very helpful when we're
- 2:06:25analyzing important documents, right? We
- 2:06:27need to focus on the important keywords
- 2:06:29in the documents instead of all of these
- 2:06:31commonly used words. Example of stop
- 2:06:34words include the, how, when, why, not,
- 2:06:38yes, no. All of these are stop words,
- 2:06:41right? So, in order to better analyze
- 2:06:43our data, we need to get rid of stop
- 2:06:45words. Now, the last terminology I'm
- 2:06:47going to discuss is document term
- 2:06:49matrix. It is important to create
- 2:06:51something known as the document term
- 2:06:53matrix in natural language processing. A
- 2:06:56DTM or a document term matrix is
- 2:06:59basically a matrix that shows the
- 2:07:01frequency of words in a particular
- 2:07:03document. Let's say that we're trying to
- 2:07:05understand if the sentence this is fun
- 2:07:09is available in one of my documents.
- 2:07:12So, if it is there in my document one,
- 2:07:14I'm going to put a one corresponding to
- 2:07:16each of the words that is available in
- 2:07:18my document. For example, in document
- 2:07:20two, I have this is, but I do not have
- 2:07:23the word fun.
- 2:07:24Similarly, in document four, I have the
- 2:07:27word this, but I do not have the word is
- 2:07:29and fun. So, basically, a document term
- 2:07:31matrix is like the frequency matrix of a
- 2:07:34document. So, during text analysis, you
- 2:07:37always begin by building a document term
- 2:07:39matrix, right? Here, you try to
- 2:07:40understand which words frequently occur
- 2:07:43and which words are important and not
- 2:07:45important in the document. So, guys,
- 2:07:47these were a couple of terminologies in
- 2:07:49natural language processing.
- 2:07:57Artificial intelligence and machine
- 2:07:59learning are not just trending
- 2:08:00technologies anymore.
- 2:08:02They're becoming the backbone of every
- 2:08:04industry. And right now, the demand for
- 2:08:06AI and ML engineers is exploding
- 2:08:09worldwide. In India, AI engineers earn
- 2:08:12anywhere from 8 lakhs to 36 lakhs per
- 2:08:15year, depending on skills and
- 2:08:17experience. In the United States, the
- 2:08:19same roles can start from $120,000 and
- 2:08:23can go all the way up to $250,000 for
- 2:08:26senior and specialized positions.
- 2:08:28So, if you have been thinking about
- 2:08:30getting into AI and ML, switching
- 2:08:32careers, or upskilling for higher-paying
- 2:08:35opportunities, there has never been a
- 2:08:37better time. And in this video, I am
- 2:08:39giving you a complete,
- 2:08:41beginner-friendly, and deeply practical
- 2:08:43AI and ML engineer roadmap that shows
- 2:08:46you exactly what to learn and how to
- 2:08:48grow in this booming field.
- 2:08:50And now, the first step of your AI
- 2:08:53journey starts with a strong
- 2:08:55foundations.
- 2:08:56And no, you don't need to be a
- 2:08:57mathematician. You simply need the
- 2:09:00essentials. So, start with Python,
- 2:09:02because Python is the language that
- 2:09:04powers almost every modern AI system.
- 2:09:07So, focus on the basics like variables,
- 2:09:10loops, functions, lists, dictionaries,
- 2:09:14file handling, and how to work with
- 2:09:16APIs.
- 2:09:17So, these skills are enough to write
- 2:09:18simple programs and understand AI code.
- 2:09:21Next, learn data handling, because AI is
- 2:09:25built on data. Use Pandas to clean and
- 2:09:27organize data, NumPy to perform
- 2:09:30calculation, and Matplotlib or Seaborn
- 2:09:33to visualize patterns. Even simple tasks
- 2:09:36like removing missing values or
- 2:09:38analyzing sales trends will prepare you
- 2:09:40for the real AI projects.
- 2:09:42Then, learn the essential math behind
- 2:09:44AI. Not heavy equations, just the basic
- 2:09:47understanding. Understand mean, median,
- 2:09:50variance, probability basic,
- 2:09:52correlations, and what vectors and
- 2:09:54matrices are. Learn what gradient
- 2:09:57descent means conceptually, so you
- 2:09:59understand how models learn without
- 2:10:01getting buried in complex math. So, once
- 2:10:03your foundations are ready, move to
- 2:10:05machine learning. ML is simply teaching
- 2:10:08computers to learn from examples.
- 2:10:10Instead of writing instructions, show
- 2:10:12the model real-world data and let it
- 2:10:14find patterns.
- 2:10:16So, learn key ML concepts like training
- 2:10:18and testing, accuracy and precision,
- 2:10:21underfitting and overfitting, cross
- 2:10:23validation, and feature engineering. So,
- 2:10:26these concepts helps you understand how
- 2:10:28to build, tune, and improve models.
- 2:10:31Then, learn the core email algorithms
- 2:10:33that companies use every single day. So,
- 2:10:36you need to start with linear
- 2:10:38regression, logistic regression,
- 2:10:40decision trees, random forest, SVM,
- 2:10:44naive Bayes, K-means clustering, and
- 2:10:46PCA.
- 2:10:47These algorithms cover most practical
- 2:10:49business problems like predicting sales,
- 2:10:52detecting fraud, segmenting customers,
- 2:10:55and identifying patterns in large data
- 2:10:57sets.
- 2:10:58Then, build small email projects such as
- 2:11:00house price predictor, spam email
- 2:11:03classifier, credit score predictor, or
- 2:11:06customer segmentation model. So, these
- 2:11:08projects give you confidence and make
- 2:11:10your portfolio job ready. After ML, move
- 2:11:13into deep learning, the technology
- 2:11:15behind ChatGPT, self-driving cars, and
- 2:11:18medical AI.
- 2:11:19Start by understanding how neural
- 2:11:21networks work. Learn what neurons,
- 2:11:24layers, activation functions, and loss
- 2:11:26functions are.
- 2:11:28You don't need to memorize formulas,
- 2:11:30just understand how the network adjusts
- 2:11:32itself to improve predictions.
- 2:11:34Pick either TensorFlow or PyTorch as
- 2:11:37your deep learning framework because
- 2:11:39both are used by companies in
- 2:11:40production, so you only need to choose
- 2:11:42one. And then, build deep learning
- 2:11:44projects like digit recognition, image
- 2:11:46classification, sentiment analysis, or
- 2:11:49go for fake news classification. So,
- 2:11:51these projects teach you how to use
- 2:11:53neural networks in real scenarios.
- 2:11:56So, AI is no longer about learning
- 2:11:59everything. It's about choosing your
- 2:12:01specialization.
- 2:12:02So, here we have options. So, option A
- 2:12:05is NLP and LLMs, which has the highest
- 2:12:08demand. So, if you want to work with
- 2:12:10chatbots, smart assistants, or language
- 2:12:13models like ChatGPT, choose NLP and
- 2:12:16LLMs. And all you need to learn is
- 2:12:19tokenization, embeddings, transformers,
- 2:12:22BERT, GPT models, prompt engineering,
- 2:12:25fine-tuning, RAG, and agentic
- 2:12:28architectures.
- 2:12:29And you can build projects like AI
- 2:12:31chatbots, document search tools,
- 2:12:34question and answer systems, or customer
- 2:12:36support bots.
- 2:12:37Option B is computer vision. If you like
- 2:12:40working with images and videos, choose
- 2:12:42computer vision. Learn CNNs, YOLO,
- 2:12:45object detection, and segmentation. And
- 2:12:48you can build real-world projects like
- 2:12:50face detection, medical image analysis,
- 2:12:53CCTV monitoring systems, or vehicle
- 2:12:55counting tools. Option C is generative
- 2:12:58AI. If you enjoy creativity, choose
- 2:13:01generative AI. Learn GANs, VAEs, and
- 2:13:05diffusion models. Build applications
- 2:13:07like AI art generators, product design
- 2:13:10tools, image-to-image systems, or video
- 2:13:12generation models. Option D is MLOps. If
- 2:13:16you prefer infrastructure and
- 2:13:18deployment, then choose MLOps. Learn
- 2:13:21Docker, Kubernetes, MLflow, CI/CD
- 2:13:23pipelines, cloud deployment, and model
- 2:13:26monitoring. And you can build projects
- 2:13:28that focus on deploying ML and LLM
- 2:13:31models into real environments. All
- 2:13:33right. So, agentic AI is the biggest
- 2:13:36trend of 2026.
- 2:13:38These are not just models, these are the
- 2:13:41intelligent agents that can reason,
- 2:13:43plan, use tools, and take actions. So,
- 2:13:46learn how agents work with frameworks
- 2:13:48like LangChain, LangGraph, and
- 2:13:51crew-based agent architectures.
- 2:13:53Understand tool calling, memory systems,
- 2:13:56planning, and multi-agent collaboration.
- 2:13:59Also, build agentic projects like an AI
- 2:14:02research assistant, an autonomous email
- 2:14:04automation agent, a financial analysis
- 2:14:07agent, or a customer service automation
- 2:14:10agent. So, these projects stand out in
- 2:14:13the interviews because companies want
- 2:14:15people who can build intelligent
- 2:14:17workflows and not just models.
- 2:14:19So, now that you have skills, you need
- 2:14:21projects that prove it. So, your
- 2:14:23portfolio should have two machine
- 2:14:25learning projects, two deep learning
- 2:14:27projects, two specialization projects,
- 2:14:29and one real end-to-end AI system. And
- 2:14:32this final project could be a chatbot
- 2:14:34with rack, a vision-based attendance
- 2:14:36system, an AI assistant with memory, or
- 2:14:40a complete ML pipeline deployed on
- 2:14:43cloud. And finally, upload your work on
- 2:14:46GitHub, write clear documentation, and
- 2:14:48add deployment links so employers can
- 2:14:51test your work instantly. So, with this
- 2:14:53roadmap, you can apply for the most
- 2:14:55in-demand roles in 2026,
- 2:14:58such as AI engineer, machine learning
- 2:15:00engineer, LLM engineer, NLP engineer,
- 2:15:04generative AI engineer, computer vision
- 2:15:06engineer, MLOps engineer, or AI
- 2:15:09automation specialist.
- 2:15:11AI is not just the future, it's the
- 2:15:14career shift of today. If you follow
- 2:15:16this roadmap step by step, you will
- 2:15:18build the skills, the projects, and the
- 2:15:20confidence to enter the world of AI and
- 2:15:23machine learning.
- 2:15:27>> [music]
- 2:15:30>> So guys, let's see what we are going to
- 2:15:32explore today. So today, we are going to
- 2:15:34explore real-world examples of machine
- 2:15:36learning, starting your journey, how you
- 2:15:38can start your journey to machine
- 2:15:40learning, and key concepts and impacts
- 2:15:42of machine learning in your day-to-day
- 2:15:44life. So guys, to help you navigate
- 2:15:47through this video, here is a quick
- 2:15:49rundown for you guys. Table of contents.
- 2:15:51What is machine learning? How does
- 2:15:53machine learning works? Five features of
- 2:15:55machine learning, types of machine
- 2:15:57learning, what skills one should have to
- 2:15:59learn machine learning, machine learning
- 2:16:01applications, and last but not the least
- 2:16:04guys, that is future of machine
- 2:16:06learning.
- 2:16:07So guys, I'm going to amaze you with
- 2:16:10this best example of Google Translator
- 2:16:12and AI which converts from one language
- 2:16:15to the another. I have chosen one
- 2:16:17language for you guys since many of you
- 2:16:19watch animes and all. So, I'm going to
- 2:16:21convert from Japanese language to the
- 2:16:24English language. Buckle up. Let's get
- 2:16:26started. So guys, I'm going to write a
- 2:16:29word in Japanese that is "Ohayo
- 2:16:32gozaimasu".
- 2:16:34If you know what "Ohayo gozaimasu"
- 2:16:36means, then please comment down below.
- 2:16:38So, I'm going to convert "Ohayo
- 2:16:40gozaimasu" from Japanese to English. So,
- 2:16:43let's copy this and paste this in here.
- 2:16:46So, "Ohayo gozaimasu" in Japanese, but
- 2:16:50in English it means good morning. A
- 2:16:52Google Translator works on a neural
- 2:16:54network. And a neural network is a
- 2:16:57machine learning algorithm which learns
- 2:16:59from Japanese language, keeps on
- 2:17:01improving itself, and gives you the
- 2:17:04optimal output that is good morning in
- 2:17:06English. This is how a language
- 2:17:08translator works. Let's see one more
- 2:17:11example. If I write here one more thing
- 2:17:14that is "Hajimemashite".
- 2:17:18"Hajimemashite" in Japanese, if you
- 2:17:20know, then please comment down below.
- 2:17:22Let me see what does it mean.
- 2:17:25So, "Hajimemashite" in Japanese, but in
- 2:17:27English it means nice to meet you. Here
- 2:17:30also the same. The neural network is
- 2:17:32learning from Japanese language and
- 2:17:34showing you what does it mean in English
- 2:17:37language. So, let me properly explain
- 2:17:40you what it actually does. So guys, a
- 2:17:43translator is nothing, but it's an AI
- 2:17:45machine that converts from one language
- 2:17:47to the another using a machine learning
- 2:17:50algorithm called neural network. Now,
- 2:17:52what neural network does is it learns
- 2:17:54from one language, makes some mistake or
- 2:17:56errors, then keeps on improving itself
- 2:17:58so that it can convert to to language
- 2:18:00such as from Japanese to English or from
- 2:18:03English to Japanese.
- 2:18:04So guys, the translator which you are
- 2:18:06using works on neural networks and this
- 2:18:09is how a translation works.
- 2:18:13So guys, you must be thinking then what
- 2:18:15is machine learning? So moving on to
- 2:18:17what is machine learning we have a
- 2:18:19machine learning is nothing but a
- 2:18:21training from historical data or
- 2:18:23experiences or in layman terms you can
- 2:18:25say learning from data to predict future
- 2:18:28or required output is called machine
- 2:18:30learning.
- 2:18:31Guys, a simple machine learning
- 2:18:33algorithm works in a way that you have a
- 2:18:35data of any kind and you give it to a
- 2:18:37machine. Now what does machine does is
- 2:18:40it learns from it in different ways
- 2:18:42using some different algorithms and
- 2:18:44gives you the required amount of future
- 2:18:47or predicts the required amount of
- 2:18:48output you wanted.
- 2:18:51Now guys, you must have understood what
- 2:18:53is machine learning. Let's deep dive a
- 2:18:55bit. Let's see what are the features of
- 2:18:57machine learning. We have five features
- 2:18:59of machine learning that is predictive
- 2:19:01modeling, automation, scalability,
- 2:19:05generalization, and adaptiveness. These
- 2:19:07are the five main features of any
- 2:19:09machine learning.
- 2:19:11So guys, tighten your seat belts. Let's
- 2:19:13move to these one by one. Predictive
- 2:19:15model. In the predictive model, what it
- 2:19:17does it it uses some mathematical
- 2:19:19functions and statistical techniques on
- 2:19:21the historical data and gives you the
- 2:19:24future predictions. The best example I
- 2:19:26can give you guys is the stock market
- 2:19:28prediction app where it uses some
- 2:19:29graphs, straight line graphs, and charts
- 2:19:32to show you the prediction based on the
- 2:19:34historical data using some statistical
- 2:19:36and mathematical functions. This is how
- 2:19:39a predictive model works.
- 2:19:42Moving on to our next topic that is
- 2:19:44automation. Automation is one of the
- 2:19:46best feature to save money. You know
- 2:19:48why? Because the companies which are
- 2:19:51having the less domains and cannot hire
- 2:19:53the employees, they can automate
- 2:19:55different machines to do the same work
- 2:19:58as the employee does. Such as if you
- 2:20:00want a developer, but you don't have the
- 2:20:02cost to pay, then you can automate a
- 2:20:04machine that can develop for you. This
- 2:20:06is the best example I can give you for
- 2:20:08the automation feature.
- 2:20:11So guys, moving on to our next feature,
- 2:20:13that is scalability. Scalability has its
- 2:20:15own importance because a machine
- 2:20:17learning algorithm, if it is not
- 2:20:19scalable, then it cannot handle larger
- 2:20:21amount of data sets or bigger data sets.
- 2:20:24I can give you the best example, that is
- 2:20:26Amazon.
- 2:20:28So guys, here I am at the Amazon
- 2:20:30website, and you can see a lot of
- 2:20:32product, and not only you can see, but
- 2:20:34whole world can see who are using Amazon
- 2:20:36app. Now, this system is a scalable
- 2:20:38system. No matter how many customers are
- 2:20:41here, and they are buying, the system
- 2:20:43will never crash. It can handle that
- 2:20:45amount of larger data sets. So, this is
- 2:20:48the best example of a scalable system.
- 2:20:51So guys, moving on to our next feature,
- 2:20:53that is generalization. In
- 2:20:55generalization, what it does is it's the
- 2:20:58ability of the model to generalize
- 2:21:01things, to forecast new data. Suppose
- 2:21:03your model is trained on a data set, and
- 2:21:06you're going to test it on some another
- 2:21:08data set, which is not there in the
- 2:21:09training, but still your model is giving
- 2:21:1295% of accuracy. Means your model is
- 2:21:15generalizing, your model is summarizing,
- 2:21:17and giving you the best and optimal
- 2:21:19result.
- 2:21:20Now guys, moving on to our last feature,
- 2:21:22that is adaptiveness. You can take it as
- 2:21:25a survival of the fittest thing, because
- 2:21:26if your model is not surviving the
- 2:21:28real-time environments or the new
- 2:21:30problems, then your model is not
- 2:21:32adaptive or good. It will going to
- 2:21:33extinct. Suppose there is a model which
- 2:21:36is built on traditional model, and still
- 2:21:38giving you best and advanced solutions
- 2:21:40on the real-time problems, then your
- 2:21:42model is adaptive and is the optimal
- 2:21:44model you can have.
- 2:21:46So guys, clear your mind because we are
- 2:21:48going to go in the types of machine
- 2:21:50learning. There are four different types
- 2:21:52of machine learning. First, we have is
- 2:21:54supervised or guided machine learning.
- 2:21:57In the supervised or guided machine
- 2:21:58learning, what it does is if there is a
- 2:22:00data set and having some values and you
- 2:22:03are labeling it as a specifying the
- 2:22:05value to the machine, then it will
- 2:22:07recognize those data through the labels
- 2:22:10and giving you the optimal
- 2:22:11classification. For example, if you have
- 2:22:14the pictures of cats and dogs and you
- 2:22:16have to classify it, then you will label
- 2:22:18it as cats and dogs. Then your machine
- 2:22:20will recognize those labels and classify
- 2:22:23and give you the optimal result.
- 2:22:26So, in supervised learning, what we have
- 2:22:28is a supervised algorithm. We have some
- 2:22:30labels and we are putting those labels
- 2:22:32to a data set. And after this, the
- 2:22:35machine easily recognizes and giving you
- 2:22:37the optimal results.
- 2:22:39So guys, moving on to our next type,
- 2:22:41that is unsupervised or unguided
- 2:22:44learning. In unsupervised or unguided
- 2:22:46learning, what machine does is you are
- 2:22:48giving the data which is not labeled.
- 2:22:50And after few trainings and making some
- 2:22:52errors, the machine easily recognizes
- 2:22:54this. So, what we have is a unsupervised
- 2:22:57algorithm. We have some unlabeled data
- 2:23:00and in those unlabeled data, your
- 2:23:02machine is trying to find patterns.
- 2:23:03After few errors, it will give you the
- 2:23:06optimal results.
- 2:23:07Moving on to our next type, that is
- 2:23:09semi-supervised learning. In
- 2:23:11semi-supervised learning, the algorithm
- 2:23:13uses both unsupervised and supervised in
- 2:23:16combined form giving you the optimal
- 2:23:18result. For example, we have a
- 2:23:20semi-supervised algorithm and we are
- 2:23:22trying to find out patterns from the
- 2:23:24data sets which are both labeled and
- 2:23:26unlabeled. So, this is how a
- 2:23:28semi-supervised learning works.
- 2:23:30Moving on to our last type, that is
- 2:23:33reinforcement learning. The
- 2:23:34reinforcement learning you can
- 2:23:36understand in a way like when you are
- 2:23:37playing a game, you make some mistake,
- 2:23:39then learn from them, then again make
- 2:23:41some mistake in a level, learn from them
- 2:23:43and reach your goal. This This how
- 2:23:45reinforcement learning works in a
- 2:23:47software. The software is being trained
- 2:23:49multiple times making some errors and
- 2:23:51learning from them and giving you the
- 2:23:53optimal results. This is the best
- 2:23:55example I can give you for the
- 2:23:57reinforcement learning that is the
- 2:23:58gaming system. So guys, what we have is
- 2:24:01a reinforcement algorithm and an
- 2:24:03environment. We are taking some actions
- 2:24:05in that environment, making some errors
- 2:24:07and getting some results. Then again
- 2:24:09making some errors and getting some
- 2:24:11results. This is how a reinforcement
- 2:24:13model works in a machine learning.
- 2:24:16So guys, moving on to the examples of
- 2:24:18machine learning, we have a voice
- 2:24:19recognition system and a image
- 2:24:22recognition system. These both you are
- 2:24:24using in your phone, in your laptop
- 2:24:26every day and machine learning is being
- 2:24:28used. Now guys, you have learned so
- 2:24:30much. Now you must be thinking that how
- 2:24:33should I start my journey? What skills I
- 2:24:35should have? So these are the skills
- 2:24:38required to learn machine learning. That
- 2:24:40is SQL, structured query language,
- 2:24:42JavaScript, C++, R programming for those
- 2:24:46who are moving with machine learning to
- 2:24:47the data scientist and Python, one of
- 2:24:50the best programming language for the
- 2:24:51machine learning. And last but not the
- 2:24:53least, if you are moving a bit deep down
- 2:24:56in the machine learning towards the deep
- 2:24:57learning, then you need NLP or natural
- 2:25:01language processing. These skills are
- 2:25:03required for the one who want to learn
- 2:25:05machine learning.
- 2:25:07So guys, moving on to the real-life
- 2:25:09examples of machine learning I have,
- 2:25:11that is Google searches. Google searches
- 2:25:14uses our history, track it down, then
- 2:25:16giving you the prediction based on your
- 2:25:18histories. For example, let me show you.
- 2:25:21So guys, I came here at Google. Now I'm
- 2:25:24going to type Amazon and it's giving me
- 2:25:28that results which are already there in
- 2:25:30my history. I must have searched before
- 2:25:32a month ago, uh 2 months ago, then 4
- 2:25:34months ago and all that history combined
- 2:25:37form is being predicted in here. Now
- 2:25:39Amazon Prime is there, Amazon videos are
- 2:25:42there. So, this is how a predictive
- 2:25:44model is working behind this Google
- 2:25:45searches.
- 2:25:47Moving on to our next real-life example,
- 2:25:49guys, that is Instagram. You swipe
- 2:25:52Instagram every day, every night. So,
- 2:25:54all that is based on your past data
- 2:25:57only. If you are seeing some videos of
- 2:25:58cats, then further swipes will be the
- 2:26:00cats only. So, this is how your history
- 2:26:03is being tracked down and watched by the
- 2:26:05Instagram machine learning algorithms
- 2:26:07and giving you the results of the same.
- 2:26:10So, this is how a real feeder in
- 2:26:13Instagram works. Moving on to our next
- 2:26:15example, that is movie recommendation
- 2:26:17system on any movie watching website
- 2:26:19such as Netflix or anime websites such
- 2:26:22as Watch Anime, Anycon, etc.
- 2:26:25So, guys, if I take you to Any watch and
- 2:26:28show you the anime recommendation
- 2:26:30trending ones are based on the people's
- 2:26:32choices they are watching more based on
- 2:26:34your histories only. If I watch any of
- 2:26:36these animes and I watch them regularly,
- 2:26:39then the recommendations will show me
- 2:26:41the same. So, this is how a
- 2:26:43recommendation system works in Any watch
- 2:26:46or you can say in Netflix based on your
- 2:26:48choices in past data.
- 2:26:51So, guys, moving on to the future of
- 2:26:53machine learning, it will be using
- 2:26:55everywhere. The advanced techniques in
- 2:26:57medical field, architecture field,
- 2:26:59electrical field, and yes, in the
- 2:27:01computer science field. Machine learning
- 2:27:03will be everywhere. It will be
- 2:27:04revolutionizing the world.
- 2:27:12Now, let's explore different types of ML
- 2:27:14models.
- 2:27:16So, not all data is structured the same
- 2:27:18way. And different problems require
- 2:27:19different approaches.
- 2:27:21So, for example, predicting stock prices
- 2:27:23requires the models that learn from
- 2:27:25historical trends.
- 2:27:26And then identifying objects in images
- 2:27:29needs models that recognize patterns in
- 2:27:31visual data.
- 2:27:32Next, the chatbots and voice assistants
- 2:27:35rely on the models trained to understand
- 2:27:37and generate human language.
- 2:27:39So, to tackle these challenges, as I
- 2:27:41discussed previously, that ML is divided
- 2:27:43into different learning models, such as
- 2:27:45supervised, unsupervised, and
- 2:27:48reinforcement learning. And each has its
- 2:27:50own strengths, and it is used depending
- 2:27:52on the problem at hand.
- 2:27:54Since we know why different ML models
- 2:27:56are needed, let's see how they play a
- 2:27:58crucial role in generative AI.
- 2:28:00Well, generative AI is one of the most
- 2:28:03exciting applications of machine
- 2:28:04learning. And unlike traditional ML
- 2:28:07models that make predictions or
- 2:28:08classifications, generative models
- 2:28:10create entirely new content. And here's
- 2:28:13how ML enables AI to generate.
- 2:28:15So, first here we have text. A language
- 2:28:18models, like GPT, generate human-like
- 2:28:20text for chatbots, content writing, and
- 2:28:22coding.
- 2:28:23Next is the image.
- 2:28:25So, AI-powered tools, like DALL-E, can
- 2:28:27create realistic images from textual
- 2:28:30descriptions.
- 2:28:31Next is videos. So, advanced ML models
- 2:28:34synthesize lifelike video content,
- 2:28:37transforming media, marketing, and even
- 2:28:39filmmaking.
- 2:28:40So, these advancements in generative AI
- 2:28:42are reshaping creativity and automation,
- 2:28:45proving that machine learning is not
- 2:28:46just about making decision, it's about
- 2:28:48creating new possibilities.
- 2:28:50So, now that we have seen how ML models
- 2:28:52enable AI to create new content. So, now
- 2:28:55let us briefly understand the different
- 2:28:56types of machine learning models.
- 2:28:58So, here, the first type of machine
- 2:29:00learning model is supervised learning.
- 2:29:03Supervised learning trains a model using
- 2:29:05labeled data, where each input has a
- 2:29:07corresponding correct output. And this
- 2:29:09makes it ideal for tasks where
- 2:29:10historical data can be used to predict
- 2:29:12future outcomes.
- 2:29:14For example, let's say spam detection.
- 2:29:17Email services, like Gmail, use a
- 2:29:19supervised learning to classify emails
- 2:29:21as spam or not spam by learning from
- 2:29:24past labeled examples.
- 2:29:26The next example is the price
- 2:29:27predictions.
- 2:29:29So, real estate platforms use regression
- 2:29:31models to predict house prices based on
- 2:29:33the features like location, size, and
- 2:29:36amenities.
- 2:29:37Now, let us see some of the popular
- 2:29:39algorithms.
- 2:29:40So, first let's discuss on decision
- 2:29:42trees. These models break down the data
- 2:29:45into a tree-like structure, where each
- 2:29:47node represent a decision based on a
- 2:29:49feature.
- 2:29:50So, they are easy to interpret and work
- 2:29:52well for both classification. For
- 2:29:54example, deciding if an email is a spam
- 2:29:57or not. And regression example,
- 2:29:59predicting house price.
- 2:30:01However, they can become overly complex.
- 2:30:04Next is the support vector machines.
- 2:30:07So, SVMs are powerful for classification
- 2:30:09task, as they find the optimal boundary,
- 2:30:11also called a hyperplane. And that best
- 2:30:14separates different classes in the data.
- 2:30:17They work well for high-dimensional
- 2:30:19spaces and cases where the distinction
- 2:30:21between categories is clear, such as
- 2:30:23handwriting, facial recognition, or
- 2:30:26medical diagnosis.
- 2:30:27So, now that we have seen how labeled
- 2:30:29data is used. So, now let's explore how
- 2:30:31unsupervised learning finds patterns
- 2:30:33without labels.
- 2:30:35Well, unsupervised learning works with
- 2:30:37unlabeled data, identifying hidden
- 2:30:39patterns and relationships without
- 2:30:41predefined categories.
- 2:30:42So, here we have some of the popular
- 2:30:44algorithms. So, first is the K-means
- 2:30:47clustering. This algorithm partitions
- 2:30:49data into a predefined number of
- 2:30:51clusters.
- 2:30:52By grouping similar data points based on
- 2:30:54their attributes. It works well for
- 2:30:57tasks like customer segmentation. Where
- 2:30:59businesses can group customer based on
- 2:31:02purchasing behavior. However, it assumes
- 2:31:04clusters are spherical and may struggle
- 2:31:07with irregular shaped data. Next we have
- 2:31:10autoencoders.
- 2:31:12So, these are specialized neural
- 2:31:13networks designed to learn efficient
- 2:31:15data representations by encoding and
- 2:31:18reconstructing input data. Let us see
- 2:31:20some of the examples.
- 2:31:22So, first example here we have is
- 2:31:24customer segmentation.
- 2:31:26Where e-commerce platforms group
- 2:31:28customer based on their shopping
- 2:31:29behavior to offer personalized
- 2:31:31recommendations.
- 2:31:32The next example is market analysis.
- 2:31:35Businesses analyze purchasing trends to
- 2:31:37find associations such as which products
- 2:31:40are frequently brought together.
- 2:31:42Now we have covered both labeled and
- 2:31:43unlabeled learning. So let's see how
- 2:31:45semi-supervised learning combines the
- 2:31:47best of both worlds.
- 2:31:49So semi-supervised learning bridges the
- 2:31:52gap between the supervised and
- 2:31:53unsupervised learning by using a small
- 2:31:55amount of data along with large amount
- 2:31:58of unlabeled data.
- 2:31:59So for example, let's say AI assistant
- 2:32:01medical diagnosis.
- 2:32:03Labeled medical images such as x-rays
- 2:32:06with diagnosis are scarce, but large
- 2:32:08amounts of unlabeled images exist.
- 2:32:11Semi-supervised learning help AI learn
- 2:32:14patterns from both labeled and unlabeled
- 2:32:16data improving accuracy in disease
- 2:32:18detection. All right. Now let's explore
- 2:32:21the reinforcement learning where AI
- 2:32:23learns through trial and error.
- 2:32:25Well, reinforcement learning is inspired
- 2:32:28by the concept of learning through trial
- 2:32:30and error. So models interact with an
- 2:32:32environment, receive rewards or
- 2:32:34penalties for actions and refine their
- 2:32:37strength over time. For example, let's
- 2:32:39say gaming.
- 2:32:40Mario AI developed using reinforcement
- 2:32:42learning learns to navigate levels by
- 2:32:45optimizing actions through trial and
- 2:32:47error. The next example is robotics.
- 2:32:50Where robots learn to walk, balance, or
- 2:32:52perform tasks through reinforcement
- 2:32:54learning by maximizing positive
- 2:32:56outcomes.
- 2:32:57Also, reinforcement learning uses
- 2:32:59agents, actions, and rewards to improve
- 2:33:02decision-making.
- 2:33:03Making it ideal for tasks requiring
- 2:33:05continuous learning and adaptation.
- 2:33:08So now that we have covered all the
- 2:33:09types of machine learning models. So
- 2:33:11let's go over some of the key tips to
- 2:33:13help you choose the right one for your
- 2:33:14needs.
- 2:33:16So here are the tips.
- 2:33:17When it comes to supervised learning
- 2:33:19classifying emails as spam or not and
- 2:33:22diagnosing diseases from patient data.
- 2:33:25Next is the unsupervised learning. So,
- 2:33:27unsupervised learning is best when
- 2:33:29you're grouping shoppers by behavior and
- 2:33:31detecting fraud in banking. Next, we
- 2:33:33have semi-supervised learning.
- 2:33:35And this is best when you're improving
- 2:33:37speech recognition with limited label
- 2:33:39data and identifying fake news.
- 2:33:42And finally, the reinforcement learning.
- 2:33:45This will be best when you're training
- 2:33:46self-driving cars to navigate,
- 2:33:48optimizing AI in video games like Mario.
- 2:33:51So, whether it's supervised,
- 2:33:53unsupervised, semi-supervised, or
- 2:33:55reinforcement learning, each model plays
- 2:33:57a crucial role in shaping AI's future.
- 2:34:00So, as generative AI continues to
- 2:34:02evolve, these models are driving
- 2:34:03innovation in text, images, and video
- 2:34:06generation.
- 2:34:07So, which machine learning model do you
- 2:34:09find the most fascinating? Let me know
- 2:34:11in the comments below.
- 2:34:15>> [music]
- 2:34:18>> Let me connect you to the real life and
- 2:34:20tell you what all are the things which
- 2:34:22you can easily do using the concepts of
- 2:34:23machine learning.
- 2:34:25So, you can easily get answer to the
- 2:34:26questions like which types of house lies
- 2:34:28in this segment or what is the market
- 2:34:30value of this house? Or is this a mail a
- 2:34:33spam or not a spam? Is there any fraud?
- 2:34:36Well, these are some of the question you
- 2:34:37could ask to the machine. But for
- 2:34:38getting an answer to these, you need
- 2:34:40some algorithm. The machine need to
- 2:34:42train on the basis of some algorithm.
- 2:34:44Okay, but how will you decide which
- 2:34:46algorithm to choose and when?
- 2:34:48Okay, so the best option for us is to
- 2:34:50explore them one by one.
- 2:34:53So, the first is classification
- 2:34:54algorithm where the category is
- 2:34:56predicted using the data. If you have
- 2:34:58some question like is this person a male
- 2:35:01or a female? Or is this a mail a spam or
- 2:35:04not a spam? Then these category of
- 2:35:06question would fall under the
- 2:35:07classification algorithm.
- 2:35:09Classification is a supervised learning
- 2:35:10approach in which the computer program
- 2:35:13learns from the input given to it and
- 2:35:14then uses this learning to classify new
- 2:35:17observation. Some examples of
- 2:35:19classification problems are speech
- 2:35:21organization, handwriting recognition,
- 2:35:23biometric identification, document
- 2:35:25classification, etc.
- 2:35:28Shall we move ahead?
- 2:35:30Okay.
- 2:35:32So, next is the anomaly detection
- 2:35:34algorithm where you identify the unusual
- 2:35:37data point. So, what is anomaly
- 2:35:38detection? Well, it's a technique that
- 2:35:40is used to identify unusual pattern that
- 2:35:43do not conform to expected behavior. Or
- 2:35:45you can say the outliers.
- 2:35:47It has many application in business like
- 2:35:49intrusion detection, like identifying
- 2:35:51strange patterns in the network traffic
- 2:35:53that could signal a hack, or system
- 2:35:55health monitoring, that is spotting a
- 2:35:56deadly tumor in the MRI scan.
- 2:35:59Or you can even use it for fraud
- 2:36:01detection in credit card transaction, or
- 2:36:03to deal with fault detection in
- 2:36:04operating environment.
- 2:36:06So, next comes the clustering algorithm.
- 2:36:08You can use this clustering algorithm to
- 2:36:10group the data based on some similar
- 2:36:12condition. Now, you can get answer to
- 2:36:14which type of houses lies in this
- 2:36:16segment, or what type of customer buys
- 2:36:18this product. The clustering is a task
- 2:36:20of dividing the population or data
- 2:36:22points into a number of groups such that
- 2:36:24the data point in the same groups are
- 2:36:26more similar to other data points in the
- 2:36:28same group than those in the other
- 2:36:30groups. In simple words, the aim is to
- 2:36:33segregate groups with similar trait and
- 2:36:35assign them into cluster.
- 2:36:37Now, this clustering is a task of
- 2:36:38dividing the population or data points
- 2:36:40into a number of groups such that the
- 2:36:42data points in the X group is more
- 2:36:44similar to the other data points in the
- 2:36:46same group rather than those in the
- 2:36:48other group. In other words, the aim is
- 2:36:50to segregate the groups with similar
- 2:36:52traits and assign them into different
- 2:36:54clusters. Let's understand this with an
- 2:36:56example. Suppose you're the head of a
- 2:36:58rental store and you wish to understand
- 2:37:00the preference of your customer to scale
- 2:37:02up your business. So, is it possible for
- 2:37:04you to look at the detail of each
- 2:37:05customer and design a unique business
- 2:37:08strategy for each of them?
- 2:37:09Definitely not. Right?
- 2:37:12But what you can do is to cluster all
- 2:37:14your customer saying to 10 different
- 2:37:16groups based on their purchasing habit
- 2:37:18and you can use a separate strategy for
- 2:37:20customers in each of these 10 different
- 2:37:22groups. And this is what we call
- 2:37:24clustering.
- 2:37:26Next we have regression algorithm where
- 2:37:28the data itself is predicted. Question
- 2:37:30you may ask to this type of model is
- 2:37:32like what is the market value of this
- 2:37:34house or is it going to rain tomorrow or
- 2:37:36not?
- 2:37:37So regression is one of the most
- 2:37:39important and broadly used machine
- 2:37:40learning and statistics tool.
- 2:37:43It allows you to make prediction from
- 2:37:44data by learning the relationship
- 2:37:46between the features of your data and
- 2:37:48some observed continuous valued
- 2:37:49response. Regression is used in a
- 2:37:52massive number of application.
- 2:37:54You know what? Stock prices prediction
- 2:37:55can be done using regression.
- 2:37:57Now you know about different machine
- 2:37:59learning algorithm. How will you decide
- 2:38:01which algorithm to choose and when?
- 2:38:03So let's cover this part using a demo.
- 2:38:05So in this demo part, what we'll do,
- 2:38:07we'll create six different machine
- 2:38:08learning model and pick the best model
- 2:38:11and build the confidence such that it
- 2:38:12has the most reliable accuracy.
- 2:38:16So for our demo part, we'll be using the
- 2:38:17Iris data set. This data set is quite
- 2:38:20very famous and is considered one of the
- 2:38:22best small project to start with.
- 2:38:24You can consider this as a hello world
- 2:38:25data set for machine learning. So this
- 2:38:27data set consists of 150 observation of
- 2:38:30Iris flower.
- 2:38:31There are four columns of measurement of
- 2:38:33flowers in centimeters. The fifth column
- 2:38:35being the species of the flower
- 2:38:36observed. All the observed flowers
- 2:38:38belong to one of the three species of
- 2:38:40Iris setosa, Iris virginica and Iris
- 2:38:43versicolor.
- 2:38:44Well, this is a good project because it
- 2:38:46is so well to understand. The attributes
- 2:38:48are numeric so you have to figure out
- 2:38:49how to load and handle the data. It is a
- 2:38:51classification problem thereby allowing
- 2:38:53you to practice with perhaps an easier
- 2:38:55type of supervised learning algorithm.
- 2:38:57It has only four attributes and 150 rows
- 2:38:59meaning it is very small and can easily
- 2:39:01fit into the memory.
- 2:39:02And even all of the numeric attributes
- 2:39:04are in same unit and the same scale. It
- 2:39:07means you do not require any special
- 2:39:08scaling or transformation to get
- 2:39:10started.
- 2:39:12So, let's start coding and as I told
- 2:39:14earlier for the demo part, I'll be using
- 2:39:16Anaconda with Python 3.0 installed on
- 2:39:18it. So, when you install Anaconda, how
- 2:39:20your navigator would look like. So,
- 2:39:22there's my home page of my Anaconda
- 2:39:23navigator. On this I'll be using the
- 2:39:26Jupiter notebook, which is a web-based
- 2:39:27interactive computing notebook
- 2:39:29environment, which will help me to write
- 2:39:30and execute my Python codes on it. So,
- 2:39:32let's hit the launch button and execute
- 2:39:34our Jupiter notebook.
- 2:39:36So, as you can see that my Jupiter
- 2:39:37notebook is starting on localhost 8890.
- 2:39:41Okay? So, this is my Jupiter notebook.
- 2:39:42What I'll do here, I'll select new
- 2:39:44notebook Python 3.
- 2:39:48There's my environment where I can write
- 2:39:50and execute all my Python codes on it.
- 2:39:52So, let's start by checking the version
- 2:39:54of the libraries. In order to make this
- 2:39:56video short and more interactive and
- 2:39:57more informative, I've already done the
- 2:39:59set of code. So, let me just copy and
- 2:40:01paste it down. I'll explain you then one
- 2:40:03by one.
- 2:40:04So, let's start by checking the version
- 2:40:05of the Python libraries.
- 2:40:07Okay? So, there's the code. Let's just
- 2:40:10copy it.
- 2:40:11Copied and let's paste it. Okay. First,
- 2:40:14let me summarize things for you. What we
- 2:40:16are doing here, we are just checking the
- 2:40:17version of the different libraries.
- 2:40:19Starting with Python, we'll first check
- 2:40:20what version of Python we are working
- 2:40:22on, then we'll check what are the
- 2:40:23version of SciPy we are using, then
- 2:40:25NumPy, Matplotlib, then Pandas, then
- 2:40:27scikit-learn. Okay? So, let's execute
- 2:40:29the run button and see what are the
- 2:40:30various version of libraries which we
- 2:40:32are using. Hit the run. So, we are
- 2:40:33working on Python 3.6.4, SciPy 1.0,
- 2:40:37NumPy 1.14, Matplotlib 2.12, Pandas
- 2:40:400.22, and scikit-learn of version 0.19.
- 2:40:44Okay?
- 2:40:45So, these are the version which I'm
- 2:40:46using. Ideally, your version should be
- 2:40:48more recent or it should match. But,
- 2:40:50don't worry if you lag few versions
- 2:40:52behind as the APIs do not change so
- 2:40:54quickly. Everything in this tutorial
- 2:40:56will very likely still work for you.
- 2:40:58Okay? But, in case you're getting an
- 2:41:00error, stop and try to fix that error.
- 2:41:03In case you're unable to find the
- 2:41:04solution for the error, feel free to
- 2:41:06reach out Edureka even after this class.
- 2:41:08Let me tell you this, if you're not able
- 2:41:09to run the script properly, you will not
- 2:41:11be able to complete this tutorial, okay?
- 2:41:13So, whenever you get a doubt, reach out
- 2:41:15to Edureka and just resolve it.
- 2:41:17Now, if everything is working smoothly,
- 2:41:19then now it's the time to load the data
- 2:41:21set. So, as I said, I'll be using the
- 2:41:23Iris flower data set for this tutorial.
- 2:41:25But, before loading the data set, let's
- 2:41:27import all the modules, function, and
- 2:41:29the object which we are going to use in
- 2:41:31this tutorial. Same, I've already
- 2:41:32written the set of code, so let's just
- 2:41:34copy and paste them. Let's load all the
- 2:41:36libraries.
- 2:41:38So, these are the various libraries
- 2:41:39which we'll be using in our tutorial.
- 2:41:42So, everything should work fine without
- 2:41:43an error. If you get an error, just
- 2:41:45stop. You need to work on your SciPy
- 2:41:46environment before you continue any
- 2:41:48further. So, I guess everything should
- 2:41:50work fine. Let's hit the run button and
- 2:41:51see.
- 2:41:53Okay, it worked. So, let's now move
- 2:41:56ahead and load the data. We can load the
- 2:41:58data direct from the UCI machine
- 2:41:59learning repository. First of all, let
- 2:42:01me tell you, we are using Panda to load
- 2:42:03the data.
- 2:42:04Okay?
- 2:42:05So, let's say my URL is this. So, this
- 2:42:08is my URL for the UCI machine learning
- 2:42:09repository from where I'll be
- 2:42:10downloading the data set, okay?
- 2:42:13Now, what I'll do, I'll specify the name
- 2:42:14of each column when loading the data.
- 2:42:16This will help me later to explore the
- 2:42:18data, okay?
- 2:42:19So, I'll just copy and paste it down.
- 2:42:22Okay?
- 2:42:23So, I'm defining a variable names which
- 2:42:25consists of various parameters including
- 2:42:27sepal length, sepal width, petal length,
- 2:42:29petal width, and class. So, these are
- 2:42:31just the name of column from the data
- 2:42:32set, okay? Now, let's define the data
- 2:42:35set. So, data set equals panda.read_csv.
- 2:42:39Inside that, we are defining URL and the
- 2:42:41names, that is equal to name.
- 2:42:44As I already said, we'll be using Panda
- 2:42:46to load the data, all right?
- 2:42:49So, we are using panda.read_csv, so we
- 2:42:51are reading the CSV file, and inside
- 2:42:53that, from where that CSV is coming?
- 2:42:54From the the Which URL? So, this is my
- 2:42:56URL. Okay?
- 2:42:58And names equal names. It's just
- 2:43:00specifying the names of the various
- 2:43:01columns in that particular CSV file.
- 2:43:03Okay?
- 2:43:04So, let's move forward and execute it.
- 2:43:06So, even our data set is loaded.
- 2:43:09In case you have some network issues,
- 2:43:11just go ahead and download the Iris data
- 2:43:13file into your working directory and
- 2:43:14load it using the same method. But yeah,
- 2:43:16make sure that you change the URL to the
- 2:43:18local name, or else you might get an
- 2:43:19error. Okay. Yeah, our data set is
- 2:43:22loaded. So, let's move ahead and check
- 2:43:23our data set. Let's see how many columns
- 2:43:25or rows we have in our data set. Okay.
- 2:43:28So, let's print the number of rows and
- 2:43:30columns in our data set. So, our data
- 2:43:32set is data set.shape.
- 2:43:35What this will do, it will just give you
- 2:43:37the numbers of total number of rows and
- 2:43:39total number of column, or you can say
- 2:43:40the total number of instances or
- 2:43:42attributes in your data set. Fine?
- 2:43:44So, print data set.shape. What are you
- 2:43:46getting? 150 and 5. So, 150 is the total
- 2:43:49number of rows in your data set, and 5
- 2:43:50is the total number of columns. Fine?
- 2:43:53So, moving on ahead, what if I want to
- 2:43:55see the sample data set? Okay. So, let
- 2:43:58me just print the first 30 instances of
- 2:43:59the data set. Okay? So, print
- 2:44:03data set.head.
- 2:44:07What I want is the first 30 instances.
- 2:44:09Fine? This will give me the first 30
- 2:44:11result of my data set. Okay? So, when I
- 2:44:13hit the run button, what I'm getting is
- 2:44:15the first 30 result. Okay?
- 2:44:180
- 2:44:20to 29. So, this is how my sample data
- 2:44:22set looks like.
- 2:44:23Sepal length, sepal width, petal length,
- 2:44:25petal width, and the class. Okay?
- 2:44:28So, this is how our data set looks like.
- 2:44:31Now, let's move on and look at the
- 2:44:32summary of each attribute. What if I
- 2:44:34want to find out the count, mean, the
- 2:44:37minimum and the maximum values, and some
- 2:44:39other percentiles as well. So, what
- 2:44:40should I do then?
- 2:44:41For that, print
- 2:44:43data set.describe.
- 2:44:47What it will give,
- 2:44:48let's see.
- 2:44:50So, you can see that all the numbers are
- 2:44:52the same scales of similar range between
- 2:44:540 to 8 cm, right? The mean value, the
- 2:44:57standard deviation, the minimum value,
- 2:44:59the 25th percentile, 50th percentile,
- 2:45:0175th percentile, the maximum value, all
- 2:45:03these values lies in the range between 0
- 2:45:05to 8 cm.
- 2:45:07Okay.
- 2:45:08So what we just did is we just took a
- 2:45:10summary of each attribute. Now let's
- 2:45:13look at the number of instances that
- 2:45:14belong to each class. So for that, what
- 2:45:17we'll do print data set first of all.
- 2:45:21So let's print data set and I want to
- 2:45:24group it
- 2:45:25group by using
- 2:45:28class
- 2:45:30and I want the size of it, size of each
- 2:45:32class. Fine?
- 2:45:34And let's hit the run.
- 2:45:47Okay. So what I want to do, I want to
- 2:45:49print print what? Data set. How I want
- 2:45:52to get it? I want it by class. So group
- 2:45:54by class.
- 2:45:56Okay. Now I want the size of each class.
- 2:45:59Find the size of each class. So group by
- 2:46:01class.size. Execute the run.
- 2:46:04So you can see that I have 15 instances
- 2:46:06of Iris setosa, 15 instances of Iris
- 2:46:08versicolor, and 15 instances of Iris
- 2:46:10virginica. Okay? All are of data type
- 2:46:13integer of base 64. Fine? So now we have
- 2:46:16a basic idea of our data. Now let's move
- 2:46:18ahead and create some visualization for
- 2:46:20it. So for this we are going to create
- 2:46:22two different types of plot. First would
- 2:46:23be the univariate plot and the next
- 2:46:25would be the multivariate plot. So we'll
- 2:46:26be creating univariate plots to better
- 2:46:28understand about each attribute. And the
- 2:46:30next we'll be creating the multivariate
- 2:46:32plot to better understand the
- 2:46:33relationship between different
- 2:46:34attributes. Okay? So we start with some
- 2:46:36univariate plot. That is plot of each
- 2:46:38individual variable. So given that the
- 2:46:40input variables are numeric, we can
- 2:46:41create box and whiskers plot for it.
- 2:46:43Okay? So let's move ahead and create a
- 2:46:44box and whiskers plot. So data set.plot.
- 2:46:47What kind I want? It's a box.
- 2:46:50Okay. And do I need a subplot? Yeah, I
- 2:46:53need subplots for that. So, subplots
- 2:46:55equal true. What type of layout do I
- 2:46:57want? So, my layout structure is 2 cross
- 2:47:012.
- 2:47:02Next, do I want to share my coordinates,
- 2:47:04X and Y coordinates? No, I don't want to
- 2:47:06share it. So, share X equal false.
- 2:47:09And even share Y, that too equals false.
- 2:47:13Okay. So, we have our dataset.plot kind
- 2:47:16equal box. My subplots is true, layout 2
- 2:47:19cross 2. And then what I want to do it,
- 2:47:21I want to see it. So, plot.show.
- 2:47:23Whatever I created, show it. Okay.
- 2:47:26Execute it.
- 2:47:29Now, this gives us a much clearer idea
- 2:47:30about the distribution of the input
- 2:47:32attribute. Now, what if I had given the
- 2:47:33layout to 2 cross 2 instead of that, I'd
- 2:47:36have given it 4 cross 4. So, what it
- 2:47:39will result? Just see. Fine. Everything
- 2:47:41would be printed in just one single row.
- 2:47:43Hold on, guys. Arya has a doubt. He's
- 2:47:44asking that why we are using the share X
- 2:47:46and share Y values. What are these? Why
- 2:47:48we have assigned false values to it?
- 2:47:50Okay, Arya. So, in order to resolve this
- 2:47:52query, I need to show you what will
- 2:47:54happen if I give true values to them.
- 2:47:55Okay. So, be with me. So, share X equal
- 2:47:58true and share Y, that equals true. So,
- 2:48:01let's see what result we'll get.
- 2:48:04You're getting it. The X and Y
- 2:48:05coordinates are just shared among all
- 2:48:07the four visualization, right? So, Arya,
- 2:48:09you can see that the sepal length and
- 2:48:10sepal width has Y values ranging from
- 2:48:130.0 to 7.5 which are being shared among
- 2:48:15both the visualization. So, is with the
- 2:48:17petal length, it has shared value
- 2:48:19between 0.0 to 7.5. Okay. So, that is
- 2:48:22why I don't want to share the value of X
- 2:48:24and Y. It's just giving us a cluttered
- 2:48:26visualization. So, Arya, why I'm doing
- 2:48:28this? I'm just doing it cuz I don't want
- 2:48:31my X and Y coordinates to be shared
- 2:48:33among any visualization. Okay. That is
- 2:48:35why my share X and share Y value are
- 2:48:37false. Okay. Let's execute it.
- 2:48:40So, this is a
- 2:48:41pretty much clear visualization which
- 2:48:43gives a clear idea about the
- 2:48:44distribution of the input attributes.
- 2:48:46Now, if you want, you can also create a
- 2:48:48histogram of each input variable to get
- 2:48:50a clear idea of the distribution. So,
- 2:48:52let's create a histogram for it. So,
- 2:48:53dataset.hist, okay? I would need to
- 2:48:56proceed. So, plot.show. Let's see. So,
- 2:48:58this is my histogram and it seems that
- 2:49:00we have two input variables that have a
- 2:49:02Gaussian distribution. So, this is
- 2:49:04useful to note as we can use the
- 2:49:06algorithms that can exploit this
- 2:49:07assumption, okay? So, next comes the
- 2:49:09multivariate plot. Now that we have
- 2:49:10created the univariate plot to
- 2:49:12understand about each attribute, let's
- 2:49:14move on and look at the multivariate
- 2:49:16plot and see the interaction between the
- 2:49:18different variables. So, first let's
- 2:49:20look at the scatter plot of all the
- 2:49:21attribute. This can be helpful to spot
- 2:49:23structured relationship between input
- 2:49:24variables, okay? So, let's create a
- 2:49:26scatter matrix. So, for creating a
- 2:49:27scatter plot, we need scatter matrix and
- 2:49:31we need to pass our dataset into it,
- 2:49:33okay? And then, what I want, I want to
- 2:49:35see it. So, plot.show. So, this is how
- 2:49:37my scatter matrix looks like. It's like
- 2:49:39that the diagonal grouping of some pair,
- 2:49:41right? So, this suggests a high
- 2:49:42correlation and a predictable
- 2:49:44relationship, all right? This was our
- 2:49:45multivariate plot. Now, let's move on
- 2:49:47and evaluate some algorithm. Now, it's
- 2:49:49time to create some model of the data
- 2:49:51and estimate the accuracy on the base of
- 2:49:53unseen data, okay? So, now we know all
- 2:49:56about our dataset, right? We know how
- 2:49:58many instances and attributes are there
- 2:49:59in our dataset. We know the summary of
- 2:50:01each attribute. Now, I guess we have
- 2:50:03seen much about our dataset. Now, let's
- 2:50:05move on and create some algorithm and
- 2:50:07estimate their accuracy based on the
- 2:50:09unseen data. Okay. Now, what we'll do,
- 2:50:11we'll create some model of the data and
- 2:50:13estimate the accuracy based on the some
- 2:50:15unseen data, okay? So, for that, first
- 2:50:17of all, let's create a validation
- 2:50:18dataset. What is a validation dataset?
- 2:50:20Validation dataset is your training
- 2:50:22dataset that will be using it to train
- 2:50:24our model, fine? All right. So, how
- 2:50:26we'll create a validation dataset? For
- 2:50:28creating a validation dataset, what we
- 2:50:30are going to do is we are going to split
- 2:50:31our dataset into two part, okay? So, the
- 2:50:33very first thing we'll do is to create a
- 2:50:35validation dataset. So, why do we even
- 2:50:37need a validation data set? So, we need
- 2:50:39a validation data set to know that the
- 2:50:41model we created is any good. Later,
- 2:50:43what we'll do, we'll use the statistical
- 2:50:45method to estimate the accuracy of the
- 2:50:47model that we create on the unseen data.
- 2:50:49We also want a more concrete estimate of
- 2:50:51the accuracy of the best model on unseen
- 2:50:53data by evaluating it on the actual
- 2:50:55unseen data. Okay? Confused? Let me
- 2:50:57simplify this for you. What we'll do,
- 2:50:59we'll split the loaded data into two
- 2:51:00parts. The first 80% of the data we'll
- 2:51:03use it to train our model. And the rest
- 2:51:0520% we'll hold back as the validation
- 2:51:07data set that we'll use it to verify our
- 2:51:09trained model. Okay? Fine. So, let's
- 2:51:11define an array. This is my array. What
- 2:51:14it will consist of? It will consist of
- 2:51:16all the values from the data set. So,
- 2:51:17data set.values.
- 2:51:19Okay? Next, I'll define a variable X
- 2:51:22which will consist of all the column
- 2:51:25from the array from zero to four.
- 2:51:28Starting from zero to four. And the next
- 2:51:30variable Y which would consist of the
- 2:51:33array starting from this. So, first of
- 2:51:37all, we'll define a variable X that will
- 2:51:39consist of the values in the array
- 2:51:41starting from the beginning zero till
- 2:51:43four. Okay? So, these are the column
- 2:51:45which we'll include in the X variable.
- 2:51:47And for a Y variable, I'll define it as
- 2:51:49a class or the output. So, what I need,
- 2:51:51I just need the fourth column that is my
- 2:51:53class column. So, I'll start it from the
- 2:51:55beginning and I just want the fourth
- 2:51:57column. Okay? Now, I'll define the my
- 2:51:59validation size.
- 2:52:01validation_size.
- 2:52:04I'll define it as 0.20 and I'll use a
- 2:52:07seed.
- 2:52:08I'll define seed equals six.
- 2:52:11So, this method seed sets the integer
- 2:52:13starting value used in generating random
- 2:52:15number. Okay? I'll define the value of
- 2:52:17seed equals six. I'll tell you what is
- 2:52:19the importance of it later on. Okay? So,
- 2:52:21let me define first few variables such
- 2:52:23as X_train, test, Y_train,
- 2:52:27and Y_test.
- 2:52:30Okay? So, what we want to do is select
- 2:52:32some model. Okay. So, model underscore
- 2:52:34selection. But, before doing that, what
- 2:52:36we have to do is split our training data
- 2:52:37set into two halves. Okay. So, dot train
- 2:52:39underscore test underscore split. What
- 2:52:42we want to split is the value of X and
- 2:52:45Y. Okay. And my test size is
- 2:52:49equals to validation size.
- 2:52:52Which is a 0.20. Correct? And my random
- 2:52:55state
- 2:52:57is equal to seed. So, what the seed is
- 2:52:59doing here, it's helping me to keep the
- 2:53:01same randomness in the training and
- 2:53:03testing data set. Fine. So, let's
- 2:53:05execute it and see what is our result.
- 2:53:08Let's execute it. Next, we'll create a
- 2:53:10test harness. For this, we'll use
- 2:53:1210-fold cross-validation to estimate the
- 2:53:14accuracy.
- 2:53:16So, what it will do, it will split our
- 2:53:17data set into 10 parts. Train on the
- 2:53:20nine part and test on the one part. And
- 2:53:22this will repeat for all combination of
- 2:53:24train and test splits. Okay. So, for
- 2:53:26that, let's define again
- 2:53:29my seed that was six, already defined,
- 2:53:32and scoring
- 2:53:34equals accuracy.
- 2:53:37Fine.
- 2:53:37So, we are using the metric of accuracy
- 2:53:39to evaluate the model. So, what is this?
- 2:53:42This is a ratio of number of correctly
- 2:53:44predicted instances divided by the total
- 2:53:46number of instances in the data set
- 2:53:48multiplied by 100, giving a percentage.
- 2:53:50Example, it's 98% accurate or 99%
- 2:53:54accurate, things like that. Okay. So,
- 2:53:55we'll be using the scoring variable when
- 2:53:57we run the build and evaluate each model
- 2:54:00in the next step. So, next part is
- 2:54:02building model.
- 2:54:04Till now, we don't know which algorithm
- 2:54:05would be good for this problem or what
- 2:54:07configuration to use. So, let's begin
- 2:54:09with six different algorithm. I'll be
- 2:54:11using logistic regression, linear
- 2:54:13discriminant analysis, K-nearest
- 2:54:15neighbor, classification and regression
- 2:54:17trees, Naive Bayes, and support vector
- 2:54:19machine. Well, these algorithms which
- 2:54:21I'm using is a good mixture of simple
- 2:54:23linear or non-linear algorithms. In
- 2:54:25simple linear which included the
- 2:54:26logistic regression and the linear
- 2:54:28discriminant analysis or the non-linear
- 2:54:30part which included the KNN algorithm,
- 2:54:32the CART algorithm, the Naive Bayes, and
- 2:54:34the support vector machines. Okay. So,
- 2:54:36we reset the random number seed before
- 2:54:38each run to ensure that evaluation of
- 2:54:40each algorithm is performed using
- 2:54:42exactly the same data splits. It ensures
- 2:54:44the result are directly comparable.
- 2:54:46Okay. So, let me just copy and paste it.
- 2:54:49Okay.
- 2:54:53So, what we're doing here, we're
- 2:54:54building five different types of model.
- 2:54:56We're building a logistic regression,
- 2:54:58linear discriminant analysis, K-nearest
- 2:55:00neighbor, decision tree, Gaussian Naive
- 2:55:02Bayes, and the support vector machine.
- 2:55:04Okay. Next, what we'll do, we'll
- 2:55:05evaluate model in each turn. Okay.
- 2:55:08So, what is this? So, we have six
- 2:55:10different model and accuracy estimation
- 2:55:12for each one of them. Now, we need to
- 2:55:14compare the model to each other and
- 2:55:15select the most accurate of them all.
- 2:55:17So, running this script, we saw the
- 2:55:19following result. So, we can see some of
- 2:55:21the result on the screen. What is this?
- 2:55:23It is just the accuracy score using
- 2:55:24different set of algorithms. Okay. When
- 2:55:27we are using logistic regression, what
- 2:55:28is the accuracy rate? When we are using
- 2:55:30linear discriminant algorithm, what is
- 2:55:32the accuracy? And so on and so. Okay.
- 2:55:34So, from the output, it seems that LD
- 2:55:36algorithm was the most accurate model
- 2:55:38that we tested. Now, we want to get an
- 2:55:40idea of the accuracy of the model on our
- 2:55:42validation set or the testing data set.
- 2:55:44So, this will give us a independent
- 2:55:46final check on the accuracy of the best
- 2:55:47model. It is always valuable to keep a
- 2:55:50testing data set for just in case you
- 2:55:52made a over-fitting to the testing data
- 2:55:54set or you made a data leak. Both will
- 2:55:56result in a overly optimistic result.
- 2:55:58Okay.
- 2:55:59You can run the LD model directly on the
- 2:56:01validation set and summarize the result
- 2:56:03as a final score, a confusion matrix,
- 2:56:06and a classification report.
- 2:56:09>> [music]
- 2:56:13>> Let us understand what regression in
- 2:56:15machine learning is.
- 2:56:16So, what exactly is regression?
- 2:56:18The main goal of regression is the
- 2:56:20construction of an efficient model to
- 2:56:22predict the dependent attributes from a
- 2:56:24bunch of attribute variables.
- 2:56:26A regression problem is where the output
- 2:56:28variable is either real or a continuous
- 2:56:30value like salary, weight, area, etc.
- 2:56:33We can also define regression as a
- 2:56:35statistical means that is used in
- 2:56:36applications like housing, investing,
- 2:56:38etc. to predict the relationship between
- 2:56:40a dependent variable and a bunch of
- 2:56:42independent variables.
- 2:56:44For example, let's say in the finance
- 2:56:46application or investing, we can
- 2:56:48actually predict the values of certain
- 2:56:50stock prices or you know those values
- 2:56:52depending on the independent variables
- 2:56:55like how many years it takes for a stock
- 2:56:57to you know actually mature or how many
- 2:56:59days will it take to grow or those
- 2:57:01variables that you have in investing and
- 2:57:04depending upon that we can make a
- 2:57:05possible outcome or a possible
- 2:57:06prediction of how a stock is going to be
- 2:57:09invested in a profit state or a loss
- 2:57:11state or all those things or we can take
- 2:57:13another example like housing. We can
- 2:57:15take different parameters like number of
- 2:57:17years it's been there, how many people
- 2:57:19have used it or what is the area of the
- 2:57:22house depending on all these factors or
- 2:57:24how many rooms does the house have, we
- 2:57:26can predict the price of a house.
- 2:57:28So this is basically what regression
- 2:57:30really is.
- 2:57:31So let us take a look at the various
- 2:57:32types of regression techniques that we
- 2:57:34have.
- 2:57:35We have simple linear regression, then
- 2:57:36we have polynomial regression, support
- 2:57:38vector regression, decision tree
- 2:57:40regression, we have random forest
- 2:57:42regression and we have logistic
- 2:57:43regression as well. That is also a type
- 2:57:45of regression that we have. But for now
- 2:57:47we'll be focusing on simple linear
- 2:57:49regression.
- 2:57:50So let's talk about how or what exactly
- 2:57:52is simple linear regression first. So
- 2:57:54one of the most interesting and common
- 2:57:56regression technique is a simple linear
- 2:57:57regression. In this we predict the
- 2:57:59outcome of a dependent variable Y based
- 2:58:02on the independent variables X. So the
- 2:58:04relationship between the variables is
- 2:58:06linear, hence the word linear
- 2:58:08regression.
- 2:58:09Then comes the polynomial regression.
- 2:58:11So in this regression technique, we
- 2:58:13transform the original features into a
- 2:58:15polynomial feature of a given degree and
- 2:58:17then perform regression on it. So, this
- 2:58:19is basically polynomial regression.
- 2:58:22After this, we have support vector
- 2:58:23machine regression or we can also call
- 2:58:25it SVR. We identify a hyperplane with
- 2:58:28maximum margin such that the maximum
- 2:58:31number of data points are within those
- 2:58:33margins.
- 2:58:34It is also quite similar to the support
- 2:58:35vector machine classification algorithm.
- 2:58:38Then we have decision tree regression.
- 2:58:41A decision tree can be used for both
- 2:58:42regression and classification. But, in
- 2:58:44this case of regression, we use the ID3
- 2:58:47algorithm, which is iterative
- 2:58:48dichotomizer 3, to identify the
- 2:58:51splitting node by reducing the standard
- 2:58:53deviation.
- 2:58:54After this, we have a random forest
- 2:58:55regression, which is basically an
- 2:58:57ensemble of predictions of several
- 2:58:59decision tree regressions.
- 2:59:01So, this is all about the types of
- 2:59:02regressions for now. We're going to
- 2:59:04focus on simple linear regression.
- 2:59:06So, let's take a look at what exactly is
- 2:59:08a simple linear regression.
- 2:59:10Simple linear regression is a regression
- 2:59:12technique in which the independent
- 2:59:14variable has a linear relationship with
- 2:59:16the dependent variable.
- 2:59:18The straight line in the diagram is the
- 2:59:19best fit line, and the main goal of the
- 2:59:21simple linear regression is to consider
- 2:59:24the given data points and plot the best
- 2:59:25fit line to fit the model in the best
- 2:59:27way possible.
- 2:59:28So, if you talk about a real-life
- 2:59:30analogy to explain linear regression, we
- 2:59:32can take an example of a car resale
- 2:59:34value. So, we have different parameters,
- 2:59:36you know, when we are talking about
- 2:59:37resale value of a car. Like how many
- 2:59:40years the car has been there in the
- 2:59:41market, and how many kilometers it has
- 2:59:44been driven,
- 2:59:45the kind of mileage the car gives, and
- 2:59:47then we have different parameters we can
- 2:59:49focus upon. And all these independent
- 2:59:51variables somehow are linearly connected
- 2:59:53or interconnected to the price of the
- 2:59:55car.
- 2:59:56So, that is one example to understand
- 2:59:58linear regression. We'll be doing that
- 2:59:59in the use case. I'll be telling you
- 3:00:01about how you can predict the price of
- 3:00:02car.
- 3:00:03Now, talking about linear regression
- 3:00:05terminologies, there are a few
- 3:00:07terminologies that you have to be
- 3:00:08thorough with to begin with linear
- 3:00:10regression.
- 3:00:11So, first of all, we have to talk about
- 3:00:13cost function.
- 3:00:14So, the best fit line can be based on
- 3:00:16the linear equation that is given here.
- 3:00:18So, in this, the dependent variable that
- 3:00:20is to be predicted is denoted by Y.
- 3:00:22A line that touches the Y axis is
- 3:00:24denoted by the intercept B0. The B1 is
- 3:00:27the slope of the line, and X represents
- 3:00:29the independent variables that determine
- 3:00:31the prediction of Y.
- 3:00:33The error in the resultant prediction is
- 3:00:35denoted by E.
- 3:00:36Now, talking about cost function, the
- 3:00:38cost function provides the best possible
- 3:00:40values for B0 and B1 to make the best
- 3:00:43fit line for the data points.
- 3:00:45We do this by converting this problem
- 3:00:46into a minimization problem to get the
- 3:00:49best values for B0 and B1.
- 3:00:51So, with this, the error is minimized in
- 3:00:53this problem between the actual value
- 3:00:55and the predicted value, and we choose
- 3:00:57the function above to minimize.
- 3:00:59Now, we square the error difference and
- 3:01:01sum the error over all the data points.
- 3:01:04The division between the total number of
- 3:01:05data points and the produced value
- 3:01:07provides the average square error for
- 3:01:09all the data points.
- 3:01:11It is also known as mean squared error,
- 3:01:13and we can change the values of B0 and
- 3:01:15B1 so that the MSE or the mean squared
- 3:01:17error value is settled at the minimum.
- 3:01:20So, this is one terminology that is cost
- 3:01:22function that we use in linear
- 3:01:23regression.
- 3:01:24Then, we have the gradient descent.
- 3:01:27So, the next important terminology to
- 3:01:28understand linear regression is gradient
- 3:01:30descent, of course, and it is a method
- 3:01:32of updating B0 and B1 value to reduce
- 3:01:35the MSE, which is the mean squared
- 3:01:36error.
- 3:01:37The idea behind this is to keep
- 3:01:39iterating the B0 and B1 values until we
- 3:01:41reduce the MSE to the minimum.
- 3:01:44Now, to update B0 and B1, we take the
- 3:01:45gradients from the cost function, and to
- 3:01:48find these gradients, we take partial
- 3:01:50derivatives with respect to B0 and B1.
- 3:01:53And these partial derivatives are the
- 3:01:55gradients and are used to update the
- 3:01:57values of B0 and B1.
- 3:01:59I'm sure guys, this is might be a little
- 3:02:00confusing for you guys if you are new to
- 3:02:03this, like gradient descent and cost
- 3:02:04function, but you don't have to worry
- 3:02:06about this because in Python when we're
- 3:02:08using linear regression, we're going to
- 3:02:09be using the scikit-learn or the
- 3:02:11scikit-learn library, so you don't have
- 3:02:12to worry about this. You just have to
- 3:02:14integrate your model with the linear
- 3:02:15regression model that we have already
- 3:02:17over there, and you'll be done with it.
- 3:02:19And when I'm implementing the linear
- 3:02:21regression model, you'll see how easy it
- 3:02:23is to actually implement linear
- 3:02:25regression in Python.
- 3:02:26So, after this, let's talk about a few
- 3:02:28advantages and disadvantages of linear
- 3:02:30regression.
- 3:02:32So, talking about the advantages first,
- 3:02:34linear regression performs exceptionally
- 3:02:36well for linearly separable data. And it
- 3:02:38is actually very easy to implement,
- 3:02:40interpret, and very efficient to train
- 3:02:43as well.
- 3:02:44And even though the linear regression is
- 3:02:45prone to overfitting, it handles it
- 3:02:48pretty well using dimension reduction
- 3:02:49techniques, regularization, and
- 3:02:51cross-validation. And one more advantage
- 3:02:54is that the extrapolation beyond a
- 3:02:56specific data set.
- 3:02:58So, these are all the advantages that we
- 3:02:59have with linear regression. Let's talk
- 3:03:01about a few disadvantages as well.
- 3:03:03So, one of the most common disadvantage
- 3:03:05with linear regression is that it takes
- 3:03:07the assumption of linearity between
- 3:03:09dependent and independent variables. The
- 3:03:11next disadvantage is it is often very
- 3:03:14prone to noise and overfitting as well,
- 3:03:16which is not a very good sign for any
- 3:03:18model if you're doing regression or
- 3:03:19classification in machine learning.
- 3:03:22The next disadvantage is it is very
- 3:03:24quite sensitive to outliers as well.
- 3:03:27And the last one is that it is very
- 3:03:28prone to multicollinearity.
- 3:03:30So, these are all the advantages and
- 3:03:32disadvantages of linear regression.
- 3:03:35>> [music]
- 3:03:39>> So, let's understand the what and why of
- 3:03:41logistic regression. Now, this algorithm
- 3:03:44is most widely used when the dependent
- 3:03:46variable, or you can say the output, is
- 3:03:47in the binary format. So, here you need
- 3:03:50to predict the outcome of a categorical
- 3:03:52dependent variable. So, the outcome
- 3:03:54should be always discrete or categorical
- 3:03:56in nature. Now, by discrete, I mean the
- 3:03:58value should be binary, or you can say
- 3:04:00you just have two values. It can either
- 3:04:02be zero or one. It can either be yes or
- 3:04:05a no. Either be true or false. Or high
- 3:04:07or low. So only these can be the
- 3:04:09outcomes. So the value which you need to
- 3:04:12predict should be discrete or you can
- 3:04:13say categorical in nature. Whereas in
- 3:04:16linear regression we have the value of Y
- 3:04:18or you can say the value you need to
- 3:04:19predict is in a range. So that is how
- 3:04:21there's a difference between linear
- 3:04:22regression and logistic regression. Now
- 3:04:24you must be having a question, why not
- 3:04:26linear regression? Now guys, in linear
- 3:04:28regression the value of Y or the value
- 3:04:30which you need to predict is in a range.
- 3:04:32But in our case, as in the logistic
- 3:04:34regression, we just have two values. It
- 3:04:36can be either zero or it can be one. It
- 3:04:39should not entertain the values which is
- 3:04:40below zero or above one. But in linear
- 3:04:43regression we have the value of Y in the
- 3:04:45range. So here, in order to implement
- 3:04:47logistic regression, we need to clip
- 3:04:48this part. So we don't need the value
- 3:04:51that is below zero or we don't need the
- 3:04:52value which is above one. So since the
- 3:04:54value of Y will be between only zero and
- 3:04:57one, that is the main rule of logistic
- 3:04:58regression, the linear line has to be
- 3:05:00clipped at zero and one. Now once we
- 3:05:02clip this graph, it would look somewhat
- 3:05:04like this. So here you're getting a
- 3:05:06curve which is nothing but three
- 3:05:07different straight lines. So here we
- 3:05:09need to make a new way to solve this
- 3:05:11problem. So this has to be formulated
- 3:05:13into equation and hence we come up with
- 3:05:15logistic regression. So here the outcome
- 3:05:17is either zero or one, which is the main
- 3:05:20rule of logistic regression. So with
- 3:05:21this our resulting curve cannot be
- 3:05:23formulated. So hence our main aim to
- 3:05:25bring the values to zero and one is
- 3:05:26fulfilled. So that is how we came up
- 3:05:28with logistic regression. Now here, once
- 3:05:31it gets formulated into an equation, it
- 3:05:33looks somewhat like this.
- 3:05:35So guys, this is nothing but a S curve
- 3:05:36or you can say the sigmoid curve or
- 3:05:38sigmoid function curve. So this sigmoid
- 3:05:41function basically converts any value
- 3:05:43from minus infinity to infinity to your
- 3:05:45discrete values which a logistic
- 3:05:47regression wants or you can say the
- 3:05:48values which are in binary format,
- 3:05:50either zero or one. So if you see here
- 3:05:53the values are either zero or one. And
- 3:05:55this is nothing but just a transition of
- 3:05:57it. But guys, there's a catch over here.
- 3:05:59So, let's say I have a data point that
- 3:06:01is 0.8. Now, how can you decide whether
- 3:06:04your value is zero or one? Now, here you
- 3:06:07have the concept of threshold, which
- 3:06:09basically divides your line. So, here
- 3:06:11threshold value basically indicates the
- 3:06:13probability of either winning or losing.
- 3:06:16So, here by winning I mean the values
- 3:06:18equals to one, and by losing I mean the
- 3:06:20values equals to zero. But how does it
- 3:06:22do that? Let's say I have data point
- 3:06:24which is over here. Let's say my cursor
- 3:06:26is at 0.8. So, here I'll check whether
- 3:06:28this value is less than my threshold
- 3:06:30value or not. Let's say if it is more
- 3:06:33than my threshold value, it should give
- 3:06:34me the result as one. If it is less than
- 3:06:36that, then it should give me the result
- 3:06:38as zero. So, here my threshold value is
- 3:06:400.5. Now, I need to define that if my
- 3:06:43value, let's say 0.8, it is more than
- 3:06:450.5, then the value shall be rounded off
- 3:06:48to one. And let's say if it is less than
- 3:06:500.5, let's say I have a value 0.2, then
- 3:06:52it should reduce it to zero. So, here
- 3:06:55you can use the concept of threshold
- 3:06:56value to find the output. So, here it
- 3:06:59should be discrete, it should be either
- 3:07:00zero or it should be one.
- 3:07:02So, I hope you caught this curve of
- 3:07:03logistic regression. So, the guys, this
- 3:07:05is the sigmoid S curve.
- 3:07:08So, to make this curve, we need to make
- 3:07:10an equation. So, let me address that
- 3:07:11part as well.
- 3:07:13So, let's see how an equation is formed
- 3:07:14to imitate this functionality. So, over
- 3:07:17here we have an equation of a straight
- 3:07:18line, which is Y is equals to MX + C.
- 3:07:21So, in this case, I just have only one
- 3:07:23independent variable. But let's say if
- 3:07:25we have many independent variable, then
- 3:07:27the equation becomes M1 X1 + M2 X2 + M3
- 3:07:30X3 and so on till MN XN. Now, let us put
- 3:07:34in B and X. So, here the equation
- 3:07:36becomes Y is equals to B1 X1 + B2 X2 +
- 3:07:39B3 X3 and so on till BN XN + C.
- 3:07:44So, guys, the equation of the straight
- 3:07:45line has a range from minus infinity to
- 3:07:47infinity. But in our case, or you can
- 3:07:50say in logistic equation, the value
- 3:07:52which we need to predict or you can say
- 3:07:53the Y value, it can have the range only
- 3:07:55from zero to one. So, in that case, we
- 3:07:57need to transform this equation. So, to
- 3:08:00do that, what we had done, we had just
- 3:08:02divide the equation by 1 - Y. So, now
- 3:08:04when Y is equals to zero, so zero over 1
- 3:08:07- 0 which is equals to 1. So, zero over
- 3:08:091 is again zero. And if we take Y is
- 3:08:12equals to 1, then 1 over 1 - 1 which is
- 3:08:15zero. So, 1 over zero is infinity. So,
- 3:08:17here my range is now between zero to
- 3:08:19infinity. But, again we want the range
- 3:08:21from minus infinity to infinity. So, for
- 3:08:24that, what we'll do, we'll have the log
- 3:08:25of this equation. So, let's go ahead and
- 3:08:27have the logarithmic of this equation.
- 3:08:29So, here we have just transform it
- 3:08:31further to get the range between minus
- 3:08:33infinity to infinity. So, over here we
- 3:08:35have log of Y over 1 - 1 and this is
- 3:08:38your final logistic regression equation.
- 3:08:40So, guys, don't worry, you don't have to
- 3:08:42write this formula or memorize this
- 3:08:44formula. In Python, you just need to
- 3:08:46call this function which is logistic
- 3:08:47regression and everything will be
- 3:08:49automatically for you. So, I don't want
- 3:08:51to scare you with the maths and the
- 3:08:52formulas behind it, but it's always good
- 3:08:54to know how the formula was generated.
- 3:08:57Moving ahead, let us see the various use
- 3:08:58cases wherein logistic regression is
- 3:09:00implemented in real life.
- 3:09:03So, the very first is weather
- 3:09:04prediction.
- 3:09:05Now, logistic regression helps you to
- 3:09:06predict your weather. For example, it is
- 3:09:09used to predict whether it is raining or
- 3:09:10not, whether it is sunny, is it cloudy
- 3:09:13or not. So, all these things can be
- 3:09:15predicted using logistic regression.
- 3:09:17Whereas, you need to keep in mind that
- 3:09:19both linear regression and logistic
- 3:09:20regression can be used in predicting
- 3:09:22weather. So, in that case, linear
- 3:09:24regression helps you to predict what
- 3:09:25will be the temperature tomorrow.
- 3:09:27Whereas, logistic regression will only
- 3:09:29tell you whether it's going to rain or
- 3:09:30not or whether it's cloudy or not,
- 3:09:32whether it's going to snow or not. So,
- 3:09:34these values are discrete. Whereas, if
- 3:09:36you apply linear regression, you're
- 3:09:37predicting things like what is the
- 3:09:39temperature tomorrow or what is the
- 3:09:41temperature day after tomorrow and all
- 3:09:43those things. So, these are the slight
- 3:09:44differences between linear regression
- 3:09:46and logistic regression. Now moving
- 3:09:47ahead, we have classification problem.
- 3:09:50So Python performs multi-class
- 3:09:51classification. So here it can help you
- 3:09:53tell whether it's a bird or it's not a
- 3:09:55bird. Then you classify different kind
- 3:09:57of mammals. Let's say whether it's a dog
- 3:09:59or it's not a dog. Similarly, you can
- 3:10:01check it for reptile whether it's a
- 3:10:03reptile or not a reptile. So in logistic
- 3:10:05regression, it can perform multi-class
- 3:10:07classification. So this point I've
- 3:10:09already discussed that it is used in
- 3:10:10classification problems. Next, it also
- 3:10:13helps you to determine the illness as
- 3:10:14well. So let me take an example. Let's
- 3:10:17say a patient goes for a routine checkup
- 3:10:19in hospital. So what doctor will do it
- 3:10:21it will perform various tests on the
- 3:10:22patient and will check whether the
- 3:10:24patient is actually ill or not. So what
- 3:10:26will be the features? So doctor can
- 3:10:29check the sugar level, the blood
- 3:10:30pressure, then what is the age of the
- 3:10:32patient? Is it very small or is it a old
- 3:10:34person? Then what is the previous
- 3:10:36medical history of the patient? And all
- 3:10:38of these features will be recorded by
- 3:10:40the doctor. And finally, doctor checks
- 3:10:42the patient data and determines the
- 3:10:44outcome of the illness and the severity
- 3:10:46of illness. So using all the data, a
- 3:10:48doctor can identify whether a patient is
- 3:10:51ill or not. So these are the various use
- 3:10:53cases in which you can use logistic
- 3:10:54regression. Now I guess enough of theory
- 3:10:57part, so let's move ahead and see some
- 3:10:59of the practical implementation of
- 3:11:00logistic regression.
- 3:11:02So over here I'll be implementing two
- 3:11:04projects wherein I have the data set of
- 3:11:06a Titanic. So over here we'll predict
- 3:11:08what factors made people more likely to
- 3:11:10survive the sinking of the Titanic ship.
- 3:11:12And in my second project, we'll see the
- 3:11:14data analysis on the SUV cars. So over
- 3:11:16here we have the data of the SUV cars,
- 3:11:18who can purchase it, and what factors
- 3:11:21made people more interested in buying
- 3:11:23SUV.
- 3:11:24So these will be the major questions as
- 3:11:25to why you should implement logistic
- 3:11:27regression and what output will you get
- 3:11:29by it. So let's start by the very first
- 3:11:31project that is Titanic data analysis.
- 3:11:33So some of you might know that there was
- 3:11:35a ship called as Titanic which basically
- 3:11:37hit an iceberg and it sank to the bottom
- 3:11:39of the ocean. And it was a big disaster
- 3:11:42at that time because it was the first
- 3:11:44voyage of the ship and it was supposed
- 3:11:45to be really, really strongly built and
- 3:11:47one of the best ships of that time. So,
- 3:11:49it was a big disaster of that time and
- 3:11:51of course there's a movie about this as
- 3:11:53well. So, many of you might have watched
- 3:11:55it. So, what we have we have data of the
- 3:11:57passengers, those who survived and those
- 3:11:59who did not survive in this particular
- 3:12:00tragedy. So, what you have to do you
- 3:12:02have to look at this data and analyze
- 3:12:04which factors would have been
- 3:12:05contributed the most to the chances of a
- 3:12:08person's survival on the ship or not.
- 3:12:10So, using the logistic regression we can
- 3:12:12predict whether the person survived or
- 3:12:14the person died. Now, apart from this we
- 3:12:16also have a look with the various
- 3:12:17features along with that. So, first let
- 3:12:19us explore the data set. So, over here
- 3:12:21we have the index value. Then the first
- 3:12:24column is passenger ID. Then my next
- 3:12:26column is survived. So, over here we
- 3:12:28have two values, a zero and a one. So,
- 3:12:31zero stands for did not survive and one
- 3:12:33stands for survived. So, this column is
- 3:12:35categorical where the values are
- 3:12:37discrete. Next we have passenger class.
- 3:12:39So, over here we have three values, one,
- 3:12:41two, and three. So, this basically tells
- 3:12:43you that whether a passenger is
- 3:12:45traveling in the first class, second
- 3:12:47class, or third class. Then we have the
- 3:12:49name of the passenger, we have the sex
- 3:12:51or you can say the gender of the
- 3:12:52passenger, whether passenger is a male
- 3:12:54or female. Then we have the age, we have
- 3:12:56the sib SP. So, this basically means the
- 3:12:59number of siblings or the spouses aboard
- 3:13:01the Titanic. So, over here we have
- 3:13:03values such as 1, 0, and so on. Then we
- 3:13:06have parch. So, parch is basically the
- 3:13:09number of parents or children aboard the
- 3:13:11Titanic. So, over here we also have some
- 3:13:13values.
- 3:13:14Then we have the ticket number, we have
- 3:13:16the fare, we have the cabin number, and
- 3:13:18we have the embarked column. So, in my
- 3:13:20embarked column we have three values, we
- 3:13:22have S, C, and Q. So, S basically stands
- 3:13:25for Southampton, C stands for Cherbourg,
- 3:13:27and Q stands for Queenstown.
- 3:13:30So, these are the features that we'll be
- 3:13:31applying our model on. So, here we'll
- 3:13:33perform various steps and then we'll be
- 3:13:35implementing logistic regression. So,
- 3:13:37now these are the various steps which
- 3:13:39are required to implement any algorithm.
- 3:13:41So now in our case we are implementing
- 3:13:43logistic regression. So very first step
- 3:13:45is to collect your data or to import the
- 3:13:47libraries that are used for collecting
- 3:13:49your data and then taking it forward.
- 3:13:51Then my second step is to analyze your
- 3:13:53data. So over here I can go through the
- 3:13:55various fields and then I can analyze
- 3:13:57the data. I can check did the females or
- 3:13:59children survive better than the males
- 3:14:01or did the rich passengers survive more
- 3:14:03than the poor passenger or did the money
- 3:14:05matter as in who paid more to get into
- 3:14:08the ship were they evacuated first and
- 3:14:10what about the workers? Does the worker
- 3:14:12survive or what is the survival rate if
- 3:14:15you were the worker in the ship and not
- 3:14:16just a traveling passenger? So all of
- 3:14:18these are very very interesting
- 3:14:20questions and you would be going through
- 3:14:21all of them one by one. So in this stage
- 3:14:24you need to analyze your data and
- 3:14:25explore your data as much as you can.
- 3:14:28Then my third step is to wrangle your
- 3:14:29data. Now data wrangling basically means
- 3:14:32cleaning your data. So over here you can
- 3:14:34simply remove the unnecessary items or
- 3:14:36if you have a null values in the data
- 3:14:38set you can just clear that data and
- 3:14:40then you can take it forward. So in this
- 3:14:42step you can build your model using the
- 3:14:44train data set and then you can test it
- 3:14:46using the test. So over here you will be
- 3:14:48performing a split which basically split
- 3:14:50your data set into training and testing
- 3:14:52data set and finally you will check the
- 3:14:54accuracy so as to ensure how much
- 3:14:56accurate your values are. So I hope you
- 3:14:58guys got these five steps that we're
- 3:15:00going to implement in logistic
- 3:15:01regression. So now let's go into all
- 3:15:03these steps in detail. So number one we
- 3:15:05have to collect your data or you can say
- 3:15:07import the libraries. So let me show you
- 3:15:09the implementation part as well. So I'll
- 3:15:11just open my Jupiter notebook and I'll
- 3:15:13just implement all of these steps side
- 3:15:15by side.
- 3:15:17So guys this is my Jupiter notebook. So
- 3:15:19first let me just rename Jupiter
- 3:15:21notebook to let's say Titanic data
- 3:15:23analysis.
- 3:15:27Now our first step was to import all the
- 3:15:29libraries and collect the data. So let
- 3:15:31me just import all the libraries first.
- 3:15:33So, first of all, I'll import pandas.
- 3:15:35So, pandas is used for data analysis.
- 3:15:38So, I'll say import pandas as pd. Then,
- 3:15:40I'll be importing NumPy. So, I'll say
- 3:15:42import NumPy as np. So, NumPy is a
- 3:15:45library in Python which basically stands
- 3:15:47for numerical Python. And it is widely
- 3:15:49used to perform any scientific
- 3:15:51computation. Next, we'll be importing
- 3:15:53seaborn. So, seaborn is a library for
- 3:15:55statistical plotting. So, I'll say
- 3:15:57import seaborn as sns. I'll also import
- 3:16:00matplotlib.
- 3:16:01So, matplotlib library is again for
- 3:16:03plotting. So, I'll say import
- 3:16:05matplotlib.pyplot
- 3:16:07as pld.
- 3:16:09Now, to run this library in Jupyter
- 3:16:10Notebook, all I have to write in is
- 3:16:12percentage matplotlib inline.
- 3:16:15Next, I'll be importing one module as
- 3:16:17well. So, as to calculate the basic
- 3:16:20mathematical functions. So, I'll say
- 3:16:22import maths. So, these are the
- 3:16:23libraries that I'll be needing in this
- 3:16:25Titanic data analysis. So, now let me
- 3:16:27just import my dataset. So, I'll take a
- 3:16:29variable, let's say Titanic data. And
- 3:16:32using the pandas, I will just read my
- 3:16:34CSV. Or you can say the dataset.
- 3:16:37I'll write the name of my dataset, that
- 3:16:38is titanic.csv.
- 3:16:40Now, I have already showed you the
- 3:16:42dataset. So, over here, let me just
- 3:16:43print the top 10 rows. So, for that,
- 3:16:45I'll just say I'll take the variable
- 3:16:47Titanic data. head and I'll say the top
- 3:16:5010 rows. Now, I'll just run this. So, to
- 3:16:52run this, I just have to press shift
- 3:16:54plus enter. Or else, you can just
- 3:16:56directly click on the cell.
- 3:16:58So, over here, I have the index. We have
- 3:17:00the passenger ID, which is nothing but
- 3:17:02again the index which is starting from
- 3:17:03one. Then, we have the survived column
- 3:17:05which has the categorical values or you
- 3:17:07can say the discrete values, which is in
- 3:17:09the form of zero or one. Then, we have
- 3:17:11the passenger class. We have the name of
- 3:17:13the passenger, sex, age, and so on. So,
- 3:17:15this is the dataset that I'll be going
- 3:17:17forward with. Next, let us print the
- 3:17:18number of passengers which are there in
- 3:17:20this original dataset. So, for that,
- 3:17:22I'll just simply type in print. I'll say
- 3:17:25number of passengers.
- 3:17:31And using the length function, I can
- 3:17:32calculate the total length. So, I'll say
- 3:17:34length and inside this I'll be passing
- 3:17:36this variable which is Titanic data. So,
- 3:17:38I'll just copy it from here. I'll just
- 3:17:40paste it {dot} index.
- 3:17:42And next, let me just print this one.
- 3:17:45So, here the number of passengers which
- 3:17:46are there in the original data set we
- 3:17:48have is 891. So, around this number were
- 3:17:51traveling in the Titanic ship. So, over
- 3:17:54here my first step is done. We have just
- 3:17:56collected data, imported all the
- 3:17:57libraries, and find out the total number
- 3:17:59of passengers which are traveling in
- 3:18:01Titanic. So, let me just go back to
- 3:18:03presentation and let's see what is my
- 3:18:04next step.
- 3:18:05So, we're done with the collecting data.
- 3:18:07Next step is to analyze your data. So,
- 3:18:09over here we'll be creating different
- 3:18:11plots to check the relationship between
- 3:18:13variables as in how one variable is
- 3:18:15affecting the other. So, you can simply
- 3:18:17explore your data set by making use of
- 3:18:19various columns and then you can plot a
- 3:18:21graph between them. So, you can either
- 3:18:23plot a correlation graph, you can plot a
- 3:18:25distribution graph. It's up to you guys.
- 3:18:27So, let me just go back to my Jupiter
- 3:18:29notebook and let me analyze some of the
- 3:18:30data. Over here my second part is to
- 3:18:32analyze data. So, I'll just put this in
- 3:18:34header two.
- 3:18:36Now, to put this in header two, I just
- 3:18:37have to go on code, click on markdown,
- 3:18:39and I'll just run this.
- 3:18:41So, first let us plot a count plot where
- 3:18:43you can compare between the passengers
- 3:18:44who survived and who did not survive.
- 3:18:46So, for that I'll be using the seaborn
- 3:18:47library. So, over here I have imported
- 3:18:50seaborn as sns. So, I don't have to
- 3:18:52write the whole name. I'll simply say
- 3:18:53sns.count plot.
- 3:18:58I'll say x is equal to survive and the
- 3:19:00data that I'll be using is the Titanic
- 3:19:01data. Or you can say the name of
- 3:19:03variable in which you have stored your
- 3:19:04data set. So, now let me just run this.
- 3:19:07So, over here as you can see I have
- 3:19:09survived column on my x-axis and on the
- 3:19:11y-axis I have the count. So, zero
- 3:19:13basically stands for did not survive and
- 3:19:15one stands for the passengers who did
- 3:19:17survive. So, over here you can see that
- 3:19:19around 550 of the passengers who did not
- 3:19:22survive and there were around 350
- 3:19:24passengers who only survived. So here
- 3:19:26you can basically conclude that there
- 3:19:28are very less survivors than
- 3:19:29non-survivors. So this was the very
- 3:19:32first plot. Now let us plot another plot
- 3:19:34to compare the sex as to whether out of
- 3:19:36all the passengers who survived and who
- 3:19:38did not survive, how many were men and
- 3:19:40how many were female. So to do that I'll
- 3:19:42simply say sns.countplot.
- 3:19:47I'll add the hue as sex.
- 3:19:49So I want to know how many females and
- 3:19:51how many males survived.
- 3:19:53Then I'll be specifying the data. So I'm
- 3:19:54using Titanic data set.
- 3:19:57And let me just run this.
- 3:19:59Okay, I've done a mistake over here.
- 3:20:01So over here you can see I have survived
- 3:20:02column on the x-axis and I have the
- 3:20:04count on the y.
- 3:20:06Now so here your blue color stands for
- 3:20:07your male passengers and orange stands
- 3:20:09for your female.
- 3:20:11So as you can see here the passengers
- 3:20:13who did not survive, that has a value
- 3:20:14zero.
- 3:20:15So we can see that majority of males did
- 3:20:18not survive. And if we see the people
- 3:20:20who survived, here we can see the
- 3:20:22majority of females survived. So this
- 3:20:24basically concludes the gender of the
- 3:20:25survival rate. So it appears on average
- 3:20:28women were more than three times more
- 3:20:30likely to survive than men. Next let us
- 3:20:32plot another plot where we have the hue
- 3:20:34as the passenger class. So over here we
- 3:20:36can see which class that the passenger
- 3:20:38was traveling in, whether it was
- 3:20:39traveling in class one, two or three.
- 3:20:42So for that I'll just write the same
- 3:20:44command. I'll say
- 3:20:45sns.countplot.
- 3:20:49I'll keep my x-axis as only. I'll change
- 3:20:52my hue to passenger class.
- 3:20:54So my variable named as pclass.
- 3:20:57And the data set that I'll be using is
- 3:20:58Titanic data. So this is my result. So
- 3:21:01over here you can see I have blue for
- 3:21:03first class, orange for second class and
- 3:21:05green for the third class.
- 3:21:07So here the passengers who did not
- 3:21:09survive were majorly of the third class
- 3:21:11or you can see the lowest class or the
- 3:21:12cheapest class to get into the Titanic.
- 3:21:15And the people who did survive majorly
- 3:21:16belong to the higher classes. So here
- 3:21:18one and two has more rise than the
- 3:21:20passenger who were traveling in the
- 3:21:21third class. So here we have concluded
- 3:21:24that the passengers who did not survive
- 3:21:26are majorly of third class or you can
- 3:21:27say the lowest class. And the passengers
- 3:21:30who were traveling in first and second
- 3:21:31class would tend to survive more. Next
- 3:21:34let us plot a graph for the age
- 3:21:35distribution. Over here I can simply use
- 3:21:37my data. So we'll be using pandas
- 3:21:39library for this. I'll declare a array
- 3:21:42and I'll pass in the column that is age.
- 3:21:45So I plot and I want a histogram so I'll
- 3:21:47say plot.hist.
- 3:21:51So you can notice over here that we have
- 3:21:53more of young passengers or you can see
- 3:21:55the children between the ages zero to
- 3:21:5710. And then we have the average age
- 3:21:59people. And if you go ahead lesser would
- 3:22:01be the population. So this is the
- 3:22:03analysis on the age column. So we saw
- 3:22:06that we have more young passengers and
- 3:22:08more mediocre age passengers who are
- 3:22:10traveling in the Titanic.
- 3:22:11So next let me plot a graph of fare as
- 3:22:13well. So I'll say Titanic data.
- 3:22:17I'll say fare.
- 3:22:18And again I'll plot a histogram so I'll
- 3:22:20say hist.
- 3:22:23So here you can see the fare size is
- 3:22:25between zero to 100. Now let me add the
- 3:22:27bin size so as to make it more clear.
- 3:22:30So over here I'll say bin is equals to
- 3:22:32let's say 20 and I'll increase the
- 3:22:34figure size as well. So I'll say fix
- 3:22:36size. Let's say I'll give the dimensions
- 3:22:39as 10 by 5.
- 3:22:41So it is bins. So this is more clear
- 3:22:43now. Next let us analyze the other
- 3:22:45columns as well.
- 3:22:47So I'll just type in Titanic data. And I
- 3:22:50want the information as to what all
- 3:22:51columns are left.
- 3:22:54So here we have passenger ID which I
- 3:22:56guess it's of no use. Then we have to
- 3:22:58see how many passengers survived and how
- 3:23:00many did not. We also see the analysis
- 3:23:02on the gender basis. We saw whether
- 3:23:03female tend to survive more or the men
- 3:23:05tend to survive more. Then we saw the
- 3:23:07passenger class where the passenger is
- 3:23:09traveling in in first class, second
- 3:23:10class or third class. Then we have the
- 3:23:12name, so in name we cannot do any
- 3:23:14analysis. We saw the sex, we saw the age
- 3:23:17as well. Then we have sibsp. So this
- 3:23:20stands for the number of siblings or the
- 3:23:22spouses which are aboard the Titanic. So
- 3:23:24let us do this as well. So I'll say
- 3:23:26sns.countplot.
- 3:23:30I'll mention x as sibsp.
- 3:23:33And I'll be using the Titanic data.
- 3:23:36So you can see the plot over here. So
- 3:23:38over here you can conclude that it has
- 3:23:40the maximum value on zero. So you can
- 3:23:42conclude that neither a children nor a
- 3:23:44spouse was on board the Titanic. The
- 3:23:46second most highest value is one. And
- 3:23:49then we have very less values for two,
- 3:23:51three, four, and so on.
- 3:23:53Next, if I go above, we saw this column
- 3:23:55as well. Similarly, you can do for
- 3:23:56parch.
- 3:23:57So next we have parch, or you can say
- 3:23:59the number of parents or children which
- 3:24:00were aboard the Titanic. So you
- 3:24:02similarly can do this as well. Then we
- 3:24:04have the ticket number. So I don't think
- 3:24:06so any analysis is required for ticket.
- 3:24:08Then we have fare. So fare we have
- 3:24:10already discussed as in the people who
- 3:24:12tend to travel in the first class
- 3:24:13usually pay the highest fare. Then we
- 3:24:15have the cabin number, and we have
- 3:24:17embarked. So these are the columns that
- 3:24:19we'll be doing data wrangling on.
- 3:24:21So we have analyzed the data, and we
- 3:24:22have seen quite a few graphs in which we
- 3:24:24can conclude which variable is better
- 3:24:27than the another or or what are the
- 3:24:28relationship they hold.
- 3:24:29So third step is my data wrangling. So
- 3:24:31data wrangling basically means cleaning
- 3:24:33your data. So if you have a large data
- 3:24:36set, you might be having some null
- 3:24:37values, or you can say NaN values. So
- 3:24:39it's very important that you remove all
- 3:24:41the unnecessary items that are present
- 3:24:43in your data set. So removing this
- 3:24:45directly affects your accuracy. So I'll
- 3:24:47just go ahead and clean my data by
- 3:24:49removing all the NaN values and
- 3:24:51unnecessary columns which has a null
- 3:24:52value in the data set. So next I'll be
- 3:24:54performing data wrangling.
- 3:25:02So first of all, I'll check whether my
- 3:25:03data set is null or not. So I'll say
- 3:25:06Titanic data, which is the name of my
- 3:25:07data set, and I'll say is null. So, this
- 3:25:10will basically tell me what all values
- 3:25:12are null, and it will return me a
- 3:25:13Boolean result. So, this basically
- 3:25:15checks the missing data, and your result
- 3:25:17will be in Boolean format, as in the
- 3:25:19result will be in true or false. So,
- 3:25:20false mean if it is not null, and true
- 3:25:23means if it is null. So, let me just run
- 3:25:25this.
- 3:25:27Over here, you can see the values as
- 3:25:29false or true. So, false is where the
- 3:25:31value is not null, and true is where the
- 3:25:34value is null. So, over here, you can
- 3:25:35see in the cabin column, we have the
- 3:25:37very first value, which is null. So, we
- 3:25:39have to do something on this.
- 3:25:41So, you can see that we have a large
- 3:25:43data set.
- 3:25:44The counting does not stop, and we can
- 3:25:47actually see the sum of it. We can
- 3:25:48actually print the number of passengers
- 3:25:50who have the NaN value in each column.
- 3:25:52So, I'll say Titanic_data
- 3:25:55is null, and I want the sum of it. So,
- 3:25:57I'll say dot sum. So, this will
- 3:25:59basically print the number of passengers
- 3:26:01who have the NaN values in each column.
- 3:26:03So, we can see that we have missing
- 3:26:05values in age column, that is 177. Then,
- 3:26:07we have the maximum value in the cabin
- 3:26:09column, and we have very less in the
- 3:26:11embarked column, that is two.
- 3:26:13So, here, if you don't want to see these
- 3:26:15numbers, you can also plot a heat map,
- 3:26:17and then you can visually analyze it.
- 3:26:19So, let me just do that as well. So,
- 3:26:20I'll say sns.heatmap
- 3:26:26and say yticklabels
- 3:26:31to false. So, I'll just run this. So, as
- 3:26:33we have already seen that there were
- 3:26:35three columns in which missing data
- 3:26:36value was present. So, this might be
- 3:26:38age. So, over here, almost 20% of age
- 3:26:41column has a missing value. Then, we
- 3:26:43have the cabin columns. So, this is
- 3:26:44quite a large value, and then we have
- 3:26:46two values for embarked column as well.
- 3:26:49Add a cmap for color coding. So, I'll
- 3:26:51say cmap
- 3:26:55So, if I do this, so the graph becomes
- 3:26:57more attractive. So, over here, your
- 3:26:59yellow stands for true, or you can say
- 3:27:01the values are null.
- 3:27:03So, here we have concluded that we have
- 3:27:05the missing value of age. We have a lot
- 3:27:07of missing values in the cabin column
- 3:27:09and we have very less value, which is
- 3:27:11not even visible in the embark column as
- 3:27:13well.
- 3:27:14So, to remove these missing values, you
- 3:27:16can either replace the values and you
- 3:27:18can put in some dummy values to it or
- 3:27:20you can simply drop the column.
- 3:27:22So, here let us first pick the age
- 3:27:24column. So, first let me just plot a box
- 3:27:26plot and then we analyze with having a
- 3:27:28column as age. So, I'll say sns.
- 3:27:31boxplot
- 3:27:33I'll say x is equals to passenger class.
- 3:27:36So, it's pclass. I'll say y is equals to
- 3:27:38age.
- 3:27:39And the data set that I'll be using is
- 3:27:41Titanic set. So, I'll say data is equals
- 3:27:43to Titanic data.
- 3:27:45You can see the age in first class and
- 3:27:47second class tends to be more older
- 3:27:49rather than we have it in the third
- 3:27:50class. Well, that depends on the
- 3:27:52experience, how much you earn or might
- 3:27:54be there any number of reasons.
- 3:27:56So, here we concluded that passengers
- 3:27:57who were traveling in class one and
- 3:27:59class two attend to be older than what
- 3:28:01we have in the class three.
- 3:28:03So, we have found that we have some
- 3:28:05missing values in M.
- 3:28:06Now, one way is to either just drop the
- 3:28:08column or you can just simply fill in
- 3:28:10some values to that. So, this method is
- 3:28:12called as imputation.
- 3:28:14Now, to perform data wrangling or
- 3:28:16cleaning, let us first print the head of
- 3:28:17the data set. So, I'll say Titanic.head.
- 3:28:20Sorry, it's Titanic_data.
- 3:28:23Let's say I just want the five rows.
- 3:28:25So, here we have survived, which is
- 3:28:27again categorical. So, in this
- 3:28:28particular column, I can apply logistic
- 3:28:30regression. So, this can be my y value
- 3:28:33or the value that you need to predict.
- 3:28:35Then, we have the passenger class. We
- 3:28:36have the name. Then, we have ticket
- 3:28:38number, fare, cabin. So, over here we
- 3:28:41have seen that in cabin we have a lot of
- 3:28:43null values or you can say the NaN
- 3:28:44values, which is quite visible as well.
- 3:28:46So, first of all, we'll just drop this
- 3:28:48column. So, for dropping it, I'll just
- 3:28:50say Titanic_data.
- 3:28:52And I'll simply type in drop and the
- 3:28:53column which I need to drop. So, I have
- 3:28:56to drop the cabin column.
- 3:28:58I mention the axis equals to one and
- 3:29:00I'll say in place also to true.
- 3:29:04So, now again I'll just print the head
- 3:29:05and let us see whether this column has
- 3:29:07been removed from the data set or not.
- 3:29:09So, I'll say Titanic .head.
- 3:29:12So, as you can see here we don't have
- 3:29:14cabin column anymore.
- 3:29:16Now, you can also drop the NA values.
- 3:29:18So, I'll say Titanic data
- 3:29:20.drop all the NA values or you can say
- 3:29:22NaN which is not a number and I'll say
- 3:29:25in place is equals to true.
- 3:29:27Let's Titanic
- 3:29:29So, over here let me again plot the heat
- 3:29:31map and let's see all the values which
- 3:29:33were before showing a lot of null values
- 3:29:35has it been removed or not. So, I'll say
- 3:29:37sns.heatmap I'll pass in the data set.
- 3:29:41I'll check if this null.
- 3:29:43I'll say white labels is equals to
- 3:29:45false.
- 3:29:47And I don't want color coding, so again
- 3:29:49I'll say false.
- 3:29:51So, this will basically help me to check
- 3:29:53whether my values has been removed from
- 3:29:55the data set or not. So, as you can see
- 3:29:56here I don't have any null values. So,
- 3:29:59it's entirely black.
- 3:30:01Now, you can actually know the sum as
- 3:30:02well, so I'll just go above.
- 3:30:05So, I'll just copy this part and I'll
- 3:30:07just use the sum function to calculate
- 3:30:09the sum.
- 3:30:10So, here that tells me the data set is
- 3:30:12clean as in the data set does not
- 3:30:14contain any null value or any NaN value.
- 3:30:18So, now we have wrangled our data or you
- 3:30:20can say clean our data.
- 3:30:21So, here we have done just one step in
- 3:30:23data wrangling that is just removing one
- 3:30:25column out of it. Now, you can do a lot
- 3:30:27of things. You can actually fill in the
- 3:30:28values with some other values or you can
- 3:30:31just calculate the mean and then you can
- 3:30:32just fit in the null values.
- 3:30:34But, now if I see my data set So, I'll
- 3:30:37say Titanic data.head.
- 3:30:39But, now if I see over here I have a lot
- 3:30:41of string values. So, this has to be
- 3:30:43converted to a categorical variables in
- 3:30:45order to implement logistic regression.
- 3:30:47So, what we will do we will convert this
- 3:30:49to categorical variable into some dummy
- 3:30:51variables and this can be done using
- 3:30:53pandas because logistic regression just
- 3:30:55take two values.
- 3:30:57So whenever you apply machine learning,
- 3:30:58you need to make sure that there are no
- 3:31:00string values present because it won't
- 3:31:02be taking these as your input variables.
- 3:31:05So using string, you don't have to
- 3:31:06predict anything. But in my case, I have
- 3:31:08the survived columns, so I need to
- 3:31:09predict how many people tend to survive
- 3:31:11and how many did not. So zero stands for
- 3:31:13did not survive and one stands for
- 3:31:15survive.
- 3:31:16So now let me just convert these
- 3:31:17variables into dummy variables.
- 3:31:20So let's use pandas and I'll say
- 3:31:22pd.get_dummies.
- 3:31:24You can simply press tab to auto
- 3:31:26complete. I'll say Titanic data.
- 3:31:29And I'll pass the sex.
- 3:31:30So you can just simply click on shift
- 3:31:32plus tab to get more information on
- 3:31:34this.
- 3:31:35So here we have the type data frame and
- 3:31:37we have the passenger ID survived and
- 3:31:39passenger class.
- 3:31:40So if you run this, you'll see that zero
- 3:31:42basically stands for not a female and
- 3:31:44one stand for it is a female. Similarly
- 3:31:46for male, zero stands for it's not male
- 3:31:48and one stand for it's male. Now we
- 3:31:50don't require both these columns because
- 3:31:52one column itself is enough to tell us
- 3:31:55whether it's male or you can say female
- 3:31:56or not. So let's say if I want to keep
- 3:31:58only male, I'll say if the value of male
- 3:32:01is one, so it is definitely a male and
- 3:32:03it is not a female. So that is how it
- 3:32:05you don't need both of these values. So
- 3:32:07for that, I'll just remove the first
- 3:32:09column, let's say a female.
- 3:32:10So I'll say drop first
- 3:32:13and true.
- 3:32:15So over here it has given me just one
- 3:32:17column which is male and has the value
- 3:32:19zero and one. Let me just set this as a
- 3:32:22variable, let's say sex. So over here I
- 3:32:24can say sex.head.
- 3:32:26I just want to see the first five rows.
- 3:32:29Sorry, it's dot.
- 3:32:31So this is how my data looks like.
- 3:32:34Now here we have done it for sex, then
- 3:32:35we have the numerical values in age, we
- 3:32:37have the numerical values in spouses,
- 3:32:39then we have the ticket number, we have
- 3:32:41the fare and we have embarked as well.
- 3:32:43So in embarked, the values are in S, C
- 3:32:45and Q. So here we can apply this get
- 3:32:48dummy function.
- 3:32:49So, let's say I'll take a variable,
- 3:32:51let's say embarked.
- 3:32:53I'll use the pandas library.
- 3:32:57I'll enter the column name, that is
- 3:32:59embarked.
- 3:33:03So, let me just print the head of it.
- 3:33:04So, I'll say embarked.head.
- 3:33:07So, over here we have C, Q, and S. Now,
- 3:33:09here also we can drop the first column
- 3:33:11because these two values are enough
- 3:33:13whether the passenger is either
- 3:33:15traveling from Q, that is Queenstown, S
- 3:33:17for Southampton. And if both the values
- 3:33:18are zero, then definitely the passenger
- 3:33:20is from Cherbourg, that is the third
- 3:33:22value. So, you can again drop the first
- 3:33:24value. So, I'll say drop
- 3:33:27and true.
- 3:33:28Let me just run this. So, this is how my
- 3:33:29output looks like. Now, similarly you
- 3:33:31can do it for passenger class as well.
- 3:33:33So, here also we have three classes,
- 3:33:35one, two, and three.
- 3:33:36So, I'll just copy the whole statement.
- 3:33:42So, let's say I want the variable name,
- 3:33:44let's say PCL.
- 3:33:47I'll pass in the column name, that is P
- 3:33:48class, and I'll just drop the first
- 3:33:51column. So, here also the values would
- 3:33:53be one, two, or three, and I'll just
- 3:33:55remove the first column. So, here we
- 3:33:57just left with two and three. So, if
- 3:33:59both the values are zero, then
- 3:34:00definitely the passenger is traveling in
- 3:34:02the first class. Now, we have made the
- 3:34:04values as categorical. Now, my next step
- 3:34:06would be to concatenate all these new
- 3:34:09rows into a data set.
- 3:34:11I can say Titanic data. Using the
- 3:34:13pandas, we'll just concatenate all these
- 3:34:15columns. So, I'll say pd.concat.
- 3:34:18And I'll say we have to concatenate sex,
- 3:34:21we have to concatenate embarked and PCL.
- 3:34:24And then I'll mention the axis to one.
- 3:34:26I'll just run this.
- 3:34:28Okay, I need to print the head.
- 3:34:30So, over here you can see that these
- 3:34:32columns have been added over here. So,
- 3:34:34we have the male column which basically
- 3:34:36tells whether a person is male or it's a
- 3:34:38female. Then we have the embarked which
- 3:34:40is basically Q and S. So, if it's
- 3:34:43traveling from Queenstown, value would
- 3:34:44be one, else it would be zero. And if
- 3:34:46both of these values are zero, it is
- 3:34:48definitely traveling from Cherbourg.
- 3:34:50Then we have the passenger class as two
- 3:34:52and three. So, if the value of both
- 3:34:54these is zero, then the passenger is
- 3:34:56traveling in class one.
- 3:34:58So, I hope you got this till now.
- 3:35:00Now, these are the irrelevant columns
- 3:35:02that we have it over here. So, we can
- 3:35:04just drop these columns. We're dropping
- 3:35:05P class,
- 3:35:07the embarked column,
- 3:35:09and the sex column.
- 3:35:10So, I'll just type in Titanic data
- 3:35:13{dot} drop. I'll mention the columns
- 3:35:15that I want to drop. So, I'll say
- 3:35:22I'll even delete the passenger ID
- 3:35:24because it's nothing but just the index
- 3:35:25value, which is starting from one.
- 3:35:27So, I'll drop this as well.
- 3:35:29Then I don't want name as well, so I'll
- 3:35:31delete name as well.
- 3:35:32Then what else we can drop? We can drop
- 3:35:34the ticket as well.
- 3:35:37And then I'll just mention the axis.
- 3:35:40I'll say in place is equals to true.
- 3:35:43Okay, so the my column name starts from
- 3:35:45upper case.
- 3:35:47So, these have been dropped. Now, let me
- 3:35:48just print my data set again.
- 3:35:51So, this is my final data set, guys. We
- 3:35:52have the survived column, which has the
- 3:35:54value zero and one. Then we have the
- 3:35:56passenger class. Oh, we forgot to drop
- 3:35:58this as well. So, no worries. I'll drop
- 3:36:00this again.
- 3:36:05So, now let me just run this.
- 3:36:08So, over here we have the survived, we
- 3:36:10have the age, we have the sibsp, we have
- 3:36:12the parch, we have fare, male, and these
- 3:36:15we have just converted.
- 3:36:17So, here we have just performed data
- 3:36:18wrangling, or you can say clean the
- 3:36:20data. And then we have just converted
- 3:36:22the values of gender to male, then
- 3:36:25embarked to Q and S, and the passenger
- 3:36:27class to two and three. So, this was all
- 3:36:29about my data wrangling, or just
- 3:36:30cleaning the data.
- 3:36:32Then my next step is training and
- 3:36:33testing your data. So, here we will
- 3:36:35split the data set into train subset and
- 3:36:37test subset. And then what we'll do
- 3:36:39we'll build a model on the train data
- 3:36:41and then predict the output on your test
- 3:36:43data set. So let me just go back to
- 3:36:45Jupiter and let us implement this as
- 3:36:46well.
- 3:36:47Over here I need to train my data set.
- 3:36:49So I'll just put this in date heading
- 3:36:51three.
- 3:36:54So over here you need to define your
- 3:36:55dependent variable and independent
- 3:36:57variable.
- 3:36:58So here my Y is the output or you can
- 3:37:00say the value that I need to predict.
- 3:37:03So over here I'll write Titanic data.
- 3:37:06I'll take the column which is survive.
- 3:37:09So basically I have to predict this
- 3:37:11column whether the passenger survived or
- 3:37:12not. And as you can see we have the
- 3:37:14discrete outcome which is in the form of
- 3:37:16zero and one. And rest all the things we
- 3:37:19can take it as a features or you can say
- 3:37:21independent variable. So I'll say
- 3:37:22Titanic data
- 3:37:24.drop.
- 3:37:26So we'll just simply drop the survive
- 3:37:28and all the other columns will be my
- 3:37:29independent variable.
- 3:37:31So everything else are the features
- 3:37:32which leads to the survival rate. So
- 3:37:34once we have defined the independent
- 3:37:36variable and the dependent variable,
- 3:37:38next step is to split your data into
- 3:37:39training and testing subset. So for that
- 3:37:42we'll be using SK learn. I'll just type
- 3:37:44in from SK learn.cross_validation
- 3:37:48import train_test_split.
- 3:37:54Now here if you just click on shift and
- 3:37:56tab, you can go to the documentation and
- 3:37:59you can just see the examples over here.
- 3:38:02I'll click on plus to open it. And then
- 3:38:04I'll just go to examples and see how you
- 3:38:06can split your data. So over here you
- 3:38:08have X_train, X_test, Y_train, Y_test.
- 3:38:12And then using this train_test_split you
- 3:38:13can just pass in your independent
- 3:38:15variable and dependent variable and just
- 3:38:17define a size and a random state to it.
- 3:38:19So let me just copy this.
- 3:38:21And I'll just paste it over here.
- 3:38:24Over here we'll train_test. Then we have
- 3:38:27the dependent variable train and test.
- 3:38:29And using the split function we'll pass
- 3:38:30in the independent and dependent
- 3:38:32variable and then we'll set a split
- 3:38:34size. So, let's say I'll put it at 0.3.
- 3:38:37So, this basically means that your data
- 3:38:38set is divided in 0.3. That is in 70/30
- 3:38:41ratio. And then I can add any random
- 3:38:43state to it. So, let's say I'm applying
- 3:38:45one. This is not necessary. If you want
- 3:38:48the same result as that of mine, you can
- 3:38:49add the random state. So, this will
- 3:38:51basically take exactly the same sample
- 3:38:53every time.
- 3:38:55Next, I have to train and predict by
- 3:38:57creating a model. So, here logistic
- 3:38:59regression will grab from the linear
- 3:39:01regression. So, next I'll just type in
- 3:39:03from sklearn.linear_model
- 3:39:07import LogisticRegression.
- 3:39:09Next, I'll just create the instance of
- 3:39:11this logistic regression model. So, I'll
- 3:39:13say log model
- 3:39:15is equals to LogisticRegression.
- 3:39:17Now, I just need to fit my model. So,
- 3:39:19I'll say log model.fit and I'll just
- 3:39:22pass in my X_train
- 3:39:25and Y_train.
- 3:39:28All right. So, here it gives me all the
- 3:39:30details of logistic regression.
- 3:39:32So, here it gives me the class weight,
- 3:39:34dual, fit intercept and all those
- 3:39:35things.
- 3:39:36Then, what I need to do, I need to make
- 3:39:38prediction. So, I'll take a variable,
- 3:39:40let's say predictions, and I'll pass on
- 3:39:42the model to it. So, I'll say log
- 3:39:44model.predict
- 3:39:46and I'll pass in the value that is
- 3:39:48X_test. So, here we have just created a
- 3:39:50model, fit that model, and then we have
- 3:39:52made predictions. So, now to evaluate
- 3:39:54how my model has been performing, so you
- 3:39:56can simply calculate the accuracy or you
- 3:39:58can also calculate the classification
- 3:40:00report. So, don't worry, guys. I'll be
- 3:40:02showing both of these methods. So, I'll
- 3:40:04say from sklearn.metrics
- 3:40:08import classification_report.
- 3:40:11So, over here I'll use
- 3:40:12classification_report and inside this
- 3:40:14I'll be passing in Y_test and the
- 3:40:16predictions.
- 3:40:21So, guys, this is my classification
- 3:40:22report. So, over here I have the
- 3:40:24precision, I have the recall, we have
- 3:40:26the F1 score, and then we have support.
- 3:40:29So, here we have the value of precision
- 3:40:31as 75, 72, and 73, which is not that
- 3:40:34bad. Now, in order to calculate the
- 3:40:36accuracy as well, you can also use the
- 3:40:38concept of confusion matrix. So, if you
- 3:40:40want to print the confusion matrix, I'll
- 3:40:42simply say from SK learn {dot} metrics
- 3:40:46import confusion matrix first of all,
- 3:40:48and then we'll just print this.
- 3:40:51So, here my function has been imported
- 3:40:52successfully, so I'll say confusion
- 3:40:54matrix.
- 3:40:56And I'll again pass in the same
- 3:40:57variables, which is Y test and
- 3:40:59predictions.
- 3:41:01So, I hope you guys already know the
- 3:41:03concept of confusion matrix. So, can you
- 3:41:05guys give me a quick confirmation as to
- 3:41:07whether you guys remember this confusion
- 3:41:09matrix concept or not? So, if not, I can
- 3:41:11just quickly summarize this as well.
- 3:41:14Okay, Jagriti says a yes.
- 3:41:16Okay, Swati is not clear with this. So,
- 3:41:18I'll just tell you in a brief what
- 3:41:19confusion matrix is all about.
- 3:41:22So, confusion matrix is nothing but a 2
- 3:41:24by 2 matrix, which has a four outcomes.
- 3:41:27This basically tells us that how
- 3:41:28accurate your values are. So, here we
- 3:41:30have the column as predicted no,
- 3:41:32predicted yes,
- 3:41:34and we have actual no and an actual yes.
- 3:41:38So, this is the concept of confusion
- 3:41:40matrix. So, here let me just feed in
- 3:41:42these values which we have just
- 3:41:43calculated. So, here we have 105,
- 3:41:48105, 21, 25, and 63.
- 3:41:53So, as you can see here, we have got
- 3:41:55four outcomes. Now, 105 is the value
- 3:41:58where our model has predicted no, and in
- 3:42:00reality it was also a no. So, here we
- 3:42:03have predicted no and an actual no.
- 3:42:05Similarly, we have 63 as a predicted
- 3:42:07yes, so here the model predicted yes,
- 3:42:09and actually also it was a yes. So, in
- 3:42:12order to calculate the accuracy, you
- 3:42:13just need to add the sum of these two
- 3:42:15values and just divide the whole by the
- 3:42:18sum. So, here these two values tells me
- 3:42:20where the model has actually predicted
- 3:42:22the correct output. So, this value is
- 3:42:24also called as true negative. This is
- 3:42:26called as false positive. This is called
- 3:42:28as true positive and this is called as
- 3:42:30false negative. Now, in order to
- 3:42:31calculate the accuracy, you don't have
- 3:42:33to do it manually. So, in Python, you
- 3:42:35can just import accuracy score function
- 3:42:38and you can get the results from that.
- 3:42:39So, I'll just do that as well. So, I'll
- 3:42:41say from sklearn.metrics
- 3:42:44import accuracy score
- 3:42:46and I'll simply print the accuracy.
- 3:42:49I'll pass in the same variables, that is
- 3:42:51Y test and predictions.
- 3:42:53So, over here it tells me the accuracy
- 3:42:54as 78, which is quite good. So, over
- 3:42:57here if you want to do it manually, you
- 3:42:59have to plus these two numbers, which is
- 3:43:01105 + 63. So, this comes out to almost
- 3:43:04168.
- 3:43:06And then you have to divide it by the
- 3:43:07sum of all the four numbers. So, 105 +
- 3:43:1063 + 21 + 25. So, this gives me a result
- 3:43:13of 214. So, now if you divide these two
- 3:43:16number, you'll get the same accuracy
- 3:43:18that is 78% or you can say 0.78.
- 3:43:21So, that is how you can calculate the
- 3:43:23accuracy. So, now let me just go back to
- 3:43:25my presentation and let's see what all
- 3:43:27we have covered till now.
- 3:43:28So, here we have first split our data
- 3:43:30into train and test subset. Then we have
- 3:43:32built our model on the train data and
- 3:43:34then predicted the output on the test
- 3:43:36data set. And then my fifth step is to
- 3:43:38check the accuracy. So, here we have
- 3:43:40calculated accuracy to almost 78%, which
- 3:43:43is quite good. You cannot say that
- 3:43:45accuracy is bad.
- 3:43:46So, here it tells me how accurate your
- 3:43:48results are. So, here my accuracy score
- 3:43:50defines that and hence we got a good
- 3:43:52accuracy.
- 3:43:54So, now moving ahead, let us see the
- 3:43:55second project, that is SUV data
- 3:43:57analysis.
- 3:43:58So, in this a car company has released
- 3:44:00new SUV in the market. And using the
- 3:44:03previous data about the sales of their
- 3:44:05SUV, they want to predict the category
- 3:44:07of people who might be interested in
- 3:44:08buying this. So, using the logistic
- 3:44:10regression, you need to find what
- 3:44:12factors make people more interested in
- 3:44:14buying this SUV.
- 3:44:15So, for this let us see our data set
- 3:44:17where I have user ID, I have gender as
- 3:44:19male and female.
- 3:44:21Then we have the age, we have the
- 3:44:22estimated salary.
- 3:44:24And then we have the purchased column.
- 3:44:26So this is my discrete column, or you
- 3:44:28can say the categorical column. So here
- 3:44:30we just have the value that is zero and
- 3:44:31one. And this column we need to predict
- 3:44:33whether a person can actually purchase a
- 3:44:35SUV or not. So based on these factors,
- 3:44:38we will be deciding whether a person can
- 3:44:40actually purchase a SUV or not.
- 3:44:42So we know the salary of a person, we
- 3:44:43know the age. And using these, we can
- 3:44:46predict whether a person can actually
- 3:44:47purchase a SUV or not. So let me just go
- 3:44:49to my Jupiter notebook and let's
- 3:44:51implement logistic regression. So guys,
- 3:44:53I will not be going through all the
- 3:44:54details of data cleaning and analyzing
- 3:44:56the part. So that part, I'll just leave
- 3:44:58it on you. So just go ahead and practice
- 3:45:00as much as you can.
- 3:45:02All right. So my second project is SUV
- 3:45:04predictions.
- 3:45:07All right. So first of all, I have to
- 3:45:08import all the libraries. So I say
- 3:45:10import NumPy as NP.
- 3:45:13And similarly, I'll do the rest of it.
- 3:45:21All right.
- 3:45:22So now let me just print the head of
- 3:45:24this data set.
- 3:45:26So this we have already seen that we
- 3:45:27have columns as user ID, we have gender,
- 3:45:30we have the age, we have the salary, and
- 3:45:32then we have to calculate whether a
- 3:45:33person can actually purchase a SUV or
- 3:45:35not.
- 3:45:36So now let us just simply go on to the
- 3:45:38algorithm part. So we'll directly start
- 3:45:40off with the logistic regression and how
- 3:45:42you can train a model. So for doing all
- 3:45:44those things, we first need to define
- 3:45:46your independent variable and dependent
- 3:45:47variable. So in this case, I want my X,
- 3:45:50that is our independent variable. I say
- 3:45:52dataset.iloc.
- 3:45:54So here I'll be specifying all the rows.
- 3:45:56So colon basically stands for that. And
- 3:45:58in the columns, I want only two and
- 3:46:01three. dot values. So here it should
- 3:46:04fetch me all the rows and only the
- 3:46:05second and third column, which is age
- 3:46:07and estimated salary. So these are the
- 3:46:09factors which will be used to predict
- 3:46:11the dependent variable that is purchase.
- 3:46:13So, here my dependent variable is
- 3:46:15purchase.
- 3:46:16And the dependent variable is of age and
- 3:46:17salary. So, I'll say
- 3:46:20dataset.iloc
- 3:46:21I'll have all the rows and I just want
- 3:46:24fourth column that is my purchase
- 3:46:25column. dot values. All right, so I just
- 3:46:28forgot one
- 3:46:29one square bracket over here. All right.
- 3:46:31So, over here I have defined my
- 3:46:33independent variable and dependent
- 3:46:34variable. So, here my independent
- 3:46:37variable is age and salary and dependent
- 3:46:39variable is the column purchase.
- 3:46:41Now, you must be wondering what is this
- 3:46:42iloc function. So, iloc function is
- 3:46:44basically an indexer for pandas data
- 3:46:46frame and it is used for integer-based
- 3:46:48indexing or you can also say selection
- 3:46:50by index.
- 3:46:52Now, let me just print these independent
- 3:46:53variables and dependent variables. So,
- 3:46:56if I print the independent variable, I
- 3:46:57have the age as well as the salary.
- 3:47:00Next, let me print the dependent
- 3:47:01variable as well. So, over here you can
- 3:47:03see I just have the values in zero and
- 3:47:05one. So, zero stands for did not
- 3:47:07purchase. Next, let me just divide my
- 3:47:10dataset into training and test subset.
- 3:47:12So, I'll simply write in from
- 3:47:13sklearn.cross_split
- 3:47:15dot cross_validation
- 3:47:18import train_test.
- 3:47:20Next, I'll just press shift and tab. And
- 3:47:23over here I'll go to the examples and
- 3:47:25just copy the same line.
- 3:47:27So, I'll just copy this.
- 3:47:30I'll remove the points.
- 3:47:31Now, I want the test size to be let's
- 3:47:33say 25. So, I have divided the train and
- 3:47:35test split in 75 25 ratio.
- 3:47:38Now, let's say I'll take the random
- 3:47:40state as zero. So, random state
- 3:47:42basically ensures the same result or you
- 3:47:44can say the same samples taken whenever
- 3:47:46you run the code. So, let me just run
- 3:47:48this. Now, you can also scale your input
- 3:47:50values for better performing and this
- 3:47:52can be done using standard scalar. So,
- 3:47:54let me do that as well. So, I'll say
- 3:47:56from sklearn.preprocessing
- 3:48:00import standard scalar.
- 3:48:02Now, why do we scale it? Now, if you see
- 3:48:04our dataset we are dealing with large
- 3:48:07numbers.
- 3:48:08Well, although we are using a very small
- 3:48:10data set, so whenever you're working in
- 3:48:12a broad environment, you'll be working
- 3:48:13with large data set where you'll be
- 3:48:15using thousands and hundred thousands of
- 3:48:17tuples. So, there scaling down will
- 3:48:19definitely affect the performance by a
- 3:48:20large extent. So, here let me just show
- 3:48:22you how you can scale down these input
- 3:48:24values. And then the pre-processing
- 3:48:26contains all your methods and
- 3:48:27functionality which is required to
- 3:48:29transform your data.
- 3:48:30So, now let us scale down for test as
- 3:48:32well as our training data set. So, I'll
- 3:48:34first make an instance of it. So, I'll
- 3:48:36say
- 3:48:37standard scalar.
- 3:48:39Then I'll have X_train. I'll say sc.fit
- 3:48:43fit_transform.
- 3:48:46I'll pass in my X_train variable.
- 3:48:52And similarly, I can do it for test
- 3:48:53wherein I'll pass the X_test.
- 3:48:57All right.
- 3:48:58Now, my next step is to import logistic
- 3:49:00regression. So, I'll simply apply
- 3:49:01logistic regression by first importing
- 3:49:03it. So, I'll say from sklearn
- 3:49:06from sklearn.linear_model
- 3:49:09import logistic regression.
- 3:49:12Now, over here I'll be using classifier.
- 3:49:14So, I'll say classifier. is equals to
- 3:49:16logistic regression.
- 3:49:18So, over here I'll just make an instance
- 3:49:19of it. So, I'll say logistic regression.
- 3:49:22And over here I'll just pass in the
- 3:49:23random state which is zero.
- 3:49:26And now I'll simply fit the model.
- 3:49:31And I'll simply pass in X_train and
- 3:49:32Y_train.
- 3:49:34So, here it tells me all the details of
- 3:49:36logistic regression.
- 3:49:39Then I have to predict the values. So,
- 3:49:41I'll say Y_pred
- 3:49:42is equals to classifier
- 3:49:45then predict function
- 3:49:47and then I'll just pass in X_test.
- 3:49:50So, now we have created the model, we
- 3:49:52have scaled down our input values, then
- 3:49:53we have applied logistic regression, we
- 3:49:55have predicted the values, and now we
- 3:49:57want to know the accuracy. So, to know
- 3:49:59the accuracy, first we need to import
- 3:50:01accuracy score. So, I'll say from
- 3:50:03sklearn.metrics
- 3:50:06import accuracy score.
- 3:50:08And using this function we can calculate
- 3:50:09the accuracy. Or you can manually do
- 3:50:12that by creating a confusion matrix.
- 3:50:14So, I'll just pass in my Y_test and my
- 3:50:16Y_predicted.
- 3:50:19All right. So, over here I get the
- 3:50:20accuracy as 89%. So, if we want to know
- 3:50:23the accuracy in percentage, so I just
- 3:50:24have to multiply it by 100. And if I run
- 3:50:26this,
- 3:50:27so it gives me 89%. So, I hope you guys
- 3:50:30are clear with whatever I have taught
- 3:50:31you today. So, here I have taken my
- 3:50:33independent variables as age and salary.
- 3:50:35And then we have calculated that how
- 3:50:37many people can purchase the SUV. And
- 3:50:40then we have calculated our model by
- 3:50:42checking the accuracy. So, over here we
- 3:50:44get the accuracy as 89, which is great.
- 3:50:51Let's compare the two models.
- 3:50:54So, first of all, let's look at the
- 3:50:55definition of linear regression and
- 3:50:57logistic regression. So, the main aim of
- 3:51:00linear regression is to predict a
- 3:51:01continuous dependent variable based on
- 3:51:04the values of the independent variables.
- 3:51:07But when it comes to logistic
- 3:51:08regression, the aim is to predict a
- 3:51:10categorical dependent variable based on
- 3:51:13the values of independent variables.
- 3:51:15These are the main aim of each of these
- 3:51:17models. Now, let's look at the variable
- 3:51:20type. Now, in linear regression, the
- 3:51:22dependent variable is always continuous.
- 3:51:25All right. This is very important to
- 3:51:26remember, because this is the main
- 3:51:28objective of linear regression. Okay, it
- 3:51:30makes use of continuous dependent
- 3:51:32variables to predict continuous values.
- 3:51:34Similarly, when it comes to logistic
- 3:51:36regression, you're going to use
- 3:51:38categorical dependent variable to
- 3:51:40predict a categorical value. All right.
- 3:51:43Now, let's look at the estimation
- 3:51:44method. So guys, linear regression is
- 3:51:47based on the least square estimation,
- 3:51:49which basically says that the regression
- 3:51:51coefficients should be chosen in such a
- 3:51:54way that it minimizes the sum of the
- 3:51:56squared distance of each observed
- 3:51:58response. Okay? Now, this is in-depth
- 3:52:01about linear regression, so that's why
- 3:52:02I'm going to leave a link in the
- 3:52:03description. Now, logistic regression on
- 3:52:06the other hand is based on maximum
- 3:52:08likelihood estimation. Okay? This
- 3:52:10basically says that the coefficients
- 3:52:12should be chosen in such a way that it
- 3:52:15maximizes the probability of Y given
- 3:52:18some value of X. Next is the equation.
- 3:52:21Earlier, we discussed this equation
- 3:52:22where for linear regression, we have Y
- 3:52:24is equal to B not plus B1 into X plus E.
- 3:52:28Okay? Similarly, this is the equation
- 3:52:30for logistic regression.
- 3:52:32So, the next difference is a best fit
- 3:52:33line. So, guys, linear regression aims
- 3:52:36at finding the best fitting straight
- 3:52:38line, which is also called the
- 3:52:40regression line. All right? But, when it
- 3:52:42comes to logistic uh regression, if you
- 3:52:44try and uh map the relationship between
- 3:52:46the dependent and independent variable,
- 3:52:48you're going to get a curve, which is
- 3:52:50also known as the sigmoid curve. All
- 3:52:52right? So, in linear regression, the
- 3:52:53relationship between the dependent and
- 3:52:55independent variable is represented
- 3:52:58using a straight line, but when it comes
- 3:52:59to logistic regression, the relationship
- 3:53:01between the dependent and independent
- 3:53:03variable is represented using a sigmoid
- 3:53:06curve. Now, let's look at the
- 3:53:07relationship between dependent and
- 3:53:09independent variable. Now, when it comes
- 3:53:11to linear regression, there has to be a
- 3:53:13linear relationship between the two. So,
- 3:53:16when I say linear, I mean that the
- 3:53:17variables have to vary linearly. Okay?
- 3:53:20That's how the straight line is formed
- 3:53:22in the first place. All right? When it
- 3:53:24comes to logistic regression, it's not
- 3:53:26necessary to have a linear relationship.
- 3:53:29Now, the output of linear regression is
- 3:53:31always going to be a predicted integer
- 3:53:33value or basically a continuous value.
- 3:53:36So, when it comes to logistic
- 3:53:37regression, the output has to be a
- 3:53:39binary value. Okay? So, it should either
- 3:53:41be class A or class B or it should be
- 3:53:43zero or one, something like that.
- 3:53:45Finally, we have applications. Now,
- 3:53:47linear regression is mainly used to
- 3:53:49predict outcomes like the expected
- 3:53:52number of sales. And you know, it's
- 3:53:54always used to predict some continuous
- 3:53:55value. All right, but when it comes to
- 3:53:57logistic regression, it's mainly used in
- 3:53:59classification. So, when you want to
- 3:54:01classify a data set into two different
- 3:54:03classes, then you use logistic
- 3:54:05regression. You can find a lot of
- 3:54:07applications of linear regression in the
- 3:54:09business domain. And logistic regression
- 3:54:12is mainly used in the cybersecurity,
- 3:54:14image processing, and classification
- 3:54:16domain.
- 3:54:17>> [music]
- 3:54:22>> What is classification?
- 3:54:25And if I have to tell you about
- 3:54:26classification,
- 3:54:29like for example, what happens is like
- 3:54:32we have two type of when we talk about
- 3:54:34machine learning, machine learning is
- 3:54:36nothing but you know, like a series of
- 3:54:38instructions you give it to the computer
- 3:54:41so that it can learn the patterns from
- 3:54:42your data set, right? To give you an
- 3:54:44example, imagine that there is a
- 3:54:46trending topic, for example, you found
- 3:54:48it you want to find it out whether
- 3:54:51Prime Minister Modi will be the second
- 3:54:53Prime Minister once again the Prime
- 3:54:55Minister for the country or not, okay?
- 3:54:57So, now what you will do is you will
- 3:54:59collect the data set from multiple
- 3:55:00different sources. And you will you will
- 3:55:04actually build it a like a algorithm
- 3:55:06where you you will get a label as yes or
- 3:55:09no. Yes, he will continue as a next
- 3:55:11Prime Minister or no, he will not
- 3:55:12continue as a next Prime Minister. So,
- 3:55:14you will collect the data set and you
- 3:55:16will feed this data set to the computer
- 3:55:17and this process is called as a machine
- 3:55:20learning, right? So, now in this case
- 3:55:22what happens is um
- 3:55:24So, this this is about the
- 3:55:26classification, right? So, now machine
- 3:55:29learning is basically of two type. One
- 3:55:31is called as a supervised machine
- 3:55:33learning, another is called as a
- 3:55:34unsupervised machine learning, and the
- 3:55:36third one is a reinforcement machine
- 3:55:38learning. So, when we speak about
- 3:55:40supervised machine learning, as the name
- 3:55:42suggests, it provides some supervision,
- 3:55:44right? For example, the teacher teaching
- 3:55:46the kid is a supervised machine
- 3:55:48learning. Right? So, we will give the
- 3:55:50trained examples. We will give the
- 3:55:52trained data set with the pure label on
- 3:55:54top of that. Uh this is called as a
- 3:55:57supervised machine learning. So, if I
- 3:55:58draw in front of you, this type of
- 3:56:00machine learning look like this,
- 3:56:02supervised machine learning. Where what
- 3:56:04happens is like you would have the data
- 3:56:06set, which is a structured data set, and
- 3:56:08you would have one column, which is
- 3:56:10called as a label. What you want to
- 3:56:12predict, okay? And you would have a uh
- 3:56:15various predictors by which you want to
- 3:56:17predict. To give you an example, imagine
- 3:56:19that you want to predict the pricing of
- 3:56:21a community. Okay? You want to predict
- 3:56:23what would be the pricing of apartment
- 3:56:25in a particular community, right? Now,
- 3:56:27these can be the variable like you can
- 3:56:29see that uh what would be the number of
- 3:56:31how many floors it has. You can have a
- 3:56:33variable like what is a pollution level,
- 3:56:35how how many educational institution are
- 3:56:37nearby, right? So, based upon that, the
- 3:56:39pricing will change. But this type of
- 3:56:42supervised machine learning, why it is
- 3:56:43called supervised machine learning
- 3:56:45because we provide the independent
- 3:56:47variable or we provide the predictors.
- 3:56:50Also, we provide the label data set.
- 3:56:52Okay? Now, the supervised machine
- 3:56:54learning is basically of two type. This
- 3:56:57supervised machine learning, one type
- 3:56:59one first is called as the, you know,
- 3:57:02regression based supervised machine
- 3:57:03learning.
- 3:57:05Okay? And the second one is called as a
- 3:57:07classification based supervised machine
- 3:57:09learning.
- 3:57:10Now, what is a difference between
- 3:57:11regression based and a classification
- 3:57:13based supervised machine learning?
- 3:57:15Regression based supervised machine
- 3:57:17learning is that machine learning where
- 3:57:19what you want to predict
- 3:57:22is continuous in nature. Okay? Imagine
- 3:57:25that you want to predict the
- 3:57:27community prices, right? Which is a
- 3:57:29continuous value. If it is a continuous
- 3:57:32value, then we will go ahead with
- 3:57:33regression based supervised machine
- 3:57:35learning. Okay? Whereas, if you want to
- 3:57:38predict something which is the discrete
- 3:57:40outcome, to give you an example, you
- 3:57:42want to predict that whether I will win
- 3:57:44the match or not.
- 3:57:46Okay? I want to predict whether the
- 3:57:48particular employee will turn out from
- 3:57:50the company or not. You want to predict,
- 3:57:53you know, like whether the person will
- 3:57:55have a cancer or not.
- 3:57:57You're getting my point, right? So, if
- 3:57:59you have the output which you want to
- 3:58:01predict is in the form of yes or no or
- 3:58:04true or false, right? This is called as
- 3:58:07the supervised machine learning, but a
- 3:58:10classification-based supervised machine
- 3:58:12learning. Okay? So, classification-based
- 3:58:15supervised machine learning is the
- 3:58:16process of dividing the data set into
- 3:58:19different categories or group by adding
- 3:58:21a label. Okay, so always remember that
- 3:58:24whenever you guys want to predict the
- 3:58:27classes in the data set, whenever you
- 3:58:29want to predict, you know, like whether
- 3:58:31this will happen or not, whether the
- 3:58:33whether a person will do a credit card
- 3:58:36fraud or not. You're getting my point,
- 3:58:38right? Whether the employee will turn
- 3:58:39out from the company or not, right?
- 3:58:42Whether the particular person will have
- 3:58:43a diabetes as a disease or not. All
- 3:58:45these questions, wherever you want to
- 3:58:47find it out yes or no or true or false
- 3:58:50or you want to predict classes in the
- 3:58:51data set, this is called as a
- 3:58:53classification-based supervised machine
- 3:58:56learning. Okay? So, this is what is
- 3:58:58called as a classification-based
- 3:59:00supervised machine learning and
- 3:59:02today I will teach you, you know, I will
- 3:59:05tell you about various form of
- 3:59:06classification-based supervised machine
- 3:59:08learning. Although we will do a deep
- 3:59:10dive into decision tree. Right? Now, you
- 3:59:13will be able to understand that decision
- 3:59:15tree how decision tree is connected to
- 3:59:17the network what we are learning today
- 3:59:19with classification-based supervised
- 3:59:22machine learning. Okay? So, now
- 3:59:25we have various algorithms. Algorithm is
- 3:59:28nothing but a set of mathematical
- 3:59:30equations for classification-based
- 3:59:33supervised machine learning. We first of
- 3:59:36all, we have something called as
- 3:59:37decision tree. Then we would have
- 3:59:39something called as random forest, then
- 3:59:41naive base, and then KNN, which is
- 3:59:43called as a K nearest neighbors. Okay?
- 3:59:46So, let me give you few statements about
- 3:59:48this algorithm, and we have many others,
- 3:59:50but today we will focus on one of them,
- 3:59:52which is decision tree.
- 3:59:54So, now first let's start with decision
- 3:59:56tree. Now, what is a decision tree? What
- 3:59:58I was telling you is, believe me or not,
- 4:00:00decision tree is something you use every
- 4:00:02day in your um in your daily life. For
- 4:00:05example, you take decisions, and for ex-
- 4:00:08today also you took a decision to attend
- 4:00:10this webinar, right? But, how do you
- 4:00:13decide a decision based on various
- 4:00:15further decisions, right? For example,
- 4:00:18for today joining the webinar, you have
- 4:00:20seen that, okay,
- 4:00:21uh when this webinar is about, okay? So,
- 4:00:24you said it is weekday or weekend. Then
- 4:00:26you might have said check it out, right?
- 4:00:28What is a time of the webinar? Then you
- 4:00:30might have checked it out, what is a
- 4:00:31topic of the webinar, right? So, and who
- 4:00:34is conducting this webinar? So, based
- 4:00:36upon this, you took a decision, shall I
- 4:00:38go ahead or not go ahead? You're getting
- 4:00:40my point, right? So, this is what is
- 4:00:42called as a decision tree. We call it as
- 4:00:44decision tree because it is a graphical
- 4:00:46representation of all the possible
- 4:00:48solution to a decision. It's like a
- 4:00:49tree. Like, decision tree is like a
- 4:00:52tree. Why? Because a tree also start
- 4:00:55with a root, and then it emerge into
- 4:00:57various branches. Similarly, you have a
- 4:00:59decision tree, which I'm showing to you
- 4:01:01here as a simplest algorithm, which is
- 4:01:03being used for machine learning
- 4:01:05purposes.
- 4:01:06Now, it is something like this. Imagine
- 4:01:08that you want to find it out that, you
- 4:01:11know, like, do you want to go to a
- 4:01:13restaurant, or do you want to buy an
- 4:01:15hamburger? Okay? So, you have two
- 4:01:17choices. Either you can go for a
- 4:01:19restaurant, or either you can buy a
- 4:01:21hamburger. Now, how would you decide
- 4:01:23which one you would follow? So, you will
- 4:01:25start with what is called as a root
- 4:01:27node. You will start with the root node
- 4:01:28that whether I'm hungry or not, right?
- 4:01:32If I am hungry, right? Then only I will
- 4:01:35go for all these activity. If I am not
- 4:01:38hungry at all, then simply go and sleep.
- 4:01:41You got my point, right? So, this is how
- 4:01:43you will start with the first node,
- 4:01:45which is to find it out whether I'm
- 4:01:47feeling hungry or not. If I'm feeling
- 4:01:48hungry, then I will decide that, "Well,
- 4:01:51I Do I have money, which is around $25
- 4:01:54worth?" If I have money, then I will go
- 4:01:56for a restaurant. If I don't have a
- 4:01:58money, then I will buy a hamburger.
- 4:02:00You understood this? This is simplest
- 4:02:03representation of decision tree.
- 4:02:05Basically, you decide something on the
- 4:02:07basis of the previous outcomes. And you
- 4:02:10can imagine any sort of example here.
- 4:02:12Imagine that you want to find it out
- 4:02:14whether the person will do a credit card
- 4:02:16fraud or not. Again, it will depend upon
- 4:02:18the previous circumstances that, you
- 4:02:20know, for example, it will depend upon
- 4:02:23how much is the salary of the person.
- 4:02:24You It will depend upon what is a job
- 4:02:26profile of the person. It will depend
- 4:02:28upon the fact that, you know, like, for
- 4:02:29example, how many fraud or how many card
- 4:02:32this person have, right? So, on [snorts]
- 4:02:34basis of you will decide that whether
- 4:02:35this will do a credit card fraud or not.
- 4:02:39So, this is the simplest form of
- 4:02:41classification-based algorithm.
- 4:02:43Then we have the next algorithm, which
- 4:02:45is called as a random forest. Now, what
- 4:02:48is a random forest? As the name
- 4:02:50suggests, you built one decision tree in
- 4:02:52the last example, right? But
- 4:02:55one decision tree can sometime be
- 4:02:57over-fitting as a case, right? So,
- 4:02:59people said that, you know, like, "Why
- 4:03:01should I only trust one decision tree?"
- 4:03:03For example, whenever you take a bold
- 4:03:05decision in your life, you don't trust
- 4:03:08only a single voice, right? You want to
- 4:03:10hear [clears throat] it from multiple
- 4:03:12different people to make your decision
- 4:03:14more stronger, right? Go to a doctor. If
- 4:03:16the doctor says to you that you have
- 4:03:18this type of a disease, you don't
- 4:03:21believe in that. What you do is you also
- 4:03:23talk to second doctor, third doctor,
- 4:03:25fourth doctor to confirm that this is
- 4:03:27true or not. Okay? So, this is what is
- 4:03:30called as a random forest. As the name
- 4:03:32suggests, why it is called random
- 4:03:33forest? It is called random forest
- 4:03:36because now you're building the various
- 4:03:38number of decision trees here. It's like
- 4:03:40a forest of all the trees, right? It is
- 4:03:43no longer a single tree, it is a forest
- 4:03:45of complete trees.
- 4:03:47So, for example, you can imagine that if
- 4:03:49it is my training data set, I will I
- 4:03:51will split my training data set into
- 4:03:53multiple examples, multiple decision
- 4:03:55trees will be built, and then based upon
- 4:03:58the majority, I will decide whether
- 4:04:00should I do this or not. Okay? So, it is
- 4:04:03also called as a bagging sort of
- 4:04:04methodology where we bring the outcome
- 4:04:07of various models, or we bring the
- 4:04:09outcome of various trees all together to
- 4:04:12make the
- 4:04:13uh to make a powerful decision. Okay?
- 4:04:16So, this is another example This is
- 4:04:18another algorithm which is called as a
- 4:04:20random forest.
- 4:04:23Then, after random forest, we have
- 4:04:25something called as naive base. Okay?
- 4:04:28Now, what is a naive base algorithm?
- 4:04:31Naive base is a also a simplest
- 4:04:33algorithm, but naive base is basically
- 4:04:36based on the base theorem. Okay? So,
- 4:04:38this algorithm is basically based on the
- 4:04:40base theorem, and base theorem is based
- 4:04:42on the conditional probability.
- 4:04:45Right? So, it is based on the
- 4:04:47conditional probability, which is your
- 4:04:49naive base. Now, in naive base, what
- 4:04:52happens is like we decide that whether
- 4:04:56something will happen or not on the
- 4:04:58basis of probability. So, let me
- 4:05:01illustrate you with this example.
- 4:05:03Imagine that you want to find it out
- 4:05:06whether I would have a disease or not.
- 4:05:09Okay? So, first of all, the probability
- 4:05:12of having a disease is 0.10. And the
- 4:05:16probability of not having a disease is
- 4:05:180.90. Okay? So, if there is a
- 4:05:21probability of having a disease is 0.10,
- 4:05:24you will further find it out that what
- 4:05:26is a probability that my test will be
- 4:05:30positive given I'm diseased.
- 4:05:33What is the probability that my test to
- 4:05:35diagnose the disease is negative given I
- 4:05:37have a disease? Right? Similarly, if you
- 4:05:40go into this direction, if there is a
- 4:05:42probability of not having a disease is
- 4:05:440.90, you will find it out that what is
- 4:05:47a probability of [snorts] having a
- 4:05:49disease or test being positive with
- 4:05:51having with no disease. And what is a
- 4:05:54probability of no disease given I don't
- 4:05:57have the disease, which is 0.90. So,
- 4:05:59basically, here you check the outcomes.
- 4:06:03Here you check the outcome of all the
- 4:06:06all the possible combinations. This is
- 4:06:08what is called as the naive base
- 4:06:10algorithm or base theorem or
- 4:06:13you know, conditional probability base
- 4:06:15theorem. Okay?
- 4:06:17Then we also have something which is
- 4:06:19called as a K nearest neighbor, which is
- 4:06:22the one of the finest algorithm
- 4:06:25which also helps you in deciding the
- 4:06:28classification, right? Now, what is K
- 4:06:31nearest neighbor? K nearest neighbor, as
- 4:06:33the name suggests, what happens in this
- 4:06:35case is we try to build, you know, we
- 4:06:38try to build
- 4:06:40basically, it's like a neighbor. So,
- 4:06:42it's something like this. Like if I give
- 4:06:44you the data set, say I give you the
- 4:06:46customer data set. Okay?
- 4:06:49And I have It is a transaction data set.
- 4:06:51So, customer number one
- 4:06:53has bought a product number one, right?
- 4:06:56From a particular vendor at a particular
- 4:06:58rate, and this is a profit this customer
- 4:07:00has given to us. And this is the
- 4:07:02revenues. Okay? This is my class
- 4:07:04customer number one. Similarly, you
- 4:07:07would have customer number two, customer
- 4:07:09number three, customer number four,
- 4:07:11right? Now, if I ask you that which
- 4:07:13customer profiling is same
- 4:07:15is nearby same, right? So, what you can
- 4:07:18do is you can group your customers or
- 4:07:20you will get to know that which customer
- 4:07:23behave in the similar way. So, you can
- 4:07:25say that customer one, customer three,
- 4:07:27and customer four, they behave in a
- 4:07:29similar way because they are giving us
- 4:07:30the high profit and high revenue
- 4:07:32margins.
- 4:07:34You understood? So, this is what happens
- 4:07:37in the case of K nearest neighbor. So, K
- 4:07:39nearest neighbor, what you generally do
- 4:07:41is you find it out that what uh I would
- 4:07:45be able to, you know,
- 4:07:47uh find it out who is my nearest
- 4:07:49neighbor, like what are the similarities
- 4:07:51in the patterns we have, right? This is
- 4:07:53what we use for K nearest neighbors. And
- 4:07:56this you can see that, for example, the
- 4:07:58algorithm based on the distance-based
- 4:08:00mechanism find it out that how many
- 4:08:02people or how what type of audience is
- 4:08:05similar to the outcomes, right?
- 4:08:08So, then moving further here
- 4:08:11the with the second topic, let's go into
- 4:08:13detail about what is decision tree all
- 4:08:15about, right? So, so far I have touched
- 4:08:18base on uh classification-based
- 4:08:20algorithm, right? And I was teaching you
- 4:08:22different type of classification-based
- 4:08:24algorithm, but since the focus of
- 4:08:26today's class is decision tree, let's
- 4:08:28take a deep dive into the decision tree.
- 4:08:31Okay? So, let's get started with
- 4:08:33decision tree.
- 4:08:34A decision tree is a graphical
- 4:08:36representation of all possible solution
- 4:08:39to a decision based on a certain
- 4:08:40condition. What does it mean? It means
- 4:08:43that it is as simple as that. Imagine
- 4:08:45that this is a tree, right? So, it is
- 4:08:47like a tree where
- 4:08:49you have a problem statement that should
- 4:08:51I accept a job offer or not. Imagine
- 4:08:54that you want to find it out that should
- 4:08:56I accept a job offer or not. This is a
- 4:08:58problem statement which is there in your
- 4:09:00mind. Now, how would you solve this
- 4:09:01problem statement with the help of a
- 4:09:03decision tree?
- 4:09:04First of all, you will start with what
- 4:09:06we call as a root node. We will start
- 4:09:08with a root node starting with, you
- 4:09:11know, to find it out what is the salary.
- 4:09:13Okay, what I'm what I what is the salary
- 4:09:15I'm getting. So if my salary is equal to
- 4:09:19or greater than equal to 50,000, I'll go
- 4:09:22here. If my salary is not greater than
- 4:09:25equal to 50,000, I will say I will not
- 4:09:28accept this offer.
- 4:09:29Okay? Now imagine that you say that your
- 4:09:32salary is greater than 50,000, then you
- 4:09:34will check another
- 4:09:36another variable here. You will check it
- 4:09:38out whether I have to commute more than
- 4:09:401 hour. If I have to commute more than 1
- 4:09:43hour, I will decline the offer.
- 4:09:46You're getting my point, right? Then you
- 4:09:48will if you don't have to commute more
- 4:09:50than 1 hour, then you will still
- 4:09:51consider this option and then you will
- 4:09:53check it out. For example, in this case
- 4:09:54we're checking it out whether you're
- 4:09:56getting the free offers also like coffee
- 4:09:59or some snacks or other things like
- 4:10:00that. If yes, then you will accept
- 4:10:03finally the offer, otherwise you will
- 4:10:05decline the offer.
- 4:10:07You got my point. This is how our
- 4:10:09decision tree works. Basically, decision
- 4:10:11tree will keep on splitting, keep on
- 4:10:13splitting unless and until you are able
- 4:10:16to find it out your decision.
- 4:10:18Okay? So here the decision was shall I
- 4:10:21accept this or not, right? So it will
- 4:10:23keep on splitting, keep on splitting
- 4:10:25unless and until you get your decision
- 4:10:27whether you do this or not. Like this
- 4:10:30example which I have to I have explained
- 4:10:32you. Okay?
- 4:10:34Now with this what happens is like let
- 4:10:37me if I go further and explain you
- 4:10:39further on this decision tree, let's
- 4:10:41understand this more importantly, right?
- 4:10:43Let's understand this
- 4:10:45one by one. So imagine that this is my
- 4:10:48data set. Okay? Now in my data set you
- 4:10:51can see that
- 4:10:53I have various [clears throat] colors
- 4:10:54given to you. It's like a green color,
- 4:10:55yellow color, red color, red color and a
- 4:10:58yellow color. I'm saying if this is a if
- 4:11:01there is a green color fruit with a
- 4:11:03diameter of three, right? And I can say
- 4:11:06it is a mango.
- 4:11:08Whereas I'm saying if it is a yellow
- 4:11:10color and a diameter three, still will
- 4:11:12be called as a mango. If it is a red
- 4:11:14color with one diameter, then it is a
- 4:11:17grape.
- 4:11:18And it can be a red color with a
- 4:11:20diameter of one still can be a grape,
- 4:11:22but even a yellow with a diameter of
- 4:11:25three can be lemon. Now, what is
- 4:11:27happening in this case is if you have
- 4:11:29this type of a data set, where you want
- 4:11:31to predict the label of the fruit,
- 4:11:33right? You want to predict whether a
- 4:11:35fruit will be mango, whether a fruit
- 4:11:37will be lemon, right? Or whether a fruit
- 4:11:40in this case is mango and lemon or
- 4:11:42grape.
- 4:11:43I have to build a classifier, which is a
- 4:11:45decision tree classifier on the basis of
- 4:11:47this data set. Imagine this is a problem
- 4:11:50statement given to you. Now, first of
- 4:11:52all, what you will do here is you will
- 4:11:54take this data set, right? You will take
- 4:11:56this data set and you will start with a
- 4:11:59root node. Root node imagine here is
- 4:12:01like, is my diameter of the fruit
- 4:12:03greater than equal to three or not?
- 4:12:06Okay? Is my diameter of the fruit
- 4:12:08greater than equal to three or not? If
- 4:12:11it is greater than equal to three,
- 4:12:13right? Now, if it is greater than equal
- 4:12:16to three, you can see that you have
- 4:12:18three fruits here, three rows of the
- 4:12:19data set here. One is green with three,
- 4:12:22mango. Yellow with three, lemon. Yellow
- 4:12:25with three, mango. And wherever it fails
- 4:12:28in the condition, you are left with
- 4:12:30where the diameter is not greater than
- 4:12:32equal to three, it is less than equal to
- 4:12:34three, then what which are the data set
- 4:12:36we have? Red one, grape and red one,
- 4:12:39grape.
- 4:12:41Okay? So, I hope you are understanding
- 4:12:43this, right? How we have starting with
- 4:12:45the decision tree with the base on one
- 4:12:47condition here, right? Which is
- 4:12:49diameter. Now, based upon this, what I
- 4:12:52will do is I have to split it further.
- 4:12:54Here in this case, I don't have to split
- 4:12:56it because I already got the result. So,
- 4:12:58if my diameter is greater than equal to
- 4:13:00three in my data set, if the diameter is
- 4:13:03not equal to greater than equal to
- 4:13:05three, I know it is a grape. Right?
- 4:13:08Whereas, if the diameter is greater than
- 4:13:10equal to three, then it can be a mango
- 4:13:13or it can be a lemon. I'm not sure on my
- 4:13:15decision. So, what I will do here is I
- 4:13:17will split it further. So, I have to
- 4:13:20split this further.
- 4:13:22Here, the splitting is not required.
- 4:13:24But, if I split this further, I may
- 4:13:26check it out now. Is the color equal to
- 4:13:29yellow or not? Now, if the color is
- 4:13:31equal to yellow, then I have two fruit
- 4:13:34which is
- 4:13:35your this row, which is, you know, um
- 4:13:38your mango or lemon. And if the color is
- 4:13:42not equal to yellow, then you will have
- 4:13:44the another row of the data set, which
- 4:13:46is, you know, which you're left with,
- 4:13:48right? So, this is how you have done
- 4:13:50this work
- 4:13:52in the case of your decision tree.
- 4:13:55Okay? And how you will find it out at
- 4:13:57which type of the criteria or which type
- 4:14:00of the algorithm I should or which type
- 4:14:02of criteria or variable I should choose
- 4:14:04here to split the tree is on the basis
- 4:14:06of Gini index and the information gain,
- 4:14:09which I will illustrate and show you in
- 4:14:11the few slides from now. Okay? So, I
- 4:14:14hope with this you understand the how
- 4:14:17your decision tree works. Basically, on
- 4:14:18the basis of the condition, and if you
- 4:14:20get a pure subset, then no need to split
- 4:14:23it further. If there is no pure subset,
- 4:14:25keep on splitting, keep on splitting
- 4:14:27unless and until you get a pure subset.
- 4:14:30Okay?
- 4:14:32So, in this case, what has happened is
- 4:14:34you got 100% mango. Here, you got 50%
- 4:14:37mango and 50% lemon. Okay?
- 4:14:41So, now
- 4:14:43this is what I have told you in the
- 4:14:45previous slide, like how does it work,
- 4:14:48right?
- 4:14:50Now, this all depend upon the Gini index
- 4:14:53basis, right? Now, in the next
- 4:14:55subsequent slide, let me tell you and
- 4:14:58explain you about Gini index and
- 4:15:00information gain. How does it work?
- 4:15:02Right? How does that something works on
- 4:15:05the basis of uh your Gini index?
- 4:15:08So, imagine that, you know, like imagine
- 4:15:11that I have a feature here is the color
- 4:15:14green or not, right? Basis of whether
- 4:15:16the feature whether you have a color
- 4:15:18green or not, what would happen here is
- 4:15:21if the color is green,
- 4:15:22you will get this row here. If it is
- 4:15:24not, you will get this these two rows
- 4:15:26here, right? In the next, it may decide
- 4:15:29on the other basis, right? It can be is
- 4:15:31the diameter which is greater than equal
- 4:15:33to three or not and many other things.
- 4:15:36Okay?
- 4:15:37Now, let's go ahead the example and the
- 4:15:40questions which is coming to your mind
- 4:15:42which is related to, you know, decision
- 4:15:44tree terminologies. I'm pretty sure you
- 4:15:47would have these questions in your mind
- 4:15:49and you would be thinking that how
- 4:15:50should I decided which feature I should
- 4:15:52use and which feature shouldn't I be
- 4:15:54using? This is the question which you
- 4:15:55had. And this is a excellent question
- 4:15:57for understanding the decision tree. But
- 4:16:00before I go on to that, I have to tell
- 4:16:02you about some of the terminologies
- 4:16:05which we commonly use while we build the
- 4:16:07decision tree. So, first of all,
- 4:16:09decision tree looks like this type of a
- 4:16:11tree-based structure, okay? So, every
- 4:16:13decision tree will have its root node.
- 4:16:16So, root node is where the decision tree
- 4:16:18will start. So, it represent the entire
- 4:16:20population or sample and it is further
- 4:16:23get divided or into two or more
- 4:16:25homogeneous sets. So, as you know that
- 4:16:26this will be the first feature on the
- 4:16:28basis of your tree will start. Like tree
- 4:16:31start with a root, your in this case,
- 4:16:33your decision tree will also start with
- 4:16:35a root node.
- 4:16:37Okay? Now, once you have the root, after
- 4:16:39that in a tree, what happens is you will
- 4:16:41start getting the branches, right? I
- 4:16:43hope you understand what is branches. In
- 4:16:45this case, we will keep on splitting,
- 4:16:47keep on splitting unless and until we
- 4:16:49get a decision whether this will happen
- 4:16:50or not. So, it's like a branches.
- 4:16:53Okay? Then, we would also have a parent
- 4:16:55or a child node. So, child node is
- 4:16:57nothing but when you have a branches,
- 4:17:00then the branches can also have the
- 4:17:01outcomes. So, in the previous example,
- 4:17:04like for example, we have whether the
- 4:17:06diameter is greater than equal to three
- 4:17:09or not. You remember? Whether the
- 4:17:10diameter is greater than equal to three
- 4:17:12or not. If it is true, what is
- 4:17:14happening? If it is a false, what is
- 4:17:15happening? So, this is a child or child
- 4:17:18node. Basically, this is a intermediate
- 4:17:20node. This is not the final decision
- 4:17:22which is being made. So, this is called
- 4:17:24as a parent or the child node.
- 4:17:27Then, we also have other terminology
- 4:17:29here, which is splitting, which you know
- 4:17:31that we will keep on splitting unless
- 4:17:33and until you get a desired node. And
- 4:17:35finally, the tree end with a leaf node.
- 4:17:37So, always remember one thing, you will
- 4:17:39start your tree with a root node. Root
- 4:17:42node is a node with which we will start
- 4:17:44the decision tree. And you will end your
- 4:17:46decision as decision tree at a leaf node
- 4:17:49where you will get a decision that you
- 4:17:51should do this or not.
- 4:17:53Okay? And now,
- 4:17:56uh pruning is a activity where
- 4:17:58uh you you will cut down the decision
- 4:18:00tree if it is pruned lot amount of
- 4:18:03times. I will even explain you this. It
- 4:18:05is a case of overfitting. You shouldn't
- 4:18:08build thousand You can shouldn't build
- 4:18:09thousand uh branches of the tree when it
- 4:18:12is not even required. So, I'll I'll
- 4:18:13explain you this point.
- 4:18:17Now, let's move further here and let's
- 4:18:20see this with the help of an example,
- 4:18:23which was your
- 4:18:24uh Gini index and your information gain.
- 4:18:28Right?
- 4:18:29So, this is where you were asking me
- 4:18:32that which question to ask and when,
- 4:18:34right? How would you decide that which
- 4:18:37feature has to be taken first? Right?
- 4:18:39Let's take up an example here and I
- 4:18:42would encourage everyone of you to hear
- 4:18:45me really fine here because this is on
- 4:18:47the basis how would you decide or how
- 4:18:49the algorithm decide to break this
- 4:18:51further. And and I'm explaining you with
- 4:18:53the help of one
- 4:18:55simplest example, and we will do some
- 4:18:57maths over it.
- 4:18:58So, imagine that
- 4:19:00there is a data set where I want to find
- 4:19:03it out whether I will play the match or
- 4:19:06not.
- 4:19:08Okay, there is a cricket match or there
- 4:19:10is a football match or whatever it is. I
- 4:19:12want to find it out whether I will play
- 4:19:14the match or not. Okay? Now,
- 4:19:17how do the decision tree look like?
- 4:19:19Decision tree look like this. If the
- 4:19:21outlook if the outlook is humid
- 4:19:25if the outlook is humid
- 4:19:27and if the outlook is humid and humidity
- 4:19:30is very high, then I will not play the
- 4:19:33match. On the other hand, if the outlook
- 4:19:36is humid and humidity is normal, I may
- 4:19:38play the match.
- 4:19:40Okay?
- 4:19:41If the outlook is
- 4:19:42absolutely clear, then I will always
- 4:19:44play. On the other hand, if outlook is
- 4:19:47windy and the winds are very strong, I
- 4:19:49may not play the match. On the other
- 4:19:51hand, if the outlook is windy and the
- 4:19:54winds are weak, I may play the match.
- 4:19:56So, now what is happening again is that
- 4:19:58you are not sure how you decided with
- 4:20:00outlook, how you decided with these
- 4:20:03nodes here, how you decided with these
- 4:20:05features here. Let me try to show you
- 4:20:07this entire data set. So, this entire
- 4:20:09data set look like this.
- 4:20:11Okay? So, this is basically your 14 days
- 4:20:14data set. And this exactly happens in
- 4:20:17the case of your classification based
- 4:20:18algorithm, you will get a data set like
- 4:20:20this.
- 4:20:21So, you can imagine that you want to
- 4:20:23find it out I would like whether I will
- 4:20:26play or not. Right? This is what you
- 4:20:29want to essentially find it want to find
- 4:20:31it out whether I would play or not. On
- 4:20:33basis of what? On basis of four
- 4:20:35different variables you have. Four
- 4:20:37different variables or features in the
- 4:20:39data set is outlook
- 4:20:41temperature, humidity, and wind. Right?
- 4:20:45So, now
- 4:20:46it can be like if the outlook is sunny,
- 4:20:49temperature is hot, humidity is high,
- 4:20:52wind is not there, I will not play the
- 4:20:55match. This is how you will read this
- 4:20:57one data point, right? Similarly, I have
- 4:20:59multiple data point. Now, I have to
- 4:21:01decide how can I build a decision tree
- 4:21:04and out of these four features, which
- 4:21:07feature should I use first as my root
- 4:21:09node? Right? So, this is what I have to
- 4:21:12decide and
- 4:21:14let's do it
- 4:21:16accordingly, right? So, now what I'll do
- 4:21:18from here is
- 4:21:21uh we I will illustrate you what is Gini
- 4:21:24index, what is information gain, and how
- 4:21:27does this happens, right? And this is
- 4:21:29basically used to build your decision
- 4:21:31tree.
- 4:21:32So, before I make you understand about
- 4:21:35Gini index or information gain, one
- 4:21:38thing which you should always remember
- 4:21:39is to understand the concept of
- 4:21:42impurity. What is impurity? Impurity is
- 4:21:44nothing but you can see there is a
- 4:21:46basket where you have apples, right?
- 4:21:49Now, if you have a basket which is of
- 4:21:50apple and another
- 4:21:52tray, it is written the label as apple.
- 4:21:55Now, in this case, you will never make a
- 4:21:57mistake. You will never make a mistake
- 4:22:00because
- 4:22:01here you have an apple, here you have
- 4:22:03only one label. So, everything will be
- 4:22:05perfect. It will be 100%. Basically,
- 4:22:08there is no impurity, there is no
- 4:22:10problem in your data set, right? On the
- 4:22:13other hand, let me flip the story. In
- 4:22:16this case, imagine that I have different
- 4:22:18fruits in the basket, which is like
- 4:22:20apple, you have a banana, you have a
- 4:22:22grapes, you have a you know, cherries,
- 4:22:25and many other things, right? And you
- 4:22:26have many apples, many labels here. In
- 4:22:29this case, try imagining you have to
- 4:22:32match each fruit with its label.
- 4:22:36Right? Now, in this case, the impurity
- 4:22:38cannot be equal to zero. The impurity
- 4:22:41will not be equal to zero in this case
- 4:22:43because what would happen here is that
- 4:22:45you have a chances of misclassification.
- 4:22:48Right? This is a very important concept,
- 4:22:50right? When you have perfect thing, the
- 4:22:53misclassification will not happen. But,
- 4:22:55whereas, if you have a multiple labels
- 4:22:57with multiple fruit, misclassification
- 4:22:59or impurities will not be equal to zero.
- 4:23:03Right? So, this is associated with a
- 4:23:06term called as entropy. Maybe in your
- 4:23:08childhood days, you have learned about
- 4:23:10entropy in your chemistry class. In a
- 4:23:12simple sense, what is entropy? Entropy
- 4:23:15is a randomness of the space sample
- 4:23:18space. Whenever you are not sure on your
- 4:23:21decision, then entropy will be more.
- 4:23:24Imagine that I'm giving you the data set
- 4:23:26where it is like 51% of doing this
- 4:23:29thing, 49% of not doing this thing. 51%
- 4:23:33chance that the employee may leave the
- 4:23:35organization, 49% chance that employee
- 4:23:37may not leave the organization. So,
- 4:23:39basically, you are not sure on your
- 4:23:40decision. If you are not sure on your
- 4:23:42decision, then the entropy will be very
- 4:23:45high. On the other hand, if you are very
- 4:23:47sure on your decision, then the entropy
- 4:23:50will be very low. So, basically, we need
- 4:23:53that feature which can provide us lowest
- 4:23:56entropy rather than the highest entropy
- 4:24:00uh to select as that as a good feature.
- 4:24:04So, we generally find it out entropy by
- 4:24:06the help of this formula. What we simply
- 4:24:09do is don't get scared with this formula
- 4:24:11because everything happens automatically
- 4:24:13in R or Python, right? Imagine just look
- 4:24:17at this formula. What is this formula?
- 4:24:18This formula says that what is the
- 4:24:21probability that something will happen
- 4:24:23multiplied by log base two probability
- 4:24:27that something will happen subtract this
- 4:24:29with probability that something will not
- 4:24:31happen into log base two probability
- 4:24:34that something will not happen. Okay,
- 4:24:36let's take up an example. Don't worry
- 4:24:38about it. Let's take up an example that
- 4:24:41probability that I will win the match
- 4:24:43you will apply here and probability that
- 4:24:46I will not play the match or win the
- 4:24:47match you will apply here and then you
- 4:24:49will calculate the entropy. Let's take
- 4:24:53you know an example how you can do this
- 4:24:56in our case.
- 4:24:58So, in our case let me show you.
- 4:25:05Yeah, how we will build the decision
- 4:25:07tree in our case. In our case you just
- 4:25:10see here what is happening is that we
- 4:25:12have
- 4:25:1414 instances
- 4:25:16or 14 [clears throat] rows of the data
- 4:25:17set where nine times if you see it I
- 4:25:20will play the match and five times I
- 4:25:23will not play the match. Okay, so if you
- 4:25:25see carefully there are nine labels
- 4:25:27where I'm playing the match and there
- 4:25:29are five labels where I'm not playing
- 4:25:31the match. Right? So, first of all I
- 4:25:33have to find it out the total entropy.
- 4:25:36How I will find the total entropy? This
- 4:25:38is being determined by probability that
- 4:25:41I will play the match sub multiply this
- 4:25:44with log base power two probability that
- 4:25:47I will play the match. What are the
- 4:25:49chances that I will play the match? Nine
- 4:25:51out of 14.
- 4:25:53All of you will be with me, right? This
- 4:25:55is nine out of 14. Multiply with log
- 4:25:58base two nine out of 14.
- 4:26:00Subtract this with what is the
- 4:26:01probability that I will not play the
- 4:26:03match? Five out of 14. Multiply this
- 4:26:06with log base two five out of 14.
- 4:26:09So, once you calculate this you will
- 4:26:11find it out the entropy of this entire
- 4:26:14system, entropy of this entire system is
- 4:26:170.94. Okay, so this is the entropy of
- 4:26:21your entire sample space. This is the
- 4:26:23first thing. Now, how this entropy will
- 4:26:26help you in selecting which features you
- 4:26:28will take or not? So, let's go further.
- 4:26:32So, now we will take each feature one by
- 4:26:34one. Whether I should take outlook,
- 4:26:36whether I should take temperature,
- 4:26:38whether I should pick up humidity, or
- 4:26:40whether should I pick up windy, right?
- 4:26:42Let's go one by one. Now, first I'm
- 4:26:44plotting for outlook. Imagine for
- 4:26:47outlook, how many times is what are the
- 4:26:50distinct value of outlook? Outlook can
- 4:26:51be sunny,
- 4:26:53outlook can be overcast, or outlook can
- 4:26:55be rainy. Right? These are the three
- 4:26:57different combinations you can have for
- 4:26:59outlook. Now, if the outlook is sunny,
- 4:27:02two times I'm playing, three times I'm
- 4:27:04not playing the match. If the outlook is
- 4:27:06overcast, I'm always playing the match.
- 4:27:08If the outlook is rainy, three times I'm
- 4:27:11playing, two times I'm not playing the
- 4:27:12match.
- 4:27:13Right? This is how I have bifurcated it.
- 4:27:16What I will do in the next iteration is,
- 4:27:18let me find it out the entropy of
- 4:27:21outlook.
- 4:27:22Okay? So, if we start with outlook,
- 4:27:26remember this formula which is, you
- 4:27:28know, probability that I will play
- 4:27:30multiply with the probability that I
- 4:27:32will play, right? And subtract with
- 4:27:35probability which I will not play, and
- 4:27:37log base two of not playing. So, 2 by 5
- 4:27:40is a chances that I will not I will play
- 4:27:42into log base two 2 by 5. This is a
- 4:27:45subtraction here, right? With I will
- 4:27:48play and log 3 by 3 by 5 I will play.
- 4:27:52Right? So, you will calculate this
- 4:27:54entropy when outlook is sunny.
- 4:27:57Yeah? So, you got this entropy when the
- 4:27:59outlook is sunny is 0.971.
- 4:28:02Accordingly, you will proceed with
- 4:28:04calculating the entropy when the outlook
- 4:28:07is overcast. Outlook is overcast, always
- 4:28:10you are playing.
- 4:28:11Right? If the outlook is overcast, every
- 4:28:13time you are playing, so it means that
- 4:28:16you will get a probability of zero. If
- 4:28:18you apply in that formula, you will get
- 4:28:20zero.
- 4:28:20And third, what what what would happen
- 4:28:23if the
- 4:28:24if the outlook is sunny? In that case,
- 4:28:27you will again multiply and you will put
- 4:28:29the formula and you will get it out
- 4:28:310.971, right? So, you will find it out
- 4:28:34entropy for each and every distinct
- 4:28:37combination of your feature.
- 4:28:40You got my point, right? Outlook being
- 4:28:42sunny, outlook being overcast, outlook
- 4:28:44being sunny
- 4:28:46uh you know,
- 4:28:47overcast, sunny, and rainy. This should
- 4:28:49be replaced here, right? And then
- 4:28:52what you will do is you will finally
- 4:28:54calculate the information gain.
- 4:28:56Information gain is nothing but what you
- 4:28:58will do is you will pick it up the
- 4:29:00chances, you will pick it up the entire
- 4:29:03chances when you are playing, right?
- 4:29:05Which is five out of 14. You remember
- 4:29:07five out of 14 were the total chances
- 4:29:09that I will play into if it is sunny,
- 4:29:13plus four out of 14 if it is overcast,
- 4:29:16five out of 14 if it is rainy, right?
- 4:29:20And you will calculate the information
- 4:29:22from this outlook. And once you
- 4:29:24calculate the information, you will
- 4:29:25subtract this information from your
- 4:29:28total entropy which I found it out in
- 4:29:30the last slide which was 0.94. You will
- 4:29:33subtract this and you will get the
- 4:29:35information gain. Or this is also called
- 4:29:37as a information gain from a particular
- 4:29:40feature. So, basically these type of a
- 4:29:43calculation first of all, don't get
- 4:29:44scared away that you have to do this
- 4:29:46calculation. But what I'm trying to
- 4:29:48explain you is this is how your
- 4:29:50algorithm will work for each and every
- 4:29:53feature. It will calculate the
- 4:29:55information gain from your feature.
- 4:29:59Right? If you have the more information
- 4:30:01gain, it means that this variable is of
- 4:30:04very very important in predicting that
- 4:30:07something will happen or not. Okay? So,
- 4:30:10for outlook, I got the information gain
- 4:30:13as 0.247 with all the calculation.
- 4:30:16Remember then we will proceed with wind
- 4:30:19if the wind is there or not, right? And
- 4:30:23then I will proceed with wind and I will
- 4:30:25calculate the information gain and say
- 4:30:27information gain I found it out is
- 4:30:290.048, right? Similarly, we will
- 4:30:31calculate for all the four of them. Let
- 4:30:34me put it together for all of you. So,
- 4:30:37this is what happens here. Now, if I put
- 4:30:40in front of all of you, these were the
- 4:30:42four different variable I have, outlook,
- 4:30:45temperature, humidity and wind, right?
- 4:30:47I'm calculating the information gain for
- 4:30:50each one of them and the information
- 4:30:52gain I got for outlook is 0.247.
- 4:30:56So, if the information gain is highest
- 4:30:58in a particular feature, that feature
- 4:31:01will become your root node.
- 4:31:04You got my point? So, therefore, we will
- 4:31:06pick it up outlook as our root node.
- 4:31:09Similarly, for when the tree get
- 4:31:11started, later on as a branch node also,
- 4:31:14your information gain will be calculated
- 4:31:17and wherever whichever feature is giving
- 4:31:19you more information gain, that will be
- 4:31:20picked up later in the
- 4:31:23your tree also.
- 4:31:25So, this is how you will build your
- 4:31:26decision tree and you will finally get a
- 4:31:29decision tree like this, okay? So, now
- 4:31:32what I will do is
- 4:31:34you know, like I will quickly show you
- 4:31:36how do we do the decision tree
- 4:31:39in your Python, okay? So, I'll show you
- 4:31:42how do you do this in Python and then I
- 4:31:44will summarize for all of you that why
- 4:31:47decision trees or tree based algorithms
- 4:31:50are better than your
- 4:31:52other algorithm, okay? So, how do you
- 4:31:55choose
- 4:31:56basically that which algorithm you will
- 4:31:58select and when, okay? So, I'll I'll
- 4:32:00describe you this, but before let's jump
- 4:32:03on to Python and go there.
- 4:32:06So, what I have done here is like I had
- 4:32:08built the decision tree in front like
- 4:32:12you know already. So, quickly I will
- 4:32:15walk you through the commands. Okay? So,
- 4:32:17what happens here is that
- 4:32:19in Python
- 4:32:21as you would be well versed with this
- 4:32:22that we generally import packages in
- 4:32:25Python. So, I'm importing NumPy. I'm
- 4:32:27importing Matplotlib for plotting the
- 4:32:29chart. I'm importing various packages
- 4:32:32from your scikit-learn which is
- 4:32:35for machine learning purposes, right?
- 4:32:36So, I'm importing your label encoder,
- 4:32:39your decision tree classifier,
- 4:32:40classification report, and I also I'm
- 4:32:43importing your
- 4:32:44tree, right? So, we have all this which
- 4:32:48I'm importing, right? After I import,
- 4:32:50what I will do is I'm reading my my data
- 4:32:53set. So, I'm showing you this with the
- 4:32:55help of a Iris data set which is one of
- 4:32:58the very popular data set for building
- 4:33:00the a decision tree, right? For any
- 4:33:03particular source. So, imagine that this
- 4:33:04is my data set and I'm just showing you
- 4:33:06six rows of the data set where you know
- 4:33:08like
- 4:33:09I want to find it out whether a
- 4:33:11particular flower species will be
- 4:33:13setosa, versicolor, or virginica. I have
- 4:33:16three different flowers which I want to
- 4:33:18predict and on the basis of sepal
- 4:33:20length, petal length, sepal width, and
- 4:33:22petal width. So, basically I have
- 4:33:24different dimensions of flowers length
- 4:33:27and width and based on that I want to
- 4:33:29find it out whether the particular
- 4:33:31species will be setosa, versicolor, or
- 4:33:33virginica. You got my point, right? So,
- 4:33:35this is a data set. So, what I will do
- 4:33:37is I as you know with every machine
- 4:33:39learning data set we play with the data
- 4:33:41set. So, I'm checking the information
- 4:33:43here
- 4:33:44like what type of the data type it is.
- 4:33:46So, sepal length it is a float, petal
- 4:33:48length it is a float, and species is a
- 4:33:51object.
- 4:33:52Why object? Because this is a
- 4:33:54categorical column.
- 4:33:56So, after this I'm also checking whether
- 4:33:58there is any null value present in the
- 4:34:00data set or not because if there is any
- 4:34:02null value, then we have to get rid of
- 4:34:04that null value or we have to replace
- 4:34:06that null value with some imputed value,
- 4:34:09right? This is what I'm checking here.
- 4:34:11Once I do that, I'm also plotting this
- 4:34:13because we usually do visualization,
- 4:34:15right? To understand the data set
- 4:34:17better. So, what I'm doing here in the
- 4:34:18with the help of SNS, which is your
- 4:34:20SNS.pairplot, I'm plotting all the
- 4:34:23possible plots. So, basically, sepal
- 4:34:25length with sepal width, what type of
- 4:34:27the combination look like. So, setosa,
- 4:34:29versicolor, and virginica, there are
- 4:34:31three different species you can see in
- 4:34:32the data set. And this is what I'm
- 4:34:35getting a trend between your sepal
- 4:34:36length and sepal width.
- 4:34:38Similarly, this is basically from sepal
- 4:34:40width to petal length.
- 4:34:42Right? So, this is how I'm understanding
- 4:34:45the patterns or I'm understanding the
- 4:34:46relationship between the variables in
- 4:34:48the data set. So, this is what I'm
- 4:34:50understanding here. I'm also checking
- 4:34:52whether there's a correlation or not.
- 4:34:54Higher the shade, it means there will be
- 4:34:56a strong correlation. So, all of these
- 4:34:58things is being done as a part of
- 4:35:00exploratory data analysis before even
- 4:35:02you start your machine learning model,
- 4:35:04right? Once you do this, after that what
- 4:35:06I have to do is after that what I'm
- 4:35:08trying to do is I'm taking your species
- 4:35:11column, what I want to predict as my
- 4:35:13target variable, which is your dependent
- 4:35:15variable. And what I So, this is my the
- 4:35:18dependent variable and with the help of
- 4:35:20what I want to predict will be my
- 4:35:22independent variable. So, I'm calling X
- 4:35:25all my independent variable and I'm
- 4:35:27calling target as my dependent variable.
- 4:35:30Once I do this, then you will also think
- 4:35:32about it that your
- 4:35:35uh the variable which I want to predict
- 4:35:36is the flower species, right? But I want
- 4:35:39to convert this into zero and one class,
- 4:35:42zero, one, two class because your
- 4:35:44computer cannot understand text, right?
- 4:35:46Computer can only understand the
- 4:35:48numbers. So, what I'm doing here is I'm
- 4:35:51uh changing it to the
- 4:35:52uh I'm changing it to the class.
- 4:35:55Right? So, what I'm trying to do here is
- 4:35:57I'm saying, "Wherever it is setosa, it
- 4:35:59will will zero. Wherever it is
- 4:36:01virginica, it will become one. And it is
- 4:36:04versicolor as a third category, it will
- 4:36:06become two. So, now imagine my data set,
- 4:36:08my label to predict becomes zero, one,
- 4:36:11or two instead of three flower which was
- 4:36:13setosa, versicolor, and virginica. This
- 4:36:15is what I'm doing with the encoder here.
- 4:36:18Once I convert this, and this become my
- 4:36:20target, what I will do is as I told you
- 4:36:22that I will split this data set into
- 4:36:24training and test data set. So,
- 4:36:26basically 80% of the data is going into
- 4:36:29the training data set and remaining 20%
- 4:36:32I'm taking as a test data set.
- 4:36:34Okay?
- 4:36:36Now, I will call decision tree
- 4:36:37classifier.
- 4:36:38You know, like I want to make a decision
- 4:36:40tree, and I want to fit this on my
- 4:36:42training data set and the test data set.
- 4:36:45And I will start building the decision
- 4:36:48tree and with the help of, you know,
- 4:36:50from So, here is when I have created the
- 4:36:53decision tree, and here is when I'm
- 4:36:55checking the prediction, how accurate my
- 4:36:57decision tree is. So, I got my
- 4:37:00precision, which is good. I got my
- 4:37:01recall. I got my F1 score, and I also
- 4:37:04get the support, right? So, all these
- 4:37:06matrices are being used to calculate
- 4:37:08your
- 4:37:09how accurate your predictions are. So,
- 4:37:12higher the precision, better the results
- 4:37:14would be. Right? And finally, I'm
- 4:37:16showing to you that how does the tree
- 4:37:18look like? So, I had tried to plot this
- 4:37:21decision tree in front of you. So,
- 4:37:22basically it start with petal length,
- 4:37:24and you can see the gini index coming up
- 4:37:26here or the information gain, which is
- 4:37:28information gain and gini index are
- 4:37:30reciprocal to each other. If you want to
- 4:37:31use information gain, you will get that
- 4:37:33score. If you don't want to use
- 4:37:35information gain, you will get gini
- 4:37:36index. So, they are both reciprocal of
- 4:37:38each other. It is one in the same thing.
- 4:37:40You use gini index or you use
- 4:37:42information gain. They are the two
- 4:37:44different metrics to build your decision
- 4:37:46tree.
- 4:37:47So, you can see that
- 4:37:49based on a petal length, the tree gets
- 4:37:51splitted like this. Then based on the
- 4:37:53petal length
- 4:37:54of different dimensions, your tree
- 4:37:56further split, then it further split,
- 4:37:59then it further split, and finally you
- 4:38:01get to know that whether the particular
- 4:38:03species will be versicolor or virginica.
- 4:38:07Okay? So, this is how your entire
- 4:38:10decision tree is built in Python. Okay?
- 4:38:14So, this is how it's so simple. Uh I
- 4:38:16know it takes time to build this thing,
- 4:38:19but once you are a good data scientist
- 4:38:21and you understand all these things, it
- 4:38:23is very simple to build all these things
- 4:38:26very easily in Python or R.
- 4:38:29Right? So, now
- 4:38:31uh finally going back how would you
- 4:38:34decide how would you decide that which
- 4:38:36algorithm, you know, which algorithm
- 4:38:39will be taken when?
- 4:38:41So, uh what happens is this is on the
- 4:38:44basis of scikit-learn. So, it it starts
- 4:38:48something like this, okay? So, it is on
- 4:38:50the basis of like this that first of all
- 4:38:52you see here that uh
- 4:38:56whether how many samples you have, how
- 4:38:58many data points you have in the data
- 4:39:00set, right? If you have more than 50
- 4:39:02samples, then you will go here. If you
- 4:39:05have less than 50 data point, then you
- 4:39:07know, you will go here. So, you will get
- 4:39:10more data set, right? So, if you have
- 4:39:12more data If you have
- 4:39:13greater than 50 data point, then you
- 4:39:15will go further. You split on your data
- 4:39:17set. If it is not, then you will you
- 4:39:19just the kind of advice that get more
- 4:39:21data set, okay? So, then you will decide
- 4:39:24that what you want to predict. If you
- 4:39:25have a labeled data set, then you will
- 4:39:27go here, right? And do clustering. If
- 4:39:30you don't have the labeled data set,
- 4:39:31then you will see whether you want to
- 4:39:33predict a quantity. If yes, you will go
- 4:39:35in regression. If you want to predict uh
- 4:39:37if you just want to do exploratory
- 4:39:39analysis, you can do dimensional
- 4:39:40reduction. If you have a labeled data
- 4:39:42set, uh you know, for classification,
- 4:39:45you can do all this classification. So,
- 4:39:47basically, this is a cheat sheet which
- 4:39:49we generally use to decide that what we
- 4:39:51have to do with the data set and when.
- 4:39:59Let's understand what is a random
- 4:40:01forest.
- 4:40:02A random forest is constructed by using
- 4:40:05multiple decision trees and the final
- 4:40:08decision is obtained by majority votes
- 4:40:11of these decision trees. So, let me make
- 4:40:13things very simple for you by taking an
- 4:40:16example. Now, suppose we have got three
- 4:40:18independent decision trees. Here we are
- 4:40:20just taking three decision trees and
- 4:40:22I've got an unknown fruit and I want
- 4:40:24that these trees would give me a result
- 4:40:27of what exactly this fruit is. So, I
- 4:40:29pass this fruit to the first decision
- 4:40:31tree, the second decision tree, and the
- 4:40:33third decision tree. Now, a random
- 4:40:35forest is nothing but a combination of
- 4:40:37these decision trees. So, the results
- 4:40:40are being fed into the random forest
- 4:40:42algorithm. So, what it sees is that,
- 4:40:45okay, the first decision tree classifies
- 4:40:47it as peach, the second decision tree
- 4:40:49says that it is an apple, and the third
- 4:40:51one says that it is a peach. So, random
- 4:40:54forest classifier says that, okay, I've
- 4:40:56got the result as two peach and one for
- 4:41:01an apple. So, I would say that the
- 4:41:03unknown fruit is an peach.
- 4:41:06All right, so this is based on the
- 4:41:08majority voting of the decision trees
- 4:41:10and that is how a random forest
- 4:41:12classifier comes to a decision of
- 4:41:14predicting the unknown value.
- 4:41:16Okay, so this was a classification
- 4:41:18problem, so it took the majority vote.
- 4:41:20Now, suppose if it was in regression
- 4:41:22problem, it would have taken mean of it,
- 4:41:24okay? So, now let's move on further to
- 4:41:27understanding what is a decision tree.
- 4:41:29But before that, we should understand
- 4:41:30that random forest the building blocks
- 4:41:33are decision trees and that's why
- 4:41:35studying decision tree becomes important
- 4:41:37because if we understand one decision
- 4:41:40tree, we can apply the same concept to
- 4:41:42random forest, okay? So, now let's move
- 4:41:45on forward and understand the important
- 4:41:47terms in random forest. And this will
- 4:41:49also help us consolidate whatever we
- 4:41:51have learned so far. So, we have taken
- 4:41:53the same small decision tree of the
- 4:41:55previous example, and let's understand
- 4:41:57these are also the important terms which
- 4:41:59will be relevant to random forest also.
- 4:42:01So, the first is the root node. Now,
- 4:42:03here what happens is that the entire
- 4:42:05training data has been fed to the root
- 4:42:07node. And then we've got here that each
- 4:42:10node will ask either true or false
- 4:42:12question with respect to one of the
- 4:42:14feature. And then in response to that
- 4:42:16question, it will partition the data set
- 4:42:18into different subsets. That's what it
- 4:42:20is it is doing here based on the
- 4:42:23condition that it if the mass body mass
- 4:42:25is greater than equal to 2500, it ask a
- 4:42:27question either yes or no. And based on
- 4:42:30that, again further partition is done.
- 4:42:32And if not, then it just classifies the
- 4:42:34species. And then again, what happens is
- 4:42:37that the splitting Now, this is very
- 4:42:39important here. The splitting takes
- 4:42:41place either with the help of a genie or
- 4:42:43entropy methods. And these helps to
- 4:42:45decide the optimal split.
- 4:42:48And we will be discussing about
- 4:42:49splitting methods very soon, right?
- 4:42:51Okay. And then we've got the decision
- 4:42:53nodes which provide the link to the leaf
- 4:42:56nodes. And these are really important
- 4:42:57because then only the leaf nodes will
- 4:43:00tell us what actually the real
- 4:43:02predictions are to which class does the
- 4:43:05species belong. So, now coming to the
- 4:43:08leaf node and these are the end points
- 4:43:10where no further division will take
- 4:43:11place and we will obtain our
- 4:43:13predictions. Okay? So, now coming up to
- 4:43:17another important thing here is working
- 4:43:19of random forest. So, now for working of
- 4:43:22random forest, we will have to
- 4:43:23understand a few important concepts like
- 4:43:26random sampling with replacement,
- 4:43:28feature selection, and also the ensemble
- 4:43:30technique which is used in random forest
- 4:43:33and that is bootstrap aggregation which
- 4:43:35is also known as bagging. So, we will
- 4:43:37understand this with the help of an
- 4:43:39example which will be very simple. And
- 4:43:42then we will go on understanding how
- 4:43:44feature selection is done in both the
- 4:43:46classification and the regression
- 4:43:48problem. Actually, how random forest
- 4:43:50select features for the construction of
- 4:43:53decision trees. Well, in random forest,
- 4:43:55the best split is chosen based on Gini
- 4:43:57impurity or information gain methods.
- 4:44:00So, this also we will understand. Now,
- 4:44:03let us first understand random sampling
- 4:44:05with replacement. Now, what happens here
- 4:44:07is that we have got a small subset of
- 4:44:09the same penguin data set, wherein we
- 4:44:11have got some six rows and four
- 4:44:14features, that means four columns. And
- 4:44:16the arrows that you can see is that now
- 4:44:18we will be creating three subsets from
- 4:44:21this small subset, right? And these
- 4:44:23three subsets will become our decision
- 4:44:25trees. And then we'll be constructing
- 4:44:27decision trees from these subsets. So,
- 4:44:29let us create our first subset. And you
- 4:44:32can see here that the subset is randomly
- 4:44:34being created. And for convenience'
- 4:44:36sake, let me just also show you the
- 4:44:39different subsets here. Okay. So, now
- 4:44:41for better understanding, let us
- 4:44:42understand this that in the first
- 4:44:44subset, if we focus, we've got certain
- 4:44:47random rows here, and we have got
- 4:44:49certain feature. But we do not know how
- 4:44:52this feature has been selected. We got
- 4:44:54island and we got body mass. But in the
- 4:44:56second subset, we got island and flipper
- 4:44:59length. And in the third subset, we got
- 4:45:02body mass and flipper length, right?
- 4:45:04Now, let's look at the rows. Now, when I
- 4:45:07am talking about these features, I will
- 4:45:09say this is feature selection, and
- 4:45:11remember this term. Now, coming to the
- 4:45:13second concept, that is random sampling.
- 4:45:16Now, random sampling is nothing but
- 4:45:17selecting randomly from your subset. So,
- 4:45:21I'm selecting randomly certain rows from
- 4:45:24my subset and creating further subset,
- 4:45:27okay? So, what is replacement here?
- 4:45:30Replacement is can be seen here and can
- 4:45:32be understood with the second subset. We
- 4:45:35see here that the Gentoo species, this
- 4:45:37is being repeated again. And this is
- 4:45:40replacement. That means that when we are
- 4:45:43working with repeated rows, and this row
- 4:45:46can be repeated again in the second or
- 4:45:49the third subset, then this is random
- 4:45:51sampling with replacement. That means my
- 4:45:53random forest can use a row multiple
- 4:45:56times in multiple decision trees, right?
- 4:45:59So, this is the basic concept of random
- 4:46:01sampling with replacement and feature
- 4:46:03selection in random forest. Another
- 4:46:06important term which I would like to
- 4:46:08bring into the notice is that
- 4:46:10when we are working with these type of
- 4:46:13small subsets, these are also known as a
- 4:46:16bootstrap data sets. And when we
- 4:46:18aggregate the results of all these data
- 4:46:20set, it becomes bootstrap aggregation.
- 4:46:23So, just filling in the gaps so that
- 4:46:25later on the concepts become more clear.
- 4:46:28So, now let's move on to drawing
- 4:46:30decision trees of these subsets, okay?
- 4:46:33So, let's draw the decision tree of the
- 4:46:35first subset. Again, we are taking body
- 4:46:38mass as the first root node, and then
- 4:46:40based on a decision like if the mass is
- 4:46:43greater than equal to 3,500, then take a
- 4:46:45decision either yes or no. If it is no,
- 4:46:47then the specie is chinstrap. And if it
- 4:46:50is yes, then again you partition based
- 4:46:52on island. And if it is Torgersen, then
- 4:46:55it is Adélie. And if it is Biscoe, then
- 4:46:58it is Gentoo species. Okay? So, this is
- 4:47:01how we will construct two more decision
- 4:47:03trees of the remaining subsets. So, in
- 4:47:05the second subset, let us just again
- 4:47:07create decision tree. And here now we
- 4:47:10are taking flipper length, and then
- 4:47:12based on a condition that if the flipper
- 4:47:15length is greater than equal to 190,
- 4:47:17then make a split. If it is yes, then
- 4:47:19the specie become Gentoo. And if it is
- 4:47:22no, that means again make a decision
- 4:47:25based on island. And if it is Torgersen,
- 4:47:27it is Adélie. And if the island is Dream
- 4:47:31Island, then it is a chinstrap species.
- 4:47:32So, this is how the decision tree of the
- 4:47:35second subset has been created and this
- 4:47:37is how it will take decisions, right?
- 4:47:40Based on the tree length, depth, and
- 4:47:42also the features it is selecting, okay?
- 4:47:46So, now let's create the third decision
- 4:47:47tree of the third subset and we get a
- 4:47:49decision tree something like this
- 4:47:51wherein body mass if it is greater than
- 4:47:534,000 and if it is yes, then clearly it
- 4:47:56is a Gentoo species and if it is no,
- 4:47:59then again make a partition with the
- 4:48:01with respect to flipper length, another
- 4:48:03feature here, and then if it is again
- 4:48:06greater than equal to 190, then the
- 4:48:08species would be Adélie, else it would
- 4:48:10be chinstrap. So, this is how decision
- 4:48:13tree three will make a decision.
- 4:48:15Now, let's just keep these decision
- 4:48:17trees with us, okay? And we will make
- 4:48:20sense of these trees just in a while.
- 4:48:23Okay?
- 4:48:24But before that, let us understand how
- 4:48:26feature selection is done in a random
- 4:48:28forest. How am I selecting the columns?
- 4:48:31So, for classification, by default the
- 4:48:33feature selection is taken as the square
- 4:48:35root of total number of all the
- 4:48:36features. Now, suppose I've got here
- 4:48:39four features, so it is a classification
- 4:48:41problem, I will take the square root of
- 4:48:43these four features, which becomes two.
- 4:48:45So, decision tree would be constructed
- 4:48:46based on two features each. If suppose I
- 4:48:49had 16 features, then it would be square
- 4:48:51root of 16, that would be four. So, four
- 4:48:53features would be taken in each decision
- 4:48:55tree. All right? And suppose if this
- 4:48:58would have been a regression problem,
- 4:49:00then by default what would happen? The
- 4:49:01features would be selected by taking the
- 4:49:04total number of features and dividing
- 4:49:06them by three, okay? So, this is how by
- 4:49:08default the feature selection is being
- 4:49:10done by a random forest. Okay, now let
- 4:49:13us move on forward to consolidating our
- 4:49:15learning.
- 4:49:16So, now we are coming to ensemble
- 4:49:18techniques, that is also known as
- 4:49:20bootstrap aggregation.
- 4:49:22Random forest uses ensemble techniques.
- 4:49:25And what is ensembling? It just means
- 4:49:27that you're aggregating the result of
- 4:49:29the decision trees and taking the
- 4:49:31majority vote in case of classification
- 4:49:34and the mean in case of regression
- 4:49:35problems and giving the output. Okay. So
- 4:49:38now we have again plotted all our
- 4:49:41decision trees here. And below we can
- 4:49:44see that there's an unknown data and I
- 4:49:47want to predict the species of this
- 4:49:48data. So what will happen is that again
- 4:49:51let us just feed this problem to each of
- 4:49:54the decision trees. And let's see what
- 4:49:57each decision tree makes the prediction.
- 4:49:59So I just feed this unknown data to
- 4:50:01decision tree one and it says that okay
- 4:50:03the species seems to be chinstrap. Okay.
- 4:50:06And then decision tree two says that
- 4:50:08based on the data it has been found that
- 4:50:11the species Adélie. And then decision
- 4:50:14tree three says that no I I with my
- 4:50:16decision tree this species is chinstrap.
- 4:50:19Okay. Now all these data is being fed to
- 4:50:23random forest classifier. And it says
- 4:50:25that okay for chinstrap I've got two
- 4:50:28votes for Adélie it's got one vote. So
- 4:50:31the new species would be chinstrap,
- 4:50:33right? So this is how the bootstrap
- 4:50:35aggregation is done based on the
- 4:50:38majority voting and the decisions taken
- 4:50:41by different decision trees they have
- 4:50:43been combined together aggregated and we
- 4:50:46get an ensemble result in the random
- 4:50:49forest. Okay. So this was very simple
- 4:50:51concept of ensemble techniques which has
- 4:50:53been used in random forest.
- 4:50:56Okay. So now let's move on forward to
- 4:50:58splitting methods. So what are the
- 4:51:00splitting methods that we use in random
- 4:51:02forest? So splitting methods are many
- 4:51:04like Gini impurity, information gain or
- 4:51:07chi-square. So let's discuss about Gini
- 4:51:09impurity. So Gini impurity is nothing
- 4:51:12but it is used to predict the likelihood
- 4:51:14that a randomly selected example would
- 4:51:16be incorrectly classified by a specific
- 4:51:19node. And it is called impurity metric
- 4:51:21because it shows how the model differs
- 4:51:24from a pure division, right? And another
- 4:51:26interesting fact about Gini impurity is
- 4:51:28that the impurity ranges from zero to
- 4:51:31one with zero indicating that all of the
- 4:51:33elements belong to a single class and
- 4:51:36one indicates that only one class exist.
- 4:51:39Now value which is like 0.5, this
- 4:51:42indicates that the elements they are
- 4:51:44uniformly distributed across some
- 4:51:46classes, right? Now moving on forward to
- 4:51:49information gain. Now this is another
- 4:51:51method which random forest can use and
- 4:51:54information gain utilizes entropy. So
- 4:51:57entropy is nothing but it is a measure
- 4:51:59of uncertainty. So information gain
- 4:52:01let's talk about that first. So the
- 4:52:04features they are selected that provide
- 4:52:06most of the information about a class,
- 4:52:08right? And this utilizes the entropy
- 4:52:10concept. So let's see what is entropy.
- 4:52:14This is a measure of randomness or
- 4:52:16uncertainty in the data, right? So we
- 4:52:19will understand this entropy with the
- 4:52:20help of a small example. So don't worry
- 4:52:22about it. So let's understand this
- 4:52:24entropy. Now suppose there's a fruit
- 4:52:26fruit tray with four different fruits,
- 4:52:28right? And what do you feel about the
- 4:52:31entropy here? That means the randomness
- 4:52:33of the data. Is it really easy to
- 4:52:36classify these fruits into the
- 4:52:38respective class? So this becomes really
- 4:52:40uncertain and the data looks messy here.
- 4:52:43But what if we just split here these
- 4:52:45into two trays where in the first tray
- 4:52:48would have peaches and oranges and the
- 4:52:50second tray will have apples and lemons.
- 4:52:53So now this becomes a little more
- 4:52:55certain. We get low randomness here and
- 4:52:58this is called as low entropy. So when
- 4:53:01we move down from tree, that means from
- 4:53:03root node to the leaf nodes, the entropy
- 4:53:06reduces and we can also calculate
- 4:53:09information gain from this entropy. That
- 4:53:12is the difference in entropy before and
- 4:53:14after to split that is known as
- 4:53:16information gain. Okay? So, once we move
- 4:53:19down the tree and start reducing the
- 4:53:21randomness from the data, the entropy
- 4:53:23becomes lower and that is what we want
- 4:53:26in our data. If there's low entropy,
- 4:53:28that means we are likely that the
- 4:53:30predictions would be more accurate and
- 4:53:33we can make predictions very easily as
- 4:53:35compared to very messy data which has
- 4:53:38high entropy. Okay? So, that was about
- 4:53:40entropy and now let us just move on to
- 4:53:44the practical demonstration or a
- 4:53:45hands-on on random forest.
- 4:53:48Okay. So, now it's time for a hands-on
- 4:53:50on random forest. So, let us just import
- 4:53:53a few basic libraries of Python in our
- 4:53:55Jupyter notebook. And we will run this.
- 4:53:58We will import pandas as pd, numpy as
- 4:54:00np, and seaborn as sns. Now, seaborn is
- 4:54:04needed here because we want to load a
- 4:54:05data set, that is a penguins data set
- 4:54:08with the help of seaborn. And this has
- 4:54:10already been preloaded in seaborn. This
- 4:54:12is already loaded data set and seaborn
- 4:54:15has got multiple data sets, you know,
- 4:54:16for practice for beginners. So, it is a
- 4:54:18good way to practice for data sets. Now,
- 4:54:21we can see this asterisk sign that means
- 4:54:23it is telling us to wait. So, let us
- 4:54:25just let it get loaded. So, we get got
- 4:54:27our data in an object called df and we
- 4:54:29can see the first five entries here. And
- 4:54:32this uh data frame is shown in the form
- 4:54:34of a table, rows and columns. And we see
- 4:54:37here some species, island, bill length,
- 4:54:39bill depth, flipper length, body mass,
- 4:54:41and the sex of the penguin. So, our task
- 4:54:43is to specify or to classify these
- 4:54:46species of penguins into their
- 4:54:48respective correct species, right? So,
- 4:54:51we see the shape of our data and we see
- 4:54:53that it is like 344 rows and seven
- 4:54:55columns.
- 4:54:56And we will see the info. So, we see
- 4:54:58df.info and this gives us, along with
- 4:55:01the non-null count, we also get the data
- 4:55:04type of the values. So, we have got
- 4:55:07species, island as the object data type,
- 4:55:09whereas the bill length, bill depth,
- 4:55:11flipper length, and body mass are in
- 4:55:13floating point, or or you can say
- 4:55:15floating data type. And the sex is in
- 4:55:18object data type, right? So, now moving
- 4:55:20on forward to calculating how many null
- 4:55:22values are there with the help of
- 4:55:24df.isnull.sum.
- 4:55:26So, we get certain like some around two
- 4:55:29null values in all these columns, as you
- 4:55:32can see the features like bill length,
- 4:55:34bill depth, flipper length, and body
- 4:55:36mass. Whereas there are 11 null values
- 4:55:38in sex feature, right? So, what we do is
- 4:55:40since they are very small null values,
- 4:55:42we can just drop it, or you can also
- 4:55:44ignore them. So, here in this data
- 4:55:46frame, what I'm doing is I'm just
- 4:55:47dropping these null values, and let us
- 4:55:50just check whether they are they are
- 4:55:51being dropped or not with the help of
- 4:55:53again the same function {dot} is null
- 4:55:55{dot} sum. And then we see that yes,
- 4:55:58they are being dropped from our data
- 4:55:59frame. Now, let us do some feature
- 4:56:01engineering with our data. Now, we have
- 4:56:03seen that we have got some object data
- 4:56:05type in our data frame. And before
- 4:56:07feeding it into algorithm that is random
- 4:56:10forest, we have to transform the
- 4:56:13categorical data or the object data type
- 4:56:15into the numeric. So, we are using here
- 4:56:17one-hot encoding to convert the
- 4:56:19categorical data into numeric. Now,
- 4:56:21there are various ways in Python which
- 4:56:22we can do that, like one-hot encoding or
- 4:56:25you can also use mapping function in
- 4:56:27Python, but here we are using one-hot
- 4:56:29encoding. So, let us just do that.
- 4:56:31And we find here first of all, let us
- 4:56:33apply it on the sex column. And here we
- 4:56:36see that we have got two unique values
- 4:56:37in sex, that is male and female. And we
- 4:56:40use pandas here to get dummies, that is
- 4:56:42how we will apply this one-hot encoding
- 4:56:45because this is how get dummies work.
- 4:56:47So, what happens is here is that the new
- 4:56:50unique values are converted into the
- 4:56:52respective columns in the data frame.
- 4:56:54So, we see here we have got two unique
- 4:56:56values, males and females, and they are
- 4:56:58being converted into the columns. Okay?
- 4:57:00So, one thing to note here is that we
- 4:57:03also get a problem of dummy trap because
- 4:57:06here we see only two unique values. Now,
- 4:57:08suppose if I had six or seven unique
- 4:57:10values and I do this one hot encoding, I
- 4:57:13would have lots of features in my data
- 4:57:16frame and that would lead to several
- 4:57:18complexities. So, what I do is
- 4:57:21to keep things simple, I can use one hot
- 4:57:23encoding when my data frame or my unique
- 4:57:25counts are low, when my unique values
- 4:57:27are less. So, since I had just two or
- 4:57:30three, I can use it. So, I'm using here.
- 4:57:33So, what I do is again, now one row, one
- 4:57:35column as we can see here that it is
- 4:57:37redundant, giving me extra information,
- 4:57:39so I will just drop it. So, I drop this
- 4:57:42first column and what I get in this data
- 4:57:44frame is only male. So, let us just
- 4:57:47infer whether I can also infer females
- 4:57:49from this or not. So, if the value is
- 4:57:51one, that means the penguin is a male
- 4:57:53and if the value is zero, that means the
- 4:57:55penguin is a female. Okay? So, only one
- 4:57:58column is needed for this data frame.
- 4:58:01So, I just kept one and dropped the
- 4:58:02other one. Okay, now apply again one hot
- 4:58:05encoding to the island feature. So, in
- 4:58:08island if we check the unique values,
- 4:58:10we've got three unique values here.
- 4:58:12Torgersen, Biscoe and Dream Island and
- 4:58:14the object is the data type, right? So,
- 4:58:17again we will use pandas, pd.get_dummies
- 4:58:21and we will use apply it on the feature
- 4:58:24island and let's get the head of it. So,
- 4:58:27we get here again the unique values were
- 4:58:29converted into columns and we get here
- 4:58:31expected three columns. And then again
- 4:58:33we will just drop the first column to
- 4:58:35get the remaining two columns. So, here
- 4:58:38also we can infer that if the island is
- 4:58:40Torgersen, if it is one, then it is not
- 4:58:43Dream, neither Biscoe, right? So, this
- 4:58:45is how we can read it from the data
- 4:58:46frame and understand that. Now, remember
- 4:58:48this thing that these two island and
- 4:58:50here sex, these are two independent data
- 4:58:53frames. These are not yet included in
- 4:58:55the main data frame. So, what we will do
- 4:58:57now is we will concatenate the above two
- 4:58:59data frames into the original data
- 4:59:01frame. So, what we do, we again create a
- 4:59:03new data frame that is new data and let
- 4:59:05us just concat with the help of
- 4:59:06pd.concat function and we will concat
- 4:59:09what? df.island and sex. And axis is one
- 4:59:13that means in the column. Okay. So, when
- 4:59:15we will run this, let's see the head of
- 4:59:17it. So, everything gets concatenated in
- 4:59:20a single data frame which is good for
- 4:59:22the feeding this data into or splitting
- 4:59:24the data into test and train data. So,
- 4:59:26now we have this new data frame and
- 4:59:29we've got some repeated columns here
- 4:59:31which needs to be deleted. So, what we
- 4:59:33do is we will delete sex and island here
- 4:59:36which are just repeating because we've
- 4:59:37got here male and we have also got here
- 4:59:40dream and togerson. So, we do not
- 4:59:42require this island column neither this
- 4:59:44sex. So, we just drop it with the help
- 4:59:46of new data.drop and the column names x
- 4:59:49is one in place equals to true, right?
- 4:59:51And let's see the head of this data
- 4:59:53frame. Head of the data frame gives me
- 4:59:55five unique values, right? And now it is
- 4:59:58time to create a separate target
- 5:00:00variable. And what we'll do is we will
- 5:00:03store in a variable called y only
- 5:00:05species. So, what we do is from this new
- 5:00:08data.species, we will just store the
- 5:00:10species in this y. And we see this
- 5:00:12y.head that is the first five species
- 5:00:15and we got the values here. That means
- 5:00:17another target variable has been created
- 5:00:19now.
- 5:00:20So, and you can also see the y.unique
- 5:00:23values as Adelie, Chinstrap and Gentoo.
- 5:00:25So, now we see here three unique values
- 5:00:28of the penguin that is Chinstrap, Adelie
- 5:00:30and Gentoo. And the data type is object
- 5:00:32here. So, again we need to convert this
- 5:00:34object into the numeric data type. So,
- 5:00:36now what we are doing is we are using
- 5:00:38the map function in Python and what we
- 5:00:40do is we map Adelie to zero, Chinstrap
- 5:00:42to one and Gentoo to two. So, this is
- 5:00:45how we see then all the values are being
- 5:00:48mapped to numeric. This is another way
- 5:00:50to convert a categorical value into a
- 5:00:52numeric value in Python. Now, what we do
- 5:00:54is let us just drop the target value
- 5:00:57species from our main data frame. So, we
- 5:00:59just drop it and let's see our new data
- 5:01:01frame. So, we see that we don't have any
- 5:01:04target species here, right? Okay. So, in
- 5:01:07X, let's store this new data and perform
- 5:01:10the splitting of the data. So, what we
- 5:01:13do is from sklearn.model_selection,
- 5:01:15we will import our train_test_split
- 5:01:18and we will split our training data into
- 5:01:2070% and 30%. So, test data becomes 30%
- 5:01:23and training data is some 70%. And this
- 5:01:26random state is zero, which means that
- 5:01:28I'm not fixing any random state. And
- 5:01:31this is also useful for the code
- 5:01:33reproducibility. Now, suppose if I again
- 5:01:35run this code, I will get the same
- 5:01:36result. It will not change. You can set
- 5:01:39this random state to any of the random
- 5:01:40number as per your choice and result
- 5:01:43would differ. Okay. So, now let us print
- 5:01:45the shape of X_train, Y_train, X_test
- 5:01:48and Y_test. So, we see here that it has
- 5:01:50been splitted into 70 and 30% and we get
- 5:01:53X_train as 233 values here and seven
- 5:01:56features. And X_test has 100 values and
- 5:01:59seven features. Similarly, Y_train you
- 5:02:01can see 233 values and Y_test has 100
- 5:02:04values. That means the species. Okay.
- 5:02:07So, that has been perfectly splitted
- 5:02:09into 70 and 30%. Now, what we do is we
- 5:02:12will train the random forest classifier
- 5:02:14on the training set. How do we do it? We
- 5:02:17will import the random forest classifier
- 5:02:19from sklearn.ensemble.
- 5:02:21So, we've already dealt with what is
- 5:02:23ensemble. And then in classifier, we
- 5:02:26will store this random forest and this
- 5:02:28n_estimators is nothing but decision
- 5:02:30tree. So, we are creating some five
- 5:02:31decision trees here. And the criteria is
- 5:02:33entropy. And again, random state is set
- 5:02:36to zero. So, let's see. And then we will
- 5:02:38fit this X_train and Y_train. So, this
- 5:02:40has been fitted and the criteria is
- 5:02:42entropy here. All right. So, now let's
- 5:02:44make some predictions and let's create a
- 5:02:46variable called Y_predict and we will
- 5:02:49just predict it on X_test. And we've
- 5:02:51also printed this Y prediction and now
- 5:02:54let's bring the confusion matrix to
- 5:02:56check the accuracy of random forest
- 5:02:58algorithm. And what we do is from
- 5:03:00matrices as sklearn matrices, we will
- 5:03:02import classification report and
- 5:03:04confusion matrix and also the accuracy
- 5:03:06score. So, we will just import them and
- 5:03:09then in CM variable, we will print the
- 5:03:12confusion of Y test and Y predictions.
- 5:03:15So, we will print it and we see here the
- 5:03:17accuracy score also, which is 98%. So,
- 5:03:21our random forest classifier is giving
- 5:03:23us a very good accuracy of 98% and you
- 5:03:26can see your confusion matrix that only
- 5:03:27two cases have been misclassified. Rest
- 5:03:30all the cases have been correctly
- 5:03:32classified by random forest classifier.
- 5:03:34Okay. So, now let's move on to printing
- 5:03:36the classification report of Y test and
- 5:03:38Y prediction. Let's see and we get the
- 5:03:41precision as 96%. That means
- 5:03:43the two predictions by the algorithm is
- 5:03:4596%. The recall or the true prediction
- 5:03:48rate is 100% which is very nice and Evan
- 5:03:51score is also good which is 98%. So,
- 5:03:54this is giving us a good result. But
- 5:03:56what if if we change the criteria from
- 5:03:58entropy to gini? So, let's just
- 5:04:00experiment with that too. So, let's try
- 5:04:03this with the different number of trees
- 5:04:04and change the criteria to gini
- 5:04:07coefficient. So, now again from
- 5:04:09sklearn.ensemble, we will import random
- 5:04:11forest classifier and fit it, okay? And
- 5:04:14here what we are doing is just we are
- 5:04:16using seven trees. Previously we used
- 5:04:18five and now in the criteria, we will
- 5:04:20use gini coefficient and random state is
- 5:04:23zero. So, let's run this and see whether
- 5:04:25there's a change in accuracy or not and
- 5:04:27let's predict this and let's check the
- 5:04:30accuracy score. What is the accuracy
- 5:04:32score for this random forest classifier
- 5:04:34with seven trees? So, we get 99%
- 5:04:36accuracy with changing the criteria and
- 5:04:39changing the number of trees. So, you
- 5:04:40can just experiment with different
- 5:04:42number of trees and different number of
- 5:04:44decision trees. Let's just experiment
- 5:04:45with, you know, 12 decision trees and
- 5:04:48see what happens.
- 5:04:50So, you can see the accuracy reduced to
- 5:04:5198%. Okay? With seven, we were getting
- 5:04:5599. So, let's just keep seven because it
- 5:04:57is giving us really good accuracy. So,
- 5:04:59this is about random forest classifier
- 5:05:02and how it works with several trees and
- 5:05:05different criteria to give us very good
- 5:05:07accuracy on our training and test data.
- 5:05:15>> [music]
- 5:05:15>> The case that we are discussing is
- 5:05:19basically the KNN algorithm, which is
- 5:05:22an algorithm which we used for mostly
- 5:05:25machine
- 5:05:26learning, which is where you have
- 5:05:28labeled data.
- 5:05:30Okay, so
- 5:05:33KNN algorithm. KNN algorithm is K
- 5:05:36nearest neighbors algorithm.
- 5:05:38And it is an example of supervised
- 5:05:40learning algorithm where
- 5:05:42basically, you try to classify a new
- 5:05:45data point based on the neighbors of
- 5:05:48that data point, which is basically
- 5:05:50which data points are closer to it. For
- 5:05:53example, here, as you can see,
- 5:05:56you have on one side couple of cats and
- 5:05:59on the other side you have couple of On
- 5:06:02one side you have dogs
- 5:06:04and on the other side you have cats.
- 5:06:07Right? Now, if a new data's point is
- 5:06:11given to us, there is a
- 5:06:13picture of a new animal.
- 5:06:15And
- 5:06:17if it is lying somewhere here, right?
- 5:06:20Then, we know that it is nearer to the
- 5:06:23cats, right? And therefore, we will
- 5:06:24classify it as cat.
- 5:06:26Whereas,
- 5:06:27if it is
- 5:06:29sort of
- 5:06:30uh
- 5:06:31nearer to the dogs, then we classify it
- 5:06:33as dog, right? So, that's the like, you
- 5:06:36know, neighborhood for the dog. And
- 5:06:38therefore, we sort of uh classify it
- 5:06:42that new animal or the new picture as as
- 5:06:45being
- 5:06:46all that of a dog. And this is quite
- 5:06:50you know, this is something which is
- 5:06:52even seen our in our regular day-to-day
- 5:06:54life, right? You know, we have had
- 5:06:57examples where our parents keep telling
- 5:06:59us, "Okay, don't play with those kinds
- 5:07:02of you know, children or something
- 5:07:04because they are not good in their
- 5:07:06studies or probably they are not
- 5:07:08so good in their behavior because you
- 5:07:10would become like them, right?" So, it's
- 5:07:12again an example from real life of
- 5:07:14classifying a particular person based on
- 5:07:16the company that they keep, right? Or
- 5:07:18from the
- 5:07:20with the kind of people that they are.
- 5:07:22So, that's that's sort of
- 5:07:26the example and now let me actually go
- 5:07:28back.
- 5:07:30Okay.
- 5:07:31What are the features of of K nearest
- 5:07:33neighbors algorithm?
- 5:07:36Okay, so
- 5:07:37let's talk about
- 5:07:39the features of KNN. So, as as I said,
- 5:07:42KNN is a supervised learning algorithm.
- 5:07:45Typically, it is you know,
- 5:07:47used for supervised learning kind of
- 5:07:49problems.
- 5:07:50It's very simple as we mentioned.
- 5:07:52Intuitively, you can you know, it's
- 5:07:55about you know, what kind of neighbors
- 5:07:56do you have? So, your class is predicted
- 5:07:59based on the your nearest neighbors as
- 5:08:01the name suggests. And then it's a
- 5:08:04non-parametric technique. So, I would
- 5:08:06like to spend a couple of minutes here
- 5:08:09to discuss about what we mean by by
- 5:08:12non-parametric. So, typically, you know,
- 5:08:15the supervised machine learning
- 5:08:16algorithms
- 5:08:18are of two kinds, right? One is the
- 5:08:20parametric types and the second one is
- 5:08:22the non-parametric type. When we say
- 5:08:24parametric, what we mean is basically
- 5:08:27that the machine or the algorithm
- 5:08:29assumes
- 5:08:31that there is an underlying function or
- 5:08:34a distribution that is known of the
- 5:08:36particular data set. Like for example,
- 5:08:39the linear regression would assume that
- 5:08:40the relationship between
- 5:08:43two two things X and Y is is linear in
- 5:08:46nature, right?
- 5:08:47Um and and and similarly for a
- 5:08:50distribution a Gaussian distribution it
- 5:08:52would assume a normal distribution of
- 5:08:54the data points and so on. So,
- 5:08:57a lot of the supervised machine learning
- 5:08:59algorithms they assume some kind of
- 5:09:01function association between the
- 5:09:04predictor which is basically the root
- 5:09:06causes which help you in predicting and
- 5:09:09the variable that you're trying to
- 5:09:10predict.
- 5:09:12Whereas there are certain algorithms
- 5:09:14like the KNN
- 5:09:15or the Parzen window or the linear
- 5:09:18discriminant analysis which is of the
- 5:09:20kind which is called non-parametric
- 5:09:22because it does not assume
- 5:09:24any particular kind of distribution or
- 5:09:26any particular kind of functional
- 5:09:28relationship of the data that you're
- 5:09:30trying to you know predict or you're
- 5:09:33trying to
- 5:09:34learn the the pattern of.
- 5:09:37So, KNN is one of the kind as I said
- 5:09:39Parzen window is another one
- 5:09:42uh which is basically where you uh in in
- 5:09:46Parzen window essentially, you know, the
- 5:09:49the volume of the data set uh or the
- 5:09:53area that the data set covers is known
- 5:09:55whereas you're trying to find the K
- 5:09:57there where uh which is the number of
- 5:09:59data points within that area or the
- 5:10:02volume, right? Whereas in case of
- 5:10:03non-parametric technique like KNN
- 5:10:06it's the opposite where the K is known
- 5:10:09which is basically you would like to
- 5:10:12associate the class of of
- 5:10:14the data point that you're trying to
- 5:10:15predict based on the K number of
- 5:10:19neighbors around it. Now, K can be
- 5:10:21three, four, five, whatever, right? 10,
- 5:10:2420, and so on.
- 5:10:26Essentially, the difference between
- 5:10:27Parzen window and KNN is that in KNN you
- 5:10:31already know the K and then from the K
- 5:10:33you try to find out the volume and
- 5:10:35therefore then you try to find the the
- 5:10:37probability of the density underlying
- 5:10:39distribution.
- 5:10:40And then there is the the discriminant
- 5:10:43analysis which is off again two kinds
- 5:10:45linear discriminant and multiple
- 5:10:46discriminant analysis where basically
- 5:10:48what you do is you you transform the
- 5:10:52underlying data the features into a
- 5:10:54higher dimension and in such a way that
- 5:10:57in the new feature space after you have
- 5:10:59transformed the data
- 5:11:01you now try to apply a parametric
- 5:11:03approach like for example you will try
- 5:11:05to project the features onto a line or
- 5:11:09if it is on a sub sub space which is
- 5:11:13higher dimension than a line then
- 5:11:15essentially it becomes multiple
- 5:11:16discriminant analysis. So
- 5:11:18basically
- 5:11:19um those are the three kinds of
- 5:11:21non-parametric techniques. So even if
- 5:11:23you were not able to sort of get the
- 5:11:25full hang of
- 5:11:27what these three types are
- 5:11:29what you need to keep in mind is that
- 5:11:31non-parametric technique does not assume
- 5:11:34any kind of distribution or any kind of
- 5:11:36functional relationship of the
- 5:11:38underlying data
- 5:11:39and therefore it gives us a lot of
- 5:11:40flexibility whereas the parametric
- 5:11:43techniques they basically assume some
- 5:11:45kind of functional relationship
- 5:11:48between the data points
- 5:11:50or they assume some kind of
- 5:11:51distribution.
- 5:11:53So where would you use a non-parametric
- 5:11:55versus a parametric technique right?
- 5:11:57So basically you would use
- 5:11:59non-parametric technique you know where
- 5:12:02you do not know about the functional
- 5:12:03relationship that is one
- 5:12:05second is that you know there is
- 5:12:09maybe let's say large amount of data and
- 5:12:12so on. And thirdly
- 5:12:14non-parametric technique like KNN does
- 5:12:16not work in very high dimensional data.
- 5:12:20So you would also use parametric
- 5:12:22techniques in that case.
- 5:12:24Whereas if you have you know smaller
- 5:12:26data sets, you would use, uh, you know,
- 5:12:29typically non-parametric techniques. And
- 5:12:32then also, if you know understand that
- 5:12:34the relationship between the data points
- 5:12:36might be, for example, linear or
- 5:12:39something like that, you would use a
- 5:12:40parametric technique. So, you're
- 5:12:42assuming that there is some kind of
- 5:12:43relationship, uh, like a linear
- 5:12:45relationship or something like that in
- 5:12:47between the data points.
- 5:12:48So, that's where you will use the
- 5:12:50parametric technique. So,
- 5:12:51again, just to summarize, basically
- 5:12:53non-parametric is an approach where you
- 5:12:55do not assume any kind of distribution
- 5:12:57or functional relationship, whereas
- 5:12:59parametric assumes a functional
- 5:13:00relationship or basically a distribution
- 5:13:02between the data points.
- 5:13:05The other feature of KNN is that it is a
- 5:13:07lazy algorithm. So, what lazy algorithm?
- 5:13:10Actually, in most of the supervised, uh,
- 5:13:13learning algorithms,
- 5:13:15uh, you basically train your model on
- 5:13:17the training data set,
- 5:13:19then you have your model,
- 5:13:21and then you apply this particular model
- 5:13:23to the test data set to then classify or
- 5:13:26predict.
- 5:13:27Uh, you know, for example, whether a new
- 5:13:29image is that of a cat or a dog, right?
- 5:13:33So, this can be, you know, some kind of
- 5:13:35algorithm like support vector machine
- 5:13:37or regression or logistic regression or
- 5:13:39whatever, right? So,
- 5:13:40you run this logistic regression or you
- 5:13:42run this regression or support vector
- 5:13:45machine on the training data set,
- 5:13:47and it learns the features or it learns
- 5:13:50the parameters of the model from that
- 5:13:52data set, and then applies this learned
- 5:13:55model on the test data set.
- 5:13:58Whereas, in case of KNN, actually, there
- 5:14:01is no training step at all. That's why
- 5:14:04it's called a lazy algorithm.
- 5:14:06Because what it does is,
- 5:14:07at the time that you are actually now
- 5:14:10want to predict,
- 5:14:12at that point in time, actually, it will
- 5:14:13go and it will do all the calculations
- 5:14:15of the distance of the new data point,
- 5:14:18like, for example, the new image of the
- 5:14:20cat or dog, from all the other data
- 5:14:23points that you have.
- 5:14:25So, it will calculate all the distances
- 5:14:27and then will check for the those data
- 5:14:29points which are or the K data points
- 5:14:31which are nearest to this.
- 5:14:33So, that's why it's called the lazy
- 5:14:35algorithm because nothing happens
- 5:14:38till the point or no calculations happen
- 5:14:40till the point you are actually trying
- 5:14:41to predict something.
- 5:14:43So, there is no training step involved.
- 5:14:45Okay, and then
- 5:14:47it's used for both classification and
- 5:14:49regression as we just mentioned. So, it
- 5:14:51can be used to predict the values as
- 5:14:53well as be able to classify something
- 5:14:56like, you know, okay, whether it is a
- 5:14:57cat or dog or if you're trying to, let's
- 5:15:00say, predict some value
- 5:15:02some forecast or something that for that
- 5:15:04also you can use it. And then it is
- 5:15:06based on feature similarity, which is
- 5:15:07basically what do we mean by feature
- 5:15:10similarity? So, feature can be things
- 5:15:11like if, for example, you know, you are
- 5:15:14looking at classifying cats versus dog,
- 5:15:16right? So, is the eyes like a dog? That
- 5:15:20can be one of the features. Is the what
- 5:15:22how do the ears look? That can be one of
- 5:15:24the features.
- 5:15:25What about the tongue, the face, and so
- 5:15:27on? So, there can be multiple such
- 5:15:28features.
- 5:15:30And how similar it these features are
- 5:15:33between two data points,
- 5:15:35which is used basically by the KNN
- 5:15:38algorithm.
- 5:15:39And then
- 5:15:40as I said, there is no training step
- 5:15:41involved. So, these are the features of
- 5:15:44KNN algorithm.
- 5:15:45And therefore, now let's look at
- 5:15:47actually just some simple examples of
- 5:15:49how it works.
- 5:15:52So, as you can see here in this slide,
- 5:15:54we have
- 5:15:56two classes of data. So, one is the all
- 5:15:58these blue data points.
- 5:16:00And there's another one which is the
- 5:16:02orange data points.
- 5:16:04Now, if you have a new data point, which
- 5:16:06is this
- 5:16:08pink one here, which class should it
- 5:16:10belong to?
- 5:16:12Should it belong to class A or should it
- 5:16:14belong to class B? So, what you would do
- 5:16:16is you would actually start calculating
- 5:16:18the distance of this pink data point
- 5:16:21from every square or blue triangle data
- 5:16:24point. And then
- 5:16:26you will decide you will have to assume
- 5:16:28a particular K. Let's say
- 5:16:30K is three.
- 5:16:32Right? Which is I'm looking at the
- 5:16:34nearest three data points. And in that
- 5:16:38case, basically, as you can see if we
- 5:16:41draw the circle, right?
- 5:16:42Then we see that two of the nearest data
- 5:16:46points within that circle is of the
- 5:16:49square orange kind. So, basically, we
- 5:16:51will predict that this particular new
- 5:16:53data point belongs to class A.
- 5:16:55Whereas, K value was seven, right? As in
- 5:16:59this particular example now.
- 5:17:01You would see that four out of the seven
- 5:17:03is actually of the blue triangle kind.
- 5:17:05And therefore, we will now classify it
- 5:17:07as belonging to class B.
- 5:17:09So, essentially, this prediction
- 5:17:12changes, as you can see here, depending
- 5:17:15on what is the K value.
- 5:17:17So,
- 5:17:18therefore, the question is what should
- 5:17:20be the value of K, right? And typically,
- 5:17:22what happens is you run a trial and
- 5:17:24error, and basically, you will come up
- 5:17:27with okay, what is the best K value. But
- 5:17:29essentially, what one needs to
- 5:17:31understand is that as the value of K
- 5:17:35increases, basically, the the partition
- 5:17:38line starts moving towards becoming more
- 5:17:41and more linear. So, it starts becoming
- 5:17:43less flexible, and it starts assuming
- 5:17:45some kind of a linear dividing line or
- 5:17:48something like that. So,
- 5:17:50what happens is that in that as you
- 5:17:53increase the K,
- 5:17:54your bias
- 5:17:56basically, increases, but your variation
- 5:18:00reduces. So, we know that in
- 5:18:03classification problems or in machine
- 5:18:04learning problems, bias and variance are
- 5:18:07two things that we are trying to manage,
- 5:18:09right? Bias is basically how close you
- 5:18:12are to the actual class or how or to the
- 5:18:15actual
- 5:18:16value. Whereas variation is how much
- 5:18:19variability is there in your prediction.
- 5:18:22So,
- 5:18:22as the K increases, the bias
- 5:18:25sort of increases, but the variance
- 5:18:27reduces. And then it is vice versa. So,
- 5:18:29if your K decreases, let's say
- 5:18:31at K is equal to 1, where you are only
- 5:18:34looking at just one nearest neighbor and
- 5:18:36then
- 5:18:37predicting based on that, actually the
- 5:18:40bias is the least, which means it is the
- 5:18:42most flexible.
- 5:18:43K is equal to 1 is the will give you the
- 5:18:45most flexible sort of demarcating line
- 5:18:49or function.
- 5:18:50Whereas the variability will be the
- 5:18:52maximum.
- 5:18:54So, that that's the sort of the
- 5:18:55trade-off.
- 5:18:57And that's how we actually determine K.
- 5:19:00So, we have to get a K value in such a
- 5:19:02way based on trial and error that
- 5:19:04sort of maximizes our sort of or reduces
- 5:19:07the bias as well as the variance. And
- 5:19:10and that's the kind of optimization we
- 5:19:11are trying to do.
- 5:19:13Okay, so how do we calculate the
- 5:19:14distance itself, right? And distance
- 5:19:17typically can be of many kinds. So, you
- 5:19:20know, the example here is of the
- 5:19:21Euclidean distance, but you can have
- 5:19:23other kinds of distances like Manhattan
- 5:19:25distance or Mahalanobis distance and you
- 5:19:28can look up references for other kinds
- 5:19:31of distances.
- 5:19:32Now, Euclidean distance is calculated
- 5:19:35for the point P1 and P2 as given here.
- 5:19:37Essentially, Euclidean distance is
- 5:19:40nothing but the you know, square root of
- 5:19:42sum of the X coordinates of these two
- 5:19:45points P1 and P2 and then Y coordinate
- 5:19:48square of
- 5:19:49of of these two points P1 and P2.
- 5:19:52And
- 5:19:53this is just an example of one kind of
- 5:19:55distance and
- 5:19:56other kinds of distances like Manhattan
- 5:19:59or Mahalanobis are also there.
- 5:20:01And then this calculating this distance
- 5:20:03becomes quite challenging,
- 5:20:05especially in cases where
- 5:20:07you know, you are trying to for example
- 5:20:09calculate, let's say, how close two
- 5:20:11LinkedIn profiles are, right? Or trying
- 5:20:13to classify uh the category of
- 5:20:16electrocardiogram and so on so forth.
- 5:20:18So, there we have to bring in more
- 5:20:20creativity to just decide what kind of
- 5:20:22distance to use.
- 5:20:24Okay. Now, let's move ahead. We will now
- 5:20:27talk of some use cases where KNN can be
- 5:20:29used and this is an example of how KNN
- 5:20:32can be used for book recommendation. So,
- 5:20:34if you have purchased books on Amazon or
- 5:20:36whatever, right?
- 5:20:38Some of these recommendations are based
- 5:20:39on on KNN algorithm. And then, you know,
- 5:20:43as we said, you know, KNN is like based
- 5:20:45on features, right? So, maybe let's say
- 5:20:48what will be the nearest neighbors of a
- 5:20:50particular book. It can be based on who
- 5:20:51is the author, what is the topic, and so
- 5:20:54on so forth.
- 5:20:55And then there are other use cases like
- 5:20:58I mentioned. So, for classifying
- 5:20:59satellite images, for classifying
- 5:21:01handwritten digits,
- 5:21:03uh on on image analytics or or for
- 5:21:07classifying electrocardiograms,
- 5:21:09um etc. Uh you know, typically KNN can
- 5:21:12be used.
- 5:21:13Okay. So, now actually we will get into
- 5:21:15some hands-on.
- 5:21:18Okay. So, to start the hands-on session,
- 5:21:21I'll go to this Jupyter notebook that I
- 5:21:25already have installed on my system.
- 5:21:28And I have a certain
- 5:21:31code written, which uh we will take two
- 5:21:33examples.
- 5:21:34Both the examples are based on data sets
- 5:21:37which are available in the open source.
- 5:21:39So, you can easily get access to that
- 5:21:42data.
- 5:21:43So, what we do is we start by importing
- 5:21:47the necessary libraries.
- 5:21:49So, we import pandas, seaborn, numpy,
- 5:21:54and matplotlib. Basically, pandas and
- 5:21:56numpy are there for doing the data
- 5:21:59manipulation and also for storing data
- 5:22:02as matrices or as arrays and and be able
- 5:22:05to perform some mathematical procedures
- 5:22:08on them.
- 5:22:09And then seaborn is basically used for
- 5:22:11plotting and matplotlib for plotting as
- 5:22:13well.
- 5:22:14And this line here get IPython just
- 5:22:17helps us to run the images that we'll be
- 5:22:20creating in line with Jupiter notebook
- 5:22:22instead of opening up a new window.
- 5:22:25So let's run this.
- 5:22:27And what it will do is it will import
- 5:22:28all these packages for us which we are
- 5:22:30going to use.
- 5:22:32And then we will first import the breast
- 5:22:35cancer data that is available in your
- 5:22:38scikit-learn datasets.
- 5:22:40So we import that. And then let's just
- 5:22:44initialize that data into a variable
- 5:22:47here called cancer. So cancer here
- 5:22:49represents all the load the breast
- 5:22:51cancer data.
- 5:22:53And now we will let's actually look at
- 5:22:55what this data is.
- 5:22:58It's a bunch of attributes in this
- 5:23:01dictionary here. So you have data, the
- 5:23:03target which is basically nothing but
- 5:23:05whether it is a cancer or not. So
- 5:23:08whether it is malignant or benign.
- 5:23:09Malignant means it's a bad cancer and
- 5:23:12benign means well it's just a tumor,
- 5:23:13it's not cancerous.
- 5:23:15Target name description, feature names
- 5:23:18which is basically
- 5:23:19the features that will tell us whether a
- 5:23:21particular
- 5:23:23case belongs to cancerous or or
- 5:23:25non-cancer or malignant or benign.
- 5:23:28And then
- 5:23:29actually let's just print the
- 5:23:30description of this particular data
- 5:23:32here.
- 5:23:34So as you can see we can use this
- 5:23:36command to print the description here.
- 5:23:39And then we see that there are 569 data
- 5:23:42points with about 30 attributes.
- 5:23:45And these attributes are radius,
- 5:23:46texture, perimeter, etc.
- 5:23:48And
- 5:23:50the the max and min values of those are
- 5:23:53given here.
- 5:23:55And then now let's look at some of the
- 5:23:57feature names.
- 5:23:59So, these are the feature names, radius,
- 5:24:03texture, and so on.
- 5:24:06And now, let's actually set up a data
- 5:24:08frame of this particular data here
- 5:24:13using pandas, this function here. So,
- 5:24:15there are 569
- 5:24:18data points.
- 5:24:19And all these are basically your
- 5:24:21features, as we talked about.
- 5:24:25And let's look at the target variable,
- 5:24:28which is nothing but whether it is
- 5:24:29telling us whether a particular data of
- 5:24:31point belongs to malignant or benign.
- 5:24:33So, zero is cancerous and one is
- 5:24:35non-cancerous.
- 5:24:37And then,
- 5:24:38we convert the target into a data frame
- 5:24:41as well.
- 5:24:42And then, let's look at the couple of
- 5:24:45examples of how the data points look
- 5:24:48like. This is, you know, one row of the
- 5:24:51data points which with all the several
- 5:24:53feature values that we have.
- 5:24:56So, basically, we use this um package
- 5:25:00called standard scalar from scikit-learn
- 5:25:02for pre-processing and for standardizing
- 5:25:04the variables.
- 5:25:06And we initialize this standard scalar
- 5:25:09into a variable called scalar.
- 5:25:11So, standardizing is nothing but, you
- 5:25:13know, basically, bringing all the
- 5:25:15samples to essentially the same range,
- 5:25:18right? Because
- 5:25:20uh what might happen is some of the data
- 5:25:22point, like, for example, temperature
- 5:25:23might be from zero to 100 and some price
- 5:25:26might be from, let's say, 1,000 to
- 5:25:30100,000 or whatever, right? So, the
- 5:25:32absolute values can can lead to some
- 5:25:34issues with respect to the prediction.
- 5:25:36Therefore, we have to standardize it or
- 5:25:37bring it between, let's say, minus one
- 5:25:40and one. So, and then, a mean of zero,
- 5:25:42right? So, we have to bring everything
- 5:25:44to the same scale to be able to compare
- 5:25:46the samples.
- 5:25:47So, we first of all, we fit the this
- 5:25:51standardization um or normalization on
- 5:25:54the data set we have.
- 5:25:56And that is we calculate the the
- 5:25:59variance and and the means and then we
- 5:26:02actually apply it on the data set to
- 5:26:04transform it to the actual values.
- 5:26:07And then
- 5:26:09if we look at the scale values now, so
- 5:26:11let's look at the scale values.
- 5:26:14And this will give an example of the top
- 5:26:17five rows here.
- 5:26:18So we can see now the values are between
- 5:26:20minus one and one.
- 5:26:22Or rather it is standardized.
- 5:26:24Essentially with a normal distribution.
- 5:26:27And then we divide this data into
- 5:26:30test and train. So basically we will
- 5:26:34train the model and then we will test it
- 5:26:36on a separate data set. If you use the
- 5:26:39same random state, you should be able to
- 5:26:41get the same result. Otherwise you may
- 5:26:43get a different result here. And
- 5:26:44essentially we are keeping the testing
- 5:26:47size to 30 which means that we are
- 5:26:48dividing the entire data set into two
- 5:26:50parts. The train part which is having
- 5:26:5270% of the data and the test part which
- 5:26:55is having 30% of the data. And again we
- 5:26:56are using this package called train test
- 5:26:59split from the scikit-learn package.
- 5:27:02So we get the X and the Ys which are
- 5:27:05basically nothing but your train and the
- 5:27:09X's are your predictors and Y is your
- 5:27:11predicted variable whether it is
- 5:27:13cancerous or not.
- 5:27:15And then now let's import the K nearest
- 5:27:17neighbors classifier. This is the actual
- 5:27:19algorithm
- 5:27:20which we are importing from scikit-learn
- 5:27:23package.
- 5:27:24And now we initialize this particular
- 5:27:28algorithm.
- 5:27:30And then we fit it on the
- 5:27:33data.
- 5:27:37And some of the parameters as you can
- 5:27:40see
- 5:27:41is basically what is the leaf size and
- 5:27:42so on so forth.
- 5:27:44The nearest N neighbors we are taking.
- 5:27:46So we are taking K is equal to one here
- 5:27:48basically as of now. We will see the
- 5:27:50results based on that and then we will
- 5:27:52change it and see how the results vary.
- 5:27:56And we now
- 5:27:59run it on the We now try to predict it.
- 5:28:03And then we will now try to evaluate
- 5:28:06what the results look like.
- 5:28:08So, we have imported the classification
- 5:28:10report and confusion matrix,
- 5:28:12which is basically trying to see whether
- 5:28:15we were able to correctly classify the
- 5:28:18cancerous as cancerous and non-cancerous
- 5:28:20as non-cancerous as or not.
- 5:28:22So, we can see that this is the actual
- 5:28:24and this is the predicted. So, basically
- 5:28:26some data points here five and four are
- 5:28:29classified wrongly, otherwise all the
- 5:28:31others are classified well. So, if we
- 5:28:33look at the accuracy calculated accuracy
- 5:28:37actually, so
- 5:28:39we see that the precision, which is true
- 5:28:42alarm, right? Which is basically from
- 5:28:44the cancerous
- 5:28:45samples, how many were you able to
- 5:28:47actually predict as cancerous?
- 5:28:49If we see the accuracy is quite high,
- 5:28:50almost 94 95%.
- 5:28:54And then the recall, which is from all
- 5:28:57of the cancerous samples, how many were
- 5:28:59you able to actually predict accurately
- 5:29:01is about again 94 95%. And F1 score is
- 5:29:04nothing but a combination of both
- 5:29:06precision as well as recall. And that's
- 5:29:08quite good as well. So, with K is equal
- 5:29:10to one, you're able to get some already
- 5:29:12some good results. Now, let's try to see
- 5:29:15how to choose the K value, right? So,
- 5:29:17this is basically nothing but a
- 5:29:20a bunch of code that actually runs the K
- 5:29:23value from one to 40 and then tries to
- 5:29:26check the accuracy.
- 5:29:28And this is just like doing a trial and
- 5:29:30error to see where we get the best
- 5:29:32results so that we can then use the best
- 5:29:34K value. So, if you can see this
- 5:29:36particular plot here after we plot the
- 5:29:38result from the running the trial and
- 5:29:41error from one to 40, we see that the
- 5:29:43error actually starts decreasing and
- 5:29:45somewhere around this K is going to 21,
- 5:29:48we get the minimum value of error. So,
- 5:29:50for us, the best K value is
- 5:29:5221. So, now if we compare the results
- 5:29:55between K is equal to 1 and K is equal
- 5:29:57to 21, we we should be able to see the
- 5:29:59prediction results. So, as you can see,
- 5:30:00this was the result with K is equal to
- 5:30:021, which is
- 5:30:03we get about 94 95% accuracy.
- 5:30:06And then with K is equal to 21,
- 5:30:09we will see whether the accuracy
- 5:30:10improves, right? So, we see that yes,
- 5:30:12the accuracy has now gone up to almost
- 5:30:1599%, which is we earlier had nine
- 5:30:19misclassified data points
- 5:30:21out of all of the points. And then here
- 5:30:24we have just two data points which are
- 5:30:26misclassified from the test data set.
- 5:30:28So, now this was one example of applying
- 5:30:31KNN on the cancer data set, which is
- 5:30:33available freely.
- 5:30:35And now let's look at another example,
- 5:30:37which is the Iris data set.
- 5:30:39And
- 5:30:40again, available freely as well.
- 5:30:43So, Iris is a type of flower.
- 5:30:45And we will see what flower is it. So,
- 5:30:48just give me a moment here. So, we again
- 5:30:49start by importing
- 5:30:51the necessary libraries. And then
- 5:30:54we'll look at what this Iris data set
- 5:30:57is. So, the Iris data set comprises of
- 5:31:0050 samples of three species of Iris
- 5:31:03flower, which is Iris
- 5:31:04setosa, Iris virginica, and Iris
- 5:31:06versicolor. These are
- 5:31:08the three types of Iris flowers. And if
- 5:31:10you run this, basically we will find
- 5:31:12that
- 5:31:16Okay.
- 5:31:17Sorry, we did not copy the entire the
- 5:31:20code here. So,
- 5:31:21it was giving an issue.
- 5:31:23Let's just run it again.
- 5:31:28Okay. So, we see that it is this
- 5:31:29particular flower, which is Iris setosa.
- 5:31:32So, the Iris setosa, you can see.
- 5:31:35And
- 5:31:36now let's look at the other two kinds of
- 5:31:38flowers here. So, which is
- 5:31:40Iris versicolor.
- 5:31:44So, this is Iris versicolor. And then
- 5:31:47you have the Iris virginica.
- 5:31:55So, we see that this one is Iris
- 5:31:57virginicas. So, essentially we now will
- 5:32:01import the sort of data set. We have
- 5:32:04already done that. And we will now use
- 5:32:06the seaborn package to actually plot
- 5:32:09some of this data and see
- 5:32:12how it looks like.
- 5:32:14Which is basically do some kind of
- 5:32:15exploratory data analysis. So, if we
- 5:32:18look at the data itself, so this is the
- 5:32:20top five rows from the data, right? So,
- 5:32:22essentially the data consists of
- 5:32:25basically what is the sepal length,
- 5:32:27sepal width,
- 5:32:28petal length, and petal width.
- 5:32:30And this is nothing but basically your
- 5:32:33petal is your this colored part of the
- 5:32:35flower and the sepal is basically your
- 5:32:37green part here, right? So, it is
- 5:32:39talking about what what is the sepal
- 5:32:40length, sepal width, and petal length,
- 5:32:42and petal width of each of the species,
- 5:32:44whether it is setosa, virginica, or
- 5:32:47versicolor. And we have we will see
- 5:32:49whether we can use KNN to actually
- 5:32:52classify these
- 5:32:54these flowers into the data points into
- 5:32:56these categories of flowers.
- 5:32:58So, let's do some quick exploratory data
- 5:33:01analysis
- 5:33:02on this. So, we are running a pair plot
- 5:33:05on the data set. And uh
- 5:33:08in the meantime, I'll just copy another
- 5:33:11part of the code here.
- 5:33:13Okay, so now the plot has come up. So,
- 5:33:16as we can see that the green is
- 5:33:18basically your setosa flower. And
- 5:33:22we can see that this pair plot actually
- 5:33:24just plots
- 5:33:25the sepal length, sepal width,
- 5:33:28and petal length, petal width of each of
- 5:33:30the samples of setosa, verse- color and
- 5:33:33virginica. And we see that
- 5:33:34the green dots which are the setosa
- 5:33:36flower is actually quite separable from
- 5:33:38the others. It's when you plot
- 5:33:40let let's say for example sepal length
- 5:33:42and
- 5:33:43petal length, right? We see that this is
- 5:33:45quite separate from the other data
- 5:33:47points. So, let's see whether you know,
- 5:33:49we can actually
- 5:33:50be able to classify it using KNN
- 5:33:53or not. And here we are running a kernel
- 5:33:56density estimation function on the
- 5:33:58setosa flower to check
- 5:34:01what kind of distribution it has. So,
- 5:34:03this is the kernel density estimation
- 5:34:06plot using the SNS package. So, only for
- 5:34:09the setosa flower. So, if we plot the
- 5:34:12sepal length and sepal width, we get
- 5:34:13something distribution like this. So,
- 5:34:15essentially we see that the maximum
- 5:34:17centered around here and then there is a
- 5:34:19distribution as you can see here. So,
- 5:34:21there's some kind of a linear
- 5:34:22relationship here.
- 5:34:24Okay, so now we will again do the same
- 5:34:27standardization of the variables
- 5:34:30that we had done in the cancer data set
- 5:34:32case. So, we are importing the standard
- 5:34:34scalar function
- 5:34:36from the scikit-learn preprocessing. So,
- 5:34:39we will initialize that.
- 5:34:41So, we are again basically doing the
- 5:34:42standardization or normalization of the
- 5:34:45data.
- 5:34:46And we will do the standardization on
- 5:34:48everything except the species which is a
- 5:34:50categorical value, right? So, it's
- 5:34:52categorical whether it is which kind of
- 5:34:54flower it is. So, we have removed that
- 5:34:56and then we have
- 5:34:58done the standardization or
- 5:35:00normalization on rest of the data.
- 5:35:02So, we now convert this into a data
- 5:35:05frame, pandas data frame. And if we look
- 5:35:08at the top five rows, now it's all
- 5:35:09converted or transformed. So, the values
- 5:35:12are now normally distributed basically.
- 5:35:15Okay, so now we
- 5:35:18divide the data again into train and
- 5:35:20test.
- 5:35:22And we again have training of about 70%
- 5:35:27and test data set of about 30%. Um so we
- 5:35:31are dividing that entire data set into
- 5:35:32these two buckets.
- 5:35:34And we will now use KNN
- 5:35:38to see if we can use KNN to classify
- 5:35:40them.
- 5:35:42Again, the same and K is equal to 1.
- 5:35:46And we will check the results and then
- 5:35:47we will do a trial and error
- 5:35:50to check what is the best value of K.
- 5:35:52So here K is equal to 1.
- 5:35:57And now we are going to predict on the
- 5:36:00test data set
- 5:36:02and look at the results.
- 5:36:05So we are now importing the
- 5:36:06classification report and the confusion
- 5:36:08matrix.
- 5:36:12So if you look at the confusion matrix,
- 5:36:18we see that
- 5:36:19kind of already we are getting quite
- 5:36:21good
- 5:36:22uh prediction. So just two misclassified
- 5:36:24points.
- 5:36:27And uh if we look at the accuracy,
- 5:36:32we see that the accuracy is quite high,
- 5:36:34around 96%
- 5:36:37already.
- 5:36:38Now we choose we have to see what is the
- 5:36:41best value of K.
- 5:36:43So
- 5:36:45essentially we will
- 5:36:49we will again run K is equal to 1 to 40
- 5:36:52and check which is the best value.
- 5:36:56So let's plot the errors when we vary
- 5:36:59the K from 1 to 40. And we see that the
- 5:37:02actually the error decreases and then
- 5:37:04increases. So basically the error is
- 5:37:07minimum with K is equal to let's say
- 5:37:08three or even five or maybe 11. So let's
- 5:37:12choose one of these values. So let's say
- 5:37:14K is equal to three.
- 5:37:16And let's see how the results look like.
- 5:37:19Does it improve the
- 5:37:21accuracy or not? So we now see that
- 5:37:24even the two data points which are
- 5:37:25misclassified earlier is now classified
- 5:37:27properly. So, the accuracy improves to
- 5:37:29100%.
- 5:37:31So, that's the
- 5:37:32example of how you can choose K.
- 5:37:36>> [music]
- 5:37:40>> What is Naive Bayes?
- 5:37:42Let us understand Naive Bayes with an
- 5:37:44example. Here, I just cannot seem to
- 5:37:47figure out which are the best days to
- 5:37:49play football with my friend. Can you
- 5:37:52please help us out?
- 5:37:54All possible conditions are given to us.
- 5:37:56There is
- 5:37:58summer, monsoon, and winter, which is
- 5:38:00nothing but the outlook.
- 5:38:03Am I correct in saying that?
- 5:38:04Summer, monsoon, and winter is nothing
- 5:38:06but the outlook. Then we have sunny or
- 5:38:09not sunny. So, that basically is the
- 5:38:12humidity, right? And then we have windy
- 5:38:15or no windy. That speaks about
- 5:38:18the winds. How are the winds?
- 5:38:20Right?
- 5:38:21So, if you look at these combinations,
- 5:38:24okay, we will look at this using Naive
- 5:38:25Bayes on how do we decide whether we can
- 5:38:28play or not.
- 5:38:29So, if I have noted down all the days it
- 5:38:31was good, bad to play football, and the
- 5:38:33combination of weather matrices on that
- 5:38:35day, that will be perfect, right? That
- 5:38:37is perfect, and we will be able to do
- 5:38:39Naive Bayes classifiers using that. Now,
- 5:38:42Naive Bayes classifier comes from the
- 5:38:44Naive Bayes theorem, and Naive Bayes
- 5:38:46theorem is purely and purely based on
- 5:38:50the assumption of independence.
- 5:38:53So, what does it mean? When I say
- 5:38:55independence, what it means is that
- 5:38:59this variable has no relationship, no
- 5:39:03association with this variable. Now,
- 5:39:06when I talk about this in a linear
- 5:39:09context, obviously in summer, you will
- 5:39:12see that we have more sunny days.
- 5:39:16Yes? So, if you look at it from the
- 5:39:18correlation area, from the linear
- 5:39:21algebra concepts, linearly these two are
- 5:39:25correlated to each other. Am I correct
- 5:39:28in saying that?
- 5:39:29Obviously, in monsoon we have less sunny
- 5:39:32days. In winter we further have less
- 5:39:34sunny days.
- 5:39:36Yeah?
- 5:39:36So,
- 5:39:37although there is a relation,
- 5:39:39Naive Bayes theorem
- 5:39:42says that all these variables are
- 5:39:45independent of each other.
- 5:39:49What does the that mean? If this is
- 5:39:51causing any kind of an effect, if this
- 5:39:54is causing any kind of an impact,
- 5:39:57this should not matter.
- 5:40:00Okay? There's going to be no
- 5:40:01relationship. Summer, monsoon, winter,
- 5:40:04it has its own weightage. Okay? And it
- 5:40:07has nothing to do with the other
- 5:40:09conditions. Every condition is equally
- 5:40:12significant.
- 5:40:14All right? So, what happens in Naive
- 5:40:15Bayes is we estimate the posterior
- 5:40:18probability of every event happening.
- 5:40:21Here, we calculate the posterior
- 5:40:24probability of an event happening.
- 5:40:27Okay? So, here if you see, if you look
- 5:40:29at the sunny conditions, what we have
- 5:40:31done, sunny conditions we have this
- 5:40:33distribution. That there is no play
- 5:40:35happening in summer, there is play
- 5:40:36happening in monsoon, there is play
- 5:40:38happening in winter.
- 5:40:39Right? So, base of the season, we are
- 5:40:42figuring that out. Similarly for windy
- 5:40:44conditions, we are doing that.
- 5:40:46Right? Again, then we do it for a
- 5:40:49combination.
- 5:40:50Right? Whether when windy conditions are
- 5:40:53yes and no, what happens to play?
- 5:40:56Okay? So, here
- 5:40:57what at the end of the day, what gets
- 5:41:00selected is the one which has a
- 5:41:02posterior probability of greater than
- 5:41:04five. Now, when I talk about posterior
- 5:41:07probability,
- 5:41:08what do I mean by posterior probability?
- 5:41:10Let me have a blank slate. Here you go.
- 5:41:13Actually, it's given. So, I need not
- 5:41:15show you that. What is the simplistic
- 5:41:17probabilistic classifier here? What is
- 5:41:19the probability of an event A happening
- 5:41:22given
- 5:41:23B.
- 5:41:24So, if you look at our problem context,
- 5:41:27what is it that we are trying to figure
- 5:41:29out?
- 5:41:29We are trying to figure out what is the
- 5:41:32probability of
- 5:41:36play happening
- 5:41:38given
- 5:41:45the outlook is sunny,
- 5:41:50{comma}
- 5:41:53the
- 5:41:54uh
- 5:41:57winds
- 5:42:03are normal,
- 5:42:07and there is no rain.
- 5:42:13On any given day,
- 5:42:15on any given day when there is no rain,
- 5:42:19there is no wind,
- 5:42:21and the outlook is sunny,
- 5:42:23whether play will happen or not. So,
- 5:42:26what we end up doing is we calculate the
- 5:42:28posterior probability of play happening
- 5:42:30given these conditions. We also
- 5:42:32calculate the posterior probability of
- 5:42:35play not happening
- 5:42:38given these conditions, and then we
- 5:42:40normalize these probabilities.
- 5:42:43Mathematically, we do all of these
- 5:42:45calculations to figure out naive Bayes.
- 5:42:47Okay? Now, here
- 5:42:49see,
- 5:42:50here we are talking about one event.
- 5:42:53Here, we have three events. We have
- 5:42:55outlook,
- 5:42:57we have winds,
- 5:42:59and we have rains.
- 5:43:00So, what this becomes is
- 5:43:04probability of three independent events.
- 5:43:07So, what we will do, we figure out what
- 5:43:10is the probability of play happening
- 5:43:16given
- 5:43:18outlook is sunny.
- 5:43:22We also figure out
- 5:43:26what is the probability of
- 5:43:29play happening.
- 5:43:34Multiply this with the probability of
- 5:43:37wind as no.
- 5:43:41Given no, then probability of play
- 5:43:44happening
- 5:43:49given
- 5:43:51rain is no.
- 5:43:53We figure out all the in three
- 5:43:56independent probabilities multiplied.
- 5:43:59All right? This is what we do
- 5:44:01mathematically.
- 5:44:03This is what is done mathematically in
- 5:44:06Naive Bayes theorem.
- 5:44:08Ultimately, the posterior probability
- 5:44:12the posterior probability probability of
- 5:44:14an event A happening given
- 5:44:17B conditions is calculated by first
- 5:44:21calculating the class probability.
- 5:44:24What is the class probability here?
- 5:44:25Probability of it raining given play was
- 5:44:29happening in that day.
- 5:44:31This is multiplied by the total
- 5:44:33probability of play happening and
- 5:44:35divided by the total probability of
- 5:44:37sunny conditions.
- 5:44:41Here, as you see, this is the posterior
- 5:44:43probability.
- 5:44:44This is the class probability.
- 5:44:47Here is the predictors probability. This
- 5:44:49is the outlook. X is the outlook. C is
- 5:44:53what we are trying to predict.
- 5:44:55Right?
- 5:44:56And here is the likelihood.
- 5:44:59So, how is this formulated into our
- 5:45:01table? Now, if you correlate this to our
- 5:45:03graph,
- 5:45:04what will be the likelihood?
- 5:45:07What will be the likelihood of sunny
- 5:45:10conditions given play happens?
- 5:45:13Sunny conditions play happens.
- 5:45:152 / 3
- 5:45:18out of
- 5:45:19Sorry, how many sunny conditions do we
- 5:45:21have? Six
- 5:45:22conditions.
- 5:45:24In six con-
- 5:45:25ditions, how many days does play happen?
- 5:45:27Two days.
- 5:45:292 / 6 1 / 3. What is the class
- 5:45:32probability? So, of all the events that
- 5:45:34are given to us, how many days does play
- 5:45:36happen?
- 5:45:37Okay? This is how this is calculated.
- 5:45:40So, here you see what we have done is
- 5:45:43Let me go back to the previous slide.
- 5:45:45Here you go.
- 5:45:46Okay? Here, what we have done is we have
- 5:45:48calculated this table. Now, it speaks
- 5:45:51about a data set. Where can you get this
- 5:45:54data set?
- 5:45:55You can look for golf play days data set
- 5:45:59online. You can look for golf play days
- 5:46:02data set.
- 5:46:03Okay?
- 5:46:04In this data set, you will find all the
- 5:46:07data
- 5:46:08which is required for this particular
- 5:46:11example to be done. So, what I will be
- 5:46:13doing is I will be doing this example,
- 5:46:15this Naive Bayes classification, with
- 5:46:18you in Python. All right? So, whatever
- 5:46:20calculations are being done here,
- 5:46:23okay? I will do the same activity in
- 5:46:26Python with pen and paper. And instead
- 5:46:28of doing this
- 5:46:32in a numeric way where I'm doing lot of
- 5:46:34probabilistic calculations,
- 5:46:37I will achieve this simply in
- 5:46:40very limited lines of code.
- 5:46:43Very limited lines of code with Python.
- 5:46:47Once I'm I have done that, I will come
- 5:46:49and explain this prob- probability table
- 5:46:52to all of us.
- 5:46:53I'm going to use some basic libraries.
- 5:46:56All right. So, here, what I will be
- 5:46:58doing is using some very basic libraries
- 5:47:01for this activity. All right?
- 5:47:57Done. Now, let me quickly go and uh
- 5:48:01read the data set. So, for that what I
- 5:48:04will do is quickly
- 5:48:08change my working directory.
- 5:48:43And now, let me quickly go and read my
- 5:48:44data. So, my data frame is pd.
- 5:49:09>> Here you go. This is my data set.
- 5:49:12Right? So, if you look at this data set,
- 5:49:14in this data set, you have 13 14 days.
- 5:49:18In these 14 days, you have the outlook,
- 5:49:21overcast,
- 5:49:22rainy, and sunny. You have temperature,
- 5:49:26temperature is hot, cool, and mild. You
- 5:49:29have humidity, you have wind, and you
- 5:49:32have play. Right? So, I will not be
- 5:49:35using uh
- 5:49:36Okay, let us use all four. In this
- 5:49:39example, they're using only three
- 5:49:40variables, but in our
- 5:49:43hands-on, okay? In this hands-on, what I
- 5:49:46will be doing is I will be uh
- 5:49:49using all four variables. Let us do
- 5:49:51that.
- 5:49:52Okay?
- 5:49:53So, before I do that, let me convert
- 5:49:56everything into a category.
- 5:49:58If you look at your data frame right
- 5:49:59now, it's not everything is not into a
- 5:50:02categorical variable.
- 5:50:04Here you go, see.
- 5:50:05Okay? So, let me quickly go and convert
- 5:50:07everything into a category.
- 5:50:26And once I have done this, uh
- 5:50:29let me create a new data frame in which
- 5:50:31I have everything as a category code.
- 5:50:35So, that I have numbers. I'll show you
- 5:50:37what What do I mean by this?
- 5:50:52>> So, let me execute this. Here you go.
- 5:50:55See, now I have two data frames. In the
- 5:50:57first data frame I have all these
- 5:50:59values. These are now categorical
- 5:51:01variables, but in the second data frame
- 5:51:03I have all ones and zeros. So, wherever
- 5:51:06you see there is sunny conditions, now I
- 5:51:08have a code two.
- 5:51:09Rainy conditions, code one.
- 5:51:11Similarly, when play happens I have a
- 5:51:15one. When play does not happen I have a
- 5:51:17zero.
- 5:51:19This is what I have done.
- 5:51:21This data frame is available online.
- 5:51:24Okay? You can get this data frame
- 5:51:27online.
- 5:51:29All right? Now, my data frame is ready.
- 5:51:32So, now what I'm going to do is I'm
- 5:51:34going to divide my data frame
- 5:51:37into training and testing. I have 14
- 5:51:39records.
- 5:51:41So, let's take 10 records for training.
- 5:51:44I will give 10 records as an input.
- 5:51:47And I will give four records, last four
- 5:51:49records
- 5:51:57as my test data frame. All right?
- 5:52:00Uh
- 5:52:00so, now I will need to create my X and
- 5:52:03my Y.
- 5:52:04So, how I will do that is
- 5:52:06I'll say Y {underscore} train
- 5:52:09is equal to from train
- 5:52:12I don't want the play variable.
- 5:52:14That play variable should be my Y. As
- 5:52:16simple as that.
- 5:52:18And I will say X {underscore} train is
- 5:52:21equal to train.
- 5:52:23Okay? And I will do the same thing for
- 5:52:25my test data frame also.
- 5:52:29Now, those who are new to Python will
- 5:52:32find this a little bit strange. Please
- 5:52:36bear with me.
- 5:52:37But these are the only calculations,
- 5:52:39only steps which need to be performed
- 5:52:41every time.
- 5:52:42You are trying to achieve maybe bias
- 5:52:45algorithm or any kind of an algorithm.
- 5:52:49All right. So, here now you see this is
- 5:52:52my training data frame in which play
- 5:52:54variable is not there.
- 5:52:56Play variable is not there. This is my Y
- 5:53:00in which only play variable is there.
- 5:53:02This is my training data set. So, both
- 5:53:04of them have 10 records with the
- 5:53:06matching index.
- 5:53:08Similarly, test data frame four records
- 5:53:12four records with the matching index.
- 5:53:15Right? So, that we know which data frame
- 5:53:18is where.
- 5:53:20Now, multinomial naive bias. Very
- 5:53:23simple, three lines of code and my model
- 5:53:26will be done.
- 5:53:28Okay? First, I initialize my model.
- 5:53:32Here you go. I have initialized my
- 5:53:33model.
- 5:53:34In this model, I fit my data.
- 5:53:39In this model, I will fit my data. So,
- 5:53:42to do that, what I say is fit
- 5:53:45X underscore
- 5:53:50train comma
- 5:53:53comma Y underscore train.
- 5:53:56Done.
- 5:53:57Your model object is now ready.
- 5:54:00And now you can simply get the
- 5:54:03classification outcomes. So, we have in
- 5:54:06our
- 5:54:07test data frame, if you look at our test
- 5:54:09data frame, this is our test data frame.
- 5:54:11We have three four conditions. All four
- 5:54:14are sunny,
- 5:54:16high temperature, low humidity, and
- 5:54:19windy. Right? And if you look at their
- 5:54:22outcomes, these are their outcomes.
- 5:54:24On the first two days, play is not
- 5:54:26happening. On the next two days, play is
- 5:54:28happening. Let us look at what is the
- 5:54:30prediction of our model for this. So, to
- 5:54:33do do that, what I simply do is
- 5:54:37X out is equal to
- 5:54:40model.predict
- 5:54:46To this I give my X {underscore} test.
- 5:54:50Here you go. Okay? Now you have your Y Y
- 5:54:54out variable, so this is the prediction
- 5:54:56for all the
- 5:54:59four inputs that you give, and this is
- 5:55:00the prediction.
- 5:55:02First day
- 5:55:03first day we say play does not happen.
- 5:55:06Let us match it.
- 5:55:08Let us try to match it with our
- 5:55:10here.
- 5:55:11See?
- 5:55:12Out of four records, three records we
- 5:55:14are predicting correctly.
- 5:55:16Three records we are predicting
- 5:55:18correctly. If you want to check the
- 5:55:20accuracy, what is the accuracy of your
- 5:55:23model? What you can simply do is print
- 5:55:25Let us print the accuracy on both
- 5:55:27training and testing.
- 5:55:34Training accuracy. How do I get the
- 5:55:36training accuracy? Very simple. model.
- 5:55:39score
- 5:55:43And here I give my X {underscore} train
- 5:55:47{comma} Y {underscore} train.
- 5:55:50And then we do the
- 5:55:52testing accuracy also.
- 5:56:06Here you go. So here you can see for our
- 5:56:10model training we have 80% accuracy, and
- 5:56:13for testing we have 75% accuracy. Okay?
- 5:56:18So this is the advantage of doing this
- 5:56:20activity in Python. But what is
- 5:56:22happening in the back end?
- 5:56:23Now let us go and also understand that
- 5:56:25in terms of naive Bayes classifier.
- 5:56:28We have successfully
- 5:56:30we have successfully implemented the
- 5:56:33Naive Bayes classifier in Python
- 5:56:35programming language. But, here let us
- 5:56:38try to understand Bayes theorem, what is
- 5:56:41happening. So, from this data set all
- 5:56:43the tabulated data frequency tables are
- 5:56:45calculated.
- 5:56:47Once the frequency tables are
- 5:56:48calculated, they are substituted in our
- 5:56:51formula to calculate the probabilistic
- 5:56:53scores. So, what is the probability of
- 5:56:56summer given it is playing conditions?
- 5:57:00Total how many playing conditions are
- 5:57:02there? Total there are nine playing
- 5:57:03conditions. Nine days play happened.
- 5:57:06That becomes our denominator. Out of
- 5:57:08those days, how many days was summer is
- 5:57:10our numerator. That is how for this we
- 5:57:13get a probability of 0.33.
- 5:57:16Then we calculate the class probability
- 5:57:18where we look at how many days was it
- 5:57:20summer? Out of total 14 days, five days
- 5:57:23was summer, so that's the probability
- 5:57:25and the class probability is 0.64. Put
- 5:57:28everything into our equation.
- 5:57:30Put everything into our equation and
- 5:57:32this is what we get.
- 5:57:34Okay? So, we do this for each and every
- 5:57:37condition.
- 5:57:38So, here we calculate it for winter.
- 5:57:42All right? Once we have done it for all
- 5:57:44three days,
- 5:57:47winter, sunny and windy days,
- 5:57:49we substitute those here
- 5:57:51and that gives us the probability which
- 5:57:53is
- 5:57:54more than 0.5. Thus, now we can say that
- 5:57:58if
- 5:57:59it is winter, sunny and conditions are
- 5:58:02sunny and there are winds. Conditions
- 5:58:04are not sunny and there are winds. Play
- 5:58:07can happen.
- 5:58:09Look at another example. If a single
- 5:58:11card is drawn from a standard deck of
- 5:58:14playing cards, the probability that card
- 5:58:16is a king is 4/52
- 5:58:18since there are four kings in a standard
- 5:58:21deck.
- 5:58:22King is the event. This card is a king.
- 5:58:25This is the event. The prior probability
- 5:58:28of this is 1 by 13. If evidence is
- 5:58:31provided, for instance, someone looks at
- 5:58:33the card that the single card is a face
- 5:58:35card, then the posterior probability can
- 5:58:38be calculated using Bayes' theorem.
- 5:58:47Okay? Since every king is also a face
- 5:58:49card, the probability of face happening
- 5:58:52given you getting a face card given it's
- 5:58:54a king is one. Since there are three
- 5:58:57face cards in each suit,
- 5:58:59all right? It's actually four. Ace is
- 5:59:01also a face card. So, it's jack, king,
- 5:59:04queen, and uh ace. The probability of
- 5:59:06the face card is 4 by 13.
- 5:59:08Okay? So, if you combine these three
- 5:59:10likelihoods, what you get is 13 by 4.
- 5:59:13So, using Bayes' theorem, this is the
- 5:59:15probability that you get.
- 5:59:18>> [music]
- 5:59:23>> What is support vector machine?
- 5:59:25Support vector machine comes under
- 5:59:27supervised machine learning.
- 5:59:30And we use it specifically for
- 5:59:32performing the task of classification.
- 5:59:35So, support vector machine is a
- 5:59:36discriminative classifier
- 5:59:38that is formally designed by a separate
- 5:59:41hyperplane.
- 5:59:43Okay? It is a representation of examples
- 5:59:45as points in a space that are mapped so
- 5:59:48that the points of different categories
- 5:59:50are separated by a gap as wide as
- 5:59:53possible.
- 5:59:55So, in this case of the support vector
- 5:59:56machine,
- 5:59:57let's say I have some data points. So,
- 6:00:00there are some data points of X, and
- 6:00:02there are data points of circle.
- 6:00:05Now, this support vector machine
- 6:00:07is a type of machine learning algorithm
- 6:00:10where if I have the collection of
- 6:00:12points, so here in this data points, I
- 6:00:14have two classes. One is X, and the
- 6:00:17another one is circle.
- 6:00:19Now, given this kind of data points,
- 6:00:22okay? Given this kind of binary
- 6:00:23classification problem,
- 6:00:26so the expectation is
- 6:00:29in case of support vector machine, I'm
- 6:00:31going to draw a hyperplane
- 6:00:33which separates as much as possible.
- 6:00:37Okay? So, I'm going to draw a hyperplane
- 6:00:40which separates these two classes as
- 6:00:43much as possible.
- 6:00:46So, it says that the I'm going to draw
- 6:00:49draw hyperplane
- 6:00:51and it it will be separated by a gap as
- 6:00:54wide as possible.
- 6:00:56So, that is the intuition behind support
- 6:00:58vector machine.
- 6:01:00Okay. Now that you have an intuition
- 6:01:02behind what is support vector machine,
- 6:01:05let's understand as how does this SVM,
- 6:01:08that is support vector machine, would
- 6:01:09work.
- 6:01:10So, in case of support vector machine,
- 6:01:13so here there is one more example. I
- 6:01:15have the set of
- 6:01:17points which is green green color and I
- 6:01:19have another set of points which are in
- 6:01:21red color. So, these two
- 6:01:24points are belonging to the different
- 6:01:26different classes.
- 6:01:28Now, what I'm going to do is I'm going
- 6:01:30to draw a hyperplane which separates
- 6:01:34these two classes data points as much as
- 6:01:36possible. And when I'm drawing the
- 6:01:38hyperplane, I'll make sure that this
- 6:01:40hyperplane is
- 6:01:42as this hyperplane is equidistant from
- 6:01:46my support vectors.
- 6:01:48Now, the support vectors is nothing but
- 6:01:51the point which is closer to my
- 6:01:52hyperplane.
- 6:01:54Now, here in this example that you're
- 6:01:55seeing,
- 6:01:57the this data point and this data point
- 6:02:01are called as support vectors because
- 6:02:03these are the data points which are
- 6:02:05nearest from my hyperplane that I've
- 6:02:06just drawn.
- 6:02:09In if I'm trying to make use of this SVM
- 6:02:12model, it is going to draw this kind of
- 6:02:14hyperplane
- 6:02:16to make sure that it is separating two
- 6:02:18classes. The two classes that we have
- 6:02:20over here in this example is red and
- 6:02:22green. It's going to separate these two
- 6:02:24classes as much as possible and it will
- 6:02:28be equidistant from my support vectors
- 6:02:31and the support vectors are nothing but
- 6:02:33the nearest point to my hyperplane.
- 6:02:36And that is how I'm going to separate
- 6:02:39between two classes when it comes to
- 6:02:40support vector machines.
- 6:02:44Now, here in this example that you're
- 6:02:46currently seeing,
- 6:02:47the hyperplane that I've just drawn, so
- 6:02:49this is a simple linear hyperplane.
- 6:02:53Just like a straight line that I'm
- 6:02:54trying to draw if I want to separate two
- 6:02:57classes of data points.
- 6:02:59Now, apart from drawing this straight
- 6:03:01line, we also have other kind of lines
- 6:03:05as well which we can draw.
- 6:03:07So, let's see how we can do that.
- 6:03:11So, the types of line that we can draw
- 6:03:13or the hyperplane that we can draw is
- 6:03:16called as SVM kernels, that is support
- 6:03:18vector machine kernels. The example that
- 6:03:20we have seen, it's an example for linear
- 6:03:23SVM kernels.
- 6:03:25So, let's see what are the other types
- 6:03:27of kernels that we have. So, when it
- 6:03:29comes to SVM kernels, we have linear
- 6:03:31kernels,
- 6:03:33radial basis function kernel and along
- 6:03:35with that, we also have polynomial
- 6:03:38kernel.
- 6:03:40Now, in case of linear kernel, I'm going
- 6:03:42to draw a hyperplane which is like a
- 6:03:44straight line.
- 6:03:45In case of polynomial kernel, I can draw
- 6:03:48my hyperplane on the basis of polynomial
- 6:03:50function that I have created on the
- 6:03:52basis of number of variables that I have
- 6:03:54and the degree that I have over there in
- 6:03:57case of polynomial. And in case of
- 6:03:59radial basis function, so I'll make use
- 6:04:01of radial basis to separate my data
- 6:04:04points.
- 6:04:07Okay, so these three are the important
- 6:04:10kernels that we have in SVM and this is
- 6:04:12one of the commonly asked interview
- 6:04:13question when it comes to the topic of
- 6:04:15support vector machines.
- 6:04:18Now, let's look at some of the use cases
- 6:04:21or the way we can do where we can use
- 6:04:24this SVM to uh
- 6:04:27work or let's look at some of the use
- 6:04:29cases where we can use this SVM.
- 6:04:32Okay.
- 6:04:34So, we can use this SVM
- 6:04:38on many of the use cases. So, to name a
- 6:04:41few, we can use it in face detection.
- 6:04:44We can use it in text and hypertext
- 6:04:46categorization. We can use the SVM if
- 6:04:49I'm trying to classify any images. I can
- 6:04:52make use in bioinformatics.
- 6:04:55And if I'm trying to detect something,
- 6:04:57so I can
- 6:04:59in the in an example here, remote
- 6:05:01homology detection, handwriting
- 6:05:03detection. Or in general, we can make
- 6:05:06use of this generalized predictive
- 6:05:08control. So, wherever we are dealing
- 6:05:10with the task of classification, we can
- 6:05:13use this SVM model. Okay. Now that we
- 6:05:17have a theoretical understanding as what
- 6:05:19is SVM and how it is actually going to
- 6:05:22look like and how it will be,
- 6:05:25let's have a quick walk through as how
- 6:05:28we can implement this SVM. Now, to
- 6:05:31implement this SVM,
- 6:05:33these are the common steps that we are
- 6:05:34going to follow.
- 6:05:36We are going to load the data.
- 6:05:39We'll explore the data.
- 6:05:41And once we have explored the data, we
- 6:05:43are going to split the data into two
- 6:05:45parts. The reason is simple. One, I have
- 6:05:47training, so I'll be making use of my
- 6:05:50training data.
- 6:05:51And once my training is complete, I'll
- 6:05:54check how my model has been trained with
- 6:05:56the help of my test data. So,
- 6:05:58I'm going to split the data.
- 6:06:00Now, once that is complete, we are going
- 6:06:02to train this SVM model. And finally, we
- 6:06:05can evaluate the model and observe as
- 6:06:08how model is working.
- 6:06:11So, this is the overview of the
- 6:06:13implementation of support vector
- 6:06:15machines.
- 6:06:17So, let's do one thing. Let's
- 6:06:21work it out and let's create the
- 6:06:23notebook in Google Colab and let's see
- 6:06:25it in action as how we can implement
- 6:06:27this SVM.
- 6:06:28I'll come back to my Google Colab.
- 6:06:31So, this is the notebook that I have
- 6:06:32already prepared and I'll give you a
- 6:06:35walk-through as we proceed along.
- 6:06:37Now, here in my first cell, I'm
- 6:06:39importing my NumPy library, Pandas
- 6:06:41library, and along with that, for
- 6:06:43creation of plots, I'm importing my
- 6:06:45Matplotlib library. Now, if you're
- 6:06:48comfortable with Seaborn, you can use
- 6:06:50the Seaborn library as well. So, in my
- 6:06:52example, I'm just making use of
- 6:06:54Matplotlib because we are not interested
- 6:06:57in creation of visualization, but we
- 6:06:59want to understand as how model is being
- 6:07:02working.
- 6:07:05Okay.
- 6:07:06And I'm going to execute this cell.
- 6:07:09So, this is going to take care of
- 6:07:10necessary imports. I'm importing my
- 6:07:12necessary libraries.
- 6:07:14And once that is done,
- 6:07:16here I'm importing this SVM. So, this
- 6:07:20SVM model is available inside my
- 6:07:22scikit-learn library. So, I've mentioned
- 6:07:24as
- 6:07:25scikit-learn .svm
- 6:07:29and from scikit-learn.svm, I'm importing
- 6:07:32my SVC. Okay? So, I'm importing my SVC.
- 6:07:35So, I'll show you what is this SVC.
- 6:07:38Um SVM SVC
- 6:07:45So, it's C means support vector
- 6:07:47classification.
- 6:07:48Okay? Now, here when I'm instantiating
- 6:07:52this SVC, I can mention what is the
- 6:07:54kernel that I want to use. And if I'm
- 6:07:57working with any polynomial kernel, then
- 6:07:59I can also mention what is the degree of
- 6:08:01polynomial that I want to use while
- 6:08:03performing the fit for my data set.
- 6:08:06So, I'm importing my SVC. And along with
- 6:08:09that, I'm also importing the data sets.
- 6:08:12So, in the scikit-learn library itself,
- 6:08:14we have a data set. So, it the
- 6:08:17scikit-learn host already like it it
- 6:08:19actually scikit-learn has many toy data
- 6:08:22set which will actually help us in our
- 6:08:24learning journey. So, we are going to
- 6:08:25use one of the data set, the famous Iris
- 6:08:28data set. We use that for multi-class
- 6:08:31classification.
- 6:08:32So, I'm going to load that Iris data
- 6:08:34set.
- 6:08:35And I'm just going to extract only two
- 6:08:37features. So, the two features that I'm
- 6:08:39extracting is petal length and petal
- 6:08:41width because I don't want to complicate
- 6:08:43it. I just want to visualize the data.
- 6:08:45So, in order to help in visualization, I
- 6:08:48I'm just getting only two features of my
- 6:08:50given data.
- 6:08:52And I'm separating my Y as
- 6:08:55Iris target. So, whatever the target
- 6:08:56variable that I had, I'm assigning to my
- 6:08:59variable of Y.
- 6:09:01Then,
- 6:09:02I'm going to uh
- 6:09:04do this check whether it is setosa or
- 6:09:07versicolor.
- 6:09:08That means this default data set, which
- 6:09:11is in multi-class classification, I'm
- 6:09:13just going to convert it into a binary
- 6:09:15classification task.
- 6:09:17You'll get a better understanding once I
- 6:09:19execute this next cell. So, this going
- 6:09:21to prepare my data set and once the data
- 6:09:24set is prepared, if I create a scatter
- 6:09:26plot, so I'm just creating the scatter
- 6:09:28plot to show us
- 6:09:30what and how my data set looks like. So,
- 6:09:33this is how my data set looks like.
- 6:09:36On my X axis, I think I'm having petal
- 6:09:38length. On my Y axis, I'm having petal
- 6:09:40width.
- 6:09:41And here,
- 6:09:42the blue points refers to the class zero
- 6:09:45and the orange points refers to the
- 6:09:48class of one.
- 6:09:50Okay? So, this is how my data set looks
- 6:09:54like.
- 6:09:55You can clearly see that I have one set
- 6:09:57of points in one region and I have
- 6:09:59another set of points in another region.
- 6:10:01Now, this is a classic example to
- 6:10:03understand about the SVM. How does it uh
- 6:10:06draw a hyperplane?
- 6:10:09So, we have the data set ready.
- 6:10:11And as I mentioned already, in order to
- 6:10:15fit this model, so when I say support
- 6:10:17vector machine, I'm going to draw a
- 6:10:18line.
- 6:10:20This line that I have drawn, it will be
- 6:10:23equidistant from my support vectors.
- 6:10:25Now, here in this example, the support
- 6:10:27vector is this because this is the only
- 6:10:30point which is nearest to my line. And
- 6:10:32here, I think this is the data point
- 6:10:34which is nearest to my SVM SVM line,
- 6:10:37that is this twisted line hyperplane
- 6:10:38line.
- 6:10:39So, I'll be placing this hyperplane such
- 6:10:42that it is equidistant from the support
- 6:10:45vectors.
- 6:10:47That is how I'll be drawing this support
- 6:10:49vector line.
- 6:10:51So, we now have an intuition. Let's see
- 6:10:53whether we get the same outcome as we
- 6:10:55are expecting.
- 6:10:57So, here I'm initializing my model. So,
- 6:11:00for initialization, I'm saying it as
- 6:11:02SVC. Use the kernel as linear because
- 6:11:06I'm able to draw a line effectively. We
- 6:11:08were We just seen. And I'm using the C
- 6:11:11as infinity, that means it should be a
- 6:11:13hard classifier. So, hard classifier
- 6:11:15means I make I want the 100% result. I
- 6:11:18mean, I don't want any loosens. I want
- 6:11:20to draw a line which passes which
- 6:11:23clearly separates two classes. So, I'm
- 6:11:25saying it as C as infinity to mention
- 6:11:27this as a hard classifier.
- 6:11:30And once I initialize any model,
- 6:11:32here in this scenario, SVM model, I'm
- 6:11:35performing the fit on my data set. Now,
- 6:11:36this is the common flow that we follow
- 6:11:39whenever we are performing the fit. So,
- 6:11:41we'll initialize the model and then we
- 6:11:43perform the fit on a data set.
- 6:11:46Now, since this SVM being a supervised
- 6:11:49machine learning model, I have to
- 6:11:51specify both my input X as well as my
- 6:11:55output Y.
- 6:11:56Hence,
- 6:11:57SVM classifier.fit
- 6:12:00X, Y.
- 6:12:01So, this is going to perform the fit for
- 6:12:03my data set.
- 6:12:05I'll just execute this. So, this has
- 6:12:08performed the fit and here it is giving
- 6:12:10me the confirmation as this is the
- 6:12:13parameter that are being used to perform
- 6:12:15the fit.
- 6:12:17Okay. Now, once I have drawn and once I
- 6:12:20have found this fit,
- 6:12:22next,
- 6:12:23if I want to display the weight terms,
- 6:12:26so, I can say it as SVM
- 6:12:28classifier.coefficients.
- 6:12:30So, these are the weight terms. And if I
- 6:12:32want to display my bias term or the
- 6:12:34intercept, it is minus 3.78.
- 6:12:38Now, this means the line that I've just
- 6:12:40drawn, so that line has the
- 6:12:44uh
- 6:12:45that line has the C term as or the W not
- 6:12:48term as minus 3.78 and W1, W2 are 1.29
- 6:12:53and 0.82, respectively.
- 6:12:56So, that's how the data is distributed
- 6:12:59for us. That's how the values has been
- 6:13:02formed for our scenario.
- 6:13:05Next, in order to get the better
- 6:13:07visualization,
- 6:13:08here I have created a function that is
- 6:13:10called as plot SVC decision boundary and
- 6:13:14this takes my SVM model,
- 6:13:17the X min and the X max.
- 6:13:21Now, W and B I'm extracting from the
- 6:13:24coefficient and the intercept parameter
- 6:13:26that we have over here.
- 6:13:28So, we are extracting from this
- 6:13:30uh
- 6:13:31at from this attributes that we have
- 6:13:33from this model.
- 6:13:35And now, if I want to draw a decision
- 6:13:37boundary,
- 6:13:38so,
- 6:13:39I need the set of points. So, in order
- 6:13:42to get the points, I'm saying it as X
- 6:13:43not is equal to np.linspace X max, X
- 6:13:46min, X max, 200. And I'm specifying as
- 6:13:50how does my decision boundary should
- 6:13:52look like.
- 6:13:53My decision boundary is given by w
- 6:13:56naught into x naught plus w one into x
- 6:13:58one plus b is equal to zero. So, this is
- 6:14:00what my
- 6:14:01decision boundary would look like. So, I
- 6:14:03know what is x naught. I know w naught.
- 6:14:06I also have w one and I also have b. So,
- 6:14:09the only term that I do not have is my
- 6:14:12x one.
- 6:14:13Okay? So, the only term that I do not
- 6:14:15have over here in this example is x one.
- 6:14:18And the x one if I want it, so I just
- 6:14:20have to substitute it. So, x one is
- 6:14:22equal to minus w zero divided by w one
- 6:14:25into x naught minus b divided by w one.
- 6:14:30Now, I'm specifying the same equation
- 6:14:33over here for my x two.
- 6:14:35So, my x naught and the decision
- 6:14:37boundary will give me the pair of input
- 6:14:40and output. Okay?
- 6:14:42Now, along with this
- 6:14:44there is a property in SVM. Okay? So,
- 6:14:47the property is given by whenever I have
- 6:14:49a margin, so that margin is given by one
- 6:14:53over w one.
- 6:14:55Okay? So, the margin is nothing but the
- 6:14:58distance between my hyper plane and the
- 6:15:01support vector. So, that is given by ma
- 6:15:04one by w one.
- 6:15:06Hence, I have mentioned as gutter up and
- 6:15:08down. Gutter up means one line or the
- 6:15:11one line where the support vector lies.
- 6:15:13So, that is given by decision boundary
- 6:15:16plus margin.
- 6:15:17And one line below my
- 6:15:19one line below my hyper plane.
- 6:15:22That is where another support vector
- 6:15:24would lie. So, I mentioned as decision
- 6:15:26boundary minus margin.
- 6:15:29Okay?
- 6:15:30Then
- 6:15:31I'm defining where exactly my support
- 6:15:34vectors are present.
- 6:15:37My support vectors uh I can access the
- 6:15:40support vectors coordinates by saying it
- 6:15:42as
- 6:15:43by accessing the attribute of my train
- 6:15:45model support underscore vectors
- 6:15:47underscore.
- 6:15:49Now, I'm specifying where exactly those
- 6:15:51support vectors are present with the
- 6:15:53help of a simple scatter plot by
- 6:15:55highlighting my support vectors.
- 6:15:57And I'm specifying where exactly my
- 6:15:59decision boundary is present. And I'm
- 6:16:02also mentioning where is my line that is
- 6:16:04gutter up and gutter down. So, let's do
- 6:16:06one thing. I'll just execute this. This
- 6:16:08is going to create me a function.
- 6:16:11I'm going to call my function
- 6:16:13support vector machine classification.
- 6:16:16And I will specify my range of X and Y
- 6:16:18as
- 6:16:19here, yeah, X min and X max as 0 {comma}
- 6:16:225.5.
- 6:16:24I'll just execute this.
- 6:16:28So,
- 6:16:29what we have done just now is we have
- 6:16:32created this hyperplane.
- 6:16:36So, the middle one, the solid line that
- 6:16:38you're seeing over here, so this solid
- 6:16:40line is called as your hyperplane.
- 6:16:43And these points that you're seeing over
- 6:16:45here, so these two points which are
- 6:16:47highlighted, these two points are
- 6:16:49actually called as support vectors.
- 6:16:54Okay? So, this dotted line that you're
- 6:16:57seeing, so this dotted line refers to my
- 6:17:00gutter up and gutter down which I've
- 6:17:02found right here.
- 6:17:05Let's do one thing. I'll add some label
- 6:17:07so that you'll get some more
- 6:17:09visualization in the plot itself. I'll
- 6:17:11say label and I'll mention it as
- 6:17:14hyperplane.
- 6:17:19Okay.
- 6:17:21And
- 6:17:23there is one more, yeah.
- 6:17:36These are support vectors.
- 6:17:38And I'll say
- 6:17:42plt.legend.
- 6:18:00So, this clearly says which are all my
- 6:18:04hyperplane and which are all my support
- 6:18:06vectors.
- 6:18:09So, this is the intuition behind support
- 6:18:11vector machines.
- 6:18:13So, we'll be drawing a hyperplane which
- 6:18:16separates the points that we have.
- 6:18:18Okay? And whichever the point which is
- 6:18:20nearest to my hyperplane, we call that
- 6:18:22point as a support vector. Now, to
- 6:18:25access that support vector, we make use
- 6:18:27of the attribute. So, let's do one
- 6:18:29thing. Let's explore the same the
- 6:18:31attributes.
- 6:18:32svm.
- 6:18:33support_vectors_.
- 6:18:36So, this is going to tell me where
- 6:18:37exactly my support vectors are present.
- 6:18:40So, one point is given by 1.9.0.4.
- 6:18:43I think this is the point that I'm
- 6:18:45talking about.
- 6:18:46And the another support vector that we
- 6:18:48have is at the location 3, 1.1. So, 3
- 6:18:51and 1.1. This is where we have another
- 6:18:54support vector.
- 6:18:56So, using all these attributes, we have
- 6:18:58been able to create this visualization.
- 6:19:03Okay.
- 6:19:05Now,
- 6:19:06whenever we are working with the support
- 6:19:07vector machines,
- 6:19:09it's very important that we scale the
- 6:19:11data first. If I do not scale the data,
- 6:19:14I'll not be able to get a better fit of
- 6:19:17my SVM model.
- 6:19:19So, here I've given one more example
- 6:19:22where I have my X
- 6:19:24uh is given as 1, 55, 23, 80. As you can
- 6:19:28clearly see, it's it's not scaled. Okay?
- 6:19:32So, I'm going to execute this cell. So,
- 6:19:35this is going to tell me and give me a
- 6:19:37visualization as how the
- 6:19:39fit will be in case of scaled and
- 6:19:42unscaled.
- 6:19:44See, if it is unscaled
- 6:19:47I'll If it is not scaled, okay? That
- 6:19:49means if it is unscaled, we can clearly
- 6:19:51see that the hyperplane that I'm drawing
- 6:19:55and the distance from my hyperplane,
- 6:19:57it's very close to each other.
- 6:20:01And whenever I'm working, it's It's It
- 6:20:03will be difficult for me to separate
- 6:20:04those two data points.
- 6:20:08But, if I scale them correctly
- 6:20:11Now, here for scaling, I have made use
- 6:20:12of a scalar standard scalar. Now, if I
- 6:20:15scale it correctly, then in that
- 6:20:17scenario, it will be easier for me and
- 6:20:20it would actually work better when I
- 6:20:22have scaled data.
- 6:20:26Okay? So
- 6:20:28this is about using the linear SVM model
- 6:20:32to perform the fit on my given data set.
- 6:20:36Now, if I go below, we have some more
- 6:20:38examples about non-linear classifiers as
- 6:20:41well.
- 6:20:42Now, in order to test out the same
- 6:20:45here
- 6:20:46I'm creating an example data set and
- 6:20:48that data set that I'm generating is
- 6:20:50called as make moons data set and this
- 6:20:53has been generated with the help of a
- 6:20:55scalar data set generator.
- 6:20:57Now, as you can clearly see, I cannot
- 6:21:00make use of linear classifier. So,
- 6:21:02linear classifier is nothing but a
- 6:21:05classifier, okay? Which is an SVM model
- 6:21:08where I'm drawing or where I'm using a
- 6:21:10straight line to split my data points. I
- 6:21:13can clearly see that I Wherever I I join
- 6:21:16or wherever I try to draw a line over
- 6:21:19here, I cannot split the data in an
- 6:21:21effective manner.
- 6:21:23Now, this brings us the challenge. Now,
- 6:21:24if I have a data set in this way where I
- 6:21:27cannot linearly separate it, how can we
- 6:21:30go about and fit uh perform the fit on
- 6:21:33our SVM model?
- 6:21:34So, in order to save us, we have a model
- 6:21:37that is called as uh SVM model, and from
- 6:21:40that SVM model, we can actually create a
- 6:21:43polynomial uh
- 6:21:45polynomial kernel. So, we can make use
- 6:21:46of polynomial kernel, and using that
- 6:21:48polynomial kernel, I can actually uh
- 6:21:52create it like this. I mean, using
- 6:21:53polynomial kernel, I can perform
- 6:21:55polynomial regression.
- 6:21:57Or I can draw a line like this. Now, to
- 6:21:59show you how it works,
- 6:22:01I'm getting some data like this. So,
- 6:22:03this is some uh random data.
- 6:22:06And I'm making use of pipeline.
- 6:22:08So, this pipeline is going to take care
- 6:22:10of my stan- standard scaler as well as
- 6:22:13kernel.
- 6:22:14I'll do one thing, I'll just come below.
- 6:22:16So, this is what we are currently
- 6:22:17interested in.
- 6:22:23So, here,
- 6:22:26I'm importing the polynomial features,
- 6:22:29and I'm generating the polynomial
- 6:22:31features for my data.
- 6:22:33I'm performing the fit and transform my
- 6:22:35polynomial data. That means, I'm just
- 6:22:37modifying my existing data, and I am
- 6:22:39sending it
- 6:22:41for my uh
- 6:22:43X, okay? So, this is how my pair of
- 6:22:45input X and Y looks like.
- 6:22:47Now, I'll use my X. I'm going to
- 6:22:49transform it with the help of my
- 6:22:51polynomial features,
- 6:22:53and then, I'm going to scale it with the
- 6:22:56help of my standard scaler,
- 6:22:58and I'm going to send it inside my
- 6:23:01classifier, that is SVM classifier.
- 6:23:08Okay? So, I'm going to
- 6:23:11combine it together like this.
- 6:23:14Now, observe what would happen.
- 6:23:17Now, once that is complete,
- 6:23:19see?
- 6:23:20With the help of my polynomial uh
- 6:23:24polynomial features that have applied on
- 6:23:26my given linear data.
- 6:23:28So, I have increased the degrees by
- 6:23:32which my model can learn.
- 6:23:36Now, instead of straight line, my model
- 6:23:37is also having the ability to learn this
- 6:23:40complex representation as well.
- 6:23:43Because I have increased the model
- 6:23:45complexity by adding my polynomial
- 6:23:47features.
- 6:23:50And while doing it, to make sure that we
- 6:23:51follow a
- 6:23:53clear path, so I have defined this is
- 6:23:56scalar's pipeline. So, if you're new to
- 6:23:58data science machine learning, I highly
- 6:24:00recommend you to learn this concept of a
- 6:24:02scalar pipeline. Now, this is scalar
- 6:24:04pipeline helps us to combine multiple
- 6:24:07operations in a single call.
- 6:24:10So, here we have created a pipeline.
- 6:24:12This pipeline is going to add some
- 6:24:15polynomial features for my input data.
- 6:24:18And on top of it, this is going to
- 6:24:19perform scaling. And then I'm going to
- 6:24:21perform this binomial classification
- 6:24:24using this SVM.
- 6:24:27And finally,
- 6:24:29I'm performing the fit on my data set.
- 6:24:31See, when I perform the fit, it takes my
- 6:24:33input X and it's going to do all these
- 6:24:36activities. It is going to chain all
- 6:24:38these activities together, and then it
- 6:24:41is going to perform the fit for my data
- 6:24:43Y.
- 6:24:45Once the fit has been complete, so we
- 6:24:47can validate how my model is performing.
- 6:24:50>> [music]
- 6:24:55[music]
- 6:24:56>> So, what is a clustering technique?
- 6:24:57Clustering technique is something that
- 6:24:59we will use it for grouping purpose.
- 6:25:03So, especially there's a very easy way
- 6:25:06to understand what is clustering
- 6:25:07technique. You would have seen such a
- 6:25:09while we are going through an
- 6:25:11COVID-19 situation, the governments has
- 6:25:13came up with creating some containment
- 6:25:15zones.
- 6:25:16As all of you must be knowing.
- 6:25:19So, how on what criteria government has
- 6:25:21taken that okay, which area supposed to
- 6:25:23be a containment zone or which area
- 6:25:25supposed to be applied with some some
- 6:25:26restrictions and which areas can be can
- 6:25:29be considered as normal? On what
- 6:25:30criteria that they have created? So,
- 6:25:32that's what using clustering technique.
- 6:25:35Which means if the governments or when I
- 6:25:37say government means that the people who
- 6:25:38will be taking the final decision in
- 6:25:40such criteria, either prime minister
- 6:25:41either either the chief ministers of
- 6:25:43that particular state will be be taking
- 6:25:45decisions whether to go for lockdown
- 6:25:47whether to not to go for lockdown or
- 6:25:49which areas has to be considered as
- 6:25:51containment zones or non-containment
- 6:25:52zones.
- 6:25:53So, those high-level decisions are
- 6:25:55something which will be taken based on
- 6:25:57the clustering technique output which is
- 6:25:59generated by the these algorithms.
- 6:26:02Based on a number of inputs, okay, what
- 6:26:04is the population in a particular area?
- 6:26:06How many number of people are affected?
- 6:26:08How many number of hospitals which are
- 6:26:10present?
- 6:26:11How many number of
- 6:26:13people who are been recovered? So,
- 6:26:16likewise based on this these multiple
- 6:26:18criteria, people will do some clustering
- 6:26:20technique on top of the data and
- 6:26:22according to that people will be
- 6:26:24segregated or the areas will be
- 6:26:26segregated.
- 6:26:27So, that saying that okay, these are the
- 6:26:28observations which belong to one
- 6:26:29cluster, these are the observations
- 6:26:31which belong to one cluster like that so
- 6:26:32that people can cluster them which can
- 6:26:34make organizations to take decisions on
- 6:26:38a very high level.
- 6:26:40Okay? That's what is all clustering
- 6:26:41technique.
- 6:26:42Which clustering technique output will
- 6:26:45contain the different different groups.
- 6:26:46It itself will group the different
- 6:26:47different components
- 6:26:49based on whatever the number of clusters
- 6:26:51that you want to generate. That may not
- 6:26:53give the direct output. On top of the
- 6:26:55generated output, people will be taking
- 6:26:56business related decisions. That's what
- 6:26:58is all about clustering techniques.
- 6:27:01Okay. So, now what we will do?
- 6:27:03Let's take an example of within
- 6:27:06clustering techniques, what are the
- 6:27:07different types of clustering techniques
- 6:27:08we have?
- 6:27:09So, what What different types of
- 6:27:10clustering that we have?
- 6:27:13So, there are multiple types of there
- 6:27:14are multiple types of ways based on the
- 6:27:17type of output that we want to produce.
- 6:27:18There are multiple different types of
- 6:27:20clustering techniques we have, but out
- 6:27:21of which the let's try to understand
- 6:27:23about what are the very famous and most
- 6:27:25widely used clustering technique
- 6:27:26algorithm. Out of which we have
- 6:27:28something called K-means clustering
- 6:27:30algorithm is one of the very famous and
- 6:27:32most widely used. More than 90% of the
- 6:27:35people will end up with using K-means
- 6:27:36clustering algorithm, which is very very
- 6:27:38famous in clustering techniques.
- 6:27:40Right? Which is very very famous in
- 6:27:42clustering techniques.
- 6:27:43So, what are these clustering
- 6:27:44techniques? As I said, how this
- 6:27:46clustering technique will work.
- 6:27:48So, K-means clustering is nothing but
- 6:27:49always remember one thing. If any one of
- 6:27:51you were going to work in machine
- 6:27:52learning or anywhere in anywhere,
- 6:27:55wherever you see a notation called K,
- 6:27:57by default K is nothing but you are you
- 6:28:01are supposed to as a user, you are
- 6:28:03supposed to provide what is the input of
- 6:28:06K. Which means wherever you see there
- 6:28:08are multiple techniques that we have in
- 6:28:09machine learning like K-means clustering
- 6:28:11technique, K nearest neighbor is one of
- 6:28:13the algorithm, K-fold cross validation,
- 6:28:15likewise. Wherever you see a notation
- 6:28:17called K, what is what does a K means?
- 6:28:19It's an input that you are supposed to
- 6:28:21provide. Always remember this.
- 6:28:24It's an input that you are supposed to
- 6:28:25provide
- 6:28:27to your algorithm. Your algorithm cannot
- 6:28:29identify that K value. Of course,
- 6:28:30everything else will be taken care by
- 6:28:31your algorithm, but whenever you see K,
- 6:28:33which means in K-means clustering, what
- 6:28:35is the meaning of K-means clustering?
- 6:28:37How many number of clusters that you
- 6:28:38want to provide? That is something that
- 6:28:40you have to input it to your algorithm.
- 6:28:43That is something that you have to input
- 6:28:44to your algorithm.
- 6:28:45Right?
- 6:28:47Here, the meaning of K is how many
- 6:28:50clusters that you want to generate.
- 6:28:52So, how many clusters that you want to
- 6:28:54generate? Okay, when you have 1,000
- 6:28:55observations which are present, when you
- 6:28:57have 1,000 in input column input records
- 6:28:59which are present in your historical
- 6:29:01data, how many number of clusters that
- 6:29:03you want to provide? Do you want to go
- 6:29:04for one cluster? Obviously, one cluster
- 6:29:06means that the entire data set will be
- 6:29:07considered as is.
- 6:29:09Do you want to create two clusters out
- 6:29:11of the data?
- 6:29:12Do you want to create three clusters out
- 6:29:14of the data? Four clusters, five
- 6:29:15clusters, or 10 clusters?
- 6:29:17So, how this can be done?
- 6:29:18There are multiple steps that are
- 6:29:20involved in generating K-means
- 6:29:22clustering algorithm. So, you can see,
- 6:29:23choose the number of clusters. This is
- 6:29:25what is nothing but your first step. It
- 6:29:26means you need to decide what is your
- 6:29:29K is nothing but number of clusters that
- 6:29:31you want to produce.
- 6:29:32And then, there is an initialization of
- 6:29:35centroids will happen as a one-time
- 6:29:37activity.
- 6:29:38Right? So, there is an initialization of
- 6:29:40centroids which will be which will be
- 6:29:42declared that will that will be used as
- 6:29:44your initial step for your machine. And
- 6:29:45then, assign the clusters, move the
- 6:29:47centroids, and optimization, and then
- 6:29:49converge the the
- 6:29:51all the clusters into one component.
- 6:29:53Yes, I know it will be very difficult to
- 6:29:54understand by looking at this thing. So,
- 6:29:56let me show you a very simple example
- 6:29:58how exactly it will be done. Maybe let
- 6:29:59me take a simple diagram for you to show
- 6:30:01how exactly that's going to work.
- 6:30:05Okay? So, let's say for example, I'm
- 6:30:07going to take some historical data just
- 6:30:09to explain you on how exactly the
- 6:30:11K-means clustering algorithm will work.
- 6:30:13So, what is that it is written in the
- 6:30:14first step?
- 6:30:16What is the data is written in the first
- 6:30:17step?
- 6:30:18So, the choose the number of clusters.
- 6:30:20Choose number of clusters. Now, it means
- 6:30:22say for example, we need to take some
- 6:30:24historical data. I'm considering some
- 6:30:25historical data here. Let's assume this
- 6:30:27is the historical data.
- 6:30:29Let's assume this is the historical data
- 6:30:31that we have.
- 6:30:32So, as you can see, there are multiple
- 6:30:35historical data points. As you can see,
- 6:30:37the first step, choose number of
- 6:30:38clusters, which means let's assume to
- 6:30:40make this thing simple, I'm going to
- 6:30:42choose that we want to have two clusters
- 6:30:43generated. Okay? So, this is which means
- 6:30:46we want to have two clusters created.
- 6:30:47So, because my number of clusters that I
- 6:30:49want to generate is two, I'm going to
- 6:30:51consider that there are two centroids
- 6:30:52which are present. This is my first
- 6:30:54step. How this algorithm will do?
- 6:30:56How algorithm will come to come with the
- 6:30:58number of clusters? So, this is how it
- 6:30:59will happen. So, the first step is to
- 6:31:01choose the number of centroid and then
- 6:31:03initialize your centroid. That's the
- 6:31:04second step. So, now what is the third
- 6:31:06step?
- 6:31:07Let's assume this is observation number
- 6:31:09one. This is our data point one.
- 6:31:12So, now what what is the next step?
- 6:31:14We will take the distance from every
- 6:31:16observation to every centroid and every
- 6:31:18observation to every centroid, which
- 6:31:19means now tell me which observation
- 6:31:23For this observation number one, which
- 6:31:24centroid is more closer?
- 6:31:26Is the green color centroid is more
- 6:31:27closer or the red color centroid is more
- 6:31:29closer for this?
- 6:31:31Centroid means that data points, the one
- 6:31:33which I have highlighted, that's what we
- 6:31:34call a centroid, data points.
- 6:31:37We have chosen two centroids only
- 6:31:39because that is the as a user you're
- 6:31:41supposed to input what's supposed to be
- 6:31:42your K.
- 6:31:43What's supposed to be your K value,
- 6:31:45that's what I said. You need to know how
- 6:31:47what is your K is nothing but how many
- 6:31:48number of clusters that you want to
- 6:31:49provide. Usually your business users are
- 6:31:51going to provide that. In case if you're
- 6:31:53going to work in this kind of
- 6:31:54algorithms, they will provide that. Or
- 6:31:56else there are some other methods
- 6:31:58available like
- 6:31:59elbow method available other things.
- 6:32:01Yeah, we'll talk about that.
- 6:32:03Now, this observation is more closer to
- 6:32:05red color. So, now what happens? What
- 6:32:07what is the next step? The algorithm
- 6:32:08will assign this particular observation
- 6:32:10one as a red color for now.
- 6:32:12As considering that this is belong to
- 6:32:13red color. Likewise for the second
- 6:32:15observation, when the second observation
- 6:32:16appears, what is the distance from the
- 6:32:18second observation to both the
- 6:32:20centroids? Now, which one is more
- 6:32:21closer? I see green color is more
- 6:32:23closer. So, now I'll mark this as green
- 6:32:25color.
- 6:32:26If both if what if the distance is same
- 6:32:28equal? So, then the algorithm will force
- 6:32:31any of the observation to get into any
- 6:32:33of the centroid.
- 6:32:35So, number of clusters number of
- 6:32:37centroids that you will choose based on
- 6:32:38the K value as I said.
- 6:32:41Now, you're going to Likewise, you will
- 6:32:43repeat the same process and whatever the
- 6:32:45observation which is more closer to
- 6:32:47whatever the centroid it is, you will
- 6:32:48mark them as with their so-called mark
- 6:32:51like this. Now, you're going to
- 6:32:52initially mark them as observations into
- 6:32:54either into green color either into red
- 6:32:56color.
- 6:32:56So, now, these observations are now
- 6:32:59considered as green color observations,
- 6:33:00and these observations are now
- 6:33:02considered as red color observations.
- 6:33:05This is step number one.
- 6:33:07What is the step number two?
- 6:33:09Step number two is segregate all these
- 6:33:11red color observations and take the
- 6:33:13average value of X and Y coordinates,
- 6:33:15and repeat the same process for
- 6:33:17calculating average coordinates of X and
- 6:33:18Y coordinates for the green color
- 6:33:20observations.
- 6:33:21Take the average of all these
- 6:33:22observations, calculate average, and
- 6:33:24take the all these observations, take
- 6:33:25the average. You end up with getting a
- 6:33:27new centroid positions called XY. Which
- 6:33:30means, now you ended up with getting a
- 6:33:32new centroids in the initial step that
- 6:33:35we have taken random centroid. Now, you
- 6:33:37got the centroids that you can use based
- 6:33:40on the previous iteration. Now, you
- 6:33:41ended up with getting a new centroids.
- 6:33:44You repeat the same process again.
- 6:33:46Again, you repeat the same process. Take
- 6:33:47the distance from every observation to
- 6:33:49every centroid and assign the
- 6:33:50observation based on the nearest nearest
- 6:33:52to distance, and continue to mark every
- 6:33:55observation either into red color or
- 6:33:56either green color or whatever it is,
- 6:33:58and you repeat the process until you
- 6:34:00will be able to see there will be no
- 6:34:02change applicable for your clusters.
- 6:34:05You repeat the process. Which means, in
- 6:34:07every step, you might end up with
- 6:34:08changing your centroid. Every step will
- 6:34:10continue to change your centroid. Now,
- 6:34:12the centroid might become like this.
- 6:34:13Then, later, your centroid will become
- 6:34:15like this.
- 6:34:17Likewise, your centroids will be keep on
- 6:34:19moving. Initially, you have taken it
- 6:34:20like this, but it it might continue to
- 6:34:22move like this. Somewhere, it will be
- 6:34:23fixed. And after that, there will be no
- 6:34:25change that you will notice if you are
- 6:34:26repeating the same process. At this
- 6:34:28particular stage, whatever the
- 6:34:30observations which are marked into which
- 6:34:32are
- 6:34:33grouped into green color, you say like
- 6:34:35these are the green color observations.
- 6:34:37Whatever the observations which are
- 6:34:39marked into red color, you will you will
- 6:34:40mark them as okay, these are the
- 6:34:41observations which belong to red color.
- 6:34:44These are the observations which belong
- 6:34:45to red color. Likewise, you will
- 6:34:47segregate all these observations either
- 6:34:49into red color or green color.
- 6:34:51Right? So, so, we can generate these
- 6:34:54cluster techniques. That's how the
- 6:34:55K-means clustering algorithm will
- 6:34:56generate these algorithms
- 6:34:58output of using this algorithm.
- 6:35:00Okay? So, that's how K-means clustering
- 6:35:02algorithm will work.
- 6:35:04Okay?
- 6:35:05So, now what likewise there are multiple
- 6:35:06algorithms that we have. So, like
- 6:35:09when we talk about machine learning, so
- 6:35:10there are multiple types of algorithms
- 6:35:12that we have. So, like how we have the
- 6:35:14how does the K value K value will occur
- 6:35:16as you can see on the PPTs which are
- 6:35:17also mentioned. So, you're going to
- 6:35:19choose some randomly generated K value
- 6:35:21and you will be choosing the number of
- 6:35:22case over here and you can see that it
- 6:35:24will assign them based on the number of
- 6:35:26these easy
- 6:35:27most nearest distances and according to
- 6:35:29that you will change your centroids.
- 6:35:31Once you change your centroids, you will
- 6:35:33repeat the same process until you're
- 6:35:34able to change that your centroids don't
- 6:35:37move further and then once it has been
- 6:35:39finalized, you will say like this is the
- 6:35:40final centroid.
- 6:35:42Final cluster output that we can
- 6:35:44generate out of this.
- 6:35:46Right?
- 6:35:47So, likewise we also have different
- 6:35:48types of cluster technique and the
- 6:35:49second type of cluster technique that we
- 6:35:51have is a fuzzy or C-means clustering.
- 6:35:53So, what is fuzzy or C-means clustering?
- 6:35:54The output will remain same. So, C-means
- 6:35:56clustering means that there are places
- 6:35:58that one or two observations can belong
- 6:36:00to one or two different clusters.
- 6:36:03Like one or two different clusters,
- 6:36:05which means usually the primary
- 6:36:06difference between
- 6:36:09the primary difference between your
- 6:36:10C-means and K-means clustering technique
- 6:36:12is in case if there are any observations
- 6:36:15which are having equal amount of
- 6:36:16distance, usually in K-means clustering
- 6:36:18what we will do, we will force this
- 6:36:20observation to be part of any of the
- 6:36:22cluster. But in C-means clustering,
- 6:36:25based on the distance that we see, there
- 6:36:27are chances that an observation can go
- 6:36:29to or can belong to one or two clusters.
- 6:36:33So, it purely depends on the business
- 6:36:34use case who is going to decide to
- 6:36:36either to go for K-means clustering or
- 6:36:38C-means clustering based on the the
- 6:36:39business use case.
- 6:36:41Based on our business outcome, so if
- 6:36:43people are going to decide whether to go
- 6:36:44for C-means clustering or K-means
- 6:36:45clustering. So, there are n-number of
- 6:36:48observations which might It's not
- 6:36:50mandatory. Which might can belong to one
- 6:36:52or more clusters. That can happen.
- 6:36:55So, the third is agglomerative
- 6:36:57clustering, so which which is the third
- 6:36:58type of clustering technique that we
- 6:37:00have. Okay? So, which is also one of one
- 6:37:03of the clustering technique that we
- 6:37:04have.
- 6:37:06Okay?
- 6:37:07So, now which can also be used here.
- 6:37:09Okay? So, now what is this agglomerative
- 6:37:12clustering? So, this is what we also
- 6:37:14call it as hierarchical clustering. So,
- 6:37:16in times you'll also call them as
- 6:37:18hierarchical clustering. These
- 6:37:19clustering techniques are built using
- 6:37:20H-clustering. In short, we call it as
- 6:37:22H-clustering.
- 6:37:23And also people will also call it as
- 6:37:24hierarchical clustering techniques. So,
- 6:37:25what are this? So, based on the type of
- 6:37:28algorithm
- 6:37:29they will try to There is a There is a
- 6:37:31mathematical expression which are
- 6:37:32involved in it. But, then considering
- 6:37:33that the limited time that we have, I'm
- 6:37:35not going to take you through all the in
- 6:37:36in detail depth of it. So, considering
- 6:37:38the way how the data points are being
- 6:37:39segregated, we'll it will build a kind
- 6:37:41of dendrogram. So, on top of this
- 6:37:43dendrogram, your observations can be
- 6:37:45classified here like this, as you can
- 6:37:46see on the screen. Is K-means clustering
- 6:37:48sensitive to outlier?
- 6:37:50You need to understand one thing.
- 6:37:52When we are talking about unsupervised
- 6:37:54learning algorithms, as I said, you may
- 6:37:56or may not have clarity on the data.
- 6:37:59Which means
- 6:38:01your assumption is that at the least
- 6:38:02level
- 6:38:04you don't have clarity on the data.
- 6:38:06Then if you don't When you don't have
- 6:38:08clarity on the data, how can you say
- 6:38:09that This is an outlier or this is not
- 6:38:11an outlier?
- 6:38:12When you have clarity on the data,
- 6:38:13that's especially when you're working on
- 6:38:14supervised learning algorithms, you can.
- 6:38:16But, you don't have output column also.
- 6:38:18How you will be able to evaluate how the
- 6:38:20so-and-so-called output column can be
- 6:38:22can be evaluated because this is not
- 6:38:24being classified properly, this is not
- 6:38:25being clustered properly based on the
- 6:38:27historical data because there is no
- 6:38:28output column.
- 6:38:29So, those type of concepts are something
- 6:38:31which you don't need to worry about when
- 6:38:32you are working on supervised learning.
- 6:38:34They will be primarily they'll be
- 6:38:35constrained when you are working on
- 6:38:37supervised learning algorithms.
- 6:38:38And of course, in case if you see that
- 6:38:40there are outliers which are present,
- 6:38:41obviously it has to be It will be
- 6:38:43considered It will be considered as one
- 6:38:44of the cluster in of any any of any of
- 6:38:46the cluster it belongs to the data. But
- 6:38:48anyhow, that's the characteristic of the
- 6:38:49data. You don't need to
- 6:38:51You don't need to take care because you
- 6:38:52might be killing the actual original
- 6:38:54values which are present in the data.
- 6:38:55But you cannot expect that outlier can
- 6:38:57be recognized
- 6:38:58for all the cases that you have in
- 6:38:59unsupervised learning.
- 6:39:01So, the next type of clustering
- 6:39:02technique that we have is division
- 6:39:04clustering. So, what is this division
- 6:39:05clustering? As a division clustering is
- 6:39:07also creates the data in a form of
- 6:39:09dendrogram. But the difference is you
- 6:39:11can see that the starts with all data
- 6:39:12points in one cluster, splits the root
- 6:39:14into child recursively based on the
- 6:39:16dendrogram, and stops when there is a
- 6:39:18single term clusters which are created.
- 6:39:19Which means that for every cluster there
- 6:39:21will be one observation which belong to.
- 6:39:23So, likewise the clustering technique
- 6:39:24will work.
- 6:39:25Now, there is one more technique that we
- 6:39:27have in terms of building a clustering
- 6:39:28technique, that's what we call it the
- 6:39:29mean shift clustering. So, what we will
- 6:39:31do we'll take an average of every
- 6:39:33cluster that we have. What we will do
- 6:39:35we'll end up with reducing their means
- 6:39:36into the into a single density of items,
- 6:39:39and then we'll continue you to repeat
- 6:39:40the process to see to that which
- 6:39:42observations mean will belong to the
- 6:39:43same cluster, and according to that
- 6:39:45we'll continue to identify which
- 6:39:47observation will belong to a cluster a
- 6:39:48particular cluster.
- 6:39:50So, likewise we can apply different
- 6:39:52types of clustering technique that can
- 6:39:53help us to cluster the data which is
- 6:39:55part of unsupervised learning.
- 6:39:57Right? So, now let's take a let's jump
- 6:40:00into something called a small hands-on.
- 6:40:01Okay, let's let's take a small
- 6:40:03Python example, and we will see how do
- 6:40:05we build that particular a clustering
- 6:40:07technique on top of the data using one
- 6:40:09of the simple data that we have.
- 6:40:11Jupiter notebook, let me open uh how
- 6:40:15we can use scikit-learn using one of the
- 6:40:17example. So, meanwhile let me show you
- 6:40:19some of the data set also.
- 6:40:22Let me show you a data set that I'm
- 6:40:23going to use as well.
- 6:40:25So, I'm What I'm going to do is I'm
- 6:40:27going to take this movie metadata
- 6:40:28information. Let me open this data set.
- 6:40:32I'm going to take this example of movie
- 6:40:34metadata information where this this
- 6:40:37data set has got a number of
- 6:40:38observations which are present.
- 6:40:41Okay, let me open this. Okay, you can
- 6:40:43see that there are a number of movies
- 6:40:44related information as you can see the
- 6:40:45movie names. So, Avatar, Pirates of the
- 6:40:48Caribbean, Spectre, The Dark Knight,
- 6:40:50Star Wars, etc. etc. John Carter,
- 6:40:52Spider-Man 3, Tangled, or etc. etc. We
- 6:40:55have a lot of movies.
- 6:40:56And about every movie we got a lot of
- 6:40:58information which is present like who is
- 6:40:59the director, who is the actor, what are
- 6:41:01the director Facebook likes, what are
- 6:41:03the actor Facebook likes, likewise we
- 6:41:04got we got a lot of information which is
- 6:41:06present as part of these particular
- 6:41:08every observation.
- 6:41:10So, now what we will do, we will try to
- 6:41:12cluster these data points. You can see a
- 6:41:13lot of observations which are given.
- 6:41:15What is the gross of the movie? What is
- 6:41:16the number of reviewers? What is the
- 6:41:17IMDb rating? What is the so-and-so
- 6:41:19called IMDb score? What is the movie
- 6:41:21span? What is the gross? What is the
- 6:41:22budget? And everything.
- 6:41:24So, now what I'll be doing, I'll be
- 6:41:26reading this data set using one of the
- 6:41:27Pandas library that we have. Okay, so
- 6:41:30I'll be reading this data set. Let me
- 6:41:31open the Jupiter notebook.
- 6:41:33I'll be using something called Pandas.
- 6:41:35So, as you can see, read.pandas.csv.
- 6:41:37I'll be reading this data set where you
- 6:41:38can see that this is the data set that
- 6:41:39I'm able to read. I got a number of
- 6:41:41observations that are present. I can see
- 6:41:43that director Facebook likes and actor
- 6:41:45three Facebook likes.
- 6:41:47So, now what I what is it I want to do
- 6:41:48is instead of building this custom
- 6:41:50technique on top of every observation,
- 6:41:52so what I will do is I will take this
- 6:41:54columns called
- 6:41:55called number of Facebook likes on
- 6:41:57director and number of actor Facebook
- 6:41:58likes versus director Facebook likes.
- 6:42:00So, where I can see that if I want to
- 6:42:02select a director Facebook likes alone,
- 6:42:03I'll be able to choose this director
- 6:42:05Facebook likes alone. You can see for
- 6:42:07every movie you got the number of
- 6:42:08director Facebook likes that are
- 6:42:09present. So, if there are more number of
- 6:42:11Facebook likes, what does it mean?
- 6:42:13The director is famous person.
- 6:42:15Correct?
- 6:42:17Or rather if there are more number of
- 6:42:18famous Facebook likes that the actor has
- 6:42:20got, which means that the person is or
- 6:42:22the actor is very famous person. That's
- 6:42:24what you can understand. So, now you can
- 6:42:25see that I'll try to extract these
- 6:42:27independent component that we talking
- 6:42:29about. I will extract all the records in
- 6:42:31all the director Facebook likes versus
- 6:42:33actor Facebook likes where I'm just
- 6:42:35going to form an object called new data
- 6:42:37by applying some I location as a filter.
- 6:42:39I location I LOC stands for index
- 6:42:41location where I can filter out what are
- 6:42:43the records that I wanted to what are
- 6:42:45the column that I wanted to using this I
- 6:42:47location function.
- 6:42:49Now I got all the so and so called
- 6:42:51number of director Facebook likes versus
- 6:42:52actor Facebook likes which are present
- 6:42:54as part of this. Okay, so this has been
- 6:42:56loaded into an object called new data.
- 6:42:59So now what is that I'll be doing? After
- 6:43:00that I'm importing something called SK
- 6:43:02learn K means cluster. So this algorithm
- 6:43:04is available as part of this
- 6:43:05scikit-learn algorithm. What are the
- 6:43:07number of steps that we have discussed
- 6:43:09you are not supposed to execute all
- 6:43:10these steps manually and of course if
- 6:43:12you want you can also write such a
- 6:43:13program as well, but what is that we'll
- 6:43:15be doing in scikit-learn
- 6:43:17library there is a Python library called
- 6:43:19scikit-learn which has got most of the
- 6:43:20algorithm present and we'll be importing
- 6:43:22this K means clustering algorithm. And
- 6:43:25for this K means clustering I'm
- 6:43:26providing the C is equal to number of
- 6:43:28clusters is equal to five.
- 6:43:29So here if you are if you are aware of
- 6:43:31uh
- 6:43:32object-oriented programming using Python
- 6:43:34you'll be able to correlate. I'm
- 6:43:35importing this K means clustering which
- 6:43:37is implemented as a class here where I'm
- 6:43:39creating an object called K means
- 6:43:40clustering by providing an input called
- 6:43:42N number of scores clusters is equal to
- 6:43:44five which means that what is the
- 6:43:45meaning of five? I want to build a five
- 6:43:47clusters out of this. So where once you
- 6:43:49create an object using K means I'm
- 6:43:51calling this method called fit the
- 6:43:53method by providing new data as my
- 6:43:55independent variables.
- 6:43:57I'm for calling this fit method which
- 6:43:59means fit is a method that we're going
- 6:44:00to invoke what supposed to be the
- 6:44:01process that needs to be executed that's
- 6:44:04going to build my clustering technique
- 6:44:05algorithm by taking number of clusters
- 6:44:07is equal to five.
- 6:44:08So now I'm able to generate my algorithm
- 6:44:10where by looking at my model I'll be
- 6:44:12able to extract what are the centroids
- 6:44:13that I got final centroids because
- 6:44:15initially you'll be taking some random
- 6:44:16centroids, but at the end you'll end up
- 6:44:18with getting a final centroid position
- 6:44:20somewhere fixed to it that is nothing
- 6:44:22but a center point for every cluster
- 6:44:24that we got. These are the final
- 6:44:25centroids that we got.
- 6:44:27Even if I want to print the what is the
- 6:44:29labels which are generated labels is
- 6:44:30nothing but it will extract the outcome.
- 6:44:32You can see that these are the label
- 6:44:33numbers which are added here. Out of
- 6:44:355,000 movies we got the labels which are
- 6:44:37added as an array. But we
- 6:44:39[clears throat] won't be able to see it
- 6:44:39like that. What is it I'm trying to do?
- 6:44:41I'm trying to get all the unique values
- 6:44:43present in this labels with respect to
- 6:44:45two counts. Now you can see that my data
- 6:44:47is now clustered into five different All
- 6:44:49the movies are clustered into five
- 6:44:50different clusters as you can see.
- 6:44:52Cluster number zero has got 4,700 Most
- 6:44:55of the observations are moved into
- 6:44:56cluster number zero.
- 6:44:57104 movies went into observation number
- 6:45:00one. 11 movies are moved into cluster
- 6:45:02cluster number two. And 87 movies are
- 6:45:04moved into cluster number three and 67
- 6:45:06into cluster number four.
- 6:45:08That is how the data properties are
- 6:45:09being distributed and that's how the
- 6:45:11clustering technique has divided the
- 6:45:12data into five clusters.
- 6:45:14Now you can see what is it I'm trying to
- 6:45:15do? I'm trying to put this into a new
- 6:45:17data of cluster which means what are the
- 6:45:19labels that are generated here.
- 6:45:21This is my output column. I'm going to
- 6:45:23create this into as a new column in my
- 6:45:25new data. Where I'm using this Allen
- 6:45:27plot that will print whatever the
- 6:45:29columns that I have called director
- 6:45:31Facebook likes versus actor Facebook
- 6:45:32likes.
- 6:45:33And I'm choosing this data is equal to
- 6:45:35new data that will help me to that will
- 6:45:37help me to identify what column can be
- 6:45:39considered as few so that I'll be
- 6:45:42printing it in a cool warm type chart.
- 6:45:44You can see that palette type is equal
- 6:45:46to cool warm type which will help me to
- 6:45:49identify based on the cluster that you
- 6:45:52have created.
- 6:45:53Now if I print this you can see that
- 6:45:54pretty much the every observation is now
- 6:45:56categorized into individual cluster you
- 6:45:58can see. These are the movies which are
- 6:45:59now graphical representation. This is
- 6:46:01the graphical representation that we are
- 6:46:03using to see how these movies are being
- 6:46:05segregated. You can see these movies are
- 6:46:07nothing but cluster number zero.
- 6:46:09You can see there are a few more movies
- 6:46:11which are extracted over here. This is
- 6:46:12the nothing but based on the color
- 6:46:14indication this is cluster number three.
- 6:46:16And these movies are created as a
- 6:46:17cluster number one.
- 6:46:19Cluster number two. These movies are
- 6:46:21something which are created as cluster
- 6:46:22number one.
- 6:46:24And these movies are created as cluster
- 6:46:25number four.
- 6:46:27Zero.
- 6:46:31One.
- 6:46:32This is nothing but two. Cluster number
- 6:46:34three. Cluster number
- 6:46:35three and then cluster number four like
- 6:46:37this.
- 6:46:38Now you can clearly see that how is the
- 6:46:39cluster came clustering has came up. The
- 6:46:41movies which are made by new people with
- 6:46:44the new directors, new actors are making
- 6:46:46films with new directors.
- 6:46:48Or new directors are making films with
- 6:46:50new actors.
- 6:46:51These are all You will see more number
- 6:46:53of movies will fall into this category
- 6:46:54because you'll end up with getting new
- 6:46:56people into the into the into the film
- 6:46:57industry most of the cases.
- 6:47:00Lot of movies are being made with new
- 6:47:02directors with new actors.
- 6:47:06And you can see that these movies are
- 6:47:07the clusters which are being segregated
- 6:47:09very clearly. The
- 6:47:11famous actors are making films with new
- 6:47:13directors. Very famous actors are making
- 6:47:15films with new directors.
- 6:47:17And these movies are something where
- 6:47:18very famous directors are less I mean
- 6:47:21average paying directors are making
- 6:47:23films with
- 6:47:24some new actors. You can see these
- 6:47:26movies are nothing but very famous
- 6:47:27actors are making directors are making
- 6:47:29films with very
- 6:47:31new actors.
- 6:47:33You can see very famous directors are
- 6:47:35making films with very famous actors.
- 6:47:37Like Liber be making a film with a Tom
- 6:47:40Cruise or something like that.
- 6:47:42So likewise you can clearly see that
- 6:47:44where instead of if you do this activity
- 6:47:46manually it might take a little longer
- 6:47:48time for you to segregate each and every
- 6:47:49component. But within five minutes we
- 6:47:51are able to cluster this activity.
- 6:47:53Right? So that's what the beauty of
- 6:47:54algorithms. You don't need to manually
- 6:47:56do this activity where it can provide
- 6:47:58your data automatically based on the
- 6:47:59properties of the data your algorithm
- 6:48:01itself will will this clustering output
- 6:48:02out of the data.
- 6:48:04All right. So that's how we will be able
- 6:48:05to build a clustering technique on top
- 6:48:07of the given data. So that's what we
- 6:48:09have as part of one of the hands on
- 6:48:10example.
- 6:48:15>> [music]
- 6:48:17>> What is hierarchical clustering?
- 6:48:20So, hierarchical clustering, it is also
- 6:48:21known as HCA or hierarchical cluster
- 6:48:24analysis, and this is a method of
- 6:48:26cluster analysis, as we have seen. So,
- 6:48:28what happens here is that this
- 6:48:30clustering allows us to build the tree
- 6:48:33structure from data similarities, like
- 6:48:35we have built X and Y, and we have
- 6:48:38created trees, and these trees are
- 6:48:40actually called known as dendrogram. So,
- 6:48:42the way you represent a hierarchical
- 6:48:44cluster or a hierarchical clustering is
- 6:48:47through dendrogram. So, we actually drew
- 6:48:49a dendrogram, okay? So, this is how the
- 6:48:52clustering is being formed, and this is
- 6:48:54how the clusters are being made, and
- 6:48:56this is how the relationship among the
- 6:48:58clusters is being shown by a dendrogram.
- 6:49:01So, based on these things, now we will
- 6:49:03go on further to understanding what is
- 6:49:05agglomerative clustering. So, when this
- 6:49:08hierarchical clustering follows a
- 6:49:09bottom-up approach, this is called
- 6:49:12agglomerative clustering. And when this
- 6:49:14is following up a top-down approach,
- 6:49:16this is used in divisive clustering. So,
- 6:49:19now let us understand what is
- 6:49:20agglomerative clustering. So, the types
- 6:49:22of hierarchical clustering are two, that
- 6:49:24is agglomerative and divisive. So, now
- 6:49:26moving on further, what is agglomerative
- 6:49:28clustering? So, agglomerative
- 6:49:30hierarchical clustering, this is also
- 6:49:32known as AGNES, which means
- 6:49:33agglomerative nesting hierarchical
- 6:49:35clustering, and it follows a bottom-up
- 6:49:37approach, which means that clustering or
- 6:49:40clusters, they are formed from the
- 6:49:41bottom and are again clustered till a
- 6:49:44complete single cluster is formed. And
- 6:49:47what happens then is that the clustering
- 6:49:50continues until we obtain a single
- 6:49:52cluster, and we will see how we obtain a
- 6:49:54single cluster, and we also represent
- 6:49:56it. So, individual data points, they are
- 6:49:58clustered based on similarity, and we go
- 6:50:01on clustering until there is only one
- 6:50:03single cluster left. So, let us just
- 6:50:05plot this agglomerative clustering and
- 6:50:07make things really simple for us.
- 6:50:10Now, suppose I have got these data
- 6:50:12points scattered here A to G and these
- 6:50:15data points have to be clustered. So
- 6:50:17another important thing is that now we
- 6:50:19will form clusters. So how clusters are
- 6:50:21formed? So we can see that based on some
- 6:50:23similarity like because of the distance
- 6:50:26nearby distance A and B can be grouped
- 6:50:28together in a single cluster. So I'm
- 6:50:31just doing that.
- 6:50:32C and D I form another cluster because
- 6:50:34they are near. So I just club them and
- 6:50:37again I would just club E and F based on
- 6:50:39their distance and G is separate so I
- 6:50:42will just form a separate cluster. Now
- 6:50:44what happens in agglomerative clustering
- 6:50:47is that I have to plot all these data
- 6:50:49points like A B C right? D
- 6:50:54E F and G. So these are separate
- 6:50:57clusters. The clustering starts from the
- 6:50:59bottom and each data point is treated as
- 6:51:02a single cluster which we will also
- 6:51:04understand with the help of an example
- 6:51:06further but now for simplicity let's
- 6:51:08take A B C D. So then what happens is
- 6:51:11that since A and B are grouped as one
- 6:51:14cluster so this is how I just group them
- 6:51:17all right? And C and D is grouped as one
- 6:51:19cluster this is how I group them. E and
- 6:51:21F are grouped in one single cluster.
- 6:51:24This is how now
- 6:51:26which is not been grouped into any of
- 6:51:28the cluster. So now clustering I said
- 6:51:30that it is it continues until a single
- 6:51:33cluster is left. So now I would have to
- 6:51:36have another level of clustering. That
- 6:51:38means that E F G since G is very close
- 6:51:42to E F I will cluster it in one single
- 6:51:44cluster okay? And since I can see that
- 6:51:47both these A B and C D pairs these
- 6:51:49clusters are again together. So I will
- 6:51:52just cross this line and I will make one
- 6:51:55single cluster of these four points
- 6:51:58right? So what I do is since these two
- 6:52:00are connected I connect them with the
- 6:52:02help of this line figure that is tree
- 6:52:04structured dendrogram and this G E and
- 6:52:08F, they are connected somehow. I connect
- 6:52:09them. All right? Now, what happens is
- 6:52:12that I have got two big clusters, and
- 6:52:14clustering continues until a single
- 6:52:16cluster is obtained. So, in the end, I
- 6:52:19will have to cluster everything into a
- 6:52:22single cluster, and this is how I do
- 6:52:25that. And to join it, I will again join
- 6:52:27this entire graph. So, this is when it
- 6:52:31follows a bottom-up to up approach. This
- 6:52:34is called as agglomerative hierarchical
- 6:52:37clustering.
- 6:52:38Okay? So, now we will go and see an
- 6:52:41example of this hierarchical
- 6:52:43agglomerative clustering, right? Okay.
- 6:52:46So, let us understand what is
- 6:52:48agglomerative clustering with this
- 6:52:50example. Now, we see that here the
- 6:52:53clustering takes place from bottom to
- 6:52:55up, and we have taken an example of
- 6:52:57population, wherein we go on clustering
- 6:52:59until we get population. So, here from
- 6:53:02the bottom, the individual professions
- 6:53:04are being plotted, and we see that let's
- 6:53:07let's take for a convenience the
- 6:53:09left-hand side. And on the left-hand
- 6:53:11side, the red dots, as you see, this is
- 6:53:14individual profession in public sector.
- 6:53:16And on the another side, which we see in
- 6:53:19the brown circles, is the private sector
- 6:53:21employment. So, somehow there's
- 6:53:23similarity between private sector, so uh
- 6:53:25they are being clustered as one single
- 6:53:27cluster, and the private sector as a
- 6:53:29another cluster.
- 6:53:31These again are being clustered into one
- 6:53:33single cluster, and that is employment
- 6:53:35cluster. Whereas on the another side, we
- 6:53:37can see another cluster, which is
- 6:53:39different, and that is unemployed
- 6:53:41section of cluster of people. Now, they
- 6:53:44again share one similarity, and that is
- 6:53:46that they all belong to a single gender,
- 6:53:49that is male. So, everything is being
- 6:53:51clustered into one single cluster, that
- 6:53:53is male. And then again, male and female
- 6:53:56clusters are being clustered together to
- 6:53:58form one single cluster, that is
- 6:54:00population. Similarly, we also divide on
- 6:54:02the right-hand side the individual
- 6:54:05professions of women clubbed into
- 6:54:07private and public sector. And then
- 6:54:09again, we have separate clusters of
- 6:54:11employed and unemployed women. And they
- 6:54:13have been grouped into one single
- 6:54:15cluster and that is women. Again, we
- 6:54:17merge the two big clusters into one
- 6:54:20single cluster that is population. So,
- 6:54:23this is how the clustering is taking
- 6:54:25place. The levels are being increasing
- 6:54:27from bottom to up. So, when we are using
- 6:54:30this bottom-up approach, this is
- 6:54:31agglomerative clustering.
- 6:54:35>> [music]
- 6:54:39>> Now, many of us have visited retail
- 6:54:41shops such as Walmart or Target for our
- 6:54:43household needs. Or let's say that we
- 6:54:45are planning to buy the new iPhone from
- 6:54:47Target.
- 6:54:48What we would typically do is search for
- 6:54:50the model by visiting the mobile section
- 6:54:52of the store and then select the product
- 6:54:54and head towards the billing counter.
- 6:54:57But in today's world, the goal of the
- 6:54:59organization is to increase the revenue.
- 6:55:01Can this be done by just pitching one
- 6:55:03product at a time to the customer? Now,
- 6:55:05the answer to this is clearly no.
- 6:55:07Hence, organization began mining data
- 6:55:10relating to frequently bought items.
- 6:55:13So, market basket analysis is one of the
- 6:55:15key techniques used by large retailers
- 6:55:18to uncover associations between items.
- 6:55:21Now, examples could be the customers who
- 6:55:23purchase bread have a 60% likelihood to
- 6:55:26also purchase jam.
- 6:55:28Customers who purchase laptops are more
- 6:55:30likely to purchase laptop bags as well.
- 6:55:33They try to find out associations
- 6:55:35between different items and products
- 6:55:37that can be sold together, which gives
- 6:55:40assisting in the right product
- 6:55:41placement.
- 6:55:43Typically, it figures out what products
- 6:55:45are being bought together and
- 6:55:47organizations can place products in a
- 6:55:49similar manner.
- 6:55:50For example, people who buy bread also
- 6:55:53tend to buy butter, right?
- 6:55:55And the marketing team at retail stores
- 6:55:57should target customers who buy bread
- 6:56:00and butter and provide an offer to them
- 6:56:02so that they buy a third item, suppose
- 6:56:05eggs.
- 6:56:06So, if a customer buys bread and butter
- 6:56:08and sees a discount offer on eggs, he
- 6:56:10will be encouraged to spend more and buy
- 6:56:12the eggs. And this is what market basket
- 6:56:15analysis is all about. This is what we
- 6:56:17are going to talk about in this session,
- 6:56:19which is association rule mining and the
- 6:56:21a priori algorithm.
- 6:56:23Now, association rule can be thought of
- 6:56:25as an if-then relationship. Just to
- 6:56:28elaborate on that, we have come up with
- 6:56:31a rule, suppose if an item A is being
- 6:56:33bought by the customer, then the chances
- 6:56:35of item B being picked by the customer
- 6:56:38too under the same transaction ID is
- 6:56:40found out. You need to understand here
- 6:56:43that it's not a causality, rather it's a
- 6:56:46co-occurrence pattern that comes to the
- 6:56:48force.
- 6:56:49Now, there are two elements to this
- 6:56:50rule. First is the if and second is the
- 6:56:53then.
- 6:56:54Now, if is also known as antecedent.
- 6:56:57This is an item or a group of items that
- 6:57:00are typically found in the item set. And
- 6:57:02the later one is called the consequent.
- 6:57:06This comes along as an item with an
- 6:57:08antecedent group or the group of
- 6:57:10antecedents are purchased.
- 6:57:13Now, if you look at the image here A
- 6:57:14arrow B, it means that if a person buys
- 6:57:17an item A, then he will also buy an item
- 6:57:19B. Or he will most probably buy an item
- 6:57:21B.
- 6:57:22Now, the simple example that I gave you
- 6:57:24about the bread and butter and the eggs
- 6:57:27is just a small example. But what if you
- 6:57:29have thousands and thousands of items?
- 6:57:32If you go to any professional data
- 6:57:34scientist with that data, you can just
- 6:57:36imagine how much of profit you can make
- 6:57:39if the data scientist provides you with
- 6:57:41the right examples and the right
- 6:57:42placement of the items which you can do.
- 6:57:45And you can get a lot of insights. That
- 6:57:47is why association rule mining is a very
- 6:57:50good algorithm which helps the business
- 6:57:52make profit. So, let's see how this
- 6:57:55algorithm works.
- 6:57:56So, association rule mining is all about
- 6:57:58building the rules. And we have just
- 6:58:01seen one rule that if you buy A, then
- 6:58:05there's a slight possibility or there's
- 6:58:07a chance that you might buy B also. This
- 6:58:10type of relationship in which we can
- 6:58:12find the relationship between these two
- 6:58:14items is known as single cardinality.
- 6:58:17But what if the customer who bought A
- 6:58:20and B also wants to buy C?
- 6:58:23Or if a customer who bought A, B, and C
- 6:58:25also wants to buy D? Then in these
- 6:58:28cases, the cardinality usually increases
- 6:58:30and we can have a lot of combination
- 6:58:33around these data.
- 6:58:35And if you have around 10,000 or more
- 6:58:37than 10,000 data or items, just imagine
- 6:58:40how many rules you're going to create
- 6:58:42for each product. That is why
- 6:58:44association rule mining has such
- 6:58:47measures so that we do not end up
- 6:58:49creating tens of thousands of rules.
- 6:58:52Now, that is where the a priori
- 6:58:54algorithm comes in. But before we get
- 6:58:56into the a priori algorithm, let's
- 6:58:58understand what's the maths behind it.
- 6:59:01Now, there are three types of matrices
- 6:59:03which help to measure the association.
- 6:59:06We have support, confidence, and lift.
- 6:59:09So, support is the frequency of item A
- 6:59:11or the combination of item A or B.
- 6:59:14It's basically the frequency of the
- 6:59:16items which we have bought and what are
- 6:59:18the combination of the frequency of the
- 6:59:20item we have bought. So, with this, what
- 6:59:22we can do is filter out the items which
- 6:59:25have been bought less frequently.
- 6:59:28This is one of the measures which is
- 6:59:29support.
- 6:59:30Now, what confidence tells us?
- 6:59:32So, confidence gives us how often the
- 6:59:34items A and B occur together given the
- 6:59:37number of times A occur.
- 6:59:39Now, this also helps us solve a lot of
- 6:59:41other problems because if somebody is
- 6:59:43buying A and B together and not buying
- 6:59:45C, we can just rule out C at that point
- 6:59:47of time.
- 6:59:48So, this solves another problem is that
- 6:59:51we obviously do not need to analyze the
- 6:59:53products which people just buy barely.
- 6:59:56So, what we can do is according to the
- 6:59:58sales, we can define our minimum support
- 7:00:01and confidence. And when you have set
- 7:00:03these values, we can put these values in
- 7:00:05the algorithm and we can filter out the
- 7:00:08data and we can create different rules.
- 7:00:11And suppose even after filtering, you
- 7:00:13have like 5,000 rules. And for every
- 7:00:16item, we create these 5,000 rules. So,
- 7:00:19that's practically impossible. So, for
- 7:00:22that, we need the third calculation
- 7:00:24which is the lift. So, lift is basically
- 7:00:26the strength of any rule.
- 7:00:28Now, let's have a look at the
- 7:00:29denominator of the formula given here.
- 7:00:32And if you see here, we have the
- 7:00:34independent support values of A and B.
- 7:00:37So, this gives us the independent
- 7:00:39occurrence probability of A and B. And
- 7:00:42obviously, there's a lot of difference
- 7:00:44between this random occurrence and
- 7:00:46association. And if the denominator of
- 7:00:49the lift is more, what it means is that
- 7:00:53the occurrence of randomness is more
- 7:00:55rather than the occurrence because of
- 7:00:58any association.
- 7:00:59So, lift is the final verdict where we
- 7:01:01know whether we have to spend time on
- 7:01:03this particular rule what we have got
- 7:01:06here or not. Now, let's have a look at a
- 7:01:08simple example of association rule
- 7:01:10mining.
- 7:01:11So, suppose we have a set of items A, B,
- 7:01:14C, D, and E and a set of transactions
- 7:01:17T1, T2, T3, T4, and T5.
- 7:01:20And as you can see here, we have the
- 7:01:21transactions T1 in which we have ABC, T2
- 7:01:24ACD, T3 BCD, T4 ADE, and T5 BCE.
- 7:01:30Now, what we generally do is create some
- 7:01:33rules or association rules such as A
- 7:01:36gives D or C gives A, A gives C, B and C
- 7:01:40gives A. What this basically means is
- 7:01:43that if a person buys A, then he's most
- 7:01:45likely to buy D. And if a person buys C,
- 7:01:48then he's most likely to buy A. And if
- 7:01:50you have a look at the last one, if a
- 7:01:51person buys B and C, he's most likely to
- 7:01:54buy the item A as well.
- 7:01:56Now, if we calculate the support,
- 7:01:58confidence, and lift using these rules,
- 7:02:00as you can see here in the table, we
- 7:02:02have the rule and the support,
- 7:02:04confidence, and the lift values.
- 7:02:06Now, let's discuss about a priori.
- 7:02:09So, a priori algorithm uses the frequent
- 7:02:12item sets to generate the association
- 7:02:14rule.
- 7:02:15And it is based on the concept that a
- 7:02:17subset of a frequent item set must also
- 7:02:20be a frequent item set itself.
- 7:02:23Now, this raises the question, what
- 7:02:24exactly is a frequent item set?
- 7:02:26So, a frequent item set is an item set
- 7:02:29whose support value is greater than the
- 7:02:30threshold value. Now, just now we
- 7:02:32discussed that the marketing team,
- 7:02:34according to the sales, have a minimum
- 7:02:36threshold value for the confidence as
- 7:02:39well as the support.
- 7:02:41So, frequent item set is that item set
- 7:02:43whose support value is greater than the
- 7:02:44threshold value already specified.
- 7:02:47Now, example, if A and B is a frequent
- 7:02:49item set, then A and B should also be
- 7:02:52frequent item sets individually.
- 7:02:55Now, let's consider the following
- 7:02:56transaction to make the things a little
- 7:02:59easier. Suppose we have transactions 1 2
- 7:03:023 4 5 and these items are there.
- 7:03:04So, T1 has 1 3 and 4, T2 has 2 3 and 5,
- 7:03:08T3 has 1 2 3 5, T4 2 5, and T5 1 3 and
- 7:03:125. Now, the first step is to build a
- 7:03:15list of item sets of size one by using
- 7:03:18this transactional data. And one thing
- 7:03:20to note here is that the minimum support
- 7:03:23count, which is given here, is two.
- 7:03:25Let's suppose it's two.
- 7:03:27So, the first step is to create item
- 7:03:29sets of size one and calculate their
- 7:03:31support values.
- 7:03:32So, as you can see here, we have the
- 7:03:33table C1 in which we have the item sets
- 7:03:361 2 3 4 5, and the support values.
- 7:03:39If you remember the formula of support,
- 7:03:41it was frequency divided by the total
- 7:03:43number of occurrence.
- 7:03:45So, as you can see here, for the item
- 7:03:46set 1, the support is three.
- 7:03:49As you can see here, the item set 1
- 7:03:51appears in T1, T3, and T5.
- 7:03:54So, as you can see, its frequency is 1,
- 7:03:562, and 3.
- 7:03:57Now, as you can see here, the item set 4
- 7:04:00has a support of 1 as it occurs only
- 7:04:02once in transaction 1.
- 7:04:04But, the minimum support value is two.
- 7:04:06That's why it's going to be eliminated.
- 7:04:09So, we have the final table, which is
- 7:04:10the table F1, in which we have the item
- 7:04:13sets 1, 2, 3, and 5, and we have the
- 7:04:15support values 3, 3, 4, and 4.
- 7:04:19Now, the next step is to create item
- 7:04:20sets of size two and calculate the
- 7:04:22support values.
- 7:04:24Now, all the combination of the item
- 7:04:25sets in the F1, which is the final table
- 7:04:29in which you discarded the four, are
- 7:04:31going to be used for this iteration. So,
- 7:04:33we get the table C2. So, as you can see
- 7:04:35here, we have 1, 2, 1, 3, 1, 5, 2, 3, 2,
- 7:04:395, and 3, 5.
- 7:04:40Now, if you calculate the support here
- 7:04:42again, we can see that the item set 1, 2
- 7:04:45has a support of 1, which is again less
- 7:04:48than the specified threshold. So, we're
- 7:04:51going to discard that.
- 7:04:52So, if we have a look at the table F2,
- 7:04:55we have 1, 3, 1, 5, 2, 3, 2, 5, and 3,
- 7:04:595.
- 7:05:00Again, we're going to move forward and
- 7:05:02create the item set of size three and
- 7:05:05calculate the support values.
- 7:05:07Now, all the combinations are going to
- 7:05:08be used from the item set F2 for this
- 7:05:11particular iterations.
- 7:05:13Now, before calculating support values,
- 7:05:15let's perform pruning on the data set.
- 7:05:18Now, what is pruning? Now, after the
- 7:05:20combinations are being made, we divide
- 7:05:21C3 item sets to check if there is
- 7:05:23another subset whose support is less
- 7:05:26than the minimum support value.
- 7:05:28That is what frequent item set means.
- 7:05:31So, if you have a look here, the item
- 7:05:33sets we have is 1 2 3, 1 2, 1 3, 2 3 for
- 7:05:38the first one. Because as you can see
- 7:05:40here, if we have a look at the subsets
- 7:05:42of 1 2 3, we have 1 {comma} 2 as well.
- 7:05:46So, we are going to discard this whole
- 7:05:48item set.
- 7:05:49Same goes for the second one. We have 1
- 7:05:512 5. We have 1 2 in that, which was
- 7:05:53discarded in the previous set or the
- 7:05:55previous step. That's why we're going to
- 7:05:57discard that also.
- 7:05:59Which leaves us with only two factors,
- 7:06:01which is 1 3 5 item set and the 2 3 5.
- 7:06:05And the support for this is two and two
- 7:06:07as well.
- 7:06:08Now, if we create the table C4 using
- 7:06:12four elements, we're going to have only
- 7:06:14one item set, which is 1 2 3 and 5. And
- 7:06:18if we have a look at the table here, the
- 7:06:20transaction table, 1 2 3 and 5 appears
- 7:06:23only once. So, the support is one.
- 7:06:25And since C4, the support of the whole
- 7:06:28table C4 is less than two, so we're
- 7:06:30going to stop here and return to the
- 7:06:32previous item set that is three.
- 7:06:34So, the frequent item sets are 1 3 5 and
- 7:06:372 3 5.
- 7:06:38Now, let's assume our minimum confidence
- 7:06:40value is 60%.
- 7:06:42For that, we're going to generate all
- 7:06:43the non-empty subsets for each frequent
- 7:06:46item sets.
- 7:06:47Now, for I equals 1 {comma} 3 {comma} 5,
- 7:06:50which is the item set, we get the subset
- 7:06:521 3, 1 5, 3 5, 1, 3, and 5.
- 7:06:58Similarly, for 2 3 5, we get 2 3, 2 5, 3
- 7:07:015, 2, 3, and 5.
- 7:07:04Now, this rule states that for every
- 7:07:06subset S of I, the output of the rule
- 7:07:09gives something like S gives I to S.
- 7:07:13That implies S recommends I of S.
- 7:07:16And this is only possible if the support
- 7:07:18of I divided by the support of S is
- 7:07:20greater than equal to the minimum
- 7:07:22confidence value.
- 7:07:24Now, applying these rules to the item
- 7:07:26set of F3, we get rule one, which is 1,3
- 7:07:30gives 1,3,5
- 7:07:32and 1,3.
- 7:07:33It means one and three gives five.
- 7:07:36So, the confidence is equal to the
- 7:07:39support of 1,3,5 / support of 1,3, that
- 7:07:44equals 2 / 3, which is 66% and which is
- 7:07:47greater than the 60%.
- 7:07:49So, the rule one is selected.
- 7:07:51Now, if we come to rule two, which is
- 7:07:531,5, it gives 1,3,5
- 7:07:56and 1,5. It means if we have one and
- 7:07:59five, it implies we also going to have
- 7:08:02three. Now, if we calculate the
- 7:08:03confidence of this one, we're going to
- 7:08:05have support 1,3,5 / support 1,5, which
- 7:08:08gives us 100%, which means rule two is
- 7:08:11selected as well. But again, if you have
- 7:08:13a look at rule five and rule six over
- 7:08:15here, similarly, if it select three
- 7:08:18gives 1,3,5 and three, it means if you
- 7:08:21have three, we also get one and five.
- 7:08:23So, the confidence for this comes at
- 7:08:2550%, which is less than the given 60%
- 7:08:29target, so we're going to reject this
- 7:08:31rule. And same goes for the rule number
- 7:08:33six.
- 7:08:34Now, one thing to keep in mind here is
- 7:08:37that although the rule one and rule five
- 7:08:39look a lot similar, they are not.
- 7:08:42So, it really depends what's on the
- 7:08:44left-hand side of the arrow and what's
- 7:08:45on the right-hand side of the arrow.
- 7:08:47It's the if-then possibility.
- 7:08:49I'm sure you guys can understand what
- 7:08:52exactly these rules are and how to
- 7:08:54proceed with the rules.
- 7:08:55So, let's see how we can implement the
- 7:08:57same in Python, right?
- 7:09:00So, for that, what I'm going to do is
- 7:09:01create a new Python file and
- 7:09:06I'm going to use the Jupyter Notebook.
- 7:09:08You're free to use any sort of IDE.
- 7:09:11I'm going to name it as Apriori.
- 7:09:14So, the first thing what we're going to
- 7:09:16do is we'll be using the online
- 7:09:19transactional data of a retail store for
- 7:09:21generating association rules. So,
- 7:09:23firstly, what we need to do is get the
- 7:09:25pandas and mlxtend libraries imported
- 7:09:27and read the file.
- 7:09:30So, as you can see here, we are using
- 7:09:32the online retail.xlsx
- 7:09:34format file.
- 7:09:36And from mlxtend, we're going to import
- 7:09:38a priori and association rules. It all
- 7:09:40comes under mlxtend.
- 7:09:43So, as you can see here, we have the
- 7:09:45invoice, the stock code, the
- 7:09:47description, the quantity, the invoice
- 7:09:49data, unit price, customer ID, and the
- 7:09:52country.
- 7:09:53Now, next in this step, what we're going
- 7:09:54to do is do data cleanup, which includes
- 7:09:57removing the spaces from some of the
- 7:09:59descriptions, and drop the rows that do
- 7:10:01not have invoice numbers, and remove the
- 7:10:03credit card transactions, because that
- 7:10:06is of no use to us.
- 7:10:13So, as you can see here, the output in
- 7:10:15which we have like 532,000
- 7:10:19rows with eight columns.
- 7:10:21So, after the cleanup, we need to
- 7:10:23consolidate the items into one
- 7:10:25transaction per row with each product.
- 7:10:27For the sake of keeping the data set
- 7:10:29small, we are only looking at the sales
- 7:10:31for France.
- 7:10:36So, as you can see here, we have
- 7:10:38excluded all the other sales. We are
- 7:10:40just looking at the sales for France.
- 7:10:42Now, there are a lot of zeros in the
- 7:10:44data, but we also need to make sure any
- 7:10:46positive values are converted to one,
- 7:10:48and anything less than zero is set to
- 7:10:49zero.
- 7:10:52So, as you can see here, we have still
- 7:10:53392 rows.
- 7:10:56We're going to encode it and see.
- 7:10:59Check again.
- 7:11:00Now that you have structured the data
- 7:11:02properly, in this step, what we're going
- 7:11:03to do is generate frequent item sets
- 7:11:05that have support at least 7%.
- 7:11:08Now, this number is chosen so that you
- 7:11:10can get close enough and generate the
- 7:11:12rules with the corresponding support,
- 7:11:13confidence, and lift.
- 7:11:19So, guys, as you can see here, the
- 7:11:20minimum support is 0.7. And what if we
- 7:11:23add another constraint on the rules,
- 7:11:25such as the lift is greater than six and
- 7:11:28the confidence is greater than 0.8?
- 7:11:31So as you can see here, we have the
- 7:11:33left-hand side and the right-hand side
- 7:11:34of the association rule, which is the
- 7:11:36antecedent and the consequence.
- 7:11:39We have the support, we have the
- 7:11:40confidence, the lift, the leverage, and
- 7:11:42the conviction.
- 7:11:43So guys, that's it for this session.
- 7:11:45That is how you create association rules
- 7:11:48using the Apriori algorithm, which helps
- 7:11:50a lot in the marketing business.
- 7:11:53It runs on the principle of market
- 7:11:55basket analysis, which is exactly what
- 7:11:57big companies like Walmart, you have
- 7:11:59Reliance, and Target too. Even IKEA does
- 7:12:03it.
- 7:12:04>> [music]
- 7:12:10>> But the first thing that we need to
- 7:12:11focus on is why artificial intelligence.
- 7:12:13Why do we need artificial intelligence?
- 7:12:16This with an example. So nowadays, if
- 7:12:18you have noticed, if your car exceeds
- 7:12:20the speed limit, so you'll get a letter
- 7:12:22or basically a challan at your home. How
- 7:12:24do you think that happens? Do you think
- 7:12:26that there's a person who is sitting in
- 7:12:27a chair and actually noting down all the
- 7:12:29number plates that crosses the speed
- 7:12:31limit? Well, that is not possible
- 7:12:33because there might be millions of cars
- 7:12:35that pass through that road. And at
- 7:12:36once, there might be many cars that will
- 7:12:38be passing through that road. So for a
- 7:12:40human being to actually do this task is
- 7:12:42next to impossible. Now, let us see
- 7:12:44another approach to this particular
- 7:12:46problem. So what we can do, we can
- 7:12:48actually make use of cameras that will
- 7:12:50click the picture of the car that
- 7:12:52exceeds the speed limit. And then we
- 7:12:54could convert that picture into a text.
- 7:12:56For example, we have a UK plate.
- 7:12:59So in this way, the human error, the
- 7:13:01risk of human error has been reduced.
- 7:13:03And at the same time, machines, they
- 7:13:05never get tired. So because of that, you
- 7:13:07can capture all the images of cars that
- 7:13:09actually crosses the speed limit.
- 7:13:11Similarly, you can think of uh many
- 7:13:13other examples as well. It is used in
- 7:13:15order to recognize a sign that is in
- 7:13:16banks if you want to authenticate
- 7:13:18whether that person is the bank customer
- 7:13:20or not. Yeah, apart from that, it is
- 7:13:22even used for self-driving cars as well.
- 7:13:24So, in US around 30,000 people die every
- 7:13:27year because of road accidents. So, that
- 7:13:29can be completely removed if we use the
- 7:13:31self-driving cars, which is based on the
- 7:13:33concept of artificial intelligence. And
- 7:13:35let me tell you guys, you might find it
- 7:13:36very fascinating that people in MIT are
- 7:13:39using artificial intelligence in order
- 7:13:41to predict the future. So, you can
- 7:13:43imagine why we need artificial
- 7:13:44intelligence. If you have any questions,
- 7:13:46any doubts, you can ask me. It is even
- 7:13:48used in places where humans can't reach.
- 7:13:50For example, uh deep oceans or
- 7:13:52navigation in Mars. So, in those places
- 7:13:55we need machines which are smart enough
- 7:13:57to carry out tasks.
- 7:13:58So, this is why we need artificial
- 7:14:00intelligence. Let us move forward and
- 7:14:01understand what exactly is artificial
- 7:14:03intelligence.
- 7:14:05Now, artificial intelligence, I know the
- 7:14:06word sounds pretty complex and there are
- 7:14:08a lot of Hollywood movies that are based
- 7:14:10on artificial intelligence. If you have
- 7:14:12seen Terminator or Matrix, all these
- 7:14:14movies are based on artificial
- 7:14:15intelligence. But, you don't need to
- 7:14:17worry about it because till now we
- 7:14:18haven't reached that level as they have
- 7:14:20shown in movies like Terminator.
- 7:14:22But, yeah, the concept is pretty
- 7:14:23similar. So, basically, we want systems
- 7:14:26and softwares in such a way that they
- 7:14:28can imitate the human behavior.
- 7:14:30Now, what happens in artificial
- 7:14:31intelligence? Artificial intelligence is
- 7:14:34accomplished by studying how human brain
- 7:14:36thinks and how human brain learns,
- 7:14:38decide, and work while trying to solve a
- 7:14:41problem. And then we use outcome of this
- 7:14:43study as the basis of development of
- 7:14:45intelligent software and systems.
- 7:14:48So, our major goal is to have systems or
- 7:14:50softwares that can imitate the human
- 7:14:53behavior. The way they think, the way
- 7:14:55they decide, the way they solve a
- 7:14:57problem. So, in that similar fashion, we
- 7:14:59want our machines to do that. So, this
- 7:15:02is basically artificial intelligence in
- 7:15:03a nutshell.
- 7:15:05Let us move forward and look at various
- 7:15:06applications of artificial intelligence.
- 7:15:09So, this slide basically talks about the
- 7:15:11application of artificial intelligence.
- 7:15:13Now, I've listed only three of them, but
- 7:15:15there are millions of applications. For
- 7:15:17example, it is used in speech
- 7:15:18recognition. So, whenever you search
- 7:15:20something on Google, so you can just
- 7:15:21tell Google and it'll search it for you.
- 7:15:23Similarly, it is used for understanding
- 7:15:24natural language as well as for image
- 7:15:26recognition as well. And there are many,
- 7:15:28many other applications in which
- 7:15:30artificial intelligence finds its use.
- 7:15:32For example, it can be used in
- 7:15:34self-driving cars. It can be used in
- 7:15:36Siri for recommending some products. And
- 7:15:38even when you go to websites like
- 7:15:39YouTube or Pandora, YouTube knows which
- 7:15:41video you want to next. Pandora knows
- 7:15:43which song you want to listen to. How do
- 7:15:45you think this happens? It happens all
- 7:15:47because of artificial intelligence. So,
- 7:15:49all of these are a few examples of
- 7:15:51artificial intelligence, but nowadays,
- 7:15:53it is used almost everywhere, guys.
- 7:15:55Trust me on that. Now, let us move
- 7:15:57forward and understand how to achieve
- 7:15:59artificial intelligence.
- 7:16:01Now, in order to achieve artificial
- 7:16:02intelligence, there were a few
- 7:16:04technologies that First came a machine
- 7:16:06learning. Now, there were certain
- 7:16:08limitations of machine learning. In
- 7:16:10order to overcome those limitations,
- 7:16:12came a deep learning. Now, let me tell
- 7:16:14you guys, the concept of artificial
- 7:16:16intelligence is not new. It was first
- 7:16:18coined in 1956,
- 7:16:20but it was just a theoretical concept.
- 7:16:22Then in '80s and '90s, we were talking
- 7:16:24about neural networks. But since we
- 7:16:26didn't have enough computational power,
- 7:16:29so we couldn't utilize it properly. But
- 7:16:31in late '90s and 2000s, we started using
- 7:16:33neural networks for machine learning.
- 7:16:35Then in 2006, the term deep learning was
- 7:16:38coined for the first time that overcame
- 7:16:40the limitations of machine learning. And
- 7:16:42from 2010, deep learning was used
- 7:16:44commercially as well. So, this was just
- 7:16:47a small history about artificial
- 7:16:48intelligence, machine learning, and deep
- 7:16:50learning. Now, in order to understand
- 7:16:52this deep learning, we need to first
- 7:16:54look at machine learning and what were
- 7:16:55the biggest limitations of machine
- 7:16:57learning that led to the evolution of
- 7:16:58deep learning. So, we'll move forward
- 7:17:00and understand what exactly is a machine
- 7:17:03learning.
- 7:17:04Now, what is machine learning? So,
- 7:17:05machine learning is nothing but a type
- 7:17:07of artificial intelligence or you can
- 7:17:09say a subset of artificial intelligence.
- 7:17:11And it provides computers with the
- 7:17:13ability to learn without being
- 7:17:15explicitly programmed. So, you don't
- 7:17:16need to hardcode your machine for that.
- 7:17:19Let us understand this with an example.
- 7:17:21So, we have a problem statement in which
- 7:17:23whenever you give a certain input, we
- 7:17:25need to determine the species of the
- 7:17:26flower. And what is that input? That
- 7:17:28input will be sepal length, sepal width,
- 7:17:31petal length, and petal width. So,
- 7:17:33whenever we get these four parameters or
- 7:17:35these four variables, our machine should
- 7:17:37be able to predict what sort of a flower
- 7:17:39it is. Now, how do you think that will
- 7:17:41happen? First, what we need to do, we
- 7:17:44need to train our machine on the basis
- 7:17:46of the data that we have. So, in this
- 7:17:49data, we have sepal length, sepal width,
- 7:17:51petal length, and petal width, and we
- 7:17:52have species. So, our machine will learn
- 7:17:55from this data. It'll determine what
- 7:17:57should be the length and width of the
- 7:17:59sepal and petal in order to classify it
- 7:18:01as setosa or versicolor or other species
- 7:18:04of flowers as well. Now, what happens
- 7:18:06next? So, you have trained your data.
- 7:18:08So, you have trained your machine from
- 7:18:09the data set. Then, what happens?
- 7:18:11Whenever you give a new input to this
- 7:18:13particular machine, it'll predict the
- 7:18:15species of the flower.
- 7:18:17So, this is how machine learning works.
- 7:18:18It is nothing but machine learning in a
- 7:18:19nutshell.
- 7:18:21So, basically, I'll just revise it once
- 7:18:23more.
- 7:18:24So, you have a data set. So, you split
- 7:18:26that data into training and testing
- 7:18:28data. So, what happens with the help of
- 7:18:30training data? You train your particular
- 7:18:32machine. And after that, you test it in
- 7:18:35order to determine the accuracy. And
- 7:18:37once it is done, whenever you give the
- 7:18:39new input, it'll predict the outcome or
- 7:18:41the desired outcome. So, this is how
- 7:18:43machine learning works, guys. Let us
- 7:18:45move forward and understand various
- 7:18:47types of machine learning.
- 7:18:49So, the first type is called supervised
- 7:18:51learning. Now, in supervised learning,
- 7:18:54what happens? You have input variables X
- 7:18:56and an output variable Y.
- 7:18:58And you can use an algorithm to learn
- 7:19:00mapping function from the input to the
- 7:19:02output. Now, let me simplify it for you.
- 7:19:04So, what happens in supervised learning,
- 7:19:06the data that you have already contain
- 7:19:09the classification. Now, let me talk
- 7:19:11about the previous example itself. So,
- 7:19:13from our data set, we knew that if we
- 7:19:15have this width, this length of our
- 7:19:17sepal and petal,
- 7:19:19so that will be the species of flower.
- 7:19:21So, the classifications are already
- 7:19:23defined. So, that will be under
- 7:19:25supervised learning. Now, let me tell
- 7:19:27you how it actually works.
- 7:19:28So, you have data. You divide that data
- 7:19:30into training data as well as test data.
- 7:19:33So, on the basis of this training data,
- 7:19:35you train your machine. And after that,
- 7:19:37you create a model. So, as you can see
- 7:19:39that this phase is called training
- 7:19:40phase. And after that, you create a
- 7:19:42model. Now, in order to check this model
- 7:19:45to get the accuracy, you have test data.
- 7:19:47So, you'll pass that test data and
- 7:19:49you'll see the accuracy. That is nothing
- 7:19:51but the actual output minus the output
- 7:19:54that is present in the test data. So,
- 7:19:56with that, you can get the accuracy. So,
- 7:19:58this is nothing but uh supervised
- 7:20:00learning. And if you have any questions,
- 7:20:02you can ask me right now.
- 7:20:04So, we'll move forward and understand
- 7:20:06unsupervised learning.
- 7:20:07Now, in unsupervised learning, unlike
- 7:20:09supervised learning, you don't have any
- 7:20:11predefined classes. So, what happens,
- 7:20:12you have data. So, on the basis of that
- 7:20:15data, you try to create your own class.
- 7:20:18You try to make sure that whatever class
- 7:20:20you create has high intra-class
- 7:20:21similarities and have a low inter-class
- 7:20:24similarities. That means, if I've
- 7:20:26created two class like this, class one
- 7:20:28and class two, so the elements of this
- 7:20:31particular class should have high
- 7:20:33similarity, but at the same time, it
- 7:20:35should have low similarity with the
- 7:20:37elements of class two.
- 7:20:39So, you can think of examples as well of
- 7:20:40unsupervised learning. For example, if I
- 7:20:43have a data about my customers. So, if I
- 7:20:46have a website and there are millions of
- 7:20:47visitors on my website, and I want to
- 7:20:49make sure that I group people on various
- 7:20:52criteria. For example, I can group
- 7:20:54people on the basis of willingness to
- 7:20:56purchase a product that is there on my
- 7:20:57website or where they are coming from,
- 7:21:00what is the source, all those things.
- 7:21:01So, I want to group my customers and I
- 7:21:04want to make sure that I have certain
- 7:21:05high priority customers and I have low
- 7:21:07priority customers and I have medium
- 7:21:09priority customers. So, with the help of
- 7:21:10unsupervised learning, I can actually do
- 7:21:13that. I can make a certain classes of
- 7:21:14people on whom I should focus more on as
- 7:21:17compared to the other class. So, this
- 7:21:18was just an example, guys. You can use
- 7:21:20it in various other fields as well. So,
- 7:21:22in marketing, this is how you can use
- 7:21:24unsupervised learning.
- 7:21:25So, this brings us to our next uh type
- 7:21:27of machine learning, which is called a
- 7:21:29reinforcement learning. Now, this is
- 7:21:31reinforcement learning, guys. Now, what
- 7:21:33happens in reinforcement learning, the
- 7:21:34machine learns by interacting with space
- 7:21:36or an environment. So, it learns with
- 7:21:38its experience, with its past
- 7:21:40experience, and also by new choice
- 7:21:42exploration. Now, I'll take the analogy
- 7:21:44of dogs. So, if you have any dog or a
- 7:21:46pet at your home, so if you have trained
- 7:21:48your dog in order to get the newspaper,
- 7:21:50and if it gets it, then you reward it
- 7:21:51with some chocolate or things that the
- 7:21:54dog likes, right? So, the dog will know
- 7:21:56whatever he has done, he's actually
- 7:21:57rewarded for that. So, it'll continue
- 7:21:59doing that. But, apart from that, if he
- 7:22:01does something else, if instead of the
- 7:22:02newspaper, he brings something else. So,
- 7:22:04what do you do? You might even punish
- 7:22:05it. So, because of that, the dog will
- 7:22:07come to know that it has to get a
- 7:22:09newspaper every morning.
- 7:22:10Now, the same example is there in front
- 7:22:12of your screen. So, you have this
- 7:22:14machine. So, it has two choices, either
- 7:22:16to touch the fire or touch the water.
- 7:22:19Now, first what it does, it goes on and
- 7:22:21touch the fire. So, because of that, it
- 7:22:23gets some burning sensation. Now, it has
- 7:22:24only other option, that is to touch the
- 7:22:27water. So, when it touch the water, it
- 7:22:29gets some reward. So, because of that,
- 7:22:31it'll understand that it does not have
- 7:22:32to touch fire ever again.
- 7:22:35Now, there's a diagram that is there in
- 7:22:36front of your screen. So, what happens
- 7:22:38you have an agent, all right? That agent
- 7:22:40performs some action. And on the basis
- 7:22:42of that action, it'll be exposed to some
- 7:22:44sort of an environment. Now, if that
- 7:22:46action is correct, then it'll be
- 7:22:48rewarded with that. But if it is not,
- 7:22:50then it will change its choice, and it
- 7:22:52will again perform some action. So, this
- 7:22:54process will keep on repeating. So, this
- 7:22:56is how reinforcement learning works.
- 7:22:59So, let us move forward and understand
- 7:23:01when we have machine learning, why do we
- 7:23:03need deep learning? That is, we'll look
- 7:23:05at various uh limitations of machine
- 7:23:07learning.
- 7:23:08Now, the first limitation is high
- 7:23:09dimensionality of the data. Now, the
- 7:23:11data that is now generated is huge in
- 7:23:14size. So, we have a very large number of
- 7:23:16inputs and outputs. So, due to that,
- 7:23:18machine learning algorithms fail. So,
- 7:23:20they cannot deal with high
- 7:23:21dimensionality of data, or you can say
- 7:23:22data with large number of inputs and
- 7:23:25outputs.
- 7:23:26Now, there's another problem as well, in
- 7:23:27which it is unable to solve the crucial
- 7:23:29AI problems, which can be natural
- 7:23:31language processing, image recognition,
- 7:23:32and uh things like that.
- 7:23:34Now, one of the biggest challenges with
- 7:23:35machine learning models is feature
- 7:23:37extraction. Now, let me tell you what
- 7:23:39are features. So, in statistics, we
- 7:23:40consider features as variables, but when
- 7:23:42we talk about artificial intelligence,
- 7:23:44these variables are nothing but the
- 7:23:45features.
- 7:23:46Now, what happens because of that? The
- 7:23:48complex problems such as object
- 7:23:50recognition or handwriting recognition
- 7:23:52becomes a huge challenge for machine
- 7:23:53learning algorithms to solve. Now, let
- 7:23:55me give you an example of this uh
- 7:23:56feature extraction. Suppose, if you want
- 7:23:58to predict that whether there'll be a
- 7:24:00match today or not. So, it depends on a
- 7:24:02various features. It depends on the
- 7:24:03whether the weather is sunny, whether it
- 7:24:05is windy, all those things. So, we have
- 7:24:07provided all those features in our data
- 7:24:09set. But, we have forgot one particular
- 7:24:11feature that is humidity. And now, our
- 7:24:13machine learning models are not that
- 7:24:14efficient that they will automatically
- 7:24:16generate that particular feature. So,
- 7:24:18this is one huge problem, or you can say
- 7:24:20limitation, with machine learning. Now,
- 7:24:22obviously, we have limitation, and it
- 7:24:23won't be fair that if I don't give you
- 7:24:25the solution to this particular problem.
- 7:24:27So, we'll move forward and understand
- 7:24:28how deep learning solves these kind of
- 7:24:29problems.
- 7:24:31Now, as you can see that the first line
- 7:24:32on your slide, which says that deep
- 7:24:34learning models are capable to focus on
- 7:24:36the right features by themselves,
- 7:24:37requiring little guidance from the
- 7:24:38programmer. So, with the help of little
- 7:24:40guidance, what these deep learning
- 7:24:42algorithms can do. They can generate the
- 7:24:44features on which the outcome will
- 7:24:46depend on. And at the same time, it also
- 7:24:48solves the dimensionality problem as
- 7:24:50well. If you have very large number of
- 7:24:51inputs and outputs, you can make use of
- 7:24:53a deep learning algorithm. Now, what
- 7:24:55exactly is deep learning? Again, since
- 7:24:57we know that it has been evolved by
- 7:24:59machine learning and machine learning is
- 7:25:01nothing but a subset of artificial
- 7:25:02intelligence. And the idea behind
- 7:25:03artificial intelligence is to imitate
- 7:25:05the human behavior. The same idea is for
- 7:25:07the deep learning as well is to build
- 7:25:09learning algorithms that can mimic
- 7:25:11brain.
- 7:25:12Now, let us move forward and understand
- 7:25:14deep learning what exactly it is.
- 7:25:16Now, the deep learning is implemented
- 7:25:18with the help of neural networks. And
- 7:25:19the idea or the motivation behind neural
- 7:25:21networks are nothing but neurons. What
- 7:25:23are neurons? These are nothing but your
- 7:25:24brain cells. Now, here's a diagram of
- 7:25:26neuron. So, we have dendrites here,
- 7:25:28which are used to provide input to a
- 7:25:30neuron. As you can see, we have multiple
- 7:25:32dendrites here. So, these many inputs
- 7:25:34will be provided to a neuron. Now, this
- 7:25:35is called cell body and inside the cell
- 7:25:37body, we have a nucleus, which performs
- 7:25:39some function. After that, that output
- 7:25:41will travel through axon and it will go
- 7:25:44towards the axon terminals. And then,
- 7:25:46this neuron will fire this output
- 7:25:48towards the next neuron. Now, the
- 7:25:50studies tell us that the next neuron now
- 7:25:52or you can say the two neurons are never
- 7:25:54connected to each other. There's a gap
- 7:25:55between them. So, that is called a
- 7:25:57synapse. So, this is how basically a
- 7:26:00neuron works like. And on the right hand
- 7:26:02side of your slide, you can see an
- 7:26:03artificial neuron. Now, let me explain
- 7:26:05you that. So, over here, similar to
- 7:26:07neurons, we have multiple inputs. Now,
- 7:26:09these inputs will be provided to a
- 7:26:11processing element like a cell body. And
- 7:26:14over here in the processing element,
- 7:26:15what will happen? Summation of your
- 7:26:17inputs and weights. Now, when it moves
- 7:26:20on, then what will happen? This input
- 7:26:22will be multiplied with our weights. So,
- 7:26:24in the beginning, what happens? These
- 7:26:25weights are randomly assigned. So, what
- 7:26:27will happen if I take the example of X1?
- 7:26:29So, X1 multiplied by W1 will go towards
- 7:26:32the processing element. Similarly, X2
- 7:26:34and W2 will go towards the processing
- 7:26:36element. And, similarly, the other
- 7:26:38inputs as well. And, then summation will
- 7:26:40happen which will generate a function of
- 7:26:41S, that is f of S.
- 7:26:43After that comes the concept of
- 7:26:45activation function. Now, what is
- 7:26:47activation function? It is nothing but
- 7:26:49in order to provide a threshold. So, if
- 7:26:51your output is above the threshold, then
- 7:26:52only this neuron will fire, otherwise it
- 7:26:54won't fire. So, you can use a step
- 7:26:56function as an activation function, or
- 7:26:57you can even use a sigmoid function as
- 7:26:59your activation function. So, this is
- 7:27:01how an artificial neuron it looks like.
- 7:27:03So, a network will be multiple neurons
- 7:27:05which are connected to each other will
- 7:27:06form an artificial neural network. And,
- 7:27:08this activation function can be a
- 7:27:10sigmoid function or a step function,
- 7:27:12that totally depends on your
- 7:27:13requirement.
- 7:27:14Now, once it exceeds the threshold, it
- 7:27:16will fire. After that, what will happen?
- 7:27:18It will check the output. Now, if this
- 7:27:20output is not equal to the desired
- 7:27:22output, so these are the actual outputs,
- 7:27:24and we know the real output. So, we'll
- 7:27:26compare both of that, and we'll find the
- 7:27:28difference between the actual output and
- 7:27:30the desired output. On the basis of that
- 7:27:32difference, we are again going to update
- 7:27:34our weights. And, this process will keep
- 7:27:36on repeating until we get the desired
- 7:27:39output as our actual output. Now, this
- 7:27:41process of updating weight is nothing
- 7:27:43but your back propagation method.
- 7:27:45So, this is neural networks in a
- 7:27:47nutshell. So, we'll move forward and
- 7:27:49understand what are deep networks. So,
- 7:27:51basically, deep learning is implemented
- 7:27:53by the help of deep networks, and deep
- 7:27:54networks are nothing but neural networks
- 7:27:57with multiple hidden layers. Now, what
- 7:27:59are hidden layers? Let me explain you
- 7:28:00that. So, you have inputs that comes
- 7:28:03here. So, this will be your input layer.
- 7:28:05After that, some process happens, and
- 7:28:07it'll go to the next node, or you can
- 7:28:09say to the hidden layer nodes. So, this
- 7:28:11is nothing but your hidden layer one.
- 7:28:13So, every node is interconnected if you
- 7:28:16can notice. After that, you have one
- 7:28:18more hidden layer where some function
- 7:28:19will happen. And, as you can see that
- 7:28:22again these nodes are interconnected to
- 7:28:23each other. After this hidden layer two
- 7:28:26comes the output layer, and this output
- 7:28:28layer again we are going to check the
- 7:28:30output whether it is equal to the
- 7:28:31desired output or not. If it is not, we
- 7:28:33are again going to update the weights.
- 7:28:35So, this is how a deep network looks
- 7:28:37like. Now, there can be multiple hidden
- 7:28:39layers. There can be hundreds of hidden
- 7:28:41layers as well. But, when we talk about
- 7:28:43machine learning, that was not the case.
- 7:28:45We were not able to process multiple
- 7:28:47hidden layers when we talk about machine
- 7:28:49learning. So, because of deep learning,
- 7:28:51we have multiple hidden layers at once.
- 7:28:54Now, let us understand this with an
- 7:28:55example. So, we'll take an image which
- 7:28:57has four pixels. So, if you can notice,
- 7:28:59we have four pixels here, among which
- 7:29:01the top two pixels are bright, that is
- 7:29:03they are black in color, whereas bottom
- 7:29:04two pixels are white. Now, what happens?
- 7:29:07We'll divide these pixels and we'll send
- 7:29:08these pixels to each and every node. So,
- 7:29:11for that, we need four nodes. So, this
- 7:29:13particular pixel will go to this node,
- 7:29:14it will go to this node, this pixel will
- 7:29:16go to this node, and finally this pixel
- 7:29:18will go to this particular node that I'm
- 7:29:19highlighting with my cursor. Now, what
- 7:29:21happens? We provide them random weights.
- 7:29:24So, these white lines actually represent
- 7:29:25the positive weights, and these black
- 7:29:27lines represents the negative weights.
- 7:29:29Now, this particular brightness, when we
- 7:29:31display high brightness, we'll consider
- 7:29:32it as negative. Now, what happens? When
- 7:29:35you see the next output or the next
- 7:29:36hidden layer, it'll be provided with the
- 7:29:38input with this particular layer. So,
- 7:29:40this will provide an input with positive
- 7:29:42weight to this particular node, and the
- 7:29:44second input will come from this
- 7:29:45particular node. Since both of them are
- 7:29:47positive, so we'll get this kind of a
- 7:29:49node. Similarly, this node as well. Now,
- 7:29:51when I talk about these two nodes, the
- 7:29:52first node over here, so this is getting
- 7:29:54input from this node as well as from
- 7:29:56this node. Now, over here we have a
- 7:29:58negative weight. So, because of that,
- 7:30:00the value will be negative, and we have
- 7:30:02represented that with black color.
- 7:30:04Similarly, over here as well, we're
- 7:30:06getting one input from here which has a
- 7:30:07negative weight, and the another input
- 7:30:09from here which has again has a negative
- 7:30:10weight. So, accordingly, we get again a
- 7:30:13negative value here. So, these two
- 7:30:14becomes black in color. Now, if you
- 7:30:17notice what'll happen next, we'll
- 7:30:19provide one input here, which will be
- 7:30:21negative and a positive weight, which
- 7:30:23will be again negative, and this will be
- 7:30:25also negative and a positive weight. So,
- 7:30:27that will again come out to be negative.
- 7:30:29So, that is why we have got this kind of
- 7:30:31a structure. If you notice this this is
- 7:30:33nothing but the inverse of this
- 7:30:34particular image. When I talk about this
- 7:30:36node over here, we are getting the
- 7:30:38negative value with a positive weight,
- 7:30:40which is negative, and a negative value
- 7:30:41with a negative weight, which is
- 7:30:42positive. So, we are getting something
- 7:30:44which is positive here.
- 7:30:46Now, obviously, I want this particular
- 7:30:47image to get inverse. I want these black
- 7:30:50strips to come up. So, what I'll do,
- 7:30:52I'll actually calculate the inverse by
- 7:30:54providing a negative weight like this.
- 7:30:55Over here, I've provided a negative
- 7:30:57weight, it'll come up. So, when I
- 7:30:58provide a positive weight, so it'll stay
- 7:31:00wherever it is. After that, it'll
- 7:31:02detect, and the output you can see will
- 7:31:04be a horizontal image, not a solid, not
- 7:31:06a vertical, not a diagonal, but a
- 7:31:08horizontal. And after that, we are going
- 7:31:10to calculate the difference between the
- 7:31:11actual output and the desired output,
- 7:31:13and we are going to update the weights
- 7:31:14accordingly. Now, this is just an
- 7:31:16example, guys. So, guys, this is one
- 7:31:18example of deep learning, where what
- 7:31:19happens, we have images here. We provide
- 7:31:22these raw data to the first layer to the
- 7:31:24input layer.
- 7:31:25Then, what happens, these input layers
- 7:31:27will determine the patterns of local
- 7:31:28contrast, or it'll fixate those patterns
- 7:31:30of local contrast, which means that
- 7:31:32it'll differentiate on the basis of
- 7:31:34colors and luminosity and all those
- 7:31:36things. So, it'll differentiate those
- 7:31:37things. And after that, in the following
- 7:31:39layer, what will happen, it'll determine
- 7:31:41the face features, it'll fixate those
- 7:31:43face features. So, it'll form nose,
- 7:31:45eyes, ears, all those things. Then, what
- 7:31:47will happen, it'll accumulate those
- 7:31:49correct features for the correct face,
- 7:31:51or you can say that and fixate those
- 7:31:52features on the correct face template.
- 7:31:55So, it'll actually determine the faces
- 7:31:56here, as you can see it over here. And
- 7:31:58then, it'll be sent to the output layer.
- 7:32:00Now, basically, you can add more hidden
- 7:32:02layers to solve more complex problem.
- 7:32:04For example, if I want to find out a
- 7:32:06particular kind of face, for example, a
- 7:32:08face which has large eyes, or which has
- 7:32:10light complexion. So, I can do that by
- 7:32:12adding more hidden layers. And I can
- 7:32:14increase the complexity also at the same
- 7:32:16time, if I want to find which image
- 7:32:18contains a dog. So, for for also, I can
- 7:32:20have one more hidden layer. So, as and
- 7:32:22when hidden layer increases, we are able
- 7:32:23to solve more and more complex problem.
- 7:32:25So, this is just a general overview of
- 7:32:27how a deep network looks like. So, we
- 7:32:29have first patterns of local contrast in
- 7:32:31the first layer. Then what happens, we
- 7:32:33fixate these patterns of local contrast
- 7:32:35in order to form the face features such
- 7:32:37as eyes, nose, ears, etc. And then we
- 7:32:39accumulate these features for the
- 7:32:41correct face and then we determine the
- 7:32:43image. So, this is how uh deep learning
- 7:32:46network or you can say deep network
- 7:32:47looks like.
- 7:32:48So, we'll move forward and I'll give you
- 7:32:50some applications of deep learning. So,
- 7:32:52here are few applications of deep
- 7:32:53learning. It can be used in self-driving
- 7:32:55cars. So, you must have heard about
- 7:32:57self-driving cars. So, what happens,
- 7:32:59it'll capture the images around it.
- 7:33:00It'll process that huge amount of data
- 7:33:02and then it'll decide what action should
- 7:33:04it take. Should it take left, right?
- 7:33:05Should it stop? So, accordingly it'll
- 7:33:07decide what action should it take and
- 7:33:09that will reduce the amount of accidents
- 7:33:10that happens every year. Then when we
- 7:33:12talk about voice control assistants, I'm
- 7:33:14pretty sure you must have heard about
- 7:33:15Siri. All the iPhone users know about
- 7:33:17Siri, right? So, you can tell Siri
- 7:33:19whatever you want to do. It'll search it
- 7:33:20for you and display for you.
- 7:33:22Then when we talk about automatic image
- 7:33:24caption generation. So, what happens in
- 7:33:26this, whatever image that you upload,
- 7:33:27the algorithm is in such a way that
- 7:33:29it'll generate the caption accordingly.
- 7:33:31So, for example, if you have say blue
- 7:33:33colored eyes, so it'll display a blue
- 7:33:35colored eye caption uh at the bottom of
- 7:33:37the image.
- 7:33:38Now, when I talk about automatic machine
- 7:33:39translation, so we can convert English
- 7:33:42language into Spanish. Similarly,
- 7:33:44Spanish to French. So, basically
- 7:33:46automatic machine translation, you can
- 7:33:47convert one language to another language
- 7:33:49with the help of deep learning. And
- 7:33:51these are just few examples, guys. There
- 7:33:52are many, many other examples of deep
- 7:33:54learning. It can be used in game
- 7:33:56playing. It can be used in many other
- 7:33:58things. And let me tell you one very
- 7:33:59fascinating thing that I've told you in
- 7:34:00the beginning as well. With the help of
- 7:34:02deep learning, MIT is trying to predict
- 7:34:04future. So, yeah, I know it is growing
- 7:34:06exponentially right now, guys.
- 7:34:14So, this is the problem statement, guys.
- 7:34:15We need to figure out if the bank notes
- 7:34:17are real or fake. And for that, we'll be
- 7:34:19using artificial neural network. And
- 7:34:21obviously, we need some sort of data in
- 7:34:23order to train our network. So, let us
- 7:34:25see how the data set looks like. So,
- 7:34:27over here I've taken a screenshot of the
- 7:34:29data set with few of the rows. In it,
- 7:34:31data were extracted from images that
- 7:34:33were taken from genuine and forged bank
- 7:34:35note-like specimens.
- 7:34:37After that, wavelet transform tools were
- 7:34:39used to extract features from those
- 7:34:41images. And these are few features that
- 7:34:43I'm highlighting with my cursor. And the
- 7:34:45final column or the last column actually
- 7:34:47represents the label.
- 7:34:48So, basically, label tells us to which
- 7:34:50class that pattern represents, whether
- 7:34:52that pattern represents a fake note or
- 7:34:54it represents a real note. Let us
- 7:34:56discuss these features and labels one by
- 7:34:58one.
- 7:34:59So, the first feature or the first
- 7:35:00column is nothing but variance of a
- 7:35:02wavelet transformed image. The second
- 7:35:04column is about skewness. The third is
- 7:35:06kurtosis of wavelet transformed image.
- 7:35:08And finally, fourth one is entropy of
- 7:35:10the image. After that, when I talk about
- 7:35:12label, which is nothing but my last
- 7:35:13column, over here if the value is one,
- 7:35:15that means the pattern represents a real
- 7:35:17note. Whereas, when value is zero, that
- 7:35:19means it represents a fake note. So
- 7:35:21guys, let's move forward and we'll see
- 7:35:22what are the various steps involved in
- 7:35:24order to implement this use case.
- 7:35:26So, over here we'll first begin by
- 7:35:28reading the data set that we have. We'll
- 7:35:29define features and labels.
- 7:35:32After that, we are going to encode the
- 7:35:33dependent variable. And what is a
- 7:35:35dependent variable? It is nothing but
- 7:35:36your label.
- 7:35:37Then, we are going to divide the data
- 7:35:39set into two parts, one for training,
- 7:35:41another for testing.
- 7:35:42After that, we'll use TensorFlow data
- 7:35:44structures for holding features, labels,
- 7:35:46etc. And TensorFlow is nothing but a
- 7:35:48Python library that is used in order to
- 7:35:50implement deep learning models or you
- 7:35:52can say neural networks.
- 7:35:53Then, we'll write the code in order to
- 7:35:55implement the model. And once this is
- 7:35:57done, we will train our model on the
- 7:35:58training data. We'll calculate the
- 7:36:00error. The error is nothing but your
- 7:36:02difference between the model output and
- 7:36:04the actual output.
- 7:36:05And we'll try to reduce this error. And
- 7:36:07once this error becomes minimum, we'll
- 7:36:09make prediction on the test data and
- 7:36:11we'll calculate the final accuracy.
- 7:36:13So guys, let me quickly open my PyCharm
- 7:36:15and I'll show you how the output looks
- 7:36:16like.
- 7:36:18So this is my PyCharm, guys. Over here,
- 7:36:19I've already written the code in order
- 7:36:21to execute the use case. I'll go ahead
- 7:36:23and run this and I'll show you the
- 7:36:24output.
- 7:36:29So over here, as you can see, with every
- 7:36:30iteration, the accuracy is increasing.
- 7:36:33So let me just stop it right here.
- 7:36:35All right. Till now, any questions, any
- 7:36:37doubts with respect to what is our use
- 7:36:38case, what is the data set about? Any
- 7:36:41questions, guys? You can go ahead and
- 7:36:42ask me.
- 7:36:44Okay, there's a question from Arpan.
- 7:36:46He's asking, "Can you explain the code?"
- 7:36:47Definitely, Arpan. I'll be doing that at
- 7:36:49the end of this class when you are done
- 7:36:51with all the fundamentals of neural
- 7:36:52networks. I'll explain you the entire
- 7:36:53code, how I've written that, and how
- 7:36:55I've used TensorFlow in order to
- 7:36:56implement a neural network.
- 7:36:58I hope I you are satisfied with the
- 7:37:00answer. Okay, he's fine with it. Any
- 7:37:02other questions, any other doubts, guys?
- 7:37:03Just go ahead and ask me. Over here, you
- 7:37:05don't need to worry about code right
- 7:37:06now, guys, because I'll explain this
- 7:37:08later in the session. So what I'll do,
- 7:37:10I'll open my slides once more and we'll
- 7:37:12discuss the fundamentals of neural
- 7:37:13networks that are required in order to
- 7:37:15implement this use case.
- 7:37:16So in order to understand why we need
- 7:37:18neural networks, we are going to compare
- 7:37:20the approach before and after neural
- 7:37:21networks. And we'll see what were the
- 7:37:23various problems that were there before
- 7:37:25neural networks. So earlier,
- 7:37:26conventional computers use an
- 7:37:28algorithmic approach. That is, the
- 7:37:30computer follows a set of instructions
- 7:37:33in order to solve a problem. And unless
- 7:37:35the specific steps that the computer
- 7:37:37needs to follow are known, the computer
- 7:37:39cannot solve the problem. So obviously,
- 7:37:42we need a person who actually knows how
- 7:37:44to solve that problem and then he or she
- 7:37:45can provide the instructions to the
- 7:37:47computer as to how to solve that
- 7:37:48particular problem, right? So we first
- 7:37:50should know the answer to that problem,
- 7:37:52or we should know how to overcome that
- 7:37:54challenge or problem which is there in
- 7:37:55front of us. Then only we can provide
- 7:37:57instructions to the computer.
- 7:37:59So this restricts the problem-solving
- 7:38:00capability of conventional computers to
- 7:38:03problems that we already understand and
- 7:38:05know how to solve. But what about those
- 7:38:07problems whose answer we have no clue
- 7:38:09of? So, that's where our traditional
- 7:38:11approach was a failure. So, that's why
- 7:38:14neural networks were introduced. Now,
- 7:38:15let us see what was the scenario after
- 7:38:17neural networks.
- 7:38:18So, neural networks basically process
- 7:38:20information in a similar way the human
- 7:38:22brain does.
- 7:38:23And these networks, they actually learn
- 7:38:25from examples. You cannot program them
- 7:38:27to perform a specific task. They will
- 7:38:29learn from their examples, from their
- 7:38:31experience. So, you don't need to
- 7:38:33provide all the instructions to perform
- 7:38:34a specific task, and your network will
- 7:38:36learn on its own with its own
- 7:38:38experience.
- 7:38:39All right. So, this is what basically
- 7:38:40neural network does.
- 7:38:42So, even if you don't know how to solve
- 7:38:43a problem, you can train your network in
- 7:38:45such a way that with experience, it can
- 7:38:47actually learn how to solve the problem.
- 7:38:50So, that was a major reason why neural
- 7:38:52networks came into existence.
- 7:38:54We'll move forward and we'll understand
- 7:38:56what is the motivation behind neural
- 7:38:58networks.
- 7:38:59So, these neural networks are basically
- 7:39:01inspired by neurons, which are nothing
- 7:39:02but your brain cells.
- 7:39:04And the exact working of the human brain
- 7:39:06is still a mystery, though.
- 7:39:08So, as I've told you earlier as well
- 7:39:09that neural networks work like human
- 7:39:10brain and so the name.
- 7:39:13And similar to a newborn human baby, as
- 7:39:15he or she learns from his or her
- 7:39:17experience, we want our network to do
- 7:39:19that as well. But, we want it to do very
- 7:39:21quickly.
- 7:39:22So, here's a diagram of a neuron.
- 7:39:24Basically, a biological neuron receives
- 7:39:26input from other sources, combines them
- 7:39:29in some way, perform a generally
- 7:39:31non-linear operation on the result, and
- 7:39:33then outputs the final result. So, here
- 7:39:36if you notice these dendrites, these
- 7:39:37dendrites will receive signals from the
- 7:39:39other neurons. Then, what will happen?
- 7:39:41It will transfer it to the cell body.
- 7:39:43The cell body will perform some
- 7:39:44function. It can be summation, it can be
- 7:39:46multiplication. So, after performing
- 7:39:48that summation on the set of inputs, via
- 7:39:50axon it is transferred to the next
- 7:39:52neuron.
- 7:39:53Now, let's understand what exactly are
- 7:39:55artificial neural networks.
- 7:39:58It is basically a computing system that
- 7:40:00is designed to simulate the way the
- 7:40:02human brain analyzes and process the
- 7:40:04information. Artificial neural networks
- 7:40:06has self-learning capabilities that
- 7:40:08enable it to produce better results as
- 7:40:11more data becomes available. So, if you
- 7:40:13train your network on more data, it will
- 7:40:14be more accurate.
- 7:40:16So, these neural networks, they actually
- 7:40:17learn by example.
- 7:40:19And you can configure your neural
- 7:40:20network for specific applications. It
- 7:40:22can be pattern recognition or it can be
- 7:40:24data classification, anything like that,
- 7:40:26all right?
- 7:40:27So, because of neural networks, we see a
- 7:40:29lot of new technology has evolved.
- 7:40:31From translating web pages to other
- 7:40:33languages to having a virtual assistant
- 7:40:35to order groceries online to conversing
- 7:40:37with chatbots. All of these things are
- 7:40:40possible because of neural networks.
- 7:40:43So, in a nutshell, if I need to tell
- 7:40:44you, artificial neural network is
- 7:40:46nothing but a network of various
- 7:40:48artificial neurons.
- 7:40:50All right? So, let me show you the
- 7:40:51importance of neural network with two
- 7:40:53scenarios, before and after neural
- 7:40:55network.
- 7:40:56So, over here we have a machine and we
- 7:40:58have trained this machine on the four
- 7:41:00types of dogs, as you can see where I'm
- 7:41:02highlighting with my cursor.
- 7:41:03And once the training is done, we
- 7:41:05provide a random image to this
- 7:41:06particular machine which has a dog. But
- 7:41:09this dog is not like the other dogs on
- 7:41:11which we have trained our system on.
- 7:41:13So, without neural networks, our machine
- 7:41:15cannot identify that dog in the picture,
- 7:41:17as you can see it over here. Basically,
- 7:41:19our machine will be confused. It cannot
- 7:41:21figure out where the dog is. Now, when I
- 7:41:23talk about neural networks, even if you
- 7:41:25have not trained our machine on this
- 7:41:26specific dog, but still it can identify
- 7:41:29certain features of the dogs that we
- 7:41:31have trained on and it can match those
- 7:41:33features with the dog that is there in
- 7:41:34this particular image and it can
- 7:41:36identify that dog. So, this happens all
- 7:41:39because of neural networks. So, this is
- 7:41:41just an example to show you how
- 7:41:42important are neural networks. Now, I
- 7:41:44know you all must be thinking how neural
- 7:41:47networks work.
- 7:41:48So, for that, we'll move forward and
- 7:41:50understand how it actually works.
- 7:41:52So, over here I'll begin by first
- 7:41:54explaining a single artificial neuron
- 7:41:56that is called as perceptron.
- 7:41:58So, this is an example of a perceptron.
- 7:42:00Over here we have multiple inputs X1,
- 7:42:02X2, {dash} {dash} {dash} till Xn. And we
- 7:42:05have corresponding weights as well. W1
- 7:42:07for X1, W2 for X2, similarly Wn for Xn.
- 7:42:10Then what happens, we calculated the
- 7:42:12weighted sum of these inputs. And after
- 7:42:15doing that, we pass it through an
- 7:42:16activation function. This activation
- 7:42:19function is nothing but it provides a
- 7:42:20threshold value. So, above that value my
- 7:42:23neuron will fire, else it won't fire.
- 7:42:26So, this is basically an artificial
- 7:42:27neuron. So, when I talk about a neural
- 7:42:29network, it involves a lot of these
- 7:42:31artificial neurons with their own
- 7:42:33activation function and their processing
- 7:42:35element.
- 7:42:36Now, we'll move forward and we'll
- 7:42:38actually understand various modes of
- 7:42:40this perceptron or single artificial
- 7:42:42neuron. So, there are two modes in a
- 7:42:44perceptron. One is training, another is
- 7:42:45using mode. In training mode, the neuron
- 7:42:48can be trained to fire for particular
- 7:42:50input patterns, which means that we'll
- 7:42:52actually train our neuron to fire on
- 7:42:54certain set of inputs and to not fire on
- 7:42:56the other set of inputs. That's what
- 7:42:58basically training mode is. When I talk
- 7:43:00about using mode, it means that when a
- 7:43:01taught input pattern is detected at the
- 7:43:03input, its associated output becomes the
- 7:43:05current output, which means that once
- 7:43:07the training is done and we provide an
- 7:43:09input on which the neuron has been
- 7:43:11trained on, so it will detect the input
- 7:43:14and will provide the associated output.
- 7:43:16So, that's what basically using mode is.
- 7:43:18So, first you need to train it, then
- 7:43:19only you can use your perceptron or your
- 7:43:21network.
- 7:43:23So, these were the two modes, guys. And
- 7:43:24next up we'll understand what are the
- 7:43:25various activation functions available.
- 7:43:28So, these are the three activation
- 7:43:29functions, although there are many more,
- 7:43:30but I've listed down three. Step
- 7:43:32function. So, over here the moment your
- 7:43:34input is greater than this particular
- 7:43:35value, your neuron will fire, else it
- 7:43:37won't. Similarly for sigmoid and sign
- 7:43:39function as well. So, these are three
- 7:43:41activation functions. There are many
- 7:43:42more that I've told you earlier as well.
- 7:43:44So, yeah, these are the three majorly
- 7:43:45used activation functions. Next up what
- 7:43:48we are going to do, we are going to
- 7:43:49understand how a neuron learns from its
- 7:43:51experience. So, I'll give you a very
- 7:43:53good analogy in order to understand
- 7:43:55that. And later on when we talk about
- 7:43:57the neural networks or you can say
- 7:43:58multiple neurons in a network, I'll
- 7:44:00explain you the math behind it. I'll
- 7:44:02explain you the math behind learning how
- 7:44:04it actually happens. So, right now I'll
- 7:44:05explain you with an analogy. And guys,
- 7:44:07trust me that analogy is pretty
- 7:44:09interesting.
- 7:44:10So, I know all of you must have guessed
- 7:44:12it. So, these are two beer mugs and all
- 7:44:14of you who love beer can actually relate
- 7:44:15to this analogy a lot.
- 7:44:17And I know most of you actually love
- 7:44:19beer, so that's why I've chosen this
- 7:44:20particular analogy so that all of you
- 7:44:22can relate to it.
- 7:44:24All right, jokes apart. So, fine guys,
- 7:44:26so there's a beer festival happening
- 7:44:27near your house.
- 7:44:29And you want to badly go there. But your
- 7:44:31decision actually depends on three
- 7:44:32factors. First is how is the weather,
- 7:44:34whether it is good or bad. Second is
- 7:44:37your wife or husband is going with you
- 7:44:38or not. And the third one is any public
- 7:44:40transport is available. So, on these
- 7:44:43three factors your decision will depend
- 7:44:44whether you will go or not. So, we'll
- 7:44:46consider these three factors as inputs
- 7:44:49to our perceptron. And we'll consider
- 7:44:51our decision of going or not going to
- 7:44:53the beer festival as our output. So, let
- 7:44:55us move forward with that. So, the first
- 7:44:57input is how is the weather, we'll
- 7:44:58consider it as X1. So, when weather is
- 7:45:00good, it'll be one and when it is bad,
- 7:45:02it'll be zero.
- 7:45:03Similarly, your wife is going with you
- 7:45:05or not, so that'd be your X2. If she is
- 7:45:08going then it's one, if she's not going
- 7:45:10then it's zero. Similarly for public
- 7:45:11transport, if it is available then it is
- 7:45:13one, else it is zero.
- 7:45:14So, these are the three inputs that I'm
- 7:45:15talking about. Let's see the output. So,
- 7:45:18output will be one when you're going to
- 7:45:19the beer festival and output will be
- 7:45:21zero when you want to relax at home. You
- 7:45:23want to have beer at home only, you
- 7:45:25don't want to go outside. So, these are
- 7:45:27the two outputs, whether you're going or
- 7:45:28you're not going.
- 7:45:30Now, what a human brain does. Over here,
- 7:45:32okay, fine. I need to go to the beer
- 7:45:34festival, but there are three things
- 7:45:36that I need to consider. But will I give
- 7:45:38importance to all these factors equally?
- 7:45:41Definitely not. There'll be certain
- 7:45:43factors which will be of high priority
- 7:45:45for me. I'll focus on those factors
- 7:45:47more. Whereas few factors won't affect
- 7:45:50that much to me. All right. So, let's
- 7:45:52prioritize our inputs or factors. So,
- 7:45:54here our most important factor is
- 7:45:56weather. So, if weather is good, I love
- 7:45:58beer so much that I don't care even if
- 7:45:59my wife is going with me or not or if
- 7:46:01there is a public transport available.
- 7:46:03So, I love beer that much that if
- 7:46:05weather is good, then definitely I'm
- 7:46:06going there. That means when X1 is high,
- 7:46:08output will be definitely high.
- 7:46:11So, how we do that? How we actually
- 7:46:12prioritize our factors or how we
- 7:46:14actually give importance more to a
- 7:46:16particular input and less to another
- 7:46:18input in a perceptron or in a neuron?
- 7:46:21So, we do that by using weights. So, we
- 7:46:23assign high weights to the more
- 7:46:24important factors or more important
- 7:46:27inputs and we assign low weights to
- 7:46:28those particular inputs which are not
- 7:46:30that important for us.
- 7:46:31So, let's assign weights, guys. So,
- 7:46:33weight W1 is associated with input X1,
- 7:46:36W2 with X2, and similarly W3 with X3.
- 7:46:39Now, as I've told you earlier as well
- 7:46:41that weather is a very important factor,
- 7:46:42so I'll assign a pretty high weight to
- 7:46:43weather and I'll keep it at six.
- 7:46:45Similarly, W2 and W3 are not that
- 7:46:47important, so I'll keep it as two two.
- 7:46:49After that, I've defined a threshold
- 7:46:51value as five, which means that when the
- 7:46:53weighted sum of my input is greater than
- 7:46:55five, then only my neuron will fire or
- 7:46:58you can say then only I'll be going to
- 7:46:59the beer festival.
- 7:47:00All right. So, I'll use my pen and we'll
- 7:47:03see what happens when weather is good.
- 7:47:06So, when weather is good, our X1 is one.
- 7:47:09Our weight is six, we'll multiply it
- 7:47:10with six.
- 7:47:11Then,
- 7:47:14if my wife decides that she is going to
- 7:47:16stay at home and she will probably be
- 7:47:18busy with cooking and she doesn't want
- 7:47:20to drink beer with me. So, she's not
- 7:47:22coming. So, that input becomes zero.
- 7:47:24Zero into two will actually make no
- 7:47:26difference because it'll be zero.
- 7:47:29Then again, there's no public transport
- 7:47:30available also. Then also this will be
- 7:47:32zero into two.
- 7:47:35So, what output I get here?
- 7:47:37I get here as six.
- 7:47:40I notice the threshold value that is
- 7:47:42five. So, definitely six is greater than
- 7:47:44five.
- 7:47:46That means my output
- 7:47:49will be one or you can say my neuron
- 7:47:51will fire or I'll actually go to the
- 7:47:53beer festival.
- 7:47:54So, even if these two inputs are zero
- 7:47:57for me, that means my wife is not
- 7:47:58willing to go with me and there is no
- 7:48:00public transport available, but weather
- 7:48:02is good, which has very high weight
- 7:48:03value and it actually matters a lot to
- 7:48:05me.
- 7:48:06So, if that is high, it doesn't really
- 7:48:08matter whether the two inputs are high
- 7:48:09or not. I'll go to the beer festival.
- 7:48:11All right? Now, I'll explain you a
- 7:48:13different scenario. So, over here our
- 7:48:15threshold was five, but what if I change
- 7:48:17this threshold to three? So, in that
- 7:48:20scenario, even if my weather is not
- 7:48:22good, uh I'll give it a zero. So, zero
- 7:48:25into six.
- 7:48:26But, my wife and public transport both
- 7:48:30are available.
- 7:48:31All right? So, one into two
- 7:48:34plus
- 7:48:35one into two.
- 7:48:38Which is equal to four.
- 7:48:41And it is definitely greater than three.
- 7:48:45Then also my output will be one. That
- 7:48:48means I will definitely go to the beer
- 7:48:49festival even if weather is bad.
- 7:48:52And my neuron will fire. So, these are
- 7:48:54the two scenarios that I've discussed
- 7:48:56with you. All right? So, there can be
- 7:48:57many other ways in which you can
- 7:48:59actually assign weight to your problem
- 7:49:02or to your learning algorithm.
- 7:49:04So, these are the two ways in which you
- 7:49:05can assign weight and prioritize your
- 7:49:07inputs or factors on which your output
- 7:49:09will depend.
- 7:49:10So, obviously on real life all the
- 7:49:12inputs or all the factors are not as
- 7:49:14important for you. So, you actually
- 7:49:16prioritize them. And how you do that in
- 7:49:17perceptron, you provide high weight to
- 7:49:19it. This is just an analogy so that you
- 7:49:22can relate to a perceptron to a real
- 7:49:24life. We'll actually discuss the math
- 7:49:26behind it later in the session as to how
- 7:49:28a network or a neuron learns. All right?
- 7:49:31So, how the weights are actually updated
- 7:49:33and how the output is changing, that all
- 7:49:36those things we'll be discussing later
- 7:49:37in this session. But my aim is to make
- 7:49:40you understand that you can actually
- 7:49:42relate to a real life problem with that
- 7:49:44of a perceptron. All right? And in real
- 7:49:47life problems are not that easy. They
- 7:49:49are very very complex problems that uh
- 7:49:51we actually face. So in order to solve
- 7:49:53those problems, a single neuron is
- 7:49:55definitely not enough. So we need
- 7:49:57networks of neuron. And that's where
- 7:49:59artificial neural network, or you can
- 7:50:01say multi-layer perceptron, comes into
- 7:50:03the picture. Now let us discuss that.
- 7:50:06Multi-layer perceptron or artificial
- 7:50:07neural network.
- 7:50:09So this is how an artificial neural
- 7:50:10network actually looks like. So over
- 7:50:12here we have multiple neurons in present
- 7:50:14in different layers. The first layer is
- 7:50:16always your input layer. This is where
- 7:50:18you're actually feed in all of your
- 7:50:19inputs. Then we have the first hidden
- 7:50:21layer. Then we have second hidden layer,
- 7:50:24and then we have the output layer.
- 7:50:25Although the number of hidden layers
- 7:50:26depend on your application, on what are
- 7:50:28you working, what is your problem. So
- 7:50:30that actually determines how many hidden
- 7:50:32layers you'll have.
- 7:50:33So let me explain you what is actually
- 7:50:34happening here. So you provide in some
- 7:50:36input to the first layer, which is
- 7:50:38nothing but your input layer. You
- 7:50:39provide inputs to these neurons. All
- 7:50:41right? And after some function, the
- 7:50:43output of these neurons will become the
- 7:50:45input to the next layer, which is
- 7:50:46nothing but your hidden layer one. Then
- 7:50:48these hidden layers also have various
- 7:50:50neurons. These neurons will have
- 7:50:51different activation functions. So
- 7:50:53they'll perform their own function on
- 7:50:54the inputs that it receives from the
- 7:50:56previous layer, and then the output of
- 7:50:58this layer will be the input to the next
- 7:51:00hidden layer, which is hidden layer two.
- 7:51:02Similarly, the output of this hidden
- 7:51:04layer will be the input to the output
- 7:51:06layer. And finally, we get the output.
- 7:51:09So this is how basically an artificial
- 7:51:10neural network looks like. Now let me
- 7:51:12explain you this with an example.
- 7:51:14So over here I'll take an example of
- 7:51:16image recognition using neural networks.
- 7:51:19So over here what happens, we feed in a
- 7:51:21lot of images to our input layer.
- 7:51:23Now this input layer will actually
- 7:51:25detect the patterns of local contrast.
- 7:51:28And then we'll feed that to the next
- 7:51:29layer, which is hidden layer one. So in
- 7:51:31this hidden layer one, the face features
- 7:51:34will be recognized. They'll recognize
- 7:51:36eyes, nose, ears, things like that. And
- 7:51:38then, that will be again fed as input to
- 7:51:41the next hidden layer.
- 7:51:42And in this hidden layer, we'll assemble
- 7:51:44those features and we'll try to make a
- 7:51:45face. And then, we'll get the output
- 7:51:48that is the face will be recognized
- 7:51:50properly. So, if you notice here, with
- 7:51:52every layer, we're trying to get a more
- 7:51:54abstract version or the generalized
- 7:51:55version of the input. So, this is how
- 7:51:58basically an artificial neural network
- 7:52:00work, how it works. All right.
- 7:52:02And there's a lot of training and
- 7:52:03learning which is involved that I'll
- 7:52:04show you now.
- 7:52:06Training a neural network. So, how we
- 7:52:07actually train our neural network? So,
- 7:52:09basically, the most common algorithm for
- 7:52:11training a network is called back
- 7:52:12propagation.
- 7:52:14So, what happens in back propagation?
- 7:52:16After the weighted sum of inputs and
- 7:52:17passing through an activation function
- 7:52:19and getting the output, we compare that
- 7:52:21output to the actual output that we
- 7:52:22already know. We figure out how much is
- 7:52:24the difference. We calculate the error.
- 7:52:27And based on that error, what we do, we
- 7:52:28propagate backwards. And we'll see what
- 7:52:31happens when we change the weight. Will
- 7:52:33the error decrease or will it increase?
- 7:52:35And if it increases, when it increases
- 7:52:37by increasing the value of the variables
- 7:52:39or by decreasing the value of variables.
- 7:52:41So, we kind of calculate all those
- 7:52:43things and we update our variables in
- 7:52:45such a way that our error becomes
- 7:52:47minimum. And it takes a lot of
- 7:52:49iterations. Trust me, guys. It takes a
- 7:52:51lot of iterations. We get output a lot
- 7:52:53of times and then we compare it with the
- 7:52:54model with the actual output. Then
- 7:52:56again, we propagate backwards. We change
- 7:52:58the variables. Then again, we calculate
- 7:52:59the output. We compare it again with the
- 7:53:01desired output or the actual output.
- 7:53:03Then again, we propagate backwards. So,
- 7:53:05this process keeps on repeating until we
- 7:53:06get the minimum value.
- 7:53:08All right. So, there's an example that
- 7:53:10is there in front of your screen. Don't
- 7:53:11be scared of the terms that I used. I'll
- 7:53:13actually explain you with an example.
- 7:53:15So, this is the example over here. We
- 7:53:16have zero, one, and two as inputs. And
- 7:53:18our desired output or the output that we
- 7:53:20already know is zero, one, and four. All
- 7:53:22right. So, over here, we can actually
- 7:53:23figure out that desired output is
- 7:53:25nothing but twice of your input. But I'm
- 7:53:27training a computer to do that, right?
- 7:53:29The computer is not a human.
- 7:53:31So, what happens? I actually initialize
- 7:53:33my weight. I keep the value as three.
- 7:53:35So, the model output will be 3 * 0 is 0,
- 7:53:383 * 1 is 3, 3 * 2 is 6. Now, obviously
- 7:53:42it is not equal to your desired output.
- 7:53:43So, we check the error. Now, the error
- 7:53:46that we have got here is 0, 1, and 2,
- 7:53:48which is nothing but your difference.
- 7:53:49So, 0 - 0 is 0, 3 - 2 is 1, 6 - 4 is 2.
- 7:53:53Now, this is called an absolute error.
- 7:53:55After squaring this error, we get square
- 7:53:57error, which is nothing but 0, 1, and 4.
- 7:54:00All right? So, now what we need to do,
- 7:54:01we need to update the variables. We have
- 7:54:03seen that the output that we got is
- 7:54:05actually different from the desired
- 7:54:07output. So, we need to update the value
- 7:54:08of the weight. So, instead of three, our
- 7:54:11computer makes it as four. After making
- 7:54:13the value as four, we get the model
- 7:54:15output as 0, 4, and 8.
- 7:54:17And then we saw that the error has
- 7:54:19actually increased. Instead of
- 7:54:20decreasing, the error has increased. So,
- 7:54:22after updating the variable, the error
- 7:54:24has increased. So, you can see that
- 7:54:26square error is now 0, 4, and 16, and
- 7:54:28earlier it was 0, 1, and 4. That means
- 7:54:30we cannot increase the weight value
- 7:54:32right now. But if we decrease that, make
- 7:54:34it as two, we get the output, which is
- 7:54:37actually equal to desired output. But is
- 7:54:39it always the case that we need to only
- 7:54:41decrease the weight? Definitely not.
- 7:54:44So, in this particular scenario,
- 7:54:45whenever I'm increasing the weight,
- 7:54:46error is increasing, and when I'm
- 7:54:47decreasing the weight, error is
- 7:54:49decreasing. But as I've told you earlier
- 7:54:51as well, this is not the case every
- 7:54:52time. Sometimes you need to increase the
- 7:54:54weight as well. So, how we determine
- 7:54:56that? All right. Fine, guys. This is how
- 7:54:58basically a computer decide whether it
- 7:54:59has to increase the weight or decrease
- 7:55:01the weight. So, what happens here? This
- 7:55:02is a graph of square error versus
- 7:55:04weight.
- 7:55:05So, over here, what happens? Suppose
- 7:55:07your square error is somewhere here.
- 7:55:09And your computer, it starts increasing
- 7:55:11the weight in order to reduce the square
- 7:55:13error. And it notices that whenever it
- 7:55:14increases the weight, square error is
- 7:55:16actually decreasing.
- 7:55:17So, it'll keep on increasing until the
- 7:55:19square error reaches a minimum value.
- 7:55:22And after that, when it tries to still
- 7:55:24increase the weight, the square error
- 7:55:26will increase. So, at that time, our
- 7:55:28network will recognize that whenever it
- 7:55:30is increasing the weight after this
- 7:55:31point, error is increasing. So,
- 7:55:33therefore, it will stop right there, and
- 7:55:34that will be our weight value.
- 7:55:36Similarly, there can be one more
- 7:55:38scenario. Suppose if we increase the
- 7:55:40weight, but then also the square error
- 7:55:41is increasing. So, at that time, we
- 7:55:44cannot increase the weight. At that
- 7:55:45time, computer will realize, "Okay,
- 7:55:46fine. Whenever I'm increasing the
- 7:55:47weight, the square error is increasing.
- 7:55:49So, it'll go in the opposite direction."
- 7:55:51So, it'll start decreasing the weight,
- 7:55:52and it'll keep on doing that until the
- 7:55:54square error becomes minimum. And the
- 7:55:56moment it decreases more, the square
- 7:55:58error is again increases. So, our
- 7:56:00network will know that
- 7:56:01whenever it decreases the weight value,
- 7:56:04the square error is increasing. So, that
- 7:56:05point will be our final weight value.
- 7:56:08So, guys, this is what basically
- 7:56:09backpropagation in a nutshell is. Fine.
- 7:56:12So, we'll move forward, and now is the
- 7:56:14correct time to understand how to
- 7:56:15implement the use case that I was
- 7:56:17talking about in the beginning. That is,
- 7:56:19how to determine whether a node is fake
- 7:56:20or real. So, for that, I'll open my
- 7:56:22PyCharm.
- 7:56:24This is my PyCharm again, guys. Let me
- 7:56:26just close this. All right.
- 7:56:28So, this is the code that I've written
- 7:56:30in order to implement the use case. So,
- 7:56:32over here, what we do, we import the
- 7:56:33first important libraries which are
- 7:56:35required. Matplotlib is used for
- 7:56:36visualization. TensorFlow, we know, in
- 7:56:38order to implement the neural network.
- 7:56:40NumPy for arrays, Pandas for reading the
- 7:56:42data set. Similarly, scikit-learn for
- 7:56:44label encoding as well as for shopping,
- 7:56:46and also to split the data set into
- 7:56:48training and testing parts.
- 7:56:49All right. Fine, guys. So, we'll begin
- 7:56:51by first reading the data set, as I've
- 7:56:52told you earlier as well when I was
- 7:56:54explaining the steps. So, what I'll do,
- 7:56:56I'll use Pandas in order to read the CSV
- 7:56:58file, which has the data set.
- 7:57:00After that, I'll define features and
- 7:57:02labels. So, X will be my feature, and Y
- 7:57:04will contain my label. So, basically, X
- 7:57:06includes all the columns apart from the
- 7:57:08last column, which is the fifth one. And
- 7:57:10because the indexing starts from zero,
- 7:57:12that's why we have written zero till
- 7:57:14fourth. So, it won't include the fourth
- 7:57:16column. All right? And so, our last
- 7:57:18column will actually be our label.
- 7:57:21Then, what we need to do, we need to
- 7:57:22encode the dependent variable.
- 7:57:24So, dependent variable, as I've told,
- 7:57:26nothing but your label. So, I've
- 7:57:28discussed encoding in TensorFlow
- 7:57:29tutorial, you can go through it, and you
- 7:57:31can actually get to know why and how we
- 7:57:32do that. Then, what we have done, we
- 7:57:35have uh read the data set. Then, what we
- 7:57:37need to do is to split our data set into
- 7:57:38training and testing. And uh these are
- 7:57:41all optional steps. You can print the
- 7:57:42shape of your training and test data. If
- 7:57:44you don't want to do it, it's still
- 7:57:45fine.
- 7:57:46Then, we have defined learning rate. So,
- 7:57:47learning rate is actually the steps in
- 7:57:50which the weights will be updated, all
- 7:57:52right? So, that is what basically
- 7:57:53learning rate is. Then, when we talk
- 7:57:55about epochs means iterations.
- 7:57:58Then, we have defined cost history, that
- 7:57:59will be an empty NumPy array, and its
- 7:58:02shape will be one, and it will include
- 7:58:03the float type object. Then, we have
- 7:58:05defined N dim, which is nothing but your
- 7:58:07X shape of axis one, which means your
- 7:58:09column. Then, we'll print that. After
- 7:58:12that, we have defined the number of
- 7:58:13classes. So, there can be only two
- 7:58:14class, whether the note can be fake or
- 7:58:16it can be real. And this model path I've
- 7:58:19given in order to save my model. So,
- 7:58:21I've just given a path where I need to
- 7:58:23save it. So, I'll just save it here
- 7:58:24only, in the current working directory.
- 7:58:26Now is the time to actually define our
- 7:58:29neural network. So, we'll first make
- 7:58:31sure that we have defined the important
- 7:58:33parameters like hidden layers, number of
- 7:58:35neurons in hidden layers. So, I'll take
- 7:58:3610 neurons in every hidden layer, and
- 7:58:37I'm taking four layers like that. Then,
- 7:58:40X will be my placeholder, and the shape
- 7:58:42of this particular placeholder is none,
- 7:58:43{comma} N {underscore} dim. N
- 7:58:45{underscore} dim value I'll get it from
- 7:58:47here, and none can be at any value. I'll
- 7:58:49define one variable W, and I'll
- 7:58:51initialize it with zeros, and this will
- 7:58:53be the shape of my weight. Similarly,
- 7:58:56for bias as well, this will be the
- 7:58:57particular shape. And there will be one
- 7:58:59more placeholder Y dash, which will
- 7:59:01actually be used in order to provide us
- 7:59:03with the actual output of the model.
- 7:59:05There'll be one model output, and
- 7:59:07there'll be one actual output, which we
- 7:59:08use in order to calculate the
- 7:59:09difference, right? So, we'll feed in the
- 7:59:11actual values of the labels in this
- 7:59:13particular placeholder Y dash.
- 7:59:16And now we'll define the model. So, over
- 7:59:18here we have named the function as
- 7:59:20multilayer perceptron, and in it we'll
- 7:59:22first define the first layer. So, the
- 7:59:24first hidden layer, and we are going to
- 7:59:26name it as layer underscore one, which
- 7:59:28will be nothing but the a matrix
- 7:59:30multiplication of X and weights of H1,
- 7:59:33that is the hidden layer one.
- 7:59:35And that'll be added to your biases B1.
- 7:59:37After that, we'll pass it through a
- 7:59:39sigmoid activation function. Similarly,
- 7:59:40in layer two as well, matrix
- 7:59:42multiplication of layer one and weights
- 7:59:45of H2. So, if you can notice, layer one
- 7:59:47was the network layer just before the
- 7:59:50layer two, right? So, the output of this
- 7:59:52layer one will become input to the layer
- 7:59:53two. And that's why we have written
- 7:59:55layer underscore one. It'll be
- 7:59:56multiplied by weights H2, and then we'll
- 7:59:58add it with the bias.
- 8:00:00Similarly, for this particular hidden
- 8:00:01layer as well, and this particular layer
- 8:00:03as well. But, over here we are going to
- 8:00:05use a ReLU activation function instead
- 8:00:07of sigmoid. Then, we are going to define
- 8:00:09the weights and biases. So, this is how
- 8:00:12we basically define weights. This is how
- 8:00:13we basically define weights. So, weights
- 8:00:15H1 will be a variable which will be a
- 8:00:18truncated normal with the shape of N
- 8:00:20underscore dim and N underscore hidden
- 8:00:22underscore one. So, these are nothing
- 8:00:24but your shapes. All right.
- 8:00:26And after that, what we have done, we
- 8:00:27have defined biases as well. Then, we
- 8:00:29need to initialize all the variables.
- 8:00:31So, all the guys, in brief, let's talk
- 8:00:34about TensorFlow.
- 8:00:36Since in TensorFlow, we need to
- 8:00:37initialize a variable before we use it.
- 8:00:42That's how we
- 8:00:43do it. We first initialize it.
- 8:00:46And then we need to run it. That's when
- 8:00:48your variables will be initialized.
- 8:00:51After that, we are going to
- 8:00:52create a stateful object, and then
- 8:00:55finally, I'm going to call my model.
- 8:00:58And then comes it
- 8:01:00part where the training happens. Cost
- 8:01:02function. Cost function
- 8:01:04is nothing but you can say an error that
- 8:01:07will be calculated between the actual
- 8:01:09output
- 8:01:10and the model output.
- 8:01:13All right, so Y is nothing but our model
- 8:01:15output and
- 8:01:17that is nothing but actual output or the
- 8:01:19output that we already know.
- 8:01:21All right, and then we are going to use
- 8:01:23a gradient
- 8:01:24descent optimizer to reduce the error.
- 8:01:27Then
- 8:01:28we are going to create a session object
- 8:01:30and uh finally we are going to run the
- 8:01:32session.
- 8:01:33So,
- 8:01:34this is how we basically for every
- 8:01:35calculated change as
- 8:01:38well as the accuracy that comes after
- 8:01:40every
- 8:01:41the epoch on the training data.
- 8:01:44After we have calculated the accuracy on
- 8:01:46the training data, we are going to plot
- 8:01:47it for every
- 8:01:49accuracy is.
- 8:01:51And after
- 8:01:52after plotting that we have accuracy
- 8:01:53with our tell using the same prediction
- 8:01:55on the test and after the print the and
- 8:01:57the mean score. So, let's do this, guys.
- 8:02:00All right, so training and what
- 8:02:01See, accuracy epochs see has 99%. So,
- 8:02:05with every epoch it is actually
- 8:02:06increasing apart from a couple of
- 8:02:08instances, it is actually keep on
- 8:02:09increasing. So, the more data you train
- 8:02:11your model on, it will be more accurate.
- 8:02:14Let me just close it. So, now the model
- 8:02:16has also been saved where I wanted it to
- 8:02:18be. This is my final test accuracy and
- 8:02:21this is the mean squared error. All
- 8:02:23right, so these are the files that will
- 8:02:24appear once you save your model.
- 8:02:26These are the four files that I've
- 8:02:27highlighted. Now, what we need to do is
- 8:02:29restore this particular model and I've
- 8:02:32explained this in detail how how to re-
- 8:02:35restore a model that you have already
- 8:02:37saved. So, over here what I'll take I've
- 8:02:39taken before to 768.
- 8:02:41So, all the values in the row of 754 and
- 8:02:45768 will be fed to our model and our
- 8:02:48model will make prediction on that. So,
- 8:02:50let us go ahead and run this.
- 8:02:53So, when I'm restoring my model, it
- 8:02:55seems that my model is 100% I'll use a
- 8:02:57value
- 8:02:58fed in. So, whatever values that I have
- 8:03:00actually given as input to my model, it
- 8:03:02has correctly identified its class,
- 8:03:04whether it's a
- 8:03:06fake note or a real note, because fake
- 8:03:09note and one stands for real note, okay?
- 8:03:11So, original class is nothing but a set,
- 8:03:13so it is zero already. And what
- 8:03:15prediction my model has made is zero,
- 8:03:17that means it is fake. percent.
- 8:03:19Similarly, for other values as well.
- 8:03:23Fine, guys. So, this is how we basically
- 8:03:24implement the use case that we saw in
- 8:03:26the beginning.
- 8:03:27So, in this slide you can notice that
- 8:03:29I've listed down only two applications,
- 8:03:30although there are many more.
- 8:03:32So, neural networks in medicine.
- 8:03:34Artificial neural networks are currently
- 8:03:36a very hot research area in medicine,
- 8:03:38and it is believed that they will
- 8:03:39receive extensive application to
- 8:03:42biomedical systems in the next few
- 8:03:43years. And currently, the research is
- 8:03:46mostly on modeling parts of human body
- 8:03:48and uh recognizing diseases from various
- 8:03:50scans. For example, it can be
- 8:03:51cardiograms, CAT scans, ultrasonic
- 8:03:53scans, etc.
- 8:03:55And uh currently, the research is going
- 8:03:57uh mostly on uh two major areas. First
- 8:03:59is modeling and diagnosing the
- 8:04:00cardiovascular system.
- 8:04:02So, neural networks are used
- 8:04:03experimentally to model the human
- 8:04:05cardiovascular system.
- 8:04:07Diagnosis can be achieved by building a
- 8:04:08model of the cardiovascular system of an
- 8:04:10individual and comparing it with the
- 8:04:12real-time physiological measurements
- 8:04:14taken from the patient. And trust me,
- 8:04:16guys, if this routine is carried out
- 8:04:18regularly, potential harmful medical
- 8:04:21conditions can be detected at an early
- 8:04:23stage and thus, make the process of
- 8:04:25combating disease much easier.
- 8:04:27Apart from that, it is currently being
- 8:04:29used in electronic noses as well.
- 8:04:31Electronic noses have several potential
- 8:04:33applications in telemedicine. Now, let
- 8:04:36me just give you an introduction to
- 8:04:37telemedicine. Telemedicine is a practice
- 8:04:39of medicine over long distance via a
- 8:04:41communication link. So, what the
- 8:04:43electronic noses will do, they would
- 8:04:45identify odors in the remote surgical
- 8:04:47environment. These identified odors
- 8:04:50would then be electronically transmitted
- 8:04:52to another site, wherein odor generation
- 8:04:54system would recreate them.
- 8:04:57Because the sense of the smell can be an
- 8:04:58important sense to the surgeon,
- 8:05:00tele-smell would enhance tele-present
- 8:05:02surgery.
- 8:05:04So, these are the two ways in which you
- 8:05:05can use it in medicine. You can use it
- 8:05:08in business as well, guys. So, business
- 8:05:10is basically a diverted field with
- 8:05:12several general areas of specialization
- 8:05:14such as accounting or financial
- 8:05:16analysis. Almost any neural network
- 8:05:18application would fit into one business
- 8:05:20area or financial analysis.
- 8:05:22Now, there is some potential for using
- 8:05:24neural networks for business purposes
- 8:05:25including resource allocation and
- 8:05:27scheduling. I've listed down two major
- 8:05:29areas where it can be used. One is
- 8:05:31marketing.
- 8:05:32So, there is a marketing application
- 8:05:33which has been integrated with a neural
- 8:05:35network system.
- 8:05:37The airline marketing tactician is a
- 8:05:39computer system made of various
- 8:05:41intelligent technologies including
- 8:05:43expert systems. A feedforward neural
- 8:05:45network is integrated with the AMT,
- 8:05:47which is nothing but airline marketing
- 8:05:49tactician, and was trained using
- 8:05:51backpropagation to assist the marketing
- 8:05:53control of airline seat allocation.
- 8:05:56So, it has wide applications in
- 8:05:58marketing as well.
- 8:06:00Now, the second area is credit
- 8:06:01evaluation. Now, I'll give you an
- 8:06:02example here. The HNC company has
- 8:06:05developed several neural network
- 8:06:06applications, and one of them is a
- 8:06:08credit scoring system which increases
- 8:06:10the profitability of existing model up
- 8:06:12to
- 8:06:13So, these are few applications that I'm
- 8:06:15telling you guys. Neural network is
- 8:06:17actually the future.
- 8:06:19People are talking about neural networks
- 8:06:21everywhere, and especially after the
- 8:06:23intro-
- 8:06:24duction of GPUs and the amount of data
- 8:06:26that we have now, neural network is
- 8:06:28actually spreading like plague right
- 8:06:29now.
- 8:06:31>> [music]
- 8:06:37[music]
- 8:06:40>> So, this is an
- 8:06:41image of New York's this picture. So,
- 8:06:43when a human will see this image, he'll
- 8:06:45a lot of buildings in different colors
- 8:06:47and stuff like that. But, how are this
- 8:06:49image? So, So there'll be three
- 8:06:51channels. Red,
- 8:06:52another will be green and finally we
- 8:06:53have blue channel which is popularly
- 8:06:55known as RGB. So all each of these
- 8:06:58channels will they have their own
- 8:06:59respective pixel values as you can see
- 8:07:01it over here. So when I say size is B
- 8:07:04cross A cross 3, it means that there are
- 8:07:07B
- 8:07:08rows, A columns and three channels. All
- 8:07:10right? So So if somebody tells you that
- 8:07:13the size of an image is 28 cross 28
- 8:07:15cross three pixels, it means that it has
- 8:07:1728 rows, 28 columns and three channels.
- 8:07:20So this is how
- 8:07:21this is for colored images for we have
- 8:07:23only two channels. So let's move forward
- 8:07:25and we'll see why can't we use for image
- 8:07:27classification.
- 8:07:28So consider an image which has 28 three
- 8:07:30pixels.
- 8:07:32So when I feed in this image to a fully
- 8:07:33con-
- 8:07:34nected network like this, then the total
- 8:07:36number of weights required in the fully
- 8:07:38connected 2,352.
- 8:07:40You can just go ahead and multiply it
- 8:07:41your-
- 8:07:42self. All right?
- 8:07:44But in real life the images are not that
- 8:07:46small. All right? So whatever images
- 8:07:47that we have, they are definitely above
- 8:07:49200 cross 200 cross three pixels.
- 8:07:51So if I take an image which has 200
- 8:07:53cross 200 cross three pixels and I feed
- 8:07:55it to a fully connected network, that
- 8:07:57time the number of weights required
- 8:07:58itself will be 120,000 guys. So we need
- 8:08:00to deal with such huge amount of
- 8:08:02parameters and obviously we require more
- 8:08:04number of neurons. So that can
- 8:08:05eventually lead to overfitting. So
- 8:08:07that's why we can't use network for
- 8:08:09image classification. Let's see why we
- 8:08:10need convolutional neural networks.
- 8:08:12Basically in convolutional neural
- 8:08:14network in the layer will only be
- 8:08:15connected to a small region of the layer
- 8:08:17before it. So if you consider this
- 8:08:18particular neuron which I'm highlighting
- 8:08:20right now is only connected to three
- 8:08:21other neurons. Unlike the fully
- 8:08:23connected network where this particular
- 8:08:24neuron will be connected to all these
- 8:08:25five neurons. Because of this we need to
- 8:08:27handle less amount of weights and in
- 8:08:29turn we need less number of neurons as
- 8:08:31well. So let us understand what exactly
- 8:08:33is convolutional neural network. So
- 8:08:35convolutional neural networks are
- 8:08:36special type
- 8:08:38of feedforward artificial neural
- 8:08:40networks which is inspired from visual
- 8:08:43cortex. So visual cortex
- 8:08:45This but a small region in our brain
- 8:08:47brain which is present somewhere here
- 8:08:49where you can see the bulb and basically
- 8:08:52what happened
- 8:08:53was an experiment conducted and people
- 8:08:55got to know that visual cortex is small
- 8:08:57regions of cells that are sensitive to
- 8:08:58specific regions of visual field.
- 8:09:01So what I'm
- 8:09:02example some neurons in the visual
- 8:09:03cortex exposed to vertical edges. Some
- 8:09:05will fire when exposed to horizontal
- 8:09:07edges. Some will fire when exposed to
- 8:09:08diagonal edges and that is nothing but
- 8:09:10the motivation behind convolutional
- 8:09:12neural network. So now let us
- 8:09:14convolutional neural network work force.
- 8:09:15So generally a collect work has three
- 8:09:17layers convolution labeling layer and
- 8:09:18fully connected layer. We'll understand
- 8:09:20each of these layers one by one. We'll
- 8:09:22take an example of a classifier that can
- 8:09:24classify an image of an X as well as an
- 8:09:26O. So with this example we'll be
- 8:09:28understanding all these four layers. So
- 8:09:30let's begin guys. Now there are certain
- 8:09:32trickier cases. So what I mean by that
- 8:09:33is X can be represented in these four
- 8:09:36forms as well, right? So these are
- 8:09:38nothing but the deformed images of X.
- 8:09:39Similarly for O as well. So these are
- 8:09:42deformed images. So even I want to
- 8:09:43classify these images either X or O. All
- 8:09:46right, because even this is X, this is
- 8:09:47X, this is X, this is X. But all these
- 8:09:49are deformed images. But they are in
- 8:09:52turn X, right? So I want my classifier
- 8:09:53to classify them as X. So basically
- 8:09:56that's what I want. So if you can notice
- 8:09:58here this is a proper image of an X and
- 8:10:00which is actually equal to this
- 8:10:02particular X which is a deformed image.
- 8:10:03Same goes for this O as well. So now
- 8:10:05what we are going to do is we know that
- 8:10:07a computer understands an image using
- 8:10:08numbers at each pixels. So what we'll
- 8:10:10do, whatever the white pixels that we
- 8:10:12have we are going to assign a value
- 8:10:13minus one to it and whatever the black
- 8:10:15pixels we have we are going to assign a
- 8:10:16value one to it. When we use normal
- 8:10:18techniques to compare these two images,
- 8:10:20one is a proper image of X and another
- 8:10:21is a deformed image of X, we got to know
- 8:10:23that a computer is not able to classify
- 8:10:25the deformed image of X correctly. Why?
- 8:10:27Because it is comparing it with the
- 8:10:29proper image of X, right? So when you go
- 8:10:31ahead and add the pixel values of both
- 8:10:33of these images you get something like
- 8:10:35this. So basically our computer is not
- 8:10:37able to recognize whether it is an X or
- 8:10:39not. Now what we do with the help of CNN
- 8:10:41we take small patches of our image. So,
- 8:10:43these patches or these pieces are known
- 8:10:46as nothing but features or filters. So,
- 8:10:48what we do, by finding rough feature
- 8:10:50matches in roughly the same positions in
- 8:10:52two images, CNN gets a lot better at
- 8:10:54seeing the similarity between the whole
- 8:10:56image matching schemes. What I mean by
- 8:10:57that is, we have these filters, right?
- 8:10:59We have these filters that you can see.
- 8:11:01So, consider this first filter. This is
- 8:11:03exactly equal to the feature or the part
- 8:11:05of the image in the deformed image as
- 8:11:07well. So, this is our proper image and
- 8:11:08this is our deformed image, all right?
- 8:11:10Right? So, this particular feature or
- 8:11:11this particular part of the image is
- 8:11:13actually equal to this particular part
- 8:11:14of the image. Same goes for this
- 8:11:16particular feature or filter as well.
- 8:11:18And similarly, we have this filter as
- 8:11:19well, which is actually equal to this
- 8:11:21particular part of the deformed image,
- 8:11:24all right? So, let's move forward and
- 8:11:25we'll see we're taking in our example.
- 8:11:27So, we'll be considering these three
- 8:11:29features or filters. This is a diagonal
- 8:11:31filter, this is again a diagonal filter
- 8:11:32and this is nothing but a small x. So,
- 8:11:34we'll take these three filters and we'll
- 8:11:36move forward. So, what we are going to
- 8:11:37do is we are going to compare these
- 8:11:39features, the small pieces of the bigger
- 8:11:41image, we are going to put it on the
- 8:11:43input image and if it matches, then the
- 8:11:45image will be classified correctly. Now,
- 8:11:47we'll begin, guys. The first layer is
- 8:11:48convolution layer. So, these are the
- 8:11:50beginning two steps of this particular
- 8:11:51layer. First, we need to line up the
- 8:11:53feature in the image and then multiply
- 8:11:54image by the corresponding feature
- 8:11:56pixel. Now, let me explain you with an
- 8:11:57example. So, this is our first diagonal
- 8:11:59feature that we'll take. We are going to
- 8:12:01put this particular feature on our image
- 8:12:04of x, all right? And we're going to
- 8:12:05multiply the corresponding pixel value.
- 8:12:07So, one will be multiplied with one,
- 8:12:09we'll get one and we'll put it in
- 8:12:10another matrix. Similarly, we are going
- 8:12:13to move forward and we're going to
- 8:12:14multiply minus one with minus one. We're
- 8:12:16going to multiply minus one with minus
- 8:12:18one, as you can see. Similarly, we
- 8:12:19multiply this result, minus one into
- 8:12:21minus one, then again minus one into
- 8:12:22minus one. So, we are going to complete
- 8:12:24this whole process and we're going to
- 8:12:25finish up this matrix, all right? And
- 8:12:27once we are done finishing up the
- 8:12:28multiplication of all the corresponding
- 8:12:30pixels in the feature as well as in the
- 8:12:32image, we need to follow two more steps.
- 8:12:34We need to add them up and divide by the
- 8:12:36total number of the pixels in the
- 8:12:37feature. So, what I mean by that is
- 8:12:39after the multiplication of the
- 8:12:41corresponding pixel values, what we do,
- 8:12:43we add all these values, we divide by
- 8:12:45the total number of pixels, and we get
- 8:12:47some value, right? And then now our next
- 8:12:49step is to create a map and put the
- 8:12:51value of the filter at that particular
- 8:12:53place. We saw that after multiplying the
- 8:12:54pixel value of a feature with the
- 8:12:56corresponding pixel value of with that
- 8:12:58of our image, we get the output which is
- 8:13:00one. So, we place one here. Similarly,
- 8:13:03we are going to move this filter
- 8:13:05throughout the image. Next up, we are
- 8:13:06going to move this filter here. After
- 8:13:08that, we're going to move it here, here,
- 8:13:09here, everywhere on the image we are
- 8:13:11going to move it and we're going to
- 8:13:12follow the same process. All right, so
- 8:13:14yeah, this is one more example where
- 8:13:15I've moved my filter in between and
- 8:13:17after doing that, I've got the output
- 8:13:19something like this, 1 1 -1 and all. So,
- 8:13:22over here if you notice, I've got couple
- 8:13:24of times -1 as well, due to which my
- 8:13:26output that comes is 0.55, right? So,
- 8:13:29I'm going to place 0.55 here. Similarly,
- 8:13:31after moving the pixel after moving the
- 8:13:33filter throughout the image, I got this
- 8:13:35particular matrix. All right? And this
- 8:13:37is for one particular feature. After
- 8:13:39performing the same process for the
- 8:13:41other two filters as well, I've got
- 8:13:43these two values. So, we have these
- 8:13:45three values after passing through the
- 8:13:46convolution layer. Let me give you a
- 8:13:48quick recap of what happens in
- 8:13:49convolution layer. So, basically we have
- 8:13:51taken three features, all right? And one
- 8:13:53by one we'll take one feature, move it
- 8:13:55through the entire image, and when we
- 8:13:56are moving it, at that time we are
- 8:13:57multiplying the pixel value of the image
- 8:13:59with that of the corresponding pixel
- 8:14:00value of the filter, adding them up,
- 8:14:02dividing by the total number of pixels
- 8:14:04to get the output. So, when we do that
- 8:14:07for all the filters, we get we got these
- 8:14:09three outputs, all right? So, let's move
- 8:14:10forward and we'll see what happens in
- 8:14:12ReLU layer. So, this is ReLU layer,
- 8:14:14guys, and people who have gone through
- 8:14:15the previous tutorial actually know what
- 8:14:17it is. So, let me just give you a quick
- 8:14:18introduction of ReLU layer. So, ReLU is
- 8:14:20nothing but a activation function. All
- 8:14:22right? So, what I mean by that is it
- 8:14:23will only activate a node if the input
- 8:14:26is above a certain quantity. While the
- 8:14:28input is below zero, the output is also
- 8:14:30zero, all right? And when the input
- 8:14:32rises above the certain threshold, it
- 8:14:34has a linear relationship with the
- 8:14:36dependent variable. Now, I'll explain
- 8:14:38you with an example. We have a graph of
- 8:14:40ReLU function here. So, my function says
- 8:14:42that when f of x is equal to zero if x
- 8:14:45is less than zero, and it is equal to x
- 8:14:47when x is greater than zero. All right?
- 8:14:49So, whatever values that I have which
- 8:14:51are below zero will actually in turn
- 8:14:53become zero, and whatever values that
- 8:14:54are above zero, our function value will
- 8:14:57also be equal to that particular value.
- 8:14:59So, f of x will be equal to x if it is
- 8:15:01greater than or equal to zero, and it
- 8:15:03will be zero if it is less than zero.
- 8:15:05So, if I have x value as minus three, so
- 8:15:07definitely it is less than zero, so f of
- 8:15:09x becomes zero. Similarly, if I have
- 8:15:11minus five x value, then that again it
- 8:15:12is less than zero, so my f of x value
- 8:15:14becomes zero. But, when I consider three
- 8:15:16as my x value, then my f of x becomes
- 8:15:19equal to x, which is nothing but three.
- 8:15:20So, over here I'll have three. Again, if
- 8:15:22I take my x value as five, then
- 8:15:24obviously it is greater than or equal to
- 8:15:26zero, then my f of x becomes equal to x,
- 8:15:30so my f of x value becomes five. So,
- 8:15:32this is how our ReLU function works. So,
- 8:15:34why are we using ReLU function here is
- 8:15:35we want to remove all the negative
- 8:15:37values from our output that we got
- 8:15:39through the convolution layer. So, we'll
- 8:15:41only take the first output that we got
- 8:15:43by moving one feature throughout the
- 8:15:45image. So, this is the output that we
- 8:15:47have got for only one filter. All right?
- 8:15:49So, over here I'm going to remove all
- 8:15:50negative values. So, over here you can
- 8:15:52see that it it was minus point one one
- 8:15:54before, and I've converted that to zero.
- 8:15:56Similarly, I'm going to repeat the whole
- 8:15:57process for the entire matrix. And once
- 8:16:00I'm done with that, I get this
- 8:16:01particular value. Now, remember this is
- 8:16:03only for the output that we got through
- 8:16:05one feature. All right? So, when we were
- 8:16:07doing convolution at that time, we were
- 8:16:08you
- 8:16:09So, this is the output only for one
- 8:16:10filter. After doing it for the output of
- 8:16:12the other two filters as well, we have
- 8:16:14got these two values more. So, totally
- 8:16:16we have these three values after passing
- 8:16:18through ReLU activation function. Next
- 8:16:20up, we'll see what exactly is pooling
- 8:16:21layer. So, in pooling layer what we do,
- 8:16:23we take a window size of two, and we
- 8:16:25move it across the entire matrix that we
- 8:16:27have got after passing through ReLU
- 8:16:28layer. And we take only the maximum
- 8:16:31value from there so that we can shrink
- 8:16:33the image. So what we are actually doing
- 8:16:35is we are reducing the size of our
- 8:16:37image. So let me explain you with an
- 8:16:38example. So this is basically one output
- 8:16:41that we have got after passing through
- 8:16:42ReLU layer. And over here we have taken
- 8:16:44a window size of two cross two. So when
- 8:16:46we keep this window at this particular
- 8:16:48position, we see that one is the highest
- 8:16:50value. So we're going to keep one here.
- 8:16:52And we are going to repeat the same
- 8:16:54process for this particular window as
- 8:16:56well. So over here the maximum value is
- 8:16:570.33 so 0.33 will come. So if you notice
- 8:17:00here, earlier we had
- 8:17:02seven cross seven matrix and now we have
- 8:17:04reduced that to four cross four matrix.
- 8:17:06So after doing that for the entire
- 8:17:08image, we have got this as our output.
- 8:17:11This output we have got after moving our
- 8:17:13window throughout the image that we have
- 8:17:16got after passing through ReLU layer,
- 8:17:17right? And when we repeat this process
- 8:17:19for all the three outputs that we have
- 8:17:21got after the ReLU layer, then we get
- 8:17:23this particular output after pooling
- 8:17:25layer. Right? So basically we have
- 8:17:27shrinked our image to a four cross four
- 8:17:29matrix. Now comes the tricky part. So
- 8:17:31what we are going to do now is stack up
- 8:17:33all these layers. So we have discussed
- 8:17:34convolution layer, ReLU layer and
- 8:17:36pooling layer. So I'll just give you a
- 8:17:38brief recap of what all things we have
- 8:17:39discussed. In convolution layer what we
- 8:17:41did, we took three features and then
- 8:17:43after that one by one we moved each
- 8:17:45filter throughout the image. And when we
- 8:17:47were moving it, we were continuously
- 8:17:49multiplying the image pixel value with
- 8:17:51that of the corresponding filter pixel
- 8:17:53value and then we were dividing it by
- 8:17:55the total number of pixels. All right?
- 8:17:57With that we got three output after
- 8:17:58passing through the convolution layer.
- 8:18:00Then those three output we passed
- 8:18:02through a ReLU layer where we have
- 8:18:03removed the negative value. All right?
- 8:18:05And after removing negative value again
- 8:18:07we have got the three outputs.
- 8:18:09Then those three outputs we passed
- 8:18:10through pooling layer. So basically
- 8:18:11we're trying to shrink our image. And
- 8:18:13what we did, we took a window size of
- 8:18:15two cross two, moved it through all the
- 8:18:17three outputs that we have got through
- 8:18:19ReLU layer. And after doing that, we
- 8:18:21were only taking the maximum value pixel
- 8:18:23value in that particular window and then
- 8:18:26we were putting it in a different matrix
- 8:18:27so that we get a shrinked image. And
- 8:18:29after passing it through pooling layer,
- 8:18:30we have got a four cross four matrix.
- 8:18:32Since we took three features in the
- 8:18:34beginning, so therefore we have got the
- 8:18:36three outputs after passing through
- 8:18:37pooling layer. All right. Next up, we
- 8:18:39are going to stack up all the layers,
- 8:18:41all right? So, let's do that. So, after
- 8:18:43passing through convolution, relu and
- 8:18:44pooling, we have got this four cross
- 8:18:46four matrix. This was our input image.
- 8:18:48Now, when we add one more layer of
- 8:18:50convolution, relu and pooling, we have
- 8:18:52shrinked our image from four cross four
- 8:18:53to two cross two as you can notice here.
- 8:18:55Now, we are going to use fully connected
- 8:18:57layer. Now, what happens in fully
- 8:18:58connected layer? The actual
- 8:18:59classification happens here, guys, okay?
- 8:19:01So, what we are doing here is we are
- 8:19:03going to take the shrinked images and
- 8:19:05put it into a single list. So, basically
- 8:19:08this is what we have got after passing
- 8:19:10through two layers of convolution, relu
- 8:19:12and pooling and this is what we have
- 8:19:13got. So, basically we're converting into
- 8:19:15a single list or a vector. How we do
- 8:19:17that? We take the first value one, then
- 8:19:18we take 0.55, then we take 0.55, then we
- 8:19:21take one again. Then we take one, then
- 8:19:22we take 0.55, 0.55, 0.55. Then we again
- 8:19:25take 0.55, one, one and 0.55. So, this
- 8:19:29is nothing but a vector or you can say a
- 8:19:31list. If you notice here that there are
- 8:19:33certain values in my list which are high
- 8:19:35for X and similarly if I repeat the
- 8:19:37entire process that we have discussed
- 8:19:39for O, there'll be certain will be high.
- 8:19:41So, for X we have first, fourth, fifth,
- 8:19:4410th and 11th element vector values are
- 8:19:47high. For O we have second, third, ninth
- 8:19:51and 12th element vector which are high.
- 8:19:53So, basically we know now if if we have
- 8:19:55an input image
- 8:19:57which has a first, fourth, 10th and 11th
- 8:20:00element vector values high, we know that
- 8:20:03we can classify it as X. Similarly, if
- 8:20:05our input image has a list which has the
- 8:20:07second, third,
- 8:20:09ninth and 12th element vector values
- 8:20:12high, then we can classify it as zero.
- 8:20:13Now, let me explain you with an example.
- 8:20:15So, after the training is done, after
- 8:20:17the after doing the entire process for
- 8:20:19both X and O, you know that our model is
- 8:20:21trained now, okay? So, we have given one
- 8:20:23a new input image and that input image
- 8:20:25passes through all the layers. And once
- 8:20:26it has passed through all the layers, we
- 8:20:28have got this 12-element vector. Now, it
- 8:20:30has 0.9, 0.65, all these values, right?
- 8:20:33Now, how do we classify it whether it is
- 8:20:35an X or O? So, what we do, we'll compare
- 8:20:37this with a list of X and O, right? So,
- 8:20:39we have got the list in the previous uh
- 8:20:41slide, if you notice. We have got two
- 8:20:43different lists for X and O. We are
- 8:20:45going to compare this new input image
- 8:20:47list that we have got with that of X and
- 8:20:49O, right? So, first let us compare that
- 8:20:51with X. Now, as I've told you earlier as
- 8:20:54well, for X there are certain values
- 8:20:55which will be higher, which is nothing
- 8:20:57but first, fourth, fifth, 10th, and 11th
- 8:20:59value, right? So, I'm going to sum
- 8:21:01first, fourth, fifth, 10th, and 11th
- 8:21:03value and I've got five. 1 + 1 + 1 + 1
- 8:21:06and + 1. So, five times one, I've got
- 8:21:08five. And now, I'm going to sum the
- 8:21:10corresponding values of my input image
- 8:21:11vector as well. So, the first value is
- 8:21:130.9. Then, the fourth value is 0.87,
- 8:21:16fifth value is 0.96, 10th value is 0.89,
- 8:21:19and the 11th value is 0.94. So, after
- 8:21:21this doing the sum of these values, I've
- 8:21:23got 4.56. When I divide this by five, I
- 8:21:25got 0.91, right? Now, this is for X.
- 8:21:28Now, when I do the same process for O,
- 8:21:30so in O, if you notice, I have second,
- 8:21:32third, ninth, and 12th element vector
- 8:21:35values is high. So, when I sum these
- 8:21:37values, I get four.
- 8:21:38And when I do the sum of the
- 8:21:40corresponding values in my input image,
- 8:21:41I've got 2.07. When I divide that by
- 8:21:44four, I got 4 for 0.51.
- 8:21:46So, now we notice that 0.91 is a higher
- 8:21:49value compared to 0.51. So, we have when
- 8:21:51we have compared our input image with
- 8:21:52the values of X, we got a higher value
- 8:21:55than the value that we have got after
- 8:21:56comparing the input image with the
- 8:21:58values of O. So, the input image is
- 8:22:00classified as X. All right, so now let
- 8:22:02us move towards our use case. So, this
- 8:22:04is our use case, guys. So, over here,
- 8:22:06what we are going to do is we are going
- 8:22:08to train our model on different types of
- 8:22:11dogs and cats images and then we are
- 8:22:14going to provide an input and it will
- 8:22:16classify whether the input is of a dog
- 8:22:19or a cat. Now, let me tell you the steps
- 8:22:21involved in it. So, what we are going to
- 8:22:23do in the beginning is obviously first
- 8:22:24we need to download the data set. After
- 8:22:26that we are going to write a function to
- 8:22:28encode the labels. Labels are nothing
- 8:22:30but the dependent variable that we are
- 8:22:32trying to predict. So, in our training
- 8:22:34data and testing data, obviously we know
- 8:22:36the labels, right? So, on that basis
- 8:22:38only we can train our model. So, we are
- 8:22:39going to encode those labels. After that
- 8:22:41we'll resize the image to 50 cross 50
- 8:22:43pixel and we are going to read it as a
- 8:22:45grayscale image. Then we are going to
- 8:22:47split the data, 24,000 images for
- 8:22:50training and 50 for testing. Once this
- 8:22:52is done, we are going to reshape the
- 8:22:54data appropriately for TensorFlow. Now,
- 8:22:56TensorFlow, I think everyone knows about
- 8:22:58TensorFlow. It's nothing but a Python
- 8:23:00library for implementing deep learning
- 8:23:01models. Then we are going to build the
- 8:23:03model, calculate the loss. It is nothing
- 8:23:05but categorical cross entropy. Then we
- 8:23:07are going to reduce the loss by using
- 8:23:10Adam optimizer with a learning rate set
- 8:23:11up set to point double zero one. Then we
- 8:23:14are going to train the train the deep
- 8:23:15neural network for 10 epochs and finally
- 8:23:18we are going to make predictions. All
- 8:23:19right, so I'll just quickly open my
- 8:23:21PyCharm and I'll show you the code how
- 8:23:22it looks like.
- 8:23:24So, this is the code that I've written
- 8:23:25in order to implement the use case. In
- 8:23:26the beginning I need to import the
- 8:23:28libraries that I require.
- 8:23:31And once it is done, what I mean my
- 8:23:32training data and the testing data. So,
- 8:23:36train and test one contains as well as
- 8:23:39testing data respectively. Then I've
- 8:23:41taken my image size as 50 and learning
- 8:23:44rate I've defined here and I've given a
- 8:23:46name to my model. You can give whatever
- 8:23:47name you want. All right, the first
- 8:23:49thing that we saw we need to encode the
- 8:23:51dependent variable. That's what we are
- 8:23:52doing here. We are encoding our
- 8:23:54dependent variable. So, whenever the
- 8:23:57label is cat, then it will be converted
- 8:23:59to an array of one comma zero and when
- 8:24:01it is dog it will be converted to an
- 8:24:02array of zero comma one. So, why why we
- 8:24:04are actually encoding the label? Because
- 8:24:06our code cannot understand the
- 8:24:07categorical variable. So, we need to
- 8:24:09encode it. Right? Next, what I'm doing
- 8:24:11is I'm resizing my image to 50 cross 50
- 8:24:14and I am converting it to a grayscale
- 8:24:15image. Right? And once this is done, I'm
- 8:24:17going to split my data set into training
- 8:24:20and testing parts.
- 8:24:21So, yeah, we are basically splitting the
- 8:24:22data set into two parts for training and
- 8:24:25testing.
- 8:24:26And here
- 8:24:27we are defining a model. So, you can
- 8:24:29just I can just go ahead and throw in a
- 8:24:32comment here.
- 8:24:35Building the model. Yeah. So, so this is
- 8:24:38where we are building the model. So,
- 8:24:40basically what we have done here is we
- 8:24:41have resized our image to 50 cross 50
- 8:24:44cross one matrix and that is the size of
- 8:24:46the input that we are using, right?
- 8:24:48Talking about. Then, what we have done
- 8:24:50here we have defined two filters and a
- 8:24:53stride of five with an activation
- 8:24:55function. After that
- 8:24:57that we have added a pooling layer, max
- 8:24:59pool layer. Okay? What we have done, we
- 8:25:01have repeated the same process, but over
- 8:25:04here we are taking 64 filters and five
- 8:25:07passing it through a real activation
- 8:25:08function. And after that we have
- 8:25:10have a
- 8:25:11repeated the 128 filters. After that we
- 8:25:12have repeated for 64 filters, then for
- 8:25:1432 filters. Then after that we are using
- 8:25:17a fully connected layer with 1024
- 8:25:19neurons. And finally we are using the
- 8:25:21dropout layer with key probability of
- 8:25:240.8 to finish our models. This is where
- 8:25:26our model is actually finished. And then
- 8:25:28what we are doing is we are using the
- 8:25:30Adam optimizer to optimize our model.
- 8:25:33So, basically whatever the loss that we
- 8:25:35have, we are trying to reduce it. And
- 8:25:37this is basically for your TensorBoard.
- 8:25:39So, we are creating some log files and
- 8:25:41then with that log file TensorBoard will
- 8:25:43create a pretty fancy graphs for us that
- 8:25:46helps us to visualize the entire model.
- 8:25:48And then what we are doing is we are
- 8:25:50trying to fit the model. And we have
- 8:25:51defined epochs as 10, that is the number
- 8:25:53of iterations that will happen will be
- 8:25:5610. And yeah, so this is pretty much it.
- 8:25:58Model name we have given. Then input is
- 8:26:00X score to check the accuracy. Similarly
- 8:26:04uh the target will be Y test labels
- 8:26:07associated with that test data will be a
- 8:26:09Y test and which we have encoded
- 8:26:11basically. So this is how we are going
- 8:26:13to actually calculate the accuracy and
- 8:26:16we'll try to reduce the loss as much as
- 8:26:17possible in 10 epochs. So till now our
- 8:26:20model is complete. We are done with it.
- 8:26:22Next what I'm doing is I'm feeding in
- 8:26:24some random input from the test data and
- 8:26:26I'm validating whether my model is
- 8:26:28predicting it correct or not. All right.
- 8:26:30So I've already trained the model
- 8:26:32because it takes a lot of time and yeah
- 8:26:34I cannot do it here. So I've already
- 8:26:36trained the model and you can see that
- 8:26:38the loss that came after the 10th epoch
- 8:26:41is 0.2973
- 8:26:43and the accuracy is somewhere around 88%
- 8:26:45which is pretty good guys and yeah and
- 8:26:48I've done the prediction on the test
- 8:26:49data as well. So let me just show it to
- 8:26:51you that. So this is the prediction that
- 8:26:53it has done on few of the images in the
- 8:26:55test data. So yeah it is a cat predicted
- 8:26:57as cat cat predicted as a cat cat cat
- 8:26:59cat and dogs as well. There are certain
- 8:27:01dogs as well.
- 8:27:04>> [music]
- 8:27:08>> Why can't we use feed forward networks?
- 8:27:11Now let us take an example of a feed
- 8:27:13forward network that is used for image
- 8:27:14classification. So we have trained this
- 8:27:17particular network for classifying
- 8:27:18various images of animals. Now if you
- 8:27:20feed in an image of a dog it'll identify
- 8:27:23that image and will provide a relevant
- 8:27:25label to that particular image.
- 8:27:26Similarly if you feed in an image of an
- 8:27:28elephant it'll provide relevant label to
- 8:27:31that particular image as well. Now if
- 8:27:32you notice the new output that we have
- 8:27:34got that is classifying an elephant has
- 8:27:37no relation with the previous output
- 8:27:39that is of a dog. Or you can say that
- 8:27:41the output at time T is independent of
- 8:27:44output at time T minus one. As we can
- 8:27:46see that there is no relation between
- 8:27:48the new output and the previous output.
- 8:27:50So we can say that in feed forward
- 8:27:52networks outputs are independent to each
- 8:27:54other. Now But are few scenarios where
- 8:27:56we actually need the previous output to
- 8:27:57get the new output. Let us discuss one
- 8:28:00such scenario.
- 8:28:01Now what happens when you read a book?
- 8:28:03You'll understand that book only on the
- 8:28:05understanding of your previous words.
- 8:28:07All right, so if I use a feedforward
- 8:28:08network and try to predict the next word
- 8:28:10in a sentence, I can't do that. Why
- 8:28:12can't I do that? Because my output will
- 8:28:15actually depend on the previous outputs.
- 8:28:17But in the feedforward network, my new
- 8:28:20output is independent of the previous
- 8:28:21outputs. That is, output at t plus one
- 8:28:24has no relation with output at t minus
- 8:28:26two, t minus one and at t. So basically,
- 8:28:28we cannot use feedforward networks for
- 8:28:30predicting the next word in a sentence.
- 8:28:32Similarly, you can think of many other
- 8:28:34examples where we need the previous
- 8:28:36output, some information from the
- 8:28:37previous output, so as to infer the new
- 8:28:39output. This is just one small example.
- 8:28:41There are many other examples that you
- 8:28:43can think of. So we'll move forward and
- 8:28:45understand how we can solve this
- 8:28:46particular problem. So over here, what
- 8:28:48we have done, we have input at t minus
- 8:28:50one. We'll feed it to our network, then
- 8:28:53we'll get the output at t minus one.
- 8:28:55Then at the next time stamp, that is at
- 8:28:57time t, we have input at time t. That
- 8:28:59will be given to our network along with
- 8:29:01the information from the previous time
- 8:29:03stamp, that is t minus one, and that
- 8:29:06will help us to get the output at t.
- 8:29:08Similarly, at output for t plus one, we
- 8:29:11have two inputs. One is new input that
- 8:29:13we give. Another is the information
- 8:29:15coming from the previous time stamp,
- 8:29:16that is t, in order to get the output at
- 8:29:19time t plus one. Similarly, it can go
- 8:29:21on. So over here, I have just written a
- 8:29:23generalized way to represent it. There's
- 8:29:25a loop where the information from the
- 8:29:27previous time stamp is flowing. This is
- 8:29:29how we can solve this particular
- 8:29:30challenge. Now let us understand what
- 8:29:32exactly are recurrent neural networks.
- 8:29:35So for understanding recurrent neural
- 8:29:36network, I'll take an analogy. Suppose
- 8:29:38your gym trainer has made a schedule for
- 8:29:40you. The exercises are repeated after
- 8:29:42every third day.
- 8:29:44Now this is the order of your exercises.
- 8:29:46First day you'll be doing shoulder,
- 8:29:47second day you'll be doing biceps, third
- 8:29:49day you'll be doing cardio. And all
- 8:29:50these exercises are repeated in a proper
- 8:29:52order. Now, what happens when we use a
- 8:29:54feedforward network for predicting the
- 8:29:56exercise today? So, we'll provide in the
- 8:29:58input such as day of the week, month of
- 8:30:00the year, and health status. All right,
- 8:30:02and we need to train our model or our
- 8:30:04network on the exercises that we have
- 8:30:06done in the past. After that, there'll
- 8:30:07be a complex voting procedure involved
- 8:30:09that will predict the exercise for us.
- 8:30:11And that procedure won't be that
- 8:30:13accurate. So, whatever output we'll get
- 8:30:15won't be as accurate as we want it to
- 8:30:17be. Now, what if I change my inputs and
- 8:30:20I make my inputs as what exercise I've
- 8:30:22done yesterday? So, if I've done
- 8:30:23shoulder, then definitely today I'll be
- 8:30:24doing biceps. Similarly, if I've done
- 8:30:26biceps yesterday, today I'll be doing
- 8:30:28cardio. Similarly, if I've done cardio
- 8:30:30yesterday, today I'll be doing shoulder.
- 8:30:32Now, there can be one scenario where you
- 8:30:34are unable to go to gym for 1 day. Due
- 8:30:36to some personal reasons, you could not
- 8:30:38go to the gym. Now, what will happen at
- 8:30:40that time?
- 8:30:41We'll go one time step back and we'll
- 8:30:43feed in what exercise that happened day
- 8:30:45before yesterday. So, if the exercise
- 8:30:47that happened day before yesterday was
- 8:30:49shoulder, then yesterday there were
- 8:30:50biceps exercises. All right, similarly,
- 8:30:52biceps happened day before yesterday,
- 8:30:54then yesterday would have been cardio
- 8:30:56exercises. Similarly, if cardio would
- 8:30:57have happened day before yesterday,
- 8:30:59yesterday would have been shoulder
- 8:31:00exercises. All right, and this
- 8:31:02prediction, the prediction for the
- 8:31:04exercise that happened yesterday, will
- 8:31:06be fed back to our network and these
- 8:31:08predictions will be used as inputs in
- 8:31:10order to predict what exercise will
- 8:31:12happen today. Similarly, if you have
- 8:31:14missed your gym, say for 2 days, 3 days,
- 8:31:16or 1 week, so you need to roll back. You
- 8:31:19need to go to the last day when you went
- 8:31:21to the gym. You need to figure out what
- 8:31:23exercise you did on that day, feed that
- 8:31:25as an input, and then only you'll be
- 8:31:26getting the relevant output as to what
- 8:31:28exercise will happen today.
- 8:31:30Now, what I'll do, I'll convert these
- 8:31:31things into a vector. Now, what is a
- 8:31:33vector? Vector is nothing but a list of
- 8:31:35numbers. All right, so this is the new
- 8:31:37information, guys, along with the
- 8:31:39information from the prediction at the
- 8:31:40previous time step. So, we need both of
- 8:31:43these in order to get the prediction at
- 8:31:44time t. Imagine if I've done shoulder
- 8:31:47exercises yesterday, so this will be
- 8:31:49one, this will be zero, this will be
- 8:31:50zero. Now, the prediction that will
- 8:31:52happen will be biceps exercise because
- 8:31:53if I have done shoulder yesterday, it's
- 8:31:55related to biceps. So, my output will be
- 8:31:57zero, one, and zero. And this is how
- 8:31:59vectors work, guys. So, I hope you have
- 8:32:01understood this, guys. Now, this is how
- 8:32:03a neural network looks like, guys. We
- 8:32:05have new information along with the
- 8:32:08information from the previous time step.
- 8:32:10The output that we have got in the
- 8:32:11previous time step will certain
- 8:32:13information from that. We'll feed into
- 8:32:15our network as inputs, and then that
- 8:32:17will help us to get the new output.
- 8:32:19Similarly, this new output that we have
- 8:32:21got will take some information from
- 8:32:23that, feed in as an input to our network
- 8:32:25along with the new information to get
- 8:32:26the new prediction, and this process
- 8:32:28keeps on repeating.
- 8:32:29Now, let me show you the math behind the
- 8:32:31recurrent neural networks.
- 8:32:33So, this is the structure of a recurrent
- 8:32:34neural network, guys. Let me explain you
- 8:32:36what happens here. Now, consider at time
- 8:32:38t equals to zero, we have input x
- 8:32:40naught, and we need to figure out what
- 8:32:41is x naught. So, according to this
- 8:32:43equation, h of zero is equal to w i,
- 8:32:47weight matrix, multiplied by our input x
- 8:32:49of zero plus w r into h of zero minus
- 8:32:54one, which is h of minus one, and time
- 8:32:56can never be negative, so we this
- 8:32:58particular equation cannot be applied
- 8:33:00here, plus a bias. So, w i into x of
- 8:33:03zero plus b h passes through a function
- 8:33:05g of h to get h of zero over here. After
- 8:33:08that, I want to calculate y naught. So,
- 8:33:10for y naught, I'll multiply h of zero
- 8:33:12with the weight matrix w i, and I'll add
- 8:33:14a bias to it and pass it through a
- 8:33:16function g of i to get y naught. Now, in
- 8:33:18the next time step, that is at time t
- 8:33:20equals to one, things become a bit
- 8:33:22tricky. Now, let me explain you what
- 8:33:24happens here. So, at time t equals to
- 8:33:26one, I have input x one, I need to
- 8:33:27figure out what is x one. So, for that,
- 8:33:29I'll use this equation. So, I'll
- 8:33:31multiply w i, that is the weight matrix,
- 8:33:34by the input x one plus w r into h of
- 8:33:38one minus one, which is zero. H of zero,
- 8:33:40we know what we got from here. So, WR
- 8:33:42into H of zero plus the bias, pass it
- 8:33:45through a function G of H to get the
- 8:33:47output as H1. Now, this H1 will use to
- 8:33:50get Y1. We'll multiply H1 with WY plus a
- 8:33:53bias and we'll pass it through a
- 8:33:55function G of Y to get Y1.
- 8:33:57Similarly, the next time stamp, that is
- 8:33:59at time T equals to two, we have input
- 8:34:01X2. We need to figure out what will be
- 8:34:03H2. So, we'll multiply the weight matrix
- 8:34:05WI with X of two plus WR into H of one
- 8:34:08that we have got here plus B of H and
- 8:34:11pass it through a function G of H to get
- 8:34:13H of two. From H of two, we'll calculate
- 8:34:15Y of two. WY into H of two plus BY, that
- 8:34:18is the bias, pass it through a function
- 8:34:20G of Y to get Y2. And this is how
- 8:34:22recurrent neural network works, guys.
- 8:34:24Now, you must be thinking how to train a
- 8:34:26recurrent neural network.
- 8:34:28So, a recurrent neural network uses back
- 8:34:29propagation algorithm for training. But
- 8:34:31back propagation happens for every time
- 8:34:34stamp. That is why it is commonly called
- 8:34:36as back propagation through time.
- 8:34:38Over here, I won't be discussing back
- 8:34:39propagation in detail. I'll just give
- 8:34:41you a brief introduction of what it is.
- 8:34:44Now, with back propagation, there are
- 8:34:45certain issues, namely vanishing and
- 8:34:47exploding gradients. Let us see those
- 8:34:49one by one.
- 8:34:50So, in vanishing gradient, what happens?
- 8:34:52When you use back propagation, you tend
- 8:34:54to calculate the error, which is nothing
- 8:34:56but the actual output that you already
- 8:34:58know minus the model output, output that
- 8:35:01you got through your model, and the
- 8:35:02square of that.
- 8:35:03So, you figure out the error. With that
- 8:35:05error, what do you do? You tend to find
- 8:35:08out the change in error with respect to
- 8:35:10change in weight or any variable. So,
- 8:35:13we'll call it weight here. So, change of
- 8:35:15error with respect to weight multiplied
- 8:35:17by learning rate will give you the
- 8:35:18change in weight. Then you need to add
- 8:35:20that change in weight to the old weight
- 8:35:22to get the new weight. All right? So,
- 8:35:25obviously, what we are trying to do, we
- 8:35:26are trying to reduce the error. So, for
- 8:35:28that, we need to figure out what will be
- 8:35:30the change in error if my variables are
- 8:35:32changed, right? So, that way we can get
- 8:35:34the change in in variable and add it to
- 8:35:36our old variable to get the new
- 8:35:37variable. Now, over here, what can
- 8:35:39happen if the value dE by dW, that is a
- 8:35:42gradient, or you can say the rate of
- 8:35:44change of error with respect to our
- 8:35:45variable weight, becomes very small than
- 8:35:47one, like it is 0.00 something. So, if
- 8:35:50you multiply that with the a learning
- 8:35:52rate, which is definitely smaller than
- 8:35:54one, then you get the change of weight,
- 8:35:56which is negligible. All right? So,
- 8:35:59there might be certain examples where,
- 8:36:00you know, you are trying to predict,
- 8:36:02say, a next word in a sentence, and that
- 8:36:03sentence is pretty long. For example, if
- 8:36:05I say, "I went to France {dash} {dash}
- 8:36:08{dash} I went to France." Then there are
- 8:36:10certain words. Then I say, "Few of them
- 8:36:13speak {dash}." Now, I need to predict
- 8:36:15speak, what will come after speak. So,
- 8:36:18for that, I need to go back in time and
- 8:36:20check what was the context, which will
- 8:36:22be very complex. And due to that,
- 8:36:24there'll be a lot of iterations. And
- 8:36:26because of that, this error, this change
- 8:36:28in weight, will become very small, very
- 8:36:31small. So, the new weight that we'll get
- 8:36:32will be actually almost equal to your
- 8:36:35old weight. So, there won't be any
- 8:36:37updation of weight that will be
- 8:36:38happening. And that is nothing but your
- 8:36:40vanishing gradient. All right, I'll
- 8:36:42repeat it once more. So, what happens in
- 8:36:44back propagation, you first calculate
- 8:36:46the error. This error is nothing but the
- 8:36:48difference between the actual output and
- 8:36:49the model output and the square of that.
- 8:36:52With that error, we figure out what will
- 8:36:53be the change in error when we change a
- 8:36:55particular variable, say, weight. So, dE
- 8:36:58by dW, multiply it with learning rate to
- 8:37:00get the change in the variable or change
- 8:37:02in the weight. Now, we'll add that
- 8:37:03change in the weight to our old weight
- 8:37:05to get the new weight. This is back
- 8:37:07propagation, is guys, all right? I'm
- 8:37:08just giving you a small introduction to
- 8:37:10back propagation. Now, consider a
- 8:37:12scenario where you need to predict the
- 8:37:14next word in a sentence. And your
- 8:37:15sentence is something like this. "I have
- 8:37:18been to France." Then there are a lot of
- 8:37:20words. After that, few people speak. And
- 8:37:24then you need to predict what comes
- 8:37:25after speak. Now, if I need to do that,
- 8:37:27I need to go back and understand the
- 8:37:29context, what is it talking about?
- 8:37:32And that is nothing but your long-term
- 8:37:34dependencies. So, what happens during
- 8:37:35long-term dependencies if this DE by DW
- 8:37:38becomes very small? Then, when you
- 8:37:40multiply it with N, which is again
- 8:37:41smaller than one, you get delta W, which
- 8:37:44will be very, very small. That will be
- 8:37:46negligible. So, the new weight that
- 8:37:48you'll get here will be almost equal to
- 8:37:50your old weight. So, I hope you're
- 8:37:52getting my point. So, this new weight
- 8:37:54So, there will be no updation of
- 8:37:56weights, guys. This new weight will
- 8:37:58definitely be will always be almost
- 8:38:00equal to our old weight. There won't be
- 8:38:02any learning here. So, that is nothing
- 8:38:04but your vanishing gradient problem.
- 8:38:06Similarly, when I talk about exploding
- 8:38:08gradient, it is just the opposite of
- 8:38:09vanishing gradient. So, what happens
- 8:38:11when your gradient or DE by DW becomes
- 8:38:13very uh large, becomes greater than
- 8:38:15greater than one? All right? And you
- 8:38:17have some long-term dependencies. So, at
- 8:38:19that time, your DE by DW will keep on
- 8:38:22increasing. Delta W will become large.
- 8:38:24And because of that, your weights, the
- 8:38:26new weight with that will come will be
- 8:38:28very different from your old weight. So,
- 8:38:30these two are the problems with back
- 8:38:31propagation. Now, let us see how to
- 8:38:33solve these problems.
- 8:38:35Now, exploding gradients can be solved
- 8:38:37with the help of truncated BPTT, back
- 8:38:38propagation through time. So, instead of
- 8:38:40starting back propagation at the last
- 8:38:42time stamp, we can choose a smaller time
- 8:38:44stamp like 10. Or we can clip the
- 8:38:47gradients at a threshold. So, there can
- 8:38:48be a threshold value where we can, you
- 8:38:50know, clip the gradients. And we can
- 8:38:52adjust the learning rate as well. Now,
- 8:38:53for vanishing gradient, we can use a
- 8:38:55ReLU activation function. We have
- 8:38:56discussed ReLU activation function in
- 8:38:58artificial neural network tutorial,
- 8:38:59guys. Similarly, we can also use LSTM
- 8:39:02and GRUs. In this tutorial, we'll be
- 8:39:04discussing LSTMs that are long
- 8:39:06short-term memory units. Now, let us
- 8:39:09understand what exactly are LSTMs.
- 8:39:12So, guys, we saw what are the two
- 8:39:13limitations with the recurrent neural
- 8:39:15networks. Now, we'll understand how we
- 8:39:17can solve that with the help of LSTMs.
- 8:39:19Now, what are LSTMs? Long short-term
- 8:39:21memory networks, usually called as
- 8:39:23LSTMs, are nothing but a special kind of
- 8:39:25recurrent neural network. And these
- 8:39:27recurrent neural networks are capable of
- 8:39:29learning long-term dependencies. Now,
- 8:39:31what are long-term dependencies? I've
- 8:39:33discussed on the previous slide, but
- 8:39:35I'll just explain it to you here as
- 8:39:36well. Now, what happens sometimes we
- 8:39:38only need to look at the recent
- 8:39:39information to perform the present task.
- 8:39:42Now, let me give you an example.
- 8:39:43Consider a language model trying to
- 8:39:45predict the next word based on the
- 8:39:47previous ones. If we are trying to
- 8:39:49predict the last word in the sentence,
- 8:39:51say, "The clouds are in the sky." So, we
- 8:39:54don't need any further context. It's
- 8:39:55pretty obvious that the next word is
- 8:39:57going to be sky. Now, in such cases
- 8:39:59where the gap between the relevant
- 8:40:01information and the place that it's
- 8:40:03needed is small, RNNs can learn to use
- 8:40:06the past information. And at that time,
- 8:40:08there won't be such problems like
- 8:40:09vanishing and exploding gradient. But,
- 8:40:11there are few cases where we need more
- 8:40:14context. Consider trying to predict the
- 8:40:16last word in the text, "I grew up in
- 8:40:19France." Then, there are some words.
- 8:40:20After that comes, "I speak fluent
- 8:40:23French." Now, recent information
- 8:40:25suggests that word is probably the name
- 8:40:27of a language. But, if we want to narrow
- 8:40:29down which language, we need the context
- 8:40:32of France from further back. And it's
- 8:40:35entirely possible for the gap between
- 8:40:37the relevant information and the point
- 8:40:39where it is needed to become very large.
- 8:40:41And this is nothing but long-term
- 8:40:42dependencies. And the LSTMs are capable
- 8:40:45of handling such long-term dependencies.
- 8:40:47Now, LSTMs also have a chain-like
- 8:40:50structure like recurrent neural
- 8:40:51networks. Now, all the recurrent neural
- 8:40:53networks have the form of a chain of
- 8:40:54repeating modules of neural networks.
- 8:40:56Now, in standard RNNs, the repeating
- 8:40:58module will have a very simple structure
- 8:41:00such as a single tan h layer that you
- 8:41:01can see. Now, this tan h layer is
- 8:41:03nothing but a squashing function. Now,
- 8:41:05what I mean by squashing function is to
- 8:41:07convert my values between minus one and
- 8:41:10one. All right, that's why we use tan h.
- 8:41:12And this is an example of an RNN. Now,
- 8:41:15we'll understand what exactly are LSTMs.
- 8:41:17Now, this is a structure of an LSTM. All
- 8:41:20If you notice, LSTM also have a chain
- 8:41:22like structure. But the ripple has
- 8:41:24different structures. Instead of having
- 8:41:26single neural network here, there are
- 8:41:28four interacting in a very special way.
- 8:41:30Now, the key to LSTM is the cell state.
- 8:41:32Now, this particular line that I'm
- 8:41:34highlighting, this is what what is
- 8:41:36called the cell state. The horizontal
- 8:41:38line running through the top of the
- 8:41:39diagram. So, this is nothing but your
- 8:41:40cell state. Now, you can consider the
- 8:41:42cell state as a kind of a conveyor belt.
- 8:41:45It runs straight down the entire chain
- 8:41:47with only some minor linear
- 8:41:48interactions. Now, what I'll do, I'll
- 8:41:50give you a walk through of LSTM step by
- 8:41:52step, all right? So, we'll start with
- 8:41:54the first step.
- 8:41:55All right, guys. So, the first step in
- 8:41:57our LSTM is to decide what information
- 8:41:59we are going to throw away from the cell
- 8:42:01state. And you know what is the cell
- 8:42:03state, right? I've discussed in the
- 8:42:04previous slide. Now, this decision is
- 8:42:06made by the sigmoid layer. So, the layer
- 8:42:09that I'm highlighting with my cursor, it
- 8:42:10is the sigmoid layer. Called the forget
- 8:42:12gate layer. It looks at HT minus one,
- 8:42:15that is the information from the
- 8:42:16previous time step, and XT, which is the
- 8:42:19new input, and outputs a number between
- 8:42:21zeros and ones for each number in the
- 8:42:23cell state, CT minus one, which is
- 8:42:25coming from the previous time step. A
- 8:42:27one represents completely keep this,
- 8:42:29while a zero represents completely get
- 8:42:31rid of this. Now, if we go back to our
- 8:42:33example of a language model trying to
- 8:42:35predict the next word based on all the
- 8:42:37previous ones, in such a problem, the
- 8:42:39cell state might include the gender of
- 8:42:41the present subject so that the correct
- 8:42:43pronouns can be used. When we see a new
- 8:42:45subject, we want to forget the gender of
- 8:42:47the old subject, right? We want to use
- 8:42:50the gender of the new subject. So, we'll
- 8:42:52forget the gender of the previous
- 8:42:53subject here. This is just an example to
- 8:42:56explain you what is happening here.
- 8:42:58Uh now, let me explain you the equations
- 8:42:59which I've written here. So, FT will be
- 8:43:02uh combining with the cell state later
- 8:43:04on, that I'll tell you. So, currently,
- 8:43:06FT will be nothing but the weight matrix
- 8:43:09multiplied by HT minus one and XT, and
- 8:43:13uh plus the bias, and this equation is
- 8:43:15passed through a sigmoid layer. All
- 8:43:17right? And we get an output that is zero
- 8:43:19and one. Zero means completely get rid
- 8:43:21of this and one means completely keep
- 8:43:23this. All right, so this is what
- 8:43:24basically is happening in the first
- 8:43:26step. Now, let us see what happens in
- 8:43:28the next step. So, the next step is to
- 8:43:30decide what information we are going to
- 8:43:32store. In the previous step, we decided
- 8:43:34what information we are going to keep,
- 8:43:35but here we are going to decide what
- 8:43:37information we are going to store here.
- 8:43:39All right, what new information we are
- 8:43:41going to store in the cell state. Now,
- 8:43:42this has two parts. First, a sigmoid
- 8:43:45layer, this is called a sigmoid layer
- 8:43:46and which is also known as an input gate
- 8:43:48layer, decide which values will update.
- 8:43:51All right, so what values we need to
- 8:43:52update. Then there's also a tan h layer
- 8:43:54that creates a vector of the candidate
- 8:43:56values c bar of t minus one that will be
- 8:44:00added to the state later on. All right,
- 8:44:02so let me explain it to you in a simpler
- 8:44:03terms. So, whatever input that we are
- 8:44:05getting from the previous time stamp and
- 8:44:07the new input, it will be passed through
- 8:44:09a sigmoid function, which will give us i
- 8:44:11of t. All right, and this i of t will be
- 8:44:14multiplied by c t but coming from the
- 8:44:17previous time stamp and the new input
- 8:44:19with that is passed through a tan h that
- 8:44:21will result in c t. And this will be
- 8:44:23later added on to our cell state. In the
- 8:44:25next step, we'll combine these two to
- 8:44:27update the states. Now, let me explain
- 8:44:28the equations. So, i of t will be what?
- 8:44:31Weight matrix and then we have h t minus
- 8:44:33one comma x t multiplied by the weight
- 8:44:35matrix plus the bias pass it through a
- 8:44:37sigmoid function, we get i of t. c bar
- 8:44:39of t will get by passing a weight matrix
- 8:44:41h t minus one x t plus bias through a
- 8:44:44tan h square function and we'll get c
- 8:44:46bar of t. All right, so as I've told you
- 8:44:48earlier as well in the next step, we'll
- 8:44:49combine these two to update the state.
- 8:44:51Let us see how we do that. So, now is
- 8:44:54the time to update the old cell state c
- 8:44:56t minus one with the new cell state c t.
- 8:44:59All right, in the previous steps, we
- 8:45:00have already decided what to do. We just
- 8:45:02need to actually do it. So, what we'll
- 8:45:04do, we'll multiply the old cell state c
- 8:45:06t minus one with f t that we got in the
- 8:45:08first step for getting the things that
- 8:45:10we decided to forget earlier in the
- 8:45:12first step if you can recall. Then what
- 8:45:14we do, we add it to IT and CT. Then we
- 8:45:18add it by the term that will come after
- 8:45:20multiplication of IT and C bar T. And
- 8:45:22this new candidate value scaled by how
- 8:45:24much we decided to update each state
- 8:45:26value. All right? So, in the case of the
- 8:45:29language model that we are discussing,
- 8:45:30this is where we would actually drop the
- 8:45:32information about the old subject gender
- 8:45:35and add the new information as we
- 8:45:36decided in the previous steps. So, I
- 8:45:38hope you are able to follow me guys. All
- 8:45:40right? So, let us move forward and we'll
- 8:45:42see what is the next step. Now, our last
- 8:45:44step is to decide what we are going to
- 8:45:46output. And this output will depend on
- 8:45:48our cell state, but [snorts] it will be
- 8:45:50a filtered version. Now, finally what we
- 8:45:52need to do is we need to decide what we
- 8:45:53are going to output. And this output
- 8:45:55will be based on our cell state. First,
- 8:45:57we need to pass HT minus one and XT
- 8:45:59through a sigmoid activation function so
- 8:46:02that we get output that is OT. All
- 8:46:04right? And this OT will be in turn
- 8:46:06multiplied by the cell state after
- 8:46:08passing it through an NH squashing
- 8:46:10function or an activation function. And
- 8:46:12why we do that? Just to push the values
- 8:46:14between minus one and one. So, after
- 8:46:17multiplying OT, that is this value, and
- 8:46:20a tan at CT, we'll get the output H2,
- 8:46:23which will be our new output. And that
- 8:46:25will only output the part that we
- 8:46:27decided to. Whatever we have decided in
- 8:46:29the previous steps, it will only output
- 8:46:30that value. All right? Now, I'll take
- 8:46:32the example of that language model
- 8:46:34again. Since it just saw a subject, it
- 8:46:36might want to output information
- 8:46:38relevant to a verb and in case that's
- 8:46:40what is coming next.
- 8:46:42For example, it might output whether the
- 8:46:44subject is singular or plural. So, that
- 8:46:46we know what form of a verb should be
- 8:46:48conjugated into. All right? And uh you
- 8:46:51can see from the uh you can see the
- 8:46:52equations as well. Again, we have a
- 8:46:54sigmoid function. Then that uh whatever
- 8:46:57output we get from there, we multiply it
- 8:46:59with tan at CT to get the new output.
- 8:47:01All right, guys? So, this is basically
- 8:47:03uh LSTMs in a nutshell. So, in the first
- 8:47:06step, we decided what we need to forget.
- 8:47:08In the next step, we decided what are we
- 8:47:10going
- 8:47:11to our cell state, what new information
- 8:47:13going to add to cell state, and we were
- 8:47:15taking example of the gender throughout
- 8:47:17this whole process. All right? And in
- 8:47:18the third step, what we do, we actually
- 8:47:20combined it to get the new cell state.
- 8:47:22Now, in the fourth step, what we did, we
- 8:47:24finally got the output that we want. And
- 8:47:27how we did that? Just by passing HT - 1
- 8:47:29and HT through a sigmoid function,
- 8:47:31multiplying it with the tan H CT, the
- 8:47:33tan H new cell state, and we get the new
- 8:47:36output. Fine, guys? So, this is what
- 8:47:38basically LSTM is, guys. Now, we'll look
- 8:47:40at a use case where we'll be using LSTM
- 8:47:43to predict the next word in a sentence.
- 8:47:45All right? Let me show you how we are
- 8:47:46going to do that.
- 8:47:48So, this is what we are trying to do in
- 8:47:49our use case, guys. We'll feed LSTM with
- 8:47:52correct sequences from the text of three
- 8:47:54symbols. For example, had a general and
- 8:47:57a label that is counsel in this
- 8:47:59particular example. Eventually, our
- 8:48:01network will learn to predict the next
- 8:48:03symbol correctly. So, obviously, we need
- 8:48:05to train it on something. Let us see
- 8:48:06what we are going to train it on.
- 8:48:08So, we'll be training LSTM to predict
- 8:48:10the next word using a sample short story
- 8:48:12that you can see over here.
- 8:48:14All right? So, it has basically 112
- 8:48:16unique symbols. So, even comma and full
- 8:48:18stop are considered as symbols. All
- 8:48:20right? So, this is what we are going to
- 8:48:22train it on.
- 8:48:23So, technically, we know that LSTMs can
- 8:48:25only understand real numbers. All right?
- 8:48:27So, what we need to do is we need to
- 8:48:29convert these unique symbols into a
- 8:48:31unique integer value based on the
- 8:48:33frequency of occurrence. And like that,
- 8:48:35we'll create a dictionary. For example,
- 8:48:37we have had here that will have value
- 8:48:3920. A will have value six. General will
- 8:48:42have value 33. All right? And then, what
- 8:48:45happens, our LSTM will create a 112
- 8:48:48element vector that will contain the
- 8:48:50probability of each of these words or
- 8:48:53each of these unique integer values. All
- 8:48:55right? So, since 0.6 has the highest
- 8:48:57probability in this particular vector,
- 8:48:59it'll pick the index value of 0.6. Then,
- 8:49:02it will see it what symbol is attached
- 8:49:04to that particular integer value. So, 37
- 8:49:06is attached to counsel. So, this will be
- 8:49:08our prediction, which is absolutely
- 8:49:09correct as the label is also counsel
- 8:49:11according to our training data. All
- 8:49:13right. So, this is what we are going to
- 8:49:15do in our use case. So, guys, this is
- 8:49:17what we'll be doing in our today's use
- 8:49:18case. Now, I'll quickly open my PyCharm
- 8:49:21and I'll show you how you can implement
- 8:49:22it using Python. We'll be using
- 8:49:24TensorFlow, which is a popular Python
- 8:49:26library for implementing deep neural
- 8:49:28networks or neural networks in general.
- 8:49:30All right. So, I'll quickly open my
- 8:49:31PyCharm now. So, guys, this is my
- 8:49:33PyCharm and I over here I've already
- 8:49:35written the code in order to execute the
- 8:49:36use case that we have. So, first we need
- 8:49:38to do is import the libraries, NumPy for
- 8:49:41arrays, TensorFlow we know,
- 8:49:42tensorflow.contrib from that we need to
- 8:49:44import RNN and random collections and
- 8:49:46time. All right. So, this particular
- 8:49:48block of code is used to evaluate the
- 8:49:51time taken for the training. After that,
- 8:49:53we have log_path and this log_path is
- 8:49:56basically telling us the path where the
- 8:49:58graph will be stored. All right. So,
- 8:49:59there will be a graph that will be
- 8:50:00created and then that graph will be
- 8:50:02launched. Then only our RNN model will
- 8:50:04be executed. Then that's how TensorFlow
- 8:50:06works, guys.
- 8:50:07So, that graph will be created in this
- 8:50:09particular path. All right. And we are
- 8:50:11using summary writer. So, that will
- 8:50:13actually create the log file that will
- 8:50:14be used in order to display the graph
- 8:50:17using TensorBoard. All right. So, then
- 8:50:19we have defined training_file, which
- 8:50:21will have our story on which we'll train
- 8:50:23our model on. Then what we need to do is
- 8:50:25read this file. So, how are we going to
- 8:50:27do that? First is read line by line
- 8:50:29whatever content that we have in our
- 8:50:31file. Then we are going to strip it.
- 8:50:33That means we are going to remove the
- 8:50:35first and the last white space. Then
- 8:50:37again, we are splitting it just to
- 8:50:40remove all the white spaces that are
- 8:50:41there. After that, we're creating an
- 8:50:43array and then we're reshaping it. Now,
- 8:50:45in during the reshape, if you notice
- 8:50:46this minus one value tells us the
- 8:50:48compatibility. All right. So, when
- 8:50:50you're reshaping it, you need to make
- 8:50:51sure that
- 8:50:53you know, we are providing in the
- 8:50:54correct parameters to reshape it. So,
- 8:50:56you can convert a three cross two matrix
- 8:50:58to a two cross three matrix, like right?
- 8:51:01So, just to make sure that that it is
- 8:51:02compatible enough, we add this minus one
- 8:51:04and it'll be done automatically. All
- 8:51:06right? Then, return content. After that,
- 8:51:09what we are doing, we are feeding in the
- 8:51:11training data that we have, training
- 8:51:12{underscore} file. We are feeding in our
- 8:51:14story and calling the function read
- 8:51:16{underscore} data. Then, what we are
- 8:51:17doing, we are creating a dictionary.
- 8:51:19What is a dictionary? We all know, key
- 8:51:20value pairs based on the frequency of
- 8:51:22occurrences of each symbol. All right?
- 8:51:24So, from here, collections.counter
- 8:51:26words.most_common. So, most common words
- 8:51:29with their frequency of occurrence,
- 8:51:30there'll be a dictionary created. And
- 8:51:32after that, uh we'll call this dict
- 8:51:34function and this dict function will
- 8:51:36feed in word and which is equal to
- 8:51:38length of dictionary. That means
- 8:51:40whatever the length of that particular
- 8:51:42dictionary, how many time it is
- 8:51:43repeated. So, we'll have the frequency
- 8:51:45as well as a symbol. That'll be our key
- 8:51:47value pair and we're reversing it as
- 8:51:49well.
- 8:51:50Then, what we are doing, we are calling
- 8:51:51it build {underscore} data set and we're
- 8:51:54feeding in our training data there. This
- 8:51:56is our vocabulary size, which is nothing
- 8:51:57but the length of your dictionary. Then,
- 8:51:59we have defined various parameters such
- 8:52:01as learning rate, uh iterations or
- 8:52:03epochs. Then, we have display step and
- 8:52:05{underscore} input. Now, learning rate,
- 8:52:07we all know what it is, uh the steps in
- 8:52:09which our variables are updated.
- 8:52:11Training {underscore} iterations is
- 8:52:12nothing but your epochs, the total
- 8:52:14number of iterations. So, we have given
- 8:52:1550,000 iterations here. Then, we have
- 8:52:17display {underscore} step, that is
- 8:52:191,000, which is basically your batch
- 8:52:20size. So, batch size is what? After
- 8:52:23every 1,000 epochs, you'll see the
- 8:52:24output. All right? So, it'll be
- 8:52:25processing it in batches of 1,000
- 8:52:27iterations. Then, we have n {underscore}
- 8:52:29input as three. Now, the number of units
- 8:52:31in the RNN cell, we'll keep it as 512.
- 8:52:34Then, we need to define X and Y. So, X
- 8:52:37will be our placeholder that will have
- 8:52:38the input values and Y will have all the
- 8:52:41labels. All right? vocab size.
- 8:52:44So, X is a placeholder where we'll be
- 8:52:46feeding in our input dictionary.
- 8:52:47Similarly, Y is also one more
- 8:52:49placeholder and it'll have a shape of
- 8:52:51none {comma} vocab size. Vocab size we
- 8:52:53have defined earlier.
- 8:52:55As you can see, which is nothing but the
- 8:52:56length of your dictionary. Then we're
- 8:52:57defining weights as well as biases.
- 8:53:00After that, we have defined our model.
- 8:53:02All right. So, this is how we are going
- 8:53:03to define it. We'll
- 8:53:05create a function RNN when we'll have X
- 8:53:08weights and biases. And after that, we
- 8:53:10are calling in RNN.multi_rnn_cell
- 8:53:12function. And this is basically to
- 8:53:14create a two-layer LSTM. And each layer
- 8:53:17has n_hidden_units.
- 8:53:19After that, what we are doing, we are
- 8:53:20generating the predictions. But once we
- 8:53:22have generated the prediction, there are
- 8:53:24n_input_outputs,
- 8:53:25but we only want the last output. For
- 8:53:28that, we have written this particular
- 8:53:29line. And then finally, we are making a
- 8:53:30prediction. We are calling this RNN and
- 8:53:32function feeding in X weights and
- 8:53:34biases. After that, we are calculating
- 8:53:36the loss as and then we are optimizing
- 8:53:38it. For calculating the loss, we are
- 8:53:41using reduce_mean softmax_cross_entropy.
- 8:53:44And this will give us basically the
- 8:53:46probability of each symbol. And then we
- 8:53:48are optimizing it using RMS uh prop
- 8:53:50optimizer. All right. And this gives
- 8:53:52actually a better accuracy than Adam
- 8:53:54optimizer. And that's the reason why we
- 8:53:56are using it. Then we are going to
- 8:53:57calculate the accuracy. And after that,
- 8:54:00we are going to initialize the variables
- 8:54:01that we have used. As we have seen in
- 8:54:03TensorFlow, that we need to initialize
- 8:54:04all the variables, unlike constants and
- 8:54:06placeholders in TensorFlow. All right.
- 8:54:08And once we are done with that, we are
- 8:54:10feeding in our values, then calculating
- 8:54:12the accuracy, how accurate it is. And
- 8:54:14then when optimization is done, we are
- 8:54:16calculating the elapsed time as well.
- 8:54:18So, that will give us how much time it
- 8:54:20took in order to train our model. Then
- 8:54:22this is just to run the TensorBoard on
- 8:54:24our local host 6006. And yeah, and this
- 8:54:28particular block of code is is used in
- 8:54:30order to handle the exceptions. So,
- 8:54:32exceptions can be like whatever word
- 8:54:34that we are putting in might not be
- 8:54:36there in our dictionary or might not be
- 8:54:37there in our training data. So, those
- 8:54:39exceptions will be handled here. And if
- 8:54:41it is not there in our dictionary, then
- 8:54:42it will print word not in our
- 8:54:44dictionary. All right. So, fine guys.
- 8:54:46Let's
- 8:54:47input some values and we'll have some
- 8:54:49fun with this model. All right? So, the
- 8:54:51first thing that I'm going to feed in is
- 8:54:53had general. So, whenever I feed in
- 8:54:56these three values, had a general,
- 8:54:58there'll be a story that will be
- 8:54:59generated by feeding back the predicted
- 8:55:01output as the next symbol in the inputs.
- 8:55:04All right? So, when I feed in had a
- 8:55:05general, so it'll predict the correct
- 8:55:07output as counsel. And this counsel will
- 8:55:10be fed back as a part of the new input
- 8:55:12and our new input will be a general
- 8:55:14counsel. So, it'll be a general counsel.
- 8:55:16All right? So, these three words will
- 8:55:18become our new input to predict the new
- 8:55:19output, which is two. All right? And so
- 8:55:21on. So, surprisingly, LSTM actually
- 8:55:24creates a story that, you know, somehow
- 8:55:26makes sense. So, let's just read it. Had
- 8:55:28a general counsel to consider what
- 8:55:30measures they could take to outwit their
- 8:55:32common enemy, the cat. By this means, we
- 8:55:35should always know when she was about
- 8:55:37and could easily. All right? So, somehow
- 8:55:39it actually makes sense when you feed in
- 8:55:40that. So, what'll happen when you feed
- 8:55:42in these three inputs, it'll predict the
- 8:55:44next word, that is counsel. After that,
- 8:55:46it'll take counsel and it'll feed back
- 8:55:48as an input along with a general. So, a
- 8:55:50general counsel will be your next input
- 8:55:53to predict two. Similarly, in the next
- 8:55:55iteration, it'll take general counsel
- 8:55:57two and predict counsel for us. And this
- 8:55:59will keep on repeating.
- 8:56:02>> [music]
- 8:56:09>> processing and why do we even need it?
- 8:56:12You see, natural language
- 8:56:13analysis in both audible data as well as
- 8:56:16the text document. NLP system can
- 8:56:19capture meaning from an input such as
- 8:56:21sentences, paragraphs, pages, and give
- 8:56:23out a desired output based on our
- 8:56:25application. So, why do we need NLP? You
- 8:56:28see, natural language processing helps
- 8:56:29computer communicate with humans in
- 8:56:31their own language and scale other
- 8:56:33language-related task. For example, NLP
- 8:56:36makes it possible for computers to read
- 8:56:38text, hear speeches, interpret it,
- 8:56:41measure the sentiment, and then
- 8:56:42determine which part of it are
- 8:56:44important. Today's machines can analyze
- 8:56:46more language-based data than humans.
- 8:56:48That too with consistency, accuracy, and
- 8:56:51in an unbiased manner. More or less, we
- 8:56:53all know that there is a staggering
- 8:56:55amount of unstructured data that is
- 8:56:57generated every day. Be it from a
- 8:56:59medical record or to a social media.
- 8:57:01Automating NLP task will critically be
- 8:57:04helpful in future for analyzing text and
- 8:57:06speech data efficiently. So, moving
- 8:57:09ahead, let us now see the ways we can
- 8:57:10process our textual data. We can process
- 8:57:13our textual data in one of two ways. One
- 8:57:15is a machine learning way, and other one
- 8:57:17is a deep learning method. In machine
- 8:57:19learning, we can make use of algorithms
- 8:57:21such as bag of words, TF-IDF to classify
- 8:57:24and predict the desired output. But, the
- 8:57:26drawback of this is that these machine
- 8:57:28learning algorithms do not consider the
- 8:57:30context of the word or a sequence. Here,
- 8:57:32the way it works is on the base on
- 8:57:34number of times a word is repeating, a
- 8:57:36probability is derived out of it, and
- 8:57:38then performs a classification task.
- 8:57:41This is the reason why we have deep
- 8:57:42learning model for NLP tasks. Speaking
- 8:57:45about deep learning model, we have
- 8:57:46something like recurrent neural network,
- 8:57:48LSTM, transformer network, Google's BERT
- 8:57:51algorithm, and many more. This deep
- 8:57:53learning model learns the pattern of a
- 8:57:55word or a sequence, and then tries to
- 8:57:57predict the desired outcome of the task.
- 8:57:59We can perform NLP using deep learning
- 8:58:01in one of two ways. One by pre-trained
- 8:58:03models such as Google's Word2vec or
- 8:58:06global vector models. And the other way
- 8:58:08is to train our own model. If you're
- 8:58:10trying to train our own model, it would
- 8:58:12require a very huge amount of data and
- 8:58:14also a compute power to support it. In
- 8:58:16most of the cases, we'll be using
- 8:58:18pre-trained models. Moving ahead, let us
- 8:58:20now discuss recurrent neural networks.
- 8:58:23As I mentioned earlier, the bag of word
- 8:58:25or TF-IDF model for processing our text
- 8:58:28is very inefficient. As I mentioned
- 8:58:30earlier, the bag of word or TF-IDF model
- 8:58:32that was used in machine learning to
- 8:58:34process our textual data is very
- 8:58:36inefficient as it takes one word at a
- 8:58:38time and also the context of the word in
- 8:58:41which it is being spoken about is
- 8:58:42totally ignored. Although this would
- 8:58:44give us some prediction, but we can
- 8:58:45expect lot of loss. This is why we use
- 8:58:48RNN model or recurrent neural network.
- 8:58:50Here it requires a sequential data.
- 8:58:53Now you might be wondering what does
- 8:58:54this sequential data mean, right? In
- 8:58:56simple words, sequential data is
- 8:58:58dependent on the past value. What I'm
- 8:59:00trying to say here is that we can read
- 8:59:02our document, right? We can read our
- 8:59:04document or English documents only from
- 8:59:06left to right side. Or take example of
- 8:59:08stock price prediction. We cannot
- 8:59:10randomly place the dates, right? If you
- 8:59:12have to make a prediction for next 2-3
- 8:59:14months, we'll obviously refer to the
- 8:59:16data that was previously recorded. So
- 8:59:18this is what a sequential data means. So
- 8:59:21what makes RNN capable of handling
- 8:59:23sequential data? Well, you see RNN model
- 8:59:26makes use of something called a state.
- 8:59:28This is nothing but a temporary memory
- 8:59:30that stores the previous data. As you
- 8:59:32can see here in an image, this is the
- 8:59:34general architecture of a recurrent
- 8:59:36neural network. So let me now move to my
- 8:59:38canvas and show you how recurrent neural
- 8:59:40network works and what are its internal
- 8:59:42workings.
- 8:59:44All right. So as we have seen in an
- 8:59:45image, right? Back in our slide, you saw
- 8:59:47that, you know, recurrent neural network
- 8:59:49have something called as states and then
- 8:59:51we also had some boxes, right? So what
- 8:59:54does this boxes represents? So let me
- 8:59:56quickly draw over here and show you what
- 8:59:57does this box represents. So if you
- 9:00:00remember, right? It goes something like
- 9:00:02we have a box here. Okay? And then we
- 9:00:04also had another box.
- 9:00:07And then let's consider like two more
- 9:00:09boxes that would be sufficient. Okay?
- 9:00:11And then the way we provide input for
- 9:00:13our recurrent neural network is over
- 9:00:15here.
- 9:00:16Okay? Now based on our application, we
- 9:00:18can demand our recurrent neural network
- 9:00:20to provide output in either one of three
- 9:00:22ways. It can either be like we can have
- 9:00:24multiple inputs, as you can see here, or
- 9:00:26we can have single input and multiple
- 9:00:28outputs, or the other way around is we
- 9:00:30can have single input and single output.
- 9:00:32So here what we'll consider is we are
- 9:00:34using multiple inputs and also we have
- 9:00:36multiple outputs.
- 9:00:38Okay. So now let us see as this is a
- 9:00:40supervised learning model, right?
- 9:00:42Obviously it will have some kind of
- 9:00:43input and then it will also have labels
- 9:00:45to clarify the data. So let's take
- 9:00:48something like if this is our X data,
- 9:00:49right? So if this is the X data, so
- 9:00:52let's take this as a list and then over
- 9:00:53here we'll have data which would
- 9:00:55represent something like X1, then we'll
- 9:00:57have X2,
- 9:00:59then we'll have X3,
- 9:01:00and then we'll have something like Xn.
- 9:01:04Okay? So what does this X over here
- 9:01:07represents? See, here let's take an
- 9:01:09example that Uri, which was a movie, is
- 9:01:12a good movie. We can have couple more
- 9:01:13X's and we'll just put it here as good
- 9:01:16movie.
- 9:01:27Okay. So this is nothing but the inputs.
- 9:01:30Each of these words. Now what our model
- 9:01:32over here does, let's take an example
- 9:01:34that we are trying to find name entity
- 9:01:36prediction, which stands for NER, right?
- 9:01:38So what does name entity prediction does
- 9:01:40is, you know, trains the model in such a
- 9:01:42manner that it finds a pattern. So as
- 9:01:44you can see here, we this is Uri, right?
- 9:01:47Uri, it's it's something like a name,
- 9:01:48right? So this this goes as one.
- 9:01:50And is is not a name. So this would be
- 9:01:52zero. Good is this is also not a name
- 9:01:55and movie is also not a name. So the
- 9:01:57output over here would be something like
- 9:01:581000. All right? So now let's see what
- 9:02:01are the inputs and how they would look
- 9:02:03like. So first off we have inputs. So if
- 9:02:06this is our X data, so we'll have input
- 9:02:08something like Xi. This should be in
- 9:02:10lower case.
- 9:02:11Okay? So this would be X1,
- 9:02:13then we'll have X2,
- 9:02:15we'll have X3, and then we'll have X4.
- 9:02:19Similarly, the Y the output over here
- 9:02:21would be something like we'll give it Y
- 9:02:23hat of 1, [snorts]
- 9:02:25Y hat of 2,
- 9:02:27Y hat off three, at the same time we'll
- 9:02:29also have why hat off four. Now if
- 9:02:31you're wondering what does this why hat
- 9:02:33represents, why hat over here is nothing
- 9:02:34but
- 9:02:35predicted values.
- 9:02:39And why is nothing but you know the
- 9:02:41labeled value or trained values.
- 9:02:45And now as we all know that the main
- 9:02:48thing or the main feature behind
- 9:02:50recurrent neural network is nothing but
- 9:02:52the states. And the way the states are
- 9:02:54represented is by using the state vector
- 9:02:56or it can also be called as context
- 9:02:57vector. So the state vector over here
- 9:02:59starts with A. Let me give it as a
- 9:03:01different color here. Let me give it as
- 9:03:02blue. So here we'll have A and this
- 9:03:05should be zero. Okay? And now this A
- 9:03:08would be passed down to this. Okay, so
- 9:03:10this would be A of one and over here
- 9:03:12we'll have A of two,
- 9:03:14A of three and then finally A of four.
- 9:03:18I'm pretty sure you might be wondering
- 9:03:19what does this A contains, right? So
- 9:03:21over here as I have mentioned, let me
- 9:03:23take this very example. So here Uri is
- 9:03:26good movie, okay? And so what does A1
- 9:03:29will contain? A1 will contain Uri. See,
- 9:03:33if we take the normal algorithm, what's
- 9:03:34going to happen is it will just consider
- 9:03:36only this one particular block. Okay? If
- 9:03:39this was just a normal algorithm or
- 9:03:41something which is used in olden days,
- 9:03:42it will just consider this particular
- 9:03:44block and it won't be considering the
- 9:03:45previous values. So this previous value
- 9:03:48is being stored over here in the form of
- 9:03:50a memory. So now A1 will contain Uri and
- 9:03:53what A2 will contain over here? So A2
- 9:03:57will be having something like Uri is.
- 9:04:00And similarly as it goes through here,
- 9:04:02it will collect each and every word and
- 9:04:04finally A4 will be complete context
- 9:04:06vector. So it will have the entire
- 9:04:07sentence Uri is a good movie.
- 9:04:12Okay, I hope you understood what does
- 9:04:14this A signifies over here. Let me
- 9:04:16quickly erase all of these. Okay, so the
- 9:04:19another important part which goes over
- 9:04:20here in any machine learning or deep
- 9:04:22learning model is nothing but the
- 9:04:24weights.
- 9:04:25So, let's consider that we have weights
- 9:04:27which is nothing but W or let's take it
- 9:04:29as U. Let me give another color here, U,
- 9:04:32V, and W. In this RNN model, right? What
- 9:04:35we're going to do is we're going to have
- 9:04:37a single or we'll have just only one
- 9:04:39weights, right? So, the U weight is same
- 9:04:42for all of these inputs here. V weight
- 9:04:45over here is same for all the outputs.
- 9:04:47The weight W is same for all our state
- 9:04:50metrics. Okay? So, this is how the
- 9:04:52weights are determined.
- 9:04:54So, now to get a better understanding of
- 9:04:56what's happening over here and to derive
- 9:04:58a mathematical equation for feedforward
- 9:05:00network, let's take a single block over
- 9:05:02here. And let's see how this internal
- 9:05:04working of this is working, okay? So,
- 9:05:06over here we'll have a box.
- 9:05:09Okay? And this will be our input. So,
- 9:05:11let's give our input as X of T because
- 9:05:14this is a generalized model, right? And
- 9:05:16now what you're going to do over here is
- 9:05:17this would be A
- 9:05:19of T minus 1. That's because A of T is
- 9:05:22will be present over here. And now if
- 9:05:24you have an output which is present over
- 9:05:26here and this would be nothing but Y hat
- 9:05:29of T.
- 9:05:30All right? So, as I've mentioned earlier
- 9:05:31that we will be having weights. So, the
- 9:05:33weights over here is nothing but U, V,
- 9:05:35and W. So, what's happening over here?
- 9:05:38Okay. So, first off, there will be a
- 9:05:41matrix multiplication between these two
- 9:05:42values.
- 9:05:44Okay? And then there will be a matrix
- 9:05:45multiplication between these two values.
- 9:05:47And once they are done, we'll add these
- 9:05:49two values and then give a activation
- 9:05:51function to it. Okay? And the activation
- 9:05:54function that will be used over here is
- 9:05:55tan H. So, if I have to put it in an
- 9:05:58equation form over here, so the product
- 9:06:00of these two, X of T
- 9:06:02times or it would be a dot product
- 9:06:05of U, okay? And the sum of these two, so
- 9:06:09it will be T minus 1 times W. And now
- 9:06:12what we're going to do is we're going to
- 9:06:13pass an activation function which is
- 9:06:14nothing but tan H over here.
- 9:06:16So, let's just give a small dotted
- 9:06:18notation over here.
- 9:06:20So, whatever the output which comes
- 9:06:22it'll be performing an addition of these
- 9:06:25two and also give an activation function
- 9:06:27which will represent here by f.
- 9:06:29All right? So, this is how we get an
- 9:06:31activation function over here. So, what
- 9:06:33does this value signifies this? This is
- 9:06:35nothing but this output over here, a of
- 9:06:37t. So, I hope you understand what is a
- 9:06:39of t. A of t is this value. Okay, let me
- 9:06:42quickly highlight that for you. So, a of
- 9:06:44t is this value. Okay? So, a of t is
- 9:06:47represented over here. And the way we
- 9:06:49get this is by this particular equation.
- 9:06:52Now, what about the value for y hat of t
- 9:06:54or the prediction of y of t? So, for y
- 9:06:57of t, it's going to be something like
- 9:06:58this. So, y hat of t, this would be
- 9:07:01nothing but we have to multiply this
- 9:07:03particular a of t with this matrix over
- 9:07:06here. Okay? Or the weights I can say.
- 9:07:08So, it's going to be nothing but a of t
- 9:07:11times the weight matrix that is v. And
- 9:07:14then we have to pass activation
- 9:07:16function. Usually, the activation
- 9:07:17function that is going to be used over
- 9:07:19here is softmax or sigmoid. It totally
- 9:07:21depends upon the what kind of output
- 9:07:24you're expecting. All right? And this is
- 9:07:26how it works. And now, an important
- 9:07:28thing that I would like to mention over
- 9:07:30here is that we also have to add
- 9:07:32something called as bias. So, it would
- 9:07:34be b over here. I'll just represented
- 9:07:36this by a different color.
- 9:07:38So, this is an equation for our
- 9:07:40recurrent neural network in a
- 9:07:41mathematical form.
- 9:07:43So, now what's going to happen is in
- 9:07:44order to find the loss, right? We have
- 9:07:46to subtract whatever value we have
- 9:07:47predicted with the given value. So, the
- 9:07:50predicted value over here, let's say
- 9:07:51this is as capital y. Okay? And we also
- 9:07:54have been given the train value. So,
- 9:07:55this would be looking something like
- 9:07:57this.
- 9:07:58So, we'll obviously have the train
- 9:07:59values.
- 9:08:01Let's just give a random output, but
- 9:08:03there'll be four. So, it'll be 0 but the
- 9:08:05train values. And now we'll have the
- 9:08:07other, this is nothing but the given
- 9:08:08label data. So, this would be correct
- 9:08:10values. So, 1 0 0 1. As be correct
- 9:08:12values. So, 1 0 0 1. As this is a
- 9:08:15supervised learning, so we obviously
- 9:08:17will be given this label data over here.
- 9:08:19So, now what this will do is this will
- 9:08:20subtract each of these values and
- 9:08:23calculate a loss. So, how do you
- 9:08:25calculate a loss, right?
- 9:08:27Okay. So, as you can see here, if I give
- 9:08:30this L, like let me change the color
- 9:08:32here. So, if I say L is loss, so this
- 9:08:35would be nothing but Y hat of first
- 9:08:37value minus the actual value. Okay, so
- 9:08:40this is nothing but the predicted value
- 9:08:42and this is nothing but the given value.
- 9:08:44So, if I try to find out the loss of
- 9:08:46this, then what this would look like is
- 9:08:48this is just for the one value, right?
- 9:08:49So, if I have to do it for all the
- 9:08:50values, then it would be the summation.
- 9:08:52So, it would be nothing but loss is
- 9:08:54equal to summation I which ranges from 1
- 9:08:57to n, right? And then we'll have loss
- 9:09:01and then theta I. Okay, so this LI over
- 9:09:04here represent these values. Okay, so
- 9:09:06this is just for one. So, if you want me
- 9:09:07to explain you this in detail, so over
- 9:09:09here we have Y1, Y2, Y3 and Y4. How this
- 9:09:12would look like is something So, this
- 9:09:14would be for one plus Y hat of two minus
- 9:09:18actual value of Y of two, then summation
- 9:09:21predicted value of Y3 minus the actual
- 9:09:24value of Y3 and then it'll be predicted
- 9:09:28value of Y4 minus the actual value of
- 9:09:30Y4. And when I perform addition over
- 9:09:33here or when I perform summation, I can
- 9:09:35generalize this equation into this form.
- 9:09:37But we're not yet done over here. As you
- 9:09:39can see, we have something called as
- 9:09:41theta values.
- 9:09:42So, what does this theta values
- 9:09:43represents?
- 9:09:44You see, the main agenda behind finding
- 9:09:47a loss is to increase our accuracy,
- 9:09:48right? So, if I say this is my gradient
- 9:09:51descent and this is the lowest global
- 9:09:54minima, right? So, my agenda over here
- 9:09:56is to reach this global minima. So, now
- 9:09:59in order for me to do this, in order to
- 9:10:01increase my accuracy, I have to
- 9:10:02obviously change my values. So, how do I
- 9:10:05do that? How do I increase an accuracy?
- 9:10:07It's obviously by these weights. Okay?
- 9:10:09It's by this W, U, and V. So, W, U, and
- 9:10:13V are nothing but the weights. So, what
- 9:10:15this theta over here represents, let me
- 9:10:17quickly erase this.
- 9:10:19Okay, so what this theta over here
- 9:10:20represents is nothing but the values of
- 9:10:22W, U, and V. So, now what we're going to
- 9:10:25do is we're going to have a partial
- 9:10:27derivative of the loss with respect to
- 9:10:31U, and then we'll have a partial
- 9:10:33derivative of loss with respect to V,
- 9:10:35and then we'll also have partial
- 9:10:37derivative of loss with respect to W.
- 9:10:40These are nothing but weights, and we
- 9:10:41are trying to train the weights.
- 9:10:44This is done in order to increase our
- 9:10:45accuracy.
- 9:10:46Okay? And the way these models get
- 9:10:48trained is with the help of back
- 9:10:50propagation. And the way the back
- 9:10:51propagation works over here is by
- 9:10:53partial derivative. So, what I'm trying
- 9:10:55to say here is as if I get my weights
- 9:10:57over here, so let me just take another
- 9:10:59color. So, this V we know that it is
- 9:11:01totally dependent upon this value over
- 9:11:03here. This won't be V, this would be W.
- 9:11:05Okay? So, this W value will be totally
- 9:11:07dependent on the previous one. And this
- 9:11:09value will be dependent upon this one,
- 9:11:10and this value will be dependent upon
- 9:11:12this one, and this one would be finally
- 9:11:13dependent on this. So, this is in order
- 9:11:16to move from here to here to here and
- 9:11:18then to here, we'll use something called
- 9:11:20as partial derivatives, right? So, this
- 9:11:23how we update our old weights.
- 9:11:25All right. Another important concept
- 9:11:27that make neural network or recurrent
- 9:11:29neural network very important is nothing
- 9:11:31but embedding layer. So, let me quickly
- 9:11:33draw a boundary over here.
- 9:11:35Okay. So, embedding layer.
- 9:11:40All right. So, what does this embedding
- 9:11:42layer signifies? Okay, so if I take a
- 9:11:44convention way, right? So, like let's
- 9:11:46say that, you know, our X or, you know,
- 9:11:49our input over here, this is nothing but
- 9:11:51X over here, right? So, the what is the
- 9:11:52dimension or the shape of this X? The
- 9:11:54shape of this X is nothing but 1 {comma}
- 9:11:574, right? 1 {comma} 4. Similarly, it's 1
- 9:11:59{comma} 4 here, 1 {comma} 4, and 1
- 9:12:02{comma} 4. But, this is not usually the
- 9:12:04case when you're working with the real
- 9:12:06world examples. When you're working with
- 9:12:08real world examples, you won't be having
- 9:12:09four different words, right? We'll
- 9:12:11obviously have lots of values. So, for
- 9:12:13example, let's say we have X and the
- 9:12:16shape of this X is something like we
- 9:12:18have 10,000 values. Okay? So, now if I
- 9:12:21try to feed this 10,000 values into my
- 9:12:25into my network over here, obviously I
- 9:12:26would be using batch propagation. So, it
- 9:12:29would take a lot of time, right? Because
- 9:12:31it's 10,000 values after all. So, in
- 9:12:33order to overcome this, what we're going
- 9:12:34to do is we're going to reduce the size
- 9:12:36of this. Okay? So, basically what word
- 9:12:38embedding layer does is think that we
- 9:12:40have a matrix. This is our input layer,
- 9:12:43right? Our input word. So, that we call
- 9:12:45this as sparse matrix.
- 9:12:47This is nothing but, you know, an
- 9:12:49individual value. That's XI or X1, X2,
- 9:12:51X3. Let's for generalization we'll give
- 9:12:53you here as XI. And the shape of this is
- 9:12:55nothing but 1,V. V here represents the
- 9:12:58end number of dimensions. And one is
- 9:13:00because it's just going to be one word,
- 9:13:02right? So, it's going to be one. So, if
- 9:13:03I consider with respect to this, this is
- 9:13:05one and this is V. Okay?
- 9:13:07So, now what will happen is if V is is
- 9:13:1010,000 or pretty great,
- 9:13:12it obviously we won't have much
- 9:13:14we're going to lose a lot of time on
- 9:13:15computation and also take huge compute
- 9:13:18power. So, what we'll do is I'll
- 9:13:19multiply this with an embedding layer.
- 9:13:24And what this embedding matrix does is
- 9:13:26it is basically a set of features. So,
- 9:13:28this is something you know, this is a
- 9:13:30black box model. And we'll all this
- 9:13:32would do is this would attract couple of
- 9:13:34features. If you want me to give you a
- 9:13:36better analogy of what this is, and if
- 9:13:38you want me to compare this with respect
- 9:13:40to a CNN, which is nothing but another
- 9:13:41great algorithm for image processing
- 9:13:43using deep learning, right? Over there
- 9:13:45we are going to use something called as
- 9:13:46filters or kernels. So, each of those
- 9:13:48filters or kernels is responsible for
- 9:13:50extracting one specific feature, right?
- 9:13:52So, this is what embedding layer does.
- 9:13:54And what this would do, for example, now
- 9:13:57let's say that the size of embedding
- 9:13:58layer is V {comma} K. K is something
- 9:14:00that we provide an input over here. So,
- 9:14:02now what this would happen is this would
- 9:14:05give us a new matrix or embedded matrix
- 9:14:08whose size would be 1 {comma} K. I'm
- 9:14:10pretty sure you didn't understand this
- 9:14:12because over here I'm using, you know,
- 9:14:13these these letters. So, in order to
- 9:14:15make you better understand this, what
- 9:14:16I'm going to do is let me take a matrix
- 9:14:18over here. Okay? Let me take something
- 9:14:20like, you know, because this would be in
- 9:14:22an embedded form, right? So, this would
- 9:14:23be a sparse matrix. So, it would be like
- 9:14:250 0 0 0 1 then we'll have 0 0 and so on.
- 9:14:30Okay? So, now what this would do is
- 9:14:33we'll also create a matrix over here.
- 9:14:35The size of this would be something
- 9:14:36similar to that of
- 9:14:38V.
- 9:14:39This is nothing but 1 {comma} V and over
- 9:14:42here it would be K. So, the matrix shape
- 9:14:44over here would be V {comma} K, right?
- 9:14:47To give you a better analogy, let me
- 9:14:48also draw a couple of boxes here.
- 9:14:51So, this is the matrix and this is the
- 9:14:52matrix over here again. So, now what
- 9:14:54will happen over here is when I try to
- 9:14:56perform this matrix multiplication,
- 9:14:58right? What this would do, you know, as
- 9:15:00everything is zero and only one value is
- 9:15:02true, so let's say this is the one
- 9:15:04value, right? And this would be going
- 9:15:06across, you know, from left to right and
- 9:15:08this would be from top to bottom. Only
- 9:15:10one part over here would be marked and
- 9:15:12rest everything would be zero.
- 9:15:13Therefore, reducing the dimension. Let
- 9:15:16me give some random values like 0.5,
- 9:15:181.8, 0.5, just some random values. So,
- 9:15:22now this would obviously reduce the size
- 9:15:24of 1 {comma} K. So, what I'm trying to
- 9:15:26say here is now, for example, say that I
- 9:15:29have a size over here as 1 {comma}
- 9:15:3110,000. Okay, which is a very huge
- 9:15:33matrix and the shape of this, let's say
- 9:15:35that it's 10,000 {comma} 200.
- 9:15:39When I perform this embedding, right? Or
- 9:15:40embedding, the shape of the new matrix
- 9:15:43would be nothing but 1 {comma} 200. If I
- 9:15:45compare this part over here to the
- 9:15:47embedding whatever we have received over
- 9:15:49here. So, let me just give a quick
- 9:15:51brief. So, this is our embedded layer.
- 9:15:53So, this is really 1,200. So, you will
- 9:15:56see that we have decreased the
- 9:15:57dimensions by a drastic amount. Okay, so
- 9:16:00this is 10,000 and this is only 200. And
- 9:16:02this would be very efficient when we are
- 9:16:04trying to feed this to our recurrent
- 9:16:06neural network. And let me quickly show
- 9:16:08you how this would go. So, first let me
- 9:16:10draw our architecture. Let's take this
- 9:16:14blocks like this.
- 9:16:17And then we'll have an output over here.
- 9:16:19Okay. So, this would be our inputs,
- 9:16:21right? So, let me give something like
- 9:16:23this.
- 9:16:24So, initially, we used to provide X
- 9:16:26values over here, right? Now, we won't
- 9:16:28be doing that. We won't be providing any
- 9:16:30X values directly. Instead of that, what
- 9:16:32I'm going to do is I'll have an
- 9:16:33embedding layer over here.
- 9:16:38And this will have the X values. So, let
- 9:16:41me give here as X of 1, so X of 2, X of
- 9:16:453, and then we'll have X of 4, and then
- 9:16:49similarly, let's take this model to be
- 9:16:51multiple input and single output. So,
- 9:16:53here we'll be have Y hat of T. And then
- 9:16:56we'll have weights, obviously. So, this
- 9:16:58would be U, V, and W. And this is
- 9:17:02nothing but our matrix over here. So,
- 9:17:04this would be A of 0, A of 1, A of 2, A
- 9:17:09of 3,
- 9:17:10and finally A of 4.
- 9:17:12Okay, so this is our context matrix. And
- 9:17:14obviously, we'll be performing an
- 9:17:15activation function here. So, I'll just
- 9:17:17give it a F. You can put F, you can put
- 9:17:19G, it's totally up to you. So, this is
- 9:17:21how our recurrent neural network would
- 9:17:22actually work. All right?
- 9:17:25So, now that we know how RNN works, let
- 9:17:27us now understand what is LSTM. Or we
- 9:17:30can also say it as long short-term
- 9:17:31memory. You see, traditional RNNs are
- 9:17:34not good at capturing long-range
- 9:17:36dependencies. What I mean to say here is
- 9:17:38that when we tend to work with a very
- 9:17:39huge data set and multiple RNN layer, we
- 9:17:42are at the risk of vanishing gradient
- 9:17:44problem. Now, you might be wondering
- 9:17:46what is this vanishing gradient, right?
- 9:17:48Well, you see when training a very deep
- 9:17:50neural network, gradient or the
- 9:17:52derivatives decrease exponentially as it
- 9:17:54propagates down the layer. This is known
- 9:17:56as vanishing gradient problem. These
- 9:17:58gradients are actually used to update
- 9:18:00the weights of a neural network. But
- 9:18:02when the gradients vanish, these weights
- 9:18:04will not get updated. In the worst case
- 9:18:06scenario, it will completely stop the
- 9:18:08neural network from training. This
- 9:18:10vanishing gradient problem is a common
- 9:18:12issue in very deep neural networks. So
- 9:18:15to overcome this vanishing gradient
- 9:18:16problem in RNNs, long short-term memory
- 9:18:19was introduced. You see LSTM or long
- 9:18:22short memory is a modification to RNNs
- 9:18:24hidden layer. LSTM is capable of
- 9:18:26remembering RNNs weights and their
- 9:18:28inputs over a very long period of time.
- 9:18:31In LSTM, in addition to the hidden
- 9:18:32state, cell state is passed down to the
- 9:18:34next block. The way LSTM works is that
- 9:18:37it can capture long-range dependencies,
- 9:18:40that is old weights. It can have memory
- 9:18:42of previous inputs for a very extended
- 9:18:44time duration. The way LSTM cell does
- 9:18:46this is by using three main gates. First
- 9:18:49one is a forget gate. Forget gate
- 9:18:51removes the information that is no
- 9:18:52longer useful in the cell state. Then we
- 9:18:55have input gate. Additional information
- 9:18:57to the cell state is added by input
- 9:18:59gate. And finally, we have something
- 9:19:01called as output gate. Additional useful
- 9:19:03information to the cell state is also
- 9:19:05added by an output gate. This gating
- 9:19:07mechanism of LSTM has allowed network to
- 9:19:10learn the conditions for when to forget,
- 9:19:12ignore, or keep information in the
- 9:19:14memory cell.
- 9:19:15So let me now quickly move to my Jupiter
- 9:19:17notebook and show you how I can
- 9:19:19implement LSTM on name entity
- 9:19:21prediction. All right, so let me quickly
- 9:19:23move there. All right, so over here
- 9:19:25first off, I'll be opening my Google
- 9:19:28Colab.
- 9:19:32Okay, so let us give a name for our
- 9:19:34Google Colab over here.
- 9:19:36Let's give a short term, right? Name
- 9:19:38entity prediction. And let's connect our
- 9:19:40Google Colab to our server.
- 9:19:42Okay, meanwhile that's connecting. So,
- 9:19:44now you might be wondering from where am
- 9:19:46I going to use my data set? So, for me
- 9:19:48to use my data set, I'll just go for
- 9:19:49Kaggle, k a g g l e
- 9:19:52baby names. So, let me just quickly show
- 9:19:55you how this data set would look like.
- 9:19:57So, this is a CSV file over here. All
- 9:19:59right, so as you can see here, we have
- 9:20:01over 93,889
- 9:20:03unique values. Okay, so this is a very
- 9:20:06huge data set. And let's try downloading
- 9:20:09this. To download this is pretty simple.
- 9:20:11All you need to do is click this and it
- 9:20:13will get downloaded. As I've already
- 9:20:15downloaded this file, let me quickly
- 9:20:17upload this on my Jupyter notebook. So,
- 9:20:19let me go here and upload it from here.
- 9:20:23Okay, so let me go to this upload file.
- 9:20:26And yeah, so I have my CSV file here and
- 9:20:29let me open this. As this is a pretty
- 9:20:31huge data set, it will take some time.
- 9:20:32Meanwhile that's loading, let's see what
- 9:20:34we can do.
- 9:20:36So, first off let's import couple of
- 9:20:37libraries. So, we'll have import pandas
- 9:20:42as pd.
- 9:20:44And then we're going to import
- 9:20:46NumPy as np. And then we also need to
- 9:20:50have matplotlib. So, from sklearn
- 9:20:53All right, and we also need something
- 9:20:55like label encoder, but I'll show you a
- 9:20:57shortcut way to you know bypass label
- 9:20:59encoding. Okay, so let's try to load our
- 9:21:02cell here. And in order for us to read
- 9:21:04this data, so it's pretty simple. All
- 9:21:06we're going to do is let's give this as
- 9:21:08a data. This would be nothing but
- 9:21:10pandas.read_csv
- 9:21:12and then we're going to pass our file
- 9:21:15name. Let me change this to our root
- 9:21:16directory
- 9:21:18by putting a dot over here. Okay, so I
- 9:21:20won't be executing this as of now
- 9:21:21because it's trying to load our file.
- 9:21:26All right, so now that we have
- 9:21:27successfully loaded our data so, let's
- 9:21:30try running this cell over here. Okay,
- 9:21:32so let me close this and let me zoom in
- 9:21:35over here.
- 9:21:36So, now what we're going to do is let's
- 9:21:38see the shape of our data. So, let's see
- 9:21:40what's the data shape. data.shape
- 9:21:43and now let's see what it would be like.
- 9:21:45Okay, so as you can see here, we have
- 9:21:47five columns. But, the number of rows
- 9:21:50that we have is 1.8 million. That is
- 9:21:52approximately 18 lakhs, right? So, this
- 9:21:55is a pretty huge value. So, now what
- 9:21:57we're going to do is we'll just see how
- 9:21:59our data is looking like. So, we'll see
- 9:22:01data.head.
- 9:22:03And let's see what we need. So, as you
- 9:22:05can see here, we have ID, which is of no
- 9:22:08use for us. Then we have name. Okay,
- 9:22:10then this year, I don't think it's of
- 9:22:12any use for us. Then we have gender and
- 9:22:14count. Count here represents, you know,
- 9:22:17how many people have the name Mary, how
- 9:22:19many people have the name Anna, how many
- 9:22:21people have the name Emma, Elizabeth,
- 9:22:23and Minnie. This is over here, out of
- 9:22:25this if you see, right? There are a
- 9:22:27couple of things that we don't need. We
- 9:22:28can drop them out. You know, all we need
- 9:22:30is a name. Okay, and then we also need
- 9:22:32the gender. Because this is going to be
- 9:22:35our prediction. We're going to predict a
- 9:22:36we'll give our own custom name and then
- 9:22:38we'll see whether the name that is
- 9:22:40you're giving is male or a female. Okay?
- 9:22:44So, now what we're going to do is let's
- 9:22:45see how many unique values we have. So,
- 9:22:47let me quickly erase this first. Okay,
- 9:22:50so what I'm going to do is data.names.
- 9:22:53So, this should give us here name. And
- 9:22:56then we'll type here as unique.
- 9:22:58Okay, so this should give me unique
- 9:23:00values. Okay, so over here I have 93,889
- 9:23:04unique names. Okay, so now what we're
- 9:23:07going to do is we want to label encode
- 9:23:09this, right? So, we want our female, uh
- 9:23:11which is nothing but F, we want female
- 9:23:13to be zero and then male to be one or
- 9:23:15vice versa. So, in order to do that,
- 9:23:17either we can use label encoder or
- 9:23:20there's a shortcut method to this. Let
- 9:23:21me quickly show you how that works. So,
- 9:23:23first of all, we'll take our data frame,
- 9:23:25so it's data. And which column do you
- 9:23:27want to do this for? We want to do this
- 9:23:29for our gender column, right? So, let me
- 9:23:32pass this and give gender. And now what
- 9:23:35we're going to do is
- 9:23:37Okay?
- 9:23:38We'll take this as as type.
- 9:23:41Okay, this would be obviously in the
- 9:23:42form of category.
- 9:23:44And now what we'll do is this is cat
- 9:23:47dot codes.
- 9:23:50Okay, so this is nothing but panda
- 9:23:51shortcut, you know, to label encoding.
- 9:23:53Let's try to execute this and see what
- 9:23:55it would look like. So, as you can see
- 9:23:57here, we have couple of zeros and, you
- 9:23:59know, ones. This is nothing but it's
- 9:24:00representing females with one and males
- 9:24:03with zeros. Okay? So, now what we'll do
- 9:24:05is we have to update this column.
- 9:24:09So, we'll paste this. And this should be
- 9:24:11something like this over here. And let
- 9:24:13me execute this. Okay, so if you want to
- 9:24:16see how our data would look like now,
- 9:24:18let me just quickly run this once again.
- 9:24:20So, you'll see here now the values has
- 9:24:22been label encoded. Okay? So, now what
- 9:24:25we're going to do is we obviously need
- 9:24:27to take the unique names, right? And
- 9:24:30then we'll obviously group it by, right?
- 9:24:31So, what we'll do for this is we'll take
- 9:24:33something like data. We'll group this by
- 9:24:37the names. So, group by
- 9:24:39names.
- 9:24:40All right? And now what we'll do is
- 9:24:42we'll calculate the mean
- 9:24:44of the genders.
- 9:24:46We'll reset the index. The reason why we
- 9:24:47want to reset the index is because, you
- 9:24:49know, if you don't give the index then
- 9:24:50our name over here will become the
- 9:24:52index, right? So, we'll give reset
- 9:24:54{underscore} index.
- 9:24:56All right? So, let's give this to a new
- 9:24:59data frame and we'll call this as DF.
- 9:25:02Okay, let me execute this now.
- 9:25:04And let's see how this DF would look
- 9:25:05like. Okay, let me execute this right
- 9:25:08after this.
- 9:25:09So, as you can see here, it has grouped
- 9:25:11by by names, all everything in an
- 9:25:13ascending order. So, if this is all in
- 9:25:15an alphabetical manner. And yeah.
- 9:25:19And now only thing that I want to work
- 9:25:21on is this gender.
- 9:25:22Okay? The reason is because over here
- 9:25:24I'm getting a floating point value. I
- 9:25:26don't want this floating point value. I
- 9:25:28want to change this to integer value,
- 9:25:30right? So, what I'll do is
- 9:25:32DF gender
- 9:25:34This would be nothing but
- 9:25:36DF gender. Then I'll all I'm going to do
- 9:25:39is as type.
- 9:25:40I'll just put here as int. So, let's now
- 9:25:42see what this value would look like.
- 9:25:45Fantastic. We over here have now, you
- 9:25:47know, ones and zeros, which is nothing
- 9:25:48but an integer value. Okay? So, if you
- 9:25:51want to see this, so I either I can
- 9:25:53write DF or I can also put as head.
- 9:25:56Okay, so these are the first five
- 9:25:57values.
- 9:25:58Okay. So, now the way our neural network
- 9:26:01is going to work or the recurrent neural
- 9:26:02network is going to work is that, you
- 9:26:04know, I hope you remember these boxes,
- 9:26:06right? So, when I was talking or when I
- 9:26:09was explaining this RNN, I was saying
- 9:26:11that I would be passing around the
- 9:26:13words. But here in this project or in
- 9:26:16this program, we won't be passing words
- 9:26:18over here. You know, we won't be passing
- 9:26:20like Abba or Abida or Adam. We won't be
- 9:26:23passing these words. Instead of that,
- 9:26:25we'll be passing letters.
- 9:26:27So, over here it's going to be like
- 9:26:29alphabets. So, A, B. It can be any
- 9:26:32alphabet. It can be Z here. So,
- 9:26:34basically it depends upon whatever the
- 9:26:36value is coming here. So, in order to do
- 9:26:38that, we have to find number of unique
- 9:26:40alphabets. So, we know how many unique
- 9:26:41alphabets we have, right? So, we it's
- 9:26:4326. So, in order to get these alphabets,
- 9:26:46what we'll do is let me first quickly
- 9:26:47erase this.
- 9:26:49Erase all drawing.
- 9:26:50So, now we have 26 alphabets. We have to
- 9:26:53create our own vocabulary. So, what I'm
- 9:26:54going to do is I'm going to import
- 9:26:56string. So, now I need letters, right?
- 9:26:59So, l e t t e r s. This would be nothing
- 9:27:01but list of string
- 9:27:05.ascii.
- 9:27:06Okay? And if you want to see what this
- 9:27:07would give me, this would be nothing but
- 9:27:10the list of alphabets, which are in
- 9:27:11lower cases.
- 9:27:13Okay?
- 9:27:14And now what we'll do is we'll try to
- 9:27:15create a label encoding or we have to
- 9:27:18create a vocabulary, right? So, we'll
- 9:27:19have something like vocab. This would be
- 9:27:21nothing but I'll be using dictionary.
- 9:27:24And now what I want is zip. The way I
- 9:27:26want over here is, you know, for every
- 9:27:28individual values of this A B C D, I
- 9:27:31want to label encode this to 0 1 and
- 9:27:35whatever the value it is, right? So, it
- 9:27:36would be from 1 to 27.
- 9:27:38A unique numbers, right? So, this would
- 9:27:40be nothing but letters. And then uh
- 9:27:42we'll be need something like uh range
- 9:27:451 {comma} 27. So, this would give me the
- 9:27:48matrix from 1 to 26, right? And let's
- 9:27:50now see what this would look like. So,
- 9:27:52we have vocab.
- 9:27:54And let me execute this. This should
- 9:27:56give me a dictionary, okay? So, here
- 9:27:58we'll convert A to 1.
- 9:28:00Okay? And then B would be 2, C would be
- 9:28:033, and so on, Z would be 26.
- 9:28:06And now what we're going to do is uh
- 9:28:08we'll just try to create the reverse
- 9:28:10vocabulary. And the reason is we
- 9:28:12obviously won't be needing this, but uh
- 9:28:14you know, just in case you want to use
- 9:28:16it would be something very similar to
- 9:28:17this. Let me just copy the exact same
- 9:28:19thing.
- 9:28:20And paste it over here.
- 9:28:22So, we'll just do it as reverse, right?
- 9:28:24So, it will be R {underscore}
- 9:28:26R {underscore} And here, instead of
- 9:28:28numbers being second,
- 9:28:30we'll just cut this letters and we'll
- 9:28:32pass letters over here.
- 9:28:34And now you'll see if you're trying to
- 9:28:36decode whatever we have predicted, you
- 9:28:37know, we can just pass it down like
- 9:28:39this.
- 9:28:40Okay. So, now what what will happen is
- 9:28:42we need to do something like, you know,
- 9:28:44all our data, whatever is there, we have
- 9:28:45to convert them into a lowercase.
- 9:28:48So, and then once we convert them into a
- 9:28:50lowercase, we have to encode them into a
- 9:28:52numbers.
- 9:28:54So, whatever I'm saying is this A A B A
- 9:28:57N, right? A ban. So, we this A A
- 9:28:59obviously first of we have to convert
- 9:29:00all of these into a lowercase,
- 9:29:02you know, this value. And then whatever
- 9:29:04the equivalent value of A, the numerical
- 9:29:07value of A, so it's obviously going to
- 9:29:09be one. We'll substitute that with this.
- 9:29:11And it's going to be a list, right? So,
- 9:29:12how do I do that? So, for that I'll
- 9:29:14write a function.
- 9:29:16So, we'll have DEF word to number,
- 9:29:18right? Word to
- 9:29:21So, now what I'm going to do is I'm
- 9:29:22going to have for loop for I in range.
- 9:29:26So, this would be nothing but
- 9:29:29we have to go through the entire shape,
- 9:29:31right? So, d f dot shape.
- 9:29:33This should give me a list and I just
- 9:29:35need the first index.
- 9:29:37Okay? So, now what I'm going to do is
- 9:29:39I'll create one new list sequence.
- 9:29:42This would be nothing but for letters in
- 9:29:46d f. Obviously, we want the names part.
- 9:29:50And in this we're going to pass the
- 9:29:51index value. It's going to be I.
- 9:29:53Let me just give some space here just so
- 9:29:55that you better understand this.
- 9:29:57And now what I'm going to do is, you
- 9:29:59know, I'll have this vocabulary.
- 9:30:01vocab See, every time I pass a letter
- 9:30:03it'll convert it into, you know, this
- 9:30:05individual letter it'll convert it into
- 9:30:07a list all the equivalent, you know,
- 9:30:09numerical representation. It'll be
- 9:30:11letters and obviously it has to be in
- 9:30:13lower so it'll be lower.
- 9:30:15And then we'll just close this bracket
- 9:30:16here.
- 9:30:17So, now what we're going to do is before
- 9:30:19we execute this function, we'll have to
- 9:30:22append this so it'll be d f.
- 9:30:24And this is going to be names
- 9:30:27dot I.
- 9:30:28We'll replace the name in that index
- 9:30:30with this particular sequence.
- 9:30:32Okay? So, now all we need to do is run
- 9:30:34this function over here.
- 9:30:36And yeah.
- 9:30:38This will take some time. The reason is
- 9:30:39because we have almost around 18 lakh
- 9:30:42values. So, yeah, this should take some
- 9:30:44time. Meanwhile, let me just comment
- 9:30:46this.
- 9:30:53Okay? So, in the next stage what we're
- 9:30:54going to do is let's see how our this
- 9:30:57value over here would look like. So, let
- 9:31:00us now first execute this.
- 9:31:03All right. So, let us now see how our
- 9:31:04data frame will look like. So, let me
- 9:31:06execute this block now.
- 9:31:09So, as you can see here, our names have
- 9:31:11been completely changed or converted
- 9:31:13into list of numbers. But now, only
- 9:31:15issue that we are trying to have is the
- 9:31:18imbalance in the size of the list.
- 9:31:20Because when we are trying to have the
- 9:31:22number of boxes, right? We won't be
- 9:31:24having variable number of boxes. Okay?
- 9:31:27So, what we're going to do is either we
- 9:31:28set a value like something like take an
- 9:31:30average number like 10, 20, or you can
- 9:31:33take something like, you know, something
- 9:31:35like you take you either depend on
- 9:31:36maximum number or the minimum number of
- 9:31:38list. But the thing is, if you take the
- 9:31:41maximum number, then we have to pad a
- 9:31:42lot of zeros, and this would lead to a
- 9:31:44loss.
- 9:31:45So, if I reduce the size, this would
- 9:31:46also decrease the accuracy, right? So,
- 9:31:48what we're going to do is we'll plot
- 9:31:49this name and gender in the form of a
- 9:31:52histogram. So, let's take here X. This
- 9:31:55would be DF names.
- 9:31:58And we'll give here as dot values.
- 9:32:00And then same thing we'll do it for Y,
- 9:32:02DF gender.
- 9:32:04And this should be dot values.
- 9:32:06So, what we'll do is we also need a
- 9:32:09list, okay? So, now as we are going to
- 9:32:10plot this on a histogram, and what we're
- 9:32:12going to see in the histogram is just to
- 9:32:15analyze, you know, this this is a graph.
- 9:32:17We want to analyze, you know, where does
- 9:32:19the highest number of sequence, or if
- 9:32:22suppose this is a size, if this is size
- 9:32:23eight, and this is like 8,000 words or
- 9:32:268,000 names have the size eight, then
- 9:32:28you know, we can keep our average
- 9:32:30somewhere near, and then we can also
- 9:32:32decide, you know, if if the number after
- 9:32:3410, if not many names have a longer
- 9:32:36number or the longer length of that
- 9:32:38name, you know, so we can keep our
- 9:32:40average somewhere around nine or 10,
- 9:32:42okay? So, let's now quickly see how we
- 9:32:44can do that.
- 9:32:46To get the length of our names, so
- 9:32:47length X or name length.
- 9:32:51This would be like list comprehension
- 9:32:53for I in range 0, DF.shape
- 9:32:58of 0.
- 9:33:00And now what we're going to do is we
- 9:33:01need to find the length. So, this would
- 9:33:03be length X of I.
- 9:33:06Or here, you can either give BF or you
- 9:33:08can also give this X, right? So, it
- 9:33:10would be length of X.
- 9:33:11Okay, so let me quickly execute this.
- 9:33:14Okay, so here we're getting an error.
- 9:33:15Oh, yeah. It's not O, it's going to be
- 9:33:17zero, right? So, let me execute this
- 9:33:19now.
- 9:33:19Okay, so let me show you how this would
- 9:33:21look like. Name length, and let me print
- 9:33:24this off.
- 9:33:25So, as you can see, this is giving me
- 9:33:26list of names.
- 9:33:27So, there are huge amount of names. So,
- 9:33:29as you can see, first we had five, five,
- 9:33:31and then nine. So, let's now plot this
- 9:33:34and see how it would look like. Import
- 9:33:37Matplotlib as plt.
- 9:33:39All right, this is perfect. So, now what
- 9:33:40I'm going to do is I have to plot this,
- 9:33:42right? So, all I'm going to do is
- 9:33:44plt.hist.
- 9:33:45All right, and now I'm just going to
- 9:33:47give name length.
- 9:33:49And number of bins, this would be like
- 9:33:51let's give 20, okay? And then plt.show.
- 9:33:54Okay, so what do we find from this graph
- 9:33:57over here? You see, this is nothing but
- 9:33:58the length of the names. So, two, four,
- 9:34:01six, eight, 12, all these are length of
- 9:34:03the names. And this is nothing but zero,
- 9:34:055,000, 10,000, this is nothing but
- 9:34:07number of names that have a length four
- 9:34:09or number of names that have the length
- 9:34:11six. So, as you can see, right? The
- 9:34:13there are around almost 25,000 names
- 9:34:15whose length is six. And then as I cross
- 9:34:18like 10 or as I cross 12, not many names
- 9:34:21are there whose length is greater than,
- 9:34:23you know, 12. So, what I'm going to do
- 9:34:25here now is now we have to pad, right?
- 9:34:27Now, we have to pad number of zeros. So,
- 9:34:29in order to pad zeros, we have something
- 9:34:31called as built-in function from Keras.
- 9:34:33So, from Keras or you can also set as
- 9:34:35Keras.preprocessing
- 9:34:38.sequence import pad_sequence.
- 9:34:42So, let me quickly execute this now.
- 9:34:44And now what I'm going to do is I'm
- 9:34:45going to create a new list.
- 9:34:47So, let this be X. This is in lowercase.
- 9:34:50So, pad_sequence. And the things that it
- 9:34:52this is going to take is obviously the
- 9:34:54sequence. We have to give a list of
- 9:34:55sequence.
- 9:34:57Okay, and then we'll give something like
- 9:34:59df.names.
- 9:35:01And then this is going to be values.
- 9:35:04And now we want to define the max
- 9:35:06length. So, this is going to be 10. We
- 9:35:08also have an option of providing where
- 9:35:10do you want to do the padding? So, we
- 9:35:11can also do it as pre or post. We'll
- 9:35:14obviously be doing pre. So, let's see
- 9:35:17how do we do that. This would be nothing
- 9:35:19but, you know, if you can see this
- 9:35:20sequence over here. So, we have padding
- 9:35:22is equal to pre. It's so it's by
- 9:35:24default, right? So, pre. So, let me
- 9:35:26execute this now. And let's see how this
- 9:35:28X would look like.
- 9:35:30Okay. So, as you can see here, X is a
- 9:35:31matrix whose length is 10. Okay? So,
- 9:35:34each of these like this is this is
- 9:35:36nothing but, you know, 19 million cross
- 9:35:3810. So, there are 10 columns throughout
- 9:35:40all.
- 9:35:41So, now what we're going to do is we're
- 9:35:42going to create our own model. So, for
- 9:35:46that we'll do from keras.layers
- 9:35:49import
- 9:35:51input layer.
- 9:35:52And then we have to have embedding layer
- 9:35:54cuz if you don't have embedding layer,
- 9:35:56then you know, it it would be like each
- 9:35:57input would be something like 1 comma or
- 9:36:001.9 million. That is 18 lakhs. So, it's
- 9:36:02a pretty huge value to compute. So, we
- 9:36:04don't want that. So, that's why we'll
- 9:36:05use embedding layer. Then we have dense
- 9:36:06layer and then we have LSTM.
- 9:36:09We also, you know, rather than taking
- 9:36:11this as a sequential model, we'll take
- 9:36:12it as, you know, feed forward. So, what
- 9:36:15we'll do is from keras.models
- 9:36:19import model.
- 9:36:21So, now what we're going to do is we'll
- 9:36:23have to create our input layer. So, this
- 9:36:25would be input is equal to input.
- 9:36:28And now the shape that we're going to
- 9:36:29pass over here for for this shape
- 9:36:33So, how many columns do we have? We
- 9:36:34obviously have 10 columns, right? So,
- 9:36:36it's going to be 10.
- 9:36:37So, now what we're going to do is next
- 9:36:38we're going to have embedding layer. So,
- 9:36:40let's say this is EMB and this would be
- 9:36:43embedding.
- 9:36:44So, input over here
- 9:36:47or the input dimension over here is
- 9:36:48nothing but vocab size.
- 9:36:50We haven't defined this vocab size, so
- 9:36:52let's quickly do that. So, vocab
- 9:36:55size this would be nothing but length
- 9:36:59of vocabularies
- 9:37:01plus one. The reason why I'm doing plus
- 9:37:03one is because we also have zeros over
- 9:37:05here, right?
- 9:37:06And this would be like vocab size if you
- 9:37:08want to see.
- 9:37:09And let me execute this.
- 9:37:11So, we have 27, right? So, 26 are the
- 9:37:13number of alphabets and one is because
- 9:37:15we have number of zeros. So, we'll pass
- 9:37:17this as vocab size.
- 9:37:19Okay? And now we're also going to pass
- 9:37:21output dimension. So, output dimension,
- 9:37:23this is nothing but, you know, how many
- 9:37:24dimensions we want. So, now this is
- 9:37:26going to be five. All right? And the
- 9:37:28input for this embedded layer is going
- 9:37:29to be from INT.
- 9:37:31Okay? And now we are going to have our
- 9:37:33first LSTM layer. So, it's going to be
- 9:37:34LSTM
- 9:37:36one. So, this would be LSTM layer.
- 9:37:40Number of units we have to define here.
- 9:37:42So, units, this is going to be like 32.
- 9:37:45The units over here does not represent
- 9:37:46the number of boxes.
- 9:37:48The units over here represent the A
- 9:37:49values, right? So, now we have to do
- 9:37:52return sequence and this is going to be
- 9:37:54true.
- 9:37:55And the input for this is going to be
- 9:37:56from embedded layer.
- 9:37:57Then we have LSTM second layer.
- 9:38:00And this is be LSTM units we're going to
- 9:38:03pass. So, number of units that we're
- 9:38:05going to pass now is 64.
- 9:38:07And now the input for this is going to
- 9:38:09be LSTM one. Finally, we have an output
- 9:38:12layer.
- 9:38:13So, at the end, right? We're going to
- 9:38:14have a dense layer, right? So, dense.
- 9:38:17So, the number of units or number of
- 9:38:18neurons at the end we are going to have
- 9:38:19one.
- 9:38:20And kind of activation function that I'm
- 9:38:22going to have here is going to be
- 9:38:23sigmoid because we have to predict
- 9:38:25either it's a male or a female. Okay?
- 9:38:27So, sigmoid.
- 9:38:29And the input for this is going to be
- 9:38:31LSTM two.
- 9:38:33So, finally, we have to add this to our
- 9:38:35model. So, I'll do it as my model.
- 9:38:38This would be model.
- 9:38:40So, now I have to define the inputs. So,
- 9:38:42I N P U T S, this is going to be inputs
- 9:38:44I N P
- 9:38:46and outputs.
- 9:38:48This is going to be out.
- 9:38:49Okay? So, let me quickly execute this
- 9:38:51now.
- 9:38:52Okay, so we have this error.
- 9:38:54Please provide either a shape. Okay.
- 9:38:57Let's see what's Oh, yeah. Oh, here I've
- 9:38:59given it as pass, right? So, it's not
- 9:39:00going to be this pass. It's going to be
- 9:39:01shape.
- 9:39:02So, let me quickly execute this once
- 9:39:04again.
- 9:39:05Okay, as you can see, we have
- 9:39:06successfully executed this. And now
- 9:39:08let's see the model. summary.
- 9:39:10So, as you can see here, first off, we
- 9:39:12have input layer.
- 9:39:13Okay. So, we can have a number of
- 9:39:15values, but there will be only 10
- 9:39:17features. Okay, that's 10 columns.
- 9:39:19And then we're going to have once you go
- 9:39:21through this embedding layer, then you
- 9:39:23know, instead of having 10
- 9:39:25you know, instead of having that 1
- 9:39:27million or whatever it is, 1.8 million,
- 9:39:29we'll have just 135 parameters. Okay.
- 9:39:32Similarly, over here and finally at the
- 9:39:33dense, we have 65 parameters and then we
- 9:39:35have one.
- 9:39:36The reason why we have 65 here, just for
- 9:39:39if you don't know, is because 64 + 1
- 9:39:41bias. And this will give us 65.
- 9:39:43Okay. So, now finally, we're going to
- 9:39:45train our model. But before that, we
- 9:39:47have to compile it. It's going to be my
- 9:39:48model.
- 9:39:49model.compile
- 9:39:51Okay. So, we have to find an optimizer.
- 9:39:53So, optimizer, best one that I feel is
- 9:39:56Adam.
- 9:39:57Then we have to find the loss.
- 9:39:59So, as we're going to use just two
- 9:40:01predictions, right? It's either it's
- 9:40:02male or female, we'll use binary
- 9:40:04cross-entropy.
- 9:40:06And then finally, the matrix that we
- 9:40:07want to use here is
- 9:40:10This will be accuracy.
- 9:40:12So, let me execute this now.
- 9:40:14And finally, we are going to compile
- 9:40:15this. So, we have history.
- 9:40:17So, this will be model.fit.
- 9:40:20Okay. And now we're going to pass our
- 9:40:22values. We're going to give X. We're
- 9:40:23going to give Y.
- 9:40:25As you know, X is nothing but a matrix
- 9:40:26which has a padding. And Y is nothing
- 9:40:28but you know, the the classes. They're
- 9:40:29telling us either it's ones or zeros.
- 9:40:31And number of epochs
- 9:40:34is going to be 10.
- 9:40:35Batch size, as this is a pretty used
- 9:40:37data set, we are going to keep a pretty
- 9:40:38high batch size.
- 9:40:40So, I'll give a batch size here as 256.
- 9:40:43And then finally, we need validation
- 9:40:45split.
- 9:40:46Okay, it's not going to be my model,
- 9:40:47it's going to be my model, right? So, my
- 9:40:49underscore model. So, finally we're
- 9:40:52going to have validation split here.
- 9:40:54So, let's give it as 20%. So, it's going
- 9:40:56to be 0.2.
- 9:40:58All right. So, finally it's the moment
- 9:41:00of truth. Let us now execute our code.
- 9:41:03This will take some time to execute.
- 9:41:06So, let's see how this would look like.
- 9:41:11Okay, so if you can analyze this data
- 9:41:13over here,
- 9:41:15so as you can see, right? This
- 9:41:16validation accuracy has to increase.
- 9:41:19And we cannot see see every time, you
- 9:41:21know, if our model is over fitting, the
- 9:41:23accuracy over here will keep on
- 9:41:24increasing. All right. So, the more
- 9:41:27reliable source over here to see is
- 9:41:29nothing but validation accuracy. So, if
- 9:41:31validation accuracy is increasing, that
- 9:41:33means our model is neither over fitting
- 9:41:34or under fitting. And you can also see
- 9:41:36that we have our validation loss, which
- 9:41:38is kind of decreasing. And over here as
- 9:41:40well, we can see the validation loss
- 9:41:41over here. We can see the loss of our
- 9:41:43model is decreasing from 60 then 40 then
- 9:41:4639 39 and 80 38. So, let's now wait for
- 9:41:49a few more epochs. So, we have four more
- 9:41:52to go.
- 9:41:55Okay, so as you can see here, you know,
- 9:41:57our validation accuracy has been
- 9:41:59increasing. So, this is a very healthy
- 9:42:01growth.
- 9:42:02And even over here, our accuracy of our
- 9:42:04model is also increasing. Fine? And the
- 9:42:06loss is decreasing. It's decreased from
- 9:42:0860% to 37% and over here, our validation
- 9:42:11loss decreased from 42% to 36%.
- 9:42:14So, let us now map this like whatever
- 9:42:16values we have received. So, in order to
- 9:42:18map this, we have H.
- 9:42:20You know, this model over here retrieves
- 9:42:21us the history function, right? So,
- 9:42:23we'll give hist
- 9:42:24dot history.
- 9:42:25And let me execute this.
- 9:42:27Okay. So, now what we're going to do is
- 9:42:29this is nothing but key value pairs. So,
- 9:42:31if I put H,
- 9:42:32you know, if I give something like
- 9:42:33accuracy, okay, we let's plot this.
- 9:42:36So, model dot plot, right? So, plt dot
- 9:42:40plot.
- 9:42:41Okay.
- 9:42:42We'll have accuracy. We'll compare this
- 9:42:44with respect to accuracy and then
- 9:42:46plt.plot
- 9:42:49then we'll have validation accuracy. And
- 9:42:51we want to show this, right? So,
- 9:42:52plt.show.
- 9:42:53So, as you can see here, okay, so just
- 9:42:55to give a better analogy, let's let's
- 9:42:57execute one of these first.
- 9:42:59Okay, so the blue line over here
- 9:43:00represents the accuracy of our model and
- 9:43:02then this is nothing but the accuracy of
- 9:43:04our training data, right? Or testing
- 9:43:06data. So, as you can see, our model
- 9:43:07accuracy isn't decreasing. So, it's a
- 9:43:09very good model and it has trained very
- 9:43:11well. So, now coming down to the moment
- 9:43:13of truth, so let's now, you know, take a
- 9:43:15random name and see whether it can
- 9:43:17predict whether the name is true or
- 9:43:18false. Okay, so we'll give here as
- 9:43:20test_name.
- 9:43:21So, let's give something like, you know,
- 9:43:24we'll this will be like name, right? So,
- 9:43:25we'll give name.lower.
- 9:43:28And now we're going to pass the name
- 9:43:29over here.
- 9:43:31So, this would be, let's say, Tom.
- 9:43:34Okay, so we have to convert this into
- 9:43:35letters, right? So, it'll be vocab of I
- 9:43:39for
- 9:43:40I in test name.
- 9:43:42And now we'll give the name here as
- 9:43:44X_test.
- 9:43:46So, this would be nothing but we have to
- 9:43:47pad the sequence, pad sequence. We'll
- 9:43:50pass this in a form of a tuple, then
- 9:43:51this would be seq.
- 9:43:54And then we know that we have to pad 10,
- 9:43:55right? And anyways, we don't have to say
- 9:43:57whether it's pre or post because by
- 9:43:59default it's going to be pre.
- 9:44:01Fine. So, let us now see how our text
- 9:44:03data would look like. So, X_test.
- 9:44:06String attribute has Okay.
- 9:44:08Oh, yeah. So, I have done a typo over
- 9:44:10here. It's going to be l o w e r.
- 9:44:13Let me execute this once again.
- 9:44:15So, as you can see here, we have a
- 9:44:17matrix which has a size of 10. And let's
- 9:44:19now see what this would predict. So,
- 9:44:22it's the moment of truth. So, y.predict.
- 9:44:25This would be model.predict.
- 9:44:28And I'm going to pass here as X_test.
- 9:44:30And let's see what does this predict.
- 9:44:32So, we'll give here as y_pred. And yeah,
- 9:44:35let us execute this.
- 9:44:37Okay, so we are getting this in a form
- 9:44:38of a array. So, what this tells us, you
- 9:44:41know, this tells us that, you know, this
- 9:44:43is like 70% chances that this name is
- 9:44:46Tom. In order to make this, you know,
- 9:44:48layman's stuff, so what we're going to
- 9:44:49do is we'll have if y pred Let me
- 9:44:52execute this first.
- 9:44:54So, let me go down to another block.
- 9:44:56So, this would be something like if
- 9:44:59y pred is less than 0.5, then we'll say
- 9:45:03the name is female, okay?
- 9:45:08Else print name is masculine or name is
- 9:45:13Always let let be male.
- 9:45:15Okay. So, same thing over here, we'll
- 9:45:16just give it as
- 9:45:18male.
- 9:45:19So, let's now see what this thing
- 9:45:20predicts. Okay, so this thing predicts
- 9:45:22male.
- 9:45:23So, let's take another common name.
- 9:45:26Let's take something like Let's go to
- 9:45:28Google and see what name can we take.
- 9:45:31Yeah, we can take up something like Brad
- 9:45:32Pitt.
- 9:45:33Okay, so let's execute this and let's
- 9:45:36execute this again.
- 9:45:37And then this.
- 9:45:39So, as you can see here, this is giving
- 9:45:40me a male name. And let's give a female
- 9:45:42name over here. Let's give as Mary.
- 9:45:45And let's see whether this would predict
- 9:45:47it as male or female. So, as you can
- 9:45:49see, it's female.
- 9:45:50Now, as if you have seen the data set,
- 9:45:52right? It says this is the name from the
- 9:45:54US kids, right? What about What will
- 9:45:56happen if I give a Indian-based name?
- 9:45:59So, let me give a Indian-based name like
- 9:46:02Priyanka.
- 9:46:03And let me execute this.
- 9:46:05It's giving me a female name, right? So,
- 9:46:07this is something pretty astonishing,
- 9:46:09right? So, why do you think it gave me a
- 9:46:10female name? Well, the reason is because
- 9:46:12when we are training this model like RNN
- 9:46:15using RNN, right? It's not looking at
- 9:46:17the name. It doesn't know whether the
- 9:46:18name is female or not. But, as a matter
- 9:46:21of fact, it is looking at the pattern.
- 9:46:23Okay? So, it might be looking at, you
- 9:46:25know, if the name ends with so and so,
- 9:46:26it is a female. If the name starts with
- 9:46:29this or if a name have something like
- 9:46:31this, it means, you know, it's a female
- 9:46:33or a male. So, to give you a better
- 9:46:35analogy, let's give something like
- 9:46:37Julia.
- 9:46:39So, let me execute this.
- 9:46:41This will give me a female, right? So,
- 9:46:43what if I give something like Juneid?
- 9:46:46So, it give it as male. So, all it's
- 9:46:48trying to do is it's trying to see, you
- 9:46:50know, uh recognize a pattern. That's why
- 9:46:52it's taking individual words at the same
- 9:46:54time. So, this is what makes NLP using
- 9:46:57LSTM very effective.
- 9:47:00All right. So, moving ahead, let's see
- 9:47:02some of the LSTMs use cases.
- 9:47:04You see, LSTM is a very popular deep
- 9:47:06learning algorithm for sequential
- 9:47:08models.
- 9:47:09Apple Siri and Google's voice search are
- 9:47:11some of the real-world examples that
- 9:47:13have used LSTM. And you won't believe
- 9:47:15it, LSTM is a success story for those
- 9:47:17algorithm.
- 9:47:18So, let us now have a look and see how
- 9:47:20LSTM changed that technology.
- 9:47:22Okay. So, starting off with Apple, in
- 9:47:242003, Apple was the first major tech
- 9:47:26company to integrate a smart assistant,
- 9:47:28that is Siri, into their operating
- 9:47:30system. And the Siri was actually a
- 9:47:32byproduct of some other company. So,
- 9:47:34Siri was a company's adoption of a
- 9:47:36standalone app that has been purchased
- 9:47:38along with the creators who made it. It
- 9:47:39was somewhere in 2010. The initial
- 9:47:42reviews about Siri was that it was
- 9:47:43intense. But, over the next few months
- 9:47:45or and years, the users became more
- 9:47:47impatient with the shortcomings. And all
- 9:47:49too often, it wrongly interpreted
- 9:47:51commands. And then, you know, no matter
- 9:47:54what you do, there was no fix for it.
- 9:47:56So, this is when Apple moved Siri's
- 9:47:57voice recognition to a neural-based
- 9:47:59system.
- 9:48:00Some of the previous technique remained
- 9:48:02operational, something like, you know,
- 9:48:03applying hidden Markov models. But, most
- 9:48:05of the time, you know, DNN or deep
- 9:48:07neural network using LSTM was used.
- 9:48:10Although people did not find any changes
- 9:48:12on the outside, but from within, it was
- 9:48:14a supercharged deep learning model.
- 9:48:16Speaking about Google's implementation,
- 9:48:19Google implemented Google voice search
- 9:48:20somewhere around 2009.
- 9:48:22Google voice transcription had initially
- 9:48:24used something called as Gaussian
- 9:48:25mixture model.
- 9:48:27This was This was nothing but an
- 9:48:28acoustic model and this was something
- 9:48:29considered to be a state-of-the-art
- 9:48:31speech recognition for almost 30 plus
- 9:48:33years.
- 9:48:34But it was in 2012, there was a boom in
- 9:48:36deep neural network.
- 9:48:37And when Google implemented deep neural
- 9:48:39network, that too using multiple layer
- 9:48:41networks, there was a huge performance
- 9:48:43gap.
- 9:48:44But things really improved when the
- 9:48:46recurrent neural network, especially
- 9:48:47with LSTM RNN, first launched on an
- 9:48:49Android speech recognition in May 2012.
- 9:48:52Compared to deep neural network, LSTM
- 9:48:54RNNs have additional recurrent
- 9:48:56connections and memory cell that allows
- 9:48:58them to remember the previous data.
- 9:49:00All right. So, moving ahead to our last
- 9:49:01topic of our session, let us now see
- 9:49:03some of the real-world applications of
- 9:49:04LSTM RNN networks. First off, we can
- 9:49:07perform named entity recognition. This
- 9:49:09is something which we did in our
- 9:49:10previous demo, right? So, what is this
- 9:49:12named entity recognition? You see, named
- 9:49:14entity recognition is a subtask of
- 9:49:16information extraction that seeks to
- 9:49:18locate and classify named entity
- 9:49:20mentioned in an unstructured data. Okay?
- 9:49:23Next, we have something called as
- 9:49:24sentiment analysis.
- 9:49:26Sentiment analysis is a predictive
- 9:49:28modeling task where model is trained to
- 9:49:30predict the polarity of a textual data
- 9:49:32or sentiments like positive, neutral, or
- 9:49:34negative.
- 9:49:35Sentiment analysis is performed by
- 9:49:37various businesses to understand their
- 9:49:39consumers' behavior towards the product.
- 9:49:42Then we have machine translation. The
- 9:49:44task of machine translation consists of
- 9:49:45reading text in one language and
- 9:49:47generating text in another language.
- 9:49:49When neural networks are used for this
- 9:49:51task, we talk about neural machine
- 9:49:53translation.
- 9:49:55Within neural machine translation, an
- 9:49:56encoder-decoder structure is quite a
- 9:49:58popular LSTM RNN architecture.
- 9:50:02>> [music]
- 9:50:06>> Why should we choose deep learning for,
- 9:50:08you know, various tasks?
- 9:50:10So, the big advantage of using deep
- 9:50:12learning is that we can extract more
- 9:50:14number of features. And when we have
- 9:50:16more number of features and when we can
- 9:50:17work at the same time with huge amount
- 9:50:19of data, we can perceive an object like
- 9:50:21a human being does. What I'm trying to
- 9:50:23say over here is like if you want to
- 9:50:25perform a classification task between
- 9:50:27pen and a pencil, you'll obviously know
- 9:50:29as a human being you'll know the
- 9:50:30difference because you have look at a
- 9:50:32pen and a pencil continuous number of
- 9:50:34times. And now when you're trying to
- 9:50:36actually classify it, you can do it with
- 9:50:38ease.
- 9:50:39Okay? And the reason for this is because
- 9:50:40you know the features of a pen and you
- 9:50:42know the features of a pencil. Okay?
- 9:50:44Similarly, this is how deep learning
- 9:50:46works. More the data you feed, more the
- 9:50:48dimensions it can analyze. More the
- 9:50:50dimensions it can learn. All right? So,
- 9:50:52as I've already mentioned, one of the
- 9:50:54most popular application of deep
- 9:50:56learning is image classification. And
- 9:50:58when it comes to image classification,
- 9:51:00it can be something as simple as
- 9:51:01classifying between two different
- 9:51:02animals to something as complicated as,
- 9:51:05you know, hiding data or trying to run
- 9:51:08automated cars using classification
- 9:51:10task. Okay? All right. So, next type of
- 9:51:13application using deep learning is using
- 9:51:15on sequential data. Sequential data
- 9:51:18basically refers to something like time
- 9:51:20series data or having to understand
- 9:51:22natural language.
- 9:51:23So, the reason why we call it sequential
- 9:51:25data is because here the previous word
- 9:51:28or the previous feature is dependent
- 9:51:30upon the next feature. Okay? So, as you
- 9:51:32can see over here, we have what time is
- 9:51:34it, right? So, if I just say it is like
- 9:51:37over here what time is and it are
- 9:51:39basically features, right? And in order
- 9:51:40for you to make an analogy or to
- 9:51:42understand, obviously have to know what
- 9:51:44has happened in the past. So, in order
- 9:51:45to do this, we use something called as
- 9:51:47RNNs. Okay? And there are various
- 9:51:49versions of RNN that go around in order
- 9:51:51to overcome the disadvantages which
- 9:51:53we'll look in in sometime. All right.
- 9:51:56So, moving on to the next application
- 9:51:57that is GANs. GANs, which stands for
- 9:52:00generative adversarial network, is an
- 9:52:02unsupervised part of a deep learning
- 9:52:04application. Some of the common
- 9:52:06application which you can see in recent
- 9:52:07days is nothing but deep fakes and many
- 9:52:09more. Finally, coming down to performing
- 9:52:11classification and regression task using
- 9:52:13multi-layer perceptron. If you remember
- 9:52:15or if you're well versed with machine
- 9:52:17learning, in order to perform
- 9:52:18classification in machine learning, we
- 9:52:20had algorithms like decision tree,
- 9:52:21random forest, or something very simple
- 9:52:24as linear regression or logistic
- 9:52:26regression. But, let me tell you what.
- 9:52:28When we try to perform classification
- 9:52:30using MLP, or multi-layer perceptron, we
- 9:52:32get a very high accuracy even compared
- 9:52:34to SVM and decision trees. All right.
- 9:52:37So, now that we know what exactly is
- 9:52:38deep learning and why we use it, let's
- 9:52:41now stream down to understand how can we
- 9:52:43process natural language data using
- 9:52:45RNNs.
- 9:52:46So, what are RNNs, right? Well, RNN
- 9:52:48basically stands for recurrent neural
- 9:52:50network. And we usually use this in
- 9:52:52order to deal with a sequential data.
- 9:52:54Sequential data can be something like a
- 9:52:56time series data, or a textual data of
- 9:52:58any format. So, why should one use RNN,
- 9:53:00right? Well, this is because there's a
- 9:53:02concept of internal memory here. RNN can
- 9:53:04remember important things about the
- 9:53:06input it has received. Which allows them
- 9:53:09to be very precise in predicting what
- 9:53:11can be the next outcome. So, this is the
- 9:53:13reason why they are performed or
- 9:53:14preferred on a sequential data
- 9:53:16algorithm, okay? And some of the
- 9:53:17examples of sequential data can be
- 9:53:19something like time series, speech,
- 9:53:21text, financial data, audio, video,
- 9:53:23weather, and many more. Although RNN
- 9:53:26were the state-of-the-art algorithm for
- 9:53:27dealing with sequential data, they come
- 9:53:29up with their own drawbacks. And some of
- 9:53:31the popular drawbacks over here can be
- 9:53:33like, due to the complication or the
- 9:53:35complexity of the algorithm, the neural
- 9:53:37network is pretty slow to train. And as
- 9:53:39there are a huge amount of dimensions
- 9:53:40here, the training is very long and
- 9:53:43difficult to do, okay? Apart from that,
- 9:53:45the most decisive feature for RNN, or
- 9:53:47for the improvement in RNN, is that of a
- 9:53:50vanishing gradient. What this vanishing
- 9:53:52gradient is is that, you know, when we
- 9:53:54go deeper and deeper into our neural
- 9:53:56network, the previous data is lost. This
- 9:53:59is because of a concept called as
- 9:54:01vanishing gradient. And due to this, we
- 9:54:03cannot work on a large or a longer
- 9:54:05sequence of data. Okay? To overcome
- 9:54:08this, we came up with some new or
- 9:54:10upgrades to the current recurrent neural
- 9:54:12networks or RNNs.
- 9:54:14Starting off with bidirectional
- 9:54:15recurrent neural network. You see,
- 9:54:17bidirectional recurrent neural network
- 9:54:19connect two hidden layers of opposite
- 9:54:21direction into the same output. With
- 9:54:23this form of generative deep learning,
- 9:54:25the output layer can get information
- 9:54:27from past future states simultaneously.
- 9:54:30So, as you can see here, we have two
- 9:54:32layers over here, and as they are
- 9:54:34bidirectional, what happens is when the
- 9:54:36algorithm feels that it is kind of
- 9:54:37losing its gradients or the previous
- 9:54:39data, it can go back and get the data
- 9:54:41from the past. So, why do we need
- 9:54:44bidirectional recurrent neural network?
- 9:54:45Well, bidirectional recurrent neural
- 9:54:47network duplicates RNN processing chain
- 9:54:50so that the input process both forward
- 9:54:52and reverse time order,
- 9:54:53thus allowing bidirectional recurrent
- 9:54:55neural network to look into future
- 9:54:57context as well. The next one is long
- 9:54:59short-term memory. Long short-term
- 9:55:01memory or also sometime referred to as
- 9:55:03LSTM is a artificial recurrent neural
- 9:55:05network architecture used in the field
- 9:55:07of deep learning. Unlike standard
- 9:55:09feedforward neural network, LSTM has a
- 9:55:11feedback connections. It can not only
- 9:55:13process single data point, but also the
- 9:55:15entire sequence of data. So, as you can
- 9:55:17see here, from what I'm trying to say is
- 9:55:19with LSTM or long short-term memory, it
- 9:55:22has something like, you know, we can
- 9:55:23feed a longer sequence compared to what
- 9:55:25it was with bidirectional RNN or RNNs.
- 9:55:29So, why is LSTM better than RNN? We can
- 9:55:31say that when we move from RNN to LSTM,
- 9:55:34we are introducing more and more control
- 9:55:36over the sequence of the data that we
- 9:55:38can provide. The LSTM gives us more
- 9:55:40control ability and does better results.
- 9:55:43All right. So, the next type of
- 9:55:44recurrent neural network is the gated
- 9:55:46recurrent neural network or also
- 9:55:48referred to as GRUs. You see, GRU is a
- 9:55:50type of recurrent neural network that
- 9:55:52is, in certain cases, is advantageous
- 9:55:55over long short-term memory. GRU makes
- 9:55:57use of less memory and also is faster
- 9:55:59than LSTM. But thing is, LSTMs are more
- 9:56:02accurate while using longer data sets.
- 9:56:05I'm sure by now you might have got a
- 9:56:07hint about the trend that has led to the
- 9:56:09improvement, right? So, the trend over
- 9:56:11here is, you know, the model should be
- 9:56:13capable of remembering and taking in on
- 9:56:16a longer input sequence.
- 9:56:18The game-changer part for the sequential
- 9:56:20data was developed when we came up with
- 9:56:22something called as transformers. And
- 9:56:24this paper was something which is based
- 9:56:26on a concept called as attention is
- 9:56:29everything.
- 9:56:30All right. So, let's take a look at
- 9:56:31this.
- 9:56:32The paper attention is all you need
- 9:56:35introduces a novel architecture called
- 9:56:37as transformers. Like LSTM, transformers
- 9:56:40is an architecture for transforming one
- 9:56:42sequence into another while helping
- 9:56:44adapt to parts, that is encoders and
- 9:56:46decoders. But it differs from previously
- 9:56:48described sequence to sequence model
- 9:56:50because it does not work like GRUs,
- 9:56:52okay? So, it does not implements uh
- 9:56:55recurrent neural networks.
- 9:56:57Recurrent neural network until now were
- 9:56:59one of the best ways to capture the
- 9:57:00timely dependence on a sequence.
- 9:57:03However, the team presenting this paper,
- 9:57:05that is attention is all you need,
- 9:57:06proved that an architecture with only
- 9:57:08attention mechanism does not use RNN can
- 9:57:11improve its result in translation task
- 9:57:14and other NLP task. One of the best
- 9:57:16examples for transformers is Google's
- 9:57:18BERT. So, what exactly is this
- 9:57:20transformer, right? You see, here we
- 9:57:22have encoder on the top and decoder on
- 9:57:24the bottom. Both encoder and decoder are
- 9:57:26comprised of modules that can stick onto
- 9:57:29the top of each other multiple times.
- 9:57:31So, what happens here is the inputs and
- 9:57:33outputs are first embedded into
- 9:57:35N-dimension space since we cannot use
- 9:57:37this directly. So, we obviously have to
- 9:57:39encode our inputs, whatever we are
- 9:57:41providing here. One slight but important
- 9:57:43part of this model is the positional
- 9:57:45encoding of different words. Since we
- 9:57:47have no recurrent neural network that
- 9:57:49can remember how sequence are fed into
- 9:57:51the model, we need to somehow give every
- 9:57:53word or part of our sequence a relative
- 9:57:55position since the sequence depends on
- 9:57:58the order of the elements, okay? These
- 9:58:00positions are added to the embedded
- 9:58:02representation of each words. All right.
- 9:58:04So, this was the brief about
- 9:58:06transformers. So, let us now move ahead
- 9:58:08and see some of the popular language
- 9:58:10models that are available in the market.
- 9:58:12All right. So, let us now start off by
- 9:58:14understanding OpenAI's GPT-3. The
- 9:58:16successor to GPT and GPT-2 is the GPT-3
- 9:58:20and is one of the most controversial
- 9:58:22pre-trained models by OpenAI. The
- 9:58:24large-scale transformer-based language
- 9:58:26model has been trained on 175 billion
- 9:58:29parameters, which is 10 times more than
- 9:58:31any previous non-sparse language model.
- 9:58:34The model has been trained to achieve
- 9:58:36strong performance on many NLP data set,
- 9:58:38including tasks like translation,
- 9:58:41answering questions, as well as several
- 9:58:42other tasks. Then we have Google's BERT.
- 9:58:45BERT stands for bidirectional encoder
- 9:58:47representations from transformers. It is
- 9:58:50a pre-trained NLP model, which is
- 9:58:52developed by Google in 2018. With this,
- 9:58:54anyone in the world can train either
- 9:58:56their own question answering module with
- 9:58:58up to 30 minutes on a single cloud TPU
- 9:59:01or few hours using single GPU. The
- 9:59:04company then released this showcasing
- 9:59:06the performance of 11 NLP tasks,
- 9:59:08including very competitive Stanford
- 9:59:10dataset questions.
- 9:59:12Unlike other language model, BERT has
- 9:59:14only been pre-trained on 250 million
- 9:59:16words of Wikipedia and 800 million words
- 9:59:18of book corpus and has been successfully
- 9:59:21used as a pre-trained model in deep
- 9:59:23neural network. According to
- 9:59:24researchers, BERT has achieved 93%
- 9:59:27accuracy, which has surpassed any
- 9:59:28previous language models.
- 9:59:30Next, we have ELMo. ELMo, also known as
- 9:59:33embedding for language model, is a deep
- 9:59:35contextualized word representation that
- 9:59:38models syntax and semantic words, as
- 9:59:40well as their logistic context. The
- 9:59:42model developed by Allen NLP has been
- 9:59:44pre-trained on a huge text corpus and
- 9:59:47learned functions from bidirectional
- 9:59:49models, that is BiLM. ELMo can easily be
- 9:59:52added to their existing models, which
- 9:59:54drastically improves the features of
- 9:59:56functions across vast NLP problem,
- 9:59:59including answering questions, textual
- 10:00:01entailment, and sentiment analysis.
- 10:00:06>> [music]
- 10:00:09>> What are GANs?
- 10:00:11So, we're going to start with generative
- 10:00:13models.
- 10:00:14So, generative models are nothing but
- 10:00:16those models that use an unsupervised
- 10:00:19learning approach.
- 10:00:20In a generative model, there are samples
- 10:00:23in the data that is input variables X,
- 10:00:26but it lacks a output variable Y. And we
- 10:00:29use the only input variables to train
- 10:00:31the generative model, and it recognizes
- 10:00:34patterns from the input variables to
- 10:00:36generate an output that is unknown and
- 10:00:39based on the training data only.
- 10:00:41In supervised learning, we are more
- 10:00:43aligned towards creating predictive
- 10:00:45models from the input variables.
- 10:00:48And this type of modeling is also known
- 10:00:50as discriminative modeling.
- 10:00:52And in a classification problem, the
- 10:00:54model has to discriminate as to which
- 10:00:57class the example belongs to. And on the
- 10:00:59other hand, unsupervised models are used
- 10:01:01to create or generate new examples in
- 10:01:04the input distribution.
- 10:01:06To define a generative model in layman
- 10:01:09terms, we can say generative models are
- 10:01:12able to generate new examples from the
- 10:01:15sample that are not only similar to the
- 10:01:17examples, but are indistinguishable as
- 10:01:20well.
- 10:01:21And the most common example of a
- 10:01:22generative model is a naive Bayes
- 10:01:24classifier, which is more often used as
- 10:01:27a discriminative model.
- 10:01:29Other examples of generative models
- 10:01:31include Gaussian mixture model and a
- 10:01:33rather modern example, that is
- 10:01:35generative adversarial networks.
- 10:01:38So, let us try to understand what
- 10:01:39exactly are GANs, or generative
- 10:01:42adversarial networks.
- 10:01:44Generative adversarial networks, or
- 10:01:46GANs, are a deep learning-based
- 10:01:48generative model that is used for
- 10:01:50unsupervised learning.
- 10:01:52It is basically a system where two
- 10:01:54competing neural networks compete with
- 10:01:56each other to create or generate
- 10:01:58variations in the data.
- 10:02:00It was first described in a paper in
- 10:02:022014 by Ian Goodfellow and a
- 10:02:05standardized and much stable model
- 10:02:07theory was proposed by Alec Radford in
- 10:02:102016, which is also known as DCGAN.
- 10:02:15Also known as DCGAN or we can call it as
- 10:02:18deep convolutional generative
- 10:02:20adversarial networks.
- 10:02:22And most of the GANs today use deep
- 10:02:24convolutional generative adversarial
- 10:02:25networks.
- 10:02:27The GANs architecture consists of two
- 10:02:29sub models known as the generator model
- 10:02:32and the discriminator model.
- 10:02:34So, a generator network takes a sample
- 10:02:36and generates sample of data.
- 10:02:38A discriminator network decides whether
- 10:02:40the data is generated or taken from the
- 10:02:42real sample using a binary
- 10:02:44classification problem with the help of
- 10:02:46a sigmoid function that gives the output
- 10:02:49in the form or the range zero and one.
- 10:02:52So, let us go ahead and take a look at
- 10:02:54how GANs actually work.
- 10:02:57To understand how GANs work, let's break
- 10:02:59it down.
- 10:03:00So, generative means that the model
- 10:03:02follows the unsupervised learning
- 10:03:04approach and is a generative model.
- 10:03:07When we talk about adversarial, the
- 10:03:09model is trained in an adversarial
- 10:03:11setting.
- 10:03:12And network simply means for the
- 10:03:14training of the model, we use the neural
- 10:03:16networks as artificial intelligence
- 10:03:18algorithms.
- 10:03:20In GANs, there is a generator network
- 10:03:22that takes a sample and generates a
- 10:03:24sample of data.
- 10:03:26And after this, the discriminator
- 10:03:27network decides whether the data is
- 10:03:29generated or taken from the real sample
- 10:03:31using a binary classification problem
- 10:03:33with the help of a sigmoid function that
- 10:03:36gives the output in the range zero to
- 10:03:37one.
- 10:03:38The generative model analyzes the
- 10:03:40distribution of the data in such a way
- 10:03:42that after the training phase, the
- 10:03:44probability of the discriminator making
- 10:03:46a mistake maximizes. and the
- 10:03:48discriminator on the other hand is based
- 10:03:50on a model that will estimate the
- 10:03:52probability that the sample is coming
- 10:03:54from the real data or not the generator.
- 10:03:57The whole process can be formalized in a
- 10:03:59mathematical formula.
- 10:04:01So G over here is generator, D is equal
- 10:04:04to discriminator, P data X is the
- 10:04:07distribution of real data, P data Z is
- 10:04:10the distributor of generator, X is the
- 10:04:13sample from the real data, and Z is the
- 10:04:16sample from generator. Where DX is the
- 10:04:18discriminator network and GZ is a
- 10:04:21generator network.
- 10:04:23So let's take a look at the flowchart
- 10:04:24once again, guys.
- 10:04:25So we have the training data, which is
- 10:04:27going to give the real sample. And the
- 10:04:29generator network is going to generate
- 10:04:31the sample from the random noise or the
- 10:04:33examples.
- 10:04:34And then it will go to the discriminator
- 10:04:36network, where it's going to check if
- 10:04:38the sample that is coming is real or
- 10:04:41fake.
- 10:04:42So that is how a GAN actually work.
- 10:04:44Now let's take a look at the training
- 10:04:46phase, like how a generative adversarial
- 10:04:48network is actually trained.
- 10:04:50So it happens in two phases, guys.
- 10:04:53So the first phase is where we train the
- 10:04:55discriminator and we actually freeze the
- 10:04:57generator, which means that the training
- 10:05:00set for the generator is done false and
- 10:05:02the network will only do the forward
- 10:05:04pass and there will not be any back
- 10:05:06propagation.
- 10:05:08Basically, the discriminator is trained
- 10:05:10with real data and checks if it can
- 10:05:12predict them correctly.
- 10:05:14And the same with the fake data to
- 10:05:15identify them as fake.
- 10:05:18After this, there's the second part
- 10:05:20where we train the generator and freeze
- 10:05:22the discriminator.
- 10:05:24So we get the result from the first
- 10:05:25phase and we use them to make better
- 10:05:28from the previous state to try and fool
- 10:05:30the discriminator better.
- 10:05:32So to understand this in the layman's
- 10:05:33term, I'm going to tell you a few steps
- 10:05:35for training, like how you should start.
- 10:05:38So the first step is you have to define
- 10:05:40in problem.
- 10:05:41You've got to define the problem and
- 10:05:43collect the data.
- 10:05:44After this, the second step is you have
- 10:05:47to choose the architecture of GAN.
- 10:05:49So, in this step, depending on your
- 10:05:50problem, you have to choose how your GAN
- 10:05:52should look like.
- 10:05:54The third step is training the
- 10:05:55discriminator on real data. So, we train
- 10:05:58the discriminator with real data to
- 10:06:00predict them as real for n number of
- 10:06:02times, so we call it a epochs as well.
- 10:06:04And then we generate the fake inputs
- 10:06:06from the generator.
- 10:06:08So, in this step, we are going to
- 10:06:09generate the fake samples from the
- 10:06:11generator.
- 10:06:12And the next step is we train the
- 10:06:14discriminator on fake data.
- 10:06:17So, whatever samples are generated from
- 10:06:18the generator network, you're going to
- 10:06:20train the discriminator to predict the
- 10:06:21generated data as fake.
- 10:06:24So, that's how we know that
- 10:06:25discriminator is actually predicting the
- 10:06:27values as correctly.
- 10:06:28And the last step is we train the
- 10:06:30generator with the output of
- 10:06:31discriminator. So, after getting the
- 10:06:33discriminator predictions, we train the
- 10:06:36generator to fool the discriminator.
- 10:06:39So, that's how we train the GAN to
- 10:06:41actually get our solution from the
- 10:06:43problem. Which is like defining the
- 10:06:45problem.
- 10:06:46So, you'll understand this when I'm
- 10:06:47talking about the applications, guys. No
- 10:06:49worry.
- 10:06:50Now, let's go ahead and take a look at a
- 10:06:51few challenges of generative adversarial
- 10:06:54networks.
- 10:06:55So, the concept of GANs is rather
- 10:06:57fascinating, but there are a lot of
- 10:07:00setbacks that can cause a lot of
- 10:07:01hindrance in its path.
- 10:07:03Some of the major challenges faced by
- 10:07:05GANs are
- 10:07:06The first one is the stability.
- 10:07:08So, there has to be a stability that is
- 10:07:09required between discriminator and the
- 10:07:11generator network, otherwise the whole
- 10:07:13network would just fall.
- 10:07:15For example, in case, let's say if the
- 10:07:17discriminator is too powerful, the
- 10:07:20generator will fail to train altogether.
- 10:07:22Won't be able to push fake samples to
- 10:07:24that discriminator, and it will always
- 10:07:27identify them as fake.
- 10:07:29And let's say if the network is too
- 10:07:30lenient, the discriminator network is
- 10:07:32too lenient,
- 10:07:34so any image that would be generated by
- 10:07:36the generator network would make the
- 10:07:38network useless.
- 10:07:39The next challenge that is faced by GANs
- 10:07:42is GANs fail miserably in determining
- 10:07:44the positioning of the objects in terms
- 10:07:46of how many times the objects should
- 10:07:48occur at that location. Suppose we have
- 10:07:50a image in which we have, let's say,
- 10:07:52three dogs with two eyes and sometimes a
- 10:07:56GAN will fail to, you know, determine
- 10:07:58the positioning of the objects in terms
- 10:07:59of it will generate an image with like
- 10:08:01one dog and six eyes. So, that's kind of
- 10:08:04a problem that we face while working on
- 10:08:06GANs.
- 10:08:07And the next challenge is 3D perspective
- 10:08:10troubles GANs as it is not able to
- 10:08:12understand the perspective also.
- 10:08:15So, it will often give a flat image for
- 10:08:16a 3D object. So, that's one challenge
- 10:08:19that we face with GANs as well.
- 10:08:21And GANs have a problem of understanding
- 10:08:24the global objects and it cannot
- 10:08:26differentiate or understand a holistic
- 10:08:28structure.
- 10:08:29Like, if you're talking about trees or
- 10:08:31if you're talking about flowers, that's
- 10:08:33a problem that GANs will follow.
- 10:08:35And last but not least, newer types of
- 10:08:37GANs are more advanced that are brought
- 10:08:39about that is deep convolutional
- 10:08:41generative adversarial networks and are
- 10:08:44expected to overcome these shortcomings
- 10:08:46altogether. So, that we don't have to
- 10:08:47worry about these. These are the
- 10:08:49shortcomings that we face with normal
- 10:08:51GANs, uh initial generative adversarial
- 10:08:53networks. Now that they have become more
- 10:08:55advanced, they actually overcome these
- 10:08:58shortcomings, so you don't have to
- 10:08:59worry, guys.
- 10:09:00So, last but not the least, I want to
- 10:09:02talk about a few applications of
- 10:09:04generative adversarial networks.
- 10:09:06So, the first one is prediction of next
- 10:09:08frame in a video.
- 10:09:10So, let's say the prediction of future
- 10:09:11events in a video frame is made possible
- 10:09:13with the help of GANs and DVD GAN or we
- 10:09:17can call it as dual video discriminator
- 10:09:19GAN can generate a 256 by 256 videos of
- 10:09:23notable fidelity up to 48 frames in
- 10:09:26length.
- 10:09:27And this can be used for various
- 10:09:28purposes including surveillance in which
- 10:09:31we can determine the activities in a
- 10:09:32frame that gets distorted due to other
- 10:09:35factors like rain, dust, smoke, etc.
- 10:09:38So, the possibilities are immense with
- 10:09:40this if you're able to predict the next
- 10:09:42frame in a video. That actually helps in
- 10:09:44a lot of things like surveillance,
- 10:09:45security, and we can predict outcomes
- 10:09:48based on these frames that we generate
- 10:09:50from a video.
- 10:09:51After this comes the text to image
- 10:09:53generation.
- 10:09:54So, basically object-driven attentive
- 10:09:56GAN, which is also known as object GAN,
- 10:09:58performs the text to image synthesis in
- 10:10:01two steps. So, the first step is
- 10:10:03generating the semantic layout and then
- 10:10:06generating the image by synthesizing the
- 10:10:07image by using a deconvolutional image
- 10:10:10generator is the final step.
- 10:10:13So, this could be used intensively to
- 10:10:14generate images by understanding the
- 10:10:16captions, the layouts, and refine
- 10:10:18details by synthesizing the words.
- 10:10:21And there is another study about the
- 10:10:23story GANs that can synthesize the whole
- 10:10:25storyboards from mere paragraphs.
- 10:10:28So, that's actually very good idea if
- 10:10:30you're talking about GANs. So, you can
- 10:10:31just give a few layouts and captions.
- 10:10:35Based on that, it will generate image
- 10:10:36for us.
- 10:10:38Talking about the next application, we
- 10:10:39have image to image translation.
- 10:10:42So, Pix2Pix is a model which is designed
- 10:10:44for general purpose image to image
- 10:10:46translation.
- 10:10:48So, let's say we have three images.
- 10:10:50We have a real image.
- 10:10:51Then we'll be having a generated image,
- 10:10:54which is basically a fake, and then it
- 10:10:56will be reconstructed to the previous
- 10:10:58image which was real.
- 10:11:00So, this is how image to image
- 10:11:01translation work, guys.
- 10:11:03And after this, we have enhancing the
- 10:11:04resolution of an image.
- 10:11:06So, super-resolution generative
- 10:11:08adversarial network, or also known as
- 10:11:10SRGAN, is a GAN which can generate the
- 10:11:14super-resolution images from
- 10:11:16low-resolution images with finer details
- 10:11:18and better quality.
- 10:11:20So, this is actually a very good
- 10:11:22application of GANs, guys. The
- 10:11:24applications can be immense.
- 10:11:26So, you imagine a higher quality image
- 10:11:29with finer details generated from a low
- 10:11:31resolution image. The amount of help it
- 10:11:34would produce to identify details in
- 10:11:36lower resolution images can be used for
- 10:11:38wider purposes including surveillance.
- 10:11:41We can use it for documentation
- 10:11:43security. We can use it for detecting
- 10:11:45patterns, etc.
- 10:11:47And last but not least, we have
- 10:11:48interactive image generation.
- 10:11:51So, GANs can be used to generate
- 10:11:53interactive images as well.
- 10:11:55And computer science and artificial
- 10:11:56intelligence laboratory also known as
- 10:11:58CSAIL
- 10:11:59has developed a GAN that can generate 3D
- 10:12:02models with realistic lighting and
- 10:12:04reflections enabled by the shape and
- 10:12:06texture editing.
- 10:12:08And more recently, researchers have come
- 10:12:10up with a model that can synthesize a
- 10:12:13re-enacted face animated by a person's
- 10:12:16movement while preserving the appearance
- 10:12:19of the face at the same time.
- 10:12:21There are a lot more applications we can
- 10:12:23work on.
- 10:12:25>> [music]
- 10:12:30>> Evolution of AI. So, AI as we know today
- 10:12:33is entirely different from where it
- 10:12:34started. Back in the 18th century, it
- 10:12:37was entirely based on myths,
- 10:12:39speculations, and fiction. But, it did
- 10:12:41start taking shape and the real
- 10:12:43initiation in its truest essence took
- 10:12:45place in 1956.
- 10:12:47The AI search began with six major
- 10:12:49design goals. The first one was teach
- 10:12:51the machines to reason in accordance to
- 10:12:53perform sophisticated mental tasks like
- 10:12:55playing chess, providing mathematical
- 10:12:57theorems, and others. The second one is
- 10:13:00knowledge representation for machines to
- 10:13:02interact with the real world as humans
- 10:13:04do. Like machines needed to be able to
- 10:13:06identify objects, people, and languages.
- 10:13:09Programming language Lisp was developed
- 10:13:11for this very purpose.
- 10:13:13The third one is teach the machines to
- 10:13:15plan and navigate around the world we
- 10:13:17live in. With this, machines could
- 10:13:18autonomously move around by navigating
- 10:13:21themselves.
- 10:13:22The fourth one is enable the machines to
- 10:13:24process natural language so that they
- 10:13:26can understand the language,
- 10:13:28conversations, and the context of
- 10:13:29speech.
- 10:13:31The fifth one is train the machines to
- 10:13:33perceive the way humans do, like touch,
- 10:13:36feel, sight, hearing, and taste.
- 10:13:38And general intelligence that included
- 10:13:40emotional intelligence, intuition, and
- 10:13:42creativity was the sixth point. Talking
- 10:13:44about machine learning,
- 10:13:46machine learning as we know can be
- 10:13:47remotely explained with the evolution of
- 10:13:49robots in the past years. Although,
- 10:13:51machine learning isn't just a machine
- 10:13:53that is going to learn stuff. It has a
- 10:13:55lot more to it. Basically, we have data
- 10:13:57at our bay. We train and test the model,
- 10:14:00which in this case can be a robot, and
- 10:14:02then make it to do task relevant to the
- 10:14:04learning. And then again, learning can
- 10:14:05be of different types, which is
- 10:14:07supervised, unsupervised, reinforcement,
- 10:14:09etc. To know more about machine learning
- 10:14:11in detail, refer to our machine learning
- 10:14:13full course tutorial to get on speed.
- 10:14:15Now, let us go ahead and take a look at
- 10:14:17what exactly is AI and machine learning.
- 10:14:20So, what exactly is AI? According to the
- 10:14:22Merriam-Webster dictionary, artificial
- 10:14:24intelligence is a branch of computer
- 10:14:25science dealing with the simulation of
- 10:14:27intelligent behavior in computers.
- 10:14:30AI is a technique that enables machines
- 10:14:32to mimic human behavior. Artificial
- 10:14:35intelligence is the theory and
- 10:14:36development of computer systems able to
- 10:14:38perform task normally requiring human
- 10:14:40intelligence, such as visual perception,
- 10:14:43speech recognition, decision-making, and
- 10:14:45translation between languages.
- 10:14:47If you ask me, AI is the simulation of
- 10:14:49human intelligence done by machines
- 10:14:51programmed by us. The machines need to
- 10:14:54learn how to reason and do some
- 10:14:56self-correction as needed along the way.
- 10:14:58And artificial intelligence is
- 10:14:59accomplished by studying how human brain
- 10:15:02thinks, learns, and decide to work while
- 10:15:04trying to solve a problem. And then
- 10:15:06using the outcomes of this study as a
- 10:15:07basis of developing intelligent software
- 10:15:09and systems.
- 10:15:11Now, let us go ahead and take a look at
- 10:15:12what exactly is machine learning. So,
- 10:15:14machine learning is a concept which
- 10:15:16allows the machines to learn from
- 10:15:18examples and experiences.
- 10:15:20And that too without being explicitly
- 10:15:22programmed. So, instead of you writing
- 10:15:24the code, what you do is you feed the
- 10:15:26data to the generic algorithm and the
- 10:15:29algorithm or the machine builds the
- 10:15:30logic based on the given data.
- 10:15:32Machine learning algorithms are an
- 10:15:34evolution of normal algorithms and they
- 10:15:36make your program smarter by allowing
- 10:15:38them to automatically learn from the
- 10:15:40data that you provide. Now that we know
- 10:15:42what AI and ML actually is, let us go
- 10:15:44ahead and take a look at a few
- 10:15:45applications of AI and ML today.
- 10:15:47So, AI and ML can be widely used in so
- 10:15:49many applications and I have listed down
- 10:15:52a few applications to give you a wider
- 10:15:53perspective. So, first of all, I'm going
- 10:15:55to talk about the usage of AI and ML in
- 10:15:57healthcare.
- 10:15:58AI and ML in healthcare is an angel in
- 10:16:01disguise. To understand this, imagine
- 10:16:03you have the past data of millions of
- 10:16:05patients with diseases in the past. Now,
- 10:16:07all this data can be put to an effective
- 10:16:10use in a sense that we would be able to
- 10:16:12detect a disease in early stages with
- 10:16:15the help of machine learning and
- 10:16:16artificial intelligence algorithms.
- 10:16:18And since the numbers never lie, we can
- 10:16:21be pretty sure about the accuracy of the
- 10:16:22results. Even so, if we have any doubts,
- 10:16:25we can always check the accuracy in
- 10:16:27almost all the cases.
- 10:16:29Now, let's talk about the usage of AI
- 10:16:31and ML in finance. So, artificial
- 10:16:33intelligence in finance is transforming
- 10:16:35the way we interact with money.
- 10:16:37AI is helping the financial industry to
- 10:16:39streamline and optimize processes
- 10:16:42ranging from credit decisions to
- 10:16:44quantitative trading and financial risk
- 10:16:46management. And if we have that figured
- 10:16:48out, it saves us from a lot of bad and
- 10:16:51risky decisions.
- 10:16:53Now, let us go ahead and take a look at
- 10:16:54object detection in which we use AI and
- 10:16:56ML. So, object detection today is
- 10:16:58playing an important part in the IT
- 10:17:00industry. For example, surveillance has
- 10:17:02never looked more tech-savvy. Google
- 10:17:04Lens is one example that uses image
- 10:17:06recognition to identify the images on
- 10:17:08the camera in real time.
- 10:17:09Now, let us go ahead and take a look at
- 10:17:11AI and ML in risk detection and
- 10:17:13predictive analysis. So, predictive
- 10:17:15analysis has proven its metal in the
- 10:17:16industry already with almost every
- 10:17:18organization using it to derive
- 10:17:19conclusions based on previous data. So,
- 10:17:22one example is how sports franchises,
- 10:17:25such as a cricket team, would take
- 10:17:26account of the performance of a player,
- 10:17:28let's say a batsman, facing the
- 10:17:30deliveries of bouncer and short ball
- 10:17:32deliveries. So, they will be able to
- 10:17:34figure out the best possible outcomes
- 10:17:36based on the short selection and
- 10:17:38strategies using the artificial
- 10:17:40intelligence and machine learning
- 10:17:41algorithms.
- 10:17:42So, this is one way we can use
- 10:17:43predictive analysis or engine it's just
- 10:17:45one example. We can use it for many
- 10:17:47purposes. Like, we can use it to predict
- 10:17:49stock prices that we can do in finance
- 10:17:51and we can use it to predict the weather
- 10:17:53based on the hundreds and hundreds of
- 10:17:55years of data that we already have.
- 10:17:57Now, let's talk about AI and ML in
- 10:17:58marketing and advertising. So, marketing
- 10:18:01and advertising industry is the most
- 10:18:03benefited with the evolution of AI and
- 10:18:05ML. They are able to recognize the
- 10:18:07browsing patterns of users through data
- 10:18:09and target users with specific content
- 10:18:11on the internet. For example, to reach
- 10:18:13out in a subtle way, you only see those
- 10:18:15ads which interest you. And you must
- 10:18:17have felt sometimes like you keep
- 10:18:19getting ads or you know the content on
- 10:18:21the internet that you talk about or you
- 10:18:22were talking about or you were thinking
- 10:18:24about. It's basically nothing but AI and
- 10:18:26machine learning that is learning a
- 10:18:28pattern through the data and the your
- 10:18:30browsing history or your browsing
- 10:18:31pattern and reaches out to you in the
- 10:18:33form of targeted marketing.
- 10:18:35And now that we have talked about the
- 10:18:36current applications, let us talk about
- 10:18:38how AI and ML will shape up in the
- 10:18:40future and how it would look like 10
- 10:18:42years, maybe 20 years from now.
- 10:18:44So, we cannot be sure about if AI and ML
- 10:18:45would shape in the future like we have
- 10:18:47seen in the movies, but for now, we will
- 10:18:49stick to the realistic possibilities.
- 10:18:51Although when I say realistic
- 10:18:52possibilities, we are not really sure
- 10:18:54how it would look like, but we can take
- 10:18:56a guess.
- 10:18:57So, when we talk about future of health
- 10:18:59care, health care would seem pretty
- 10:19:01reachable and advanced. Even now,
- 10:19:03researchers are working on detecting
- 10:19:05diseases in the early stages based on
- 10:19:07the lifestyle and other relevant data.
- 10:19:09And in the future, we can expect more
- 10:19:10advancements in the psychological part
- 10:19:12as well, where we will be able to
- 10:19:14identify traits and warnings in the
- 10:19:16early stages and work on it before it
- 10:19:18gets better off us.
- 10:19:20Now imagine being able to cure a disease
- 10:19:22before even getting the hint of it. That
- 10:19:24is what researchers are aiming for and
- 10:19:26it looks pretty promising, guys. Let me
- 10:19:28tell you. Now let's talk about the
- 10:19:29future of self-driving cars. So
- 10:19:31self-driving cars looks like a dream
- 10:19:33come true today. But in the coming
- 10:19:35decades, we're going to see a lot of
- 10:19:37developments in the self-driving cars.
- 10:19:39We would have overcome all the
- 10:19:40challenges that we face today and I'm
- 10:19:42not saying we will have levitating cars
- 10:19:44driving you to your destinations, but it
- 10:19:47won't be less than a fascinating
- 10:19:48experience that may look like a dream
- 10:19:49today or you might have seen in the
- 10:19:50movies.
- 10:19:52Now let's talk about the AI ML in
- 10:19:53manufacturing that would take the
- 10:19:55future.
- 10:19:56So robots in manufacturing is one thing
- 10:19:58that is going to change the future for
- 10:20:00us.
- 10:20:01Manufacturing industries will have the
- 10:20:02best ever workforce and I'm not talking
- 10:20:05about the human aspect of it. The
- 10:20:06manufacturing would be so much easier
- 10:20:08with the robots and since they don't get
- 10:20:10tired and they won't even ask for leaves
- 10:20:13or they might. We We never know. And AI
- 10:20:15and robotics is a risky slope, although,
- 10:20:18but that is not entirely true.
- 10:20:19Researchers and experts are working day
- 10:20:22and night to make it as safe as
- 10:20:24possible.
- 10:20:25And then again, let me talk about future
- 10:20:26of finance with AI and ML. So managing
- 10:20:29finance and risk detection would become
- 10:20:31a piece of cake and to understand this
- 10:20:33in layman terms, you will be able to do
- 10:20:35your taxes without even lifting a pen.
- 10:20:37And the fraud detection and trading
- 10:20:38would become a lot easier and
- 10:20:40accessible.
- 10:20:41Financial advisory would take a much
- 10:20:43advanced shape as we are already seeing
- 10:20:45it with a lot of trading applications in
- 10:20:47the market.
- 10:20:48Then again, we have computer vision
- 10:20:49which is going to change the future for
- 10:20:50us. And I'm sure most of you are aware
- 10:20:52of the concept of God's eye that we have
- 10:20:54already seen in the movies. Although it
- 10:20:56is fictional, but not sure for the wrong
- 10:20:57reasons, but for the greater good, this
- 10:20:59might be the possibility in the coming
- 10:21:01years as computer vision has started to
- 10:21:03overcome a lot of challenges in the
- 10:21:04real-time image recognition.
- 10:21:06And then we have the future of NLP
- 10:21:08natural language processing and it's
- 10:21:10going to you know it would open a lot of
- 10:21:12linguistic barriers in the
- 10:21:13conversational AI. The conversational AI
- 10:21:15that we see today is limited to certain
- 10:21:17tasks, but in coming years it could be
- 10:21:19like a personal assistant or even a life
- 10:21:21guide as well.
- 10:21:23And with the recent advancements we are
- 10:21:24aiming for a very flexible interface
- 10:21:27that is going to work for everyone with
- 10:21:29any linguistic experience or any
- 10:21:31linguistic expectations.
- 10:21:34And then we have a rather fictional
- 10:21:36concept that I'm going to talk about
- 10:21:37which is immortality through AI. So
- 10:21:39there are scientists and researchers who
- 10:21:41are trying to figure out a way to map
- 10:21:43the brain simulation on a computer. So
- 10:21:45immortality isn't just living until the
- 10:21:47very eternity, it is in my opinion
- 10:21:49leaving a legacy and but in hindsight
- 10:21:52this task is pretty impossible, but we
- 10:21:53never know in the future this might be a
- 10:21:55possibility and we will be able to live
- 10:21:57through a computer where people would
- 10:21:59have figured out a way to shift all of
- 10:22:01our brain simulations onto a computer
- 10:22:04and we'll be able to think on its own
- 10:22:05like our own very image. So that is one
- 10:22:08possibility with the future in AI and ML
- 10:22:11that many researchers are actually
- 10:22:13aiming for and we might as well get
- 10:22:15through with it. So hang in there guys
- 10:22:18and now let me just talk about a key
- 10:22:20skills of AI and ML specialist. There
- 10:22:22was a lot of applications that I just
- 10:22:24told you about. Now let's take a look at
- 10:22:25what are the skill sets that are
- 10:22:27required to become a AI ML specialist in
- 10:22:29today's world.
- 10:22:30So first of all you have to be familiar
- 10:22:31with programming in Python
- 10:22:33and you must be very well aware of maths
- 10:22:35and statistics as well because it needs
- 10:22:36a lot of logic to build algorithms and
- 10:22:38understand them and there's a lot of
- 10:22:40applied mathematics behind it as well.
- 10:22:42So you have to be familiar with maths
- 10:22:43and statistics and programming language
- 10:22:45which is Python and the versioning tools
- 10:22:47like TensorFlow, Keras, PyTorch, etc.
- 10:22:50And then you must have an expertise in
- 10:22:51one of the following machine learning
- 10:22:53domains which is image processing,
- 10:22:55computer vision, language processing,
- 10:22:56speech signal processing, etc. And then
- 10:22:59you must have an advanced knowledge in
- 10:23:00deep learning algorithms as well because
- 10:23:02it is a very important aspect of AIML.
- 10:23:04And there has to be an effective
- 10:23:06communication skills and a you have to
- 10:23:07be a problem solver because in
- 10:23:09hindsight, if you get a problem, the
- 10:23:11only thing that an employer seeks from
- 10:23:13you is the solution of the problem. So,
- 10:23:15anyways, you have to be a problem solver
- 10:23:17in that skill set. And you must be
- 10:23:20experienced in data visualization tools
- 10:23:21and methods like Tableau, Matplotlib,
- 10:23:23Power BI, etc. And there has to be a
- 10:23:25knowledge in ensemble and online
- 10:23:27learning. And you must have an
- 10:23:29experience with SQL and other database
- 10:23:31related languages. And you have to be
- 10:23:33familiar with parallel computing using
- 10:23:35GPUs and unique big systems. So, these
- 10:23:37are the key skills that are required to
- 10:23:39become an AIML specialist, guys.
- 10:23:41Now, let me just walk you through the
- 10:23:42market trends that looks pretty
- 10:23:44promising for now. So, first of all, I'm
- 10:23:46going to talk about the market trends
- 10:23:47and it looks pretty solid, guys, with
- 10:23:49the amount of data flowing in each year
- 10:23:51that it is pretty obvious that it will
- 10:23:53be opening a lot of doors for skilled
- 10:23:54professionals. And the correct approach
- 10:23:56is to get skilled since it is still in
- 10:23:58the evolution phase and we have to scale
- 10:24:01a lot more possibilities in these
- 10:24:02domains. So, if you have the skill set
- 10:24:04for it, it's going to be a very bright
- 10:24:06future for you guys. And to be specific,
- 10:24:08there are a lot of opportunities in
- 10:24:09health care, finance, conversational AI
- 10:24:12or we can call it chatbots, and object
- 10:24:14detection, etc. Almost every industry
- 10:24:16would move to automating their
- 10:24:17processes. And what else than machine
- 10:24:19learning and AI to do your job? So, it
- 10:24:21is the best time to learn AI and ML if
- 10:24:23you are looking for a bright career
- 10:24:25right now. And let us go ahead and take
- 10:24:27a look at the salary trends as well so
- 10:24:28you get the perspective of how much
- 10:24:29you're going to get paid. So, I have
- 10:24:31categorized the salary trends in a few
- 10:24:33job profiles in AI and ML. For a machine
- 10:24:35learning engineer, the takeaway fruits
- 10:24:37of your labor would look around $114,000
- 10:24:40a year. And for a machine learning
- 10:24:41scientist, the average salary looks
- 10:24:43around $120,000 a year and can go as
- 10:24:46high as $150,000 a year.
- 10:24:48And for an AI engineer, the average
- 10:24:50salary is around a $90,000, but it can
- 10:24:53go as high as $140,000 a year as well.
- 10:24:56And for an AI researcher, it goes from a
- 10:24:58$125,000 average to as high as $150,000
- 10:25:03a year. And now let us go ahead and take
- 10:25:04a look at a few companies that are
- 10:25:06hiring right now for AI and ML
- 10:25:07specialists. And although there are a
- 10:25:09lot more companies that are hiring for
- 10:25:11AI and ML specialists right now, I've
- 10:25:12just listed down a few over here. So, we
- 10:25:14have Ford Motors, we have Capgemini,
- 10:25:16Accenture, Dell, Deloitte, Google,
- 10:25:18Amazon. And there are a lot of startups
- 10:25:20as well which are actually artificial
- 10:25:22intelligence and machine learning based.
- 10:25:24So, there's a lot of opportunity for you
- 10:25:26guys. And let me just tell you the best
- 10:25:28approach to actually get a job in AI and
- 10:25:30ML industry. So, the best approach to
- 10:25:33find a job as an AI and ML specialist,
- 10:25:35even if you are a beginner or an
- 10:25:37experienced professional, this works for
- 10:25:39everyone. So, first of all, you have to
- 10:25:41start with a programming language,
- 10:25:42preferably Python, because it works best
- 10:25:44with AI and ML algorithms. And after you
- 10:25:47are done mastering these basics, start
- 10:25:49with machine learning and artificial
- 10:25:50intelligence algorithms. And before
- 10:25:52that, make sure you are sophisticated
- 10:25:53enough to work with data. I mean, you
- 10:25:55can analyze the data, clean it, prepare
- 10:25:57it for model building, and etc. And try
- 10:25:59to learn all of them with the
- 10:26:00implementations on unique data instead
- 10:26:03of the generic data that you find on the
- 10:26:04internet. The next step would be to
- 10:26:06learn the advanced concepts in AI and ML
- 10:26:08like TensorFlow for object detection,
- 10:26:10speech recognition, image processing,
- 10:26:11etc. And after you have mastered the
- 10:26:14versioning tools, you must make sure
- 10:26:16that you have a credibility in order to
- 10:26:17get a job. Because as an employer,
- 10:26:20anyone would look for credible person
- 10:26:22proficient enough to do the job. And
- 10:26:24where will you get that? If you have a
- 10:26:26relevant master's degree, it is well and
- 10:26:28good. But if you don't have a degree,
- 10:26:30you can always go for a certification,
- 10:26:32which will give you the credibility, and
- 10:26:34it is going to be the best option to
- 10:26:35prove your mettle. And then there's one
- 10:26:38way to look at it. I mean, you can take
- 10:26:39up the Edureka's post graduate program
- 10:26:41in artificial intelligence and machine
- 10:26:43learning, which is going to be a a good
- 10:26:45deal for you because it is an
- 10:26:46affiliation with a top college and then
- 10:26:49you are good to go. You can apply for
- 10:26:51jobs and you'll get the job easily if
- 10:26:52you have all the skill sets and you have
- 10:26:54the experience in making relevant, you
- 10:26:56know, you are working on real-time
- 10:26:58projects as well. So, these are going to
- 10:26:59be very useful for you.
- 10:27:02>> [music]
- 10:27:07>> Let's get started with our basic level
- 10:27:09questions.
- 10:27:10So, first we have what is the difference
- 10:27:12between AI, machine learning, and deep
- 10:27:14learning? I'm sure all of you have this
- 10:27:16question at the top of your mind because
- 10:27:18there's a huge confusion between AI,
- 10:27:19machine learning, and deep learning. So,
- 10:27:21let's try to understand how they are
- 10:27:23different. Now, first of all, AI came
- 10:27:25into existence at around 1950s. All
- 10:27:28right, this was followed by machine
- 10:27:29learning and then deep learning was
- 10:27:31introduced. Now, AI basically represents
- 10:27:34simulated intelligence in machines,
- 10:27:36which means that it represents any robot
- 10:27:39or any machine that can mimic the
- 10:27:41behavior of a human being. Machine
- 10:27:43learning on the other hand is a practice
- 10:27:46of getting machines to make decisions
- 10:27:48without being explicitly programmed to
- 10:27:50do so. Now, if you don't program a
- 10:27:52machine, how are you going to let it
- 10:27:54make decisions? Now, the way machines
- 10:27:57learn is through data. So, the most
- 10:27:59important thing in machine learning is
- 10:28:01the data. All right, you're going to
- 10:28:02train machines using data so that they
- 10:28:04can make their own decisions. Next, we
- 10:28:06have deep learning. Now, deep learning
- 10:28:08is basically the process of using
- 10:28:10artificial neural networks to solve
- 10:28:12complex problems. So, basically you can
- 10:28:15think of deep learning as a field that
- 10:28:17tries to mimic our brain. Okay, so how
- 10:28:20we have neural networks in our brain,
- 10:28:22that's exactly how deep learning uses
- 10:28:24the concepts of artificial neural
- 10:28:26networks in order to solve problems.
- 10:28:28Now, AI is a subset of data science. So
- 10:28:31guys, first of all, data science is the
- 10:28:32process of deriving useful insights from
- 10:28:35data. All right, it's a process of
- 10:28:37extracting information from data that
- 10:28:39will help you solve problems. So, AI is
- 10:28:42a subset of data science. Now, on the
- 10:28:44other hand, machine learning is a subset
- 10:28:46of AI and data science because machine
- 10:28:48learning comes after AI. So, basically
- 10:28:51in AI, you're going to make use of
- 10:28:53techniques and concepts of machine
- 10:28:55learning in order to solve problems.
- 10:28:57Then, we have deep learning. So, it's
- 10:28:59sort of a hierarchy. First, we have data
- 10:29:01science, then we have AI, then we have
- 10:29:03machine learning, and then we have deep
- 10:29:04learning. Deep learning is a subset of
- 10:29:06machine learning, AI, and data science.
- 10:29:09Okay, I hope this is clear. Now, the
- 10:29:11main aim of artificial intelligence is
- 10:29:13to build machines in such a way that
- 10:29:15they're capable of thinking like human
- 10:29:17beings. All right, so basically, they
- 10:29:19must be able to mimic the behavior of a
- 10:29:22human being. Now, the aim of machine
- 10:29:24learning on the other hand is to make
- 10:29:26machines learn by providing them a lot
- 10:29:28of data. Okay, once you make a machine
- 10:29:30learn through data, it's going to be
- 10:29:32able to solve complex problems and find
- 10:29:34solutions. Now, the aim of deep learning
- 10:29:37is to build neural networks that are
- 10:29:40able to solve more advanced and complex
- 10:29:42problems. Okay, now like I mentioned,
- 10:29:44deep learning is like an artificial
- 10:29:47brain. All right, you're basically
- 10:29:48building an artificial brain that is
- 10:29:50able to think exactly like how we do.
- 10:29:53Okay, that's what deep learning is. It's
- 10:29:55a little more advanced than machine
- 10:29:57learning. Now guys, in short, AI,
- 10:29:59machine learning, and deep learning are
- 10:30:01used to solve problems through data. So,
- 10:30:03basically, AI makes use of techniques
- 10:30:06and methods of machine learning and deep
- 10:30:08learning to solve problems or to draw
- 10:30:10useful insights from data. So, this is
- 10:30:13the difference between AI, machine
- 10:30:14learning, and deep learning. I hope all
- 10:30:16of you are clear with this. Now, let's
- 10:30:18look at our question number two. The
- 10:30:20question is, "What is artificial
- 10:30:22intelligence? Give an example of where
- 10:30:25AI is used on a daily basis." So, there
- 10:30:27are a lot of definitions of AI on the
- 10:30:29internet. A few of them are, "Artificial
- 10:30:32intelligence is an area of computer
- 10:30:34science that emphasizes on the creation
- 10:30:37of intelligent machines that work and
- 10:30:39react like humans. So, like I said,
- 10:30:41basically, a machine that is able to
- 10:30:43mimic the behavior of a human being is
- 10:30:46known as artificial intelligence.
- 10:30:48Another such definition is the
- 10:30:49capability of a machine to imitate the
- 10:30:51intelligent human behavior. All right?
- 10:30:54So, artificial intelligence, in short,
- 10:30:56is basically a machine that we created
- 10:30:58who can act and think like a human
- 10:31:00being. Now, where do you think AI is
- 10:31:02used on a daily basis? There are tons of
- 10:31:05applications that make use of AI, but
- 10:31:07one of the most popular applications of
- 10:31:09AI is a Google search engine. Now, if
- 10:31:12you just open up Google search and you
- 10:31:13start typing anything, immediately you
- 10:31:15get recommendations. These
- 10:31:17recommendations you derive by using
- 10:31:19machine learning algorithms, by using
- 10:31:21deep neural networks, and so on. So, on
- 10:31:24the top of my head, the most general
- 10:31:25example of AI is a Google search engine.
- 10:31:28All of us use Google search engine, and
- 10:31:30we know how quick it is with its results
- 10:31:32and how relevant searches it gives us.
- 10:31:35All this is because of AI. All right?
- 10:31:38Now, let's look at our next question,
- 10:31:39which states, "What are the different
- 10:31:41types of AI?" Now, a lot of people might
- 10:31:43not be aware of this because there are a
- 10:31:46couple of types of AI or a couple of
- 10:31:48types of machines which are
- 10:31:50hypothetical. Okay, we haven't actually
- 10:31:52implemented these machines in the real
- 10:31:54world. We just have a theoretical
- 10:31:56definition of these. Okay, let's look at
- 10:31:58what I'm talking about. So, first of
- 10:32:00all, we have reactive machines AI. Now,
- 10:32:03these machines are all based on the
- 10:32:05present actions. Okay, they have no
- 10:32:07memory or they have no concept of
- 10:32:09storing memory so that they can learn
- 10:32:11from that experience. They just react at
- 10:32:14the moment. Okay, so they're based on
- 10:32:15present actions, and they cannot use
- 10:32:18previous experiences to form current
- 10:32:20decisions and update their memory. Then,
- 10:32:22we have limited memory AI. Now, this
- 10:32:25type of AI has some temporary storage of
- 10:32:27memory in it. Now, if we have some
- 10:32:29memory stored in a machine, we know that
- 10:32:31it can look back into the memory and it
- 10:32:34can try to make decisions based on
- 10:32:36previous or past experiences.
- 10:32:38So, limited memory AI makes use of that
- 10:32:40concept. We have temporary memory here.
- 10:32:42We do not have permanent memory, but one
- 10:32:45of the top applications of limited
- 10:32:47memory AI is the self-driving cars. I'm
- 10:32:50sure all of you have heard of
- 10:32:51self-driving cars. They make use of
- 10:32:53limited memory AI in order to run. Then
- 10:32:56we have theory of mind AI. Now, like I
- 10:32:59mentioned earlier, there are a couple of
- 10:33:01types of artificial intelligent machines
- 10:33:03which are not actually implemented in
- 10:33:05the real world. An example of that is
- 10:33:07theory of mind AI. Okay, this is
- 10:33:09basically an advanced machine which will
- 10:33:12have the ability to understand emotions,
- 10:33:15people, and other things in the real
- 10:33:16world. We might have come close to this
- 10:33:19type of AI, but we haven't actually
- 10:33:20developed something that can understand
- 10:33:22emotions. Next, we have self-aware AI.
- 10:33:25Now, this is another such example of a
- 10:33:27machine that is not built in the real
- 10:33:29world. This basically includes any
- 10:33:31machine that has consciousness or that
- 10:33:34can react just like a human being. Okay,
- 10:33:36so basically a machine that can take own
- 10:33:38decisions, that can form own
- 10:33:40conclusions, and these are machines that
- 10:33:42have the capability of making their own
- 10:33:44decisions without any human
- 10:33:45intervention. Now, this kind of AI is
- 10:33:47not developed, like I mentioned, because
- 10:33:49it's going to take up a lot of resources
- 10:33:51and we still haven't reached that peak
- 10:33:54of evolution yet. Then we have
- 10:33:56artificial narrow intelligence. Now,
- 10:33:58these are the general purpose AI that we
- 10:34:00see on a daily basis. I'm sure all of
- 10:34:03you have used Google Assistant, you've
- 10:34:05used Siri. All of that comes under
- 10:34:07artificial narrow intelligence. After
- 10:34:09that, we have artificial general
- 10:34:11intelligence. Now, these are a little
- 10:34:13more advanced than the artificial narrow
- 10:34:15intelligence.
- 10:34:16Then we have artificial superhuman
- 10:34:18intelligence. Now, these are one of the
- 10:34:21most advanced type of AIs that are
- 10:34:23there. Now, like I mentioned earlier,
- 10:34:25there are a couple of types of
- 10:34:27artificial intelligent machines which
- 10:34:29are not actually implemented in the real
- 10:34:31world. An example of that is artificial
- 10:34:34super human intelligence. So guys, these
- 10:34:36were the different types of AI. Now,
- 10:34:39let's look at the next question which
- 10:34:41says explain the different domains of
- 10:34:43artificial intelligence. Now, AI covers
- 10:34:45a lot of different domains starting with
- 10:34:48machine learning. Okay, so machine
- 10:34:50learning like I mentioned earlier is the
- 10:34:52science of getting computers to act by
- 10:34:54feeding them data and by letting them
- 10:34:56learn a few tricks on their own without
- 10:34:59being programmed to do so. Okay, so
- 10:35:01you're not explicitly programming the
- 10:35:03machine, instead you're feeding it a lot
- 10:35:04of data so that it understands the data
- 10:35:07and it makes its own decisions. Then we
- 10:35:09have neural networks. Now, neural
- 10:35:11networks are basically a set of
- 10:35:13algorithms or you can say a set of
- 10:35:14techniques which are modeled in
- 10:35:17accordance with a human brain. Okay,
- 10:35:19like I mentioned earlier, deep learning
- 10:35:21or neural networks is almost the same
- 10:35:23thing. Deep learning makes use of neural
- 10:35:25networks in order to solve complex
- 10:35:27problems. Now, we have robotics. Now,
- 10:35:29robotics is a subset of AI which
- 10:35:32includes different branches and
- 10:35:33applications of robots. These robots are
- 10:35:36basically artificial agents which act in
- 10:35:39a real world environment. Okay, so an AI
- 10:35:41robot works by manipulating the objects
- 10:35:43in its surrounding by perceiving,
- 10:35:45moving, and taking relevant actions.
- 10:35:48Then we have expert systems. Now, an
- 10:35:50expert system is basically a computer
- 10:35:52system that mimics the decision-making
- 10:35:54ability of a human being. Now, I know
- 10:35:56all of these domains sound very similar,
- 10:35:59but they have a very different approach
- 10:36:01with which they solve the problem. All
- 10:36:03right, that's the main difference
- 10:36:04between these domains. Next, we have
- 10:36:06fuzzy logic systems. Now, traditional
- 10:36:09systems usually give out output in the
- 10:36:11form of binary. So, usually if you feed
- 10:36:14something to a machine, it's always in
- 10:36:15the binary form. The output is also
- 10:36:17usually in the form of yes, no, true,
- 10:36:19false, and so on. But when it comes to
- 10:36:21fuzzy logic, it tries to give an output
- 10:36:23in the form of degrees of truth. Okay,
- 10:36:26so it's very different when compared to
- 10:36:28the traditional computer systems or the
- 10:36:29traditional programs. Next, we have
- 10:36:32natural language processing. Now, this
- 10:36:34is a field of AI that analyzes natural
- 10:36:37human language to derive useful insights
- 10:36:40so that it can solve problems. Now, NLP
- 10:36:42is used majorly in social media
- 10:36:44platforms. So, Twitter sentimental
- 10:36:46analysis is done via NLP. Even Facebook
- 10:36:49uses NLP in a lot of things. All right,
- 10:36:52so NLP, fuzzy logic, expert systems,
- 10:36:54machine learning, neural networks, and
- 10:36:56robotics are the different domains of
- 10:36:58AI. I hope all of you are clear with the
- 10:37:00domains. Now, let's look at our next
- 10:37:03question. Okay, so how is machine
- 10:37:05learning related to artificial
- 10:37:07intelligence? There is a huge confusion
- 10:37:10between machine learning and AI. A lot
- 10:37:12of people tend to believe that AI and
- 10:37:13machine learning is one in the same
- 10:37:15thing. All right, I would say that you
- 10:37:17cannot compare AI and machine learning
- 10:37:19because machine learning is a subset of
- 10:37:21AI. So, basically, AI makes use of
- 10:37:24machine learning algorithms and machine
- 10:37:26learning concepts to solve problems.
- 10:37:28That's the basic difference or that is
- 10:37:30where the confusion ends. Machine
- 10:37:32learning is a technique which is
- 10:37:34implemented in artificial intelligence
- 10:37:36in order to solve problems. I hope this
- 10:37:39is clear. Now, let's look at what are
- 10:37:41the different types of machine learning.
- 10:37:43So, there are three types of machine
- 10:37:45learning. We have supervised,
- 10:37:46unsupervised, and reinforcement
- 10:37:48learning. Now, supervised learning is
- 10:37:51the type of learning in which the
- 10:37:52machine learns by using labeled data.
- 10:37:55Now, to make you understand, let's look
- 10:37:56at an example. Okay, let's say that
- 10:37:58you've input images of apples and
- 10:38:01oranges to your machine and you've
- 10:38:03labeled them. You've told the machine
- 10:38:05like, "Listen, this is the apple, this
- 10:38:07is an orange, and the output should also
- 10:38:09look like this." Okay, so you're
- 10:38:11labeling the input as apple and an
- 10:38:13orange, and then you're asking the
- 10:38:15machine to output an apple and an
- 10:38:17orange. But, when it comes to
- 10:38:18unsupervised learning, you're not going
- 10:38:20to label them. You're just going to give
- 10:38:22them images of apple and oranges, and it
- 10:38:24has to figure out on its own. It has to
- 10:38:26try and understand the difference
- 10:38:28between apple and oranges, try and
- 10:38:30understand how they look different, or
- 10:38:32how they have a different color. So,
- 10:38:34basically in unsupervised learning, you
- 10:38:35don't have a labeled data set. Okay,
- 10:38:37you're going to give it an unlabeled
- 10:38:38data set, and you're going to ask it to
- 10:38:40find out and classify which is an apple
- 10:38:43and which is an orange. Okay, that's the
- 10:38:44difference between supervised and
- 10:38:46unsupervised. Now, reinforcement
- 10:38:47learning is comparatively different.
- 10:38:50Let's imagine that you were put off in
- 10:38:52an island. Okay, let's say that you were
- 10:38:53left in an isolated island. What would
- 10:38:56you do? Now, initially, we'll all panic,
- 10:38:58and we won't know what to do. But, after
- 10:39:00a point, you'll start exploring the
- 10:39:02island. You'll start adapting to the
- 10:39:04change in the climate conditions, you'll
- 10:39:06start looking for food, and then you'll
- 10:39:08try and understand which food is right
- 10:39:09for you and which food is wrong for you.
- 10:39:11You know, you'll learn from your
- 10:39:12experience. So, in reinforcement
- 10:39:15learning, basically, an agent interacts
- 10:39:17with its environment by producing
- 10:39:19actions and discovers errors or rewards.
- 10:39:22Now, the type of problems that
- 10:39:23supervised learning is used to solve is
- 10:39:25regression and classification. When it
- 10:39:27comes to unsupervised, it is association
- 10:39:29and clustering. And in reinforcement
- 10:39:31learning, it's all the reward-based
- 10:39:33problems. The type of data for
- 10:39:35supervised learning is labeled data. For
- 10:39:37unsupervised, it is unlabeled. And for
- 10:39:39reinforcement, it is no predefined data.
- 10:39:42Now, when I say no predefined data, I
- 10:39:43mean that the reinforcement learning
- 10:39:45agent has to start collecting the data.
- 10:39:48So, basically, in reinforcement
- 10:39:50learning, from data collection to model
- 10:39:52evaluation, it does everything. In terms
- 10:39:54of training, supervised learning
- 10:39:56provides external supervision in the
- 10:39:58form of labeled data set. In
- 10:40:00unsupervised learning, there's no
- 10:40:02supervision. That's why it's called
- 10:40:03unsupervised learning. Again, in
- 10:40:05reinforcement learning, there's no
- 10:40:06supervision at all. The agent has to
- 10:40:08figure everything out. Now, how
- 10:40:10supervised learning works is uh you map
- 10:40:13the labeled input to the known output.
- 10:40:15So, basically you teach the machine like
- 10:40:17you tell it that this is the input and
- 10:40:19this has to be the output. When it comes
- 10:40:21to unsupervised learning, you just
- 10:40:22provide data to the machine and it has
- 10:40:24to understand patterns and it has to
- 10:40:26discover the output. Now, in
- 10:40:28reinforcement learning, it has to follow
- 10:40:30the trial and error method. Okay,
- 10:40:32there's no particular way in which the
- 10:40:35agent learns. It just has to explore the
- 10:40:37environment, try out a few things, and
- 10:40:39learn from that experience. Popular
- 10:40:41supervised learning algorithms include
- 10:40:43linear regression, logistic regression.
- 10:40:46For unsupervised, we have K-means. And
- 10:40:47for reinforcement learning, we have
- 10:40:49Q-learning. So, guys, these were the
- 10:40:51different types of machine learning and
- 10:40:53I also discussed the difference between
- 10:40:55the three. Now, let's move on and look
- 10:40:57at our next question, which is what is
- 10:40:59Q-learning?
- 10:41:00In the previous slide itself, I told you
- 10:41:02that a type of reinforcement learning
- 10:41:03algorithm is Q-learning. So, basically,
- 10:41:06here what happens is an agent tries to
- 10:41:08learn the optimal policy from its past
- 10:41:11experience with the environment. The
- 10:41:13past experience of an agent are a
- 10:41:15sequence of action, state, and rewards.
- 10:41:18So, what happens is, first of all, you
- 10:41:20take an agent and you put it in state
- 10:41:22zero. Okay, let's say there's some state
- 10:41:24known as state zero. Now, this agent is
- 10:41:27going to perform some action A0. On
- 10:41:30performing this action, it is going to
- 10:41:32get a reward R1. And if it gets a reward
- 10:41:35R1, then it's going to move to state S1.
- 10:41:38But in case the action is wrong, then
- 10:41:40it's going to get a negative reward, as
- 10:41:43in some points are going to be reduced.
- 10:41:45So, guys, think of Q-learning as a game.
- 10:41:47You're in state zero, and then you do
- 10:41:49some action, and either you get a reward
- 10:41:51and go to the next state, or else you
- 10:41:53lose and you go back to the same state.
- 10:41:56So, until you learn, you're going to be
- 10:41:57in the same state. But if you keep
- 10:41:59learning and if you keep receiving
- 10:42:01positive rewards, then you're going to
- 10:42:02move on to state one, and similarly, you
- 10:42:04move on to state two, three, and so on.
- 10:42:06This is what Q learning is about.
- 10:42:08Now, the next question is what is deep
- 10:42:10learning? Now, deep learning, like I
- 10:42:12mentioned earlier, basically mimics the
- 10:42:15way our brain works. Okay, it learns
- 10:42:17from experience. Now, the main concept
- 10:42:19behind deep learning is neural networks.
- 10:42:22In our brain also, we have neural
- 10:42:23networks. [clears throat] So, what deep
- 10:42:24learning tries to do is it tries to use
- 10:42:26the concept of neural networks in order
- 10:42:29to solve complex problems. So,
- 10:42:30basically, we're trying to mimic our
- 10:42:32brain. Any deep neural network will have
- 10:42:35three types of layers. The first is the
- 10:42:37input layer. Now, this layer will
- 10:42:39basically receive all the input, and it
- 10:42:41will forward them to the hidden layer.
- 10:42:43Now, in the hidden layer, all the
- 10:42:45analysis and the computation takes
- 10:42:47place. All right, once the computation
- 10:42:49is done, the result is transferred to
- 10:42:51the output layer. Now, there can be n
- 10:42:53number of hidden layers depending on the
- 10:42:55type of problem you're trying to solve.
- 10:42:57Then, we have the output layer. So,
- 10:42:59basically, this layer is responsible for
- 10:43:01transferring the information from the
- 10:43:03neural network to the outside world. So,
- 10:43:05it's as simple as that. It's pretty
- 10:43:07obvious. Input layer will take in the
- 10:43:09input, hidden layer will perform the
- 10:43:11computations, and the output layer will
- 10:43:13give out the output. This is a small
- 10:43:15explanation of what deep learning is.
- 10:43:17Now, of course, this is much more
- 10:43:18complex than this, but in short, this is
- 10:43:21exactly what deep learning is. Now,
- 10:43:23let's look at our next question, which
- 10:43:25is explain how deep learning works. So,
- 10:43:28basically, deep learning is a concept
- 10:43:30based on something known as neuron.
- 10:43:33Okay, neuron is a basic unit of the
- 10:43:35brain. Inspired from this neuron, they
- 10:43:37came up with something known as
- 10:43:39perceptrons or artificial neurons. Now,
- 10:43:42in this image on the left-hand side, you
- 10:43:44can see that there is something known as
- 10:43:46dendrite. These are modules which
- 10:43:48receive the input. It basically receives
- 10:43:50all the signals that we send to our
- 10:43:52brain. Okay, similar to the dendrites
- 10:43:55are the input layer in our artificial
- 10:43:57neural networks. Now, in the previous
- 10:43:59slide, we discussed that the input layer
- 10:44:00takes in all the input from the outside.
- 10:44:03That's exactly what a dendrite does. So,
- 10:44:05basically a perceptron receives multiple
- 10:44:08inputs. It applies various
- 10:44:10transformations and functions, and then
- 10:44:12it provides an output.
- 10:44:14So, basically guys, just like how our
- 10:44:16brain contains multiple connected
- 10:44:18neurons called neural networks, we also
- 10:44:21have a network of artificial neurons
- 10:44:23called perceptrons to form a deep neural
- 10:44:25network. So, basically an artificial
- 10:44:27neuron or a perceptron, it models a
- 10:44:30neuron which has a set of inputs, each
- 10:44:33of which is assigned some specific
- 10:44:35weight. Okay, all of these inputs will
- 10:44:37have a specific weight, and the neuron
- 10:44:39will compute some function on these
- 10:44:41weighted inputs and give you the output.
- 10:44:44So, the neuron will basically perform
- 10:44:45analysis and all of that on these
- 10:44:47weighted inputs to give you some output.
- 10:44:50This is a basic concept of deep
- 10:44:52learning. So, there are inputs which
- 10:44:54have some weight on it, and these inputs
- 10:44:56are then formulated and analyzed in
- 10:44:59order to give you an output. Now, let's
- 10:45:01look at our next question, which is
- 10:45:03explain the commonly used artificial
- 10:45:05neural networks. Now, this is a very
- 10:45:07theoretical question because in order to
- 10:45:09make you understand how each of them
- 10:45:11work will take a lot of time. Okay, so
- 10:45:13I'm just going to briefly tell you what
- 10:45:15each of these networks are and what they
- 10:45:17do. Now, feedforward neural network is
- 10:45:19the most basic kind of artificial neural
- 10:45:21network. So, basically the feedforward
- 10:45:23neural network is unidirectional. The
- 10:45:26data passes through the input nodes and
- 10:45:28leaves through the output nodes. In
- 10:45:30feedforward neural network, usually the
- 10:45:32number of hidden layers depends on the
- 10:45:34complexity of the problem. Coming to
- 10:45:36convolutional neural networks, here
- 10:45:39basically the input features are taken
- 10:45:42in small sets. Okay, or they're taken in
- 10:45:44batches. This will help the network
- 10:45:47remember better because you're feeding
- 10:45:49batches of images or you're feeding
- 10:45:51batches of input to the neural network.
- 10:45:54Now, this type of neural network is
- 10:45:56mainly used for signal and image
- 10:45:57processing. Next, we have recurrent
- 10:46:00neural networks. These are also known as
- 10:46:03long short-term memory networks. So,
- 10:46:05this basically works on the principle of
- 10:46:07feeding the output of a layer back into
- 10:46:10the input layer in order to predict the
- 10:46:12outcomes. Okay, this way it's more
- 10:46:14precise and it is a little more complex
- 10:46:16when compared to convolutional networks.
- 10:46:18Now, one main important point of
- 10:46:20recurrent neural networks is that they
- 10:46:22have something known as memory. So,
- 10:46:24basically each neuron will have some
- 10:46:26information or some memory stored in
- 10:46:28them so that if they can use this memory
- 10:46:31in order to take actions in the future.
- 10:46:33So, they have some experiences stored in
- 10:46:36the form of memory so that they can make
- 10:46:37their decisions based on previous
- 10:46:39actions. Now, finally, we have
- 10:46:41autoencoders. Now, autoencoders are
- 10:46:44mainly used in dimensionality reduction
- 10:46:46for learning generative models. Okay,
- 10:46:49and one more important thing about
- 10:46:50autoencoders is that the number of units
- 10:46:53in the output layer and the input layer
- 10:46:55is the same. This is because the output
- 10:46:57layer has to reconstruct its own inputs.
- 10:46:59So, these were the different types of
- 10:47:02artificial neural networks. Now, let's
- 10:47:04look at our next question, which is what
- 10:47:06are Bayesian networks? Okay, so a
- 10:47:08Bayesian network is a statistical model
- 10:47:11that represents a set of variables and
- 10:47:13the conditional dependencies in the form
- 10:47:16of a directed acyclic graph. Now,
- 10:47:18basically, on the occurrence of any
- 10:47:20event, a Bayesian network can be used to
- 10:47:22predict the likelihood that any one of
- 10:47:25several possible known causes was a
- 10:47:27contributing factor. An example of this
- 10:47:30is a Bayesian network could be used to
- 10:47:32study the relationship between diseases
- 10:47:34and symptoms. So, given a set of
- 10:47:36symptoms, the Bayesian network can be
- 10:47:38used to find out the probability of the
- 10:47:41presence of any diseases. All right, so
- 10:47:43the next question is explain the
- 10:47:45assessment that is used to test the
- 10:47:47intelligence of a machine. Now, guys,
- 10:47:49this is a a common question and it is
- 10:47:52sort of a general knowledge-based
- 10:47:53question. All right, I'm hoping that
- 10:47:55most of you know the answer to this. So,
- 10:47:57let's look at what the answer is. I'm
- 10:48:00not sure how many of you have heard of
- 10:48:01Alan Turing. So, Alan Turing was the one
- 10:48:04who came up with the Turing test. Now,
- 10:48:06this test is basically to determine
- 10:48:08whether or not a computer is capable of
- 10:48:11thinking like a human being.
- 10:48:13So, if a machine or if a computer passes
- 10:48:15this exam, it means that that machine is
- 10:48:18capable of thinking like a human being.
- 10:48:20It means that it is successfully an
- 10:48:22artificial intelligent machine, meaning
- 10:48:24that it can make its own decisions and
- 10:48:26interpret data and form their own
- 10:48:28formulations or form their own
- 10:48:30conclusions about the data. Now, sadly,
- 10:48:32I don't think there are a lot of
- 10:48:33machines that have passed the Turing
- 10:48:35test. In fact, I'm not sure if there is
- 10:48:37any machine that's passed the Turing
- 10:48:39test as of now, but in the near future,
- 10:48:41I'm sure that we'll see machines who are
- 10:48:44more smarter than human beings and who
- 10:48:46have passed this test. Now, for a
- 10:48:48machine, it might be very easy to do
- 10:48:50computations, but it might be very hard
- 10:48:52for a machine to just get up and walk
- 10:48:54around. All right, the simple things
- 10:48:56that us humans can do is very
- 10:48:58complicated for a machine. They can do
- 10:49:00computations which we can do in probably
- 10:49:02a year, they can do those computations
- 10:49:04in maybe a week or less than a week.
- 10:49:07But, doing simple things such as walking
- 10:49:09up to the fridge or walking up to the
- 10:49:11kitchen is very hard for the machines.
- 10:49:14So, to achieve that level of
- 10:49:15intelligence, we're going to take a
- 10:49:16while, but in the near future, I'm sure
- 10:49:18we'll see machines which are way more
- 10:49:20capable than human beings. Now, let's
- 10:49:22move on to our next level. Now, here
- 10:49:25I'll basically be discussing
- 10:49:27intermediate level artificial
- 10:49:28intelligence questions. So, let's look
- 10:49:30at the first question. All right, the
- 10:49:32first question is how does reinforcement
- 10:49:35learning work? Explain with an example.
- 10:49:38Okay, so first of all, reinforcement
- 10:49:40learning is a type of machine learning.
- 10:49:42We discussed about reinforcement
- 10:49:44learning earlier. Reinforcement learning
- 10:49:46is a type of machine learning wherein
- 10:49:48there's an agent and you put this agent
- 10:49:50in an unknown environment. All right,
- 10:49:52now the agent has to figure out actions,
- 10:49:54what sort of actions it must take, and
- 10:49:56how it's going to get rewards so that it
- 10:49:58can move from state zero to state one.
- 10:50:01It's sort of like a video game. If
- 10:50:02you're in a video game, let's say if
- 10:50:04you're playing Counter-Strike, you're in
- 10:50:06level zero or state zero. Now, if you
- 10:50:09perform some action and if you get some
- 10:50:11rewards, you're going to move to state
- 10:50:12one. That's exactly how reinforcement
- 10:50:14learning works. If you perform the
- 10:50:16relevant actions and the correct
- 10:50:17actions, you're going to get a reward
- 10:50:19and you'll move on to the next state.
- 10:50:21But in case you perform a wrong action,
- 10:50:23you'll get negative rewards and you'll
- 10:50:25stay in the same state unless and until
- 10:50:27you don't learn. All right, so if you
- 10:50:29learn and achieve, then you'll move to
- 10:50:31the next state. So, basically a
- 10:50:33reinforcement learning system will have
- 10:50:35two main components. It'll have an agent
- 10:50:38and an environment. Now, the agent I've
- 10:50:40been repetitively saying an agent An
- 10:50:42agent is basically the reinforcement
- 10:50:44learning algorithm. It is the model. The
- 10:50:47model has to learn everything on its
- 10:50:48own. It has to collect data on its own.
- 10:50:51It has to draw useful insights on its
- 10:50:53own. Okay, you're not going to feed any
- 10:50:55predefined data to this reinforcement
- 10:50:57learning agent. All right, he has to
- 10:50:59figure out everything on its own. So,
- 10:51:01let's look at an example of
- 10:51:02Counter-Strike. Okay, I'm not sure how
- 10:51:04many of you play the game, but yeah.
- 10:51:06What happens here is the reinforcement
- 10:51:08learning agent or the player one
- 10:51:10collects a state S0 from the
- 10:51:12environment. Okay, so let's suppose that
- 10:51:14you're playing Counter-Strike and you're
- 10:51:16in state zero. Now, you'll perform some
- 10:51:19action A0. All right, initially it's
- 10:51:21going to be a random action. So,
- 10:51:23obviously if you're put in an unknown
- 10:51:24environment, your first action is going
- 10:51:26to be random, correct? Because you don't
- 10:51:28know what's right, you don't know what's
- 10:51:29wrong. So, in your state zero, you'll
- 10:51:31take an action A0. This will result in a
- 10:51:34new state S1. And on achieving state S1,
- 10:51:38the agent will get a reward R1, okay,
- 10:51:40from the environment. Now, in the case
- 10:51:43of Counter-Strike games, if you've
- 10:51:45observed, whenever you win a state or
- 10:51:47you pass a level, you're going to get
- 10:51:49some rewards. Maybe you'll get more
- 10:51:51weapons or you'll get more points. Okay,
- 10:51:53just like that, in reinforcement
- 10:51:55learning problem, you'll get some reward
- 10:51:57R1. Okay, it's basically a plus point.
- 10:52:00You might get a negative reward or a
- 10:52:02positive reward based on the action that
- 10:52:04you take. Now, this loop will go on
- 10:52:07until the agent is dead or it reaches
- 10:52:09the destination. So, in Counter-Strike,
- 10:52:11until you have failed the level, you
- 10:52:14will keep playing the game, right?
- 10:52:15You'll keep moving from state one, state
- 10:52:17two, state three, and so on. Or, if
- 10:52:19you've reached the destination, then
- 10:52:20it's the end game. That's exactly how it
- 10:52:23works in reinforcement learning. If the
- 10:52:25agent has explored the entire
- 10:52:26environment and reached the end state,
- 10:52:29that's when the loop will end. All
- 10:52:31right, that's exactly how reinforcement
- 10:52:33learning works. It is very similar to
- 10:52:35the games that we play. All right, it's
- 10:52:37very understandable. Now, let's move on
- 10:52:39and discuss the next question. So, the
- 10:52:41next question is explain Markov decision
- 10:52:44process with an example. Now, the
- 10:52:46solution for a reinforcement learning
- 10:52:48problem is achieved through the Markov
- 10:52:50decision process. It's basically a
- 10:52:53mathematical approach that maps the
- 10:52:55solution in reinforcement learning.
- 10:52:57Okay, so now to understand this, there
- 10:52:59are a couple of parameters in a Markov
- 10:53:02decision process. They're going to be a
- 10:53:04set of actions called A. Okay, you can
- 10:53:06name them A. A set of states, there's
- 10:53:09going to be reward, there's going to be
- 10:53:11policy, and there's going to be value.
- 10:53:13To sum it up, what exactly happens in a
- 10:53:15Markov decision process is that the
- 10:53:18agent takes an action A to transition
- 10:53:21from the start state to the end state.
- 10:53:23Now, while doing so, the agent receives
- 10:53:26some reward R for each action that he
- 10:53:28takes. The series of actions taken by
- 10:53:30the agent will define a policy or an
- 10:53:33approach. And the rewards collected will
- 10:53:36define the value. So, the main goal in a
- 10:53:39Markov decision process is to maximize
- 10:53:41the rewards by choosing the most optimum
- 10:53:44policy. Meaning that you're going to
- 10:53:46choose the best path or the best
- 10:53:47solution in order to get the most number
- 10:53:50of rewards. Now, in order to make you
- 10:53:52all understand this better, let's solve
- 10:53:54the shortest path problem by using
- 10:53:56Markov decision process. I'm sure all of
- 10:53:59you have heard of shortest path problem.
- 10:54:01This was I think taught to us when we
- 10:54:03were in 11th or 12th, I'm not sure. Look
- 10:54:06at the diagram that is over here. This
- 10:54:08is basically a representation of our
- 10:54:11problem. Given this representation, our
- 10:54:13goal here is to find the shortest path
- 10:54:16between the node A and node D. All
- 10:54:18right, you can see nodes A, B, C, and D.
- 10:54:21We have to find the shortest path
- 10:54:23between node A and node D. Now, the link
- 10:54:25between these two nodes has a number on
- 10:54:28it. Okay, for example, between A and C
- 10:54:30you can see there's a number 15. Okay,
- 10:54:32this basically denotes the cost to
- 10:54:34traverse that edge. So, if you want to
- 10:54:36go from A to C, you'll spend around 15
- 10:54:39points. So, our end goal here is to
- 10:54:41travel between node A and node D with
- 10:54:44minimal possible cost. We should travel
- 10:54:47between A to D in such a way that our
- 10:54:49cost is minimal. Now, in this problem if
- 10:54:51you notice that we have a set of states.
- 10:54:54Okay, these are denoted by the nodes A,
- 10:54:56B, C, D. Now, like I mentioned earlier,
- 10:54:59a Markov decision process has a set of
- 10:55:01states. Similarly, in this problem the
- 10:55:03set of states are A, B, C, D. The action
- 10:55:06is to traverse from one node to the
- 10:55:08other. So, going from A to B is
- 10:55:10basically an action. Going from A to C
- 10:55:12is another action. Going from A to D is
- 10:55:15another action, and so on. Now, reward
- 10:55:17is represented by the cost on each of
- 10:55:20these links. And the policy is the path
- 10:55:22which is taken to reach the destination.
- 10:55:25So, our aim here is to choose a policy
- 10:55:28that gets us to node D in the minimum
- 10:55:30cost possible. So, how do you think you
- 10:55:33can solve this problem? All right, you
- 10:55:34can start off at node A and you can take
- 10:55:37baby steps to your destination. Now,
- 10:55:39initially only the next possible node is
- 10:55:41visible to you. Like I mentioned
- 10:55:43earlier, the initial action taken in a
- 10:55:46reinforcement learning problem is always
- 10:55:48random. So, at random you'll choose any
- 10:55:50node. Let's say you take A to B. Now, if
- 10:55:53you go from A to B, you can go B to D
- 10:55:55and you'll reach the destination. So,
- 10:55:57policy is the path which is taken to
- 10:56:00reach the destination. All right, so it
- 10:56:02can go from A to B to D or you can go
- 10:56:04from A to C to D or you can go A C B D.
- 10:56:08All right, now it's up to you to figure
- 10:56:09out which is the shortest path. All
- 10:56:11right, you have to choose a path in such
- 10:56:13a way that the cost between A to D is
- 10:56:16minimized. So guys, this was a simple
- 10:56:18problem of how Markov decision process
- 10:56:21is used to solve the shortest path
- 10:56:22problem. Now, let's move on and look at
- 10:56:25our next question. All right, now the
- 10:56:27next question is explain reward maximiza
- 10:56:30tion in reinforcement learning. So,
- 10:56:32basically a reinforcement learning agent
- 10:56:35works based on the theory of reward
- 10:56:37maximization. Okay, in the previous
- 10:56:39question itself I told you that the main
- 10:56:41aim of reinforcement learning is to
- 10:56:43maximize the reward. So, that's why a
- 10:56:46reinforcement learning agent must be
- 10:56:48trained in such a way that he takes the
- 10:56:50best action so that the reward is
- 10:56:52maximum. Okay, this is exactly what
- 10:56:54reward maximization means. He has to
- 10:56:57choose the best policy in such a way
- 10:56:59that the reward is maximum. Now, let me
- 10:57:02explain this with a small game. So, in
- 10:57:04the figure you can see a fox, you can
- 10:57:06see some meat and you can see a tiger.
- 10:57:08Now, our reinforcement learning agent is
- 10:57:10the fox. His end goal is to eat the
- 10:57:13maximum amount of meat before being
- 10:57:15eaten by the tiger. Okay, so he has to
- 10:57:18explore around, eat the maximum number
- 10:57:20of meat that he can eat before the tiger
- 10:57:22kills him. Since the fox is a clever
- 10:57:25fellow, he eats the meat that is closer
- 10:57:27to him. Okay, so rather than eating the
- 10:57:30meat which is close to the tiger, he
- 10:57:32eats the meat which is only close to
- 10:57:33him. This is because the closer he gets
- 10:57:35to the tiger, the higher are his chances
- 10:57:38of getting killed. So, as a result of
- 10:57:40this, the rewards near the tiger, even
- 10:57:43if they are bigger meat chunks, will be
- 10:57:45discounted. So, because the fox is not
- 10:57:48going closer to the tiger and eating the
- 10:57:50meat chunks closer to the tiger, this
- 10:57:52reward will get discounted. Now, I know
- 10:57:55you're all wondering what discounted is.
- 10:57:57Now, this is done because of the
- 10:57:59uncertainty factor that the tiger might
- 10:58:01kill the fox. So, what is discounting of
- 10:58:04reward? Okay, how does it work? To
- 10:58:06understand this, we define a discount
- 10:58:09rate called gamma. Okay, this is a
- 10:58:10parameter and the value of gamma always
- 10:58:14ranges between zero and one. So, the
- 10:58:16smaller the gamma, the larger the
- 10:58:18discount and so on. So, guys, this was
- 10:58:20reward maximization. So, here basically
- 10:58:23the fox will try to get as much as meat
- 10:58:25chunks as he can and he'll also try to
- 10:58:28avoid getting killed because that will
- 10:58:30end the reinforcement learning loop. We
- 10:58:32also discussed the discounted factor.
- 10:58:34All right, now that is not needed to
- 10:58:36understand reward maximization, but I
- 10:58:38just thought I'll add on some extra
- 10:58:39info.
- 10:58:40Now, let's look at the next question
- 10:58:42which is what is exploitation and
- 10:58:44exploration trade-off?
- 10:58:46So, basically exploration is, like the
- 10:58:49name suggests, it is about exploring and
- 10:58:52capturing more information about an
- 10:58:53environment. Now, on the other hand,
- 10:58:56exploitation is about using the already
- 10:58:58known exploited information to heighten
- 10:59:01the rewards. So, consider the same
- 10:59:03example that we discussed in the
- 10:59:04previous question. Here, the fox only
- 10:59:07eats the meat chunks which are close to
- 10:59:09him. Okay, he does not eat the bigger
- 10:59:11chunks because even though the bigger
- 10:59:13chunks would give him more rewards, it
- 10:59:15would get him killed. Okay, he does not
- 10:59:17go towards the tiger itself. Now, if the
- 10:59:19fox only focuses on the closest reward,
- 10:59:23he will never reach the big chunks of
- 10:59:24meat. Okay, this is what exploitation
- 10:59:27is. He's sticking only to the
- 10:59:29information that he knows and he's
- 10:59:31trying to get the most number of rewards
- 10:59:32from it. But if the fox decides to
- 10:59:35explore a bit, it can find the bigger
- 10:59:37rewards. Okay, the bigger rewards are
- 10:59:39basically the big chunks of meat which
- 10:59:41are near the tiger. And this is exactly
- 10:59:43what exploration is. Okay, exploitation
- 10:59:46is about using the already known
- 10:59:48information to heighten your rewards.
- 10:59:51Exploration on the other hand is about
- 10:59:53exploring and capturing more information
- 10:59:55about an environment.
- 10:59:57All right, so that was about
- 10:59:58exploitation and exploration. Now, let's
- 11:00:01move on to our next question. So, this
- 11:00:04is a difference question which ask the
- 11:00:07difference between parametric and
- 11:00:09non-parametric models. So, a parametric
- 11:00:12model basically uses a fixed number of
- 11:00:14parameters to build the model. Now,
- 11:00:16first of all guys, what are parameters?
- 11:00:18Now, parameters are basically predictor
- 11:00:21variables that are used to build a
- 11:00:23machine learning model or build any
- 11:00:25predictive analytics model. Now that we
- 11:00:27know what parameters are, let's try to
- 11:00:29understand the difference between a
- 11:00:31parametric and a non-parametric model.
- 11:00:34Now, parametric model basically uses a
- 11:00:35fixed number of parameters to build the
- 11:00:37model. A non-parametric model uses
- 11:00:40flexible number of parameters to build
- 11:00:42the model. When it comes to a parametric
- 11:00:44model, the assumptions about the data
- 11:00:46are very strong. In a non-parametric
- 11:00:49model, there are fewer assumptions about
- 11:00:51the data. A parametric model has the
- 11:00:53fixed number of parameters. Everything
- 11:00:55is defined over here. So, the
- 11:00:57computation is very fast. Okay, you know
- 11:00:59what sort of variables you'll need to
- 11:01:01predict the outcome. Okay, you have a
- 11:01:03defined set of variables or a defined
- 11:01:05set of predictor variables that will
- 11:01:07compute your outcome. So, that's why the
- 11:01:09computation is a bit faster. When you
- 11:01:12compare to non-parametric models, there
- 11:01:14are a lot of parameters taken into
- 11:01:16account. Now, when it comes to a
- 11:01:18non-parametric model, you do not have a
- 11:01:20fixed number of parameters. All right,
- 11:01:22you do not have a fixed number of
- 11:01:24predictor variables that will help you
- 11:01:26get to the outcome. So, the computation
- 11:01:28is a bit slower. Now, parametric models
- 11:01:31require lesser data and non-parametric
- 11:01:34require more data. Example of parametric
- 11:01:37models include logistic regression and
- 11:01:39naive bias. And for non-parametric
- 11:01:41models, we have KNN and decision tree
- 11:01:43models. Now, logistic and naive bias
- 11:01:46models are very strong models because
- 11:01:49they have a fixed number of parameters
- 11:01:51or a fixed number of predictor variables
- 11:01:53and they will give you an immediate
- 11:01:55output. Okay, when it comes to
- 11:01:57non-parametric models like decision tree
- 11:01:58models and KNN, you might even observe a
- 11:02:01little bit of overfitting. Okay, this
- 11:02:03happens because you have a fewer number
- 11:02:06of assumptions about the data and also
- 11:02:08because your parameters are not fixed.
- 11:02:10Now, that's not the reason for
- 11:02:12overfitting, but it's seen that in some
- 11:02:14of the non-parametric models,
- 11:02:16overfitting occurs more often. Now,
- 11:02:18let's discuss the next question, which
- 11:02:21is what is the difference between
- 11:02:22hyperparameters and model parameters?
- 11:02:25Now, model parameters are the predictor
- 11:02:27variables that I was speaking about
- 11:02:29earlier. Hyperparameters, let's discuss
- 11:02:32what they are. Okay, model parameters
- 11:02:34are the features of training data that
- 11:02:36will learn on its own during training.
- 11:02:38Whereas model hyperparameters are the
- 11:02:40parameters that determine the training
- 11:02:42process.
- 11:02:43Now, let's see that you want to
- 11:02:45determine the height of an individual
- 11:02:47depending on his weight. The height and
- 11:02:50weight will become your model
- 11:02:52parameters. But, your hyperparameter is
- 11:02:55basically the learning rate. It's the
- 11:02:57rate at which your model is going to
- 11:02:59learn this correlation between the
- 11:03:01height and the weight. So, this is the
- 11:03:03difference between model parameters and
- 11:03:05hyperparameters. Model parameters are
- 11:03:07the ones that you find in your data.
- 11:03:10These are all the variables that you use
- 11:03:12to predict your outcomes.
- 11:03:13Hyperparameters will define your
- 11:03:15training process. There is a huge
- 11:03:17difference between model and
- 11:03:18hyperparameters. Another difference is
- 11:03:21that they are internal to the model and
- 11:03:23their value can be estimated from the
- 11:03:24data. Hyperparameters are external to
- 11:03:27the model and their value cannot be
- 11:03:29estimated from data. Now, like I said,
- 11:03:31model parameters are derived from your
- 11:03:33data itself. Okay, these are the
- 11:03:35parameters that are there in your data.
- 11:03:37Hyperparameters are the ones that you
- 11:03:39define in order to train your entire
- 11:03:42data. So, that is the difference between
- 11:03:44hyperparameters and model parameters.
- 11:03:47So, next question is what are
- 11:03:48hyperparameters in deep neural networks?
- 11:03:51So, guys, like I mentioned in the
- 11:03:53previous example, hyperparameters are
- 11:03:56variables such as the learning rate.
- 11:03:58This will define how your entire data
- 11:04:00training process goes. For those of you
- 11:04:02don't know, in order to build a model,
- 11:04:04you first need to train the model and
- 11:04:06then you need to test it. Okay, now
- 11:04:08while training the model, you're going
- 11:04:09to make the model learn a lot of things.
- 11:04:12You're going to give it a lot of data.
- 11:04:14It has to figure out relations between
- 11:04:16various variables and how these
- 11:04:18variables are affecting the output. All
- 11:04:21of this training will depend on a few
- 11:04:23variables such as the learning rate.
- 11:04:25Okay, these are basically called
- 11:04:26hyperparameters in deep neural networks.
- 11:04:29So, these parameters will define the
- 11:04:31number of hidden layers that are present
- 11:04:33in the network. Okay, and more the
- 11:04:35number of hidden layers, the more
- 11:04:36accurate your network is going to be.
- 11:04:39Whereas, if you have less a number of
- 11:04:40units, you may cause underfitting in
- 11:04:42your data. Underfitting will also result
- 11:04:45in inaccurate predictions. So, that's
- 11:04:47why you need to make sure that the
- 11:04:49number of hidden units in your hidden
- 11:04:50layers are perfect, are ideal. Okay, and
- 11:04:53this is determined by your learning rate
- 11:04:56or by your hyperparameters. Not by your
- 11:04:58learning rate specifically, but by the
- 11:05:00number of hyperparameters you have.
- 11:05:02These number of hidden layers are
- 11:05:04determined by the hyperparameters. Okay,
- 11:05:07that's why hyper parameters are very
- 11:05:08important in deep neural networks. I
- 11:05:11hope you all are clear with this. Now,
- 11:05:13let's look at our next question, which
- 11:05:14is explain the different algorithms used
- 11:05:17for hyperparameter optimization. We'll
- 11:05:20discuss the three methods, which are
- 11:05:21grid search, random search, and Bayesian
- 11:05:24optimization. Okay, now grid search
- 11:05:26basically will train the network only on
- 11:05:29the two sets of hyper parameters, which
- 11:05:31are learning rate and the number of
- 11:05:32layers. Okay, so it's going to use every
- 11:05:34combination of these two sets in order
- 11:05:37to train the network. Okay, after that
- 11:05:39it'll evaluate the efficiency of the
- 11:05:41model by using the cross-validation
- 11:05:43techniques. Cross-validation is the best
- 11:05:46improvement method. Okay, it's the best
- 11:05:48way to check if your model is optimal or
- 11:05:51not. Then we have random search. Now,
- 11:05:54this will randomly select samples, and
- 11:05:56it will evaluate sets for a particular
- 11:05:58probability distribution. Now, in random
- 11:06:01search there is no fixed number of hyper
- 11:06:03parameters that it's going to evaluate.
- 11:06:05So, it'll randomly select a set of hyper
- 11:06:08parameters. Okay, for example, now
- 11:06:10instead of checking your entire sample
- 11:06:13or your entire Let's say that you have
- 11:06:1510,000 samples. Instead of checking all
- 11:06:17of these samples, it'll randomly select
- 11:06:19100 parameters that can be checked.
- 11:06:21Okay, and then it'll use this to build
- 11:06:23the model.
- 11:06:24After that, we have Bayesian
- 11:06:25optimization.
- 11:06:27Okay, now Bayesian optimization
- 11:06:29basically uses something known as a
- 11:06:31Gaussian process. Basically, the
- 11:06:33Gaussian process will help in model
- 11:06:35tuning. Okay, model tuning or you can
- 11:06:37also say parameter tuning. So, parameter
- 11:06:39tuning will help you tweak the
- 11:06:41parameters a little bit in order to
- 11:06:43improve the efficiency of the model.
- 11:06:45Okay, so Bayesian optimization basically
- 11:06:47makes use of the Gaussian process, which
- 11:06:50will provide model tuning to your
- 11:06:51algorithm and thus improve the
- 11:06:53efficiency. Now guys, the one of the
- 11:06:55most important ways to improve the
- 11:06:57efficiency of a model is by
- 11:06:59hyperparameter optimization. Okay, if
- 11:07:01you're tuning your hyper parameters and
- 11:07:03if you're trying to check in which way
- 11:07:05these hyper parameters will give you the
- 11:07:07most accurate outcome, that's when your
- 11:07:10result will be very good. Okay, so
- 11:07:12that's the best way to improve the
- 11:07:13efficiency of the model. All right, now
- 11:07:15let's look at our next question. The
- 11:07:18next question is how does data
- 11:07:20overfitting occur and how can it be
- 11:07:22fixed? Now guys, this is a very common
- 11:07:24question in a machine learning or in an
- 11:07:27artificial intelligence interview. Okay,
- 11:07:29people expect you to understand what
- 11:07:31data overfitting is and how you can fix
- 11:07:34these problems. Okay, because data
- 11:07:36overfitting occurs pretty often,
- 11:07:37especially if you're using decision
- 11:07:40trees or if you're using random forest.
- 11:07:43Okay, random forest actually reduces
- 11:07:44overfitting, but sometimes with these
- 11:07:46complex models you can get data
- 11:07:48overfitting. Now, to answer this
- 11:07:49question, first of all, let's understand
- 11:07:51what overfitting really is. So,
- 11:07:54overfitting occurs when a machine
- 11:07:56learning algorithm captures the noise of
- 11:07:58the data. Okay, this causes an algorithm
- 11:08:01to show low bias, but high variance in
- 11:08:03the outcome. Now, what overfitting
- 11:08:05really means is you have trained your
- 11:08:08model way too many times on the training
- 11:08:10data. Okay, so basically the model has
- 11:08:13memorized the training data. It has
- 11:08:15memorized the noise in the training
- 11:08:17data. Okay, so if you feed new data to
- 11:08:20the model during the testing stage, it
- 11:08:22will not be able to recognize the noise
- 11:08:24or it will not be able to recognize any
- 11:08:26sort of correlation in that data. Okay,
- 11:08:28that's why it won't be able to get a
- 11:08:30proper outcome. Okay, that's when
- 11:08:32overfitting happens. You have trained
- 11:08:34the model way too much with the training
- 11:08:36data and this has resulted in inaccurate
- 11:08:39outcome during the testing phase. Okay,
- 11:08:42that's what overfitting is about. Now,
- 11:08:44how do you avoid overfitting? First of
- 11:08:46all, is cross-validation. Now, before
- 11:08:49this also I mentioned that
- 11:08:50cross-validation is the best way to
- 11:08:52obtain a more optimal solution. Now, the
- 11:08:55general idea behind cross-validation is
- 11:08:58to split the training data in order to
- 11:09:00generate multiple mini train test
- 11:09:03splits. Okay, these splits can be used
- 11:09:05to tune your model. Okay, so you're
- 11:09:07basically splitting the training data in
- 11:09:09such a way that, you know, the model
- 11:09:11does not just use the entire training
- 11:09:13data and memorize it. Instead, it's
- 11:09:15going to check the different sets in the
- 11:09:17training data and the different sets in
- 11:09:18the testing data and learn from it.
- 11:09:20Okay, so cross validation is one of the
- 11:09:22best ways to prevent overfitting.
- 11:09:24Another method to prevent overfitting is
- 11:09:27by training the model with more data.
- 11:09:30So, feeding more data to the machine
- 11:09:31learning model will help in better
- 11:09:33analysis and classification. However,
- 11:09:35this method is not always going to work,
- 11:09:38but yeah, this is also one of the ways
- 11:09:40to prevent overfitting. Okay, next we
- 11:09:42have removing features. Now, many times
- 11:09:45the data set contains irrelevant
- 11:09:47features or predictor variables, which
- 11:09:49are not needed for analysis. Such
- 11:09:51features will only increase the
- 11:09:53complexity of the model. Therefore, it
- 11:09:55lead to possibilities of data
- 11:09:57overfitting. Okay, so if you have
- 11:09:59irrelevant data, like for example, if
- 11:10:02you're trying to understand the weight
- 11:10:04of a person depending on its height, and
- 11:10:06you have another variable, let's say,
- 11:10:08you have a variable like the name of the
- 11:10:10person. Okay, now the name of the person
- 11:10:12is not relevant in understanding the
- 11:10:14height of an individual. So, if you have
- 11:10:17irrelevant predictor variables, then it
- 11:10:20will just increase the complexity of the
- 11:10:21model because you have an extra
- 11:10:23irrelevant variable. All right, this
- 11:10:25will only increase the complexity of the
- 11:10:27model. It will not help the model in any
- 11:10:29way. So, make sure you remove irrelevant
- 11:10:31features or you remove redundant
- 11:10:33features. Okay, the next method is early
- 11:10:36stopping. Now, a machine learning model
- 11:10:38is trained iteratively. This will allow
- 11:10:41us to check how well each iteration of
- 11:10:43the model performs. But, after a certain
- 11:10:46number of iterations, the model's
- 11:10:48performance starts to saturate. Further
- 11:10:51training will only result in
- 11:10:52overfitting. Okay, so like I mentioned,
- 11:10:55if you train the model with the same
- 11:10:57data and you make the model memorize the
- 11:11:00data, then it'll just saturate. It won't
- 11:11:02be able to predict any outcomes after a
- 11:11:04point. What you have to do is you have
- 11:11:06to understand where you need to stop
- 11:11:09training the model. So, this can be
- 11:11:11achieved by using a mechanism known as
- 11:11:13early stopping. So, at this point you
- 11:11:15know that you have to stop training the
- 11:11:16model because this might result in
- 11:11:18overfitting. Now, regularization is one
- 11:11:21of the most common ways to prevent
- 11:11:22overfitting. Regularization can be done
- 11:11:25in n number of ways. Okay, the method
- 11:11:27will always depend on the type of
- 11:11:29learner you're implementing. For
- 11:11:30example, pruning is performed on
- 11:11:32decision trees. Now, pruning is a type
- 11:11:35of regularization. Similarly, the
- 11:11:37dropout technique can be used on neural
- 11:11:39networks. And also, there are other
- 11:11:40methods like parameter tuning which can
- 11:11:42help to solve overfitting.
- 11:11:44The next way to prevent overfitting is
- 11:11:46by using ensemble models. Now, ensemble
- 11:11:49learning is a technique that is used to
- 11:11:51create multiple machine learning models
- 11:11:53which are then combined to produce more
- 11:11:56accurate results. So, basically if you
- 11:11:58have one problem statement in machine
- 11:12:00learning, you're going to use like five
- 11:12:03to 10 different models and then you're
- 11:12:05going to calculate the accuracy
- 11:12:07depending on the average of the result
- 11:12:09from each of these models. By this way,
- 11:12:11you will reduce overfitting. Now,
- 11:12:13ensemble models is one of the best ways
- 11:12:16to prevent overfitting. An example is
- 11:12:18the random forest. Random forest uses
- 11:12:21ensemble of decision trees to make more
- 11:12:23accurate predictions and to avoid
- 11:12:25overfitting. So, basically random forest
- 11:12:28is a set of decision trees. So, here
- 11:12:30you're going to train the model by using
- 11:12:32a set of decision trees and this way
- 11:12:34you'll have different data sets and on
- 11:12:36each of these data sets you'll have a
- 11:12:38different decision tree model. Okay,
- 11:12:40this will reduce overfitting to a very
- 11:12:42large extent. That's why in most of the
- 11:12:44cases when you see a decision tree
- 11:12:46having overfitting issues, you'll be
- 11:12:48asked to use random forest. So guys,
- 11:12:51those were the different ways to prevent
- 11:12:52overfitting. Now the next question is
- 11:12:55mention a technique that helps to avoid
- 11:12:57overfitting in a neural network. Now the
- 11:13:00most famous method to prevent
- 11:13:02overfitting in neural networks is
- 11:13:04dropout technique. Okay, now dropout is
- 11:13:06a type of regularization technique which
- 11:13:08is used to avoid overfitting in a neural
- 11:13:10network. So here what you do is you
- 11:13:12randomly select neurons and you drop
- 11:13:15them during the training phase. Right?
- 11:13:17So the dropout value also has to be
- 11:13:19chosen very carefully because a higher
- 11:13:21dropout value will result in under
- 11:13:23learning by the network. So if you're
- 11:13:25dropping out too many predictor
- 11:13:26variables or if you're dropping out too
- 11:13:28many neurons in a neural network, then
- 11:13:30the model will not learn enough. Okay,
- 11:13:33because there's not enough predictor
- 11:13:34variables or not enough neurons. But if
- 11:13:37you have too much of a low rate for a
- 11:13:39dropout value, then this might have a
- 11:13:41very minimal effect. So make sure your
- 11:13:43dropout value is very optimal depending
- 11:13:45on the problem you're trying to solve.
- 11:13:47Okay, so dropout is the technique which
- 11:13:49is used to avoid overfitting in a neural
- 11:13:51network. Next question is what is the
- 11:13:54purpose of deep learning framework such
- 11:13:56as Keras, TensorFlow, and PyTorch? So
- 11:13:59Keras is basically an open-source neural
- 11:14:01network library which is written in
- 11:14:03Python. So basically it is designed to
- 11:14:06enable fast experimentation with deep
- 11:14:08neural networks. Now TensorFlow is
- 11:14:10another open-source software library for
- 11:14:12data flow programming. TensorFlow is
- 11:14:14mainly used in machine learning
- 11:14:16applications. Similarly, PyTorch is
- 11:14:18again an open-source machine learning
- 11:14:20library for Python. Its applications are
- 11:14:23mainly in the field of natural language
- 11:14:24processing. Now I'd say that these three
- 11:14:27deep learning frameworks are the most
- 11:14:29important when it comes to machine
- 11:14:30learning and deep learning because they
- 11:14:32have a varied set of functions in them
- 11:14:35which help in building a better machine
- 11:14:36learning model or a better deep learning
- 11:14:39network. Now let's look at question
- 11:14:41number 24, which is differentiate
- 11:14:43between NLP and text mining. So guys,
- 11:14:46NLP stands for natural language
- 11:14:47processing for those of you who don't
- 11:14:49know. Now first of all, let me clear out
- 11:14:51a confusion between text mining and
- 11:14:53natural language processing. A lot of
- 11:14:55people tend to think that text mining
- 11:14:57and NLP are the same thing, but text
- 11:14:59mining is the broader field and NLP is
- 11:15:02basically an application of text mining
- 11:15:04or it's basically a technique used in
- 11:15:06text mining. So the aim of text mining
- 11:15:09is to extract useful insights from
- 11:15:10structured and unstructured text.
- 11:15:12Whereas the aim of NLP is to understand
- 11:15:15what is conveyed in these texts. Now
- 11:15:17text mining can be done using text
- 11:15:19processing languages like Perl and NLP
- 11:15:22can be achieved using advanced machine
- 11:15:24learning models such as deep neural
- 11:15:25networks. Now the outcome for text
- 11:15:28mining is you'll calculate the frequency
- 11:15:30of words, you'll understand the patterns
- 11:15:32between different words, you'll
- 11:15:33understand the correlations between two
- 11:15:35different words and you'll see how these
- 11:15:37two words occur together more frequently
- 11:15:39and why they occur together more
- 11:15:41frequently. So text mining basically
- 11:15:43will give you a more understanding about
- 11:15:45the words that are used in a document.
- 11:15:48Whereas in NLP, you'll understand the
- 11:15:50grammar behind the text. You'll
- 11:15:52understand in more depth about the
- 11:15:54language that is used in the document or
- 11:15:57in whatever you're trying to analyze. So
- 11:15:59that is the difference between NLP and
- 11:16:01text mining. NLP is a little more
- 11:16:03advanced field because you use deep
- 11:16:05neural networks to perform this. Text
- 11:16:07mining on the other hand makes use of
- 11:16:09NLP. Next question is what are the
- 11:16:12different components of NLP? Now there
- 11:16:14are two components of natural language
- 11:16:16processing, which is natural language
- 11:16:18understanding and natural language
- 11:16:20generation. In natural language
- 11:16:22understanding, you'll basically map your
- 11:16:24input to some useful representation.
- 11:16:26This means that you'll try to understand
- 11:16:28the correlations in your language and
- 11:16:30it'll also include analyzing different
- 11:16:32aspects of the language. All right, so
- 11:16:34this is majorly about understanding your
- 11:16:36text. When it comes to natural language
- 11:16:38generation, here you'll understand how
- 11:16:40to generate text by having a brief plan
- 11:16:43about the text. You'll have sentence
- 11:16:45planning and you'll have text
- 11:16:46realization. Now, natural language
- 11:16:48generation will basically break down
- 11:16:50sentences or will break down text in
- 11:16:52order to understand it better. Okay,
- 11:16:54that's what natural language generation
- 11:16:56is. Natural language understanding is
- 11:16:58more about analyzing your language or
- 11:17:00analyzing the text that you have at hand
- 11:17:02and predicting some useful outcome out
- 11:17:04of it. Generation is more focused on the
- 11:17:07planning aspect of your text. So, these
- 11:17:10are the different components of natural
- 11:17:11language processing. Now, let's look at
- 11:17:13what is stemming and lemmatization in
- 11:17:16natural language processing. Now, what
- 11:17:18is stemming? It is an algorithm which
- 11:17:20works by cutting off the end or the
- 11:17:22beginning of the word and only taking
- 11:17:25into account a list of common prefixes
- 11:17:27and suffixes that can be found in
- 11:17:29inflicted words. Now, for example, on
- 11:17:32the screen you can see that there is a
- 11:17:34detections, detected, detection, and
- 11:17:36detecting.
- 11:17:38Now, if you apply stemming on these four
- 11:17:40words, it will lead to detect. Okay,
- 11:17:43because at the end of the day,
- 11:17:44detections, detected, detection, and
- 11:17:46detecting is the same thing as detect.
- 11:17:48So, stemming will help you remove all of
- 11:17:50these unwanted prefixes and suffixes.
- 11:17:53This way you can analyze the importance
- 11:17:55of the word. All right, you don't have
- 11:17:57to have extra suffix or prefix before
- 11:17:59the word. Now, sometimes during
- 11:18:01stemming, cutting off the ends of the
- 11:18:03words will form an inaccurate result.
- 11:18:05Okay, that's why we have lemmatization.
- 11:18:08In lemmatization, the most important
- 11:18:10thing is the morphological analysis of
- 11:18:12the word. Okay, so here, in order to
- 11:18:14perform lemmatization, you have to have
- 11:18:17a detailed dictionaries which the
- 11:18:18algorithm can look through and it can
- 11:18:21form back to its lemma.
- 11:18:22So, the main difference between stemming
- 11:18:24and lemmatization is that stemming will
- 11:18:26just crop the prefix and the suffix,
- 11:18:28whereas lemmatization will try to
- 11:18:30understand the word in a grammatical way
- 11:18:33and give you an actual word as the
- 11:18:34output. Next is to explain the fuzzy
- 11:18:37logic architecture. All right, so the
- 11:18:39fuzzy logic architecture looks like what
- 11:18:42is shown on the screen. Okay, so
- 11:18:44basically the input is fed into
- 11:18:46something known as the fuzzifier. Okay,
- 11:18:48the fuzzifier or the fuzzification
- 11:18:50module will transform the system's input
- 11:18:53into a number of fuzzy sets. Okay, after
- 11:18:55that it's fed to the controller. Now,
- 11:18:57the controller will have knowledge base
- 11:18:59and the inference engine. Knowledge base
- 11:19:01is basically a set of rules or you can
- 11:19:03say it's an algorithm which is provided
- 11:19:06by experts. Inference engine, like the
- 11:19:08name suggests, will basically infer
- 11:19:10meaning out of these rules. Okay, so
- 11:19:12once you've applied the rules to your
- 11:19:14input, you'll have to draw some useful
- 11:19:16insights or you'll have to infer these
- 11:19:18inputs. Okay, for that you use the
- 11:19:20inference engine. After that, whatever
- 11:19:23inferences and analysis you've formed
- 11:19:25from your inference engine is passed on
- 11:19:26to the defuzzification module. Now, the
- 11:19:29defuzzification will just give you a
- 11:19:31crisp output. All right, it'll give you
- 11:19:33a clear and cut output. That is the
- 11:19:36whole fuzzy logic architecture.
- 11:19:38Now, let's understand the components of
- 11:19:40an expert system. Now, there are three
- 11:19:42important components in an expert
- 11:19:44system, which is knowledge base,
- 11:19:46inference engine, and user interface.
- 11:19:48Now, like I mentioned in fuzzy logic,
- 11:19:50the knowledge base and inference engine
- 11:19:52will play the same part. The user
- 11:19:54interface is basically to provide
- 11:19:56interaction between the users of the
- 11:19:58expert system and the expert system.
- 11:20:01Okay, the expert system is basically a
- 11:20:03program that helps in decision-making
- 11:20:05process. Okay, so here the knowledge
- 11:20:07base will contain some high-quality
- 11:20:09knowledge or it contain rules and
- 11:20:11algorithms. The inference engine will
- 11:20:13acquire all the knowledge that is needed
- 11:20:16to solve the problem. And the user
- 11:20:18interface is just for the users to
- 11:20:19interact with the expert system. Okay,
- 11:20:21this is the whole expert system
- 11:20:23component. Now, obviously this This a
- 11:20:25little more complex than this, but uh
- 11:20:27stick to how this works. All right, I'm
- 11:20:29just going to tell you the working of
- 11:20:30expert systems and fuzzy logic. If I
- 11:20:33start to explain each and everything,
- 11:20:35it's going to take a lot of time. All
- 11:20:36right, so let's move on to our next
- 11:20:38question, which is how is computer
- 11:20:40vision and AI related? Now, computer
- 11:20:43vision is a field of artificial
- 11:20:45intelligence that is used to obtain
- 11:20:47information from images or
- 11:20:49multi-dimensional data. Now, computer
- 11:20:51vision is basically the concept behind
- 11:20:53the self-driving cars that you see these
- 11:20:55days. All right, computer vision
- 11:20:57involves a lot of image processing. So,
- 11:20:59machine learning algorithms like K-means
- 11:21:01can be used in image segmentation.
- 11:21:03Support vector machines can be used for
- 11:21:05image classification. Okay, that's how
- 11:21:07computer vision and AI are related.
- 11:21:09Because most of the things that happen
- 11:21:11in computer vision like image processing
- 11:21:13and segmentation make use of machine
- 11:21:15learning algorithms like K-means and
- 11:21:17support vector machines. So, to sum it
- 11:21:19up, computer vision makes use of
- 11:21:21artificial intelligence technologies to
- 11:21:24solve complex problems such as object
- 11:21:26detection, image processing, and so on.
- 11:21:28That is the relationship between
- 11:21:30computer vision and AI. Now, question
- 11:21:32number 30 is which is better for image
- 11:21:35classification? Is it supervised or
- 11:21:38unsupervised classification? So, guys,
- 11:21:40earlier in the session we discussed what
- 11:21:42supervised learning is and what
- 11:21:43unsupervised learning is. In supervised
- 11:21:46learning, the images are interpreted
- 11:21:48manually by the machine learning expert
- 11:21:50to create feature classes. Now, what
- 11:21:52this means is you're manually going to
- 11:21:54feed a labeled set of data to the
- 11:21:56supervised learning model. All right,
- 11:21:58that's how supervised learning works.
- 11:22:00You're manually going to feed a set of
- 11:22:02images which are labeled to the
- 11:22:04classifier. In unsupervised learning,
- 11:22:06the machine learning software creates
- 11:22:08feature classes based on image pixel
- 11:22:10values. So, basically in unsupervised
- 11:22:12classification, the model itself has to
- 11:22:15figure out what to do and what not to
- 11:22:17do. Okay, so it'll create a own feature
- 11:22:19class based on some values such as image
- 11:22:22pixels or it can also use the image
- 11:22:24color or it can use intensity factors in
- 11:22:27order to classify. So, if you ask me it
- 11:22:29is better to opt for supervised
- 11:22:31classification because you're manually
- 11:22:33inputting images with a lot more
- 11:22:35information. Okay, whereas in
- 11:22:37unsupervised learning you're totally
- 11:22:38letting the model perform everything.
- 11:22:40Okay, so in image classification, I
- 11:22:42think it's better to go for supervised
- 11:22:44learning. Now, let's look at question
- 11:22:46number 31. The next question is finite
- 11:22:50difference filters in image processing
- 11:22:51are very susceptible to noise. To cope
- 11:22:54up with this, which method can you use
- 11:22:56so that there would be minimal
- 11:22:58distortions by noise? Now, the noise in
- 11:23:01an image can be due to high intensity or
- 11:23:03high contrast. Okay, so if you increase
- 11:23:06the contrast and increase the intensity
- 11:23:08of an image, you won't be able to
- 11:23:10understand each pixel. Okay, so each
- 11:23:12pixel will have a value associated to it
- 11:23:15and if the intensity and the contrast of
- 11:23:17that pixel is a little too much, it'll
- 11:23:19be hard for us to understand the image
- 11:23:21properly. It'll be hard to perform image
- 11:23:24analysis because we don't have a clear
- 11:23:26image. Contrast and intensity will just
- 11:23:28cause noise in an image. So, the best
- 11:23:31method to remove this is image
- 11:23:32smoothing. Okay, it is used for reducing
- 11:23:35noise by forcing pixels to be more like
- 11:23:38their neighbors. Okay, this way you'll
- 11:23:40have a faded image or you'll have a more
- 11:23:42equalized image. Now, the next question
- 11:23:45is how is game theory and AI related? So
- 11:23:48guys, AI is actually applied in a vast
- 11:23:51number of fields. Okay, so a lot of
- 11:23:53fields from computer vision to game
- 11:23:55theory to machine learning, AI is always
- 11:23:58a concept behind these fields. Most of
- 11:24:01the game examples that we see make use
- 11:24:03of reinforcement learning or deep neural
- 11:24:05networks. Now, deep neural networks and
- 11:24:07reinforcement learning are very closely
- 11:24:09related to AI because they are branches
- 11:24:11of machine learning. So, machine
- 11:24:13learning is majorly involved in game
- 11:24:15theory. An example of this is in Dota 2
- 11:24:18also they make use of machine learning.
- 11:24:20So, game theory is just a very logical
- 11:24:23approach to solving a problem. And
- 11:24:25machine learning is the best way to
- 11:24:27implement game theory. Now, question
- 11:24:29number three is what is the minimax
- 11:24:31algorithm? Explain the terminologies
- 11:24:33involved in the problem.
- 11:24:35Now guys, minimax is one of the main
- 11:24:37algorithms which is used in game theory.
- 11:24:39All right, it is used to choose an
- 11:24:41optimal move for a player assuming that
- 11:24:43the other player is also playing
- 11:24:45optimally. Meaning that both of these
- 11:24:47players are playing in order to win and
- 11:24:50you're going to use the minimax
- 11:24:51algorithm on one of these players so
- 11:24:53that they choose the optimal move. In
- 11:24:56order to understand the minimax
- 11:24:57algorithm, you need to know what are the
- 11:24:59components in a game. Okay, there's
- 11:25:01something known as game tree. It is
- 11:25:03basically a tree structure which
- 11:25:05contains all the possible moves in a
- 11:25:07game. If it's up, down, right, left, any
- 11:25:09strategy, everything is mentioned in the
- 11:25:11game tree. Now, initial state is
- 11:25:13obviously the initial position of the
- 11:25:15player on the board. All right, the
- 11:25:17successor function it defines all the
- 11:25:19possible moves that a player can make.
- 11:25:22We'll understand this in the next
- 11:25:23question itself, so don't worry if you
- 11:25:25haven't understood this properly.
- 11:25:27Terminal state is obviously the end of
- 11:25:29the game. It's basically the state which
- 11:25:31will lead to the end game or it will
- 11:25:33lead to your destination. Utility
- 11:25:35function is a numerical value for the
- 11:25:37output of the game. So guys, these were
- 11:25:39the terminologies and this is what the
- 11:25:41minimax algorithm is. It is basically a
- 11:25:44game theory algorithm which helps a
- 11:25:46player choose the best optimal policy in
- 11:25:49order to win a game. I'll explain this
- 11:25:51in more depth in the upcoming slides.
- 11:25:54So, let's move on. Now, the next couple
- 11:25:56of questions are going to be
- 11:25:57scenario-based questions. Now, such
- 11:25:59questions are very important in an
- 11:26:01interview because this is where the
- 11:26:03interviewer will understand how well you
- 11:26:05know the concepts. So, the first
- 11:26:07question is show the working of the
- 11:26:09minimax algorithm using the tic-tac-toe
- 11:26:11game. Now, one of the major applications
- 11:26:14of the minimax algorithm is the
- 11:26:16tic-tac-toe game. Okay, you can
- 11:26:18understand and analyze all the possible
- 11:26:20outcomes of the tic-tac-toe game by
- 11:26:22using the minimax algorithm. Let's see
- 11:26:24how this happens. Now, first of all, in
- 11:26:27a minimax algorithm or in a game, there
- 11:26:29are two players involved. Okay, the max
- 11:26:32is the player that tries to get the
- 11:26:33highest possible score, and min is the
- 11:26:36player that tries to get the lowest
- 11:26:37possible score. So, this algorithm is
- 11:26:40designed in such a way that assuming
- 11:26:42that there going to be two players, and
- 11:26:44obviously one player is going to win the
- 11:26:45game, and that is the max player, and
- 11:26:48min is the player which loses the game
- 11:26:50and has the lowest possible score. Now,
- 11:26:52the first step in the minimax algorithm
- 11:26:54is to generate the entire game tree.
- 11:26:57Okay, the game tree is all the possible
- 11:26:59outcomes that can happen in tic-tac-toe.
- 11:27:01Okay, in the figure you can see that
- 11:27:02first X is aligned in the first box,
- 11:27:04then in the second box, third box, and
- 11:27:06so on. All the possible actions that you
- 11:27:09can take in a tic-tac-toe game are put
- 11:27:11in this game tree. And then, step number
- 11:27:14two is to apply the utility function to
- 11:27:16get the utility values from all the
- 11:27:18terminal states. Getting utility value
- 11:27:21is important because this is how you'll
- 11:27:23understand your outcome. Okay, you'll
- 11:27:24understand if you're going to win or
- 11:27:25lose. Now, in the terminal states,
- 11:27:28whatever numbers you see over here,
- 11:27:30these are the utility values. Now, step
- 11:27:32three is determine the utilities of the
- 11:27:34higher nodes with the help of utilities
- 11:27:36of the terminal nodes. Now, in this
- 11:27:39diagram, you can see that in the
- 11:27:40terminal nodes, we have the utility
- 11:27:42values. The step three is to get utility
- 11:27:45values in the higher stages, which is
- 11:27:48the min stage. All right, these two
- 11:27:50circles, you need to fill in the utility
- 11:27:51values by using the utility values which
- 11:27:54are in the terminal state.
- 11:27:55Now, how do you calculate the utility
- 11:27:57value? Let's start by calculating the
- 11:28:00utility value of the left node. Okay,
- 11:28:02this red color node, we'll start by
- 11:28:05calculating this.
- 11:28:06Now, you calculate that by finding the
- 11:28:08minimum of the three nodes that it's
- 11:28:11leading to. Now, this red node is
- 11:28:13leading to three, five, and 10. And the
- 11:28:15minimum out of three, five, 10 is three.
- 11:28:17So, the utility value for this red node
- 11:28:20is going to be three. Okay, similarly
- 11:28:22for this green node, it's going to be
- 11:28:23two because the minimum value between
- 11:28:25two and two is still two. Now, step four
- 11:28:27is to fill in these utility values that
- 11:28:29you've calculated. So, now we have a
- 11:28:32minimax algorithm which has all the
- 11:28:34utility values filled in. Now, the only
- 11:28:36utility value which isn't filled is the
- 11:28:38one with max. Okay, the one on the root
- 11:28:41node. Here, we haven't filled the
- 11:28:43utility value. Again, to fill this
- 11:28:45value, you're going to check the nodes
- 11:28:46which are directly connected to it,
- 11:28:48which is three and two. You'll find the
- 11:28:50maximum between these two because this
- 11:28:52is the max function. All right, so here
- 11:28:54you'll get a value of three. So, that's
- 11:28:57why the best opening move for max is the
- 11:28:59left node. Okay, you can make use of the
- 11:29:02left node in order to win the game. This
- 11:29:04is the first step that the max player
- 11:29:06has to take in order to get to the path
- 11:29:08of winning the game. So guys, by doing
- 11:29:11this for each and every step, you can
- 11:29:13win the game. Okay, so you'll have to
- 11:29:15calculate the utility value at the
- 11:29:17terminal nodes. You'll have to move up
- 11:29:19to the other hierarchical nodes above
- 11:29:21it, calculate the utility values there
- 11:29:23until you reach the root node. Okay,
- 11:29:25once you reach the root node, you'll get
- 11:29:26a utility value and that utility value
- 11:29:29will be connected to some move or some
- 11:29:31node. You'll have to take that node or
- 11:29:34you'll have to take that move in the
- 11:29:36game in order to win the game. So, this
- 11:29:38way you'll have to calculate the utility
- 11:29:40value for each and every move that the
- 11:29:42player makes so that the player will win
- 11:29:44the game.
- 11:29:45So guys, minimax algorithm is quite easy
- 11:29:47and it's very understandable. All you
- 11:29:49need to know is a little bit of math in
- 11:29:51order to solve this problem. Question
- 11:29:52number 35 is which method is used for
- 11:29:56optimizing a minimax based game? Now,
- 11:29:58this is not a scenario-based question,
- 11:30:00but this question is usually asked if an
- 11:30:02interviewer asks you about a minimax
- 11:30:05game.
- 11:30:05Now, the best way to optimize a minimax
- 11:30:08game is by using something known as
- 11:30:10alpha-beta pruning. Now, the main thing
- 11:30:12about alpha-beta pruning is that it'll
- 11:30:14remove all the nodes that are not
- 11:30:16affecting the final decision. It's just
- 11:30:18a faster way to reach your outcome.
- 11:30:21That's what alpha-beta pruning is all
- 11:30:23about. So, let's look at an example to
- 11:30:26understand this. Okay, let's say there
- 11:30:27was another node over here. Okay, here
- 11:30:29you can see that this is going down to a
- 11:30:31terminal state with utility value two.
- 11:30:34Okay, now you don't know the value of
- 11:30:36the other two nodes, but if you use
- 11:30:38minimax to calculate the utility of the
- 11:30:41other two nodes, you'll get a value of
- 11:30:43three. So, in this example again, we'll
- 11:30:45start at the terminal nodes. So, three,
- 11:30:47five, 10 are the utility values here.
- 11:30:50So, this will give us a value of three
- 11:30:52because we're calculating the minimum
- 11:30:53over here. Now, here you have two and
- 11:30:56you have two unknown values. You have A
- 11:30:58or B. Okay, I've named them as A and B.
- 11:31:01Okay, let's leave this for now. Let's go
- 11:31:02to the next node. Okay, here the
- 11:31:04possibilities are two, seven, and three.
- 11:31:07So, the minimum between two, seven,
- 11:31:08three is two.
- 11:31:10Okay, so here there's going to be three.
- 11:31:11There's going to be a value, let's say
- 11:31:13C, and here there's going to be a value,
- 11:31:15let's say two. Now, we know that the
- 11:31:18maximum between three, C, and something
- 11:31:20else will be three. Okay, that's because
- 11:31:23two is the minimum value over here, and
- 11:31:25the maximum will obviously be three. So,
- 11:31:27the hint here is in the two AB node. We
- 11:31:30know that the value or the utility value
- 11:31:32will obviously be equal to two or it'll
- 11:31:35be less than two because you're
- 11:31:36calculating the minimum in this step.
- 11:31:38Now, if you calculate the max out of
- 11:31:40these three values, we'll obviously get
- 11:31:42the answer as three.
- 11:31:44So, this way this entire node itself is
- 11:31:46removed because you don't need it to get
- 11:31:48to the final answer. Okay, that's what
- 11:31:50alpha-beta pruning is all about. It'll
- 11:31:52identify the nodes which are not going
- 11:31:54to affect the final decisions, and it'll
- 11:31:56just remove those nodes. So guys, this
- 11:31:58is how the optimization for a minimax
- 11:32:01game is done. It's done using the
- 11:32:03alpha-beta pruning.
- 11:32:04The next question is which algorithm
- 11:32:06does Facebook use for face verification?
- 11:32:09Now guys, even though this might seem
- 11:32:11like a general knowledge question, this
- 11:32:14is actually a very important sort of
- 11:32:16question in artificial intelligence.
- 11:32:18Okay, even if you don't know the answer
- 11:32:20to this, you should have an idea of how
- 11:32:22the algorithm might work. Okay, that's
- 11:32:24exactly what the interviewer wants to
- 11:32:26know. He wants to know whether you know
- 11:32:28how the algorithm works step-by-step.
- 11:32:30You might not know the final answer or
- 11:32:32you might not know the exact algorithm
- 11:32:34which Facebook uses because obviously
- 11:32:35Facebook uses more than one algorithm to
- 11:32:38achieve this, but you must know the
- 11:32:40steps in which the face verification
- 11:32:42works. Okay, that's the main goal behind
- 11:32:44this question. Now anyway, the algorithm
- 11:32:47used by Facebook is the deep face. Okay,
- 11:32:49deep face makes use of a lot of neural
- 11:32:51networks and a lot of algorithms. Okay,
- 11:32:53so it works on artificial intelligence
- 11:32:56techniques, like I mentioned earlier.
- 11:32:58Now how would a face verification work?
- 11:33:00How do you think it works? Now it starts
- 11:33:02by an input. So the idea here is you
- 11:33:05have to scan a huge number of photos and
- 11:33:07you'll have to feed it to the algorithm.
- 11:33:10Okay, now these photos can have a lot of
- 11:33:12disturbance, a lot of distortions and it
- 11:33:14can have different angles or anything
- 11:33:16like that. Okay, you have to feed any
- 11:33:18sort of photos that are possible. Okay,
- 11:33:20even they are complex to understand, but
- 11:33:22you have to still feed the model with
- 11:33:24all the possible photos that you can
- 11:33:25get. Now the next step is the main
- 11:33:27process. Here there are a few important
- 11:33:30things which is detect, align, represent
- 11:33:32and classify. Detect is basically you'll
- 11:33:34detect facial features. All right,
- 11:33:37you'll try to understand the distance
- 11:33:38between the eyes and the nose of a
- 11:33:40person, the way the lips is aligned or
- 11:33:43anything like that. That's what aligning
- 11:33:45is about. You'll align and compare the
- 11:33:46various features in the face in order to
- 11:33:49understand the facial features. You'll
- 11:33:51represent the key patterns by using some
- 11:33:533D graphs or 3D models. Okay, it's very
- 11:33:55important to visualize whatever you get
- 11:33:58because visualization will help you
- 11:34:00understand the correlation. It'll help
- 11:34:02you understand that okay, the eyes are
- 11:34:03at this distance, the nose is at this
- 11:34:05distance, and so on. Finally, you'll
- 11:34:07classify the images based on the
- 11:34:09similarity. All right, that's how the
- 11:34:11output comes out. And basically, the
- 11:34:13output is you need to detect whether two
- 11:34:15images represent the same person or not.
- 11:34:18Okay, so by studying the facial features
- 11:34:20and by using image processing and by
- 11:34:22using computer vision, Facebook's
- 11:34:24achieves face verification. You need to
- 11:34:26know the basic concept behind face
- 11:34:28verification. You need to know that it
- 11:34:30starts with image collection or data
- 11:34:32acquisition. After that, you're going to
- 11:34:34perform image processing or
- 11:34:36pre-processing. All right, and this
- 11:34:38might involve performing conversions
- 11:34:40from RGB to any other state like YCbCr.
- 11:34:44Okay, I'm not going to go in depth of
- 11:34:45this because the video will get to about
- 11:34:472-3 hours. So, there are a lot of ways
- 11:34:49in which you can convert an image and
- 11:34:51you know, you can understand the image
- 11:34:52more properly. Also, an important thing
- 11:34:54in image processing is it's not done
- 11:34:57just based on the image. All right.
- 11:34:59You're going to take the image, you're
- 11:35:00going to form a matrix, and you're going
- 11:35:02to have pixel values in these matrix.
- 11:35:04So, it is a very in-depth approach. All
- 11:35:06right, it's not a very simple approach.
- 11:35:08When I'm speaking about it, it might
- 11:35:10seem simple, but image analysis is very
- 11:35:13in-depth. After image analysis, you can
- 11:35:15perform image segmentation. All right,
- 11:35:17image segmentation is basically dividing
- 11:35:19the image into different segments and
- 11:35:21studying each image segment separately.
- 11:35:24Then after that, you can do feature
- 11:35:25extraction. Here, you'll try to
- 11:35:27understand the features and how they are
- 11:35:29related to each other. Finally, you'll
- 11:35:31classify the images and see whether two
- 11:35:34images represent the same person or not.
- 11:35:37So, the main idea behind the Facebook
- 11:35:39algorithm is image processing, neural
- 11:35:41networks, machine learning, and computer
- 11:35:43vision. All right, and all of this comes
- 11:35:45down to artificial intelligence. Next,
- 11:35:48we have explain the logic behind
- 11:35:50targeted marketing and how can machine
- 11:35:52learning help with this? Now, target
- 11:35:54marketing is something that we see very
- 11:35:56often. All right, let's say that you
- 11:35:58were looking for some shoe on Amazon. In
- 11:36:01a day or two, you just open up YouTube
- 11:36:03and Facebook. You'll see that you'll get
- 11:36:05ads of shoes from Amazon. Okay, this is
- 11:36:08targeted marketing. So, basically Amazon
- 11:36:10knows that we've been looking for a
- 11:36:12particular type of shoe, so it's going
- 11:36:14to target you with that particular ad.
- 11:36:16Okay, this is what targeted marketing is
- 11:36:18in short. Targeted marketing can be done
- 11:36:20in different ways. For example, it can
- 11:36:23be done depending on your geography or
- 11:36:25it can be done depending on your social
- 11:36:27economic profile. Okay, let's say that
- 11:36:30Amazon has details about your age, it
- 11:36:33has details about what sport you like to
- 11:36:35play. Let's say that you've been
- 11:36:36browsing through a lot of sports. Okay,
- 11:36:39you've been browsing through a lot of
- 11:36:40sport equipments or something like that.
- 11:36:43Amazon will know that you're interested
- 11:36:44in this by using machine learning, of
- 11:36:46course, and it will send you ads based
- 11:36:49on what you're interested in. Okay, this
- 11:36:51is what target marketing really is.
- 11:36:53Now, how does machine learning come into
- 11:36:55target marketing? Okay, so there's
- 11:36:57something known as text analytics
- 11:36:59systems. Now, the applications for text
- 11:37:01analytics ranges from search
- 11:37:03applications, text classification, named
- 11:37:06entity recognition, or pattern search.
- 11:37:09Okay, so it's basically a way to
- 11:37:10understand what you're looking for.
- 11:37:12Okay, they'll try to understand your
- 11:37:14search history and they'll try to target
- 11:37:16you by using your interests. Clustering
- 11:37:19is another way of targeted marketing.
- 11:37:21All right, you'll cluster customers who
- 11:37:23have similar interests and you'll send
- 11:37:25them similar ads or you'll send them
- 11:37:26similar offers. Classification is
- 11:37:28another method used for targeted
- 11:37:30marketing. Now, here you'll make use of
- 11:37:33algorithms like decision trees and
- 11:37:34neural networks. Now, recommender
- 11:37:36systems is what I spoke about earlier.
- 11:37:38When it comes to Amazon, they recommend
- 11:37:40items to you based on your interest or
- 11:37:43based on people who have similar
- 11:37:44interests like you. Market basket
- 11:37:47analysis is another method that machine
- 11:37:49learning uses for marketing. Okay, here
- 11:37:51basically you'll understand a
- 11:37:53combination of products that are
- 11:37:54frequently bought. Okay, by
- 11:37:56understanding what two products are
- 11:37:58frequently bought, you can give some
- 11:37:59offers or you can give some discounts on
- 11:38:01those products so that people buy more
- 11:38:03and more. Okay, that's how market basket
- 11:38:05analysis also works. So guys, this is
- 11:38:07what targeted marketing is.
- 11:38:10Now next is how can AI be used to detect
- 11:38:13fraud? AI is used in a lot of ways in
- 11:38:16credit card fraud detection. It's used
- 11:38:18in detecting anomalies and all of that.
- 11:38:20It basically makes use of machine
- 11:38:22learning algorithms to do this. Now
- 11:38:24let's try to understand how this process
- 11:38:26works. Okay, first it begins with data
- 11:38:28extraction or data collection. So at
- 11:38:30this stage data is either collected
- 11:38:32through a survey or through web
- 11:38:34scraping. Okay, if you're trying to
- 11:38:35detect credit card fraud then
- 11:38:37information about the customer's
- 11:38:39collected. All right, this includes any
- 11:38:41transactional or any shopping and
- 11:38:43personal details. Next is data cleaning.
- 11:38:45So at this stage the redundant data must
- 11:38:48be removed. Any inconsistencies or any
- 11:38:50missing values that you have in your
- 11:38:52data, it has to be removed because they
- 11:38:54lead to wrongful prediction. Okay, so
- 11:38:56therefore you have to get rid of any
- 11:38:58inconsistencies in this stage. Next we
- 11:39:01have data exploration and analysis. Now
- 11:39:03this is the most important step in AI.
- 11:39:06Okay, here you study the relationship
- 11:39:08between various predictable variables.
- 11:39:10For example, let's say that a person has
- 11:39:12spent an unusual sum of money on a
- 11:39:15particular day. Now the chances for a
- 11:39:17fraudulent occurrence is very high
- 11:39:19because usually the person is not used
- 11:39:21to spending this much money. So such
- 11:39:23patterns have to be detected and
- 11:39:24understood in data exploration and
- 11:39:26analysis. This is followed by building a
- 11:39:29machine learning model. Now here there
- 11:39:31are any machine learning algorithms that
- 11:39:33can be used for fraud detection or
- 11:39:35anomaly detection. One such example is
- 11:39:37logistic regression, okay, which is a
- 11:39:39classification algorithm and it can be
- 11:39:42used to classify events into two
- 11:39:44classes. Okay, you can use them to
- 11:39:46classify a person or classify an event
- 11:39:49as either fraudulent and non-fraudulent.
- 11:39:52Then comes model evaluation. Here you'll
- 11:39:54basically test the efficiency of the
- 11:39:56machine learning model. Okay, so if
- 11:39:58there's any room for improvement, then
- 11:40:00you can perform parameter tuning and you
- 11:40:02can improve the model. This will just
- 11:40:04improve the accuracy of the model. So
- 11:40:06guys, all of these complex problems like
- 11:40:09fraud detection or object detection, all
- 11:40:12of this is done through a process. Okay,
- 11:40:14and in general the process is data
- 11:40:16collection, data cleaning, exploration
- 11:40:18and analysis, building a model and model
- 11:40:21evaluation. Most of these complex
- 11:40:23problems can be solved by using this
- 11:40:25approach. Now let's look at our next
- 11:40:27question.
- 11:40:28Okay, a bank manager is given a data set
- 11:40:31containing records of thousands of
- 11:40:33applicants who have applied for loan.
- 11:40:35How can AI help the manager understand
- 11:40:37which loans he can approve?
- 11:40:39To be more specific, this problem
- 11:40:41statement can easily be solved by using
- 11:40:43the KNN algorithm. Okay, KNN is
- 11:40:46basically stands for K nearest neighbor.
- 11:40:49All right, this is a classification and
- 11:40:51a regression algorithm.
- 11:40:53So if you use a KNN algorithm, it'll
- 11:40:55form two classes. One is the loan is
- 11:40:57approved and the other is applicants
- 11:40:59whose loan is not been approved. So like
- 11:41:02I said, K nearest neighbor is a
- 11:41:03supervised learning algorithm that
- 11:41:05classifies a new data point into the
- 11:41:08target class depending on the features
- 11:41:10of its neighboring data points. All
- 11:41:12right, so KNN basically focuses on the
- 11:41:14neighbors and it understands that if a
- 11:41:16new data point is similar to one of its
- 11:41:18neighbors, then it has to classify that
- 11:41:20new data point into that neighbor's
- 11:41:22class. Now again, the methodology for
- 11:41:24solving this problem is same. You start
- 11:41:27by data collection, data cleaning,
- 11:41:29exploration and analysis, building a
- 11:41:31model and model evaluation. All right.
- 11:41:34So, while data collection, you can
- 11:41:35collect data like account balance,
- 11:41:37credit amount, age, occupation, loan
- 11:41:40records, and all of that. So, by using
- 11:41:42this data, you can predict whether or
- 11:41:43not to approve the loan of an applicant.
- 11:41:46Data cleaning, again, you have to remove
- 11:41:48any variables which will not help the
- 11:41:50model. Okay, any variables which will
- 11:41:52just increase the complexity of the
- 11:41:54model. Okay, so you'll remove such
- 11:41:55variables at this stage. In data
- 11:41:58exploration and analysis, you will
- 11:42:00understand the patterns in your data.
- 11:42:02Okay, let's see that a person has a
- 11:42:04history of unpaid loans. Okay, if any
- 11:42:07person or any applicant has a history of
- 11:42:09unpaid loans, then the chances are that
- 11:42:11he might not get approval on his loan
- 11:42:13application. Okay, this is obvious
- 11:42:15because the manager is going to see that
- 11:42:17his previous loans are still due. So,
- 11:42:20that's why he won't be able to approve
- 11:42:21the application. So, these are the kind
- 11:42:23of patterns that are detected in
- 11:42:25exploration and analysis. Now, building
- 11:42:28a machine learning model, you can use n
- 11:42:30number of models when it comes to
- 11:42:31predicting whether an applicant loan
- 11:42:34request is approved or not. Now, like I
- 11:42:36mentioned, one of the easy algorithms
- 11:42:38that you can implement is the K nearest
- 11:42:39algorithm. Okay, it can be used for both
- 11:42:42classification and regression. It
- 11:42:44classify the applicant's loan request
- 11:42:46into approved or disapproved based on
- 11:42:49the socio-economic profile of the
- 11:42:50applicant. Okay, based on variables like
- 11:42:53loan based on variables like the salary,
- 11:42:55the occupation of the applicant. Now,
- 11:42:58model evaluation, again, is the same
- 11:43:00thing. You'll basically evaluate the
- 11:43:02efficiency of the model. You'll try to
- 11:43:03improve the accuracy of the model by
- 11:43:05using parameter tuning or cross
- 11:43:07validation. So, guys, this is how a bank
- 11:43:10manager can understand whether a loan
- 11:43:12can be approved or not. Again, in this
- 11:43:14question, they are just trying to test
- 11:43:16if you know how the flow of the problem
- 11:43:18will go. If you know how this problem
- 11:43:20can be solved. You don't have to know
- 11:43:22the exact details, but you have to know
- 11:43:24how you can approach the problem.
- 11:43:26Okay, now let's move on and look at
- 11:43:28question number 40. Now the question
- 11:43:30here is place an agent in any one of the
- 11:43:32rooms and the goal is to reach outside
- 11:43:35the building. Can this be achieved
- 11:43:37through AI? If yes, explain how it can
- 11:43:39be done. Now in this question there is a
- 11:43:42diagram along with a small explanation.
- 11:43:44Okay, so basically there are four rooms
- 11:43:47in this diagram. Basically 0 1 2 3 and 4
- 11:43:50represent rooms and this 5 represents
- 11:43:54outside the building. Okay, now the goal
- 11:43:56is to place an agent in any one of these
- 11:43:59rooms in such a way that he has to reach
- 11:44:01room number five or he has to reach
- 11:44:03outside the building. Now they've also
- 11:44:05mentioned that a room number one and
- 11:44:07room number four directly lead outside
- 11:44:10the building. That's correct because
- 11:44:11room number one is directly connected to
- 11:44:13five and four is also directly connected
- 11:44:16to five. Right? Four leads outside the
- 11:44:18building. Now if you look at room number
- 11:44:21zero, if you want to go from zero to
- 11:44:23five, first from zero you'll have to go
- 11:44:25to four and then only you can go to
- 11:44:26five. Similarly, if you look at room
- 11:44:28number three, if you want to go from
- 11:44:30three to five, you'll have to take
- 11:44:32three, then you'll have to go to one and
- 11:44:34then only you'll have to go to five.
- 11:44:36These are not directly connected to
- 11:44:38outside the building, whereas room
- 11:44:39number one and four are directly
- 11:44:41connected to outside the building. All
- 11:44:43right, I hope the question is clear. Now
- 11:44:45as soon as you read the question, you
- 11:44:47must know that this is a reinforcement
- 11:44:49learning question. All right, it's
- 11:44:51pretty clear because they have mentioned
- 11:44:53that there is an agent which is going to
- 11:44:55be placed in any one of the rooms and he
- 11:44:57has to basically explore the environment
- 11:45:00and reach five, which is basically
- 11:45:02outside the building. So as soon as you
- 11:45:03read the question, the first thing that
- 11:45:05should come into your head is that this
- 11:45:06is a reinforcement learning problem. Now
- 11:45:09this problem can be solved by using the
- 11:45:11Q-learning algorithm. Now if you
- 11:45:13remember earlier in the session, I
- 11:45:15discussed what exactly Q-learning is and
- 11:45:17how it works. So Q-learning is basically
- 11:45:20a reinforcement learning algorithm which
- 11:45:22is used to solve reward-based problems.
- 11:45:24So, now let's look at how we'll solve
- 11:45:26the problem. First step would be to
- 11:45:29represent the rooms on a graph. All
- 11:45:31right, so each room over here you'll
- 11:45:33represent it as a node and each door
- 11:45:35will represent a link. So, if you look
- 11:45:37at this figure over here, this is our
- 11:45:39original figure that was given in the
- 11:45:41question and now this is the graph that
- 11:45:43we draw from this figure. So, we have
- 11:45:46node one. Let's look at how node one is
- 11:45:49connected to node three. So, basically
- 11:45:51you can go from node one to node three
- 11:45:53and you can go from node three to node
- 11:45:55one. If you look at the diagram, there
- 11:45:57is a direct connection from one to three
- 11:45:59and three to one. Okay, let's look at
- 11:46:01one and two. Now, there's no link
- 11:46:03between one and two because if you look
- 11:46:05over here, if you want to go from room
- 11:46:07number one to room number two, you
- 11:46:08cannot directly go. You'll have to go
- 11:46:11from room number one to room number
- 11:46:12three and only then you'll reach room
- 11:46:14number two. All right, that's why
- 11:46:15there's no link between room number one
- 11:46:18and node number two.
- 11:46:19Similarly, if you look at node one and
- 11:46:22four, they are directly connected to
- 11:46:24five. This is because room number one
- 11:46:26and four directly lead to this goal. Our
- 11:46:29goal is to reach room number five. So,
- 11:46:32one and four are directly connected to
- 11:46:34five, whereas the others are connected
- 11:46:36just like how they're shown in this
- 11:46:38figure. Okay, it's pretty
- 11:46:39understandable, guys. This is just
- 11:46:41logic. Now, let's look at the next step.
- 11:46:43Now, the next step is to associate a
- 11:46:46reward value to each door. What we're
- 11:46:48going to do here is we're going to build
- 11:46:50a reward matrix. Now, I'll tell you what
- 11:46:52that exactly means. So, for the doors
- 11:46:55that lead directly to the end goal,
- 11:46:57which is room number five, you'll assign
- 11:46:59a reward of 100 to those doors. So, if
- 11:47:01you're traversing from node number one
- 11:47:03to node number five, you'll get a reward
- 11:47:05of 100. Similarly, if you're traversing
- 11:47:08from four to five, you'll get a reward
- 11:47:09of 100. Five to five also you'll get a
- 11:47:12reward of 100 because your end goal is
- 11:47:14five, right? So, basically any link that
- 11:47:16leads directly to our end goal, for that
- 11:47:19link, we're going to assign a reward of
- 11:47:20100. Now, doors that are not directly
- 11:47:23connected to the target room will have a
- 11:47:25reward zero. This is because if you take
- 11:47:28room number two or if you take room
- 11:47:29number three, you won't reach room
- 11:47:31number five directly. And our goal here
- 11:47:34is to reach room number five. That's why
- 11:47:36for the other doors, we've given a
- 11:47:38reward of zero. For door number one and
- 11:47:40door number four, however, we have
- 11:47:42rewards of 100. Similarly, for door
- 11:47:44number five or room number five, also we
- 11:47:46have a reward of 100. So, basically,
- 11:47:49each action or each link will represent
- 11:47:52a reward. So, let's say that you're
- 11:47:54traversing from room number one to room
- 11:47:56number four. If you go from room number
- 11:47:59one to room number three and then three
- 11:48:01to four, your reward is going to remain
- 11:48:03zero because you're not reaching the end
- 11:48:05goal here. Only if you traverse from one
- 11:48:07to five or if you traverse from four to
- 11:48:09five or five to five, you'll get a
- 11:48:11reward of 100. All right, it's as simple
- 11:48:14as that. Now, let's see how the
- 11:48:15Q-learning algorithm works in this
- 11:48:18particular problem statement. Now, there
- 11:48:20are two main components in this
- 11:48:22algorithm. All right, there is state and
- 11:48:24there is action. Now, basically, all
- 11:48:26these rooms will represent the state and
- 11:48:28the agent's movement from one room to
- 11:48:30the other will represent an action. So,
- 11:48:33basically, 0 1 2 3 4 and 5 represent the
- 11:48:36state and let's say you're traversing
- 11:48:38from two to three. This two to three
- 11:48:40will basically represent an action. In
- 11:48:42order to make you understand how this
- 11:48:44works, let's say that you're traversing
- 11:48:46from room number two to room number
- 11:48:47five. Your initial state is going to be
- 11:48:50room number two. Your next state is
- 11:48:52going to be room number three. All
- 11:48:53right, so you're moving from two to
- 11:48:55three and you're getting a reward of
- 11:48:56zero. Remember that. Now, from state
- 11:48:58three, you'll either be going to state
- 11:49:00two or you can go to state one or state
- 11:49:03four. All right, if you choose four, you
- 11:49:06can uh directly go to five and here
- 11:49:08you'll get a reward of 100. If you
- 11:49:10choose one, again, you'll go directly to
- 11:49:12five. You'll get a reward of 100. But if
- 11:49:15you go back to two, you'll get a reward
- 11:49:16zero. So guys, let me tell you that the
- 11:49:18agent is going to explore over here,
- 11:49:20okay? He has no idea about the
- 11:49:22environment, so he's not going to go
- 11:49:23from two, three, one, five directly, all
- 11:49:26right? He's not going to know that this
- 11:49:28will lead to the output. He has to
- 11:49:30explore, he has to make mistakes, he has
- 11:49:32to learn, and he has to find out the
- 11:49:34best path to reach room number five. Now
- 11:49:36next is our reward matrix. So guys,
- 11:49:39there are two main matrices in
- 11:49:40Q-learning algorithm. One is a reward
- 11:49:42matrix, and the other is going to be the
- 11:49:44Q matrix or the memory matrix. All
- 11:49:47right, I'll be discussing the memory
- 11:49:48matrix in a while, but for now let's
- 11:49:50look at the reward matrix. So what I'm
- 11:49:53doing here is I'm basically putting all
- 11:49:55the reward values for traversing from
- 11:49:57one node to the other node in a matrix
- 11:49:59known as the reward matrix. Now the
- 11:50:01minus one will basically represent the
- 11:50:03null values. What I'm trying to say is
- 11:50:06there is no connection from zero to
- 11:50:07zero, all right? That's why I'm giving a
- 11:50:09value of minus one. But you might say
- 11:50:11that why is the reward from five to five
- 11:50:14hundred? Now this is because if you go
- 11:50:16from room number five to room number
- 11:50:18five, you're still reaching the end
- 11:50:20goal, all right? That's why there's an a
- 11:50:22reward of 100 over here. Let's look at
- 11:50:24zero {comma} four, all right? There is a
- 11:50:26link from zero to four, but the reward
- 11:50:28is going to be zero because four is not
- 11:50:31the end goal, all right? Four is not
- 11:50:33your goal room or anything, that's why
- 11:50:35your reward is going to be zero. Now the
- 11:50:37reward is 100 only for one {comma} five
- 11:50:39that is if you traverse from one to
- 11:50:41five, for four {comma} five, which is if
- 11:50:44you traverse from four to five, and five
- 11:50:46{comma} five. Now this is because
- 11:50:48through all these three actions you're
- 11:50:49going back to room number five, all
- 11:50:51right? Which is your end goal, that's
- 11:50:53why we have a reward of 100 for these
- 11:50:55three actions, all right? So like I
- 11:50:57mentioned earlier, we're going to have
- 11:50:59another matrix known as a Q matrix,
- 11:51:01which will basically represent the
- 11:51:03memory of what the agent has learned
- 11:51:05through experience. Okay, that's the
- 11:51:07only way the agent will actually learn
- 11:51:09further. If the agent forgets everything
- 11:51:11that he's learned, then there's no point
- 11:51:13because he'll have to redo everything
- 11:51:14from scratch and again he'll forget
- 11:51:16everything. That's why we have a matrix
- 11:51:18known as a Q matrix, which represents
- 11:51:20the memory of the agent. Okay, so if a
- 11:51:22agent has traveled from room number two
- 11:51:24to three, he's going to remember the
- 11:51:26reward. Okay, that reward is going to be
- 11:51:28stored in the Q matrix. And the rows of
- 11:51:31the Q matrix will represent the current
- 11:51:32state and the columns will represent the
- 11:51:35possible actions which lead to the next
- 11:51:37state. Now, the formula to calculate the
- 11:51:39Q matrix is the following. You have Q
- 11:51:42state {comma} action is equals to R
- 11:51:44state {comma} action. Okay, let's say
- 11:51:46the state is your S state number one and
- 11:51:48you're moving to state number two. Okay,
- 11:51:50so Q 1 {comma} 2 is going to be so on. R
- 11:51:53state {comma} action will represent the
- 11:51:55reward of 1 {comma} 2. Okay, let's try
- 11:51:57to understand the reward. So, the reward
- 11:52:00of 1 {comma} 2 is minus 1. Okay, so here
- 11:52:03you'll get a value of minus 1. And then
- 11:52:05you'll have plus the gamma parameter
- 11:52:08into the maximum of the next state and
- 11:52:10the next possible action that you can
- 11:52:12take. Okay, now what is a gamma
- 11:52:14parameter? So, the gamma parameter has a
- 11:52:17range of 0 to 1. Okay, so it can be
- 11:52:19between the value of 0 and 1. Now, if
- 11:52:22gamma is closer to 0, then the agent
- 11:52:24will tend to consider only immediate
- 11:52:26rewards. But if the gamma parameter is
- 11:52:29close to 1, the agent will consider
- 11:52:31future rewards with greater weight. Now,
- 11:52:33I don't know if this reminds you of
- 11:52:35something, but earlier in the session, I
- 11:52:37discussed two important concepts of
- 11:52:39reinforcement learning, which was
- 11:52:41exploration and exploitation trade-off.
- 11:52:44Now, if you're exploring, then the gamma
- 11:52:46parameter is going to be closer to 1,
- 11:52:48but if you're exploiting, then the gamma
- 11:52:50parameter is going to be closer to 0.
- 11:52:53It's better if the gamma parameter is
- 11:52:54closer to 1 because it means that you're
- 11:52:56exploring the entire environment and
- 11:52:59you're trying to get future rewards.
- 11:53:00You're going to get greater weightage
- 11:53:02and more rewards.
- 11:53:03So guys, to sum up the entire thing,
- 11:53:05let's look at how the Q-learning
- 11:53:07algorithm will solve the problem. So you
- 11:53:09begin by setting the gamma parameter and
- 11:53:12the environment rewards in reward matrix
- 11:53:14R. We already did that. After that, you
- 11:53:17set the matrix Q to zero because
- 11:53:19initially the agent will start with no
- 11:53:21knowledge of the environment. As the
- 11:53:24agent explores the environment, the Q
- 11:53:26matrix will start filling up. After
- 11:53:28that, the next step will be select a
- 11:53:30random initial state. Like I mentioned
- 11:53:33earlier, initially you'll randomly
- 11:53:34select any state because the agent has
- 11:53:37no idea about the environment. So you'll
- 11:53:39randomly select a state at step number
- 11:53:41three. Then step number four is you set
- 11:53:44the initial state as current state. Step
- 11:53:47number five, select one among all
- 11:53:49possible actions for the current state.
- 11:53:51Okay, so if you've chosen the current
- 11:53:52state as one, let's say that all the
- 11:53:55possible states that you can traverse to
- 11:53:57from one is two, three, and so on. So
- 11:54:00these are all the possible actions for
- 11:54:02the current state. Now use the possible
- 11:54:04actions and consider going to the next
- 11:54:06state. Once you know what are the
- 11:54:08possible actions from the current state,
- 11:54:10you're going to go to one of those
- 11:54:12possible actions and then you move to
- 11:54:14the next state. After that, get maximum
- 11:54:16Q value for this next state based on all
- 11:54:19possible actions. If you get the maximum
- 11:54:21Q value, it means that you've chosen an
- 11:54:24optimal policy in order to reach your
- 11:54:26end state. Finally, you'll compute the Q
- 11:54:29value by using the formula that we
- 11:54:30discussed earlier. And the last step is
- 11:54:33you have to repeat all of these states
- 11:54:35until your current state is equal to
- 11:54:37your goal state. And our goal state is
- 11:54:39room number five. Now guys, this is a
- 11:54:41very logical solution because you can
- 11:54:43easily understand what is happening over
- 11:54:45here. You have an agent, he has to
- 11:54:47explore through all the states in such a
- 11:54:49way that he reaches the end goal using
- 11:54:51the optimum policy. Okay, that's the end
- 11:54:54goal of Q-learning algorithm and that's
- 11:54:56exactly how you're going to solve this
- 11:54:58problem.
- 11:54:59And
About this transcript
This page contains the full transcript of AI & ML Full Course 2026 | Complete Artificial Intelligence and Machine Learning Tutorial | Edureka by edureka!, generated from the public captions YouTube serves with the video. The transcript has 130,062 words across 20,074 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.