AI And Machine Learning Full Course [FREE] | Learn AI And Machine Learning In 24 Hours | Simplilearn — Transcript
Full transcript
- 0:05Welcome to simply learns YouTube
- 0:07channel. Artificial intelligence and
- 0:09machine learning are transforming the
- 0:11way business operate, make decisions and
- 0:13innovate. From personalized
- 0:14recommendation on streaming platforms
- 0:17and intelligent chat bots to
- 0:19self-driving vehicles and advanced
- 0:20healthcare systems, AI and machine
- 0:22learning are powering some of the most
- 0:25impactful technologies of our time. And
- 0:27AI and machine learning engineering
- 0:29combines programming, mathematics,
- 0:31statistics, data science, machine
- 0:33learning expertise to create intelligent
- 0:35applications that deliver business
- 0:37value. In this complete AI and machine
- 0:39learning engineering course, you will
- 0:41learn everything from the fundamentals
- 0:42of AI and machine learning to advanced
- 0:45concept used in the modern intelligent
- 0:47systems. We will start with the
- 0:48programming and mathematical foundation
- 0:50then gradually move into data analysis,
- 0:52machine learning algorithms, deep
- 0:54learning and generative AI. Throughout
- 0:56the course, you will gain hands-on
- 0:57experience with industry standard tools
- 0:59and frameworks such as Python, NumPy,
- 1:02Pandas, Kikit Learn, TensorFlow,
- 1:04PyTorch, and others. You'll also learn
- 1:06how to collect and prepare data, build
- 1:08predictive models, train neural
- 1:10networks, evaluate model performance,
- 1:12and deploy machine learning solution in
- 1:14real world environments. By the end of
- 1:16this course, you'll have a strong
- 1:17understanding of complete AI and machine
- 1:20learning life cycle and practical skills
- 1:22required to pursue a career as a AI and
- 1:25machine learning engineer. Having said
- 1:27that, let's take a look at today's
- 1:28agenda. We'll start off with module one,
- 1:30which is introduction to artificial
- 1:32intelligence and machine learning.
- 1:33Module two is Python programming for AI
- 1:36and ML. Module three is mathematics,
- 1:39statistics, and probability for machine
- 1:41learning. Module four is data
- 1:43collection, cleaning, and
- 1:44pre-processing. Module five is
- 1:46exploratory data analysis and data
- 1:48visualization. Module six is machine
- 1:51learning fundamentals. Module seven is
- 1:53supervised learning algorithms. Module
- 1:55eight is unsupervised learning
- 1:56algorithms. Module 9 is model evaluation
- 1:59and feature engineering. Module 10 is
- 2:02deep learning and neural networks.
- 2:04Module 11 is natural language
- 2:05processing. Module 12 is computer vision
- 2:08fundamentals. Module 13 is generative
- 2:11AI, LLMS and AI agents. Module 14 is
- 2:15envelopes, model deployment and AI
- 2:17engineering workflows. Module 15 is real
- 2:20world AI projects. Module 16 is
- 2:23interview question and answers. Hope I
- 2:25made myself clear with that agenda.
- 2:27That's it. If these are the type of
- 2:28videos you would like to watch, then hit
- 2:29that subscribe button with the bell icon
- 2:31to get notified whenever we host. Also,
- 2:33just so that you know, if you want to
- 2:34upskill yourself, master generative AI
- 2:36and land your dream job or even grow in
- 2:38your career, then you must explore
- 2:40Simply Learn's cohort of various
- 2:42generative AI training and professional
- 2:44certification programs. Simply learn
- 2:46offers a variety of masters
- 2:47certification and post-graduate programs
- 2:48in collaboration with some of the
- 2:50world's leading universities. Through
- 2:51our courses, you will gain knowledge
- 2:53along with work ready expertise in
- 2:55skills like Python, Aentic AI, AI
- 2:57automation systems, LLMs, and over a
- 3:00dozen others. And that's not all. You
- 3:02will also get the opportunity to work on
- 3:03multiple projects led by industry
- 3:05experts working on top tier
- 3:07service-based and product companies.
- 3:09After completing these courses,
- 3:10thousands of learners have transition
- 3:12into an AI and machine learning role as
- 3:14a fresher or moved on to your higher
- 3:16paying job and profile. If you're
- 3:17passionate about making your career in
- 3:19this field, then make sure to check out
- 3:20the link in the pinned comments and in
- 3:22the description box to find an AI and
- 3:25machine learning program that fits your
- 3:27experience and areas of interest. So
- 3:29let's get started with our AI and
- 3:31machine learning engineer full course
- 3:33with a small quiz. What is NLP? Is it
- 3:36network layer processing, natural
- 3:38language processing, neural learning
- 3:40platform, or is it native language
- 3:42programming? Please let us know your
- 3:44answers in the comment section below.
- 3:45Now over to our training experts.
- 3:48>> Once upon a time in the quiet town of
- 3:50Newite, there lived a curious teenager
- 3:53named Arya. She wasn't like most kids in
- 3:56her school. While others were busy with
- 3:58sports or music, Arya was fascinated by
- 4:01machines, especially the idea of making
- 4:03machines think like humans. Her
- 4:05curiosity began one evening when she
- 4:07asked her grandfather, who used to be a
- 4:09computer engineer, "Can machines ever
- 4:11think?" Her grandfather smiled and said,
- 4:14"That's what artificial intelligence is
- 4:16all about." Arya's eyes lit up.
- 4:18Artificial intelligence? What's that? So
- 4:21he began to tell her a story, not a
- 4:24fairy tale, but a real story about the
- 4:25science and ideas behind machines that
- 4:28learn, decide, and sometimes even
- 4:31surprise their creators. Artificial
- 4:34intelligence, or AI, is the science of
- 4:37making machines that can do things that
- 4:39normally require human intelligence.
- 4:42This includes tasks like recognizing
- 4:44faces, understanding speech, making
- 4:47decisions, and even playing games. But
- 4:50AI isn't magic. It's built through
- 4:53programming, mathematics, and data. Arya
- 4:56imagined a robot that could talk like a
- 4:58human and help with homework. Her
- 5:00grandfather nodded. That's one kind of
- 5:02AI, but there are many types. He
- 5:05explained that AI isn't just about
- 5:07robots. In fact, most AI systems are
- 5:10just computer programs running inside
- 5:12machines we already use, like phones,
- 5:15laptops, or even refrigerators. Her
- 5:17grandfather told her that AI comes in
- 5:19two main types, narrow AI and general
- 5:22AI. Narrow AI is the kind we see today.
- 5:25It's designed to do one specific task.
- 5:28For example, the AI in a smartphone that
- 5:30unlocks the screen by recognizing your
- 5:32face is only good at that one job. It
- 5:35can't cook or write a story. General AI,
- 5:38on the other hand, would be as smart as
- 5:41a human, able to learn anything and do
- 5:44many tasks. But this type of AI doesn't
- 5:47exist yet. It's more of a dream for now.
- 5:50Arya asked, "How do these machines
- 5:52learn? That's where machine learning
- 5:54comes in," her grandfather replied.
- 5:57Machine learning is a type of AI that
- 5:59learns from data instead of being told
- 6:01what to do step by step. "Imagine
- 6:04teaching a dog to sit. You show it how,
- 6:07give it treats, and repeat. Over time,
- 6:10the dog learns. Machine learning works
- 6:13the same way. You feed it data and it
- 6:15finds patterns. For example, if you want
- 6:18a computer to recognize pictures of
- 6:19cats, you show it thousands of cat
- 6:22pictures. It starts to see what cats
- 6:24usually look like. Furry whiskers,
- 6:27pointy ears. Over time, it learns to
- 6:31tell a cat apart from a dog or a chair.
- 6:34The program that does this learning is
- 6:36called a model. A model is like a brain
- 6:39built by the computer using the data it
- 6:41was given. The more data it gets, the
- 6:43better it learns. But how does the
- 6:45computer know what a cat is? Arya asked.
- 6:48Her grandfather said, "That's thanks to
- 6:51something called a neural network. It's
- 6:53a method used in machine learning that's
- 6:55inspired by how our brains work. A
- 6:58neural network is made up of layers of
- 7:00tiny parts called neurons. These are not
- 7:03real brain cells, but math functions.
- 7:06Each neuron takes in numbers, does some
- 7:09math, and passes the result to the next
- 7:11layer of neurons. Imagine passing a note
- 7:14through a group of friends, and each one
- 7:17adds or changes a word before giving it
- 7:19to the next. By the end, the note may
- 7:21have transformed in a useful way. That's
- 7:24what a neural network does to data. It
- 7:28turns it into something meaningful, like
- 7:30recognizing a cat in a picture. The more
- 7:33layers a network has, the more complex
- 7:36patterns it can understand. When a
- 7:39network has many layers, it's called
- 7:41deep learning. To get a neural network
- 7:43to work, it needs to be trained.
- 7:46Training is the process where the model
- 7:48is shown lots of examples so it can
- 7:50learn. Training involves giving the
- 7:52model data and letting it guess
- 7:54something like whether a picture has a
- 7:56cat. At first, it guessed badly, but
- 7:59then it compares its guess to the
- 8:01correct answer. If it's wrong, it
- 8:03adjusts itself using a method called
- 8:06back propagation. Back propagation is
- 8:08like checking your math homework. If the
- 8:11answer is wrong, you go back, find where
- 8:14you messed up, and fix it. In AI, this
- 8:17helps the model improve step by step.
- 8:20This cycle of guessing, checking, and
- 8:22adjusting is repeated many times. The
- 8:24model slowly gets better at the task.
- 8:27Can AI make mistakes? Arya asked. Oh
- 8:31yes, her grandfather said AI is smart in
- 8:34some ways but not perfect. AI only
- 8:36learns from the data we give it. If the
- 8:39data is bad, the AI will be bad. This is
- 8:42called bias. For example, if a face
- 8:45recognition system is trained mostly on
- 8:47photos of light-kinned people, it might
- 8:49not work well on darkerkinned people.
- 8:52Also, AI doesn't really understand the
- 8:54world. It only sees patterns in numbers.
- 8:58It doesn't know what a cat feels like or
- 9:00why we love them. That's why AI can
- 9:02sometimes be fooled by simple tricks
- 9:05like weird images that a human would
- 9:07never mistake for a cat. AI is
- 9:09everywhere, her grandfather explained.
- 9:12It helps recommend videos on YouTube,
- 9:14powers voice assistants like Siri or
- 9:16Alexa, drives some cars, and even helps
- 9:20doctors find diseases and scans. But not
- 9:23all AI is harmless. It can be used for
- 9:26spying, spreading fake news, or making
- 9:29decisions that affect people's lives,
- 9:31like who gets a loan or a job? That's
- 9:34why it's important for people to
- 9:35understand how AI works so they can ask
- 9:38good questions and build it responsibly.
- 9:41Arya asked, "Will AI take over the
- 9:44world?" Her grandfather laughed. Not
- 9:46like in the movies, but it will change
- 9:48the world. The future of AI depends on
- 9:51how people choose to use it. It can help
- 9:53solve big problems like climate change
- 9:55or disease. But it also needs rules and
- 9:58careful thinking. Just like fire or
- 10:00electricity, AI is a tool, a powerful
- 10:02one. If used wisely, it can do great
- 10:05good. Arya sat back, her mind buzzing.
- 10:08She had started the day wondering if
- 10:10machines could think. Now she knows that
- 10:12while they don't think like humans, they
- 10:15can do amazing things through learning
- 10:17data, and clever programming. She smiled
- 10:20and said, "Maybe I'll build an AI
- 10:22someday." Her grandfather smiled, too.
- 10:25Just remember, it's not about making a
- 10:26machine smart. It's about making it
- 10:28useful and fair for everyone. And from
- 10:32that day on, Arya started her journey
- 10:35not just to understand AI, but to shape
- 10:38it with care, creativity, and curiosity.
- 10:41>> We know humans learn from their past
- 10:43experiences, and machines follow
- 10:46instructions given by humans.
- 10:48But what if humans can train the
- 10:51machines to learn from their past data
- 10:52and do what humans can do and much
- 10:54faster? Well, that's called machine
- 10:56learning. But it's a lot more than just
- 10:58learning. It's also about understanding
- 11:00and reasoning. So today we will learn
- 11:02about the basics of machine learning. So
- 11:05that's Paul. He loves listening to new
- 11:08songs.
- 11:10He either likes them or dislikes them.
- 11:12Paul decides this on the basis of the
- 11:14song's tempo, genre, intensity, and the
- 11:19gender of voice. For simplicity, let's
- 11:21just use tempo and intensity for now.
- 11:24So, here tempo is on the x-axis, ranging
- 11:27from relaxed to fast, whereas intensity
- 11:30is on the y-axis, ranging from light to
- 11:33soaring. We see that Paul likes the song
- 11:36with fast tempo and soaring intensity
- 11:40while he dislikes the song with relaxed
- 11:43tempo and light intensity. So now we
- 11:45know Paul's choices. Let's say Paul
- 11:47listens to a new song. Let's name it as
- 11:49song A. Song A has fast tempo and a
- 11:53soaring intensity. So it lies somewhere
- 11:55here. Looking at the data, can you guess
- 11:58whether Paul will like the song or not?
- 12:00Correct. So Paul likes this song. By
- 12:02looking at Paul's past choices, we were
- 12:05able to classify the unknown song very
- 12:07easily, right? Let's say now Paul
- 12:10listens to a new song. Let's label it as
- 12:12song B. So song B lies somewhere here
- 12:16with medium tempo and medium intensity.
- 12:19Neither relaxed nor fast, neither light
- 12:22nor soaring. Now, can you guess whether
- 12:24Paul likes it or not? Not able to guess
- 12:26whether Paul will like it or dislike it.
- 12:29Are the choices unclear? Correct. We
- 12:31could easily classify song A. But when
- 12:34the choice became complicated as in the
- 12:36case of song B. Yes. And that's where
- 12:39machine learning comes in. Let's see
- 12:40how. In the same example for song B, if
- 12:43we draw a circle around the song B, we
- 12:45see that there are four votes for like
- 12:48whereas one vote for dislike. If we go
- 12:50for the majority votes, we can say that
- 12:53Paul will definitely like the song.
- 12:54That's all. This was a basic machine
- 12:56learning algorithm also. It's called K
- 12:58nearest neighbors. So this is just a
- 13:00small example in one of the many machine
- 13:03learning algorithms quite easy right
- 13:05believe me it is but what happens when
- 13:08the choices become complicated as in the
- 13:11case of song B that's when machine
- 13:13learning comes in it learns the data
- 13:15builds the prediction model and when the
- 13:17new data point comes in it can easily
- 13:19predict for it more the data better the
- 13:22model higher will be the accuracy there
- 13:24are many ways in which the machine
- 13:26learns it could be either supervised
- 13:29learning unsupervised learning or
- 13:31reinforcement learning. Let's first
- 13:32quickly understand supervised learning.
- 13:35Suppose your friend gives you 1 million
- 13:37coins of three different currencies. Say
- 13:391 rupee, 1 and 1 dirham. Each coin has
- 13:43different weights. For example, a coin
- 13:45of 1 rupee weighs 3 g. 1 euro weighs 7 g
- 13:49and 1 dirham weighs 4 g. Your model will
- 13:51predict the currency of the coin. Here
- 13:54your weight becomes the feature of coins
- 13:56while currency becomes their label. When
- 13:58you feed this data to the machine
- 14:00learning model, it learns which feature
- 14:03is associated with which label. For
- 14:05example, it will learn that if a coin is
- 14:07of 3 g, it will be a 1 rupee coin. Let's
- 14:10give a new coin to the machine. On the
- 14:12basis of the weight of the new coin,
- 14:14your model will predict the currency.
- 14:16Hence, supervised learning uses labeled
- 14:19data to train the model. Here, the
- 14:21machine knew the features of the object
- 14:23and also the labels associated with
- 14:25those features. On this note, let's move
- 14:27to unsupervised learning and see the
- 14:29difference. Suppose you have cricket
- 14:31data set of various players with their
- 14:33respective scores and the wickets taken.
- 14:35When we feed this data set to the
- 14:37machine, the machine identifies the
- 14:39pattern of player performance. So, it
- 14:41plots this data with the respective
- 14:43wickets on the x-axis while runs on the
- 14:45y-axis. While looking at the data,
- 14:47you'll clearly see that there are two
- 14:49clusters. The one cluster are the
- 14:51players who scored high runs and took
- 14:53less wickets while the other cluster is
- 14:56of the players who scored less runs but
- 14:58took many wickets. So here we interpret
- 15:00these two clusters as batsmen and
- 15:03bowlers. The important point to note
- 15:05here is that there were no labels of
- 15:07batsmen and bowlers. Hence the learning
- 15:09with unlabeled data is unsupervised
- 15:12learning. So we saw supervised learning
- 15:13where the data was labeled and the
- 15:15unsupervised learning where the data was
- 15:17unlabeled. And then there is
- 15:19reinforcement learning which is a
- 15:21reward-based learning or we can say that
- 15:22it works on the principle of feedback.
- 15:24Here let's say you provide the system
- 15:26with an image of a dog and ask it to
- 15:28identify it. The system identifies it as
- 15:31a cat. So you give a negative feedback
- 15:33to the machine saying that it's a dog's
- 15:35image. The machine will learn from the
- 15:37feedback and finally if it comes across
- 15:39any other image of a dog, it'll be able
- 15:41to classify it correctly. That is
- 15:43reinforcement learning. To generalize
- 15:45machine learning model, let's see a
- 15:47flowchart. Input is given to a machine
- 15:49learning model which then gives the
- 15:50output according to the algorithm
- 15:52applied. If it's right, we take the
- 15:54output as our final result. Else we
- 15:57provide feedback to the training model
- 15:58and ask it to predict until it learns. I
- 16:02hope you've understood supervised and
- 16:03unsupervised learning. So let's have a
- 16:05quick quiz. You have to determine
- 16:07whether the given scenarios uses
- 16:09supervised or unsupervised learning.
- 16:11Simple, right? Scenario one. Facebook
- 16:13recognizes your friend in a picture from
- 16:15an album of tagged photographs.
- 16:19Scenario two, Netflix recommends new
- 16:21movies based on someone's past movie
- 16:23choices.
- 16:25Scenario three, analyzing bank data for
- 16:28suspicious transactions and flagging the
- 16:30fraud transactions. Think wisely and
- 16:32comment below your answers. Moving on,
- 16:34don't you sometimes wonder how is
- 16:37machine learning possible in today's
- 16:38era? Well, that's because today we have
- 16:40humongous data available. Everybody's
- 16:43online either making a transaction or
- 16:46just surfing the internet and that's
- 16:47generating a huge amount of data every
- 16:50minute and that data my friend is the
- 16:52key to analysis. Also, the memory
- 16:54handling capabilities of computers have
- 16:56largely increased which helps them to
- 16:58process such huge amount of data at hand
- 17:01without any delay. And yes, computers
- 17:04now have great computational powers. So
- 17:06there are a lot of applications of
- 17:08machine learning out there. To name a
- 17:10few, machine learning is used in
- 17:12healthcare where diagnostics are
- 17:13predicted for doctor's review. The
- 17:15sentiment analysis that the tech giants
- 17:17are doing on social media is another
- 17:19interesting application of machine
- 17:21learning. Fraud detection in the finance
- 17:23sector and also to predict customer
- 17:25churn in the e-commerce sector. While
- 17:27booking a cab, you must have encountered
- 17:29search pricing often where it says the
- 17:31fair of your trip has been updated.
- 17:33Continue booking. Yes, please. I'm
- 17:35getting late for office. Well, that's an
- 17:38interesting machine learning model which
- 17:40is used by global taxi giant Uber and
- 17:43others where they have differential
- 17:44pricing in real time based on demand,
- 17:47the number of cars available, bad
- 17:49weather, rush hour, etc. So they use the
- 17:51search pricing model to ensure that
- 17:54those who need a cab can get one. Also,
- 17:56it uses predictive modeling to predict
- 17:59where the demand will be high with a
- 18:01goal that drivers can take care of the
- 18:03demand and search pricing can be
- 18:05minimized. Great. Hey Siri, can you
- 18:07remind me to book a cab at 6 p.m. today?
- 18:10>> Okay, I'll remind you.
- 18:11>> Thanks.
- 18:12>> No problem.
- 18:13>> Artificial intelligence, machine
- 18:15learning, and deep learning represent
- 18:16the evolution of computer science
- 18:18towards creating intelligent systems. AI
- 18:21is the broader concept striving to build
- 18:23machines capable of humanlike
- 18:25intelligence. ML is a subset of AI
- 18:27emphasizing algorithms that learn from
- 18:29data to make predictions or decisions.
- 18:32DL in turn is a specialized branch of ML
- 18:35that employs deep neural networks to
- 18:37model complex patterns. Imagine an AI
- 18:39powered voice assistant like Apple Siri.
- 18:42It utilizes ML to understand and respond
- 18:45to user queries, learning from
- 18:47interactions over time. Deep learning
- 18:49comes into play when Siri recognizes
- 18:51speech patterns or interprets natural
- 18:53language using neural networks to
- 18:55process intricate features. The better
- 18:57it becomes at understanding diverse
- 18:59accents or refining responses
- 19:01exemplifying the continuous learning
- 19:03inherent in these technologies. AI seeks
- 19:06to emulate human intelligence. ML
- 19:08harness data for learning and DL employs
- 19:10deep neural networks for intricate task.
- 19:13The integration of these technologies
- 19:15manifest in everyday applications,
- 19:17transforming how we interact with and
- 19:20benefit from intelligent systems. This
- 19:22technology enables voice interaction,
- 19:24allowing the device to play music, set
- 19:26alarms, present audio books, and provide
- 19:28up-to-date information on topics like
- 19:30news, weather, sports, and traffic
- 19:32reports, etc. Let's move forward and see
- 19:35what is machine learning. Machine
- 19:37learning is a subset of artificial
- 19:38intelligence that focuses on developing
- 19:40algorithms and models capable of
- 19:42learning and making predictions or
- 19:45decisions without being explicitly
- 19:47programmed. ML systems leverage data to
- 19:49recognize patterns, adapt and improve
- 19:51their performance over time. There are
- 19:53several types of machine learning.
- 19:55Number one, supervised learning. The
- 19:57algorithm is trained on a label data set
- 19:59where each input is associated with a
- 20:01corresponding output. Number two comes
- 20:03as unsupervised learning. Unsupervised
- 20:05learning deals with unlabelled data to
- 20:07find inherent patterns or structures
- 20:09within the information. And then comes
- 20:11the reinforcement learning. This type
- 20:13involves training agents to make
- 20:15sequences of decisions by interacting
- 20:17with an environment. And then comes
- 20:19semi-supervised learning.
- 20:20Semi-supervised learning combines
- 20:22supervised and unsupervised learning
- 20:24elements typically using a small amount
- 20:26of labelled data and a larger pool of
- 20:28unlabelled data. Let us move forward and
- 20:30see what deep learning is. Deep
- 20:32learning, a branch of machine learning,
- 20:34focuses on algorithms inspired by the
- 20:36human brain structure and functionality.
- 20:38It excels in processing vast amounts of
- 20:40both structured and unstructured data.
- 20:42At the heart of deep learning are
- 20:43artificial neural networks, empowering
- 20:45machines to make decisions. The key
- 20:47distinction between deep learning and
- 20:49machine learning lies in data
- 20:50presentation. Machine learning
- 20:52algorithms typically demand structured
- 20:53data while deep learning networks
- 20:55operate through multiple layers of
- 20:57artificial neural networks allowing them
- 20:59to handle diverse data formats. So let's
- 21:02start with the difference between
- 21:03artificial intelligence, machine
- 21:04learning and deep learning. And this
- 21:07we'll show in a table form. So starting
- 21:09with the definition.
- 21:11So definition of artificial
- 21:13intelligence. So broad field of machine
- 21:16learning or creating machines with
- 21:18intelligent behavior is artificial
- 21:19intelligence. And when we talk about
- 21:21machine learning, it's the subset of AI
- 21:23focusing on algorithms learning from
- 21:25data. And then comes the deep learning
- 21:27that is specialized subset of ML using
- 21:29deep neural networks. And now we'll see
- 21:32the difference with the learning
- 21:33approach between all these three. So in
- 21:35learning approach artificial
- 21:37intelligence can include rulebased
- 21:39systems, expert system and more. And in
- 21:42machine learning, it learns from data
- 21:43patterns without explicit programming.
- 21:46And then comes the deep learning where
- 21:48it learns hierarchical representation
- 21:50using neural networks. And if we talk
- 21:52about scope, it encompasses various
- 21:54techniques beyond learning from data.
- 21:57And in machine learning, it primarily
- 21:58focus on learning patterns from data.
- 22:01And then the deep learning, it
- 22:02specifically utilizes deep neural
- 22:04networks for complex task. And now we'll
- 22:07move to the next difference. And we'll
- 22:09start with an example. So in artificial
- 22:12intelligence, the example is autonomous
- 22:14vehicles, chatboards or expert systems.
- 22:17And for machine learning, it's spam
- 22:18filters, recommendation systems, image
- 22:20recognition. And in deep learning it is
- 22:22image and speech recognition natural
- 22:25language processing. And now we'll see
- 22:27the difference for the data
- 22:29requirements. So it depends on the
- 22:31specific application and problem solving
- 22:32approach. And in machine learning it
- 22:34requires labeled or unlabelled data or
- 22:36training. And for the deep learning it
- 22:38relies on large amounts of labelled data
- 22:40for training deep networks. And now for
- 22:44the complexity artificial intelligence
- 22:46addresses a wide range of task including
- 22:48those beyond ML. And in machine
- 22:50learning, it deals with moderate to
- 22:52complex task depending on algorithms.
- 22:54And for the deep learning, it is well
- 22:56suited for intricate task often
- 22:58requiring substantial computational
- 23:00resources. And now see the flexibility.
- 23:03So for the artificial intelligence, it
- 23:06can be rule- based, evolving and
- 23:08adaptive. And for the machine learning,
- 23:10the flexibility adapts to patterns and
- 23:12the changes in data. And for the deep
- 23:14learning, it adapts to hierarchical
- 23:17representations and diverse data types.
- 23:20And then comes the training process. So
- 23:23in artificial intelligence, training
- 23:25process varies based on specific AI
- 23:27techniques used. And in machine
- 23:28learning, training involves feeding data
- 23:31and adjusting model parameters. And in
- 23:33deep learning, training involves
- 23:35optimizing neural weights and
- 23:37structures. And now we'll talk about the
- 23:39applications between all these three
- 23:40terms that is a IML and deep learning.
- 23:43So for artificial intelligence the
- 23:45applications are robotics, natural
- 23:47language processing, game playing and
- 23:49for machine learning it's predictive
- 23:51analytics, fraud detection and
- 23:53healthcare diagnostic and for the deep
- 23:55learning that is image recognition,
- 23:57speech synthesis and language
- 23:59translation.
- 24:00>> Now you guys must be thinking why should
- 24:02I consider a career in AI? Well AI is
- 24:05not just a passing trend. It's a seismic
- 24:08shift that is reshaping our world and
- 24:11creating new venues for innovation and
- 24:14discovery. Now by embracing a career in
- 24:16AI, you become a part of dynamic field
- 24:19that thrives on solving complex problem,
- 24:22pushing boundaries and making a profound
- 24:25impact on society. The demand for AI
- 24:28professionals is skyrocketing across the
- 24:30industries from healthcare, finance,
- 24:32entertainment, transportation.
- 24:34Organizations are actively seeking
- 24:37talented individuals who can harness the
- 24:39power of AI and drive their business
- 24:42forward. But what skills does it take to
- 24:44become an AI engineer? How can you
- 24:46embark on this thrilling journey? We
- 24:49have the answer to all your questions.
- 24:51Some steps are crucial to master the
- 24:53field of AI and become an AI engineer.
- 24:56Let's go through them real quick. So the
- 24:59first step is to establish a strong
- 25:01foundation in mathematics and
- 25:03programming. Start by gaining a solid
- 25:06understanding of critical mathematical
- 25:08concept such as linear algebra, calculus
- 25:12and probability theory. Additionally, it
- 25:14is crucial to become proficient in
- 25:16programming languages like Python which
- 25:19is commonly used in AI and develop
- 25:22coding skills. Next, you need to pursue
- 25:25a degree in relevant field. Earn
- 25:27bachelor's or master's degree in
- 25:29computer science, data science, AI or a
- 25:32related discipline to acquire a
- 25:34comprehensive understanding of AI
- 25:37principle and techniques and after that
- 25:39you need to acquire knowledge in machine
- 25:42learning and deep learning. Familiarize
- 25:44yourself with ML algorithms, neural
- 25:47network and deep learning frameworks
- 25:50like for example TensorFlow, PyTorch to
- 25:53train and optimize models using real
- 25:56world data sets and afterward engage in
- 25:59practical projects. Gain hands-on
- 26:02experience and demonstrate your skills
- 26:04by working on AI projects. Building a
- 26:08portfolio of projects that showcase your
- 26:10ability to solve AI problems can make a
- 26:13strong impression on potential
- 26:15employers. After that, collaborate and
- 26:18network. This is really important.
- 26:20Engage with AR communities, attend
- 26:23conferences, and participate in online
- 26:25forums to connect with professionals in
- 26:28this field. Collaborating with others
- 26:31can enhance your learning experience and
- 26:33open up new opportunities.
- 26:36Seek internships or entrylevel positions
- 26:39where you can gain practical experience
- 26:42through AI internships or entry-level
- 26:44roles in industry or research
- 26:47institution. Now this will provide
- 26:49valuable exposure and help you further
- 26:51develop your skills. After that
- 26:54continuously learn and adapt. In the
- 26:56fast-paced world of AR, it is very
- 26:59important to stay updated on new
- 27:01developments, explore specialized areas,
- 27:04and embrace emerging technologies and
- 27:06tools. Continual learning and
- 27:08adaptability are essential for pursuing
- 27:10a successful career as an AI engineer.
- 27:14Now that you're familiar with the steps
- 27:16involved in the journey of an AI
- 27:17engineer, let's discuss the essential
- 27:19skills you need to know to become an AI
- 27:22engineer. So, here's a breakdown of the
- 27:24skills needed. First one is having
- 27:26strong programming abilities. This
- 27:29typically refers to expertise in one or
- 27:32more programming languages commonly used
- 27:34in data science and machine learning
- 27:37such as Python or R language. Now,
- 27:40proficiency in programming allows you to
- 27:42write efficient and scalable code for
- 27:45data analysis, modeling and algorithm
- 27:48implementation.
- 27:49Next, you need knowledge of machine
- 27:51learning algorithms. This involves
- 27:53understanding and familiarity with wide
- 27:56range of machine learning algorithms
- 27:58including both supervised and
- 28:00unsupervised techniques. You should be
- 28:03able to select and apply appropriate
- 28:05algorithms for specific problems as well
- 28:08as evaluate and optimize their
- 28:10performance. Next skill is proficiency
- 28:13in statistics and mathematics. Sound
- 28:16knowledge of statistics and mathematics
- 28:18is fundamental for data analysis and
- 28:21machine learning. You should be
- 28:23comfortable with statistical concepts,
- 28:25hypothesis testing, regression analysis,
- 28:28probability theory, linear algebra and
- 28:31calculus.
- 28:32Now after that you have acquired a good
- 28:35amount of knowledge of these skill set,
- 28:37we'll move on to our next skill which is
- 28:39having familiarity with deep learning
- 28:42frameworks. Now deep learning has gained
- 28:44significant popularity in recent years
- 28:47and familiarity with deep learning
- 28:49frameworks like TensorFlow, PyTorch or
- 28:52Keras is valuable. Now these frameworks
- 28:55provide tools and libraries for
- 28:57building, training and deploying deep
- 28:59neural networks for tasks such as image
- 29:02recognition, natural language processing
- 29:04and time series analysis. Next, you need
- 29:08experience with big data technologies.
- 29:11Dealing with large scale data sets
- 29:13requires knowledge of big data
- 29:14technologies such as Apache, Hadoop,
- 29:17Spark or distributed computing
- 29:19frameworks. Understanding how to
- 29:22process, store and analyze data
- 29:24efficiently in distributed environments
- 29:26is very essential. Now after you have
- 29:28gotten experience with big data
- 29:30technologies, now it's the time to move
- 29:32on to our next skill which is having
- 29:34excellent problem solving and analytical
- 29:36skills. Now these skills will enable you
- 29:39to break down complex problems, identify
- 29:42key factors and develop efficient
- 29:44solution.
- 29:46Now you should be able to adapt at
- 29:48critical thinking, troubleshooting and
- 29:50debugging to handle real world
- 29:52challenges in data science and machine
- 29:54learning. So guys, remember to stay
- 29:56updated with the latest advancements in
- 29:59the field and continue learning to stay
- 30:01at the forefront of data science and
- 30:03machine learning. So that's all we had
- 30:06for you in this AI engineer road map. Do
- 30:08you know how AI has become so fast? It's
- 30:11now replacing entire teams in some
- 30:13industries. Yes, it's true. Over 50% of
- 30:17companies are already using AI to
- 30:19automate jobs. AI tools are writing
- 30:21emails, creating content, and even
- 30:24giving job interviews. And while some
- 30:26people are worried AI will take the job,
- 30:28I let you in on a secret. AI is also
- 30:31creating tons of highpaying roles. The
- 30:34catch, you need the right skills to get
- 30:37it. And that starts with learning with
- 30:39the right programming language. Now,
- 30:41I've tested a whole bunch of them.
- 30:44Python, C++, R, Java, you name it. And
- 30:48in this video, I'm breaking down the top
- 30:50five programming languages for AI that
- 30:54you need to know if you want to build a
- 30:56career, land real jobs, and actually
- 30:58stay relevant in the age of AI. We will
- 31:01cover what each language is best at, how
- 31:04to start learning, what kinds of AI jobs
- 31:06they lead to, and yes, how much you can
- 31:09earn with each one. All right, first up,
- 31:12we've got Python. And honestly, this one
- 31:14is the most valuable player of the AI
- 31:17development. Just like the star player
- 31:19in a sports team, Python is the go-to
- 31:22language that everyone relies on when it
- 31:26comes to building AI system. So, why is
- 31:28Python the AI king? Let me break it
- 31:30down. Simplicity and readability. Now,
- 31:34Python is super easy to learn. It's
- 31:36almost like writing in plain English.
- 31:38You don't have to worry about
- 31:40complicated code. If you're just
- 31:42starting out in programming, that is
- 31:44definitely the language you are going to
- 31:47feel most comfortable with. It's got
- 31:49this userfriendly vibe that makes it
- 31:52simple even for people new to coding.
- 31:54Second of all, it has got endless
- 31:56libraries. Now, Python is packed with
- 31:58tools. We call it libraries like
- 32:00TensorFlow, PyTorch and Scikitlearn.
- 32:03Think of these library as pre-made
- 32:06toolkits that make AI development way
- 32:09easier. They save you a lot of time
- 32:11because instead of building everything
- 32:13from scratch, you can use these
- 32:15libraries to quickly train your models
- 32:18and run algorithms. It's like having a
- 32:20shortcut to building AI system. It has
- 32:23also got rapid prototyping. If you need
- 32:26to test your ideas quickly, Python is
- 32:28perfect for that. You can build a model,
- 32:30a simple version of your AI system in no
- 32:33time. So whether you're working on
- 32:35machine learning models or neural
- 32:37networks, fancy word for AI system that
- 32:39learn like the brain, Python let you
- 32:42prototype or build a quick model fast.
- 32:45So I know you must be wondering now what
- 32:47kind of AI jobs can Python land me? It's
- 32:49a great question. With Python, you could
- 32:52land jobs like data scientist, machine
- 32:55learning engineer, or an AI researcher.
- 32:58Now these jobs typically pay between
- 33:00around six lakh to 15 lakh peranom.
- 33:03That's the salary range. But the best
- 33:05part is as you gain more experience and
- 33:07expertise that number will go way
- 33:10higher. So how do you start learning
- 33:12Python? You don't have to break the bank
- 33:15to learn Python. You can get started
- 33:17with free platforms and YouTube
- 33:19channels. Simply learn even offers a
- 33:22free comprehensive course in Python and
- 33:24I'll leave the link for you to check it
- 33:26out. And the best part is Python has got
- 33:29huge community. So if you ever feel
- 33:31stuck, there's always someone out there
- 33:34who's ready to help you. Next, we'll
- 33:36talk about C++. Now C++ isn't as
- 33:39beginner friendly as Python, but it's a
- 33:42beast when it comes to performance heavy
- 33:44applications. If you're working on
- 33:46realtime AI like self-driving cars or
- 33:49high frequency trading algorithms, then
- 33:52C++ is where you want to be. But why did
- 33:55we choose C++ for AI? First of all,
- 33:58because of its speed and efficiency.
- 34:00Now, C++ is all about its speed. It's
- 34:03the language you want when you're
- 34:05working with large data sets or AI
- 34:08applications that need to be super fast.
- 34:11Second of all, it has got lowlevel
- 34:13memory management. Now, C++ gives you
- 34:16full control over memory, which is
- 34:18essential when you're building AI system
- 34:20that require extensive computation and
- 34:23realtime performance. But isn't C++ more
- 34:26complex than Python? Definitely, yes.
- 34:29But if you're diving into AI
- 34:30applications that require high
- 34:32performance, think computer vision or
- 34:34robotics, C++ is unmatched. It's a bit
- 34:38trickier to learn, but if you want to
- 34:40build realtime AI systems, it's worth
- 34:43the effort. Roles like AI software
- 34:45developer or computer vision engineer
- 34:48are your goto with C++. The salary range
- 34:51for these roles is around 8 lakh to 20
- 34:54lakh peranom depending on the project's
- 34:56complexity and your experience. Third on
- 34:59a list is Java. This one's a workhorse
- 35:02in the world of AI. And if you're aiming
- 35:05to work on enterprise level AI projects,
- 35:09then Java is definitely a language you
- 35:11want to know. Now, it's not the first
- 35:13choice for small scale AI projects. But
- 35:16when it comes to big scalable systems,
- 35:19Java is untouchable. So why Java for AI?
- 35:22Because of its scalability. Now, you can
- 35:24think Java as a beast when it comes to
- 35:27handling large scale applications. If
- 35:29you're working on AI system that need to
- 35:32process huge data sets or manage complex
- 35:34computations, then Java can handle it
- 35:37all without breaking a sweat. It's
- 35:39designed to scale which makes it perfect
- 35:41for enterprise level AI projects where
- 35:44big data is involved. Platform
- 35:46independence. One of the best things
- 35:48about Java is its right ones run
- 35:51anywhere feature. It doesn't matter
- 35:53which platform you're using, whether
- 35:54it's Windows, Mac, Linux, Java can run
- 35:58all of it without issue. This is a huge
- 36:01win when you're building AI systems that
- 36:03need to operate across multiple
- 36:05platforms. Mature libraries. Java has
- 36:08been around for decades and because of
- 36:10that, it's packed with reliable
- 36:12libraries for AI. Libraries like Qua,
- 36:15H2O make implementing machine learning
- 36:17models or building AI system a lot
- 36:20smoother. These libraries we tried and
- 36:22tested so you know you're working with
- 36:24solid tools. Let's talk about what jobs
- 36:27can you actually land with Java. Now
- 36:29with Java you're looking at some big
- 36:32roles in the AI world. Think of AI
- 36:34solution architect or AI backend
- 36:37developer. These positions are not just
- 36:39highly respected but it also comes with
- 36:42a solid salary range typically between 7
- 36:45lakh to 18 lakh peranom. And with
- 36:47experience, well, let's just say that
- 36:50number can easily climb higher. Now, you
- 36:52must be thinking, how do I get started
- 36:54with Java? Now, if you're already
- 36:56familiar with object- oriented
- 36:58programming, learning Java will be a
- 37:00breeze. And if you're new to it, don't
- 37:02worry. You can start with some great
- 37:04resources like a YouTube channel or
- 37:07LinkedIn Learning. There are plenty of
- 37:09courses that will teach you how to use
- 37:11Java for AI from the ground up. Now,
- 37:14let's talk about R. This one's for all
- 37:17data science enthusiasts out there. If
- 37:20you're diving into statistical AI and
- 37:22the language built specifically for
- 37:24handling massive data and performing
- 37:26complex statistical analysis, R is your
- 37:29goto. So why R for AI? Because of its
- 37:32statistical power. R is packed with
- 37:35tools for statistical modeling. So if
- 37:37you're working on AI projects that need
- 37:39data analysis before you even start
- 37:41applying machine learning, R makes it a
- 37:44breeze. It's got everything you need for
- 37:47analyzing trends, finding patterns and
- 37:50building strong predictive models. It
- 37:53has also got a feature of its data
- 37:54exploration and visualization. One of
- 37:57the R's strength is its data exploration
- 38:00and visualization capabilities. You can
- 38:03easily plot, chart and analyze your data
- 38:05to uncover insights. This makes art
- 38:07perfect for the datadriven side of AI
- 38:10development where understanding your
- 38:12data is just as important as building
- 38:14the models. But wait, can I still work
- 38:16in AI if I learn R or is it just for
- 38:19data analysis? Absolutely. R is
- 38:22fantastic for AI projects that rely on
- 38:24statistical methods and data analysis.
- 38:27It's actually the language of choice for
- 38:29roles like AI data analyst or
- 38:31quantitative analyst where you'll be
- 38:33building predictive models or analyzing
- 38:35data trends to make decisions. Now these
- 38:38roles are in high demand and the salary
- 38:41range typically falls between 6 lakh to
- 38:4312 lakh peranom but with experience you
- 38:46can definitely push those numbers
- 38:48higher. Now to get started with art, you
- 38:50can find tons of free resources on a
- 38:52plat new kid on the block that's growing
- 38:55fast in the AI space. It's relatively
- 38:58young compared to Python or C++. But
- 39:01trust me, it's making a huge impact. And
- 39:04here's why. Now, Julia was created back
- 39:06in 2012 by a group of researchers who
- 39:09wanted a programming language that could
- 39:11handle the complex calculations required
- 39:14for scientific computing and they nailed
- 39:17it. But why did Julia grew so fast?
- 39:20Well, it's been picking up speed because
- 39:22it combines the performance of C++ with
- 39:25the readability of Python. You get the
- 39:28speed and efficiency that C++ is known
- 39:30for, but with Python's clean and easy to
- 39:33write code, it's like the best of both
- 39:35worlds. Let's talk about why did we
- 39:38choose Julia for AI? Because of its
- 39:40speed and simplicity. Now, Julia's speed
- 39:42is one of the biggest advantages. It's
- 39:44designed for high performance computing.
- 39:46So if you need to run complex AI models
- 39:49or process tons of data, Julia will do
- 39:52it in a fraction of the time it would
- 39:54take in other languages. And the syntax,
- 39:56it's also super easy to read and write.
- 39:58So you're not sacrificing convenience
- 40:00for performance. It has also got the
- 40:02feature of scientific computing. Now
- 40:04Julia is optimized for AI task like deep
- 40:07learning and numerical analysis. You can
- 40:09think AI applications in robotics, data
- 40:12science, and scientific research. Now,
- 40:14if you're working on projects that
- 40:15require heavy computations or advanced
- 40:18AI models, Julia's is your go-to. So,
- 40:21why isn't everyone using Julia yet? It's
- 40:24still growing, but Julia community
- 40:26expanding rapidly, and more libraries
- 40:29and frameworks are being developed every
- 40:31day. It's quickly becoming a top choice
- 40:33for high performance AI, and it
- 40:35continues to evolve. And of course, I
- 40:38expect to be even more popular. So, is
- 40:40Julia better than Python or C++? Now the
- 40:43answer is it depends. Now if you're
- 40:44building scientific AI applications that
- 40:47require high performance, Julia is a
- 40:49fantastic option. It's still growing but
- 40:51the community expands. Julia will
- 40:53quickly become more powerful in the AI
- 40:55space. Julia is perfect for roles like
- 40:57AI developer in the scientific or
- 40:59numerical computing space. Salaries can
- 41:01range from 7 lakh to 15 lakh peranom
- 41:04especially if you're working with
- 41:06advanced AI. So guys there you have it
- 41:08the five best programming languages for
- 41:11AI. So whether you're interested in
- 41:12machine learning, realtime AI or data
- 41:15science, there's language for you. Each
- 41:17of these will help you land AI job you
- 41:20want and give you the tools you need to
- 41:21build powerful AI system. Which one are
- 41:24you going to start with? Drop your
- 41:26thoughts in the comment section below
- 41:27and let's talk about it. And if you
- 41:29found this video helpful, hit that like,
- 41:32share, and subscribe button to get more
- 41:34AI tips and career advice by simply
- 41:36learn. get started with the onboarding
- 41:38and interface including the subscription
- 41:40plan. As you can see here, it is
- 41:42offering us three major plans. Now, now
- 41:47there is a free version of manuals you
- 41:49can use on a day-to-day basis. But make
- 41:51sure to know that everyday credits are
- 41:54six rupees.
- 41:55>> But make sure everyday credits are only
- 41:57300 to limited. But only 300 credits
- 42:00will be assigned to you on a everyday
- 42:02basis. Now, the first plan is $20 per
- 42:05month, which gives you 300 fresh credits
- 42:08every day, 4,000 credits per month,
- 42:11in-depth research for everyday task,
- 42:13professional website for standard
- 42:14outboard, insightful slides for regular
- 42:17content, task scaring, and wide
- 42:19research, early access beta features,
- 42:21and 20 concurrent task, 20 schedule
- 42:24tasks. Now again if you are working in
- 42:26an organization which where you can auto
- 42:28you have to auto too many stuffs you can
- 42:31upgrade to a $40 or $200 plan. Now $20
- 42:35is for a person single usage because
- 42:37it's only 300 fresh credits per day.
- 42:39It's total of 4,000 per month as well.
- 42:43So when it so when it comes to $40 plan
- 42:46you can consider sharing it with two to
- 42:48three people. Again it's 300 credits but
- 42:518,000 credits per month. all the other
- 42:54things plus plus you'll get an addition
- 42:57of 4,000 more credits to work on. Now
- 43:00when it comes to 200 you'll get a
- 43:03firstly you'll get free cloud computing
- 43:05where you don't have to worry about the
- 43:06storage and stuff and here it is 40,000
- 43:09credits per month an organization which
- 43:12uses automation tools a lot more can use
- 43:15this now we are going to start by
- 43:18understanding the manusi interface and
- 43:20the first thing we need to lock out at
- 43:22the hub which is basically your main
- 43:24dashboard. Now before we start giving
- 43:26task to manus AI it is very important to
- 43:29understand how credits works because for
- 43:32many users credits can be a little
- 43:34confusing in the beginning. When you
- 43:36open the dashboard you will notice that
- 43:38man's AI shows two different credit
- 43:40counters. The first one is the daily
- 43:43refresh credits. These are the credits
- 43:45that refresh every day. For example you
- 43:47may see around 300 credits per day. The
- 43:50important thing to remember is that
- 43:51these are use them or lose them credits.
- 43:55That means they reset every 24 hours and
- 43:58if you don't use them, they do not carry
- 44:00forward in the next day. So these are
- 44:02the daily credits which are good for
- 44:04regular task, quick experiments, small
- 44:07research work, testing prompt or even
- 44:09trying out different features inside
- 44:11Manusa. The second credit counter is
- 44:14your monthly pool. This is your main
- 44:17credit balance for the month. For
- 44:18example, if you're on a standard plan,
- 44:21you may need something like 4,000
- 44:23monthly credits. These credits are more
- 44:26useful for larger and more complex task.
- 44:29So, if you ask manus AI to do something
- 44:31longunning like researching a topic
- 44:34deeply, creating a report, browsing
- 44:36multiple sources, analyzing information,
- 44:38or even completing a multi-step
- 44:40workflow, then this monthly pool gives
- 44:42you the main runway to complete those
- 44:45bigger tasks. So just remember this
- 44:47simple difference. Daily credits are for
- 44:50everyday use and reset every 24 hours.
- 44:53Monthly credits are your larger credit
- 44:55pool for bigger tasks throughout this
- 44:57month. Now the next important thing in
- 44:59the dashboard is the active task window.
- 45:02This is where manusci shows the tasks
- 45:04that are currently running and this is
- 45:06one of the most powerful parts of the
- 45:08platform.
- 45:09Unlike a normal chatbot where you can
- 45:11ask one question wait for one answer,
- 45:14Manos AI can work on multiple task at
- 45:16the same time. For example, on this
- 45:18subscription you can run up to 20
- 45:21concurrent task at once. That means
- 45:23manos can work on multiple request in
- 45:26parallel. Maybe one task is researching
- 45:29a topic, another is preparing a
- 45:31document, another is analyzing a website
- 45:33and another is organizing the
- 45:35information. For the free users, the
- 45:37limit is usually lower around five
- 45:40concurrent tasks. But the important
- 45:42thing is not just the number of tasks.
- 45:45The important thing is that these tasks
- 45:47are asynchronous and cloud-based. This
- 45:49means once you start a task, Manus AI
- 45:52continuously working in cloud. You do
- 45:54not have to keep watching the screen the
- 45:56entire time. You can start a task, close
- 45:59the browser, disconnect from the
- 46:01internet, and even come back later. and
- 46:03Manus AI can still continue to work on
- 46:07that task in the background. This is
- 46:09what makes it feel less like a normal AI
- 46:12chat port and more like an AI worker.
- 46:14You're not just asking a question and
- 46:16waiting for the reply. You're assigning
- 46:18work, letting the agent process it, and
- 46:21then checking the results once the task
- 46:23is completed. So before using Minus AI
- 46:26for real workflows, always understand
- 46:28these three things. Your daily credits
- 46:30reset every day. Your monthly credits
- 46:33support bigger and longer tasks and your
- 46:36concurrent task window shows how many
- 46:38jobs Manus AI is currently handling for
- 46:41you. Once you understand this dashboard,
- 46:43it becomes much easier to manage your
- 46:45credits, plan your task properly, and
- 46:47use Manus AI more efficiently. Now that
- 46:50we have understood the dashboard and
- 46:52credits, let's move on to the next
- 46:53important part of Manus AI interface,
- 46:56which is the goal, input, and task
- 46:58planning area. Now this is where you
- 47:00actually start working with manus. In a
- 47:03normal chatbot we usually give a small
- 47:05instructions one by one. But in Manus AI
- 47:07the idea is slightly different. Here you
- 47:10give a highle goal and manus plans the
- 47:13steps needed to complete that goal. So
- 47:15in the main input box let's type a
- 47:17simple goal such as research the top AI
- 47:21tools for content creation and create a
- 47:24comparison report. So let's start.
- 47:26research the top AI tools for content
- 47:30creation and create comparison report.
- 47:35Now here we have assigned a proper goal
- 47:37to Manus AI. Now notice what happens
- 47:40after we enter this prompt. Manos does
- 47:43not directly jump into the final answer.
- 47:45First it create a task plan. This is
- 47:48where you will see a to-do list or
- 47:50step-by-step structure showing how Manus
- 47:52is planning to complete the task. For
- 47:55example, manus may break the goal into
- 47:57steps like understanding the topic,
- 47:59searching the AI content, creation
- 48:01tools, collecting useful information and
- 48:03comparing those tools and finally
- 48:05preparing the report. So here you can
- 48:07see the steps. In simple words, manus is
- 48:09taking one big goal and breaking it into
- 48:12smaller actions. This view is very
- 48:14important because it gives us a chance
- 48:15to review the plan before the agent
- 48:18starts doing heavy work. Before manos
- 48:20begins browsing, opening pages,
- 48:22analyzing sources and consuming more
- 48:25credits, we can quickly check whether
- 48:27the plan looks correct. For example, in
- 48:29this case, we should check is manners
- 48:31searching for the right type of tooth.
- 48:33Is it planning to compare them properly?
- 48:36Is it going to create a final report as
- 48:38we asked? If the plan looks correct, we
- 48:41can continue. But if the plan looks
- 48:42incomplete or slightly wrong, we can
- 48:44stop and adjust the prompt before moving
- 48:46forward. This helps us avoid wasting
- 48:49time and credits. So the key point here
- 48:51is simple. In manus AI, we don't need to
- 48:54write every step manually. We can give
- 48:57one clear goal and manus will create a
- 48:59plan for completing it. But before
- 49:02allowing the task to continue, always
- 49:04review the documentation decomposition.
- 49:07But before allowing the task to
- 49:08continue, always review the
- 49:10decomposition view. This helps you
- 49:13understand how the agent is thinking and
- 49:14whether it is moving in the right
- 49:17direction. So in this example, our goal
- 49:19was to research the top AI tools for
- 49:22content creation and create a comparison
- 49:24report. And Manus turns that single bowl
- 49:27into the structured task plan that can
- 49:29review before execution. This is what
- 49:32makes Manus air different from a regular
- 49:34chatbot. It does not just answer
- 49:36immediately. It plans the work first,
- 49:38shows the direction and then start
- 49:40completing the task. Now that Manus has
- 49:42understood our goal, the created task
- 49:45plan, the next step is execution. It's
- 49:48already executing. This is where Manus
- 49:50AI actually starts working on a task.
- 49:52You can think of this part as a hand in
- 49:54the platform. The goal input is where
- 49:56Manus understands what we want. The
- 49:59planning view is where it decide how to
- 50:01do it and the execution view is where it
- 50:04actually performs the work. Once we
- 50:06approve or continue with the task,
- 50:08manage begins completing the steps one
- 50:11by one. The interesting part is that we
- 50:13can watch this happen in real time. On
- 50:16one side, you will usually see the
- 50:17progress list or task steps. This shows
- 50:20that manus has completed what is
- 50:23currently doing and what is still
- 50:25remaining. The next is that you can see
- 50:28manus actually taking action. For
- 50:30example, if the task is repeat, for
- 50:33example, if the task requires research,
- 50:35you can see the agent opening websites
- 50:36and browsing pages. If the task requires
- 50:39collecting information, it may take
- 50:41screenshots, extract details, or even
- 50:44organize the data. If the task needs a
- 50:46structured output, manuals may update a
- 50:49spreadsheet, write content, run the
- 50:51code, or even build an interactive
- 50:53artifact. So instead of only showing the
- 50:56final result, manos shows the workflow
- 50:58while it's happening. This is useful
- 51:00because we can understand how the agent
- 51:02is working not just what answer it gives
- 51:05to the end. Now another important thing
- 51:07is to understand here is the sandbox
- 51:10environment. Manos does not directly
- 51:12operate your local computer. It works
- 51:14inside a cloud and is created for the
- 51:17task. Inside this sandbox, miners can
- 51:20browse websites, collect information,
- 51:21test the ideas, run code, fill forms and
- 51:24build outputs without affecting your
- 51:26personal system. For example, if we ask
- 51:28miners to research AI tools and prepare
- 51:30a vision report, it can browse different
- 51:32website, collect the required details,
- 51:34organize them and then create the final
- 51:36report inside this workspace. And for
- 51:39more advanced task, the sandbox can also
- 51:42help maners create things like websites,
- 51:44slide decks, spreadsheets, dashboards,
- 51:46and other interactive files. This is one
- 51:49of the major reasons maners feels
- 51:51different from the normal chatbot. A
- 51:53regular chatbot mostly gives a text
- 51:55responses. But maners can actually
- 51:58perform actions inside a controlled
- 52:00environment. So while the task is
- 52:02running, we should keep an eye on two
- 52:04things. First the progress list to
- 52:06understand which step manus is working
- 52:09on. Second is realtime action view to
- 52:12see what the agent is actually doing.
- 52:14This makes the whole process more
- 52:16transparent. You're not blindly waiting
- 52:18for the final output. You can see the
- 52:20agent browsing, checking information,
- 52:22organizing the data and building results
- 52:25step by step. So in simple terms, the
- 52:28execution view shows manus in action.
- 52:31The sidebyside workflow helps us track
- 52:33the task in real time and the sandbox
- 52:36environment gives Manus a safe cloud
- 52:38workspace where it can browse, run code,
- 52:41collect data and create useful outputs.
- 52:43This is the part where Manus moves from
- 52:45planning the work to actually doing the
- 52:47work. So we'll get back to this task
- 52:50once it is completed. Let's start with a
- 52:52new task. Now that we have seen how
- 52:54Manus works inside the browser, let's
- 52:56look at how can Manus be on normal web
- 53:00interface. Now here you can even connect
- 53:02a different apps such as Gmail, browser,
- 53:05meta and you can add other connectors as
- 53:08well if you're planning to automate any
- 53:10kind of workflow. Now here when you come
- 53:12to the desktop side you will have a
- 53:14mobile app as well as the desktop app as
- 53:16well. Now if you come to settings you
- 53:18may find an option called integration.
- 53:20This is where you can actually connect
- 53:23manos with platforms like slack,
- 53:25telegram or even line. So as you can see
- 53:28here there are connectors. This is
- 53:30useful because it allows you to interact
- 53:32with manus through a messaging apps you
- 53:34already use. For example, instead of
- 53:37opening the browser every single time,
- 53:39you can just delegate a task, check the
- 53:42progress or monitor updates from a
- 53:45messaging app. So if you're working with
- 53:46a team, Slack can be useful. If you want
- 53:49quick mobile access, WhatsApp, Telegram
- 53:52or Lion can make it easier to stay
- 53:54connected with the agent. Part two manus
- 53:57AI. Let's continue. The main benefit is
- 54:00remote control. You can start monitoring
- 54:02task even when there is no sitting in
- 54:05front of the main system. Now the next
- 54:07advanced feature is the desktop my
- 54:09computer feature. You can download the
- 54:12computer version here in the desktop
- 54:14app. This is available when you have
- 54:16Manus desktop app installed in your Mac
- 54:19or a PC. Here Manus can request access
- 54:22for your local machine for specific
- 54:23action. For example, it may need you to
- 54:26read a local file, open a folder or run
- 54:28a terminal command. But the important
- 54:31thing is to notice that manus does not
- 54:33get a fully access automatically. There
- 54:36are permissions grades. There is a
- 54:38permission gate when manus wants to
- 54:40perform an action on your computer. You
- 54:43will see prompts like allow once or
- 54:45allow always. From a safety point of
- 54:48view, allow once means you are giving
- 54:50permission only for that specific
- 54:52action. Always allow means you are
- 54:55allowing that type of action more
- 54:56regularly depending on the setup. So
- 54:59while showing this, this is especially
- 55:01useful when you want manos to work with
- 55:04files on systems, run scripts or even
- 55:06help with local development tasks. Now
- 55:08the third advanced area is the web app
- 55:11builder. This is where manage becomes
- 55:13even more powerful. In the web app
- 55:15builder interface, you can see manage
- 55:17generating a live interactive web
- 55:19application. This is not just writing a
- 55:21text or giving code snippets. It can
- 55:23actually build pages, connect the
- 55:25databases, structure the app and prepare
- 55:27it while working with the project. For
- 55:29example, if we ask manus to create a
- 55:31simple landing page or a small web app,
- 55:34it can generate a layout and add
- 55:35interactive sections, connect the
- 55:37required backend logic, and even support
- 55:40things like database setup and SEO
- 55:42optimization. The best part here is that
- 55:45you can watch the agent work step by
- 55:47step. You can see it creating files,
- 55:50updating the design, testing the pages
- 55:52and even improving the final output. So
- 55:54this part is useful for users who want
- 55:57to build something practical like
- 55:58websites, dashboard, internal tool,
- 56:00product page or even prototype without
- 56:02manually writing everything line of code
- 56:05from scratch. To summarize this section,
- 56:08so now let's test the same logic. Now
- 56:10let's ask minus AI to create a web
- 56:13landing page for a skincare brand. So
- 56:16create a brand. Now to summarize this
- 56:19section, the browser is the main place
- 56:21where you use manus AI which is this.
- 56:24The messaging integration help you
- 56:26delegate and monitor task remotely. The
- 56:28desktop app gives you manus control
- 56:30access to your local machine with
- 56:32permission prompts. And the web app
- 56:34builder helps manus create live
- 56:36interactive web project. So this is what
- 56:39takes manus from being just a web- based
- 56:42AI agent to something that can connect
- 56:44with your communication tool, your
- 56:46computer and your real project works.
- 56:48Now as you can see there are approaches
- 56:50here. This is the code for the entire
- 56:53web page. Let it generate. I'll show you
- 56:55the output since this is just running in
- 56:57the first step. There are more three
- 56:59steps involved in this. So we'll get
- 57:01back to this once this is done. Now we
- 57:03are going to see where Manus AI becomes
- 57:06really powerful which is deep research
- 57:08and data. The main idea here is very
- 57:11simple. Manus AI is not just a chatboard
- 57:13that gives one quick answer. It can work
- 57:16more like an autonomous research worker.
- 57:18That means you can give it a goal and it
- 57:21can plan the task, browse multiple
- 57:22sources, collect the information, cross
- 57:24the check details and organize the
- 57:26findings and finally create a proper
- 57:28output. So instead of manually opening
- 57:3020 tabs and copying the nodes, checking
- 57:32the resources and building the report
- 57:34yourself can handle a larger part of
- 57:36that workflow for you. Let's start with
- 57:39a we'll just use a practical prompting
- 57:42as of now. Now for this demo, you can
- 57:44just type in research the top CRM tools
- 57:47for small business and create a
- 57:49comparison report with pricing, key
- 57:51features, pros, cons, best use cases and
- 57:54source link. So I've given the exact
- 57:56same prompting. Now once we enter this
- 57:59goal, manus first creates a plan. This
- 58:01is important because the task is not
- 58:03just asking for a simple answer. We are
- 58:05asking manage to research multiple CRM
- 58:08tools, compare them and prepare a
- 58:09structured report. Once the task starts,
- 58:12notice how manners does not depend only
- 58:14on one search result. It begins visiting
- 58:17different websites and checking product
- 58:19pages, pricing pages, review platform,
- 58:22blogs, and others available sources.
- 58:24This is what we call multi-source
- 58:27research. For example, if MinusAI is
- 58:29researching CRM tools, it may check
- 58:31official websites for pricing, review
- 58:33platforms for user feedback and
- 58:35comparison articles for feature level
- 58:38difference. The important thing here is
- 58:40that Manos is not just collecting random
- 58:42information. It is trying to cross
- 58:44interface the details. So if one website
- 58:46mentions a price, Manos can compare it
- 58:49with the official pricing page. If one
- 58:52source mentions a feature, it can check
- 58:53whether the same feature is also listed
- 58:55on the product website. This helps
- 58:57improve the quality of the research.
- 59:00Now, while the agent is working, keep
- 59:02your attention on realtime interaction
- 59:04view. On one side, you can see the task
- 59:07progress. On the other side, you see
- 59:08minus browsing websites, opening pages,
- 59:11taking screenshots, reading the
- 59:12information, and updating its findings.
- 59:14This makes the process more transparent.
- 59:17You're not blind. You're not blindly
- 59:19waiting for final answer. So you are
- 59:21actually seeing how the agent is
- 59:23collecting and organizing the
- 59:24information. Another important thing is
- 59:26to notice how manage handles small
- 59:28problems during the search. Sometimes a
- 59:30page may not open. Sometimes a link may
- 59:33be broken. Sometimes a website may be
- 59:35JavaScript heavy and difficult to read.
- 59:38In manual workflow we would have stopped
- 59:40and find another source assets. But
- 59:43maners has planning layer that can
- 59:45create recovery steps. So if one source
- 59:48does not work, it can try another
- 59:50source. search again or adjust the path
- 59:52without needing constant human help.
- 59:54This is why manners is useful for
- 59:56research heavy tasks. At the end, the
- 59:58output should not just be a paragraph
- 1:00:00summary. A good result should be a
- 1:00:03structured artifact like a comparison
- 1:00:05table or a full research report. For the
- 1:00:08CRM example, the final output can
- 1:00:10include tools, names, pricing, key
- 1:00:12features, pros, cons, best use cases,
- 1:00:14and source links, which we'll check back
- 1:00:16in a few minutes. If you can move beyond
- 1:00:19one short answers and prefer a full
- 1:00:22research workflow across multiple
- 1:00:24sources, manusi is the tool. Now let's
- 1:00:27move on to which is wide search. This is
- 1:00:30more advanced credit intensive feature.
- 1:00:33So we'll get back to all the three in a
- 1:00:35minute. We'll get back to all the three
- 1:00:38outputs and I explain what was the exact
- 1:00:40steps required. So coming back to wide
- 1:00:43research. In normal research, the agent
- 1:00:45may explore sources step by step. But in
- 1:00:47wide research, the idea is very
- 1:00:49different. Wide research is designed for
- 1:00:52scaling. Instead of checking a few
- 1:00:54sources one after the other, it can
- 1:00:56explore many sources in parallel. Manus
- 1:00:59described wide research as using
- 1:01:01parallel multi-agent orchestration where
- 1:01:03many agents can work across large
- 1:01:06research space at the same time. So this
- 1:01:08is not meant for basic questions like
- 1:01:10what is CRM or even give me five tools.
- 1:01:14This feature is better for high impact
- 1:01:16research tasks like market analysis,
- 1:01:18competitive research, industry reports,
- 1:01:20investment research, product research,
- 1:01:23or even strategy planning. For example,
- 1:01:25we can use a large version of the same
- 1:01:28CRM topic. Run wide research on the CRM
- 1:01:31software market for smaller businesses.
- 1:01:33Compare major players, pricing, trends,
- 1:01:36AI features, and even customer
- 1:01:37sentiment, market positions, or even a
- 1:01:39growth opportunities. This kind of
- 1:01:41prompt is much broader. Here we're not
- 1:01:44only asking for a tool comparison. We
- 1:01:46are asking miners to understand the
- 1:01:48market from a different angles. It may
- 1:01:50explore companies, websites, review
- 1:01:53sites, market reports, competitors,
- 1:01:55pages, product documentation, user
- 1:01:57discussions, and other public sources.
- 1:02:00Now, before starting wide research,
- 1:02:02always explain the credit part clearly.
- 1:02:05This type of task can consume a lot more
- 1:02:07credits than a normal research would.
- 1:02:09Since wide research explores a large
- 1:02:12number of sources and runs a much
- 1:02:13heavier workflow, it can cost
- 1:02:15significantly more credits. So we should
- 1:02:17use it for important research work, not
- 1:02:20for a simple Q&A. This is important for
- 1:02:22learners. Think of it like hiring a full
- 1:02:25research team for one task. You would
- 1:02:27not only use them for a small
- 1:02:29definition. You would use it when the
- 1:02:31output has real business value. So the
- 1:02:33main takeaway here is use normal
- 1:02:36research for focused task. Use why
- 1:02:38research when you need a large scale
- 1:02:40high depth analysis across many sources.
- 1:02:43Now next we'll move on to it is useful
- 1:02:45because it shows how manuals can move
- 1:02:47from raw data to a finished business
- 1:02:49report. Here for example let's say let's
- 1:02:52upload a CSV file. So here I've taken a
- 1:02:55random data set from Kaggle and I've
- 1:02:57uploaded it. It says loan data set. Now
- 1:03:00let's give it a prompt saying analyze
- 1:03:02this loan data and create a report
- 1:03:04showing revenue trends, top performing
- 1:03:06products etc. So let's just say analyze
- 1:03:09this loan data and create a report
- 1:03:14showing the trends. So mind you I have
- 1:03:17already cleaned this data and executed
- 1:03:20using AI which is in collab but still it
- 1:03:24took me like proper an hour to create
- 1:03:26it. So let's just leave it. Now as you
- 1:03:29can see this is where maners becomes
- 1:03:31different from normal AI tools. It does
- 1:03:33not only look into the file and guess
- 1:03:35the answer. It can work inside a cloud
- 1:03:37sandbox. Inside this sandbox, miners can
- 1:03:40write and execute code such as Python to
- 1:03:43process the data. So if the file
- 1:03:44contains thousands of rows, miners can
- 1:03:46calculate totals, averages, trends,
- 1:03:48category performance, product
- 1:03:50performance, regional performance, and
- 1:03:51other useful metrices. While this is
- 1:03:54happening, show the executional view.
- 1:03:56You may see man is reading the file,
- 1:03:58writing the code, running analysis,
- 1:04:00checking the output and generating
- 1:04:01charts. This is manus is not producing
- 1:04:04text. It is actually performing mini
- 1:04:06data and this is workflow. After
- 1:04:07processing the data, Manus can also
- 1:04:09create visualization. For example, it
- 1:04:12can generate charts showing loan
- 1:04:14prediction data, which category will
- 1:04:17take more loan, etc. Then the final
- 1:04:19step, it can synthesize everything into
- 1:04:21a business report. Now, what does a good
- 1:04:24report include? what the data shows,
- 1:04:26which products are performing well,
- 1:04:28which areas need attention, which trends
- 1:04:30are visible and what actions the
- 1:04:32business should take next. So from one
- 1:04:34uploading of file and one prompt, Manus
- 1:04:37can complete an end to end workflow. It
- 1:04:40can pass the data, run the code, create
- 1:04:42charts, interpret the results and write
- 1:04:44a final report. This is why Minus is
- 1:04:47very useful for business users,
- 1:04:49analytics, marketers, sales teams,
- 1:04:51founders, and students learning data
- 1:04:53analysis. Now let's see how Manis AI can
- 1:04:56work on autonomous research worker. It
- 1:04:59can browse multiple sources, analyze the
- 1:05:00data, run code and prepare structured
- 1:05:02reports. Now in this module we will see
- 1:05:05manus AI as a creator. This is where
- 1:05:07manus move from just giving answers to
- 1:05:09creating finished functional artifacts.
- 1:05:12So instead of only asking manus to
- 1:05:14explain something, we can ask to build
- 1:05:16something. It can create web apps, slide
- 1:05:18text, posters, infographic, visual
- 1:05:20content and also complete project assets
- 1:05:23from a single natural language prompt.
- 1:05:25So let's get started. So here let's give
- 1:05:27manners a single prompt. Now let's ask
- 1:05:30it to create a landing page for AI
- 1:05:32productivity tools for students with
- 1:05:34sections for features, pricing,
- 1:05:35testimonials, FAQs, and call in action.
- 1:05:38Now can you notice what happens here? We
- 1:05:41are not giving miners a full design
- 1:05:43document. We are not writing code. We
- 1:05:45are not explaining very section step by
- 1:05:48step. We're only giving it an idea.
- 1:05:50Manus takes this idea, understands the
- 1:05:52goal, creates a plan, decides the page
- 1:05:55structure, writes the content, designs
- 1:05:56the layout, and starts building the
- 1:05:58page. This is important. Manus is not
- 1:06:01just giving us text response. It is
- 1:06:03creating a clickable portfolio. So, as
- 1:06:06you can see, it already started creating
- 1:06:08the This can be very useful for
- 1:06:10developers, product managers, startup
- 1:06:12founders, marketers, and business teams.
- 1:06:14If someone has an idea and want to
- 1:06:16quickly see how it might look at a
- 1:06:19website, manners can help create the
- 1:06:21first version very quickly. Instead of
- 1:06:23spending hours preparing a wireframe or
- 1:06:25explaining the idea to a designer or a
- 1:06:28developer, we can just use maners to
- 1:06:30create a rough working version. Then we
- 1:06:33can share it with the team, client or
- 1:06:35stakeholder for feedback. Depending on
- 1:06:37the tunnels can also help with more
- 1:06:39advanced parts like databases, payment
- 1:06:42flow, SEO friendly structure and
- 1:06:44deployment related steps. But for
- 1:06:46beginners, the main thing is to
- 1:06:48understand this manual can move from an
- 1:06:51idea to a functional prototype. Also,
- 1:06:54this type of task is more resource
- 1:06:56inensive than a simple chat response.
- 1:06:59Building a web page may consume hundreds
- 1:07:01of credits. Sometimes around 500 to,000
- 1:07:04or even more depending on the
- 1:07:06complexity. So before running a web app
- 1:07:09task, always check the estimated credit
- 1:07:11usage. This is because minus is not only
- 1:07:13writing text, it is planning, coding,
- 1:07:16testing, building and sometimes handling
- 1:07:19deployment steps as well. Now let's move
- 1:07:21on to the next part. Now, now let's ask
- 1:07:24manus to create a slide deck. Now let's
- 1:07:28ask the manus AI to create a text on the
- 1:07:30future of AI agents for business teams.
- 1:07:33Now once we give this prompt, Manus
- 1:07:35starts planning the slide tech. It does
- 1:07:37not randomly create slide. It first
- 1:07:39creates a proper structure. For example,
- 1:07:41it may begin with an introduction, then
- 1:07:43explain what AI agents are, why
- 1:07:45businesses are using them, their
- 1:07:47benefits, use cases, challenges, and
- 1:07:49finally a conclusion. This is what makes
- 1:07:51the output useful. It's not just a set
- 1:07:53of separate slides. It's a structured
- 1:07:56visual story for research heavy topics.
- 1:07:58Miners can also browse the web, collect
- 1:08:01useful information and include cited
- 1:08:03points. This makes it useful for
- 1:08:05business presentation, research decks,
- 1:08:07pitch decks, training models, and
- 1:08:09internal reports. While the task is
- 1:08:11running, look at the interaction view.
- 1:08:14You can see manners creating the
- 1:08:15outline, preparing slide content,
- 1:08:17improving the design, building the final
- 1:08:19deck and once the deck is ready, you can
- 1:08:22usually download in a businessfriendly
- 1:08:24format like Pex. So we can still open it
- 1:08:27in PowerPoint and make final manual
- 1:08:29changes. Now this is very important
- 1:08:31because manus gives us a strong first
- 1:08:33version but we can still fine-tune it
- 1:08:35the slides based on our brand audience
- 1:08:37or even presentation style. Now let's
- 1:08:39move on. Let's come back to this later.
- 1:08:42Let's see what are the outputs for all
- 1:08:44the prompting that we have given. So
- 1:08:46firstly I have asked it to create a
- 1:08:48landing page for a skincare brand. Now
- 1:08:51as you can see there is a skincare brand
- 1:08:53where you can also edit these. So the
- 1:08:56name is given the benefits products
- 1:08:59purifying tensor what is the cost in
- 1:09:02dollars. You can edit the landing page.
- 1:09:05This usually used to take days for an
- 1:09:07UIUX designer to design the entire page.
- 1:09:10is just done with a small prompt. So you
- 1:09:12have products, reviews, shop now and if
- 1:09:15you come here ready to transform your
- 1:09:17skin, the shops, new arrivals, colle
- 1:09:21collection, support, etc. This doesn't
- 1:09:24look like it's just done from a
- 1:09:27prompting. Now if you come to the second
- 1:09:29one, let's we had asked to compare the
- 1:09:33top CRM tools. Let's see what's the
- 1:09:36answer for that. So here the prompt was
- 1:09:39to research the top CRM tools for small
- 1:09:42businesses and create a comparison
- 1:09:44report with pricing, key features, pros,
- 1:09:46cons, best use cases and course link. So
- 1:09:48as you can see let's open this report.
- 1:09:52So here we have a summary where small
- 1:09:54business CRM section is the best
- 1:09:56approach as a trade-off among these easy
- 1:09:59tools. Now is it comparing all the
- 1:10:01things that we have given? The first one
- 1:10:03it's HubSpot sales hub starting price
- 1:10:06key features pros cons best use cases
- 1:10:09and source link all the things are
- 1:10:11present usually if you use a person they
- 1:10:15used to browse through every single
- 1:10:16website they could find and create such
- 1:10:19kind of report now it's done in just a
- 1:10:22small prompt that I've given now let's
- 1:10:24move on to the next one which is loan
- 1:10:26data analysis this is the most useful
- 1:10:28tool for data analyst because we spend
- 1:10:31hours. They spend hours cleaning the
- 1:10:34data, visualizing trends, what graphs is
- 1:10:36suitable for what kind of data,
- 1:10:38normalizing the data and so many other
- 1:10:41steps. Now, if you can just upload a
- 1:10:43file and ask it to create all the
- 1:10:45reports and all the things necessary to
- 1:10:47take a business decision, this will be
- 1:10:50the most useful tool for data analysts.
- 1:10:53Now, let's see the answer for this. As
- 1:10:55you can see the graphs are there. Let me
- 1:10:57just open. You can give a prompt where
- 1:10:59which kind of graph you want, what
- 1:11:01against what graph you want etc. You can
- 1:11:04see the credit history, marital status,
- 1:11:06property area, education,
- 1:11:08self-employment and dependency all
- 1:11:10against approval rate. Now if you come
- 1:11:13here there is a summary as well which is
- 1:11:15a report. Now the summary is that the
- 1:11:18report analyzes 614 do applications
- 1:11:20using the uploaded loan data set
- 1:11:22covering applicants demographic income
- 1:11:24co-licant income etc. The data set shows
- 1:11:27422 applications were approved
- 1:11:30presenting an overall 68% while 192
- 1:11:33applications were rejected representing
- 1:11:35a rejection rate of 31%. Now as you can
- 1:11:38see we have approval and rejected rate
- 1:11:42and the data overview. What are the
- 1:11:45data?
- 1:11:47Now here this is a very small data set
- 1:11:49and I took almost an day to work with
- 1:11:51this data set and create modeling etc.
- 1:11:54This is done within a few minutes and
- 1:11:56this is amazing because it takes a lot
- 1:11:59of time cleaning the data set knowing
- 1:12:01the data how to understand the data.
- 1:12:05This sorts out all the problem. Now
- 1:12:07coming to the next one content creation.
- 1:12:10So here I had asked manusi to research
- 1:12:13the top AI tools for content creation
- 1:12:15and create a company report. So here as
- 1:12:18you can see there is a report that is
- 1:12:20given. Let's preview it. So here the
- 1:12:23heading is there explore tools download
- 1:12:25the report again this is treating as
- 1:12:27like a website that has all the
- 1:12:29information. So here you can see tool
- 1:12:31distribution by category text generation
- 1:12:34tool is like one etc. Pricing tier
- 1:12:37distribution 63.2 to AI tools directly.
- 1:12:41The first thing is chat GPT which is
- 1:12:44probably mostly consumed I think. Next
- 1:12:46is Jasper AI. Then we have Canva AI and
- 1:12:49next Grammarly Ptory Morph AI Descript
- 1:12:54Midjourney
- 1:12:55Surfer SEO Gemini Claude Copy.ai etc. So
- 1:12:59here you can see the price also what is
- 1:13:02best suited for there is a free version
- 1:13:04also. So it's given free version as well
- 1:13:07rating. This is amazing for content
- 1:13:10creation because usually we don't get
- 1:13:12pictures which give the exact direction
- 1:13:14or exact ratio of the exact numbers that
- 1:13:17we found online. It's either we have to
- 1:13:19create from scratch. So this is amazing
- 1:13:22for content like you have ratings, you
- 1:13:25have pricings, you have to compare them,
- 1:13:28select the top tools to compare. Let's
- 1:13:30compare chat chibity and Jasper sorry
- 1:13:33Jasper and chat chibity and also Canva
- 1:13:35AI all three are equally used. Now let's
- 1:13:38deselect them and copy.ai. Now as you
- 1:13:41can see copy.ai is a little bit less on
- 1:13:43ratings. Oh my god this is too good to
- 1:13:47be a tool. This is literally AI to work.
- 1:13:50Let's check out the next one which is
- 1:13:52landing page for AI productivity. Now
- 1:13:55again this also will be a landing page.
- 1:13:58So it's basically like a website. Now as
- 1:14:01you can see we have the heading college
- 1:14:02study flow features pricing testimonials
- 1:14:04FAQs study flow AI is your personal AI
- 1:14:07tutor study planner productive companion
- 1:14:10get instead explanation organize your
- 1:14:12listings there's a free trial watch a
- 1:14:14demo powerful features for the success
- 1:14:16everything you need to excel in your
- 1:14:18studies all in one place etc. And you
- 1:14:21have the pricing as well. This looks
- 1:14:24like a legit platform website that has
- 1:14:29no flaws. There is literally a review
- 1:14:32rating also frequently asked questions
- 1:14:35which is common in most of the websites.
- 1:14:38And then you have the down at 2024 study
- 1:14:41flow AI all rights reserved. Next let's
- 1:14:44see if the slide deck is ready. Now
- 1:14:48let's play the PPT. It's about the
- 1:14:51future of AI agents for business. So as
- 1:14:54you can see first is the heading
- 1:14:57footages. The next one is core ship with
- 1:15:00this AI agents change the unit of work.
- 1:15:02And then you have what and all things
- 1:15:04are changing. Why now the agent stack is
- 1:15:09maturing. AI agents are not just smarter
- 1:15:12chat bots. What are the difference
- 1:15:13between chatbot co-pilot AI agent agent
- 1:15:16portfolio? The new team model in human
- 1:15:18agent collaboration business teams will
- 1:15:21adopt agents by functions scale agents
- 1:15:24required enterprise architecture
- 1:15:26governance adoption. It's a legit PPT to
- 1:15:29explain each and every single step of AI
- 1:15:32agents future. Manus AI is literally
- 1:15:35describing how AI is put to work not
- 1:15:38just give a text response.
- 1:15:40>> All right, so we're ready to start. uh
- 1:15:42we are going to start with this you know
- 1:15:45first course which is going to study the
- 1:15:46basics of Python. Python will be our
- 1:15:49primary focus for the entire program. Um
- 1:15:52we will use co-pilot. So there's there
- 1:15:56will be co-pilot material um later on in
- 1:15:59the program but like in this first
- 1:16:01course we're going to be focused on
- 1:16:03Python and and for most um things we
- 1:16:06will be using Python. Um even when we
- 1:16:08use co-pilot it will produce Python code
- 1:16:11everything we do will be in Python. I
- 1:16:12think one of the things is by you know
- 1:16:14by the end of the program if anything
- 1:16:17else you guys will be in a much better
- 1:16:19position with Python. You'll be better
- 1:16:21Python coders by the end by the end of
- 1:16:23the program. If you don't learn anything
- 1:16:24else you'll get better at Python. I
- 1:16:27promise. Uh because that's you know all
- 1:16:29of our examples all of our demos
- 1:16:32everything we do will be in Python. So
- 1:16:34you'll you'll get better at it. uh for
- 1:16:37sure and we'll have a lot of practice to
- 1:16:39do that. Okay. So this first lesson is
- 1:16:42all about an introduction to what Python
- 1:16:45is. So if you're completely unfamiliar
- 1:16:47with it, totally fine. We will uh get
- 1:16:51you up to speed and talk about the
- 1:16:53fundamentals and how to set everything
- 1:16:55up on your own computer and talk about
- 1:16:57the various ways to um utilize Python.
- 1:17:01that some of it will involve a setup you
- 1:17:04can do on your own computer. Some of it
- 1:17:05will involve some cloud resources um so
- 1:17:09that you don't need to set anything up
- 1:17:11on your computer if you don't want to.
- 1:17:13Um we'll have options there which will
- 1:17:15be nice. So I will show us those and
- 1:17:17walk us through those. But this first
- 1:17:19lesson all about the basics uh and
- 1:17:23getting set up. So, um what's
- 1:17:27interesting is like at the beginning of
- 1:17:28every lesson, we usually have this uh
- 1:17:31kind of um engagement or discussion. Uh
- 1:17:35but you know, we've I kind of already
- 1:17:37asked you guys about this of uh uh if
- 1:17:40you're familiar with programming, if
- 1:17:41you're familiar with Python. Um but one
- 1:17:44thing I want you to think about a little
- 1:17:45bit is that um especially as we go along
- 1:17:48and learn about what Python is is why is
- 1:17:51Python the
- 1:17:53chosen language for AI? So why is it the
- 1:17:57one that everyone uses uh to do AI? And
- 1:18:00I think what you're going to learn is
- 1:18:03that it has a really amazing ecosystem
- 1:18:08that has been around for a long time
- 1:18:10that um supports AI in particular. So,
- 1:18:15Python is the go-to for anything AI,
- 1:18:19data science, machine learning, anything
- 1:18:21in that sort. Uh, because it's been used
- 1:18:25for so long for that and it has such a
- 1:18:28uh community and ecosystem around it.
- 1:18:31That's something we're going to learn.
- 1:18:32It's also really easy to learn and use,
- 1:18:36which makes it nice to to be uh kind of
- 1:18:39an introduction to the field. It doesn't
- 1:18:42take a lot to get started in it.
- 1:18:44because it's so easy to work with. Um, I
- 1:18:47can tell you as someone who's gone
- 1:18:49through that experience, like I studied
- 1:18:51mathematics in college and in graduate
- 1:18:54school and studied like probability and
- 1:18:56statistics, but I was able to teach
- 1:18:59myself Python primarily and use that to
- 1:19:02get into kind of data science and
- 1:19:04machine learning in the industry.
- 1:19:06So, and I think that's a common story is
- 1:19:08people and I've seen that from many
- 1:19:10learners coming from uh different
- 1:19:12backgrounds. Uh they've been able to
- 1:19:14pick up Python pretty easily because
- 1:19:17it's a very easy language to understand
- 1:19:19and and syntax of it and there's so many
- 1:19:22tools within it that make it really easy
- 1:19:24to work with.
- 1:19:26So, um I promise it won't be as uh
- 1:19:31daunting as it may seem even if you're
- 1:19:33coming at it from zero experience. Uh, I
- 1:19:36think you'll find this is the perfect
- 1:19:39way to get into programming and get into
- 1:19:41data science and and AI and machine
- 1:19:44learning because it's so easy to pick up
- 1:19:46and learn and it has such a nice rich
- 1:19:48community ecosystem.
- 1:19:51So, just wanted to mention that.
- 1:19:55Okay. So, some of our objectives for
- 1:19:57this first lesson will be to talk about
- 1:20:00programming languages in general and um
- 1:20:02programming in general. So maybe you
- 1:20:05know more generic than Python just you
- 1:20:08know what are what do general programs
- 1:20:10look like? What are some of the building
- 1:20:12blocks of programs that are important?
- 1:20:15What are some of those uh key principles
- 1:20:17of programming that we will want to
- 1:20:20follow as well? Even if we're doing
- 1:20:21Python for AI purposes.
- 1:20:25Um so just talk about programming in
- 1:20:27general and then kind of zoom in on
- 1:20:29Python as we go along. One of the things
- 1:20:31we'll be interested in doing is just
- 1:20:33getting you guys set up. So talk about
- 1:20:34how we can configure Python for you to
- 1:20:37use on your own machine. Um but also
- 1:20:40have some options that don't require
- 1:20:41installing anything on your own machine.
- 1:20:43Uh which is nice. Um and then as I said,
- 1:20:46we'll kind of zoom in on Python, talk
- 1:20:47about its benefits, uh some of the nice
- 1:20:50features. I've kind of already mentioned
- 1:20:51it. Really big community around it, easy
- 1:20:53to learn. We'll just talk about those
- 1:20:55more in detail. Talk about um why it's
- 1:20:58so popular in the AI world. Um,
- 1:21:02and then we'll get into some very
- 1:21:03fundamental things specific to Python.
- 1:21:05So once we talk about the background,
- 1:21:07get you guys set up, we'll go into uh
- 1:21:12some of the syntax basics, things like
- 1:21:15identifiers, things like indentation,
- 1:21:17comments, um, some of the basics of the
- 1:21:20code that are going to be important for
- 1:21:21you to kind of get started with. Um and
- 1:21:24then talk about some of the basic data
- 1:21:26types that Python offers to manipulate
- 1:21:28and work with data which of course is
- 1:21:30important um when you know as we go
- 1:21:33forward and and do anything with data
- 1:21:35which of course with AI we will be
- 1:21:37interested in doing um but that's these
- 1:21:41are the objectives of just the this
- 1:21:42first lesson. As we go forward we're
- 1:21:45going to learn about many other basic
- 1:21:48topics within Python. So things like how
- 1:21:51to write functions, how to build
- 1:21:54objects, how to manipulate our flow of
- 1:21:57the program with like things like if
- 1:21:59else statements, things like loops.
- 1:22:02We'll learn all about that in kind of
- 1:22:04the next lessons after this one. But
- 1:22:07this is all the content for this lesson.
- 1:22:09I anticipate today
- 1:22:12we will get through all of this today
- 1:22:13and then get into the second lesson
- 1:22:15which will um get into those kind of if
- 1:22:19else and loops. So we'll we'll get we'll
- 1:22:21I'm sure by today we'll get into those.
- 1:22:24All right. Any questions on kind of what
- 1:22:26we're going to learn in this first
- 1:22:27lesson? So mainly trying to get you guys
- 1:22:29set up, give you some background on
- 1:22:30Python and then towards the end of the
- 1:22:32lesson um get into some basics of the
- 1:22:35syntax is kind of the goals I would say.
- 1:22:38Okay. Okay. So when we talk about
- 1:22:39programming um what do we mean by
- 1:22:42programming in general? It's really uh
- 1:22:44synonymous with instruction. So
- 1:22:47programming really means giving or
- 1:22:49writing instructions for a computer to
- 1:22:52perform tasks. Um so these instructions
- 1:22:56we write down in what we call code. But
- 1:23:00those those are just telling the
- 1:23:01computer what to do. And of course the
- 1:23:03computer's not going to do anything
- 1:23:05unless we write down these instructions.
- 1:23:08So these instructions can do really
- 1:23:11powerful things. They can power, you
- 1:23:13know, whole applications, things that we
- 1:23:15use every day like Microsoft Word,
- 1:23:17PowerPoint, Excel, those kind of things.
- 1:23:19Um they can automate tasks. They can um
- 1:23:22power websites. Um they can do AI,
- 1:23:26right? So we can have um things like
- 1:23:28chat GBT and Alexa and Siri, etc., etc.
- 1:23:32Um these are all powered by instructions
- 1:23:35telling the computer what to do.
- 1:23:37One of the things that we will get
- 1:23:38better at as we go along is figuring out
- 1:23:41how to write these instructions in
- 1:23:42Python. Python is going to be the
- 1:23:45language we write those instructions in
- 1:23:48um and and they will be executed by a
- 1:23:52Python um program. But we should think
- 1:23:56of programming in general as just
- 1:23:58instructing the computer what to do just
- 1:24:01at a high level. Right?
- 1:24:03So when we talk about these
- 1:24:07instructions, they have two ways of
- 1:24:10being executed by the the computer. Um
- 1:24:15and roughly these break down into what
- 1:24:17we call interpreted languages and
- 1:24:20compiled languages. So that the code
- 1:24:22that we write which is um representing
- 1:24:25the instructions that we write can be
- 1:24:28executed um in one of these two ways.
- 1:24:32Let me start with the left. So the
- 1:24:33interpreted languages.
- 1:24:35This means that the computer is
- 1:24:38literally executing the the instructions
- 1:24:41line by line by line when we run the
- 1:24:45program. So there is no
- 1:24:49translation of anything. It's just
- 1:24:51literally taking our instructions and
- 1:24:53running it line by line, instruction by
- 1:24:55instruction essentially. Um, now the
- 1:24:59advantage to doing this is that it's uh
- 1:25:03easier to debug because the instructions
- 1:25:06are going to be executed one by one. So
- 1:25:07it can hit an error pretty quick. If
- 1:25:09there's a mistake in one instruction,
- 1:25:12nothing else will run. Um, however, it's
- 1:25:15also slower because we're going to take
- 1:25:18it one instruction at a time. Um, and so
- 1:25:22the the uh this way of running programs
- 1:25:26tends to be slower, but it's also easier
- 1:25:29to work with, which is why we're so
- 1:25:32interested in Python. It's in this
- 1:25:34bucket of what we call interpreted
- 1:25:36languages. So a lot of scripting
- 1:25:38languages find themselves in this bucket
- 1:25:40of being executed one line at a time. No
- 1:25:43translation needed by the machine. It
- 1:25:45just reads our instructions and executes
- 1:25:47it. The thing that does the execution is
- 1:25:50called an interpreter.
- 1:25:52Um, and Python has an interpreter that
- 1:25:56we will get you guys set up with on your
- 1:25:59own machine that can execute Python
- 1:26:01code. So you need an interpreter. The
- 1:26:04interpreter just executes your
- 1:26:05instructions line by line by line. Um,
- 1:26:08so some examples would be like Python.
- 1:26:10That's what we're going to study in this
- 1:26:12um entire program. But there's other
- 1:26:14languages like JavaScript, Ruby,
- 1:26:17um Pearl, many others that are uh
- 1:26:21interpreted. They require an
- 1:26:23interpreter, but they execute line by
- 1:26:24line by line and there's no intermediate
- 1:26:26translation of anything. Um it's kind of
- 1:26:29executed as is. Now, contrast this with
- 1:26:34compiled languages, which are uh kind of
- 1:26:37a different piece. they these these
- 1:26:40instructions have to be translated into
- 1:26:42something the machine can understand in
- 1:26:45order to execute. So there is an
- 1:26:47intermediate step of what we call
- 1:26:50compiling the code um into uh basically
- 1:26:55a translated version of your
- 1:26:57instructions so that the machine can
- 1:26:59execute it. Now there's a trade-off
- 1:27:01there. Doing that can make it more
- 1:27:03difficult to develop and it can take
- 1:27:05longer to debug because you have to go
- 1:27:07through this translation step every
- 1:27:09single time through the compiler.
- 1:27:12But when you run the code because it's
- 1:27:15already been translated into this
- 1:27:16machine format, it's a lot faster. Um,
- 1:27:20so some examples of languages like this
- 1:27:22are C, C++, Java,
- 1:27:25um, Go,
- 1:27:27but uh, we won't really be working with
- 1:27:29those. We'll just be sticking with
- 1:27:31Python. But if you have experience with
- 1:27:32those languages, those you're probably
- 1:27:34familiar with this, you have to compile
- 1:27:36the program first before you can execute
- 1:27:38it. But we are going to be in this
- 1:27:41interpreted world. If you know and it's
- 1:27:43okay like if none of this makes sense,
- 1:27:45that's okay. Just understand that um
- 1:27:48generally interpreted languages are
- 1:27:50going to be more user friendly because
- 1:27:52they're they're easier to execute. They
- 1:27:55don't require as many moving parts as
- 1:27:58what a compiled language would require.
- 1:28:01which is nice for us, right? Nice for
- 1:28:02Python. That's what we're going to be
- 1:28:04interested in working with. Uh kind of
- 1:28:08um yeah, they're kind of rel So, so the
- 1:28:10question is are JavaScript and Java
- 1:28:12related? Kind of. Um, JavaScript is kind
- 1:28:16of like the um the the
- 1:28:19scripting version of um some of the same
- 1:28:22concepts we see in Java, but Java is the
- 1:28:25compiled um it it requires a a special
- 1:28:29kind of what's called a Java runtime,
- 1:28:31which is a a compiler to translate the
- 1:28:35Java code into um machine code that the
- 1:28:39Java runtime will execute. JavaScript is
- 1:28:42not like that at all. It can actually be
- 1:28:44ran in a web browser which is um
- 1:28:47JavaScript usually powers a lot of like
- 1:28:49front-end websites are usually powered
- 1:28:52by JavaScript and Java usually powers
- 1:28:54more like backend
- 1:28:57um applications like actual software
- 1:28:59programs are usually would be coded in
- 1:29:01Java. JavaScript is going to be used
- 1:29:04more for like building a website. But,
- 1:29:06you know, I'm not an expert on that
- 1:29:08really, but that's kind of my
- 1:29:11understanding of it. And if anyone is an
- 1:29:13expert on those differences, feel free
- 1:29:15to let us know in the chat. But, uh,
- 1:29:18that's my that's my basic summary of
- 1:29:20that. Okay. So, we have interpreted
- 1:29:24languages. That's where Python falls
- 1:29:25under. So, it just um summarizing that,
- 1:29:29it's going to be easier to work with
- 1:29:30those, which is great for us. That's
- 1:29:32another reason why Python's so easy.
- 1:29:34It's interpreted, meaning that
- 1:29:36everything executes. We don't need to
- 1:29:37worry about compiling things, which is
- 1:29:40nice. Um, but also in terms of
- 1:29:45programming, there's also uh categories
- 1:29:48of how the instructions are written that
- 1:29:51you can bucket different languages into.
- 1:29:53So for example um some language are are
- 1:29:56more um procedural in nature meaning
- 1:29:59that you write out all the instructions
- 1:30:01exactly kind of line by line by line.
- 1:30:04You don't really organize things at all
- 1:30:06in your instructions.
- 1:30:08Um so some examples would be like C and
- 1:30:10Pascal
- 1:30:11are more like that. Um then on the
- 1:30:15opposite end of the spectrum is kind of
- 1:30:16object-oriented
- 1:30:18in which case you uh build your code and
- 1:30:22organize it around the idea of
- 1:30:24everything being an object. And so some
- 1:30:26uh Python actually falls into this
- 1:30:28category where um uh most things in
- 1:30:31Python are objects and you manipulate
- 1:30:34objects and objects have data to them.
- 1:30:37They have things they can do and
- 1:30:39interact with other objects. Um, so
- 1:30:42think of it just as a way we will
- 1:30:44organize our instructions.
- 1:30:46Python allows us to organize it around
- 1:30:48the concept of an object. We'll learn
- 1:30:50about what that means as we go along,
- 1:30:52but just realizing that some programming
- 1:30:55languages break down along these um kind
- 1:31:00of buckets here. Um, Python is also a
- 1:31:03scripted language, meaning you can write
- 1:31:06out your code in a individual script and
- 1:31:09you can e that you can have an
- 1:31:11interpreter that executes that script.
- 1:31:13Um, so you don't need to organize all
- 1:31:15your code inside of an object. So for
- 1:31:18that reason, Python super flexible.
- 1:31:22That's another reason why it's so nice
- 1:31:24to use. It actually falls into both of
- 1:31:26these buckets on the right, which is
- 1:31:28very convenient. We can have basically
- 1:31:30this means we can have a lot of
- 1:31:32organization or very little organization
- 1:31:34depending on how we want to set it up.
- 1:31:37Yeah, Roberto. So even though there are
- 1:31:39different types so Java is compiled and
- 1:31:42Python is interpreted
- 1:31:45um they are both object-oriented meaning
- 1:31:48so think of the this slide as telling
- 1:31:51you how the instructions are organized.
- 1:31:55So how they are executed is different.
- 1:31:57So, Java requires a a compiler to
- 1:32:00execute things. Python requires an
- 1:32:02interpreter.
- 1:32:05This is more about how the instructions
- 1:32:06are organized. So, Java and Python both
- 1:32:10allow you to organize your code into
- 1:32:11objects.
- 1:32:13Um, but what's nice about Python is it
- 1:32:16also falls under the bucket of
- 1:32:17scripting, meaning that it allows you to
- 1:32:20organize things into scripts, which is
- 1:32:23less organization than it would be in
- 1:32:25into objects. We're actually going to
- 1:32:26learn about objects later on in a future
- 1:32:29lesson, like how to build objects and
- 1:32:31what they mean.
- 1:32:34So yeah, even though they're different,
- 1:32:35they're both object-oriented, which just
- 1:32:37means that you can organize your code
- 1:32:40into objects. Python allows that. So
- 1:32:42does Java. So does C++.
- 1:32:45Uh many many languages allow for um
- 1:32:48organizing your your code into objects.
- 1:32:52So we're going to learn about that.
- 1:32:54It's It's not that one's better. They're
- 1:32:57just um I I would put them at different
- 1:33:00So, let me draw this. I would put them
- 1:33:03at different spectrum, different ends of
- 1:33:05the spectrum on organization.
- 1:33:07So, scripting
- 1:33:10is very loose. Basically, you it's more
- 1:33:14like a an individual um uh set of
- 1:33:18instructions to do one task. you can
- 1:33:21just have and you can have many
- 1:33:22individual scripts to do many small
- 1:33:24tasks. Um, and then on the other end of
- 1:33:27the spectrum, think about it as like
- 1:33:29you've organized your cabinet into many
- 1:33:32folders and many like uh you know many
- 1:33:37pieces of organization that are we would
- 1:33:39call objects. Um so objectoriented
- 1:33:43programming OOP is kind of on the other
- 1:33:46end of the spectrum when it comes to
- 1:33:48like level level
- 1:33:51of organization.
- 1:33:56Does that make sense? So scripting very
- 1:33:58loose. It usually scripting is is um
- 1:34:01reserved for like one task and it's um
- 1:34:03you're just writing out your
- 1:34:04instructions to accomplish that one
- 1:34:06task.
- 1:34:07um which is helpful for like automation
- 1:34:09of things because you're going you're
- 1:34:11usually automating like a single task.
- 1:34:14Um so it's very loose. It's not very
- 1:34:16organized into nothing is organized
- 1:34:17necessarily into objects. Um very loose
- 1:34:20organization. Object-oriented is much
- 1:34:24more structure to it and things being
- 1:34:26put into objects um in order to
- 1:34:28manipulate and work with objects
- 1:34:30throughout the program. Yeah, it's not
- 1:34:33that one's better. I think it's more
- 1:34:35just use case dependent. Um there are
- 1:34:38times where it actually will benefit us
- 1:34:41from using objects. Um and I think the
- 1:34:45thing to pay attention to on this slide
- 1:34:46is that look at where Python falls into.
- 1:34:49It actually falls into both. Meaning
- 1:34:51that we can have things very loose and
- 1:34:54easy to work with because scripting
- 1:34:56usually will be faster and easier to
- 1:34:58just write something to to accomplish
- 1:35:00one task. But we have the flexibility to
- 1:35:03organize our code into objects if we
- 1:35:05want to. which will be better for
- 1:35:08bigger tasks that require more
- 1:35:11organization
- 1:35:13like training a neural network or
- 1:35:16building an LLM.
- 1:35:18Those bigger tasks would benefit from
- 1:35:20organization.
- 1:35:23And then uh finally on this slide um
- 1:35:26there are languages that are built on
- 1:35:27the concept of um their their entire way
- 1:35:31of writing instructions is more in a
- 1:35:33functional way meaning everything is
- 1:35:35based on operating uh functions and
- 1:35:38variables. Um and so there are some
- 1:35:41languages like that has and scholar are
- 1:35:43very popular ones. Um but that is can be
- 1:35:49very difficult to learn. It's it can be
- 1:35:51difficult but very nice in some ways
- 1:35:53because uh it can be very natural to
- 1:35:56think of um you manipulate like giving
- 1:35:59instructions to computer in a functional
- 1:36:01way. Think about it as like applying a
- 1:36:03function to a variable.
- 1:36:05Um that makes sense but writing your all
- 1:36:09of your instructions in that way can be
- 1:36:10kind of difficult to learn. So for that
- 1:36:13reason I think these languages are more
- 1:36:15difficult to learn but they can be very
- 1:36:17powerful. Um and they find themselves
- 1:36:20very useful in like operating on big
- 1:36:23data.
- 1:36:24Um so if you ever heard of like Spark um
- 1:36:27Spark operates with uh Scola for
- 1:36:30instance um but uh we won't really focus
- 1:36:34on functional. It's kind of its own
- 1:36:36paradigm.
- 1:36:38Um but uh again like Python is where our
- 1:36:42focus will be. It allows us to be really
- 1:36:45organized, loosely organized. Nice
- 1:36:48flexibility there.
- 1:36:51So, so far
- 1:36:53based on these two slides, I'm showing
- 1:36:54you that Python is interpreted, which is
- 1:36:57easier and faster to work with. Um, not
- 1:37:00faster to run, but faster to get up and
- 1:37:02running because you don't need to
- 1:37:03compile things. That's nice from our
- 1:37:06perspective.
- 1:37:08And it's also has very good flexibility
- 1:37:11when it comes to organizing our
- 1:37:12instructions, organizing our code can be
- 1:37:14very loose in scripts, could be very
- 1:37:17structured in in objects.
- 1:37:21Okay. Okay. So generally no matter how
- 1:37:23uh no matter what language it is um when
- 1:37:26you process those instructions generally
- 1:37:29things are going to be organized
- 1:37:32even if it's in a script or if it's
- 1:37:33object-oriented
- 1:37:35um you're generally going to have the
- 1:37:38very beginning of the program um kind of
- 1:37:40setting up the input then the middle of
- 1:37:42it really processing that and doing
- 1:37:45something with that. So that's usually
- 1:37:46like the bulk of the logic is in the
- 1:37:49processing phase and then generally
- 1:37:51you're producing some output. So that
- 1:37:53could be like a model prediction, that
- 1:37:56could be um a a graph that you've built
- 1:37:59from your code um whatever that output
- 1:38:02is. But generally it flows this way.
- 1:38:04This is this is makes sense, right? Of
- 1:38:06course there's input, you're
- 1:38:08manipulating that input in some way and
- 1:38:10then you're producing some output. I
- 1:38:11think that all makes sense. That's a
- 1:38:13very logical way to flow.
- 1:38:15Um
- 1:38:17now that's not to say that within this
- 1:38:19processing step there may not be
- 1:38:22um iteration like of course there may
- 1:38:25may be times where we need to as part of
- 1:38:28the processing kind of iterate and do
- 1:38:30multiple passes of processing. Um so the
- 1:38:33processing could be a lot. We could be
- 1:38:35doing a lot. We could be doing a little.
- 1:38:37Just depends on what we're actually
- 1:38:38doing. So, if we're reading in some data
- 1:38:41as the input um and then we're just
- 1:38:44doing some simple um slicing and dicing
- 1:38:47of it, that's some easy processing and
- 1:38:49maybe producing a graph or producing a
- 1:38:51metric, something of that sort, that's
- 1:38:53pretty easy to do. But if we're training
- 1:38:56a neural network or training a model,
- 1:38:59the processing step can take a while and
- 1:39:01it may, you know, be very iterative in
- 1:39:03nature. So it just depends on what we're
- 1:39:06doing and those instructions.
- 1:39:08But no matter what, most of our programs
- 1:39:11will flow in this way kind of input
- 1:39:14processing output. It makes sense. It's
- 1:39:16very logical.
- 1:39:19So what are some principles that we
- 1:39:22should abide by when we're writing our
- 1:39:24code? So this this would really be for
- 1:39:26any language, but of course for Python
- 1:39:28that we are interested in. Um so
- 1:39:31something we're going to be interested
- 1:39:32in doing is um basically avoiding
- 1:39:36repetition where we can. So instead of
- 1:39:39having copy paste everywhere, we will
- 1:39:42generally favor organizing our code to
- 1:39:45some degree. Meaning we will utilize
- 1:39:47functions where it makes sense and
- 1:39:49objects where it makes sense to organize
- 1:39:50things. And also instead of um having
- 1:39:54very repetitive code, we will favor
- 1:39:57using uh loop structures that can
- 1:40:00iterate over um things many times
- 1:40:04instead of us us having to write all
- 1:40:06those out one by one by one. So we're
- 1:40:08going to learn about these tools that we
- 1:40:10have at our disposal, but they will help
- 1:40:12us organize our code, avoid repetition
- 1:40:15all over the place. One of the things we
- 1:40:17want to avoid is having the same code
- 1:40:21repeated all over the place. If if we
- 1:40:24find ourselves doing that, we should
- 1:40:25really put that code into a function or
- 1:40:28maybe into an object so that we can
- 1:40:29reuse it. So, we're really going to
- 1:40:32favor like reusability of things,
- 1:40:36re recycle, reuse, you know. So, we're
- 1:40:39going to learn how to do that, how to
- 1:40:42build functions, how to build objects.
- 1:40:44But that's something we're going to
- 1:40:44favor uh when we're when we're
- 1:40:46programming. It's something you should
- 1:40:47be on the lookout for. If you find
- 1:40:50yourself writing the same code over and
- 1:40:52over just in different spots, um that's
- 1:40:55probably a clue you should organize that
- 1:40:57into a function so you can just call
- 1:40:58that function wherever you need to
- 1:41:00rather than copying all that code. Okay,
- 1:41:03so we're going to avoid repetition.
- 1:41:05Now, the the reason we're going to do
- 1:41:06that is to uh you know keep everything
- 1:41:11simple. We want to make sure things are
- 1:41:13clean, simple, understandable.
- 1:41:16Um, we don't want to h we don't want to
- 1:41:18have overly complex things that are very
- 1:41:21difficult to follow. So, one of the
- 1:41:23things that is going to be really nice
- 1:41:25about Python is it lends itself very
- 1:41:27well to being simple because it's going
- 1:41:31to be so easy to actually read and
- 1:41:33understand um, you know, understand
- 1:41:36what's going on. But one of the things
- 1:41:38that falls in line with this is like um
- 1:41:41for instance naming things
- 1:41:43appropriately. So instead of just
- 1:41:45calling everything in our code like X Y
- 1:41:47and Z if somebody comes along and reads
- 1:41:50oh I see your code has an X Y and Z that
- 1:41:53may not make sense. You know we would
- 1:41:56want to be more thoughtful with the
- 1:41:58names of our variables and names of our
- 1:42:00function. So instead of XYZ, maybe we
- 1:42:02would use something like name or place
- 1:42:06or you know something appropriate to
- 1:42:08identify this is what this is. So think
- 1:42:12about that when you're writing your code
- 1:42:14is try to make it understandable. Name
- 1:42:17things that somebody else reading it
- 1:42:20would understand what it is if they see
- 1:42:21that name. So that's that's a mistake I
- 1:42:24see a lot of people make when they first
- 1:42:25start. It's okay like when you're first
- 1:42:27getting started and practicing to name
- 1:42:28things like X, Y, and Z. I think that's
- 1:42:30fine. Or like ABC.
- 1:42:32Um,
- 1:42:34but does that make sense? Like if
- 1:42:35somebody else was reading it, they see
- 1:42:38XYZ in the program, that may not make
- 1:42:40sense, you know? So, but if it has a
- 1:42:43good name to it, you could say, oh, like
- 1:42:45I see this is somebody's name that this
- 1:42:47variable is referring to or this is um a
- 1:42:50particular object that this is referring
- 1:42:51to. Um, it's not just kind of an
- 1:42:54abstract X or Y or Z.
- 1:42:58Yeah, no spaghetti. Yeah, that's that's
- 1:43:02what uh that's what a lot of people
- 1:43:04refer to that as. Uh just sloppy,
- 1:43:07unorganized, um hard to understand code.
- 1:43:12One of the things that's great about
- 1:43:13Python is it's naturally very
- 1:43:15understandable. So like I don't think we
- 1:43:17will have that issue as much as if we
- 1:43:19had other languages, but it's still
- 1:43:22possible.
- 1:43:23So these are things we'll learn as we go
- 1:43:25along. I'm just trying to get it into
- 1:43:27your mind a little early here. Name
- 1:43:29things appropriately is main one of the
- 1:43:31main pieces of advice I can give here.
- 1:43:35Um
- 1:43:36so the next tip is to organize things.
- 1:43:39This goes along with avoiding
- 1:43:40repetition. So organize
- 1:43:43um let's put things into functions.
- 1:43:45Let's put things into objects where it
- 1:43:46makes sense. If we know we're going to
- 1:43:48reuse that um let's put it into a
- 1:43:51function. And so we're going to learn
- 1:43:52about how to do that. But generally this
- 1:43:55is good practice if you find yourself
- 1:43:58writing um uh code to do something and
- 1:44:02it turns out to be um
- 1:44:06it turns out to be uh something you know
- 1:44:08you're going to reuse or it turns out to
- 1:44:10be more than a handful of lines of code.
- 1:44:14Generally you want to organize that into
- 1:44:15a function so that uh it's clear
- 1:44:19this is what this code is doing. This is
- 1:44:21what it's responsible for. it's obvious
- 1:44:24um you know that it's organized into
- 1:44:27into uh that unit of work essentially.
- 1:44:32So we are going to practice this. This
- 1:44:34is something we're going to get good at
- 1:44:35I think as we go along because we're
- 1:44:37going to favor organization where it
- 1:44:40makes sense.
- 1:44:42Okay.
- 1:44:44So readability. One of the things is
- 1:44:46using meaningful names. I kind of
- 1:44:48already mentioned that. The other thing
- 1:44:50is using good comments. So, we're going
- 1:44:52to learn probably today how to make
- 1:44:54comments in our Python code, which is
- 1:44:56going to be helpful to orient yourself
- 1:44:58or another reader of it to, hey, this is
- 1:45:01what this function does. This is what
- 1:45:03this line of code is doing. Um, I can't
- 1:45:06tell you how many times, you know,
- 1:45:07people write code and then it they
- 1:45:09themselves come back to it a week later
- 1:45:12and have no idea what it's doing. That
- 1:45:15happens all the time. It's even happened
- 1:45:16to me. So, uh, comments are your friend
- 1:45:20in that regard. and that um they don't
- 1:45:22really cost you anything to put comments
- 1:45:23in there um to say to to kind of
- 1:45:27highlight this is what this piece of
- 1:45:30code is doing and you can make a note to
- 1:45:32yourself right within the code. That's
- 1:45:34what comments are. They're basically
- 1:45:36notes to yourself. Um
- 1:45:40so we're going to learn about that
- 1:45:41today. How to write comments and and
- 1:45:43what that looks like in the code. The
- 1:45:46other thing is indentation. you know,
- 1:45:48Python supports uh indent like you have
- 1:45:50to indent. So, that's not really going
- 1:45:52to be an issue. Some languages don't
- 1:45:54really support that, especially the
- 1:45:56compiled ones. They don't enforce
- 1:45:58strictly indentation. They enforce other
- 1:46:00things like braces and and semicolons
- 1:46:03and such, but um our our Python code
- 1:46:07will be properly indented uh by
- 1:46:10necessity because otherwise it won't
- 1:46:12work. So, um, that's something we're
- 1:46:13going to learn about too today is how we
- 1:46:16indent things and why that matters.
- 1:46:19We'll talk about that.
- 1:46:23Um,
- 1:46:24I see a question from Sherry. Is Python
- 1:46:26a program that can be programmed with
- 1:46:28simple language? Yes, it's very easy to
- 1:46:32uh it it's Python is a very natural
- 1:46:35language to program in because um yeah
- 1:46:38it's very simple uh simple languages
- 1:46:41used all over the place. I think it's
- 1:46:44going to be really easy to learn. I
- 1:46:46think it'll be really easy to pick up.
- 1:46:47At least that's my hope and I think it
- 1:46:49from my experience it is. As I said I
- 1:46:52was someone who did that and I've worked
- 1:46:54with many learners who've done the same.
- 1:46:57So yes, I think it'll be pretty easy to
- 1:47:00pick up, very simple.
- 1:47:04Um, and then the other thing is we can
- 1:47:08do uh we can find our errors very
- 1:47:10quickly. Now, because this is an
- 1:47:11interpreted language, we can run things
- 1:47:13one line at a time and we we will
- 1:47:15quickly hit errors
- 1:47:18uh early on in our code if if we have
- 1:47:20them. So this will be nice and Python
- 1:47:22provides really good um error messages
- 1:47:25um to say hey like this is what's wrong
- 1:47:28with your code you should fix it this
- 1:47:31way um essentially like giving you a
- 1:47:34clue into what needs to be fixed. Um so
- 1:47:37so this is something uh that we will
- 1:47:40practice with as we go along is kind of
- 1:47:42um finding errors and what to do with
- 1:47:45them. Um, but because it's interpreted,
- 1:47:48we will run across those very quickly.
- 1:47:50Unlike with compiled language, which is
- 1:47:52harder to debug because you basically
- 1:47:53have to compile everything, hope that it
- 1:47:56compiles. If it does, then you have to
- 1:47:58run things. Um, it just takes longer to
- 1:48:01get through that debugging phase. But
- 1:48:03with the with Python, it's very quick.
- 1:48:05You get a very quick feedback loop on if
- 1:48:08your code's working or not, which is
- 1:48:10nice. A lot of votes for C. I agree. AC
- 1:48:14is the correct answer here. So the
- 1:48:17interpreter is the thing that will
- 1:48:20execute the code line by line. So it
- 1:48:23doesn't do everything at once. It
- 1:48:25actually goes line by line, which is why
- 1:48:29you can stumble onto your errors quickly
- 1:48:32because if you're going line by line
- 1:48:35um and you have an error on this first
- 1:48:37line, you're never going to reach these
- 1:48:39other lines, right? You're it's just
- 1:48:40going to show you this is where your
- 1:48:42error is. it's on line 101 or whatever
- 1:48:44it is and you know it's going to show
- 1:48:47you where the error is. So it's going to
- 1:48:49go one at a time and execute those. Um
- 1:48:53it's not going to convert the code into
- 1:48:56machine language. That's what a compiled
- 1:48:58language would do, not an interpreted
- 1:49:00one. Um and uh they do require an
- 1:49:05interpreter. So D is just completely
- 1:49:07wrong. It's the opposite of that. It
- 1:49:08does require it. So the interpreter is
- 1:49:10the thing that is executing the uh code
- 1:49:13line by line. So what is Python in
- 1:49:16particular? So it is a as we've already
- 1:49:20seen an interpreted language meaning
- 1:49:23that it requires an interpreter to
- 1:49:25execute it. It's going to be executed
- 1:49:26line by line by that interpreter. Um it
- 1:49:29has capability to be object-oriented. It
- 1:49:32also has capability to be scripted.
- 1:49:34um which is just in relation to how it's
- 1:49:37organized. One of the really nice things
- 1:49:41is it is what we call dynamically typed
- 1:49:45or what you would say dynamic semantics.
- 1:49:48We will see what this means but
- 1:49:51basically it means that we don't have to
- 1:49:53declare what every piece of uh what
- 1:49:56every variable or every piece of data is
- 1:49:59inside of Python. We can let the
- 1:50:00interpreter interpret that which is
- 1:50:03nice. It makes things really easy to
- 1:50:05work with. We don't need to say okay
- 1:50:07this is an integer. This is a
- 1:50:08floatingoint number. This is an array.
- 1:50:10This is you know with a lot of program
- 1:50:14especially compiled languages
- 1:50:15programming languages you have to do
- 1:50:18that because you have to tell the
- 1:50:20compiler this is what this piece of data
- 1:50:23is. But with an interpreter the
- 1:50:25interpreter can as the name suggests
- 1:50:28interpret that. It doesn't need to know
- 1:50:30what everything is in terms of its data
- 1:50:33type, which is which makes it really
- 1:50:35easy to code. On the cons of that, it
- 1:50:39can make it more prone to error because
- 1:50:41you're not really enforcing types. So,
- 1:50:44there is somewhat of a trade-off there.
- 1:50:46But um for our purposes the dynamic
- 1:50:49semantics make make it so that um the
- 1:50:52interpreter can dynamically understand
- 1:50:55what data is um based on how it's being
- 1:50:59used which is great um for us like it
- 1:51:02makes it just quicker to get up and
- 1:51:04running and started and and working with
- 1:51:06data. We don't need to declare what its
- 1:51:08type is which is um static semantics.
- 1:51:12Um now Python itself amazing programming
- 1:51:16language that's used across many
- 1:51:18different applications um such as data
- 1:51:20science, automation, machine learning,
- 1:51:23AI. It's also used in to build software
- 1:51:27even um not sure if you guys know this
- 1:51:29but there's um some really famous
- 1:51:32software that's written in Python. Um,
- 1:51:35one of the most famous is Instagram at
- 1:51:38Meta is completely coded in Python,
- 1:51:41which is it's over like 20,000 lines of
- 1:51:43Python code, which is pretty amazing.
- 1:51:46But um so of course it's been really um
- 1:51:51heavily used in AI and machine learning
- 1:51:53and such but it's also as a programming
- 1:51:56language been used for other things like
- 1:51:58more pure software applications which is
- 1:52:00what makes Python really nice is it's so
- 1:52:02simple so easy to learn. Um so for that
- 1:52:06reason uh it is going to be great for us
- 1:52:10to get started with especially if you're
- 1:52:11coming in with basically no programming
- 1:52:13experience. The other thing about Python
- 1:52:16is it has uh as I said earlier like a
- 1:52:20really big ecosystem uh meaning that
- 1:52:22there's many different packages and
- 1:52:25modules within those package packages
- 1:52:28that do things already. So we don't
- 1:52:31what's great about Python is we won't
- 1:52:33need to reinvent the wheel on so many
- 1:52:36different things like if we need to
- 1:52:38build a plot if we need to train a model
- 1:52:41and and use a specific type of model
- 1:52:44that likely already exists in a package
- 1:52:47somewhere. And what's great is they're
- 1:52:50almost always open source meaning we
- 1:52:52don't have to pay for anything. You can
- 1:52:54just use it out of the box which is
- 1:52:56fantastic. So there's within Python
- 1:52:59there's so many ways to do things
- 1:53:01especially in the AI and machine
- 1:53:02learning world that we'll just borrow
- 1:53:05those and use them in our own code um
- 1:53:08which helps uh you know with um getting
- 1:53:12up and running very quickly. We don't
- 1:53:14need to reinvent things. We can just use
- 1:53:16things that already exist um which is
- 1:53:19fantastic. So that ecosystem really
- 1:53:22benefits machine learning AI. Um because
- 1:53:26they they already exist. We don't need
- 1:53:28to spend our time rewriting all those
- 1:53:30things. Um and so that's something we're
- 1:53:34going to learn as we go along is like
- 1:53:35how to install those, how to import
- 1:53:38those, how to use those in our own code,
- 1:53:40those those packages that already do
- 1:53:44something for us. So we don't need to
- 1:53:46come up with it on our own. we just need
- 1:53:48to use it properly. Okay, so there's a
- 1:53:51little bit of history. Python was first
- 1:53:54invented in the late 1980s by a guy
- 1:53:57named Guido Van Rossom in Amsterdam. Um,
- 1:54:01where it gets its name is after the old
- 1:54:04comedy series, you guys might be
- 1:54:05familiar with it, the Monty Python
- 1:54:07Flying Circus Show. Um, and so that's
- 1:54:11where it's got its name. um you know it
- 1:54:14was first created then but has since
- 1:54:16taken on a really big role in the
- 1:54:20especially you know I keep saying in the
- 1:54:22AI community so much so that it has its
- 1:54:25own software foundation that kind of is
- 1:54:27responsible for maintaining it they meet
- 1:54:29regularly they come up with improvements
- 1:54:33um they come up with new versions of
- 1:54:35Python
- 1:54:36uh for example Python 3.14 just released
- 1:54:40in October which is a major release. Uh
- 1:54:44they hadn't had one in a while and that
- 1:54:47one is uh 3.14. So it's kind of known as
- 1:54:50Python.
- 1:54:52Um which was a big milestone. Um but you
- 1:54:56know they have uh they've had many
- 1:54:58different versions over the years. It's
- 1:55:00been maintained and developed by this
- 1:55:01software foundation. Um and people are
- 1:55:06actively working on it at many large
- 1:55:09companies. So for instance, Meta has a
- 1:55:11big group that is working on um Python
- 1:55:14improvements. Microsoft as well, um
- 1:55:17Google, all of those guys have groups
- 1:55:19kind of working to improve Python
- 1:55:20because they all use it. And so what
- 1:55:22they typically do is work on it, open
- 1:55:25source it, and then the community gets
- 1:55:27to use those tools, those packages,
- 1:55:29those tools, those improvements. Um so
- 1:55:32it's it's actively um utilized across
- 1:55:35many big companies actively uh
- 1:55:37maintained by them or contributed to by
- 1:55:40them. So that's that's really great. Um
- 1:55:43you know Python was originally derived
- 1:55:46from other language um other languages
- 1:55:51uh as kind of a trying to find like a
- 1:55:54mixture of some of the best of all
- 1:55:56worlds. But its main like driving force
- 1:56:00in why Python came to existence from
- 1:56:02these other languages is it just its
- 1:56:05ease of use. People really wanted
- 1:56:07something like super easy to get up and
- 1:56:08running and something really natural.
- 1:56:11Um and so we will as we start learning
- 1:56:14the syntax of it I think you guys will
- 1:56:16understand why it's so easy. But um
- 1:56:18that's that's what led to the
- 1:56:19inspiration is just people wanted
- 1:56:21something easier to work with not as not
- 1:56:23as uh strenuous to kind of get up and
- 1:56:26running.
- 1:56:29What open source license is it? Um,
- 1:56:32that's a good question. I think it's the
- 1:56:34MIT license, but I could be wrong on
- 1:56:36that.
- 1:56:37You could look it up. If you go to
- 1:56:39python.org.
- 1:56:41Yeah, if you go to python.org, I think
- 1:56:43it might talk more about what the uh
- 1:56:46license structure is there. I want to
- 1:56:48say it's MIT open license, but
- 1:56:52I've I'm really not 100% sure on that.
- 1:56:57Okay. So, what are some of the benefits
- 1:56:59of working with Python? And these are
- 1:57:00things you will experience as we go
- 1:57:02along, but just wanted to call them out.
- 1:57:04Um, the flexibility of it. As I said, it
- 1:57:07can be really organized into
- 1:57:09object-oriented or it can be loosely
- 1:57:11organized into scripts. So, that
- 1:57:14flexibility alone is really awesome. um
- 1:57:17which has allowed it to power many
- 1:57:20different things like um APIs, web
- 1:57:22pages, full-blown applications like
- 1:57:24Instagram, um chat, GPTs, like actual uh
- 1:57:29AI, LLMs.
- 1:57:32Um you know, it has so much flexibility
- 1:57:35there to power so many different
- 1:57:36applications.
- 1:57:38Um probably the biggest benefit,
- 1:57:40especially to us, is its ease of use.
- 1:57:43Um,
- 1:57:45uh,
- 1:57:48oh, thank you. Some Tim just posted it.
- 1:57:50It's the the GNU,
- 1:57:53uh, public license. Yes.
- 1:57:58Oh, never mind. It's a Python software.
- 1:58:00It has its own. Okay, perfect. Thanks
- 1:58:02for sharing that. Thanks for sharing
- 1:58:04that. Yeah, I wasn't completely sure
- 1:58:06which which license it was, but it is
- 1:58:09open source. Um, and people do make
- 1:58:12their own kind of derivations of Python.
- 1:58:16But as I was saying, one of the benefits
- 1:58:17of Python is how easy it is to learn. I
- 1:58:20keep emphasizing that because it's true.
- 1:58:22Once we get into it, you will see this.
- 1:58:24I promise it'll be easy to learn, easy
- 1:58:26to pick up. Um, and it's designed in
- 1:58:30that way. Designed to be very minimalist
- 1:58:32as a language, which is great.
- 1:58:35um it has a lot of things that come with
- 1:58:38it and it's kind of built into Python, a
- 1:58:41lot of capability. So we call that the
- 1:58:43standard library. It's just the things
- 1:58:45built into Python. It has a lot of
- 1:58:46capability out of the box. Um you know,
- 1:58:50not only that, but it has a large
- 1:58:51community that's developed so many
- 1:58:52different packages that do things for
- 1:58:54us, especially in the AI world. So
- 1:58:57that's another great thing, kind of a
- 1:58:59robust community developing these
- 1:59:01packages that help us get things done.
- 1:59:05Um, readability. So because the code is
- 1:59:08so simple, it's also easy to read. So
- 1:59:11you can usually read other Python code
- 1:59:13and quickly understand what it's doing
- 1:59:15which you know makes for easy um easy
- 1:59:20understanding of other people's code
- 1:59:21easy understanding of code in the
- 1:59:23community and kind of almost like it's
- 1:59:26selfdocumenting because it's so easy to
- 1:59:28read. So that that simplicity that ease
- 1:59:32of use lends itself well to being really
- 1:59:35readable. You can usually just take a
- 1:59:37look at the code, easily read it,
- 1:59:39understand what it's doing, which is
- 1:59:41great, like great for you guys learning,
- 1:59:44great for taking a look at the demos and
- 1:59:46examples that we will do. They're very
- 1:59:48readable.
- 1:59:56Okay. So why has Python really dominated
- 2:00:02AI? So this is a valid question like
- 2:00:04even so it's used for many different
- 2:00:06things. It's a programming language. So
- 2:00:07it can build application and I've given
- 2:00:09you the example of Instagram and there's
- 2:00:10many others um that are built off of
- 2:00:14Python code. Why is it so useful for AI
- 2:00:19in particular?
- 2:00:21mainly
- 2:00:23uh some of the reasons we've already
- 2:00:25talked about mainly how easy it is to
- 2:00:28use lends itself well for AI because um
- 2:00:32that has allowed people to kind of
- 2:00:35quickly get up and running and test out
- 2:00:36their algorithms, test out their models
- 2:00:40just really quickly with Python. That's
- 2:00:42great. The other things listed on here
- 2:00:46are certainly big reasons as well. So
- 2:00:50for example, it has so many community
- 2:00:54libraries, those those packages that um
- 2:00:58have AI models and AI tools that we can
- 2:01:03reuse that people have built these up
- 2:01:05over years and years and years. Um so
- 2:01:08it's to our benefit to reuse those and
- 2:01:11not have to reinvent everything and we
- 2:01:13can get quickly up and running with
- 2:01:14those which would be great.
- 2:01:17The other thing is Python, it lends
- 2:01:18itself very very well to working with
- 2:01:20data in general. Very easy to work with
- 2:01:23data, very easy to load it in from
- 2:01:24external sources, query it, work with
- 2:01:27it, visualize it. Python is so adept at
- 2:01:31that. Um, so that's what makes it really
- 2:01:34nice at doing machine learning and AI
- 2:01:35because so much of it is manipulating
- 2:01:37data. So, um, for that reason alone,
- 2:01:41Python is so popular in the AI community
- 2:01:43just because of its ability to work with
- 2:01:45data. It's so easy. This is something
- 2:01:47we're going to really focus in on like
- 2:01:49in our next course when we talk about
- 2:01:51data science.
- 2:01:53But, um,
- 2:01:55just the ability and the power of it to
- 2:01:58work with data makes lends itself well
- 2:02:00to AI uh, capabilities.
- 2:02:03Um, the other thing is I mentioned the
- 2:02:05rapid prototyping. You can quickly build
- 2:02:07a model in Python because the code is so
- 2:02:09easy. So, and there's so many libraries
- 2:02:11already can quickly prototype. Um,
- 2:02:15it has obviously a big community around
- 2:02:18it that's building out these packages,
- 2:02:20writing documentation, maintaining it
- 2:02:22from an open source level. So, that's
- 2:02:25another reason it's very popular. Um,
- 2:02:27Python's also used with other
- 2:02:29technologies. So, it does have
- 2:02:32capability to integrate with other
- 2:02:34languages. So for instance, Python can
- 2:02:36one of the most popular integrations is
- 2:02:38Python can work with C and C++. So
- 2:02:41sometimes that's necessary to integrate
- 2:02:43with those to do certain things. Um so
- 2:02:47Python has been extended to work with
- 2:02:49other languages. So sometimes there's
- 2:02:51other uh necessary support from other
- 2:02:55like things in other languages that are
- 2:02:56necessary to power something in AI. um
- 2:02:59for example working with GPUs
- 2:03:03and doing things in deep learning. Um
- 2:03:07there's been a lot of integration with
- 2:03:09uh working with um C tools. Now will we
- 2:03:12do that? No, it's already been done for
- 2:03:15us and some of these packages. But um
- 2:03:19the pure ability of Python to do that is
- 2:03:22really powerful and it gets taken for
- 2:03:24granted honestly because you don't see
- 2:03:26that it's underneath the hood and it's
- 2:03:28abstracted away from you when you work
- 2:03:29with those Python packages. But there
- 2:03:32was a lot of work that went into it to
- 2:03:33integrate it with other kind of other
- 2:03:35programming languages.
- 2:03:39Okay.
- 2:03:41So as an example like I mentioned the
- 2:03:43Instagram one. So Netflix for instance,
- 2:03:45all of their recommendation is powered
- 2:03:48by Python. So when you open up Netflix
- 2:03:51or really any streaming service for that
- 2:03:54matter, they're going to use Python to
- 2:03:56deliver those recommendations and
- 2:03:58produce those personalized
- 2:03:59recommendations. Um Spotify as well for
- 2:04:02like music. Um nearly all recommendation
- 2:04:05algorithms are written in Python.
- 2:04:08And in this program, we are actually
- 2:04:11going to learn about recommendation
- 2:04:13systems. So that'll be pretty fun way
- 2:04:16down the road when we get into machine
- 2:04:17learning. We'll talk about how do we
- 2:04:19build a recommendation engine,
- 2:04:22but um they're all done through Python
- 2:04:25for for example. So really cool uh use
- 2:04:28cases there.
- 2:04:32So one of the things I wanted to address
- 2:04:34is how AI itself is changing coding. So
- 2:04:39you guys may be aware of this, but
- 2:04:41obviously there's been a huge um kind of
- 2:04:45explosion in generative AI tools that
- 2:04:48can help write documents and write
- 2:04:50emails and write text and all these
- 2:04:52things. One of the things they can do is
- 2:04:54write code. So um one of the big areas
- 2:04:58where AI is changing coding is it's an
- 2:05:00its ability to generate code for us. And
- 2:05:04so um throughout this program like we
- 2:05:07won't shy away from that necessarily
- 2:05:10and I encourage you guys to use AI tools
- 2:05:13as you see fit to help your own
- 2:05:16understanding and help your own
- 2:05:17productivity. Um
- 2:05:20you know we still will go through the
- 2:05:22fundamentals so you can understand it
- 2:05:24but the AI tools can definitely be a
- 2:05:27supplement to help. Um it's just that I
- 2:05:31think you guys will understand it better
- 2:05:32going through the examples that we do we
- 2:05:34do together and so that when AI
- 2:05:37generates code you will be able to
- 2:05:39understand it and also be able to debug
- 2:05:41it right because it's not always going
- 2:05:43to be perfect. So that's always the
- 2:05:45catch with AI is that you know it
- 2:05:48doesn't always produce perfect answers.
- 2:05:50Um but the at the very least we will be
- 2:05:53able to you know debug things and
- 2:05:56understand things better so that uh we
- 2:05:59can catch those errors.
- 2:06:01Um so obviously like AI is also besides
- 2:06:05flat out generating it it's also
- 2:06:07suggesting what should be there. So, uh,
- 2:06:10some of the code editors really do a
- 2:06:12good job at that, suggesting things, um,
- 2:06:15picking up on what you should produce
- 2:06:17next. That's going to be, um, very
- 2:06:20interesting as we get into, uh, some of
- 2:06:23the platforms that you guys will work
- 2:06:25with to write your Python code. They
- 2:06:27will have that ability. Um, so, uh, the
- 2:06:32other thing is like there's some cloud
- 2:06:35tools that, um, don't require writing
- 2:06:39much code at all and they can just do
- 2:06:41things. So, in other words, you can
- 2:06:42power them by prompts. You're not really
- 2:06:44writing code. You're just writing
- 2:06:45natural language and then they do
- 2:06:47something. Um, they generate the code in
- 2:06:50the background and they execute
- 2:06:52something. Um we will learn about those
- 2:06:55things uh later on in the program
- 2:06:57especially because we we will cover
- 2:06:59generative AI in the future
- 2:07:03um towards the end of our program. So if
- 2:07:05you're wondering like are we going to
- 2:07:07cover LLMs? Are we going to cover how
- 2:07:09these things get generated? Yes. It just
- 2:07:12will be um later on in the program.
- 2:07:18Okay. A lot of votes for B.
- 2:07:20Yeah, pretty unanimous on B. I think I
- 2:07:22agree with it. Yeah, B is definitely the
- 2:07:23right answer. So, all of the
- 2:07:25recommendation systems which we will
- 2:07:27learn how to build ourselves later on
- 2:07:30are written in Python and um they uh are
- 2:07:36machine learning models that make the
- 2:07:37recommendations and that machine
- 2:07:39learning is driven by data um and all of
- 2:07:43that data is manipulated in Python
- 2:07:46um and used to train uh models that do
- 2:07:49the recommendations. That's all
- 2:07:51happening in Python.
- 2:07:54So, we're going to talk about getting
- 2:07:56you guys set up on your own machine and
- 2:07:59talking about the different development
- 2:08:02environments we can use to actually work
- 2:08:04with Python code. Um, before we go into
- 2:08:08that, any questions about anything we
- 2:08:09covered so far?
- 2:08:12Everything's good so far. Yep. And you
- 2:08:15know again if you have experience in
- 2:08:17Python I recognize that it is going to
- 2:08:18be a little slow in beginning. Um it's
- 2:08:21mostly to get us really oriented to some
- 2:08:24background around Python and get us set
- 2:08:26up and then we will be doing you know uh
- 2:08:29getting into the syntax and all that uh
- 2:08:33coming up shortly. So we will actually
- 2:08:36be learning Python specifics coming up
- 2:08:38soon. But you know we're going to um get
- 2:08:41everything set up first.
- 2:08:45All right. So, let's continue then.
- 2:08:47Thank you guys for that.
- 2:08:50So, um it turns out that there are many
- 2:08:53tools in the community for developing
- 2:08:56Python code. And so, um you might hear
- 2:08:59this word ID. It is short for integrated
- 2:09:02development environment. This is a piece
- 2:09:04of software that helps you write and
- 2:09:08test Python code. So, and there's many
- 2:09:12out there. There's a bunch on this list.
- 2:09:14We are going to focus on a few options.
- 2:09:18There's even more than what's on this
- 2:09:20list, but we're going to focus on a few
- 2:09:22options. These IDs are designed to
- 2:09:24really help you write Python. They
- 2:09:26provide many tools in the background
- 2:09:29that make your life easier when you're
- 2:09:31working with Python. So, for example,
- 2:09:33they can provide syntax highlighting.
- 2:09:36They can tell you when you have a syntax
- 2:09:38error. Um, almost like a spell check for
- 2:09:41Python.
- 2:09:43Um, they can help you run Python code
- 2:09:45right within the window. Um, they can
- 2:09:48help you organize your projects. Uh,
- 2:09:50they can do a lot of different things.
- 2:09:53Um, and so there's many tools out there
- 2:09:55that can do it, and it's really a
- 2:09:58personal preference which one you use,
- 2:10:00but in this program, we're really going
- 2:10:02to focus on a few of them to to showcase
- 2:10:05those options because they're very
- 2:10:06popular options. Um, and then, uh, allow
- 2:10:11you guys the flexibility to choose which
- 2:10:13option makes the most sense for you. So,
- 2:10:15generally, that's going to be mostly a a
- 2:10:19preference.
- 2:10:20um mostly a preference as to which one
- 2:10:23you're the most comfortable with, but I
- 2:10:25want to give you guys the option to uh
- 2:10:28explore
- 2:10:30the various options that are available.
- 2:10:36Uh Roberto, is there one that stands out
- 2:10:38as an industry standard? Um there's a
- 2:10:41couple that you see like honestly the
- 2:10:45two of them that we will study uh in
- 2:10:48this coming up in the next few slides
- 2:10:50are the industry standard which are
- 2:10:51going to be VS code Microsoft VS code
- 2:10:54and then Jupyter notebooks. So these two
- 2:11:00are going to be uh ones that we will
- 2:11:03study in particular and use throughout.
- 2:11:07Um
- 2:11:10so so yes we will those will be industry
- 2:11:13standards. PyCharm's also very popular.
- 2:11:16Um so I don't want to rule out PyCharm.
- 2:11:18I know a lot of people who use it. So um
- 2:11:21I would encourage you to explore PyCharm
- 2:11:23as well if you want to but we are not
- 2:11:25going to do that uh in in these slides
- 2:11:28but um I would check it out and see if
- 2:11:32you like it. Um it's another very I'm
- 2:11:35putting a an asterisk next to it because
- 2:11:37I think it's one of the more popular
- 2:11:40uh yes uh yeah we're going to do
- 2:11:43descriptions.
- 2:11:45um requirements uh I'll try my best to
- 2:11:48give those but honestly the requirements
- 2:11:49will be given when you install them. Um
- 2:11:54so the other thing I want to say is we
- 2:11:56will have a couple options that don't
- 2:11:58require you to install anything. So I'm
- 2:12:00going to showcase those as well. So
- 2:12:03there's a couple options that are um we
- 2:12:05won't have to install anything because
- 2:12:07they're going to be cloud-based.
- 2:12:09Okay, I'll show you those.
- 2:12:16Okay. So, but yeah, VS Code, I think VS
- 2:12:18Code and Jupyter notebooks are are
- 2:12:21probably the industry standard most
- 2:12:23popular uh idees.
- 2:12:27Okay. So, what we would recommend in
- 2:12:30this program and the ones that we will
- 2:12:32use the most uh throughout are going to
- 2:12:35be these three. Visual Studio Code, also
- 2:12:37known as VS Code, Jupyter Notebooks, and
- 2:12:40Google Coll Collab, which is Google's
- 2:12:43hosted
- 2:12:45um Google's hosted version of notebooks
- 2:12:48essentially. Um so
- 2:12:52I will showcase each one of these and
- 2:12:54give you some examples of how to set it
- 2:12:56up and examples of how to work with it.
- 2:12:59Um, and that's what we'll do over the
- 2:13:01course of the next few slides and the
- 2:13:03next uh bit of time is I'm going to go
- 2:13:06through each one of these and kind of
- 2:13:07show you what you would need to do to
- 2:13:08get it set up. Um, now that being said,
- 2:13:15excuse me, these two are ones that you
- 2:13:18will install.
- 2:13:20These two you would install locally on
- 2:13:23your on your own machine.
- 2:13:26And this one is um uh cloud hosted
- 2:13:32by Google and it's free. Um all of these
- 2:13:36are free but uh the first two VS code
- 2:13:40and Jupyter notebook you would install
- 2:13:42on your own machine. Collab you would
- 2:13:43just access through your web browser. It
- 2:13:45is hosted by Google. So that's an
- 2:13:47advantage. You don't really need to
- 2:13:48install anything. And for that reason um
- 2:13:51sometimes we will favor Collab. Uh and
- 2:13:53for other reasons too. Collab has some
- 2:13:55really nice features if you've never
- 2:13:57used it. Um, but notebooks, um, Jupyter
- 2:14:03Notebook and Collab are very similar.
- 2:14:06They're very similar. Collab just has
- 2:14:08its own spin-off on on the notebook, um,
- 2:14:11type of file that Jupyter Notebooks work
- 2:14:14with. And it's um, like I said, kind of
- 2:14:16cloud hosted. So, I'm going to I'm going
- 2:14:18to walk us through each one of these and
- 2:14:21explain to you what they do, what they
- 2:14:24look like, and then we will um I'll set
- 2:14:26up each one of them uh kind of in a live
- 2:14:28demo so you guys can see. Um but uh we
- 2:14:34throughout the program, it will really
- 2:14:37be up to you which one of these you want
- 2:14:39to use. There's no hard requirement to
- 2:14:41use any one of them. It's really going
- 2:14:43to be your preference which one of these
- 2:14:46tools you want to use to work with
- 2:14:47Python. Whatever one you feel
- 2:14:49comfortable working with, that's the one
- 2:14:51you should use.
- 2:14:53All three of these are very popular in
- 2:14:55the industry. So, you're not missing out
- 2:14:56by using one versus the other. Um,
- 2:14:59they're all very popular. Even Collab, I
- 2:15:02know it wasn't on the screen, but it is
- 2:15:04widely used in in the community and the
- 2:15:06industry.
- 2:15:10Uh, no system recommendations for
- 2:15:11training LLMs. Um, no, because we don't
- 2:15:13we won't really focus on that until the
- 2:15:15end. When we get to when we get into
- 2:15:17generative AI, we'll talk about that.
- 2:15:21When we get into generative AI, we'll
- 2:15:22talk about that.
- 2:15:28So, yeah, we're not we're not focusing
- 2:15:30on LM in the beginning. That's that's an
- 2:15:33advanced topic for us.
- 2:15:36What is my personal preference? Um, I
- 2:15:40like Visual Studio Code. Um, personally
- 2:15:43I that's what I use for my day-to-day
- 2:15:45work is uh Visual Studio Code. I like
- 2:15:48Visual Studio Code and I like Collab a
- 2:15:50lot. Um, so you know, we'll talk about
- 2:15:54this, but one of the advantages to
- 2:15:56Collab is that it has free access to
- 2:15:58GPUs, which is huge for doing things
- 2:16:02like uh neural nets. Um, so we will lean
- 2:16:06on collab quite a bit later on
- 2:16:10uh later on when we um actually get to
- 2:16:15deep learning and neural nets. We'll
- 2:16:17because collab has free access to GPUs.
- 2:16:20I'll show us that. It's it's really
- 2:16:22nice.
- 2:16:24And when you do anything with neural
- 2:16:25nets, it usually benefits you to have a
- 2:16:27GPU access. Um
- 2:16:30so that'll be nice. But I usually do
- 2:16:33most Python coding inside of VS Code. It
- 2:16:36supports Python pretty pretty well.
- 2:16:42What is more commonly used in the
- 2:16:44industry? Um,
- 2:16:46the two most popular are Visual Studio
- 2:16:49Code and and Notebooks. Jupiter
- 2:16:51notebooks.
- 2:16:52They're both like you can't go wrong
- 2:16:54with either one.
- 2:16:58Those
- 2:17:01two are really popular. Jupyter
- 2:17:02notebooks and Visual Studio Code are
- 2:17:04really popular. There's there's both of
- 2:17:07those you would be okay with. Either
- 2:17:10one.
- 2:17:12Let me start with Visual Stu Studio
- 2:17:13Code. So, um now Visual Studio Code
- 2:17:18is a more general code editor. So, it's
- 2:17:22actually you can edit lots of different
- 2:17:25languages inside of VS Code. Um, so you
- 2:17:29could do Java, you could do C, you could
- 2:17:31do Scala, you can do Go, you can do all
- 2:17:35kinds of languages are supported inside
- 2:17:37of Visual Studio Code. So it's a really
- 2:17:39fantastic product for programming in
- 2:17:41general, not just Python. Um, it has
- 2:17:44built-in terminal support. It has
- 2:17:47co-pilot integrated into it, which is
- 2:17:50nice for AI, like generative AI
- 2:17:53assistance working with your code, which
- 2:17:55is nice. Um, of course it supports
- 2:17:58Python, which is what we are interested
- 2:18:00in. Um, it has it has Python tools. I
- 2:18:04will show us which ones we should
- 2:18:06install as part of VS Code so that we
- 2:18:09can work with Python files and
- 2:18:11notebooks. Um, so it's it's a really
- 2:18:15great code editor in general, which is
- 2:18:17why I like using it. Um, but in
- 2:18:20particular, it's pretty good at working
- 2:18:22with Python. it it supports Python
- 2:18:24pretty uh deeply. Um so and for that
- 2:18:28reason VS code is really really popular.
- 2:18:32But just keep in mind you can actually
- 2:18:33use it for many different types of code
- 2:18:35that uh that people write uh JavaScript
- 2:18:39um Java as I said like many languages
- 2:18:42are supported inside of Visual Studio
- 2:18:43Code. So it's a more general code
- 2:18:46editor. It happens to be really great at
- 2:18:48working with Python though.
- 2:18:52All right. So, I'm going to show us a
- 2:18:53demo on setting up VS Code. Now, we are
- 2:18:57going to do this for each one of these.
- 2:19:01For Jupiter and for Collab, I'm going to
- 2:19:03I'm going to do similar demos. So, um
- 2:19:07don't worry, we'll get to those, but I
- 2:19:09want to start with VS Code to show you
- 2:19:11kind of how to get that set up and what
- 2:19:13it looks like. Um, so where you can find
- 2:19:16this demo
- 2:19:19is inside of the demos that I mentioned
- 2:19:23earlier in the reference material. So
- 2:19:24I'm going to I'm going to jump over to
- 2:19:27that. Let me show you guys.
- 2:19:32So I'm back in the LMS. You guys will
- 2:19:35want to download the demos. I think
- 2:19:37somebody linked it earlier in case this
- 2:19:39didn't show up for you, but we're going
- 2:19:40to be inside of the demos and we're
- 2:19:42going to do demo one for lesson one. We
- 2:19:45do lesson one, demo one, which is going
- 2:19:47to be the VS Code demo.
- 2:19:51So, the main steps that we're going to
- 2:19:53do is just going to be to point you to
- 2:19:57where to install Visual Studio Code. So,
- 2:19:59it is a it is an application is a free
- 2:20:02application you can install on your
- 2:20:03machine. Um, so,
- 2:20:07uh, you will want to follow this link
- 2:20:10that is within the demo file, this
- 2:20:12code.vvisualstudio.com/d
- 2:20:14download and download it for your
- 2:20:16particular platform. So, if you're on
- 2:20:17Windows, obviously, choose the Windows.
- 2:20:20If you're on a Mac, um, choose Mac and
- 2:20:24make sure that you choose the right, one
- 2:20:27of the precautions is to choose the
- 2:20:28right Mac platform. So, if you have like
- 2:20:30an M1, 2, M3, M4 Mac, choose the Apple
- 2:20:34Silicon
- 2:20:36um button. If you're on an older Mac, um
- 2:20:39then you'll want to use the Intel chip
- 2:20:41one. Um
- 2:20:45uh if you're on if you happen to be on
- 2:20:47Linux, which I don't probably most of
- 2:20:50you are not, but if you are, um you want
- 2:20:52to download the right uh distribution uh
- 2:20:55version.
- 2:20:57But, uh follow this link first. So
- 2:20:59that's the first step. Very easy step.
- 2:21:01Just go to that site, pick your right
- 2:21:03platform and uh go ahead and download
- 2:21:06the installer. And mostly we will be
- 2:21:10walking through the steps in the
- 2:21:12installer. And then um I will show us
- 2:21:15what it looks like once it's installed
- 2:21:18and then show you a couple additional
- 2:21:20steps that are actually not mentioned in
- 2:21:21this file that I think are worth doing
- 2:21:23to get you set up.
- 2:21:29Uh yes, we will be doing Jupiter next.
- 2:21:31Yes, we we'll we're going to be covering
- 2:21:34VS Code, Jupiter, and Collab. I'm going
- 2:21:36to show us examples of all of those.
- 2:21:46Okay, let me ask you guys. Were you guys
- 2:21:48able to get to the download page and
- 2:21:50start that download and installation of
- 2:21:52VS Code?
- 2:21:56able to do that.
- 2:21:58Any issues with that?
- 2:22:05Okay. Yeah, it's just like any
- 2:22:08yet I love I love the optimism
- 2:22:12yet.
- 2:22:16Uh already having both of them
- 2:22:18installed. Okay. Yeah. No, if you
- 2:22:20already have it installed, I mean,
- 2:22:21great. I'll show so if you if you
- 2:22:23already have VS Code installed, great.
- 2:22:25You can sit tight. I will show you um a
- 2:22:29couple of extensions that you'll want to
- 2:22:31add for Python support
- 2:22:34if you have it installed already. I'll
- 2:22:37show us how you can use it with Python
- 2:22:39in particular.
- 2:22:41Okay.
- 2:22:44If you already have it installed,
- 2:22:45perfect. Looks like you have it
- 2:22:47launched.
- 2:22:51Still working on it. Okay. So, these
- 2:22:54these instructions um uh show an example
- 2:22:57of someone that would be on a Microsoft
- 2:22:59platform um walking through the
- 2:23:03installation.
- 2:23:07Uh if you're on a Windows, you probably
- 2:23:10want to create a desktop icon. You
- 2:23:11definitely want to add it to your path.
- 2:23:21And this just shows what's being
- 2:23:23installed. So this is all the install
- 2:23:24wizard on Windows. Nothing that exciting
- 2:23:27there. So this if you follow all these
- 2:23:30steps, you will have it installed. I
- 2:23:32hope you have enough disc space. Uh I
- 2:23:34don't think it's too big.
- 2:23:37I don't think it's too too massive. I
- 2:23:39forget how much space it takes. I don't
- 2:23:40think it's that much.
- 2:23:45I don't think it's that much. But um
- 2:23:47yeah, hopefully you have enough.
- 2:23:51So if if you don't
- 2:23:54uh if you do not have enough disc space
- 2:23:57um don't worry because we're going to do
- 2:23:59collab which doesn't require you
- 2:24:01installing anything. So you can always
- 2:24:03use that option. All right. So if if for
- 2:24:06some re let me just say that too just
- 2:24:08even if if it's not a dispace issue if
- 2:24:10you have an in any installation issues
- 2:24:13no worries because we will work with
- 2:24:16collab and Google that is going to be
- 2:24:18cloud hosted that you don't need to
- 2:24:20install anything you just need a Google
- 2:24:22account
- 2:24:24okay a free Google account
- 2:24:27um so no worries at all if you cannot
- 2:24:30get any of these things installed the
- 2:24:33which are going to be Jupiter and
- 2:24:36uh Jupiter and VS Code.
- 2:24:41Where do we go? I haven't said yet. It
- 2:24:42just I'm just making sure it's installed
- 2:24:44for folks.
- 2:24:46I'm going to I'm going to go over to it
- 2:24:47in a second, but did we generally get it
- 2:24:50installed and do you have it open? So,
- 2:24:52if you once you get it installed,
- 2:24:55uh once you get it installed, then open
- 2:24:57it.
- 2:25:04Yeah, you need to get it installed. Uh,
- 2:25:07it should be this first. It should be
- 2:25:09this link here.
- 2:25:13Follow this link to get it installed.
- 2:25:20Oops, I pasted the wrong link.
- 2:25:37Let me find I'll copy and paste the
- 2:25:39link. But yeah, take take a moment to
- 2:25:41get it open. Once you have it open, just
- 2:25:44sit tight
- 2:25:47if you want to.
- 2:25:50What does it say?
- 2:25:54Yeah, feel. So, for you guys seeing the
- 2:25:57co-pilot features, um,
- 2:26:00click click use AI features. I think
- 2:26:03that's okay. Yes. Um, you'll you'll
- 2:26:06likely want copilot. Yes.
- 2:26:09Click click okay on that.
- 2:26:17That's the link, by the way, for the
- 2:26:19download
- 2:26:21in case uh we needed to get to it.
- 2:26:30Okay.
- 2:26:32So, I'm going to go over to VS Code
- 2:26:36and show you what it looks like on uh my
- 2:26:38end.
- 2:26:43Okay. So, you should have something that
- 2:26:44looks roughly like this. I don't have
- 2:26:47anything open. I don't have any files
- 2:26:49open. Uh just kind of have a blank
- 2:26:52screen here. Um, but if you I would
- 2:26:55recommend uh using the AI features if
- 2:26:58you can. Um, I think that'll come in
- 2:27:01handy later on.
- 2:27:04Um, are we
- 2:27:07comfortable uh moving forward? I want to
- 2:27:09show us the extensions that support
- 2:27:11Python. So, right now when you first
- 2:27:14when you first install this, it does not
- 2:27:18work with Python out of the box. We have
- 2:27:20to install a couple extensions inside of
- 2:27:23here to get it to work with Python. I'm
- 2:27:25going to show us how to do that.
- 2:27:30Don't worry about tuning any settings.
- 2:27:32No, don't worry about doing any of that
- 2:27:34at this stage. Don't really need to tune
- 2:27:37anything. We just need to get Python
- 2:27:40support.
- 2:27:47Okay.
- 2:27:49So, you guys with me on this main page?
- 2:27:58You can use your corporate. Sure. Sure.
- 2:28:00Yeah, you can you if you have it. If you
- 2:28:02have co-pilot and want to use your
- 2:28:04corporate, you can use that. That's
- 2:28:05fine.
- 2:28:08But you guys are with me on the main
- 2:28:09page because I'm about to show us uh I'm
- 2:28:12about to show us the extensions we need
- 2:28:14to install to work with Python.
- 2:28:17Okay, really important because this
- 2:28:20isn't this is not in the documentation.
- 2:28:34Um, no, no need to reinstall. Um, you
- 2:28:39can I'll show you how to add that
- 2:28:41through the extensions. No need to
- 2:28:43reinstall.
- 2:28:46You can add it as an extension. Yeah.
- 2:28:52Okay.
- 2:28:54So, let me ask you guys on the left,
- 2:28:58do you see
- 2:29:01this
- 2:29:03little box icon that if you hover over
- 2:29:08it says extensions?
- 2:29:12Do you see that? you. There may be other
- 2:29:14things here too, but at least that one
- 2:29:17with the extensions.
- 2:29:26Okay, so we do see that one. Okay,
- 2:29:32so what we want to do,
- 2:29:38no, I wouldn't I wouldn't uninstall.
- 2:29:40That's okay because we're actually gonna
- 2:29:42install Anaconda to get Jupiter. I
- 2:29:45wouldn't un I wouldn't I would cancel
- 2:29:46that if you can because you're going to
- 2:29:48want that for Jupiter as well. I
- 2:29:51wouldn't uninstall Anaconda.
- 2:29:54I wouldn't uninstall. But I mean, if
- 2:29:56it's already going if it's already doing
- 2:29:57it, that's okay. We'll just reinstall it
- 2:29:59later. All right. So, back to the
- 2:30:01extensions. So, let's click on the
- 2:30:03extensions.
- 2:30:09Okay. So, do we see something like this
- 2:30:11that has a search bar for extensions?
- 2:30:16Do we see the search bar for the
- 2:30:18extensions?
- 2:30:21Okay. What do you think? We're going to
- 2:30:23search for
- 2:30:26Python.
- 2:30:28Python.
- 2:30:30We're going to search for Python. Yeah.
- 2:30:33So, you are going to want to install the
- 2:30:36official Python extension from
- 2:30:39Microsoft. It is this one that has the
- 2:30:41blue check mark next to Python.
- 2:30:45Uh, so there now there are other ones
- 2:30:49here,
- 2:30:50but we just want the one that says
- 2:30:53Python
- 2:30:55from Microsoft. Do we see that extension
- 2:30:57when you type in Python? Do we see that
- 2:31:00one?
- 2:31:02So just so it should just say Python. It
- 2:31:05should be Microsoft.
- 2:31:07Uh it's really popular. It has a lot of
- 2:31:10downloads. Over 192 million downloads as
- 2:31:13an extension.
- 2:31:15It's from Microsoft.
- 2:31:18Okay. Click on that.
- 2:31:21Click on that.
- 2:31:24And then you should see an install
- 2:31:26button. It I already have it installed.
- 2:31:28So it says uninstalled. Right here there
- 2:31:29should be an install button. Install the
- 2:31:32Python extension.
- 2:31:35So out of 192 million installs,
- 2:31:39really popular extension.
- 2:31:44Are you guys able to install it?
- 2:31:47You want to install that? It should be
- 2:31:50pretty quick.
- 2:31:55It should be pretty quick. It's not that
- 2:31:57big of an extension.
- 2:32:02So, what this does is
- 2:32:07just the Python Sherry. It's just a
- 2:32:09Python one. If you go into the
- 2:32:11extensions and then search for Python,
- 2:32:14it is just the one. It's just this one
- 2:32:15that says Python and it's from
- 2:32:17Microsoft.
- 2:32:19Python blue check mark Microsoft.
- 2:32:22You want that one.
- 2:32:25And then you want to click on that one
- 2:32:27and then hit the install.
- 2:32:43Um, Roberto, is that for a co-pilot?
- 2:32:51Is that for a co-pilot? I
- 2:32:54maybe try closing it and reopening it.
- 2:32:57Try closing VS Code, reopening and
- 2:32:59retrying the install.
- 2:33:08Um, no, we're not opening any folders
- 2:33:10right now. We're not opening it. We're
- 2:33:12just installing the extension.
- 2:33:14That's all. We're just installing the
- 2:33:15extension.
- 2:33:18We're not opening any project folders.
- 2:33:22just installing the extension.
- 2:33:25Were we were we able to install that?
- 2:33:46I know there's a lot by Microsoft, but
- 2:33:48there should just be one that that says
- 2:33:50Python.
- 2:33:54there. So see how the name like this
- 2:33:57name is this name here is eyesore. This
- 2:33:59name is Python debugger. This one is
- 2:34:01pilance.
- 2:34:04Just the one that says Python.
- 2:34:10Just that one.
- 2:34:12That's the one we want. Only that one
- 2:34:14right now.
- 2:34:19Okay. Perfect. Perfect.
- 2:34:22Okay, great.
- 2:34:29Okay, so one more extension for you
- 2:34:32guys. So once you install that one, I
- 2:34:34have one more for you that you want to
- 2:34:35install.
- 2:34:38Are we ready for that one? One more we
- 2:34:41want to install.
- 2:34:45Okay, we're ready for the next one. So,
- 2:34:48the next one you want to install
- 2:34:51is the Jupiter extension,
- 2:34:58which is the Jupiter.
- 2:35:01It's this one. It's the very first one
- 2:35:03here on my screen. So, it's it says
- 2:35:05Jupiter
- 2:35:06and it's from Microsoft.
- 2:35:11Okay, we want to install that one.
- 2:35:14Jupiter and it's from Microsoft. want to
- 2:35:17install that one.
- 2:35:21So, this one has 98 million uh installs.
- 2:35:26You want to install this one.
- 2:35:29Did you guys find that one? So, you want
- 2:35:31to type in Jupy
- 2:35:34Ter and it should be the Jupiter
- 2:35:39extension here
- 2:35:42that is uh from Microsoft.
- 2:35:46So you want to install that one.
- 2:35:51Great.
- 2:35:53Now what does this one do? This
- 2:35:55extension will allow you to work with
- 2:36:00Jupiter notebooks inside of VS Code if
- 2:36:03you want to.
- 2:36:06So you Jupiter notebook has its own
- 2:36:10standalone program which we will look at
- 2:36:11next.
- 2:36:14But you can open you can have those
- 2:36:17files, those Jupyter notebook files be
- 2:36:19compatible with VS Code and open them
- 2:36:21and edit them and run them inside of VS
- 2:36:23Code if you want to. So this extension
- 2:36:27gives you the flexibility to work with
- 2:36:29notebooks inside of VS Code. So you
- 2:36:31never have to leave VS Code if you want
- 2:36:32to work with notebooks. Um,
- 2:36:35so this is a good extension if you
- 2:36:38really want to work with notebooks and
- 2:36:39stay inside of VS Code.
- 2:36:48Yes. Uh when you Yeah. When you install
- 2:36:51install an extension, it might it might
- 2:36:53install a couple other dependency
- 2:36:55extensions. Yes. But that's okay. Those
- 2:36:57are required. That's okay.
- 2:37:01That's that's that's okay.
- 2:37:07All right. How do we feel? Good. Uh did
- 2:37:09we get those installed?
- 2:37:12Did we do were we able to get those
- 2:37:13installed?
- 2:37:16Okay, here is how we will test that it
- 2:37:20all worked. So, we're going to do
- 2:37:22something really simple.
- 2:37:25Here's how we will test that it worked.
- 2:37:30Let me go out of here and back to our
- 2:37:32files.
- 2:37:35So, out of the extensions, I just went
- 2:37:37to the top button where it's the little
- 2:37:39file um icon and um I am going to
- 2:37:46um
- 2:37:47go up to the very very top where it um
- 2:37:50so you guys see on your VS Code window
- 2:37:53where it says file, edit, selection,
- 2:37:55view. I'm just going to create um
- 2:38:00I'm just going to create uh a new
- 2:38:04new file.
- 2:38:08So, do you guys see that where where you
- 2:38:09say file edit selection view? Click on
- 2:38:12file and then click on new file.
- 2:38:19You should see what I see on this screen
- 2:38:21right here.
- 2:38:25If you see
- 2:38:28if you see Python and Jupyter notebook
- 2:38:31then you know those are installed
- 2:38:33correctly.
- 2:38:34Do you guys see these options text
- 2:38:36Python and Jupyter notebook?
- 2:38:42Great. So what that means is we we can
- 2:38:44now create those kind of files in the
- 2:38:47future. We can create notebooks. you can
- 2:38:50create Python files and VS Code will be
- 2:38:52able to work with those.
- 2:38:57If you don't see Python, that means your
- 2:38:59Python extension didn't install yet or
- 2:39:03you didn't install it. So, you want to
- 2:39:04go back to you want to go back to your
- 2:39:08extensions and make sure you installed
- 2:39:09Python.
- 2:39:11So go go go to this button over here,
- 2:39:13the extensions,
- 2:39:17type in Python,
- 2:39:21and then make sure you install this
- 2:39:23Python extension.
- 2:39:29Okay. So, you're going to install the
- 2:39:32Python extension and you're going to
- 2:39:35install the Jupiter extension,
- 2:39:40which is this, and install both of
- 2:39:42those.
- 2:39:44Make sure those are installed. If
- 2:39:46they're installed and you still didn't
- 2:39:48see that when you went to file um new
- 2:39:51file,
- 2:39:53if you don't see those, then um try
- 2:39:58exiting VS Code and relaunching it.
- 2:40:02Okay? Try exiting VS Code and reopening
- 2:40:04it and seeing if you can make a new
- 2:40:06file.
- 2:40:11Okay? But it should be under uh at the
- 2:40:14top file and then new file
- 2:40:19and then you should see those options
- 2:40:21Python and Jupiter.
- 2:40:25Once you have those extension installed,
- 2:40:26you may need to close out of VS Code and
- 2:40:29reopen it to see that.
- 2:40:40Okay, perfect. after you relaunched.
- 2:40:42Okay, perfect. Yeah, you may need to
- 2:40:44relaunch so that it can show the it can
- 2:40:47show the extensions.
- 2:40:51Yeah,
- 2:40:54perfect.
- 2:40:57Okay, perfect. So, that's set up for you
- 2:40:59guys. So, um Perfect. It's set up for
- 2:41:02you guys. Uh we will work with it in the
- 2:41:05future, but just wanted to make sure it
- 2:41:07was installed and set up. Once we start
- 2:41:09working with Python, um I will show you
- 2:41:12guys how to how to work with it. Um but
- 2:41:15glad that's set up for now.
- 2:41:23Uh what issue are you having uh Romero?
- 2:41:28Is it not showing? It's not showing
- 2:41:29Python or Jupiter for you when you do
- 2:41:31file new file.
- 2:41:34It's not showing those.
- 2:41:37You may need to exit VS Code and reopen
- 2:41:40it.
- 2:41:48You uh Sil, yeah, you can you can make
- 2:41:51one. We're not going to do anything with
- 2:41:52it right now.
- 2:41:56It's not going to you're not going to do
- 2:41:57anything with it right now, but um
- 2:42:06it's make sure you're searching for it
- 2:42:08with a Y. It's J U P Y T E R.
- 2:42:14You have to search. You have to So when
- 2:42:16you go when you click on the extension,
- 2:42:19search for JUP
- 2:42:23Y. It should be the first thing that
- 2:42:25shows up with Jupy Ter.
- 2:42:29It's this Jupiter one from Microsoft.
- 2:42:36I kernel I'll So the let me show us let
- 2:42:39me show us that later. The kernel you
- 2:42:41have to um you have to have a Python
- 2:42:43interpreter.
- 2:42:46So you may need to install a Python
- 2:42:48interpreter to to be able to run the
- 2:42:50kernel. So, I need to show us that. Um,
- 2:42:54but I I don't want to get into that
- 2:42:55right now.
- 2:43:02Save what to
- 2:43:10Oh, wherever you want. Wherever you want
- 2:43:13on your own machine. It's up to you. It
- 2:43:15doesn't really matter. Just wherever you
- 2:43:17want.
- 2:43:26It doesn't matter. It's up to you.
- 2:43:29All right. So, what I want to do is uh I
- 2:43:33want to take a break. Um because now,
- 2:43:36you know, I said after two hours, we'll
- 2:43:39take a longer break. Um so, we will now
- 2:43:43we'll take a 10-minute break. Now, um if
- 2:43:47you're still having any issues, um we
- 2:43:49can try to get you set up at the end of
- 2:43:51class. Um but we are going to set up.
- 2:43:54So, coming up after our break, we're
- 2:43:56going to take a 10-minute break. Coming
- 2:43:57up after that, we'll we'll go and
- 2:44:00install Jupyter Notebook. And then after
- 2:44:03that, we will look at Collab. So, you're
- 2:44:05going to have multiple options to run
- 2:44:07Python. Not So, if this wasn't working
- 2:44:09for you, that's okay. We'll try a
- 2:44:11different route.
- 2:44:13Okay? I will try a different route. Um I
- 2:44:16I know Collab will work for you because
- 2:44:18that is hosted by Google and really easy
- 2:44:21to get working with. So at the worst
- 2:44:23case scenario, Collab will work for you.
- 2:44:25I know it. Um but we'll try to get
- 2:44:28Jupyter Notebooks installed for you as
- 2:44:30well. But if you're having issues with
- 2:44:31VS Code, let me know at the end of
- 2:44:33class. We'll try to get you set up,
- 2:44:35okay?
- 2:44:37You're still having issues with it.
- 2:44:42But um what we're going to do right now
- 2:44:43is take take a 10-minute break.
- 2:44:48So let's try to be back um in about uh
- 2:44:5210 minutes. Let's call it an even um
- 2:44:56let's call it an even
- 2:45:01uh what will we be covering? Um
- 2:45:03installing the other installing the
- 2:45:05other um Python setups. So Jupyter
- 2:45:07notebook and working with collab. And
- 2:45:10then we will get into the basics of
- 2:45:11Python's the syntax. So we're going to
- 2:45:13talk about indentation, identifiers, um
- 2:45:16maybe if we have time, basic variable
- 2:45:18types, data types. Yep. So we'll get
- 2:45:20into Python.
- 2:45:22We will get into Python today.
- 2:45:25All right. So let's jump over to
- 2:45:28uh Jupiter notebooks. So um what's so
- 2:45:32special about Jupiter?
- 2:45:35Well, it turns out that uh Jupiter is a
- 2:45:39platform for running what are called
- 2:45:42notebook files. So obviously we just
- 2:45:44installed the Jupiter extension in VS
- 2:45:46Code which will allow us to run
- 2:45:48notebooks in VS code but Jupiter has its
- 2:45:52own notebook platform and that's what
- 2:45:54you will install in this setup. Um,
- 2:45:58notebooks are special. They are um
- 2:46:01really great um Python code files that
- 2:46:06give us the ability to execute isolated
- 2:46:09what are called cells of code. So we can
- 2:46:13run one cell at a time and test and
- 2:46:16debug the execution of that single cell
- 2:46:19without affecting any of the other
- 2:46:21cells. So, um, notebooks are great for,
- 2:46:26uh, running code live and interactive.
- 2:46:29When we do a lot of our demos in this
- 2:46:31program, they're all going to be in
- 2:46:33notebooks. Um, so that we can kind of
- 2:46:36run things one cell at a time. Um,
- 2:46:41uh,
- 2:46:42no. So without notebooks you either have
- 2:46:45to run you run like a Python script like
- 2:46:49a Python file um which is a py file and
- 2:46:54usually you have to either run that
- 2:46:55through a debugger or run the entire
- 2:46:58script at once. You don't really get
- 2:47:01code isolated into individual cells
- 2:47:03which is really nice with notebooks. The
- 2:47:06other thing is notebooks are easily
- 2:47:08sharable.
- 2:47:10So you can share a notebook with
- 2:47:12somebody else and they can open it and
- 2:47:13see all of your inputs and outputs in
- 2:47:16the notebook which is really nice like
- 2:47:17all of the outputs get saved into the
- 2:47:20notebook. Um which is nice. So and
- 2:47:25notebooks uh especially in the Jupiter
- 2:47:28platform are going to have all the data
- 2:47:30science libraries available to them. So,
- 2:47:32uh, if you're people usually love doing
- 2:47:34notebooks for working with data, um,
- 2:47:38really easy to work with data inside of
- 2:47:40notebooks and and build things like
- 2:47:42plots. You can display your,
- 2:47:45uh, you can display your graphs really
- 2:47:48easily inside of the notebook and then
- 2:47:50share your notebook so other people can
- 2:47:52see your graphs. Um, so notebooks are
- 2:47:55really awesome like interactive
- 2:47:57environments for running code. Um and we
- 2:48:01will favor notebooks uh as our primary
- 2:48:04way of running code throughout the
- 2:48:06program. Now where you open those
- 2:48:08notebooks is up to you. You can open
- 2:48:11them in VS Code. You can open them in
- 2:48:12the Jupyter notebook platform. Uh
- 2:48:16you can open them inside of Collab and
- 2:48:19run notebooks in Collab. Uh notebooks
- 2:48:23are very very popular.
- 2:48:28Why isn't running in notebooks the
- 2:48:29default? It's because uh not all code
- 2:48:32runs inside of cells. Like applications
- 2:48:34are not going to be well suited for
- 2:48:36notebooks. Like Instagram is not running
- 2:48:39in a notebook. Uh it's more structured
- 2:48:41into actual Python files and actual uh
- 2:48:45more structured programs are going to be
- 2:48:47not in a notebook. Notebook is more for
- 2:48:51prototyping and debugging and uh
- 2:48:55executing small chunks of code to test
- 2:48:58it out. It's not for writing larger
- 2:49:00programs like an like a
- 2:49:03an LLM application like a chatbot would
- 2:49:06generally be in not in a notebook. It'd
- 2:49:08be in like a Python file.
- 2:49:15Uh cells versus class objects. So cells
- 2:49:17are just small uh think of them as small
- 2:49:22little environments to execute our code.
- 2:49:24Um class objects are actual chunks of
- 2:49:27code that define an object. They're
- 2:49:29they're different things.
- 2:49:34Yeah, different things. We'll we'll
- 2:49:35learn about objects. Um and we will
- 2:49:38certainly see what cells are as we go
- 2:49:40through. I'm going to show you an
- 2:49:41example of a cell coming up when we
- 2:49:43install Jupiter.
- 2:49:45But uh let's talk about let's uh go
- 2:49:48through the installation of Jupyter
- 2:49:50notebook so you can see what a notebook
- 2:49:52looks like. I think that'll be helpful
- 2:49:54to orient.
- 2:49:58So let's go over to that demo. So this
- 2:50:01is going to be demo two
- 2:50:04uh demo two inside of um uh lesson one.
- 2:50:09So we're going to go over to that.
- 2:50:17Everybody has this one. Okay, perfect.
- 2:50:19Okay, so you're going to follow this
- 2:50:20instruction. Now, what this is going to
- 2:50:22do is first
- 2:50:25um
- 2:50:27No, this has not this is not going to be
- 2:50:29anything to do with VS Code. This is
- 2:50:31going to be a different platform. This
- 2:50:32is going to be Jupiter.
- 2:50:36Where is this? This is the
- 2:50:38This is the demos.
- 2:50:41This is uh demo two inside of that demos
- 2:50:44folder that we said uh
- 2:50:48to to uh grab all the demos
- 2:50:52from your LMS.
- 2:50:56Does anybody have that uh demo 2 PDF
- 2:50:59they can upload? I I think somebody
- 2:51:01uploaded all of them earlier, but if you
- 2:51:02have demo two, want to upload it real
- 2:51:05quick? I don't have the PDFs.
- 2:51:11if somebody wants to share that.
- 2:51:15So there so they're different. Um VS So
- 2:51:18what I was saying is you can open
- 2:51:21notebooks inside of VS Code and the
- 2:51:23thing that allows you to open notebooks
- 2:51:25in VS Code is the extension.
- 2:51:28So yes, if you're going to work with
- 2:51:29notebooks in VS Code, you need the
- 2:51:30extension installed. But you can use the
- 2:51:34standalone Jupiter platform
- 2:51:37to work with notebooks. It's up to you.
- 2:51:40If you like using VS Code,
- 2:51:43um if you like using VS Code, you can do
- 2:51:45it that way. If you like uh the Jupiter
- 2:51:48platform, you can do it that way. It's
- 2:51:51up to you. It's just a preference. I'm
- 2:51:54giving you guys options. That's my goal
- 2:51:57is to give you options and let you guys
- 2:51:59choose what you're most comfortable
- 2:52:00with.
- 2:52:04Okay. And we're and we're taking time to
- 2:52:06do that now in the beginning of the
- 2:52:08program, right? Because we're going to
- 2:52:10be doing a lot of Python examples coming
- 2:52:12up as we start learning Python. So, it's
- 2:52:15it's valuable to spend that time now. I
- 2:52:17know it can seem a little slow, but I
- 2:52:21promise it'll be worth it so that you
- 2:52:22guys have options for running your
- 2:52:24running your code.
- 2:52:29Yes. Thank you guys for uploading those.
- 2:52:31appreciate it. Those are the demos you
- 2:52:34want to uh follow along with.
- 2:52:38Okay, so the first step here is going to
- 2:52:40be to install Anaconda. Now, you may be
- 2:52:43wondering, what is Anaconda? I thought
- 2:52:44we were talking about Jupiter, and
- 2:52:47that's a valid question. Anaconda is a
- 2:52:51what's called a distribution of Python.
- 2:52:55So Anaconda is a program a software a
- 2:52:59collection of software programs that
- 2:53:02give you a version of Python with a
- 2:53:05bunch of packages
- 2:53:08uh with a bunch of packages already
- 2:53:11installed.
- 2:53:13Um and then
- 2:53:17uh one of those is the Jupiter package
- 2:53:21so that you can run Jupyter notebooks.
- 2:53:24And what Jupyter notebooks will be
- 2:53:28is a uh basically a web browser
- 2:53:32application that will open up a notebook
- 2:53:37editor in your web browser. So that's
- 2:53:40ultimately what we're going to do, but
- 2:53:42we are going to install it via the
- 2:53:44Anaconda distribution
- 2:53:47uh via the Anaconda distribution of
- 2:53:50Python.
- 2:53:51So that's where we're going to start is
- 2:53:53with the initial download of Anaconda.
- 2:53:58Oh, it's no no skipped registration.
- 2:54:01Okay, let me let me uh open the link.
- 2:54:04I think there I think there's a way to
- 2:54:06find it without having to do the
- 2:54:07registration.
- 2:54:11There's a way to get to it without
- 2:54:12having to do that. I'm going to find it
- 2:54:14real quick.
- 2:54:21Oh, you can't. Okay. So, if you can't
- 2:54:22install it, that's okay. We will be able
- 2:54:24to work with notebooks in collab and you
- 2:54:27can work with notebooks inside of VS
- 2:54:29Code. That's fine, too.
- 2:54:38Yeah, I'm getting I I'm going through
- 2:54:40the registration process so I can um I
- 2:54:42can show you that install.
- 2:54:46Okay, let me share my screen.
- 2:54:50Did you guys get to once you go through
- 2:54:52the like setting up your account, do you
- 2:54:55get to this page?
- 2:55:00Do you get to this page for those of you
- 2:55:02going through? Yeah, that looks right
- 2:55:04for you, Ashish. That looks right.
- 2:55:11Do you guys get to this page though when
- 2:55:13you get through your like account setup?
- 2:55:16Okay, you got to this page. Okay, so
- 2:55:18then choose your correct Windows or Mac
- 2:55:21down. You want to be over here on the
- 2:55:23left. You want to do Anaconda
- 2:55:25distribution.
- 2:55:27This is what you want to do. So, choose
- 2:55:30the right one. And if you're on an M1,
- 2:55:31M2, M3, you're going to do the silicon.
- 2:55:35If you're on an older Mac, you're going
- 2:55:37to do the 64. And then obviously, if
- 2:55:39you're on a Windows, you should be
- 2:55:40clicking over here to do Windows. But
- 2:55:43you want to do the Anaconda
- 2:55:44distribution, not Minion. Okay. So,
- 2:55:47click on the installer for Anaconda
- 2:55:51distribution.
- 2:55:56Okay? And then let that install. Now
- 2:56:00while that's installing let me explain
- 2:56:02something about the difference between
- 2:56:05uh I think it was asked earlier what's
- 2:56:07the difference between um Anaconda
- 2:56:11uh as the default Python. So Anaconda
- 2:56:16as I was saying earlier is a version of
- 2:56:19Python that has a bunch of data science
- 2:56:22and machine learning packages already
- 2:56:24installed for you. So uh it comes with a
- 2:56:29bunch of packages that are already
- 2:56:31installed. So if you use that Python
- 2:56:35um that Python has a bunch of packages
- 2:56:38built in with it that you don't need to
- 2:56:39go out and install. So, Anaconda is a
- 2:56:42very popular version of Python for
- 2:56:45people to install that are working in
- 2:56:46data science, AI, ML. Very popular
- 2:56:50version because it already comes with a
- 2:56:52bunch of packages that you would use for
- 2:56:54manipulating data for doing machine
- 2:56:57learning or doing anything in AI. So,
- 2:57:00it's it's a very um popular
- 2:57:02distribution. It also comes with
- 2:57:06Jupiter, which is why we wanted to use
- 2:57:09it because it comes with the notebook
- 2:57:11capability out of the box.
- 2:57:16Okay. So, I'm going to launch. So, when
- 2:57:19this is done installing, you want to
- 2:57:22launch the program that gets installed
- 2:57:25called the uh Anaconda Navigator.
- 2:57:29So, it should install a program on your
- 2:57:31machine called the Anaconda Navigator.
- 2:57:33Do you guys have that? Did anybody get
- 2:57:35through and and you have that program?
- 2:57:38The Anaconda Navigator.
- 2:57:41You don't need any advanced ones. You
- 2:57:43don't need any advanced options.
- 2:57:47Just the just the defaults. All the
- 2:57:49defaults
- 2:57:51should be good.
- 2:57:57Still downloading. Okay. I'm going to
- 2:57:58show you
- 2:58:00I'm going to show you what the navigator
- 2:58:02looks like once you once you have it.
- 2:58:06That's okay if it takes a little bit of
- 2:58:07time to download. That's okay. Um,
- 2:58:09basically once you download it, um, you
- 2:58:12just have to click a couple more buttons
- 2:58:13and then you can access Jupiter.
- 2:58:17Okay, let me share my screen and show
- 2:58:20you what you like. Once it installs,
- 2:58:23this is what it should look like. It's
- 2:58:24okay if it's taking a little bit of
- 2:58:25time.
- 2:58:27You should have something that kind of
- 2:58:30looks like this, which is the um
- 2:58:32dashboard that has the different
- 2:58:35programs available to you to you.
- 2:58:39Um
- 2:58:41do you guys see something like this? If
- 2:58:44you have the navigator,
- 2:58:47do you see something like this?
- 2:58:52which is the which is the like when you
- 2:58:54open the navigator program, you should
- 2:58:56see something like this that has a bunch
- 2:58:58of different um
- 2:59:04you do. Okay.
- 2:59:07It's if it's taking a little bit of time
- 2:59:08that's okay.
- 2:59:11Yes. Na Anaconda Navigator is how you
- 2:59:13launch Yes.
- 2:59:17Anaconda Navigator is how you launch it.
- 2:59:19So yeah, you want to open that. Now the
- 2:59:21the whole reason to come here
- 2:59:24is so we can launch Jupiter notebooks.
- 2:59:28So we can launch Jupyter notebooks. Um
- 2:59:33this is the program we ultimately want
- 2:59:34to launch. This is going to
- 2:59:38uh allow us to open notebook files, edit
- 2:59:41them, run code cells. I'll show you what
- 2:59:44a notebook looks like in a second. But
- 2:59:47but once you have Anaconda installed,
- 2:59:51open the navigator and then launch
- 2:59:54Jupiter notebook. It's just one extra
- 2:59:56step. Launch the Jupiter notebook.
- 3:00:00What that should do is launch the the
- 3:00:04notebook.
- 3:00:06Uh it should launch the web browser of
- 3:00:11your like whatever you have as your
- 3:00:12default web browser. It should open the
- 3:00:14notebook program in every in your web
- 3:00:17browser. So if it's Chrome, Firefox,
- 3:00:19Edge, whatever your default web browser,
- 3:00:21it's going to launch the notebook
- 3:00:23program in the browser.
- 3:00:27Okay.
- 3:00:31So, I'm going to launch it and then I'll
- 3:00:34show you what it looks like. Again, if
- 3:00:35it's taking you a little bit of time,
- 3:00:36that's okay. Whenever it's done,
- 3:00:41how did I get to these icons? Just
- 3:00:43launch. Do you have the Anaconda
- 3:00:44Navigator program?
- 3:00:46It should have got It should be
- 3:00:48installed.
- 3:00:49Open the Anaconda Navigator program.
- 3:00:53It should have been installed with the
- 3:00:54Anaconda installation.
- 3:01:05All right. Was anybody able to get to
- 3:01:06this the Jupiter this? So, it should
- 3:01:09launch in your browser. Anybody
- 3:01:13able to get to that?
- 3:01:16Fantastic. Fantastic. I'm glad some of
- 3:01:18you guys are able to get to it. And if
- 3:01:20it's it's not yet, that's okay.
- 3:01:22Remember, when it's done installing,
- 3:01:23you're going to go to Anaconda Navigator
- 3:01:26and then
- 3:01:28uh launch Jupiter Notebook. That's
- 3:01:32That's what you're going to do.
- 3:01:34That's okay, Roberto. It's okay.
- 3:01:37All right. I do want to I want to show
- 3:01:39you guys a notebook. I just want to show
- 3:01:42you what it looks like. What I'm going
- 3:01:44to do is I'm going to
- 3:01:48um open a notebook by going to new and
- 3:01:52then Python 3 notebook. So you can open
- 3:01:55a folder, you can open a terminal, you
- 3:01:57can open a text file. I'm going to open
- 3:01:59a Python 3 which is a notebook. You so
- 3:02:02it's a it's a a notebook powered by
- 3:02:04Python.
- 3:02:06So I'm going to click on that which will
- 3:02:09launch a new notebook in a new tab.
- 3:02:14And here I am in the notebook editor. So
- 3:02:17now I am in a notebook editor screen. So
- 3:02:20if you go when you first launch Jupiter
- 3:02:23you you can navigate to notice that that
- 3:02:26notebook got created here where I
- 3:02:28currently am on my machine. I could
- 3:02:30navigate to I could navigate to
- 3:02:33documents and then I could um you know
- 3:02:37create a new file there or I could make
- 3:02:39a new folder here and and do it that
- 3:02:42way. Um but uh I am uh just editing this
- 3:02:49notebook right here within this um
- 3:02:51current folder that I'm in.
- 3:02:58Okay. So, do you guys remember when I
- 3:03:00said that code gets executed in a in an
- 3:03:02isolated cell?
- 3:03:05Do you remember that?
- 3:03:08Um,
- 3:03:09this is what a cell looks like. And you
- 3:03:13can make new cells by hitting this plus
- 3:03:15button.
- 3:03:17So, you hit this plus button over here,
- 3:03:19you can make new cells. So, if you hit
- 3:03:23plus++,
- 3:03:25I'm making a bunch of cells.
- 3:03:27Now, what's really cool about cells?
- 3:03:30Yeah, Tim just discovered this. What's
- 3:03:32really cool about cells is you can
- 3:03:34change them to be text or code. So, if
- 3:03:38you change it to markdown,
- 3:03:41I can write markdown text in here to say
- 3:03:44this is my notebook. And then if I run
- 3:03:48this, it's going to display as text.
- 3:03:52So if I run that cell which uh when I'm
- 3:03:56editing it I can click run and it will
- 3:04:00render that as text because I changed
- 3:04:03the cell type to markdown. Markdown is
- 3:04:06just a flavor of text style.
- 3:04:11So otherwise we can write some Python
- 3:04:14code. Now, what I want you guys to type,
- 3:04:16I'll type this in the chat to verify
- 3:04:19everything is working is I want you to
- 3:04:21type print
- 3:04:24hello world.
- 3:04:27I want you to type that
- 3:04:30inside of a cell
- 3:04:35and then
- 3:04:37and then hit run.
- 3:04:43And it should run that code.
- 3:04:47And you should see you should be able to
- 3:04:49see uh you should be able to see that
- 3:04:58Shift enter. Yep. You can whenever
- 3:04:59you're on a cell, you can hit shift
- 3:05:01enter. It'll run the cell.
- 3:05:11You can That's okay. You can always You
- 3:05:12can go back and watch the video. So,
- 3:05:15this is being recorded. You can go back
- 3:05:16and watch the video. I I know it's a
- 3:05:19little frustrating. It's still
- 3:05:20installing for you, but go back and
- 3:05:23watch the video. And I definitely
- 3:05:24encourage once it's installed to go back
- 3:05:26and try this, which would be just
- 3:05:30launching your Anaconda Navigator
- 3:05:33and then launching Jupiter.
- 3:05:41If you don't have, by the way, if you
- 3:05:43don't have Python 3, um you may need to
- 3:05:48uh exit your navigator and reopen it.
- 3:05:53Okay, you may need to exit your your
- 3:05:55navigator, reopen it so that you can
- 3:05:56launch Jupiter again.
- 3:06:05Were you guys able to run this in a
- 3:06:07cell? For those of you that have Jupiter
- 3:06:08running, were you able to run this?
- 3:06:14Nice. And it worked for you. Okay,
- 3:06:16perfect. Perfect.
- 3:06:19So this is what I meant by this is an
- 3:06:22isolated cell. So notice that we can run
- 3:06:24this
- 3:06:26and it doesn't affect
- 3:06:30um
- 3:06:32Sure. Sure. I hear you. I I hear you.
- 3:06:34Update the doc. Uh I can How about I
- 3:06:38post it in our um our Slack channel? By
- 3:06:42the way, do you guys have access to the
- 3:06:44to the Slack channel?
- 3:06:51Okay, I can post it there. I can post
- 3:06:53the instructions to get there.
- 3:06:59Okay, I can post it in our our uh
- 3:07:01cohort's uh channel.
- 3:07:05I hear you. You don't want to search for
- 3:07:07our video. I I hear you.
- 3:07:16Uh I don't have the link on hand, but
- 3:07:19you can get to it through the LMS.
- 3:07:22So if you go to the LMS and go to
- 3:07:27uh there should be
- 3:07:31um
- 3:07:33there should be a link to get to it
- 3:07:35within there. It's should be like over
- 3:07:38here on the right.
- 3:07:43I don't have the link I don't have the
- 3:07:45link to it off hand. Yeah,
- 3:07:51but there should be a way to get to it
- 3:07:53from the LMS.
- 3:07:56Yeah, there should be a banner here. I
- 3:07:59don't know why I don't have it, but
- 3:08:01should be there.
- 3:08:05Okay. So, if you're just getting things
- 3:08:08installed, how do you get to here? Um,
- 3:08:11you open the navigator.
- 3:08:15Open the navigator.
- 3:08:18Syntax is hello world.
- 3:08:22It's just inside of it's just that print
- 3:08:26hello world.
- 3:08:30Um, open the navigator.
- 3:08:34Open the navigator which looks like
- 3:08:36this.
- 3:08:38Let me share my screen.
- 3:08:44Okay. Open the Anaconda Navigator that
- 3:08:47got installed.
- 3:08:48Then click launch on the Jupyter
- 3:08:52notebook program. So you should have
- 3:08:54this at least. You may have other ones.
- 3:08:57Click on this launch. Uh click on this
- 3:09:01launch and then you can launch the uh
- 3:09:04Anaconda Navigator.
- 3:09:14Okay.
- 3:09:16If it if it's a little stuck, that's
- 3:09:17okay. We're going to move on. We're
- 3:09:19going to go to Collab, which can run
- 3:09:21notebooks as well. So, if it seems a
- 3:09:23little stuck, that's okay. I will post
- 3:09:26in our Slack instructions on how to run
- 3:09:28this.
- 3:09:31That's okay.
- 3:09:36All right. But what I wanted to do,
- 3:09:39what I wanted to do before we move on to
- 3:09:41collab is I just wanted to show you I
- 3:09:43wanted to call out a couple things about
- 3:09:45notebooks.
- 3:09:47Um
- 3:09:49is that uh a couple things about
- 3:09:52notebooks. One is that notice that these
- 3:09:54cells are very isolated. Whatever I put
- 3:09:56here
- 3:09:57um does not affect what I had before. So
- 3:10:01I can add numbers like that and it can
- 3:10:04um compute that and this does not affect
- 3:10:07this. So so this is why notebooks are so
- 3:10:09great is you can document the notebooks
- 3:10:12with with mixing in text and code like
- 3:10:16we do here. Um you can run code in its
- 3:10:19own cells.
- 3:10:21uh you can run code in its own cells and
- 3:10:24then you can um have that very isolated.
- 3:10:26So I could jump down here and run
- 3:10:28something and that doesn't matter that
- 3:10:30there's nothing here like it doesn't
- 3:10:32need to be in order. I can um you know I
- 3:10:36can uh run stuff out of I can overwrite
- 3:10:38this
- 3:10:42um and run that and it produces the
- 3:10:44output. Uh
- 3:10:48so you know many things we could do uh
- 3:10:53inside of notebooks that are really
- 3:10:54fantastic for just quickly prototyping
- 3:10:57and running Python code inside of cells.
- 3:11:00So it's very nice that way. So notebooks
- 3:11:03are notebooks are really nice. You can
- 3:11:04also like I could share this file. So
- 3:11:07this produces a file on my machine. Uh
- 3:11:10if I go back to the navigator
- 3:11:13um
- 3:11:15Anaconda navigator Jerry Anaconda
- 3:11:18navigator
- 3:11:22um
- 3:11:24can you read value of variables from
- 3:11:26another cell? Uh you you have to store
- 3:11:29them into variables. So I could I could
- 3:11:31call this uh x
- 3:11:36and then I could refer to x later.
- 3:11:39We'll learn about that. We'll learn
- 3:11:40about that with variables.
- 3:11:44But yes, you can kind of do that with
- 3:11:46variables.
- 3:11:53All right.
- 3:11:56Um,
- 3:12:02what do the numbers after?
- 3:12:07Which numbers? these the ones in the
- 3:12:09brackets.
- 3:12:16Oh th so those are which cells we've uh
- 3:12:20executed. So I executed this one first.
- 3:12:23So it it's it's number one. And then I
- 3:12:26executed um
- 3:12:29uh this I think I did this second. So it
- 3:12:32it's text. It doesn't really get one of
- 3:12:34those. And then I did this one third.
- 3:12:37And then I did um I think I did this one
- 3:12:41fourth and then it got overwritten with
- 3:12:43the fifth. So it just tells you like how
- 3:12:45many executions you've done and what is
- 3:12:47which number execution that was. Then I
- 3:12:49did this one sixth.
- 3:12:51It just keeps track of your executions.
- 3:12:55Okay. So somebody asked about a kernel.
- 3:12:58What is a kernel? So uh the kernel is um
- 3:13:06the kernel is basically the interpreter.
- 3:13:09So it's the thing that that the kernel
- 3:13:11is just a a um a copy of the interpreter
- 3:13:16that the notebook is attaching to in
- 3:13:18order to run. So the notebook can't run
- 3:13:21anything because remember a pi Python
- 3:13:24needs
- 3:13:25Python needs an interpreter to run its
- 3:13:29code. So in notebooks we basically
- 3:13:33create like a virtual copy of the
- 3:13:35interpreter called a kernel. Um and you
- 3:13:38can actually have many kernels based on
- 3:13:41your um your base interpreter. So what's
- 3:13:45nice is Anaconda
- 3:13:49um Anaconda comes with
- 3:13:53uh an interpreter for you and then you
- 3:13:57create kernels that are virtual copies
- 3:13:59of that um that are virtual copies of
- 3:14:03that uh interpreter so that you can run
- 3:14:05your Python code against it. Remember
- 3:14:08you need an interpreter but notebooks
- 3:14:11attach to kernels. Kernels are like
- 3:14:14virtual interpreters.
- 3:14:16Um, and you can have many kernels based
- 3:14:18on the original interpreter. So the
- 3:14:21kernel is literally just think of it
- 3:14:23like the computer that's powering the
- 3:14:25notebook. That's all. It's just the
- 3:14:27compute engine that's allowing you to
- 3:14:29execute your Python code. So every
- 3:14:32notebook has an associated kernel.
- 3:14:36And what's interesting is if you restart
- 3:14:38your kernel, you lose all your data. So
- 3:14:41all of these outputs that we have um you
- 3:14:44would lose if you restarted your kernel.
- 3:14:48So if I restart um now I like since I
- 3:14:52restarted this is not going to know what
- 3:14:54x is. So if I try to print x again it's
- 3:14:57going to say I don't know what x is
- 3:14:58because I restarted my kernel. I lost
- 3:15:01all that data.
- 3:15:05But I can redefine it. And then there it
- 3:15:08is. And notice that my iterations
- 3:15:10restart
- 3:15:12um my iterations restart when I uh
- 3:15:15restart my kernel. So if if I go back
- 3:15:17and restart the kernel again
- 3:15:20and now if I run this, this will be
- 3:15:23first. So notice how that restarts to
- 3:15:25first. This will be second. This will be
- 3:15:28third.
- 3:15:31Try restarting. I don't know what that
- 3:15:33is.
- 3:15:35I don't know why that
- 3:15:40Yeah, choose the Anaconda. Either one.
- 3:15:42Choose. You want to use Anaconda as your
- 3:15:45But what that's saying is what do you
- 3:15:46want to use as your interpreter to to
- 3:15:48build your kernels. So, choose one of
- 3:15:51those. That's fine.
- 3:15:54So, yeah, Anaconda requirements, laptop
- 3:15:56requirements. Um,
- 3:15:58you need a little bit of you need a
- 3:16:00little bit of RAM. Uh, you need a little
- 3:16:04bit of RAM to run the notebooks because
- 3:16:07you need some memory. Um, you don't need
- 3:16:10a lot of it though. I'd be surprised if
- 3:16:12you didn't meet the requirements. It's
- 3:16:13not that much, but you do need some.
- 3:16:17I'm not sure the exact. I'd have to look
- 3:16:20that up on the Anaconda website.
- 3:16:24Oh, it must have been full to start with
- 3:16:26or pretty full. I'd be This doesn't take
- 3:16:28up that much space, I don't think.
- 3:16:33Was it pretty full to begin with?
- 3:16:36I would assume. I don't think this takes
- 3:16:38up that much space.
- 3:16:42Uh, what I want to do then, I want to go
- 3:16:44over to collab. Okay. How do we feel
- 3:16:46about the notebooks? I maybe if it's
- 3:16:48still installing for you, give it a
- 3:16:50little time. Open the open the
- 3:16:52navigator.
- 3:16:54Let's try Coll. I guarantee you Collab
- 3:16:56will work for you if you're still having
- 3:16:58issues with if you're having issues with
- 3:16:59Jupiter.
- 3:17:01No worries. Let's just try collab. I
- 3:17:03promise that will be a lot easier, be a
- 3:17:06million times easier, I think, than than
- 3:17:08working with Jupiter. Okay, great. So,
- 3:17:11we will continue then.
- 3:17:16All right, let me jump over to our
- 3:17:20final
- 3:17:21um
- 3:17:23demo with setting up a collab notebook.
- 3:17:26So I'm just going to jump into doing
- 3:17:28that on in the interest of time.
- 3:17:31Uh
- 3:17:34so
- 3:17:38would you be taking up additional
- 3:17:39sessions too other than uh so we're I'm
- 3:17:42going to be the instructor for all of
- 3:17:45the courses in this program. So you're
- 3:17:48you're stuck with me
- 3:17:50for all of those. Does that make sense?
- 3:17:53like all of the all of the uh AI
- 3:17:57engineer program uh courses.
- 3:18:04Yeah. Yeah. Stuck or be excited. It's
- 3:18:08going to be one or the other. Probably
- 3:18:10not an in between feeling.
- 3:18:13Hopefully. Cool. Hopefully. Hopefully
- 3:18:16good. Yeah. Like I said, I've taught
- 3:18:19this many times. Uh, I think it would be
- 3:18:24uh I think it'll be good.
- 3:18:26You are stuck. Okay. Well, we're going
- 3:18:28to get you unstuck with Collab. I would
- 3:18:30not worry about getting Jupiter set up.
- 3:18:33If it's not working for you, we're going
- 3:18:34to ditch it and we're going to use
- 3:18:35something else that works. I promise
- 3:18:38it's not a I promise getting Jupiter set
- 3:18:40up is not that important relative to
- 3:18:42getting at least one of these options
- 3:18:44that works.
- 3:18:47So, if it's not working over on Jupiter,
- 3:18:50I'm not worried in the slightest about
- 3:18:52it because there's going to be plenty of
- 3:18:53options to run run Python code. In fact,
- 3:18:56we're going to do one next which is
- 3:18:58going to be with um with Collab. So,
- 3:19:02we'll do that. Um so, let me jump into
- 3:19:07that. Let me share my screen here.
- 3:19:11Um,
- 3:19:20learning a lot already. Great. That's
- 3:19:21great. Glad to hear that. Thank you.
- 3:19:26Okay.
- 3:19:29Thank you guys. Appreciate it.
- 3:19:31All right. Let me go to the demo. Demo
- 3:19:35three.
- 3:19:37All right. So, what you want to do
- 3:19:39essentially is to go to this website,
- 3:19:43um, which I have here. I'm going to, uh,
- 3:19:46paste it in the chat. Um, so what you
- 3:19:50want to do is go to Google's website for
- 3:19:52their Collab platform, uh, which is, so
- 3:19:57Collab is a, um, notebook platform that
- 3:20:02Google hosts. So, you don't need to
- 3:20:04install anything. You just go there in
- 3:20:05your favorite web browser, log in with
- 3:20:08your Google account. In fact, I don't
- 3:20:10even think you need to be necessarily
- 3:20:12logged in. You can in order to save your
- 3:20:14notebooks to your drive,
- 3:20:16but um you go there and you basically
- 3:20:20open up a notebook and start working
- 3:20:22with it right away. And it's fantastic.
- 3:20:26Their notebook environment already has a
- 3:20:29lot of packages installed into it for AI
- 3:20:32and machine learning. So that's f that's
- 3:20:34really great. Um
- 3:20:38uh once you get to the page um you
- 3:20:41should log in though if you have a
- 3:20:43Google account. Only reason I say that
- 3:20:46is because it will save your notebooks
- 3:20:48to your drive automatically so that you
- 3:20:50it will automatically save your
- 3:20:51notebooks just like as if you're working
- 3:20:53in a Google doc. So that's great. So
- 3:20:56that um it saves your work
- 3:20:58automatically.
- 3:20:59Um, so please, you know, I would
- 3:21:01recommend getting a Google account if
- 3:21:03you don't have one for free. Logging in
- 3:21:05using Collab is completely free.
- 3:21:08Um, so it's a fantastic platform. Um,
- 3:21:12when you go to that site, uh, assuming
- 3:21:14you've logged in, you want to click on
- 3:21:16that lower left blue button where it
- 3:21:18says new notebook. You can see it in
- 3:21:20this screenshot. And I I'll open up one
- 3:21:23in a moment on on my screen. But do you
- 3:21:26guys see this screen right here that's
- 3:21:28in this screenshot that has the new
- 3:21:31notebook on the bottom in the lower
- 3:21:32left?
- 3:21:36No. From that site, what do you see?
- 3:21:45Oh, so you're already in a notebook. It
- 3:21:47you're already in a notebook. Like it
- 3:21:48says, "Welcome to Collab.
- 3:21:52Oh, okay. So, it already opened the
- 3:21:53notebook for you. Okay, that's that's
- 3:21:55fine. That's fine. I'll show you uh I'll
- 3:21:58show you what that looks like. That's no
- 3:22:00problem. That means you're already
- 3:22:02inside of it.
- 3:22:05Okay.
- 3:22:08Okay. So, then we're pretty much in the
- 3:22:10notebook environment and we can start
- 3:22:12running code. Let me let me hop over to
- 3:22:14Collab and show you guys what it looks
- 3:22:16like.
- 3:22:17Let me stop sharing that and jump over
- 3:22:19to
- 3:22:21collab here.
- 3:22:32Okay.
- 3:22:34So, if you're in the welcome to collab,
- 3:22:36um that's fine or you can start a new
- 3:22:40notebook. Let me assume that we've
- 3:22:41opened up welcome to collab. So, you're
- 3:22:43in this screen. What you want to do if
- 3:22:45you're in this screen is just go to go
- 3:22:47up to file and do new notebook.
- 3:22:52Just go to file, new notebook in drive.
- 3:22:54Just do that. File new notebook
- 3:23:00and this will create a new notebook for
- 3:23:02you which will start fresh a blank
- 3:23:04notebook.
- 3:23:06Okay.
- 3:23:09Were you able to do that? If you guys
- 3:23:13were folks able to get here to this uh
- 3:23:16blank notebook
- 3:23:18one way or the other, you clicked the
- 3:23:19blue button to hit a new one or you went
- 3:23:21up to file and did new notebook.
- 3:23:25Yes. Okay.
- 3:23:28So, there we are. Without doing all the
- 3:23:30Jupiter install steps, we're in a
- 3:23:32notebook. Look how easy that was, right?
- 3:23:34So, why didn't we just start with this?
- 3:23:37Um
- 3:23:39so yeah so we're in Google's notebook
- 3:23:42platform uh which is a fantastic
- 3:23:44platform and uh what's great about this
- 3:23:47is you can export these now these are
- 3:23:50pyb which is which is uh if you're
- 3:23:52curious what that extension means it's
- 3:23:54short for interactive python notebook
- 3:23:58okay IPIB
- 3:24:00so these are the files that you can open
- 3:24:02in Jupiter if you have uh if you notice
- 3:24:05when you open up Jupyter notebook book
- 3:24:07earlier it was a IP YMBB
- 3:24:10um inside of VS code when you work with
- 3:24:13notebooks they are IP YMBB so IPMBB is a
- 3:24:17notebook file and it can be opened in
- 3:24:20any one of these three platforms right
- 3:24:22the Jupiter from Anaconda the uh VS code
- 3:24:26can open IPMB and you can also upload
- 3:24:30your own notebooks here if you have them
- 3:24:32on your own machine you just go to file
- 3:24:34upload notebook and And then it will
- 3:24:37open up a box where you can choose which
- 3:24:39file on your machine to upload. So you
- 3:24:41can upload your own notebooks, which
- 3:24:43will be uh great when we get into um
- 3:24:46demos. We have demo notebooks for you
- 3:24:48guys that we'll work through with our
- 3:24:50code. You can upload those into Collab
- 3:24:52and work with them directly inside of
- 3:24:54here.
- 3:24:56So let's try running something. Let's do
- 3:24:59the print
- 3:25:01hello world.
- 3:25:05So, um, you want to type that in and I
- 3:25:08can paste it in the chat for you guys
- 3:25:11and then you want to you want to hit
- 3:25:13either shift enter or this play button
- 3:25:15right next to the cell.
- 3:25:21Okay, so that might take a moment
- 3:25:23because it's booting up your uh your
- 3:25:25kernel.
- 3:25:27Uh, but then it should run and you
- 3:25:29should see the output. Now, this is
- 3:25:30going to look very similar to Jupiter,
- 3:25:32just slightly different.
- 3:25:36We're inside of Collab
- 3:25:38and we started a new notebook.
- 3:25:42We're just inside of a blank notebook
- 3:25:44for now. And we uh are just within this
- 3:25:47first cell and I I'm doing hello world.
- 3:25:50Were you guys able to run that?
- 3:26:00We didn't. But we could we could open a
- 3:26:03notebook in VS Code because we installed
- 3:26:05the extension. We did that. Remember we
- 3:26:08installed the Jupiter extension. So we
- 3:26:10can open notebooks in VS Code and we can
- 3:26:12run them there. I just didn't show that
- 3:26:14to us. Uh we might do that later down
- 3:26:17the road, but you do have that
- 3:26:19flexibility to run things there if you
- 3:26:22want to.
- 3:26:26Okay. One thing I want to show you guys
- 3:26:29that's really cool. So, um, one thing I
- 3:26:33want to show you is if you go up to
- 3:26:34runtime
- 3:26:36and then go down to change runtime type.
- 3:26:40Do you do you guys have that? Change
- 3:26:41runtime type. So, if you click on
- 3:26:44runtime
- 3:26:46and then change runtime type.
- 3:26:49Do we have that? Click on that. Click on
- 3:26:53change runtime type.
- 3:26:56And look at our options. We can choose a
- 3:26:58GPU for free.
- 3:27:03So we can swap over to a GPU kernel
- 3:27:06which is fantastic for training deep
- 3:27:08learning models and we can use that GPU
- 3:27:11for free. This is one of the reasons
- 3:27:13that uh Collab is so amazing is they
- 3:27:17give free access to a GPU. So if you
- 3:27:21don't have one on your own machine um
- 3:27:24you can use the GPUs from Collab for
- 3:27:26free.
- 3:27:28Yeah, go ahead. I mean, there's no
- 3:27:30nothing wrong with it. So, uh, what it's
- 3:27:32going to ask you to do is, uh, terminate
- 3:27:35your current kernel because you're
- 3:27:37connected to a CPU basic kernel. It
- 3:27:40wants you to disable that so you can
- 3:27:41swap over the GPU. Click okay. That's
- 3:27:44okay.
- 3:27:48All right. And then we are now um, we
- 3:27:50click save. And that will swap us over
- 3:27:53and connect us to a GPU. So, if you how
- 3:27:57you know that you're connected to a GPU
- 3:27:58is if you go over to um
- 3:28:03if you go over to this box on the right.
- 3:28:05Do you guys see that one where it says
- 3:28:07RAM and disk? If we click that,
- 3:28:12it will show us our resource resource
- 3:28:14usage. And you should see GPU RAM
- 3:28:16available of 15 gigs.
- 3:28:20So, you have 15 gigabytes of GPU RAM
- 3:28:22available that you can use.
- 3:28:25So remember, you just click this RAM
- 3:28:28um
- 3:28:31you should click this RAM uh
- 3:28:35uh
- 3:28:36sorry this RAM and disk.
- 3:28:44Roberto, did you swap over the runtime
- 3:28:46to
- 3:28:47uh Yeah, it should pop up. Okay, then
- 3:28:51you should be able to click on this
- 3:28:57Yeah, it might take some time to connect
- 3:28:59to one because what Google has to
- 3:29:01allocate one to you um and then it has
- 3:29:04to like connect it over the cloud. It
- 3:29:06can take a minute. Yeah, it can take a
- 3:29:08minute. It has to allocate one to you.
- 3:29:10So the question is which one is better?
- 3:29:12Um,
- 3:29:14so for the vast majority of things,
- 3:29:17the CPU, the standard CPU runtime, which
- 3:29:20is the default, is going to be better
- 3:29:22for the vast majority of things. The
- 3:29:24only time the GPU is really going to be
- 3:29:26beneficial is when we start doing deep
- 3:29:28learning and training neural networks,
- 3:29:31then the GPU will be really beneficial.
- 3:29:34It will speed up the training time by a
- 3:29:37significant amount.
- 3:29:40I can tell you like I trained a uh
- 3:29:43neural network for images for image
- 3:29:46recognition. Uh it took me it took two
- 3:29:50hours on the CPU and then when I swapped
- 3:29:52it over to GPU it took less than a
- 3:29:54minute
- 3:29:57took less than a minute and it was
- 3:29:58taking two hours on the CPU.
- 3:30:03So yeah, training neural nets on a GPU
- 3:30:07when we get to that is going to be
- 3:30:09beneficial. So if you're not using
- 3:30:12Collab right now, that's okay, but in
- 3:30:14the future when we get into deep
- 3:30:16learning, you're likely going to want to
- 3:30:17use Collab to swap over to the GPU for
- 3:30:20free.
- 3:30:22Now, they do rate limit you,
- 3:30:26so it's free, but you can max it out in
- 3:30:28a day and then they cool you off for 24
- 3:30:31hours, which I have I have done, uh,
- 3:30:34unfortunately. So, like, if you max out
- 3:30:37that RAM and you use it too much in a
- 3:30:4024-hour period, they will, uh, not allow
- 3:30:44you to connect to a free GPU for another
- 3:30:4624 hours.
- 3:30:48So, I doubt you'll run into that
- 3:30:51situation, but I have before
- 3:30:53if you're just if you're just using it
- 3:30:56so much.
- 3:31:02No. So, unfortunately,
- 3:31:04uh, no.
- 3:31:09So, unfortunately, no. You cannot, um,
- 3:31:12Collab doesn't connect to your local
- 3:31:14resources. It only it does the cloud
- 3:31:16Google's cloud resources. So no, you
- 3:31:18can't use your own through collab. But
- 3:31:20yes, you could use your own GPU through
- 3:31:23VS Code. I will show us how to do that
- 3:31:25later when we get into deep learning.
- 3:31:28I will show you that later.
- 3:31:32We we won't need to worry about that
- 3:31:34now, but later on, yes, that'll be
- 3:31:37important.
- 3:31:42Uh it doesn't show GPU. Make sure you
- 3:31:44swap over the runtime to go to change
- 3:31:47runtime type and make sure you pick GPU.
- 3:31:51Make sure you go away from
- 3:31:56No, you should use Collab. I wouldn't
- 3:31:58You don't need to buy anything. You just
- 3:32:01use Collab. Just use Collab for sure.
- 3:32:05Collab's free. It does everything you're
- 3:32:07going to need for the class.
- 3:32:12Yeah, I I highly advocate for Collab. It
- 3:32:15So, by the way, if you're curious, like
- 3:32:17Collab came about because Google wanted
- 3:32:20the the machine learning research
- 3:32:22community to have access to GPUs for
- 3:32:25free to um develop like machine learning
- 3:32:28and uh deep learning models. So, uh it's
- 3:32:33been around for a while. I remember
- 3:32:35using Collab um probably almost 10 years
- 3:32:38ago and it used to be it used to be very
- 3:32:43lucky if you got a GPU. You used to like
- 3:32:46you used to have to click and then hope
- 3:32:48that you would get allocated one and
- 3:32:50sometimes you wouldn't and I would sit
- 3:32:52there and have to refresh and try to
- 3:32:54hope that I would get a GPU but now it's
- 3:32:57it's like readily available which is
- 3:32:59fantastic.
- 3:33:03No, you're But you're joining it at a
- 3:33:04good time because I'm telling you, the
- 3:33:06GPU was very difficult to get. I would
- 3:33:10always try to switch over to that and I
- 3:33:12would rarely be able to.
- 3:33:18So,
- 3:33:22pretty good. But like I said, like if
- 3:33:25the CPU is perfectly fine for everything
- 3:33:28we're going to do, except when we get
- 3:33:30into deep learning, you're going to want
- 3:33:31to swap that over to GPU. But that's
- 3:33:34going to be for we have a while till we
- 3:33:36get to deep learning.
- 3:33:38We have a lot to learn between now and
- 3:33:40then.
- 3:33:45Okay.
- 3:33:47What do we think? Do we like collab?
- 3:33:50We're comfortable with it. Feel free to
- 3:33:51use it. Feel free to use VS Code. Feel
- 3:33:53free to use Jupiter. Whatever you want
- 3:33:56to use, okay? There's options, right? I
- 3:34:00hopefully you have options that work for
- 3:34:02you. Um they are all used in the
- 3:34:05industry. So you're not missing out on
- 3:34:08if whatever you use, people use it of
- 3:34:11these three people use all of them.
- 3:34:15So feel free to use whatever is easiest.
- 3:34:22Yeah, I can show that real quick. Yeah,
- 3:34:28let me go back over to it.
- 3:34:31I'm going to be real quick on it though
- 3:34:32because I want to make sure we get over
- 3:34:33to our other material.
- 3:34:44Okay, let me show you how you can run a
- 3:34:46notebook. Let me show you how to run a
- 3:34:48notebook. So I'm going to go to file,
- 3:34:50new file, and open a Jupyter notebook.
- 3:34:56Okay. So it'll open a new. Now notice
- 3:34:59notice the extension of it
- 3:35:02is
- 3:35:05MB. That should be no surprise. That is
- 3:35:07the universal kind of interactive Python
- 3:35:10notebook file.
- 3:35:12Okay. So the biggest thing you have to
- 3:35:15do when you open a notebook in VS Code
- 3:35:18is you have to
- 3:35:21tell it what kernel to connect to. So
- 3:35:25have to go to select kernel
- 3:35:27and then what you have to select is the
- 3:35:31Python environment. And luckily if you
- 3:35:34installed Anaconda
- 3:35:36you have a built-in Python environment
- 3:35:39which is going to be your uh which is
- 3:35:42going to be um the
- 3:35:47you know which is going to be the uh
- 3:35:49Anaconda that you installed.
- 3:35:51So I have Anaconda here. Now I have a
- 3:35:55lot of other ones but the
- 3:35:58uh Anaconda is here. Say it's this one.
- 3:36:08Does that so when you
- 3:36:11when you uh
- 3:36:16Yeah, you have to install you have to
- 3:36:18install a Python environment. Yes. So
- 3:36:21you want to install Anaconda first and
- 3:36:23then you can run your then you can run
- 3:36:25your notebooks.
- 3:36:33And then you just uh run your code as
- 3:36:36usual
- 3:36:40and then we can run that.
- 3:36:44Yeah, it but like it's working as if you
- 3:36:47know the same kind of notebook that we
- 3:36:49have inside of Collab, the same kind of
- 3:36:52notebook we have in Jupiter.
- 3:36:57You you have to have a Python installed
- 3:37:00for this to work. So you go to you go to
- 3:37:02Python environments
- 3:37:06and then choose a Python environment.
- 3:37:10You could try to create one. I'm not
- 3:37:12sure if that'll work for you. Create
- 3:37:14Python environment. You could try that,
- 3:37:16too.
- 3:37:21Yeah, that's fine. Any anyone will work.
- 3:37:26Any Python will work. You just need to
- 3:37:27pick a Python. Anyone will work.
- 3:38:02Okay. Yeah, if it defaulted to something
- 3:38:04that's fine, too.
- 3:38:07And then we can generate more cells
- 3:38:14and run cells.
- 3:38:19But yeah, that's the thing is you're
- 3:38:20going to want to install um Anaconda
- 3:38:22most likely because you need a Python
- 3:38:25version installed on your machine in
- 3:38:28order to run this.
- 3:38:37Perfect.
- 3:38:42Okay.
- 3:38:44All right. So, what I want to do is jump
- 3:38:45back over to our notes so we can
- 3:38:47continue along. Um again feel free to
- 3:38:50use whatever
- 3:38:52platform works for you. Collab,
- 3:38:55doing notebooks in VS Code, doing
- 3:38:57Jupiter notebooks, whatever works for
- 3:39:00you, please feel free to use that. There
- 3:39:03is no wrong way of using it. Whatever is
- 3:39:06best suited to you and you're most
- 3:39:07comfortable with, please use that
- 3:39:10option.
- 3:39:14Uh, it's lowercase P. That's why
- 3:39:19lowercase P. Capital P is not a function
- 3:39:22in Python.
- 3:39:24Lowerase.
- 3:39:28Yeah, go with Collab. Yeah, if you're if
- 3:39:30you're on a machine, you can't install
- 3:39:32anything, go with Collab. That's totally
- 3:39:34fine. That's why it's there is for the,
- 3:39:38you know, convenient kind of cloud
- 3:39:40aspect to it.
- 3:39:42Um,
- 3:39:43feel free to do collab for everything.
- 3:39:45That's totally fine.
- 3:39:48I will use collab from time to time as
- 3:39:50well.
- 3:39:55All right, let me uh go back to our
- 3:39:59notes then
- 3:40:05and pick up from uh syntax. So, I'm
- 3:40:09going to go back to Let me share my
- 3:40:10screen. Go back to
- 3:40:16Can you use Collab on your phone? I've
- 3:40:18never tried it. I would be surprised.
- 3:40:21Maybe an iPad.
- 3:40:23Maybe like a tablet. It could work
- 3:40:25pretty well.
- 3:40:27Phone. I'm not so sure.
- 3:40:37Yeah. Go ahead. Try it and let me know
- 3:40:39how it works.
- 3:40:45Try it and let me know. I really don't
- 3:40:47know. I'm curious now to try that.
- 3:40:53All right. So, I'm going back over the
- 3:40:54notes. We're going to finish up today uh
- 3:40:56what the time we have left to go through
- 3:40:59some syntax. So really uh
- 3:41:03really getting into um into Python like
- 3:41:07the actual code of it so we can get
- 3:41:10started on that and start working our
- 3:41:12way through it.
- 3:41:14Uh the difference so the the difference
- 3:41:17is um you will be executing py files
- 3:41:22with the within the terminal. So you'll
- 3:41:25be running Python files instead of cells
- 3:41:28in a notebook. you're just you're
- 3:41:30running a you're running a Python script
- 3:41:34rather than individual cells.
- 3:41:39Okay, so there's a difference there.
- 3:41:44And the reason the reason we choose
- 3:41:45notebooks is to run individual cells.
- 3:41:48It's just easier.
- 3:41:50Same syntax,
- 3:41:52same syntax. It's just the code is not
- 3:41:54isolated into cells.
- 3:42:00still Python.
- 3:42:05All right.
- 3:42:08Um, let's go forward into the syntax,
- 3:42:13start learning about it.
- 3:42:15All right. So, something we need to
- 3:42:18learn about is how do we properly write
- 3:42:21Python code? What is the syntax to it?
- 3:42:24So, some things we're going to need to
- 3:42:26learn about are how to write proper
- 3:42:29identifiers, which are names.
- 3:42:32Identifiers are just names for
- 3:42:33variables. So, we need to know what's
- 3:42:36allowed, what's not allowed. We need to
- 3:42:38talk about what the indentation means
- 3:42:40and why do we need it in Python. I want
- 3:42:43to show you guys how to write comments
- 3:42:45because that's really important to
- 3:42:46leaving notes to yourself or others
- 3:42:49about the code and then talk about um
- 3:42:52generally how we produce output and how
- 3:42:54we can accept input um from a user or
- 3:42:58someone interacting with our code. I
- 3:43:01want to talk about all these things.
- 3:43:02We'll see how how much of what we get
- 3:43:04to.
- 3:43:07But let's start with the identifier. So
- 3:43:10what is an identifier in programming?
- 3:43:12This is really for any programming
- 3:43:14language. An identifier is just a name
- 3:43:18we give to something inside of our code.
- 3:43:21So it's a name we give to a variable, a
- 3:43:23name we give to a function, a name we
- 3:43:25give to an object. Um so any name we
- 3:43:30give to something in our code, like when
- 3:43:31we set something equal to x, like x
- 3:43:35equals 3 + 3. um that thing the x is the
- 3:43:41name we are giving or assigning to a
- 3:43:44result or a variable or an object. So
- 3:43:48anytime we write down a name in our code
- 3:43:52of something there are certain rules
- 3:43:54that those names have to follow and
- 3:43:57these are something we will um pick up
- 3:43:59as we go along but I wanted to call them
- 3:44:02out here. So um when we name anything in
- 3:44:06Python
- 3:44:08generally they have to follow these set
- 3:44:10of rules meaning they have to be a combo
- 3:44:13of lowercase or uppercase letters either
- 3:44:16one's allowed it can be even be a
- 3:44:19mixture of lowercase and uppercase
- 3:44:21numerical digits are allowed in the name
- 3:44:24that's okay any digit 0 to nine it's
- 3:44:27okay and underscores are okay
- 3:44:32underscores are Okay. And there's no uh
- 3:44:36minimum or maximum length. So names can
- 3:44:40be really long, they can be really
- 3:44:41short. Um of course they should be
- 3:44:45meaningful. So when we name something,
- 3:44:49it should not remember we want to kind
- 3:44:51of get away from naming everything X or
- 3:44:53Y or A or B um because those names may
- 3:44:58not mean much when we look back at the
- 3:45:00code. So even though those are valid
- 3:45:02names from an identifier perspective, we
- 3:45:05want to be really meaningful when we
- 3:45:07name something. We name a variable, name
- 3:45:09a function, name an object. Um,
- 3:45:13here's one catch.
- 3:45:16The name cannot start with a number. So
- 3:45:20I can't name something uh just the the
- 3:45:23number zero or the number one um because
- 3:45:27I can't start with that. Now, it can
- 3:45:29include that.
- 3:45:31So, if I need to include a number in the
- 3:45:34name, as long as it doesn't start with
- 3:45:37it, that's okay. But names cannot start
- 3:45:40with a digit. That's just one rule of
- 3:45:42Python. Anything that you're assigning a
- 3:45:45name to,
- 3:45:47like a variable, function, whatever,
- 3:45:50cannot start with a number or else it'll
- 3:45:52be invalid.
- 3:45:54Okay?
- 3:45:56So I'll show us examples of that later.
- 3:45:59Yes, they can start with underscores.
- 3:46:00Yes, that's okay. It can start with
- 3:46:03underscores. Of course, it can be lower,
- 3:46:05uppercase. It can start with It cannot
- 3:46:07start with a digit. It can have digits.
- 3:46:10They just can't be the first character
- 3:46:12of the name.
- 3:46:15Yes.
- 3:46:17Um, now special symbols cannot be used
- 3:46:20in the name. So you cannot have a
- 3:46:22percentage, dollar sign, exclamation
- 3:46:24point, hyphen,
- 3:46:27pound symbol, at amperand symbol, at
- 3:46:30symbol. None of those can be used in the
- 3:46:33name. So those symbols are not
- 3:46:35recognized.
- 3:46:36So if you try to include those in the
- 3:46:38name of something,
- 3:46:40that will produce an error. So we don't
- 3:46:43want to do that. The other thing we want
- 3:46:45to avoid is naming something in the same
- 3:46:49name as something that already exists
- 3:46:52internal to Python. So those things are
- 3:46:54called keywords. So there are certain
- 3:46:57keywords that have a meaning in Python.
- 3:47:00They are built into the language. We
- 3:47:03cannot reuse those. They're basically
- 3:47:05reserved. Um so something like class is
- 3:47:09reserved because it means something. It
- 3:47:11means you're declaring a class. We'll
- 3:47:13see. We'll talk about what that means
- 3:47:14later. Or something like global cannot
- 3:47:17be used because it declares something as
- 3:47:19global. Um,
- 3:47:22so
- 3:47:24you know, we'll learn what some of those
- 3:47:26keywords are. There's a list of them
- 3:47:28that are in the Python documentation,
- 3:47:31but we want to avoid naming things after
- 3:47:34builtin
- 3:47:36uh uh functions or builtin keywords. Um
- 3:47:41so so in fact we've already used one
- 3:47:45which is the print function. You know
- 3:47:47when we printed out hello world we would
- 3:47:50want to avoid naming something print
- 3:47:52because print means something. It exists
- 3:47:55as a function. We don't want to name
- 3:47:58something print
- 3:48:00right that would that would produce it
- 3:48:02because it would produce confusion. The
- 3:48:03interpreter would see that and say oh do
- 3:48:05you mean the function print or you
- 3:48:07trying to name something print? it
- 3:48:09wouldn't know. So, we want to avoid
- 3:48:12naming something that already exists
- 3:48:14inside of Python like print or like
- 3:48:17class global. Um, there's many others.
- 3:48:26Okay. Lastly, and this one always throws
- 3:48:29people off, is that when we name
- 3:48:31something uh that is case sensitive. So,
- 3:48:36if you name something lowercase A, that
- 3:48:39is a completely different variable or
- 3:48:41completely different function than if we
- 3:48:42were to name something capital A. These
- 3:48:45are different. They're treated
- 3:48:47differently. So, Python will think that
- 3:48:50those are two different uh objects or
- 3:48:54variables or whatever the case is. So be
- 3:48:56really careful with case sensitivity.
- 3:48:59Python is case sensitive.
- 3:49:03Lowercase A will not be treated the same
- 3:49:05as capital A. And whenever you're naming
- 3:49:07something, so if I have a variable and I
- 3:49:10I I set lowercase A equal to five and
- 3:49:14then I set um uh capital A equals to 10.
- 3:49:19Then if I um print A, that would produce
- 3:49:23five.
- 3:49:25But if I print capital A, that will
- 3:49:27produce 10. It's not the same. So it is
- 3:49:31case sensitive.
- 3:49:33These would be two different names.
- 3:49:35Lowerase A and capital A.
- 3:49:40Okay.
- 3:49:42So, we're going to do examples with
- 3:49:44these, but these are just some rules we
- 3:49:47have to abide by in the syntax if we're
- 3:49:50naming anything like naming our
- 3:49:53variables, naming our functions, naming
- 3:49:54our objects as we go along in in the
- 3:49:57course, right? We just cannot The main
- 3:50:00one that trips people up is we can't
- 3:50:02start with a digit and we can't use
- 3:50:05words that already exist like print.
- 3:50:11Oh, is my video stuck for people?
- 3:50:17Was it stuck?
- 3:50:19Oh, okay. Always let me know because it
- 3:50:22might it might be.
- 3:50:24Always let me know because it definitely
- 3:50:26could be. So, it's better to know than
- 3:50:29not to know.
- 3:50:36Okay.
- 3:50:37Oh, no worries. Like I said, always
- 3:50:39always feel free to to let us know cuz
- 3:50:43um it would if it is, then it's good to
- 3:50:46call it out. So, no worries about that.
- 3:50:49Any questions about these names? Do does
- 3:50:52it make sense about like how we name
- 3:50:55things matters and there just are
- 3:50:57certain rules that we have to follow. Um
- 3:51:00we want to avoid these wacky symbols.
- 3:51:03Um, you know, we want to avoid naming
- 3:51:06things that already exist. We want to
- 3:51:08avoid starting with a digit. Otherwise,
- 3:51:11it's going to be a pretty standard like
- 3:51:13lowercase, uppercase mixture,
- 3:51:16maybe occasionally with an underscore
- 3:51:18mixed in there. Um, or or digits even.
- 3:51:22As long as we don't start with one,
- 3:51:23that's okay.
- 3:51:30Yeah, Tim, that's a good reference. the
- 3:51:32PEP. So PEP
- 3:51:35are the set of guidelines that um the
- 3:51:38Python Foundation has kind of agreed
- 3:51:41upon as um here's what you should use as
- 3:51:45your style guide. Here's what here's
- 3:51:48what the community believes is the best
- 3:51:49style for Python. Those are good to
- 3:51:52read.
- 3:51:56Yeah, those are those are uh good
- 3:51:57references for really like formatting
- 3:52:00and styling your your Python code uh to
- 3:52:03be in line with kind of what the
- 3:52:05community expects.
- 3:52:10Okay.
- 3:52:16Okay. So, let me give you some examples.
- 3:52:19Um so, the ones on the left are going to
- 3:52:21be valid. So, we can name something my
- 3:52:23class. We can name something var_1
- 3:52:26that's okay.
- 3:52:29Count
- 3:52:30um that's okay. Uh
- 3:52:34but if we have
- 3:52:37um like on the right if we have
- 3:52:39something that starts with a digit that
- 3:52:41would be bad. So so this one is no good
- 3:52:44because it starts with this number.
- 3:52:47That's not good. Um this name has this
- 3:52:50wacky at symbol in it. that's no good.
- 3:52:53So, this would be a bad name for
- 3:52:55something that would produce an error.
- 3:52:57Remember, the interpreter is going to
- 3:53:00see that and reject it essentially and
- 3:53:02say, "You can't name something this.
- 3:53:05It's not valid." Um, same thing with
- 3:53:08trying to name something global. This is
- 3:53:09a keyword that already exists in the
- 3:53:11language. The interpreter is going to
- 3:53:13see that and get confused. It's not
- 3:53:15going to know if you're talking about
- 3:53:16the keyword that's built in or you're
- 3:53:19trying to come up with your own name.
- 3:53:21It's not going to know. So, it's just
- 3:53:22going to throw an error.
- 3:53:24Um, so again, we want to we want to keep
- 3:53:27things simple. We want to use
- 3:53:30underscores where it makes sense. We we
- 3:53:32don't want to start with numbers. Um, we
- 3:53:35can use a mixture of lowercase and
- 3:53:36uppercase. That's fine.
- 3:53:39Um, this is a good variable name rather
- 3:53:43than if I just called something X.
- 3:53:46That again, we're trying to avoid that's
- 3:53:49something I always see in the beginning.
- 3:53:51I think is okay in the beginning, but
- 3:53:52it's something we really want to be
- 3:53:54conscious of is naming our variables
- 3:53:57very meaningfully.
- 3:53:59Like count is going to be more
- 3:54:01meaningful if we're keeping track of a
- 3:54:03count of something. We would rather call
- 3:54:06that count than if than if I just called
- 3:54:09it X. Because if you read the code,
- 3:54:11which do you guys believe me? Like when
- 3:54:14you see it, you kind of know exactly
- 3:54:16what it means. X or count? What do we
- 3:54:19think?
- 3:54:22Which one like has more meaning to it
- 3:54:24when you see it? You know exactly what
- 3:54:26it's keeping track of. X or count?
- 3:54:30Yeah, count.
- 3:54:32I would agree with that. Count. Yep.
- 3:54:36So, it's just an like that's just a
- 3:54:37single example of trying to keep track
- 3:54:41of things in a meaningful way. That's a
- 3:54:44good name to give to a variable. That
- 3:54:46would be uh you know keeping track of
- 3:54:49something the count of something
- 3:54:54rather than if we just called it x or y
- 3:54:56or a or b.
- 3:55:02All right, I want to talk to you guys
- 3:55:05about indentation next. So now we know
- 3:55:09we have to name things appropriately and
- 3:55:11the interpreter will give us an error if
- 3:55:12we don't name things appropriately.
- 3:55:15What about indentation?
- 3:55:20So indentation
- 3:55:22is a way for Python to understand what
- 3:55:26code gets executed together.
- 3:55:30Okay.
- 3:55:31So
- 3:55:33and it it also indicates that I am
- 3:55:38breaking the flow of the code from one
- 3:55:41section to the next. So the indentation
- 3:55:43is really important to signify to the
- 3:55:46interpreter there is a new section of
- 3:55:49code that has to be considered
- 3:55:52um before I move on. So um you should
- 3:55:57always use indentation
- 3:56:00whenever you have a colon like we see a
- 3:56:05colon here with if else statements. Now
- 3:56:08we haven't learned about if else
- 3:56:09statements but we will. But notice how
- 3:56:12we have the colon there and the
- 3:56:15interpreter would be okay with this.
- 3:56:18This would work.
- 3:56:20Okay,
- 3:56:22which is a simple statement of saying is
- 3:56:24five greater than two? Yes. So in the
- 3:56:27case that it is, let's run this code.
- 3:56:30But we're only able to run it because
- 3:56:32the interpreter recognizes it's
- 3:56:34indented.
- 3:56:36So the indentation is really really
- 3:56:38critical as it makes the interpreter
- 3:56:41understand what should be next. The
- 3:56:44interpreter understands what should be
- 3:56:46next after this statement. Um like an if
- 3:56:51statement or a loop statement. Um we
- 3:56:56will always have indentation. So this
- 3:56:59would actually uh throw an error because
- 3:57:01there's this is not indented. This is at
- 3:57:04the same level and if you have collab
- 3:57:08open you could try this for yourself.
- 3:57:11Um if you had collab open it you could
- 3:57:13try it for yourself is like this would
- 3:57:16this would throw an error where it says
- 3:57:19I am expecting indentation but you did
- 3:57:21not have have any.
- 3:57:27Does it matter how many spaces?
- 3:57:30Uh you yes you want to use four spaces.
- 3:57:34This this should be four spaces here.
- 3:57:39Spaces or tabs?
- 3:57:41Uh I'm only laughing because it's a
- 3:57:45it's a kind of a controversial question
- 3:57:50in the community. Some people get really
- 3:57:53upset over
- 3:57:55one or the other. I'm not one of those
- 3:57:57people. I don't really care. they so
- 3:58:00most code editors
- 3:58:02uh set the tab automatically as four
- 3:58:08spaces. So a tab will do the same thing
- 3:58:12as if you manually just did four spaces.
- 3:58:14It doesn't really matter in that case.
- 3:58:17So either way,
- 3:58:23yeah, you're so the ide will do that for
- 3:58:26you. The IDs will generally do that for
- 3:58:29you because they know it should go on
- 3:58:30the same line.
- 3:58:35Would it work as on the same line? Yes,
- 3:58:39in some cases it will, but not all. It
- 3:58:42depends on how complex it is. But what
- 3:58:46do you think is easier to read
- 3:58:49from a readability perspective? Which is
- 3:58:51easier
- 3:58:55if it's all in one line or is it more
- 3:58:57readable and easier to follow if it's
- 3:59:00indented?
- 3:59:08Yeah, the that's the purpose. So yes, I
- 3:59:12I agree. indented makes it easier to
- 3:59:14read. So that's another reason Python
- 3:59:18really enforces indentation is to make
- 3:59:20it easier to read. There's a reason they
- 3:59:22do that. It's to make it easier to read.
- 3:59:27Okay?
- 3:59:28And that's one of the best selling
- 3:59:30points of Python is how easy it is to
- 3:59:32read and work with. The indentation
- 3:59:34really helps. So to summarize this, we
- 3:59:38are going to have to use indentation.
- 3:59:40Anytime we have
- 3:59:43uh a statement with a colon. Anytime we
- 3:59:47have a statement with a colon, we're
- 3:59:48going to have to have an indentation
- 3:59:50immediately follow it. And there are
- 3:59:52certain statements that have a colon
- 3:59:53like if, else, else if, and any loop,
- 3:59:59any looping statement. Now, all of these
- 4:00:01things we're going to learn about, we'll
- 4:00:02learn about it in our next lesson.
- 4:00:05But anytime we have a colon, this is
- 4:00:08signaling the interpreter, okay, there
- 4:00:10needs to be a block of code following
- 4:00:13that, which is, yeah, as you say,
- 4:00:15Romero, it's like a a hierarchy. Yes,
- 4:00:18that's a great way of thinking about it.
- 4:00:20It's saying, okay, I should check this,
- 4:00:23then do this if that's true.
- 4:00:26that tells the interpreter this is only
- 4:00:28going to be executed in the event that
- 4:00:30this is actually true. Otherwise, I'm
- 4:00:32going to keep going.
- 4:00:42Okay.
- 4:00:44All right. Any questions about the
- 4:00:45indentation? This is something we're
- 4:00:46going to learn about more as we go
- 4:00:48along. When do we use indentation and
- 4:00:50when do we not? We're going to learn
- 4:00:51about it when we get into the if else
- 4:00:53and the loops which we will study.
- 4:00:57But do we do we let me ask you guys
- 4:00:59this. Do we understand the idea or the
- 4:01:03intent behind indentation?
- 4:01:07Do we roughly get that idea? We don't we
- 4:01:10don't know yet when to use it. I get
- 4:01:12that. But more the intent or the purpose
- 4:01:15of using it is to really like section
- 4:01:18things off.
- 4:01:21Yeah.
- 4:01:23Okay. Good. Good. Good. Good. Glad to
- 4:01:26hear that. Okay.
- 4:01:36Okay.
- 4:01:38Let's wrap up today by talking about
- 4:01:40comments. So, uh what are comments?
- 4:01:43These are like annotations or notes to
- 4:01:47yourself that are completely ignored by
- 4:01:52the interpreter.
- 4:01:53So when your code gets executed, the
- 4:01:56comment does not play any role in what
- 4:01:59gets executed. The interpreter will
- 4:02:00actually just completely ignore it. The
- 4:02:02moment it sees the comment, it will just
- 4:02:04ignore it and go to the next line.
- 4:02:07So the purpose of it is for humans to
- 4:02:10leave a note to other humans reading the
- 4:02:12code and that is very powerful is to be
- 4:02:16able to read those to to leave those
- 4:02:18notes and not have it affect the actual
- 4:02:22code that's uh that's actually executed.
- 4:02:25So there's multiple ways to make
- 4:02:28comments inside of Python. The most
- 4:02:30basic is to use the pound symbol. So the
- 4:02:34remember we cannot use pound symbols to
- 4:02:36name anything.
- 4:02:38This is why because the pound symbol is
- 4:02:41res reserved for making comments. So you
- 4:02:44you when you have a pound symbol like
- 4:02:47this uh that immediately signals to the
- 4:02:50interpreter everything else that follows
- 4:02:53that on this line is a comment. Any
- 4:02:56other text that follows that on that
- 4:02:58line is a comment. And usually your
- 4:03:00editor like VS Code, Jupiter, Collab
- 4:03:05will color that differently, maybe like
- 4:03:08a a like you can see in here, this is a
- 4:03:10Jupiter example. You can see it's kind
- 4:03:12of a light gray,
- 4:03:14greenish gray
- 4:03:16um to signal that this is a comment. Um
- 4:03:21now you may be asking why would we have
- 4:03:23comments? Again, you are going to look
- 4:03:25back at code weeks later,
- 4:03:29especially in this program. You're going
- 4:03:30to look at code in review and be like,
- 4:03:32what what was I doing there? If you
- 4:03:36leave a comment, you'll remember what
- 4:03:38you were doing there, why you had that
- 4:03:40line. Um, and not only that, like when
- 4:03:43you share your code with others, which
- 4:03:45in the real world you would be doing,
- 4:03:47collaborating with others, right? Adding
- 4:03:50in those comments can be really
- 4:03:52beneficial to do. So
- 4:03:55you will see me
- 4:03:58throughout the program. I'm going to
- 4:04:00leave a lot of comments on our demos and
- 4:04:02our notebooks that we work on together
- 4:04:04in the live sessions. I will leave
- 4:04:06comments mainly to call out certain
- 4:04:09things like I will say this is a really
- 4:04:11important step or we are doing this
- 4:04:14because I will leave a lot of comments
- 4:04:17and I encourage you guys to do the same
- 4:04:18in your own code. Um, remember they're
- 4:04:22free. They're they get ignored by the
- 4:04:24interpreter. They don't affect anything.
- 4:04:26They are notes to yourself. So, use them
- 4:04:29accordingly. Um, and you know, there's
- 4:04:32actually multiple ways to leave
- 4:04:34comments, but this is I'll show us those
- 4:04:36as we go along. But this is the uh most
- 4:04:39basic is you just you you type in a
- 4:04:42pound symbol and then everything else
- 4:04:44that follows that is uh is your comment.
- 4:04:50Okay.
- 4:04:52All right, guys. That's it for today.
- 4:04:55Um, what a great first session. Thank
- 4:04:57you guys. Thank you guys for being
- 4:04:58patient. um going through the setup of
- 4:05:01some of those tools. I hope you landed
- 4:05:02on one that worked for you. Um you know,
- 4:05:06use that one going forward, please. If
- 4:05:08it's collab, use that. Jupiter, use
- 4:05:10that. VS Code, use that. Whatever you
- 4:05:13you uh feel most comfortable with,
- 4:05:15please use that. Um we have a lot to
- 4:05:17cover still. You know, we're going to um
- 4:05:20continue on Wednesday. Uh we were we
- 4:05:24will uh continue talking about Python.
- 4:05:26We're just getting our feet wet a little
- 4:05:28bit on on Python. A lot more to cover.
- 4:05:31We're going to get into the the more
- 4:05:33nitty-gritty of the code. So, it'll be
- 4:05:35really fun. We'll cover if else loops,
- 4:05:39um how to control the flow of our
- 4:05:40program. Um we will do that. This is
- 4:05:44where we left off was writing comments
- 4:05:47uh in Python code, which I I did want to
- 4:05:49remind us of how to do that. it is going
- 4:05:52to be using the uh pound symbol to
- 4:05:56initiate a comment. And basically the
- 4:05:58Python interpreter will ignore
- 4:05:59everything else on that line. Uh it it
- 4:06:03treats all of that text as a comment.
- 4:06:05And again like comments are free. You
- 4:06:08might as well use them to your advantage
- 4:06:10to kind of uh leave a note to yourself
- 4:06:12of hey this is what this code is doing.
- 4:06:15Um so that when you come back and read
- 4:06:17it uh you can understand it better. So I
- 4:06:20encourage you guys like when we do demos
- 4:06:24uh and we will do a lot of demos um
- 4:06:27especially today leave comments you know
- 4:06:30put comments in there so so you can make
- 4:06:34a note to yourself what this code is
- 4:06:36doing um so you will get I think it'll
- 4:06:40be good to get in the habit of leaving
- 4:06:41comments uh to kind of mark up the code
- 4:06:44to to kind of remind yourself oh this is
- 4:06:47what it was doing when you look back at
- 4:06:49uh in the future.
- 4:06:52Okay. So we had ended with that.
- 4:06:56What I wanted to do was move on into the
- 4:06:59next slide. So talk about um basically
- 4:07:02how we display output to the screen
- 4:07:05which we've already seen an example of
- 4:07:06when we did the hello world which is the
- 4:07:08print function on the right. So this is
- 4:07:11a by the way this is a Python function
- 4:07:14and you know it's a function because
- 4:07:18of these parentheses. So these
- 4:07:21parenthesis signal that this is a
- 4:07:23function because it expects some sort of
- 4:07:25input to go inside of those parentheses.
- 4:07:28And and the input that would go inside
- 4:07:30of there is going to be text like some
- 4:07:34sort of uh some sort of text that
- 4:07:36belongs inside of quotes. and whatever
- 4:07:39we put there um will display on the
- 4:07:43screen. So that's useful for us to like
- 4:07:45display information
- 4:07:47um print we would say we are printing
- 4:07:49out information to the screen. Um so if
- 4:07:53we want to know the value of a variable
- 4:07:55or the value of something that we are
- 4:07:57doing a calculation with or uh you know
- 4:08:00sanity check something in our code we we
- 4:08:02can print it out which would be using
- 4:08:05the print function and it will display
- 4:08:07that value onto the screen. So we will
- 4:08:10use the print function quite a bit. Um
- 4:08:14you know we haven't learned what
- 4:08:15functions are but uh functions in Python
- 4:08:18are you know um designed to be uh chunks
- 4:08:22of code that execute and do something
- 4:08:25and they take arguments and you know it
- 4:08:28takes an argument because of the
- 4:08:29parenthesis that is um signaling that
- 4:08:31there should be some something inside of
- 4:08:33this parenthesis here which is going to
- 4:08:36be uh text. So whatever text you want to
- 4:08:38display or maybe some variable you want
- 4:08:40to display um that would go inside of
- 4:08:43there. So we'll get the hang of using
- 4:08:45the print function as we go along but
- 4:08:47just wanted to call out that's the
- 4:08:49primary methodology of kind of um
- 4:08:52displaying something on the screen if we
- 4:08:54want to print function. Um now the
- 4:08:58reverse of that is uh asking a user to
- 4:09:03uh input some data. So that would be
- 4:09:05this input function and um this is
- 4:09:09something that uh as you can see an
- 4:09:12example below is we can put some text
- 4:09:15inside of this parenthesis. So again a
- 4:09:17function it has those parenthesis that
- 4:09:19signals it's a it's a function. Um,
- 4:09:23and we can put some text in there which
- 4:09:25would be kind of what displays in to the
- 4:09:29user as kind of a prompt like here. Uh,
- 4:09:33enter your name and that would display
- 4:09:35on the screen and then there would be a
- 4:09:37box next to it. I'm going to show us
- 4:09:39this. I'm going to actually run this
- 4:09:41inside of a notebook in a minute. But
- 4:09:44then there would be a box that displays
- 4:09:46that that would say um hey you know
- 4:09:50enter your name and then you can type
- 4:09:52input in uh and and then when you hit
- 4:09:55enter it will save that input into into
- 4:09:58this variable called name. So um and
- 4:10:01remember name this is a valid identifier
- 4:10:04because it starts with a lowercase n um
- 4:10:08which is fine and it it has all valid
- 4:10:11characters. It doesn't have any wacky,
- 4:10:13you know, uh, pound symbol or at or
- 4:10:17anything crazy. So, it's it's a decent
- 4:10:19identifier. Um,
- 4:10:22so name name would be okay. And so input
- 4:10:25is whenever you want to get whenever you
- 4:10:28want to allow the user to input
- 4:10:30something like it'll bring up a text box
- 4:10:32and they can enter some data. Um, and
- 4:10:35that will be saved in this whatever
- 4:10:38variable you name this you set equal to
- 4:10:40input. Um, and then you can see like as
- 4:10:44soon as we put that in, we immediately
- 4:10:45dis we can display it. So we we print
- 4:10:48hello and then comma name which
- 4:10:50references whatever we stored whatever
- 4:10:53the user input there. So I'm I'm going
- 4:10:54to show us an example of that. Um, but
- 4:10:57input is what get is our primary way of
- 4:11:01getting input from the from the the user
- 4:11:04in a text box so we can use that data in
- 4:11:07our program.
- 4:11:08Print is our primary way of displaying
- 4:11:11data that we already have in our code.
- 4:11:13We can print it which will display it.
- 4:11:16Um
- 4:11:17so we're going to see many examples of
- 4:11:19these as along but just wanted to call
- 4:11:20out those two. These
- 4:11:24functions by the way are built into
- 4:11:26Python. So we don't need to create them
- 4:11:29ourselves. They already exist. They're
- 4:11:31already built into Python. Um nothing
- 4:11:34special we need to do to use them. We
- 4:11:36can just use them right out of the box.
- 4:11:38So again, we'll see this in our in our
- 4:11:39code examples that we're going to do in
- 4:11:41a minute.
- 4:11:45Um, where exactly would an end user be?
- 4:11:47So maybe we ask them for some input. Um,
- 4:11:51and then we do like so we ask them for
- 4:11:53some like their name, their email, their
- 4:11:57uh date of birth, those kind of things.
- 4:12:00We can ask in the input and then we
- 4:12:01maybe we store them in a database or we
- 4:12:03do something with it in the Python code.
- 4:12:05So um whenever you want to accept input
- 4:12:09from a end user that's when you would
- 4:12:11use this input.
- 4:12:14It just depends on the application
- 4:12:17right on the application like what kind
- 4:12:18of input data you you want to accept
- 4:12:20from the from the user.
- 4:12:27Okay. Okay. So, I'm going to show us
- 4:12:28this.
- 4:12:33Um, before we do that demo,
- 4:12:36um, let me ask you guys, which of the
- 4:12:38following do do we remember from Monday?
- 4:12:40Which of the following identifier names
- 4:12:43follows Python's rules
- 4:12:46and best practices for readability?
- 4:12:50So not only so you should be looking for
- 4:12:52the answer choice here that follows the
- 4:12:54rules but also is a meaningful name.
- 4:13:00A lot of different choices. Okay.
- 4:13:04By the way,
- 4:13:06which let me ask let me I'll come back
- 4:13:08to the answer to the original question,
- 4:13:10but let me ask this alternative
- 4:13:11question. Which one of these is not
- 4:13:14valid? Meaning it would Python would
- 4:13:17throw an error how you use it.
- 4:13:20Which one of these is not valid in
- 4:13:22general?
- 4:13:25Cool. Great. It is C. You guys were
- 4:13:27right on top of that. Very good. So C is
- 4:13:28not valid. It names of things cannot
- 4:13:31start with a number. So So that's um C
- 4:13:35is completely invalid in general and
- 4:13:37that would produce uh an error.
- 4:13:41Right. Starts with a number. Exactly.
- 4:13:43Which we cannot do. Now starting with an
- 4:13:46underscore is okay. that's allowed. So
- 4:13:49that's not an issue. And having a number
- 4:13:51be second after the underscore is okay
- 4:13:55as well. So technically A and B would
- 4:13:59follow the rules. So those definitely
- 4:14:01follow the rules. Um now are they
- 4:14:06readable and meaningful is the question.
- 4:14:10I would argue that possibly not. Var 123
- 4:14:15is pretty generic. I would argue that
- 4:14:18even though it's valid, like Python
- 4:14:20would not have any errors with that uh
- 4:14:22variable name, it's not very meaningful.
- 4:14:24It's too generic. It's almost as if we
- 4:14:26just called something X. We just called
- 4:14:28something var 23, that's probably not
- 4:14:32going to be meaningful to us and we're
- 4:14:34not going to understand what that really
- 4:14:36represents. If somebody were to come
- 4:14:37along and read it and see VAR 123,
- 4:14:41that's probably not that great of a
- 4:14:43name. It's not telling us exactly what
- 4:14:45that represents. So, I would say A is
- 4:14:49likely not um
- 4:14:52A is likely not uh a good choice and C
- 4:14:56we know is invalid. So, really I think
- 4:14:57the only two options you could argue are
- 4:15:00B and D. I think D is a really good
- 4:15:02answer. It um it it follows the rules.
- 4:15:07Uh underscores are fine. Everything is
- 4:15:09lowercase. That's fine. Um so it's valid
- 4:15:13but it also is meaningful as a name
- 4:15:15right so final result value um we we
- 4:15:19should probably like in our code we
- 4:15:21would have context we would know what
- 4:15:22that means okay this is our final result
- 4:15:25um so that's a that's a pretty good name
- 4:15:27for something um you know this one is
- 4:15:32okay it's just not that readable 321
- 4:15:35customer details DB table it's okay I
- 4:15:40don't think It's um it's not the worst.
- 4:15:42It it definitely would work, but it's um
- 4:15:46kind of a clunky name. I'm sure we could
- 4:15:48come up with something better, but it
- 4:15:50would work technically. There'd be no
- 4:15:52issues with it.
- 4:15:55Okay. So, I think D is probably the best
- 4:15:58choice, but B is valid, too. I think B
- 4:16:00could work for this. D and B, I think,
- 4:16:01are okay.
- 4:16:05Okay. Good. Good. You guys are right on
- 4:16:07top of that. you have a good I think you
- 4:16:08have a good feel for what the allowed
- 4:16:11names for things are which is good.
- 4:16:15Okay.
- 4:16:17Um
- 4:16:19I'm going to then swap over to this
- 4:16:22demo. So you guys should have the uh
- 4:16:26demos um and we kind of went through
- 4:16:29some of those first few last time to get
- 4:16:31you set up on Cola and Jupiter and VS
- 4:16:34Code. Um, I am going to be using Collab
- 4:16:39for most of these, but feel free to use
- 4:16:42whatever you want to use. If you want to
- 4:16:43use Jupiter, if you want to use VS Code
- 4:16:45and run your notebooks on your own
- 4:16:47machine, feel free to use whatever you
- 4:16:49want to use. I'm going to be using
- 4:16:50Collab just for the simplicity of it.
- 4:16:54Um, so, so this demo will walk us
- 4:16:57through, um, opening up a new Collab
- 4:17:00notebook and then running those input
- 4:17:02and print. So, some examples with input
- 4:17:04and print. Um, so we'll do that
- 4:17:06together. Let me go over to that demo.
- 4:17:13So, if you're following along, we are
- 4:17:15going to be doing um demo 4. So, it
- 4:17:20should be lesson one, demo 4.
- 4:17:27Um, do you guys have this? Give you a
- 4:17:30moment to to pull that up. Lesson one,
- 4:17:32demo 4.
- 4:17:35You guys have access to this one. So, we
- 4:17:37did we did one, two, and three on
- 4:17:39Monday, which were just getting those
- 4:17:41environments set up. So, this is demo 4.
- 4:17:45Um, which again, I know step one says
- 4:17:48open collab. Feel free to open your own
- 4:17:50notebook in VS Code or open your own
- 4:17:52notebook in Jupiter as well, whatever
- 4:17:53you're most comfortable with. Um,
- 4:17:57I'm going to be using the collab to to
- 4:17:58do this, but feel free to use whatever
- 4:18:01works. You're we really just need a
- 4:18:03notebook to be able to run this code.
- 4:18:05So, however you're running notebooks,
- 4:18:07whether that's in VS Code or Jupiter or
- 4:18:09Collab, I any of those, either one is uh
- 4:18:13perfectly fine.
- 4:18:15So, so step one is to open up a
- 4:18:18notebook. I'm going to do it in Collab,
- 4:18:19which is what this says. Um, and then
- 4:18:22you can make a new notebook and then
- 4:18:24rename it to my first program. I'm going
- 4:18:26to do that in a second. And then, um, so
- 4:18:30I'm going to walk through this live with
- 4:18:32you, but just showing you some of the
- 4:18:33steps we're going to do. Um, the first
- 4:18:36thing we're going to do is is just
- 4:18:38practice doing the print hello world
- 4:18:40again so that we can um, execute a print
- 4:18:44statement. So, we'll practice that.
- 4:18:47We're going to make a second cell. Um,
- 4:18:51which we can do in Collab or VS Code or
- 4:18:55Jupiter by hitting the plus button.
- 4:18:56There's usually a plus. Uh, you can see
- 4:18:59it here. Uh, in multiple places in
- 4:19:02Collab, you can do it right below an
- 4:19:03existing cell or there's always a plus
- 4:19:06code here, which is kind of what you
- 4:19:08have in Jupiter. Usually in Jupiter, you
- 4:19:09have a plus button. So, you can just hit
- 4:19:12hit that plus button, it'll make a new
- 4:19:13cell. Um, and so we'll make a new cell
- 4:19:17so we can write some more code.
- 4:19:21Um,
- 4:19:22and then in this new one, we are going
- 4:19:24to practice doing some comments.
- 4:19:27We're going to practice doing some
- 4:19:28comments and then um see how we can do
- 4:19:32uh some more print statements. Okay, so
- 4:19:36let's do that. Let me jump over to
- 4:19:37Collab. Let's walk through these first
- 4:19:39few steps together. Um, and then uh
- 4:19:43we'll come back to this and finish out
- 4:19:45the rest of the steps because we're also
- 4:19:46going to do input. So I'm going to show
- 4:19:48you how to do these uh input which will
- 4:19:51you can see here like the input is going
- 4:19:53to create a text box where you can put
- 4:19:56input and it will you hit enter it will
- 4:19:58save it for you. So input allows you to
- 4:20:01get input from the keyboard
- 4:20:04and save that into a variable to use for
- 4:20:08later.
- 4:20:10Okay. So, let's jump over to
- 4:20:15um let's jump over to
- 4:20:18I'll show us I'll show us in a second.
- 4:20:20How do you rename it?
- 4:20:22I'll show you. Let me jump over to
- 4:20:24collab.
- 4:20:28Um
- 4:20:36okay.
- 4:20:37So, I am over lesson one, demo 4. Yep,
- 4:20:42that's the one we're doing.
- 4:20:45Okay. So, I am in Collab. I'm going to
- 4:20:46start a new notebook.
- 4:20:51Start a new notebook in Collab. Uh, so
- 4:20:53now I'm here. I'm just on a fresh
- 4:20:54notebook. Um, nothing that interesting
- 4:20:57going on. Here's how you rename it. just
- 4:20:59go up to this box on the left
- 4:21:03and almost like a Google doc just just
- 4:21:06uh click into that name and then start
- 4:21:10typing to erase it. So see how I'm like
- 4:21:12hovering over that name and then I'm
- 4:21:14clicking on it and then I can start
- 4:21:16typing to erase it. So I can we can name
- 4:21:18this my
- 4:21:20first program
- 4:21:23and then hit enter and it will save
- 4:21:25that.
- 4:21:33Oh yeah. So if you're in VS Code um do
- 4:21:37to to rename it do file and then save as
- 4:21:42and then you can give a new name to it.
- 4:21:45file, save as.
- 4:21:49Okay, that's how you can rename it in VS
- 4:21:51Code.
- 4:21:58All right, let's do let's do the first
- 4:22:01step. Um, let's do print. So, we're
- 4:22:04going to do print. So, type in print and
- 4:22:07we can uh we can do parenthesis.
- 4:22:11Um, remember this is a function. So, we
- 4:22:14need the we need the parenthesis to
- 4:22:16signal that we want to put some text
- 4:22:18inside of this print function. And then
- 4:22:20you want to do uh you want to do quotes.
- 4:22:24You want to do quotes, the double quotes
- 4:22:26there, in order to allow us to put in
- 4:22:29some text. So, Python will interpret
- 4:22:32what's inside of the quotes as text and
- 4:22:34it will display that text. So we can do
- 4:22:37hello world my first
- 4:22:43Python program.
- 4:22:46Okay. And then we can run it. So feel
- 4:22:49free to put whatever text in here. It
- 4:22:51doesn't really matter exactly what it
- 4:22:52is, but you put some text in there
- 4:22:54between the parenthesis and then hit
- 4:22:56run.
- 4:23:01And the notebook will take a second to
- 4:23:05connect. And then there it is. Right? So
- 4:23:07then you see the the text displayed on
- 4:23:09the screen.
- 4:23:12Try that out. Are you guys able to run
- 4:23:14the print
- 4:23:16in your Jupyter what whether it's
- 4:23:17Collab, whether it's uh Jupyter
- 4:23:19notebook, whether it's VS Code. Can you
- 4:23:21run the print
- 4:23:32install?
- 4:23:35Yeah, I installed that um in VS Code.
- 4:23:38Yep. Try installing that.
- 4:23:48Okay, great. You guys were able to run
- 4:23:49that. Very good. Very good. Okay.
- 4:23:57No, you don't want to save it as a JSON
- 4:23:59file. You want to save it as a pyb just
- 4:24:02like this. See how this one is uh IP
- 4:24:05YMBB?
- 4:24:07That's the format you want. Remember
- 4:24:09that is interactive Python notebook.
- 4:24:13You want that file IPY MB.
- 4:24:25Uh perfect. Yeah, you get you got it to
- 4:24:27run.
- 4:24:32You don't have any extension? No. If
- 4:24:34you're in VS Code, remember from Monday,
- 4:24:36you need to install the the Jupiter
- 4:24:39extension.
- 4:24:45If you're in VS Code, you got to install
- 4:24:47the Jupiter extension.
- 4:24:57You have to manually so manually save
- 4:24:59it.
- 4:25:03You can save it as a py. I would do ipy
- 4:25:06so you can open it in collab.
- 4:25:09Type it yourself.
- 4:25:11Type overwrite what's there and type it.
- 4:25:14Type in um my notebook whatever the name
- 4:25:18is.
- 4:25:20Type it out yourself if you can. like
- 4:25:23save as and then
- 4:25:26type out the full file name yourself.
- 4:25:32Now let's practice a comment. Let's
- 4:25:35practice a comment. So let's build let's
- 4:25:36do a new code cell. So we made a you can
- 4:25:40either do it here. If you hover over
- 4:25:42your cell, you can hit plus to build a
- 4:25:44new code cell or you can hit plus here
- 4:25:46to make a new code cell. So let's do
- 4:25:48that.
- 4:25:56You should be saving.
- 4:25:59Don't worry about the type. Just type in
- 4:26:01the namey imm. I don't think you need
- 4:26:04to.
- 4:26:08Or you could just hit what you could do
- 4:26:09is you could just hit save and then in
- 4:26:12your file explorer you could just rename
- 4:26:15it.
- 4:26:17So if you just save it will save it to
- 4:26:18the default location and then just and
- 4:26:20then just rename it.
- 4:26:23So maybe try that route. Just just do
- 4:26:24save. Just save it. And then it should
- 4:26:30it should try to save it as IPymbb.
- 4:26:38Okay. Okay. Let's practice. Um, so the
- 4:26:42next step in the demo, if you're
- 4:26:43following along in the demo document, it
- 4:26:46wants us to do, so we did the print. We
- 4:26:49want to do um a practice some comments.
- 4:26:57>> Okay, perfect. Uh, let's practice some
- 4:27:00comments. So, um, remember I told you
- 4:27:02that we can do, uh, comments with the
- 4:27:04pound sum. So, this is a
- 4:27:09comment. So practice writing a comment.
- 4:27:11Remember you start a comment with a
- 4:27:13pound symbol.
- 4:27:16Um it will get ignored
- 4:27:19by the interpreter
- 4:27:23interpreter. So feel free to type in
- 4:27:25whatever text you want. I'm just
- 4:27:27reminding us that whatever the comment
- 4:27:29is is going to be ignored and we can
- 4:27:32have whatever code below that that we
- 4:27:34want to have and that comment will get
- 4:27:36completely ignored. So let's do another
- 4:27:38print. So write a comment,
- 4:27:41hit enter. Immediately below that in a
- 4:27:44new line, let's do another print.
- 4:27:48This code
- 4:27:50gets executed.
- 4:27:54So we know this print statement is going
- 4:27:56to get executed, but this comment is
- 4:27:59going to be ignored by the interpreter.
- 4:28:02So let's run that.
- 4:28:04So this code gets executed. This comment
- 4:28:07gets completely ignored,
- 4:28:10right? That comment gets completely
- 4:28:11ignored, which is great. Try writing a
- 4:28:14comment. Are you guys able to write
- 4:28:15comments?
- 4:28:17So write a comment and then try writing
- 4:28:19a print statement right after it.
- 4:28:22And and feel free to put whatever text
- 4:28:24you want inside the comment. And feel
- 4:28:25free to
- 4:28:30uh for the comment, is the space after
- 4:28:32the pound symbol required? No, it's not.
- 4:28:34So, we could test it out. So, I removed
- 4:28:36the space. Doesn't matter. It It's just
- 4:28:39for readability. I usually like doing
- 4:28:42that so that I have some space after it.
- 4:28:44And this is a little It's just a little
- 4:28:46bit more readable, right? It's not like
- 4:28:47mixed together.
- 4:28:52It's just for readability.
- 4:28:55Great. You guys wrote a comment. Okay.
- 4:28:57Perfect. Perfect. We're able to write
- 4:28:59comments. Really great. Okay.
- 4:29:02Okay.
- 4:29:06If you put multiple code lines, do we
- 4:29:09need any separator? Like, no, they just
- 4:29:12go on new lines. So, do you mean like a
- 4:29:14second print statement? Let's We could
- 4:29:16try that. Let's do a secondary print
- 4:29:18statement. So, we can do print.
- 4:29:21Um, this one is on the next line. No
- 4:29:28separator
- 4:29:30needed.
- 4:29:34Do you see that? See how it's on its
- 4:29:35own? I did a print right below this
- 4:29:38other print. And as long as they're on
- 4:29:39their own line, that's okay. They just
- 4:29:42need to be on their own lines. They
- 4:29:43don't need any separator.
- 4:29:46If we run this, then this one gets exe.
- 4:29:50Then see how this is now printed out
- 4:29:51below it. Right there.
- 4:29:58Is there any character limit on the
- 4:29:59comments? Uh, no. There's no character
- 4:30:02limit. Um,
- 4:30:04but
- 4:30:06there's no character limit, but a good
- 4:30:08practice is to not like you don't want
- 4:30:10this to be super long and to to take up
- 4:30:13the whole screen, right? Because then
- 4:30:15it's not really readable.
- 4:30:17So, there's no limit, but you don't want
- 4:30:19to have overly
- 4:30:22long comments. You want to keep them
- 4:30:24kind of concise and short.
- 4:30:26So just so you can read them and they're
- 4:30:28they don't take up a lot of space.
- 4:30:34Not able to add print statement below.
- 4:30:36Why? Why is that?
- 4:30:40You should be able to should be able to
- 4:30:42have a print right below this print.
- 4:30:44Shouldn't be anything that make sure you
- 4:30:46close this parenthesis. Make sure every
- 4:30:48print needs to close the parenthesis
- 4:30:53and they all you also need to close the
- 4:30:55quotes. So close this quote, close this
- 4:30:58quote within within the print
- 4:31:02that needs to be done. So you should be
- 4:31:06able to run I'll paste this for you guys
- 4:31:08in the chat. Should be able to run this
- 4:31:19All right. One thing I want to show you
- 4:31:20guys is just like the demo says in the
- 4:31:22word document, um you can do multi-line
- 4:31:25comments. So if you need to do a lot of
- 4:31:28comments, all you need to do is triple
- 4:31:31quotes. So triple quote,
- 4:31:34then um triple quote, and then
- 4:31:38everything in between.
- 4:31:44That's interesting that it did that.
- 4:31:47Yeah. So, we can do a pound symbol,
- 4:31:50pound symbol, pound symbol,
- 4:31:55pound symbol, and that that should all
- 4:31:57work. So, we can do that.
- 4:32:01Yeah, I think it's a collab thing, but
- 4:32:03normally in in like Jupiter or in
- 4:32:05Python, it it will work just fine. But
- 4:32:08like in collab, I think they don't like
- 4:32:10the triple quotes.
- 4:32:12But yeah, do you guys see do you guys
- 4:32:14see how I just did it like this with the
- 4:32:15pound symbols? That's okay, too.
- 4:32:19Everything between these
- 4:32:23uh pound symbols is a comment and is
- 4:32:26ignored. So now we should be able to run
- 4:32:28that. So there we go. Everything gets
- 4:32:30ignored there. Does that make sense to
- 4:32:32us? The pound symbol comments
- 4:32:41does the does the using the pound
- 4:32:43symbols. So notice how we use that to do
- 4:32:45multiple lines of comments. So we did
- 4:32:48one here, we did one here. We can have
- 4:32:50as many we can have
- 4:32:53um as many
- 4:32:56uh comment lines as we want and they
- 4:32:59will all get ignored.
- 4:33:10What is those? It's supposed to be
- 4:33:11multi-line comments, but for some reason
- 4:33:13it's not working. Um, it so the the
- 4:33:17triple quote is supposed to be like
- 4:33:19representing that you can have a whole
- 4:33:21block of comments.
- 4:33:25I don't know why it's not working in
- 4:33:26collab for me.
- 4:33:31It's working for you. Okay. Okay. I
- 4:33:32don't know why it's not working.
- 4:33:42Single quote.
- 4:33:55It's still It still displays here, which
- 4:33:57I don't get why that's happening.
- 4:34:05It's kind of weird to me.
- 4:34:11Yeah, I don't get why inconsistency. It
- 4:34:13usually It usually works for me. I don't
- 4:34:15get that at all.
- 4:34:41Still still doesn't work for me. I don't
- 4:34:42know why that doesn't
- 4:34:56Yeah.
- 4:34:57I don't get why that's not really liking
- 4:34:59those triple quotes. Oh well. I mean,
- 4:35:02not a big deal. We can just do
- 4:35:05Okay, we can do we can try single.
- 4:35:12Still doesn't work.
- 4:35:25Yeah. Oh well, we can do a pound symbol.
- 4:35:29That will always work. Pound symbol is
- 4:35:31honestly more popular anyway. Most code
- 4:35:34that you see in the wild will have pound
- 4:35:36symbols wherever they're doing um
- 4:35:39wherever they're doing uh comments. So
- 4:35:42that it's fine. Just use a paddle for
- 4:35:44now.
- 4:35:57Uh yeah, that's correct. I don't know
- 4:35:59why that's that's correct. Um I don't
- 4:36:02know why collab doesn't seem to like
- 4:36:03that. It should be ignored
- 4:36:06um generally with the triple quotes, but
- 4:36:09uh that's okay. I'm not too concerned
- 4:36:11about it for now. I guess what you and
- 4:36:14when I do comments, you're usually going
- 4:36:15to see me using the pound symbol
- 4:36:17anyways. It' be very rare that I would
- 4:36:18need to do uh quotes.
- 4:36:27Yeah, it's weird that collab doesn't
- 4:36:29work very consistently. That's okay.
- 4:36:32All right. What I want to show us is I
- 4:36:36want to move on to the input. So, I want
- 4:36:38to I want you guys to see
- 4:36:40I want you guys to type in this code
- 4:36:42here that will take input from a text
- 4:36:45box and save it into a variable called
- 4:36:47name. So, the code we're going to do is
- 4:36:50going to be like this. It's going to be
- 4:36:52name equals input
- 4:36:56and then we'll put um please
- 4:37:00enter your name.
- 4:37:04Okay.
- 4:37:06So this, by the way, I'm going to
- 4:37:09comment this code here. Um, this code
- 4:37:13should
- 4:37:15create
- 4:37:17a text box for us to put in our name.
- 4:37:24Okay, so that's what should happen. So
- 4:37:26when we run this, um, it should pop open
- 4:37:30a text box right below this. And we can
- 4:37:33type in our name and hit enter. And when
- 4:37:35we do that, it will store that result in
- 4:37:38this variable called name, which we can
- 4:37:39use uh wherever we want to in the code.
- 4:37:42So if I hit run, there's that text box.
- 4:37:45Do you guys see that? There's the text
- 4:37:48box. And see how it says, please enter
- 4:37:51your name. And so we can type in our
- 4:37:53name.
- 4:37:56And we hit enter. And there it's stored
- 4:37:59in the name. We can even um display name
- 4:38:03by doing print and then the name which
- 4:38:07will display uh the name that we stored
- 4:38:10when we did the input.
- 4:38:12So try this one out. Try this code out
- 4:38:16for yourself. Try typing input
- 4:38:18parenthesis
- 4:38:20and then you want to have some text
- 4:38:21there. It doesn't matter exactly what it
- 4:38:22is, but something like please enter your
- 4:38:24name or enter your name.
- 4:38:27Try that out. And then it should store
- 4:38:30uh you should be able to type in the box
- 4:38:32that shows up. Hit enter on your
- 4:38:34keyboard. It should save that. And then
- 4:38:36you can um print it out. You can print
- 4:38:39out that name which will um
- 4:38:42display that whatever we typed in
- 4:38:44before.
- 4:38:52What does it look like, Roberto? What
- 4:38:54does it look like? Were
- 4:39:02you Were other people able to run this?
- 4:39:06Oh, yeah. Thank you, Melanie. Yeah, I
- 4:39:08see that. Perfect.
- 4:39:11Perfect. That looks good to me.
- 4:39:20Uh, you don't need a space. um it just
- 4:39:24looks nice, right? It's so that's a good
- 4:39:27practice to have the space so that uh
- 4:39:29this code is um evenly spaced out and it
- 4:39:33looks nicer on the on the screen.
- 4:39:41Um name equals input print hello there
- 4:39:45uh name
- 4:39:52You need Yeah. So, uh, Roberto, you need
- 4:39:55a you need a comma after after the
- 4:39:58quotes.
- 4:40:00After the quotes, you need a comma after
- 4:40:02the quotes to signal to Python that
- 4:40:05you're putting in you have you have some
- 4:40:07text and then an additional input.
- 4:40:12So, it need it needs to be more like it
- 4:40:14needs to be like this. print. Um, hello
- 4:40:17there.
- 4:40:19And then you need an extra comma.
- 4:40:24See how I have an extra comma after the
- 4:40:25quote. You need you need that.
- 4:40:32Sorry. Now, Roberto's uh Kiati.
- 4:40:36Hope I'm pronouncing that right.
- 4:40:40Okay. So, do we feel good about input
- 4:40:42and what it does?
- 4:40:45Perfect. Do we feel good about input and
- 4:40:47what it does? It It brings up a text
- 4:40:49box.
- 4:40:53It Did you hit enter, Roberto? To like
- 4:40:55Were you able to type something in and
- 4:40:57hit It's going to run until you hit
- 4:40:59enter.
- 4:41:00You have to type in the text and then
- 4:41:02hit enter into the box.
- 4:41:06So, let me rerun this. So, it See how
- 4:41:10it's still running? See how this like
- 4:41:12it's going to keep running forever until
- 4:41:15I type something in
- 4:41:21and then when I hit enter it will stop.
- 4:41:28What does your code look like?
- 4:41:52Okay, that looks right.
- 4:41:56Try try stopping it and rerunning it.
- 4:42:03Try try hitting the stop button and then
- 4:42:06rerun it.
- 4:42:18um you should so yeah you should put
- 4:42:20that in a different cell. So if you if
- 4:42:24you separate your code into individual
- 4:42:25cells so you could do like um you could
- 4:42:30do name. So we could we could separate
- 4:42:32this. So this code is the only code
- 4:42:36that's running in this cell.
- 4:42:49That doesn't make sense. Something else
- 4:42:52is
- 4:42:54that doesn't make sense cuz like this
- 4:42:56collab tab is only taking up 235
- 4:43:00megabytes.
- 4:43:02So something is
- 4:43:05chewing up your memory that's not really
- 4:43:07I I can't imagine. Are you using collab?
- 4:43:10You can see like it's not using that
- 4:43:13much. Only 240
- 4:43:15230ish.
- 4:43:17Yeah, I don't think I don't think Collab
- 4:43:19is the culprit unless you loaded in some
- 4:43:21really massive data or something.
- 4:43:25I can't imagine that's the issue.
- 4:43:29You did. You loaded in data. That's
- 4:43:32I mean Yeah. Then it's going to it's
- 4:43:34going to take in memory. Oh, okay. Okay.
- 4:43:37Okay.
- 4:43:43Uh MJ, what are you on? Are you on
- 4:43:49Yeah. Could you screenshot it?
- 4:43:52If it's not working for you, could you
- 4:43:53try collab? If Could you try collab just
- 4:43:56for the sake of like getting it running?
- 4:43:59Things should work in Collab pretty
- 4:44:01easily.
- 4:44:04You're using Collab and nothing's
- 4:44:05working. Uh, are you making sure it's a
- 4:44:07code cell and not a text cell?
- 4:44:12It's not a text cell like this,
- 4:44:15which would be like,
- 4:44:18this is where it will be blue.
- 4:44:29Did you have that? You need to make sure
- 4:44:30it's code. Yeah.
- 4:44:34And when I run that, it's going to be
- 4:44:35it's going to display text. Yeah.
- 4:44:41Okay. Great.
- 4:44:45Glad that it's working. Great.
- 4:44:49Okay.
- 4:44:51Um All right. One more. Uh one more
- 4:44:54example what I want to show you guys is
- 4:44:56how to do how to include the name in a
- 4:44:59print statement. So if we do something
- 4:45:01like print. So um we can include the
- 4:45:06name in a print statement. So if we do
- 4:45:10something like print and then we have um
- 4:45:13hello there and then we have um this and
- 4:45:18then we have welcome to Python.
- 4:45:23um this will
- 4:45:26uh this will display all of that
- 4:45:29together. So notice that we can have as
- 4:45:32many um pieces of information that we
- 4:45:34want to display kind of one after the
- 4:45:36other as long as they're separated by
- 4:45:38these commas.
- 4:45:40So we have this uh text,
- 4:45:44this text because text is stored in that
- 4:45:47variable. Um, then this text and then we
- 4:45:54print that all out and we can have this
- 4:45:56whole collection of text displayed to
- 4:45:58the screen. Try that one out.
- 4:46:11Oh, they do the same thing. They do the
- 4:46:14same thing. So, the comma and the Sorry.
- 4:46:15Yeah, I just noticed the demo does a
- 4:46:17plus. They do the same thing in Python.
- 4:46:20So, we can swap that over to a plus.
- 4:46:21Both of them work.
- 4:46:24They have the same I shouldn't say they
- 4:46:25do the same thing, but they have the
- 4:46:27same effect.
- 4:46:29They have the same effect.
- 4:46:31Actually, there's no You need a little
- 4:46:33bit more spacing here. So, the comma
- 4:46:35gives you a little bit better uh
- 4:46:37spacing.
- 4:46:41So, what the Let me break this down.
- 4:46:43what the so plus
- 4:46:46plus um adds together
- 4:46:51uh text and so what we're doing here
- 4:46:54technically is adding all our text
- 4:46:56together and then displaying it. Um, so
- 4:46:59plus as together text and then the comma
- 4:47:03um,
- 4:47:05uh, prints out multiple pieces of text.
- 4:47:11So they they have the same effect, but
- 4:47:13yeah, you can use you can use either
- 4:47:15one.
- 4:47:25Okay.
- 4:47:27Um, one thing I wanted to show you guys
- 4:47:28too, by the way, in Collab, if you're
- 4:47:31working inside of Collab, I want you to
- 4:47:34hover over your name variable.
- 4:47:37So, if you just take your mouse and
- 4:47:39hover over that,
- 4:47:42do you guys see what it says here?
- 4:47:46Do you see how it says string name and
- 4:47:49then it has the value of that, which is
- 4:47:52which is my name. So, that's something
- 4:47:54cool about Collab is if you hover over
- 4:47:57variables, it will tell you what their
- 4:47:59type is. Now, we haven't learned about
- 4:48:01types, but any text inside of quotes is
- 4:48:05a string. It's it's what we would call a
- 4:48:07string. We're going to learn about that.
- 4:48:09And
- 4:48:11um
- 4:48:13we it also displays what data we
- 4:48:15currently have stored in that variable.
- 4:48:17So all you have to do is hover over a
- 4:48:19variable um to to see what the value is.
- 4:48:24Yeah, that's yeah, that's kind of a
- 4:48:26limitation of VS Code. That's true.
- 4:48:30It doesn't show you immediately on
- 4:48:32hovering.
- 4:48:47Don't see the value on hovering. So, um,
- 4:48:50click into the cell. You have to click
- 4:48:52into the cell and then hover over it.
- 4:48:55Click into the cell and then hover over
- 4:48:57it. It should it should work. Yeah, you
- 4:49:00have to click on the cell or whatever
- 4:49:02cell you're on and then uh hover over
- 4:49:06that and it should work.
- 4:49:10Uh, Mariel asks, "How do we integrate
- 4:49:12that Python code to a client application
- 4:49:14for a user to enter a value?
- 4:49:17um we would likely have a different set
- 4:49:20of code to do that. Um there is Python
- 4:49:23code that can get a UI and uh we we will
- 4:49:26see that um later on in the in the like
- 4:49:30way later on towards the end of the
- 4:49:32program. We'll see that um we can we can
- 4:49:34write Python code to do a UI essentially
- 4:49:38to to make like a almost like a web page
- 4:49:41for someone to enter some input. We'll
- 4:49:44see that uh much later on. So, we're not
- 4:49:47going to get to that right now. It's
- 4:49:48really complex.
- 4:49:55Um what is the purpose of having
- 4:49:58multiple cells? It's so that we can run
- 4:50:00individual pieces of code within those
- 4:50:02cells. It allows us to isolate, right?
- 4:50:05Because I can run I can run code inside
- 4:50:07of these cells and they don't affect any
- 4:50:09other cell. So, it's it's just for like
- 4:50:12debugging and isolation, which is nice,
- 4:50:16right? I don't need to worry about
- 4:50:17running all of it at once. I can run one
- 4:50:19cell at a time.
- 4:50:34Okay. Any other questions?
- 4:50:44Um, can we execute multiple lines
- 4:50:46together? Yes, we did that. Here I had
- 4:50:49multiple. So, I'll I'll show you again.
- 4:50:51I can do um print
- 4:50:54um this is one statement
- 4:50:58and then I can come down and do uh print
- 4:51:02um this is another and then maybe I can
- 4:51:06do some math.
- 4:51:13So you can have as many lines as you
- 4:51:15want
- 4:51:17within a cell.
- 4:51:22Within a cell, you can have as many
- 4:51:24lines of code as you want.
- 4:51:34Is there a way to tell it the order the
- 4:51:36cells execute? Um, you no, if you if you
- 4:51:40go up to um if you go up to run all,
- 4:51:44it's going to run them all in order from
- 4:51:45top to bottom. Uh, in order to tell
- 4:51:49which cells to execute, you can
- 4:51:51rearrange them. You can always like So,
- 4:51:54I could rearrange these cells, by the
- 4:51:56way, by I think there's a way to move it
- 4:51:58down.
- 4:52:00So, I can move it down. So now I'm
- 4:52:03rearranging. So you can move cells. I
- 4:52:05think you can even drag and drop them.
- 4:52:07So notice how I took the one that's at
- 4:52:09the very top and I'm moving it down.
- 4:52:13Otherwise, you have to click, right? You
- 4:52:15just have to like I can run them in any
- 4:52:17order. If I just click like if I click
- 4:52:20here, it will run that one first. If I
- 4:52:21go back up here, it will run that one
- 4:52:24next. So you just click around which
- 4:52:27ones you want to run. Does that make
- 4:52:29sense?
- 4:52:30I can run them in any order as long as I
- 4:52:32click on whatever order I want to do it
- 4:52:34in.
- 4:52:41Okay.
- 4:52:43All right.
- 4:52:45Perfect. So, that that wraps up that
- 4:52:47demo. I hope it was informative. I hope
- 4:52:49you saw the the print statement. Um
- 4:52:51we're going to see that many times. The
- 4:52:53input statement. Um that's pretty cool.
- 4:52:56Um and you got to run you got to run
- 4:52:58some Python. So, if it's your first time
- 4:53:00ever doing programming, congratulations.
- 4:53:02You ran some Python code. That is really
- 4:53:04exciting. Um, so glad we got to do that.
- 4:53:07Um, let's go back to our notes
- 4:53:15and then we'll um
- 4:53:18let let me share the screen.
- 4:53:26Okay.
- 4:53:28So the next thing on our agenda is to
- 4:53:32cover variables and data types. So I
- 4:53:34just said like text is the string data
- 4:53:37type but let's learn about all the
- 4:53:38different data types that are going to
- 4:53:39be available to us inside of Python and
- 4:53:42let's talk about variables. Um it's
- 4:53:44going to be a good discussion. So um I
- 4:53:47think what we'll do is we'll take a
- 4:53:48fivem minute break now and we come back
- 4:53:51and we can start this uh discussion
- 4:53:53about variables and data types. Um, so
- 4:53:56let's take uh a fivem minute break.
- 4:54:00And so let's try to be back um around
- 4:54:06uh 8:30.
- 4:54:13Okay.
- 4:54:21Try to be back around 8 8:30.
- 4:54:34Okay. So, what are variables? These are
- 4:54:38um
- 4:54:40basically our way of storing data to
- 4:54:43make it easier to reference them and
- 4:54:45manipulate uh throughout our program.
- 4:54:48So, we've actually already used a
- 4:54:49variable. We we used one in our uh demo
- 4:54:52we just did where we called the input
- 4:54:54the name. Uh we stored that input into a
- 4:54:58variable called name. And so um
- 4:55:01variables just really are a reference to
- 4:55:05some data. That's all they are. They
- 4:55:07allow us to reference that data
- 4:55:09throughout the program. We can store
- 4:55:11information into a variable and then
- 4:55:13access it throughout our code. Um so on
- 4:55:16this screen are some examples of
- 4:55:17variables. Now, variables have names,
- 4:55:20which is why I said usually we want
- 4:55:23those to be meaningful. Like X is a
- 4:55:25valid name, but it's not that
- 4:55:28interesting of a name. It doesn't give
- 4:55:30us that information much information
- 4:55:32about what it's what it really means.
- 4:55:33So, probably not the best name. Um, but
- 4:55:37we have things like uh we can we can
- 4:55:40store some text inside of this variable
- 4:55:43called name. We can store a number
- 4:55:44inside of this um variable called price.
- 4:55:47we can store uh a true or a false value
- 4:55:50inside of this variable called
- 4:55:52is_active.
- 4:55:54Um and so variables will show up all
- 4:55:58over our code and uh they are basically
- 4:56:01our way to reference some values. Now
- 4:56:06these things over here are basically
- 4:56:09different types of data that we need to
- 4:56:11learn about, right? So we need to learn
- 4:56:13about what is a 10 versus what is in
- 4:56:15something inside of quotes. is it a
- 4:56:17string versus something that has
- 4:56:19decimals which is a floatingoint number
- 4:56:22versus something that is true or false
- 4:56:23which is a boolean value. We need to
- 4:56:25learn about those data types. But notice
- 4:56:28how all of these are being referenced by
- 4:56:31a um by a variable that has some name to
- 4:56:35it. Okay. So the variable is this guy.
- 4:56:38It is our reference to that data. Um and
- 4:56:42we will use variables throughout um so
- 4:56:45that we can have you know references to
- 4:56:47information in our code.
- 4:56:50So variables are fundamental um to to
- 4:56:54working with Python.
- 4:56:56Um now variables can store different
- 4:56:58kinds of data. So I just alluded to
- 4:57:00that. And so the different types of data
- 4:57:02available to us in Python kind of fall
- 4:57:04in these two different categories. one
- 4:57:06being single values or what are known as
- 4:57:09scalar values. So these are things like
- 4:57:12integers. So the number 10, the number
- 4:57:151, the number 2,323,
- 4:57:19those are all whole number integers. Um
- 4:57:22floats, which are anything with a
- 4:57:23decimal.
- 4:57:25So 32.3,
- 4:57:273.14,
- 4:57:29um 1.2, those are all floating point
- 4:57:32numbers. Um, booleans only have two
- 4:57:36options. They only have true or false.
- 4:57:38So, they represent kind of a binary uh
- 4:57:41value um which we say is true or false.
- 4:57:46And um then we also have um complex
- 4:57:50numbers which are which have imaginary
- 4:57:53uh parts to them. We won't really be
- 4:57:55dealing with complex numbers too much so
- 4:57:57I wouldn't worry about them. But in
- 4:57:59reality, Python supports working with
- 4:58:01the uh complex numbers and doing complex
- 4:58:03math. But uh so so complex numbers just
- 4:58:06have kind of a real part and an
- 4:58:08imaginary part to them. Um wouldn't
- 4:58:11worry too much about that. Again, we're
- 4:58:12not really going to work with those ever
- 4:58:14throughout throughout the program, but
- 4:58:15it does exist. Python supports it. So
- 4:58:18scalar data, single values, think
- 4:58:22numbers, think single numbers like
- 4:58:24floats, think integers, um single uh
- 4:58:27true or false values. So these kinds of
- 4:58:30data can be stored into variables.
- 4:58:33On the opposite end of the spectrum are
- 4:58:36aggregated types that we are storing
- 4:58:39multiple things.
- 4:58:41So we're going to learn about all of
- 4:58:43those, but um probably the most common
- 4:58:45and one that we've already dealt with is
- 4:58:47going to be a string. So a string is
- 4:58:49technically an aggregated type because
- 4:58:51it has multiple characters that form,
- 4:58:54you know, an overall uh string, which is
- 4:58:57a string is usually you you know it's a
- 4:59:00string because it's inside of quotes,
- 4:59:02right? It's inside of these double
- 4:59:03quotes or single quotes. Um, Python
- 4:59:07actually doesn't care about quotes
- 4:59:09really in terms of if it's single or
- 4:59:10double as long as you're consistent with
- 4:59:12it. Like if you if you start with double
- 4:59:15quotes, you should end with double
- 4:59:16quotes. If you start with single, you
- 4:59:18should end with single. Python doesn't
- 4:59:20really care either way. Um, so strings
- 4:59:25are going to represent um collections of
- 4:59:27characters. Um we are going to talk
- 4:59:30about sets which are basically like u an
- 4:59:33array of unique values. Um so we'll talk
- 4:59:37about sets we'll talk about lists which
- 4:59:39are a really important structure. It's
- 4:59:41basically an array that can hold many
- 4:59:44different types of data. Um so we'll
- 4:59:46talk about list. We'll talk about
- 4:59:48tupils. So you if you see that word
- 4:59:50tuple e that is um people some people
- 4:59:54pronounce it tuple. I I like to call it
- 4:59:56tupole, but um that is going to be very
- 4:59:59similar to an array. It's just going to
- 5:00:01have slight differences and uh if you
- 5:00:04can change it or not. Tupils you
- 5:00:06actually cannot change once you create
- 5:00:07it. Um versus list you can modify list.
- 5:00:11You can add things to it. You can remove
- 5:00:12things from it. Tupils you cannot. So
- 5:00:15we're going to learn about those
- 5:00:16differences as we go along and start
- 5:00:18working with these different types of
- 5:00:19data.
- 5:00:21Um but they are designed to hold
- 5:00:23multiple values, right? So you can see
- 5:00:25in that example that list has integers,
- 5:00:27it has strings, it can it can hold
- 5:00:29multiple types which is if you're coming
- 5:00:32from other languages is generally not
- 5:00:34the case. Um like arrays in Java, arrays
- 5:00:37in C, they can only hold one type of
- 5:00:40data in the array. They can't hold
- 5:00:42multiple.
- 5:00:43Um
- 5:00:45so uh then finally a dictionary. A
- 5:00:48dictionary is if you're coming from
- 5:00:49other languages, it's like a map, a
- 5:00:51hashmap. Basically it allows you to have
- 5:00:54uh keys mapped to values. So it's a
- 5:00:57really dictionaries are highly useful
- 5:00:59for storing information where we want to
- 5:01:02reference like this value maps to this
- 5:01:06value. So for instance in this
- 5:01:07dictionary the string a maps to one and
- 5:01:11then the string b maps to uh you know
- 5:01:14two and or whatever it maps to. And this
- 5:01:19will allow us to look up values in the
- 5:01:22dictionary. So we could look up, hey,
- 5:01:23what is the value stored at key A or
- 5:01:25what is the value stored at key B? Those
- 5:01:28kind of things. Dictionaries will be
- 5:01:30incredibly useful. We're going to
- 5:01:32explore all of those more as we go along
- 5:01:34in the lesson, but um for right now, it
- 5:01:38should be making sense that there are
- 5:01:40some data types that store multiple
- 5:01:42values like array or sorry, lists, um
- 5:01:45dictionary, strings, and then there are
- 5:01:48some data types that only have a single
- 5:01:49value like a single number like a float,
- 5:01:52integer, um boolean.
- 5:01:55Okay, so more to come on aggregated
- 5:01:57data. We're going to work with those,
- 5:01:58learn about the differences, learn about
- 5:02:00what it looks like in the code to work
- 5:02:02with the set, a dictionary, tupil, list,
- 5:02:05but those generally hold multiple values
- 5:02:08or can hold multiple values whereas um
- 5:02:12scalar data is only going to hold one.
- 5:02:18Okay.
- 5:02:22All right. So uh so as we said earlier
- 5:02:27um you know integers, floats, booleans,
- 5:02:30they only hold a single value. By the
- 5:02:33way, inside of Python, if you ever want
- 5:02:35to see what the type of a variable is.
- 5:02:38So let's say we know we have a variable
- 5:02:39called name. We can always check the
- 5:02:42what data type it is by by using the
- 5:02:45built-in type function. So we can use
- 5:02:48type and then pass in that variable
- 5:02:51and this will display what data type it
- 5:02:54is. So um if we stored the value 42 in
- 5:02:59some variable called int, if we um
- 5:03:02displayed if we did type um if we did
- 5:03:06type of this it would uh produce int
- 5:03:09which would say okay this value is an
- 5:03:12integer versus 3.14 that's going to be a
- 5:03:15float versus capital t true that's going
- 5:03:19to be uh the boolean type bool.
- 5:03:27Okay. So, uh we have integers, we have
- 5:03:30floats, we have booleans, all of which
- 5:03:33we will use throughout and we'll see
- 5:03:34where we will use those one versus the
- 5:03:37other. We'll learn about that.
- 5:03:41Um as I said, complex. So, uh just
- 5:03:44showing you here that those exist
- 5:03:46obviously. Um like I said complex has uh
- 5:03:49a real part and an imaginary part which
- 5:03:52you can access separately. So if you
- 5:03:54store a value as a complex you can uh
- 5:03:57access its real and imaginary parts
- 5:03:59separately which you may need to do for
- 5:04:01some type of uh calculations.
- 5:04:04Um again we won't really work with
- 5:04:07complex numbers in in this program. So
- 5:04:09not a big deal for us but it is
- 5:04:11supported
- 5:04:12and you know a lot of um mathematical
- 5:04:15packages in Python will use complex
- 5:04:18numbers uh if they need to but we won't
- 5:04:21really do it in this program. There's
- 5:04:24not really a need to for us.
- 5:04:28All right. So aggregated data um we have
- 5:04:32those strings which we've already seen.
- 5:04:34Those are the things inside of quotes.
- 5:04:35We have sets which are going to be
- 5:04:37collections of data um that are unique
- 5:04:42basically only allowing one uh copy of
- 5:04:45those elements inside the set. We're
- 5:04:46going to learn about that. Um list which
- 5:04:50is going to be a collection of items
- 5:04:53which we can change, we can add things
- 5:04:54to it, we can remove. Um lists are
- 5:04:56really awesome uh structure in Python.
- 5:05:00Um, what I want you to see right now
- 5:05:02though is you can start to see the
- 5:05:05syntax differences, right? So, like a
- 5:05:07set, um, a a set is where we have, uh,
- 5:05:13this brace. Notice that a set is created
- 5:05:16with a curly brace versus a list which
- 5:05:18is created with a bracket. So, right
- 5:05:20away, like when you see a brace, you
- 5:05:23should be thinking either a set or
- 5:05:25dictionary. Those are the two things
- 5:05:27that are created with a curly brace. Um,
- 5:05:30and you know it's a dictionary because a
- 5:05:31dictionary will have the colon which
- 5:05:33will map I'll show you that on the next
- 5:05:35screen. But that will map things from
- 5:05:37key to value. Um, depending on if you
- 5:05:40know left and right of the colon. Um,
- 5:05:43but do you guys see that like the syntax
- 5:05:45difference of a list? A list has a
- 5:05:48bracket set has a curly brace. Um,
- 5:05:52that's just one small difference. you
- 5:05:53know, we're going to learn like what is
- 5:05:54the actual difference between a set and
- 5:05:56a list, but that's just one I'm pointing
- 5:05:58out right now.
- 5:06:02Um, what does mutable mean? So, mutable
- 5:06:05uh means that we can change it. It's
- 5:06:08it's able to be changed. So, immutable
- 5:06:12would be we cannot change it.
- 5:06:17Yeah. And and one thing about a list
- 5:06:19that's really nice is every list has a
- 5:06:22natural ordering to it which is actually
- 5:06:24really beneficial. So a list has a
- 5:06:27notion of the first item, the second
- 5:06:30item, the third item, the fourth. That's
- 5:06:33really important for accessing data
- 5:06:35within the list. Okay. So lists are
- 5:06:39really powerful. Um
- 5:06:43yeah. So a a set the reason it shows
- 5:06:46it's in a different order is because a
- 5:06:48set does not maintain order. A set never
- 5:06:51maintains order because um it a set is
- 5:06:55not you do not access items by by order.
- 5:07:00So that's just something unique to a set
- 5:07:02is that it doesn't have a natural order.
- 5:07:03So every time you print it out, it will
- 5:07:05display in a different order.
- 5:07:06Potentially it's random. It's random
- 5:07:08order when you when you display it. A
- 5:07:11set is just meant to be a general
- 5:07:13collection. Think of it like a bucket.
- 5:07:15Like here's this bucket of items that I
- 5:07:17have.
- 5:07:18It's just a collection of items. A list
- 5:07:21actually maintains an order, a
- 5:07:23consistent order of items.
- 5:07:27This is different.
- 5:07:30So we'll talk more about that when we
- 5:07:31get into those.
- 5:07:45Okay.
- 5:07:48All right. So I wanted to show you also
- 5:07:49the tupole in the dictionary. So a
- 5:07:51tupole
- 5:07:53is also a collection of items. Now the
- 5:07:55tupil is ordered. So it's like a list.
- 5:07:58It's ordered but it is immutable.
- 5:08:02Meaning you cannot change a tupole. So
- 5:08:04once you create a tupil you cannot
- 5:08:07change it or else you'll get an error.
- 5:08:09Python will tell you hey this is
- 5:08:10immutable I can't change this. So if you
- 5:08:13try changing being if you try to add
- 5:08:15something to the tupil if you try to
- 5:08:17modify one of the entries in the tupil
- 5:08:19like if I try to if I go in and try to
- 5:08:21change this a um to a d
- 5:08:26um this would not be allowed. This would
- 5:08:28this would throw an error. The
- 5:08:30interpreter would say hey you're trying
- 5:08:31to change something that cannot be
- 5:08:32changed. So tupils are immutable but
- 5:08:36they have a benefit beyond a set of
- 5:08:38actually being ordered. So there there's
- 5:08:41a natural ordering to a tupil where this
- 5:08:43is the first item, this is the second
- 5:08:45item, this is the third and every time
- 5:08:46you display a tupil will be in a
- 5:08:48consistent order. But tupils are not
- 5:08:52like a list. You can't change it. So
- 5:08:55tupils are useful for situations where
- 5:08:58you want ordering, but you don't want
- 5:09:00anybody to change any of that data
- 5:09:02that's in the tupil. It's it's not
- 5:09:04changeable.
- 5:09:06Mutable meaning it just means changeable
- 5:09:08like you can modify it. If if something
- 5:09:11is mutable, you can modify it.
- 5:09:15Immutable like a tupole is immutable. We
- 5:09:17cannot modify it once we create it.
- 5:09:19That's what it is.
- 5:09:23Okay.
- 5:09:26Yeah. Okay. So then finally a
- 5:09:29dictionary. Now you by the way um look
- 5:09:32at the tupole. See how it's created with
- 5:09:34a parenthesis.
- 5:09:37So that's different than the curly
- 5:09:38brace. That's different than the
- 5:09:39bracket. Right? So a tupole you know
- 5:09:41it's a tupil because of the parenthesis
- 5:09:44and the items are separated by a comma
- 5:09:46just how just how they are in a set and
- 5:09:47just how they are in a list. Um
- 5:09:51so so the the parenthesis gives it away
- 5:09:53that it's a tupole. Um now look at the
- 5:09:56dictionary and the dictionary is um a
- 5:10:00collection of key value pairs. So this
- 5:10:03is a key value pair. This is a key value
- 5:10:05pair. Um this is a key value pair and on
- 5:10:08and on. We can have as many as we want.
- 5:10:10And one thing I want you to notice about
- 5:10:12this is there is no restriction on the
- 5:10:16data types of the keys and the values.
- 5:10:18So keys can be integers, keys can be
- 5:10:22strings, values can be integers, values
- 5:10:25can be strings, values could be floats,
- 5:10:28values could even be other dictionaries
- 5:10:31or lists. So value like we could have
- 5:10:35what's called a nested dictionary where
- 5:10:37we actually have something mapping over
- 5:10:39to another dictionary.
- 5:10:42That's totally possible in Python. So we
- 5:10:45can have dictionaries that part of the
- 5:10:47values inside of the dictionary actually
- 5:10:49have our dictionaries themselves and
- 5:10:52that would represent kind of a nested
- 5:10:54structure there. So for instance this
- 5:10:57name could map to a dictionary with
- 5:10:59everybody's name in it. Um or it could
- 5:11:03map to a list um you know
- 5:11:07could map to a list it can map to
- 5:11:09whatever it could map to a tupole. Uh so
- 5:11:11you there's really no restriction in
- 5:11:12what the keys and values uh are going to
- 5:11:15be.
- 5:11:18Uh Brent is it more efficient than the
- 5:11:20other uh is what more efficient than the
- 5:11:23other methods? Just want to clarify your
- 5:11:25question so I so I answer it properly.
- 5:11:35The tupole versus using an array.
- 5:11:38Yeah. Yeah. Yeah. So uh these are all
- 5:11:41good questions. So um the tupole
- 5:11:46is guaranteed not to be changed. So it
- 5:11:49is a little faster when we are looking
- 5:11:52up items like when we are referencing
- 5:11:53items. It's a little bit faster because
- 5:11:56uh we know that it's not going to be
- 5:11:58modified ever. So everything is going to
- 5:12:00be consistently in the same spot. So
- 5:12:03like whatever's first is going to stay
- 5:12:05first, whatever's second is going to
- 5:12:06stay second and on and on. So tupil is
- 5:12:09is nice in that sense. A list can be
- 5:12:12changed. So whatever is first may not
- 5:12:15guarantee to be first in the future. We
- 5:12:17can modify it. We can remove things. We
- 5:12:19can add things to the list. So we can
- 5:12:22expand. The list is very like dynamic.
- 5:12:24The list. So the list is less efficient
- 5:12:28because it's way more dynamic. Does that
- 5:12:30make sense? Like it can change. You can
- 5:12:31add you can keep expanding the list by
- 5:12:34adding things to it. You can shrink the
- 5:12:35list by removing things from it.
- 5:12:38So list is way more dynamic which for a
- 5:12:41lot of scenarios is useful,
- 5:12:44right? We want to be able to add and
- 5:12:45remove and modify things.
- 5:12:48Um but a tupole is more rigid in the
- 5:12:51sense that once you create it, you
- 5:12:52cannot change anything about it.
- 5:13:00Yeah. Yeah. So a dictionary is good for
- 5:13:03Yeah. Like a phone book would be a good
- 5:13:05example of a dictionary because you with
- 5:13:08a dictionary you're usually looking up
- 5:13:10things. So you so like a diction in a
- 5:13:12phone book you have a name that maps to
- 5:13:15a phone number.
- 5:13:17Um so yes you have a you have that a
- 5:13:21dictionary will map a key to a value
- 5:13:23just like a name would be mapped to a
- 5:13:25phone number. So yeah a phone book makes
- 5:13:28a lot of sense.
- 5:13:31Um,
- 5:13:34a list, a list is like any is like a
- 5:13:36normal like like your grocery list. Like
- 5:13:38you may add things to it, you may remove
- 5:13:39things from it, you may change things on
- 5:13:41it. It's very dynamic. Um, a tupole is
- 5:13:45kind of like a fixed um set of data
- 5:13:49that's ordered in some way. So maybe
- 5:13:52like um what you would see on on a on a
- 5:13:55letter like you have your name, you have
- 5:13:57your address, you have um your zip code,
- 5:14:00like you kind of have those and it it
- 5:14:02should stay that way in order to mail
- 5:14:04the letter kind of thing.
- 5:14:07Uh can you convert a tupil to a list?
- 5:14:09Yes, you can do vice versa. You can
- 5:14:11convert a tupil to a list and you can
- 5:14:13convert um you can convert a list to a
- 5:14:15tupole. Yes, you can convert between
- 5:14:18them.
- 5:14:20I'll show us examples of that later.
- 5:14:29Okay. So, just to recap there,
- 5:14:33tupil is not changeable, but it has an
- 5:14:38order. So, it has a natural ordering to
- 5:14:40it. Whatever is first is first. Whatever
- 5:14:43second is second, third and third. So,
- 5:14:44you can access things based on their
- 5:14:46position within the tupole. That's
- 5:14:48really nice. But you cannot modify
- 5:14:50anything about a tupil once you create
- 5:14:51it.
- 5:14:52Okay. A list has an ordering to it. You
- 5:14:57can access things based on their
- 5:14:58position. But a list is dynamic. It is
- 5:15:01mutable. Meaning you can change it. You
- 5:15:03can change values. You can add things to
- 5:15:06it. You can remove things from it. Okay?
- 5:15:08So very dynamic. That's what a list is.
- 5:15:10Um dictionary. It maps keys to values.
- 5:15:14No restrictions on what those keys and
- 5:15:16values can be.
- 5:15:18All right. And then a set. A set is
- 5:15:20think of it like a bucket. It just has
- 5:15:22things in it. A set has no order to it.
- 5:15:25So you cannot access things based on
- 5:15:26their order. And every time you uh
- 5:15:29display the set, you can get a different
- 5:15:31ordering. Um
- 5:15:34but a a set only is special in that it
- 5:15:38only allows unique items. So if you try
- 5:15:40to put multiple copies of a piece of
- 5:15:42data, it's only going to keep one of
- 5:15:43them. So a a set is like a bucket with
- 5:15:47only unique things in it.
- 5:15:50Okay. So and sometimes that's really
- 5:15:52useful is to know like what are the
- 5:15:54unique values? Uh a set would help us
- 5:15:57maintain that.
- 5:16:02Any questions about those? You know, we
- 5:16:04have to we have to work with this and
- 5:16:05see this in the code and we will. But
- 5:16:07just any questions right now about these
- 5:16:09different types of data that we're
- 5:16:10talking about.
- 5:16:21Okay.
- 5:16:23Very good.
- 5:16:26All right. Let's talk about assignment.
- 5:16:28So, what that means is um Oops. Let's
- 5:16:31talk about assignment which means that
- 5:16:34we will be um taking a variable name and
- 5:16:38assigning data to it. Now we've already
- 5:16:40seen this. We already saw it in our demo
- 5:16:42where we did input. We did name equals
- 5:16:45input.
- 5:16:46So the equals symbol is how we assign
- 5:16:51values to a variable.
- 5:16:54That makes sense, right? It's very like
- 5:16:56self-explanatory.
- 5:16:58But um what we should think about with a
- 5:17:02variable is really the fact that a
- 5:17:04variable is is a reference to that data.
- 5:17:09Okay. So when we say x= 34, we are
- 5:17:13assigning 34 to the name x. So x becomes
- 5:17:18a variable which is referencing the data
- 5:17:21which is an integer 34. Right?
- 5:17:25What's really interesting about that and
- 5:17:26this is how you can kind of test your
- 5:17:28intuition of the fact that this is a
- 5:17:31reference is if we come along and have
- 5:17:33another name Y and we set that equal to
- 5:17:36X.
- 5:17:38This is just saying that we are creating
- 5:17:41another reference that is equal to the
- 5:17:43reference we already have. Now, why
- 5:17:46would we ever do that? Probably we
- 5:17:48wouldn't. That's kind of redundant. But
- 5:17:50this just proves that they're ultimately
- 5:17:52references because when we display x, we
- 5:17:56get 34. Of course, that's what we stored
- 5:17:59the value 34
- 5:18:01uh referenced by x.
- 5:18:04And then when we print y, we get the
- 5:18:06same number, right? We get 34. And why
- 5:18:09does that happen? Because we we
- 5:18:11literally declared y equal to x. Meaning
- 5:18:15y should reference the same data that x
- 5:18:17does. Okay, so as variables they are
- 5:18:21equal meaning that um X is being
- 5:18:25assigned to Y meaning Y should reference
- 5:18:28the same data that X does. So they they
- 5:18:31uh contain the same data. Now what's
- 5:18:34interesting is if you print out the ID.
- 5:18:37So the ID is the internal
- 5:18:40um the internal memory address
- 5:18:46of of the reference.
- 5:18:49Um now it usually we don't care about
- 5:18:51that but this is just to prove the point
- 5:18:53is that you can see these are the same
- 5:18:55address. These are the same. That's by
- 5:18:57design because we're saying okay I have
- 5:19:00this reference X which is referencing
- 5:19:01this data 34. it's stored at this
- 5:19:03address. Um, and then when I come along
- 5:19:06and say, okay, y equals x, that's just
- 5:19:10the same reference. You see how it's the
- 5:19:12same exact address,
- 5:19:14same reference.
- 5:19:16So, just proving that variables are
- 5:19:19literally just references to data. They
- 5:19:21allow us to reference that data, which
- 5:19:23is really, really, you know, nice. So,
- 5:19:25we can reuse x throughout the code. Um,
- 5:19:28we can reuse name. we can you know
- 5:19:30whatever we create we can reuse.
- 5:19:35Um if you look over to the right we have
- 5:19:38an alternative example which um now
- 5:19:42resets y to store a new value. So
- 5:19:45instead of saying y equals to x we
- 5:19:47actually overwrite y and reassign it to
- 5:19:49the integer 78. That's a new piece of
- 5:19:52data right 78. So now if you look at
- 5:19:56their their uh references, they're
- 5:19:59different. These are different. And that
- 5:20:02makes sense because now they're pointing
- 5:20:03to two different uh pieces of data,
- 5:20:06right? X is pointing to 34. Y is
- 5:20:09referencing to 78. So of course they're
- 5:20:12going to be different uh different
- 5:20:13addresses. And this is a bit of a typo.
- 5:20:16This should say ID of Y
- 5:20:19because we're ref we're talking about Y.
- 5:20:21It's a bit of a typo there.
- 5:20:24Okay, so hopefully this now this example
- 5:20:27is just to reinforce the fact that when
- 5:20:29we use the equal sign, we're setting
- 5:20:31equal we're setting a variable name
- 5:20:33equal to a piece of data, right? And
- 5:20:36that is creating a reference to that
- 5:20:38piece of data.
- 5:20:40That's all we're that's all we're saying
- 5:20:41with this. So we are assigning a piece
- 5:20:45of data to that reference X or Y or
- 5:20:47whatever it is.
- 5:20:53Okay.
- 5:20:56All right. Let me ask you guys. Um, what
- 5:20:59is the default data type of a variable
- 5:21:02assigned using the input function? This
- 5:21:04is an interesting question. We didn't
- 5:21:05actually cover this, so I'm really
- 5:21:06curious to see what you guys think about
- 5:21:08this.
- 5:21:24A lot of votes for for string.
- 5:21:28Let's get a few more.
- 5:21:38Perfect. Yeah. So, water votes receipt
- 5:21:40it is a string. So, that that begs the
- 5:21:44question like what happens if we input a
- 5:21:47number? Like what happens if we put in a
- 5:21:49two? What happens to that? You know that
- 5:21:53two will actually be read in as the
- 5:21:56string two. So it would be So if we use
- 5:21:59the input and we it pulls up that text
- 5:22:01box and we put in a number like two
- 5:22:06um and we set that equal to the variable
- 5:22:09x, whatever we name that name x,
- 5:22:11whatever. What that really means is x is
- 5:22:14going to be um equal to the the um x is
- 5:22:19going to be equal to the
- 5:22:22uh string 2. So that's something to be
- 5:22:25cautious about with the input is it
- 5:22:28always assumes the input data is going
- 5:22:30to be a string. So luckily there's a way
- 5:22:33to convert between strings and numbers.
- 5:22:37So if we wanted to turn this into the
- 5:22:39actual number, what we would do is use
- 5:22:41the the data type function int, which
- 5:22:44would convert uh this would convert it
- 5:22:48over to the numerical two. Would
- 5:22:50actually convert it from a string to an
- 5:22:52integer. We just use int. Or we could
- 5:22:54use like if we if somebody put in a
- 5:22:56decimal like 2.5
- 5:22:59then um we could do a float
- 5:23:04of 2.5
- 5:23:06and that would convert that over to uh
- 5:23:09the the number.
- 5:23:12Okay, let me actually show you guys
- 5:23:15this. Let me go over to Collab real
- 5:23:16quick and show you guys this. I know
- 5:23:18it's not in a demo, but I think it'll be
- 5:23:19better if I just show you what I mean by
- 5:23:21this because this is an important point
- 5:23:23with input.
- 5:23:26So, let me uh stop sharing there. Let me
- 5:23:30go over to Collab for a second so I can
- 5:23:32show you literally what this means.
- 5:23:38So, go back into the notebook here. So
- 5:23:41what I want to show you is that um when
- 5:23:45we do input
- 5:23:47the default type
- 5:23:51is string.
- 5:23:53So for instance when I do um
- 5:23:59when I do uh uh value equals to input
- 5:24:05and let's try um enter your age.
- 5:24:12Oops. Enter your age.
- 5:24:15And then we uh run this.
- 5:24:21So we enter the age. Now this is going
- 5:24:24to be read in as a string. So even
- 5:24:28though I'm putting a number there, it's
- 5:24:30actually going to be read in as a
- 5:24:31string. So now
- 5:24:34look at what the type of value is.
- 5:24:39It's a string. Do we see that? So
- 5:24:42string. So this number even though we
- 5:24:45put in a number it gets it the the input
- 5:24:49function always converts it to a string
- 5:24:52no matter what we put there. If we put a
- 5:24:53decimal if we put a a large number it's
- 5:24:56always going to assume it's a it's a
- 5:24:58string. So luckily
- 5:25:01um we can convert to an integer
- 5:25:05by using int the int function.
- 5:25:09So um we can print sorry we can say
- 5:25:12value
- 5:25:14uh or we can do int value which which
- 5:25:18will convert that 32 string because
- 5:25:23right now if I were to um just display
- 5:25:27value it's a string 32. You can see it
- 5:25:30inside of the quotes. But now when I do
- 5:25:33this uh and I can run that now it's an
- 5:25:37integer. Do we see that now it's
- 5:25:40actually a number
- 5:25:42which is great. It no longer has those
- 5:25:43quotes. It's actually going to be
- 5:25:45treated as an actual integer which which
- 5:25:47may be useful for calculations or
- 5:25:49storing it or whatever whatever we need
- 5:25:51to do with it. So that's just one piece
- 5:25:53of caution with the input is if you're
- 5:25:56working with numerical data it's going
- 5:25:58to treat it as a string. We have to
- 5:26:00convert it.
- 5:26:02Okay
- 5:26:05questions on that. Does that make sense
- 5:26:07to us? like the input's always going to
- 5:26:09accept the input as a string. So if we
- 5:26:12want to work with it alternatively
- 5:26:15um we should convert it.
- 5:26:20Uh you can yeah so like you could
- 5:26:23convert um if I did this if I wrapped
- 5:26:27this around in the int function that
- 5:26:30would automatically
- 5:26:32take whatever we put whatever this
- 5:26:34returns would automatically be um cast
- 5:26:37over to an int. So we could do that. So
- 5:26:41let me show you that. So when I run
- 5:26:43this, I can put in 32
- 5:26:48and it it's like automatically going to
- 5:26:51be casted to an integer. So there now
- 5:26:53it's an integer. Does that make sense?
- 5:26:55Like when I wrap this int around the
- 5:26:57input, it's going to automatically
- 5:26:59convert
- 5:27:17Uh, what did you put in the input box?
- 5:27:19So, yes, you'll get an error if you
- 5:27:22don't put in a valid integer.
- 5:27:24So, let's put in like if I put in my
- 5:27:28name,
- 5:27:30this is going to be this should be an
- 5:27:31error because I don't know how to
- 5:27:34convert this string over to a number. It
- 5:27:36doesn't make sense to do that, right?
- 5:27:38So, this should be an error.
- 5:27:42Right? That will be an error because
- 5:27:44it's a string.
- 5:27:53So why did you get an error? Uh input
- 5:27:55enter your age value int value.
- 5:28:04Uh did the did the text box show up?
- 5:28:06Maybe try separating it into a different
- 5:28:08cell.
- 5:28:10Try try putting the other two lines in a
- 5:28:12different cell. Um, you need the text
- 5:28:14box to show up and then you need to
- 5:28:16enter something.
- 5:28:18Yeah.
- 5:28:22Okay.
- 5:28:24All right. Does this all make sense? Any
- 5:28:26questions about this? About the input
- 5:28:28function.
- 5:28:33Okay.
- 5:28:36Good. Okay, let me go over back to the
- 5:28:38notes then.
- 5:28:46Okay.
- 5:28:53All right. So, we have another demo.
- 5:28:55We'll do that now. Uh I was just kind of
- 5:28:57doing one, but let's go back over to
- 5:28:59this will be demo five. Let's do that.
- 5:29:01So, we're going to practice assigning
- 5:29:03different values um to variables and
- 5:29:06displaying them just so you get in the
- 5:29:08habit of being able to create your own
- 5:29:10variables and just go through that kind
- 5:29:12of one more time. We'll do this one
- 5:29:14relatively quickly um and then uh move
- 5:29:17on.
- 5:29:20So, this will be uh demo five.
- 5:29:29So, let me pull that one up for you
- 5:29:31guys.
- 5:29:51Okay, let me share my screen.
- 5:29:57All right. So, this is going to be demo
- 5:29:59five. Um,
- 5:30:03now again, like feel free to use
- 5:30:05whatever platform you've been using.
- 5:30:07Collab, Jupyter Notebook. I know this
- 5:30:09instruction says set up a Jupyter
- 5:30:11notebook. Feel free to use whatever you
- 5:30:12want. You can use Collab. Um, whatever's
- 5:30:15been working for you to build your build
- 5:30:17your notebooks. So obviously this this
- 5:30:20looks a little different than collab but
- 5:30:21it's because it's the Jupiter. Um so we
- 5:30:25create a notebook.
- 5:30:27Now what I want you to see
- 5:30:31is this takes the approach of everything
- 5:30:33we just did. Let me zoom in on this. Uh,
- 5:30:37I know that's a little small,
- 5:30:41but this is doing everything we just
- 5:30:43said we could do where we
- 5:30:46um essentially take
- 5:30:49So, I just want to zoom in on this. Um,
- 5:30:52notice that we
- 5:30:55uh take the um input and this will be
- 5:31:00saved as a string.
- 5:31:01Um,
- 5:31:03so this will be saved as a string and
- 5:31:06this will be saved into this name. And
- 5:31:08for instance, this will be saved as a
- 5:31:11string, but we convert it over to an
- 5:31:12int, which is exactly the kind of
- 5:31:14example I just did, right? Where we take
- 5:31:17take an input, we convert it over to to
- 5:31:21uh int.
- 5:31:24Does somebody have Yeah. Does somebody
- 5:31:26have the demos available? like if if
- 5:31:29somebody doesn't mind sharing those in
- 5:31:30the chat. I again I don't have the PDFs.
- 5:31:33They should be from your LMS. They
- 5:31:35should be in the reference material.
- 5:31:37There should be a demos folder that you
- 5:31:39can download. If somebody has those and
- 5:31:41doesn't mind sharing them.
- 5:31:45They have that folder of them, like a
- 5:31:47zip folder of them, that'd be fantastic.
- 5:31:54Yeah. Thanks. Thanks. This is This is
- 5:31:58the demo we're going through currently.
- 5:32:01Perfect. So, for you guys having trouble
- 5:32:04navigating the demos, please download
- 5:32:07this zip folder.
- 5:32:10Download the zip folder that that these
- 5:32:13guys are uploading. Thank you so much.
- 5:32:14Download the zip folder so you have all
- 5:32:16of them.
- 5:32:18Please take a moment to do that.
- 5:32:25Okay.
- 5:32:26Uh,
- 5:32:36copy the code and got an error at height
- 5:32:38value. Use foot, not meter. I mean, it
- 5:32:40shouldn't matter. It, you know, you
- 5:32:42should just be the point of that one is
- 5:32:44to put in a decimal.
- 5:32:50How to create a new file. Um, what
- 5:32:52platform are you on? Collab.
- 5:32:56I don't know what platform you're on.
- 5:32:58Collab. Uh, just go to file, new
- 5:33:01notebook.
- 5:33:03New notebook in drive, I think is what
- 5:33:05it's called.
- 5:33:13Do you see that? It should be like it
- 5:33:15should be at the top. There should be a
- 5:33:16file and then new notebook.
- 5:33:22Let me go over to it.
- 5:33:26Uh,
- 5:33:30this one. You don't see this
- 5:33:34file. It's at the top. The top of the
- 5:33:36notebook. Do file and then new notebook.
- 5:33:41You don't see new notebook.
- 5:33:56Uh if you if you don't see that, just go
- 5:33:58to a new tab. Just go to a new tab and
- 5:34:01go to um Google Collab.
- 5:34:05You can always do that. Just go to just
- 5:34:07start a new um just go to Google Collab
- 5:34:10and then it will let you like launch a
- 5:34:12new notebook. So just just do that. Just
- 5:34:13do a new tab if it doesn't work.
- 5:34:22Okay.
- 5:34:24So, by the way, one of those examples
- 5:34:26was entering a float. So, it looked kind
- 5:34:28of like this. So we had um our our
- 5:34:31height is equal to float and then we had
- 5:34:36uh input and then we had um enter your
- 5:34:41height and then this was um uh some sort
- 5:34:47of uh this should be some sort of
- 5:34:49decimal value. So let's say it is um I
- 5:34:54don't know uh 5.7
- 5:34:57whatever that is uh feet it doesn't it's
- 5:35:00just some decimal um and then we hit uh
- 5:35:04we hit enter that will store the height
- 5:35:07as a float so that when we um display
- 5:35:10the height uh it will be rendered as a
- 5:35:13float appropriately right that's what
- 5:35:15that that's what should happen
- 5:35:20that's the point of that It just needs
- 5:35:21to be some decimal. It should work.
- 5:35:31All right, let me go back to the demo
- 5:35:34document.
- 5:35:38All right, were you guys able to run
- 5:35:40some of these? Like, were you able to
- 5:35:42run some of the inputs and change them?
- 5:35:44So, try these out on your own real
- 5:35:46quick. like try doing int and then input
- 5:35:48for enter your age. It should convert
- 5:35:52that. You should be putting in a number
- 5:35:54or else you'll get an error and it
- 5:35:56should convert that over.
- 5:36:04I by the way I wouldn't worry about this
- 5:36:06last one uh because we haven't learned
- 5:36:08about the comparison operator yet which
- 5:36:11is this equals equals. So we'll learn
- 5:36:13about that in a in a little bit in a few
- 5:36:15minutes. So, don't worry about that one
- 5:36:17too much right now. But at least these
- 5:36:18first few should make some sense and we
- 5:36:21should be able to do.
- 5:36:25Were you guys able to run one of those
- 5:36:27and convert over the the float or int
- 5:36:32and do the input and convert it?
- 5:36:36Did that work for you?
- 5:36:41Give it a try.
- 5:36:54Let me clear that. Any questions about
- 5:36:57that?
- 5:37:05Should look something like this.
- 5:37:10Good. We're good on that on converting
- 5:37:11over the input. Okay, perfect. Sounds
- 5:37:14like Sounds like we're able to run that
- 5:37:15and uh it was okay.
- 5:37:23What are you entering for the feet?
- 5:37:28Like, are you literally entering like
- 5:37:30quotes?
- 5:37:32Yeah, that's not going to work when you
- 5:37:34do that because it's going to um there's
- 5:37:37a string f, there's a character there.
- 5:37:39it's not going to be able to convert
- 5:37:40over to.
- 5:37:42So if you did if you did 6.4 that would
- 5:37:46work.
- 5:37:48Any decimal should work. But like the f
- 5:37:50is a character. So the the float doesn't
- 5:37:54know how to convert over a character,
- 5:37:57right? Yeah. So so that's not going to
- 5:38:00work. You need to put in a decimal to to
- 5:38:03be able to convert over to the number.
- 5:38:07Okay.
- 5:38:09Very good. Very good. Let's go back over
- 5:38:12to our notes so we can continue along.
- 5:38:291.7. Yeah. If you have any if you have
- 5:38:33any character, it's not going to work.
- 5:38:36It's not going to work. You need to put
- 5:38:37in you need to put in a decimal.
- 5:38:44All right, let's talk about operators.
- 5:38:47So, these are going to be really
- 5:38:48important. Um,
- 5:38:54let's talk about operators so that we
- 5:38:57can uh
- 5:39:00uh be able to compare things and work
- 5:39:03with things. Um so let's let's talk
- 5:39:06about Python operators.
- 5:39:08So what are operators? What do we mean
- 5:39:10by that? In Python, operators are
- 5:39:13special symbols or keywords that perform
- 5:39:16operations. So as the name suggests,
- 5:39:18it's performing some level of operation.
- 5:39:21Um which means that the interpreter
- 5:39:24should do some sort of logical
- 5:39:25operation, mathematical operation,
- 5:39:27relational operation to produce a
- 5:39:30result. Um, so usually that means
- 5:39:33there's going to be multiple variables
- 5:39:35that are going to be used to do some
- 5:39:36operation between. So an example of an
- 5:39:39operation would be like adding,
- 5:39:41subtracting, multiplying. That's an
- 5:39:42operation. But we can have logical
- 5:39:45operations like taking the um logical
- 5:39:49and or logical or of things. We'll see
- 5:39:51what that means. But um in Python,
- 5:39:54there's many situations where we want to
- 5:39:56we want to be able to do operations
- 5:39:59between variables. Whether that's simple
- 5:40:01mathematical or maybe some type of
- 5:40:03relational like testing if a value is in
- 5:40:06a list. That's an important operation.
- 5:40:09Is 10 in my list? Is five in my list? Um
- 5:40:13those are important operations. So we
- 5:40:15want to learn about these operators and
- 5:40:17they're going to be really important for
- 5:40:18us going forward is because these will
- 5:40:21be very standard. um things we will use
- 5:40:24as we uh go along. So we're going to
- 5:40:27spend some time talking about operators.
- 5:40:30Um so it turns out in Python um you can
- 5:40:33kind of group operators into many
- 5:40:35different categories. Um there's going
- 5:40:38to be standard arithmetic operators.
- 5:40:40Those are your everyday things like
- 5:40:41plus, minus, um division,
- 5:40:44multiplication. Um, assignment
- 5:40:46operators, which we've already seen, is
- 5:40:48things like equals, where we're setting
- 5:40:50a reference equal to something. We've
- 5:40:52already seen that. That's an assignment.
- 5:40:54Comparison, which is things like greater
- 5:40:56than or less than. Those are important
- 5:40:58for comparing values, comparing
- 5:41:00variables. Um, logical operators are
- 5:41:03going to be something like and and or,
- 5:41:06which will um do a logical operation
- 5:41:08between two two boolean values. That'll
- 5:41:11be important. And then we have a
- 5:41:13collection of miscellaneous operators.
- 5:41:15Um those will be things like is
- 5:41:17something in a collection like is five
- 5:41:20in a list? That's an operator. So we'll
- 5:41:23talk we're going to talk about all of
- 5:41:24these but just pointing out that there's
- 5:41:26many different categories of operators
- 5:41:28in Python.
- 5:41:35Okay, let's first talk about the
- 5:41:36arithmetic operators. So these are going
- 5:41:39to be your standard everyday um uh
- 5:41:42operations between numbers. So if we
- 5:41:45have numerical values like integers or
- 5:41:47floats, we can do math between them.
- 5:41:49That makes sense. Like that should be a
- 5:41:50capability of Python and it certainly
- 5:41:52is. We can add things, we can subtract
- 5:41:55things, we can multiply things, we can
- 5:41:57divide things. So um here are all those
- 5:42:01operators. We have plus minus the
- 5:42:03asterisk is a multiplication. So x
- 5:42:06asterisk y will multiply those together.
- 5:42:09So if we have two variables, one of them
- 5:42:11is 50, one of them is four, we do x
- 5:42:14asterisk y, that's going to multiply
- 5:42:16them together to get 200. Pretty pretty
- 5:42:18straightforward. Um
- 5:42:21division is one that we should be
- 5:42:22careful of. Of course, like we don't
- 5:42:24want to divide by zero. So if you I if
- 5:42:28the uh this secondary value that we end
- 5:42:32up dividing by is zero, that'll give us
- 5:42:34an error. Um the interpreter will say,
- 5:42:36"Hey, you're trying to divide by zero."
- 5:42:38We can't do that. It'll it'll produce an
- 5:42:40error. So that's the only thing we have
- 5:42:42to be on the lookout for with division.
- 5:42:43Just don't want to divide by zero.
- 5:42:46Um
- 5:42:48so all these are pretty standard. I
- 5:42:50think they all make sense.
- 5:42:52Hopefully they do to you. I think
- 5:42:54they're all pretty standard. you know,
- 5:42:55the kinds of things you'd see on a on a
- 5:42:57basic calculator. They all make sense.
- 5:42:59They should exist. Now, here's some more
- 5:43:02exotic ones. Um, I don't know if you
- 5:43:05guys have ever seen the the modulus
- 5:43:07operator, also known as modulo. This is
- 5:43:10one that returns the remainder of a
- 5:43:13division. Okay? So the the percentage
- 5:43:15sign is a mathematical operation between
- 5:43:18two numbers that returns not the
- 5:43:21quotient like not the actual division
- 5:43:23result but the remainder. So 50 divided
- 5:43:27by four
- 5:43:29um you know four goes into 50 um it goes
- 5:43:34in there uh uh 12 times evenly but it
- 5:43:38has two left over right. So there the
- 5:43:41remainder there is two. So the result of
- 5:43:44x mod we would read this as x mod y or
- 5:43:48modulo y um returns two. So if you're if
- 5:43:53you're unfamiliar with the modulo
- 5:43:54operation that seems a little bizarre
- 5:43:55that you take these two numbers
- 5:43:58um oops it seems a little bizarre that
- 5:44:01you take these two numbers and you like
- 5:44:02do this operation and you get a
- 5:44:06remainder result but it's actually a
- 5:44:07very powerful operation. Um the reason
- 5:44:11being is that sometimes we want to know
- 5:44:13what the remainder is more than we want
- 5:44:14to know what the quotient is. For
- 5:44:16instance, things that are very like
- 5:44:18cyclic in nature. Um so maybe we cycle
- 5:44:21through a collection and we want to know
- 5:44:24like how many times do we cycle through
- 5:44:26and then we have something left over
- 5:44:28which is the remainder. Um so the modulo
- 5:44:31operation is pretty useful. You could
- 5:44:33also check like if a number is even or
- 5:44:35odd using this. Like so if you modulo by
- 5:44:38two and it returns zero, that means it's
- 5:44:40even, right? Because that means there's
- 5:44:42there's nothing left over when I divide
- 5:44:44by two. So modulo is kind of a nice way
- 5:44:46to check if a number is even or odd. Um
- 5:44:50so modulo is a pretty nice uh operation.
- 5:44:53We'll use it from time to time. Uh but
- 5:44:55that is the percent operator. So x
- 5:44:59percent y will look for that remainder
- 5:45:01of the division. Um now there is also a
- 5:45:05double slash operator which is the
- 5:45:08integer division operator. This is kind
- 5:45:11of the reverse of modulo. It takes the
- 5:45:14largest integer quotient that that uh we
- 5:45:17can do from a division perspective. So
- 5:45:20remember I said 50 / 4. We can divide 4
- 5:45:23into 50 12 times evenly and we have two
- 5:45:27left over. So the integer division will
- 5:45:29just return to us an integer always
- 5:45:32which will be that quotient.
- 5:45:35So this is the quotient
- 5:45:38um and this is the uh remainder of 50 /
- 5:45:414. So the integer division returns to
- 5:45:45you that whole number like the largest
- 5:45:47number of times that that number goes
- 5:45:50into the other. So 12 times evenly
- 5:45:53obviously there's a remainder there but
- 5:45:55um but but yeah so integer division that
- 5:45:59one's useful if we want to know like how
- 5:46:01many times can I fit a value into
- 5:46:03another value a whole number of times
- 5:46:06and that happens from from time to time
- 5:46:07we may need to know that.
- 5:46:11Okay last operation here is exponent. So
- 5:46:15the exponent is the asterisk asterisk
- 5:46:18operator. Um so that raises a number to
- 5:46:23a power. Um so for instance like x star
- 5:46:27y or asteris y would mean that we are
- 5:46:30doing an operation like 5 to the 4th
- 5:46:33power um which is 625.
- 5:46:37Okay. So asterisk pretty useful. Like
- 5:46:39probably the most common asterisk would
- 5:46:41be squaring something which would be x
- 5:46:43um star star 2 which would would would
- 5:46:47be um x squared. So I mean that's a
- 5:46:50pretty common operation there is to
- 5:46:52raise something to the second power
- 5:46:54maybe the third raising something to the
- 5:46:56fourth probably less common but um the
- 5:47:00the asterisk asterisk operator is is how
- 5:47:03we do exponents in Python.
- 5:47:07Okay,
- 5:47:08so these are all basic arithmetic
- 5:47:11operations we can do between variables
- 5:47:13in Python. All right, any questions on
- 5:47:17those? Do those kind of make sense to us
- 5:47:19from a syntax perspective?
- 5:47:23Pretty straightforward, I think.
- 5:47:25Hopefully nothing too surprising there.
- 5:47:27Um,
- 5:47:29do you used to use module all the time
- 5:47:31for date date calculation? Yeah. Yeah.
- 5:47:34like when you uh find out how many like
- 5:47:36days how many weeks or where you are in
- 5:47:38the week, you cycle through like uh
- 5:47:41modulo 7 or something.
- 5:47:44That make sense?
- 5:47:52Okay,
- 5:47:54very good.
- 5:47:56Okay, I want to talk about assignment
- 5:47:58operators now. Now we've already seen
- 5:48:00this which is the basic equal sign that
- 5:48:04is a data assignment operator right so
- 5:48:07that means that we are setting a value
- 5:48:10equal to a reference so we are storing
- 5:48:12data inside of this reference variable a
- 5:48:15we use the basic equal sign as our
- 5:48:18assignment operator so that equal sign
- 5:48:20is called the assignment operator now
- 5:48:23what's really awesome is we can combine
- 5:48:26this basic assignment operator with our
- 5:48:28arithmetic ones to update values
- 5:48:33um and modify them uh as kind of a
- 5:48:37shortcut to say uh so so for example
- 5:48:40like a plus= 5 really represents the
- 5:48:43fact that I want to reassign a to the
- 5:48:47result of a + 5. So this means take
- 5:48:52whatever it is add five to it and
- 5:48:54reassign it to the value of a. So this
- 5:48:57is the same thing as if we just shortcut
- 5:49:00it in and Python will recognize if we do
- 5:49:02plus equals 5 it's the same thing. So
- 5:49:07and actually we can do that with any of
- 5:49:09these arithmetic operators. So if we
- 5:49:11want to take a variable multiply it by
- 5:49:14two and reassign it to that variable we
- 5:49:16can use star equals. So like a asterisk
- 5:49:21equals 2 is the same thing as if we were
- 5:49:25to reassign a to the value of a * 2.
- 5:49:30Does that make sense on the
- 5:49:33reassignment portion of that? So plus
- 5:49:35equals divide equal modulo equals star
- 5:49:39star equals would exponent something and
- 5:49:41reset it back to the variable.
- 5:49:45um minus equals we'll subtract and
- 5:49:48reassign that back to the variable. So
- 5:49:51you know x minus equ= 3 we'll subtract
- 5:49:54three from x and re and basically update
- 5:49:56it right reassign it back to x.
- 5:50:01So so that's pretty useful like whenever
- 5:50:03we need to do an operation and add like
- 5:50:07um you know a a very typical
- 5:50:11reassignment is to do like a plus equals
- 5:50:131
- 5:50:16That's a very typical reassignment
- 5:50:18because what this is the same as is a
- 5:50:21equals a + one. So that's like a single
- 5:50:25increment of a. We're just updating it
- 5:50:26by one.
- 5:50:29So plus equals 1. We may see that from
- 5:50:31time to time.
- 5:50:33A loop coming on. Yeah. Yeah. These are
- 5:50:35used in like while loops. Yeah. Like you
- 5:50:38do plus equals and you increment it
- 5:50:41until you reach a certain condition.
- 5:50:43Yeah.
- 5:50:44Now, if you're coming from other
- 5:50:45languages, if you have programming
- 5:50:47programming experience, you're coming
- 5:50:48from other languages, Python does not
- 5:50:50have an increment operator like plus+. I
- 5:50:54wish it did, but it doesn't. So, like I
- 5:50:56know in in Java and I think C they have
- 5:50:59um you can do like uh a plus+ or
- 5:51:03actually reverse you can do plus a but
- 5:51:06um that does not exist in Python
- 5:51:09unfortunately. You have to do the plus
- 5:51:11equals reassignment. So they don't have
- 5:51:13an increment operator. You'd have to
- 5:51:15you'd have to do just plus equals one to
- 5:51:18do the same effect as plus+.
- 5:51:21So I I know some people ask about that,
- 5:51:23but yeah, doesn't exist unfortunately.
- 5:51:29All right. Any questions about
- 5:51:30assignment? It's just really the equal
- 5:51:32sign and we can tack on the arithmetic
- 5:51:34to do some type of basic math and
- 5:51:37reassign to the variable.
- 5:51:39Hopefully the fact that we're using a
- 5:51:41single equals makes sense. Where people
- 5:51:43get confused all the time is the
- 5:51:46difference between a single equal sign
- 5:51:47and multi and two equal signs which
- 5:51:50we're going to see. Two equal signs
- 5:51:52means something completely different
- 5:51:54than a single equal sign. Single equal
- 5:51:57sign is an assignment. We are taking
- 5:52:00data and storing it in a reference
- 5:52:02variable,
- 5:52:04right?
- 5:52:06But multiple equal signs, we're going to
- 5:52:07learn about what that means. That's
- 5:52:08actually a comparison.
- 5:52:10It's something different.
- 5:52:13All right, we'll continue. Thank you
- 5:52:14guys. All right, so we're talking about
- 5:52:17uh comparison. So, uh we're going to
- 5:52:20talk about a few operators that allow us
- 5:52:22to compare two values. Now, this is
- 5:52:24going to be useful as we go forward
- 5:52:26because sometimes we want to know when
- 5:52:27is a value bigger than something or less
- 5:52:29than something or equal to something,
- 5:52:31not equal to something. Those
- 5:52:33comparisons are going to be useful.
- 5:52:35um and we have a collection of operators
- 5:52:38to do that for us. So again, one that I
- 5:52:41think a lot of people get confused on is
- 5:52:43the um equals comparison operator which
- 5:52:47is uh the double equals symbol. So a lot
- 5:52:52of people get confused on that. What is
- 5:52:54the difference between a single equal
- 5:52:55sign and a double? This double equal
- 5:52:58sign is checking if two values are
- 5:53:02equal.
- 5:53:03Um
- 5:53:05so for instance we have uh these two
- 5:53:08numbers x and y they're both integers
- 5:53:10that are 20. We check if x equals equals
- 5:53:13to y and that returns true because
- 5:53:18uh these two values are the same. They
- 5:53:20both equal 20. So when x equals equals y
- 5:53:23that is a true statement. So these
- 5:53:26that's something to realize is that
- 5:53:27these comparisons are things that return
- 5:53:29booleans true or false because a number
- 5:53:32is going to be bigger than another yes
- 5:53:34you know true or false they they are a a
- 5:53:38uh comparison that gives us a kind of a
- 5:53:40yes or no answer. Um
- 5:53:44so the equals equals checks if two
- 5:53:46values are the same and then the uh not
- 5:53:50equals operator which is uh an
- 5:53:52exclamation point with an equals um
- 5:53:55checks to see if two values are
- 5:53:57different. So they are not equal. So for
- 5:54:00instance if we had um uh 45 and 24 we we
- 5:54:04uh do x not equals y that would return
- 5:54:07true.
- 5:54:08Um now if we had these two values as
- 5:54:11before and we checked here x not equals
- 5:54:14to y um this would be false because they
- 5:54:18are equal right so um
- 5:54:21not equals to checks if values are
- 5:54:24different so that's a simple comparison
- 5:54:26are they not equal um so so in this case
- 5:54:29that would return true
- 5:54:32so these are pretty useful if we want to
- 5:54:34compare directly is a value equal to
- 5:54:36another we use the equals equals If
- 5:54:38they're different, we use the not
- 5:54:40equals. And we're going to have
- 5:54:42different scenarios where we will use
- 5:54:43those.
- 5:54:45I also want to call out the basic, you
- 5:54:48know, greater than and less than. So the
- 5:54:50this first one is the less than
- 5:54:52operator. It is uh going to be obviously
- 5:54:54returning true when a number is less
- 5:54:56than another number. So when we have
- 5:54:59things like 20 and uh 30, this x less
- 5:55:04than y would return true because 20 is
- 5:55:06definitely smaller than 30. So this
- 5:55:08returns true. Um
- 5:55:11and then greater than checks if a number
- 5:55:13is bigger than another. So that
- 5:55:15comparison uh x bigger than y in this
- 5:55:18case would uh return true as a
- 5:55:21comparison. So again these are all
- 5:55:23operators that check uh comparison
- 5:55:26between two numbers that will be uh
- 5:55:30really useful as we go forward and start
- 5:55:32to work with data and numbers and we do
- 5:55:35comparisons.
- 5:55:36uh we will do those all the time later
- 5:55:38on.
- 5:55:42Now there's also scenarios when when we
- 5:55:44want to know is it less than or equal
- 5:55:46to. So that operator just tacks on an
- 5:55:49equal sign. So less than equals
- 5:55:52is the less than or equal to operator.
- 5:55:55So for instance 10 less than or equal to
- 5:55:5730 that is true um because 10 is
- 5:56:00certainly smaller than 30. But um it
- 5:56:03would have been true even if x was 30.
- 5:56:06That would also be true because 30 uh 30
- 5:56:10equals to um 30 would equal to 30. That
- 5:56:14would be a true statement.
- 5:56:16Um greater than or equal to same same
- 5:56:19scenario. We have a greater than and
- 5:56:21then we have an equal sign right after
- 5:56:22it. This returns true if something is
- 5:56:24bigger than or equal to another number.
- 5:56:27So here's an interesting one. We do 30
- 5:56:29bigger than or equal to 30. That returns
- 5:56:31true because 30 equals to 30. That that
- 5:56:34makes sense. So less than or equal to
- 5:56:36bigger than or equal to we can we can do
- 5:56:38with these simple operators.
- 5:56:42Uh is greater than greater than similar
- 5:56:44to the usage of brackets? No.
- 5:56:47Uh so greater than or greater than is
- 5:56:49what's called a uh a bit shift operator.
- 5:56:54It's a little bit different. Um I I
- 5:56:57would I'm going to save any explanation
- 5:57:00that just just look that one up is what
- 5:57:02I'll say. It it does like a bit um a bit
- 5:57:05manipulation which is um a bit of a bit
- 5:57:10of a hassle to deal with but we we won't
- 5:57:12ever use greater than or we won't ever
- 5:57:14use greater than greater than. It's it
- 5:57:16does some sort of a shifting operation
- 5:57:18like a bit mathematics which we we don't
- 5:57:21need to do.
- 5:57:29Okay. So those are comparisons. Um let's
- 5:57:32look at our logical operators. So now
- 5:57:35these ones are going to be really really
- 5:57:36interesting and useful when we get into
- 5:57:38controlling the flow of our program. Um
- 5:57:42so logical operators are used for
- 5:57:44combining conditional statements. So
- 5:57:47conditional statements are things that
- 5:57:49return these are statements that return
- 5:57:52um true or false. So they return a
- 5:57:54boolean and we can it's it's like we are
- 5:57:57combining them together in certain ways.
- 5:58:00Okay,
- 5:58:02so the and operator, let's look at that
- 5:58:06one first, which in Python is the
- 5:58:08literal word and. So that's very nice.
- 5:58:11It's it's literally the the keyword and.
- 5:58:14Um, and what this does is it takes the
- 5:58:18result of some boolean comparison and
- 5:58:20some other boolean comparison and
- 5:58:22returns true if both of them are true.
- 5:58:26So and will only return true as a
- 5:58:30combination if both individual
- 5:58:32statements are true. They both have to
- 5:58:34be true. The moment one of them is false
- 5:58:36and will return false.
- 5:58:38So this is useful for doing a
- 5:58:41combination of things where we want
- 5:58:43every individual thing to be true. So a=
- 5:58:471. This is a true statement because a
- 5:58:50equals 1 and then b= 2 is a true
- 5:58:52statement. So both of these would be
- 5:58:54true. So therefore when we combine them
- 5:58:57with the and this overall combination is
- 5:59:00true.
- 5:59:02So keep that in mind. These operators
- 5:59:05are ones that combine individual logical
- 5:59:09statements or conditional statements.
- 5:59:11Right?
- 5:59:13Okay. Now the one that is less
- 5:59:16restrictive than and is the or statement
- 5:59:19which um is used when you only want at
- 5:59:24least one of the statements to be true.
- 5:59:27So if we want to combine these things
- 5:59:28and only require at least a minimum of
- 5:59:31one to be true, we use the or statement.
- 5:59:34So for instance, a= 1 is true because a=
- 5:59:391. So that's true. and then B equals
- 5:59:42equals to 2 is false. So this one is
- 5:59:45false. But that doesn't matter from the
- 5:59:47perspective of or because we have a
- 5:59:49minimum of one of these statements being
- 5:59:51true. So or when we use the logical
- 5:59:53combination of or um we just need either
- 5:59:57or to be true. So a= 1 is true. So this
- 6:00:01overall returns true.
- 6:00:04So or is something that will combine
- 6:00:06conditional statements and return uh
- 6:00:08return true if at least one of them is
- 6:00:11true. If all of them are false or it
- 6:00:13would return false because none of them
- 6:00:15are none of them would be true.
- 6:00:20Okay. So we have and we have or and then
- 6:00:23we have not. So not is an interesting
- 6:00:26one. Not essentially reverses a boolean.
- 6:00:31So if we have a statement that is
- 6:00:33inherently true and we put a not in
- 6:00:35front of it, it will invert that to be
- 6:00:38false. If we if we have something that
- 6:00:40is false and we put a not in front of
- 6:00:42it, it will return uh true.
- 6:00:45One of the interesting examples in
- 6:00:47Python and this trips up people all the
- 6:00:49time is the fact that Python treats zero
- 6:00:54very specially. So the integer zero
- 6:00:58is
- 6:00:59oops the integer zero is inherently
- 6:01:02treated by Python as false.
- 6:01:06So Python treats zero as false and then
- 6:01:10every other integer as true. Basically
- 6:01:12being Python is indicating that it is
- 6:01:15something that is not zero. uh anything.
- 6:01:19So like um B equals to one would be
- 6:01:22treated as true because it's as long as
- 6:01:25it's something that's not zero
- 6:01:28then Python treats that integer as as a
- 6:01:32true boolean essentially. Um so why
- 6:01:36that's interesting is if you put a not
- 6:01:38in front of this this would actually
- 6:01:40return true because not false what is
- 6:01:44the opposite of false? It is true right?
- 6:01:46So not false it would return true. So we
- 6:01:51we will see not from time to time. Uh
- 6:01:54not shows up when we want to negate
- 6:01:56something. So when we um you know you
- 6:01:59know maybe we have a an iteration an
- 6:02:02iterative loop and we say while not
- 6:02:05finished and we you know then we will
- 6:02:08execute a bunch of statements while we
- 6:02:10continue to not be finished and then the
- 6:02:12moment that that it finishes then it
- 6:02:14then the loop would be over. So not is
- 6:02:17powerful to kind of invert uh trus to
- 6:02:20falses and falses to true.
- 6:02:23Um so maybe we want to check if
- 6:02:25something is not empty. Meaning that um
- 6:02:28if it's empty
- 6:02:30uh if it's not empty that would be
- 6:02:32false. Not empty um you know maybe it
- 6:02:36would return true. So not is something
- 6:02:39we will uh see from time to time as a
- 6:02:42negation operator logical negation.
- 6:02:48Okay.
- 6:02:50Any questions on
- 6:02:52uh any questions on these operators?
- 6:02:56These now these we're going to use these
- 6:02:58in the control of the flow of our
- 6:03:00program.
- 6:03:05One more example for not. Yeah. So a
- 6:03:07pretty typical case for not would be
- 6:03:09something like this where um
- 6:03:12uh maybe we have some code oops I always
- 6:03:16forget to swap over to this maybe we
- 6:03:18have some code that checks so if we have
- 6:03:20a list
- 6:03:23if we have a list and it's uh empty
- 6:03:27okay and it has nothing in it let's say
- 6:03:29it has nothing in it we could we could
- 6:03:30have some code that says like if um if
- 6:03:35not
- 6:03:38uh list
- 6:03:41um then we then do something. So then if
- 6:03:47not list uh meaning that it's not empty
- 6:03:50then check then grab the first value.
- 6:03:52Let's say that grab the first value. So
- 6:03:54we'd have some code like this. So uh
- 6:03:58this this not is used to like we can
- 6:04:01negate the fact that this is going to be
- 6:04:03empty and then this would be true and
- 6:04:06then we can continue to access something
- 6:04:08because that would mean it's not empty.
- 6:04:11So not empty is a pretty standard use
- 6:04:15case for not like to check that
- 6:04:16something is not empty.
- 6:04:22Uh can we also use not for checking
- 6:04:24value in the list? Yeah, that's that's
- 6:04:26what we're doing here to say like is it
- 6:04:28not empty?
- 6:04:42Okay.
- 6:04:47All right.
- 6:04:49Let's go to some miscellaneous
- 6:04:51operators. So we have now some of these
- 6:04:54are going to be incredibly useful. The
- 6:04:56one on this page not that useful. The is
- 6:05:00mainly because it's very rare that we
- 6:05:03would uh that we would check these. So
- 6:05:07is is what we call the identity operator
- 6:05:10and this is something that um checks to
- 6:05:12see if two references are the same.
- 6:05:16Okay, two references are the same. um
- 6:05:20meaning that they're referencing the
- 6:05:21same piece of data. Um so now this is a
- 6:05:26very interesting case where we have a
- 6:05:28equals to a list b equals to the list.
- 6:05:32However, when we ask the question a is b
- 6:05:35this would actually return false. Now
- 6:05:37that seems very counterintuitive but the
- 6:05:39reason that's the case is because we are
- 6:05:41creating two different references. We're
- 6:05:44saying A equals to this list, B equals
- 6:05:47to this list, which is a whole new piece
- 6:05:49of data.
- 6:05:51It's a whole new piece of data. So
- 6:05:53therefore, we can't claim that they're
- 6:05:54the same reference even though they're
- 6:05:56the Now what would be true is A equals
- 6:06:00equals B because their data is the same.
- 6:06:04That would be true, but their their
- 6:06:06references are different because they're
- 6:06:09different variables, right? A and B are
- 6:06:10different variables.
- 6:06:12Different memory location. Exactly.
- 6:06:14Different references. So A and A is B is
- 6:06:18the same as checking um if ID A equals
- 6:06:23equals IDB. Does that make sense? That's
- 6:06:26basically checking that that logical uh
- 6:06:29comparison if their addresses are the
- 6:06:32same. It's the same check. So is is
- 6:06:35basically a shorthand for doing this.
- 6:06:39And so th those would be false because
- 6:06:41they're going to be two different
- 6:06:42references, A and B.
- 6:06:45Um, however, we can use our not. So A is
- 6:06:49not B. That's actually true because it's
- 6:06:51the inverse of is, right? So that that
- 6:06:54actually would invert the false and this
- 6:06:56would be true. A is not B. That is true.
- 6:07:02Behind the scenes, yeah, like the
- 6:07:04memory, yeah, the location in the
- 6:07:06computer memory is different. Yes,
- 6:07:08because they are different variables,
- 6:07:09different references.
- 6:07:15Yeah, the data is equal. The data is
- 6:07:17equal, but the references are different,
- 6:07:20which is what this checks. You know, we
- 6:07:23have two different names, A and B. Those
- 6:07:25are different.
- 6:07:27Now, take a look at this last example.
- 6:07:29This is saying A is a list. B equals to
- 6:07:32A. Now, remember what that does? That is
- 6:07:35the assignment of we're saying B is the
- 6:07:38same reference as A. That's what this
- 6:07:40does here. The same reference. So does
- 6:07:44it make sense to us that when we ask now
- 6:07:47A is B. This should be true. And it is
- 6:07:50like this is true because um
- 6:07:56this is true because they are literally
- 6:07:57the same reference. We're setting B
- 6:07:59equals to A. So they are referring to
- 6:08:03the same data. Now
- 6:08:05um in terms of a variable reference they
- 6:08:08so now their ids are the same
- 6:08:09essentially their memory locations are
- 6:08:12the same.
- 6:08:15If you want to compare only data then uh
- 6:08:19what operators have we looked at that
- 6:08:20are comparison
- 6:08:22we should think about that I mean we
- 6:08:24just saw if we go back a couple slides
- 6:08:26we have a bunch of operators to compare
- 6:08:28data that's these guys right like equals
- 6:08:31equals greater than less than so what we
- 6:08:35could do is say does the data equal the
- 6:08:38other data which would be something like
- 6:08:39this equals equals operator
- 6:08:42when two values are equal not the
- 6:08:45reference ES. Does that make sense? This
- 6:08:47equals equals is checking if two values
- 6:08:49are the same, which is the the data, not
- 6:08:52the not the reference.
- 6:09:02Okay.
- 6:09:05Now I want to show you a really powerful
- 6:09:08um I want to show you a really powerful
- 6:09:11miscellaneous operator which is going to
- 6:09:14be uh which is going to be the in or
- 6:09:18what's called the membership operator.
- 6:09:21So this is an operator that checks if a
- 6:09:25value is a member of a collection.
- 6:09:29So this could be like uh this could be
- 6:09:33like um you know where we have a list, a
- 6:09:38tupil, a dictionary, just a collection
- 6:09:40of data and we want to know is a value a
- 6:09:44member of that collection which is
- 6:09:46really useful for testing you know do we
- 6:09:49have membership of something inside of
- 6:09:51something else. So for instance, let's
- 6:09:54say we have a list and we have a list A
- 6:09:57equals and then we have 20 45 and 10
- 6:09:59inside that list. So if we ask the
- 6:10:01question
- 6:10:0310 in A, this actually would return true
- 6:10:06because 10 is a member of A. 10 is a
- 6:10:09member of that list. So that is true.
- 6:10:12Now that's useful to know. So the N
- 6:10:15operator is a really powerful uh really
- 6:10:18powerful operator.
- 6:10:20Um, same same thing with not. So, we can
- 6:10:23use not in. So, 10 not NA would be false
- 6:10:27because there it is. We know it's a
- 6:10:28member of A. So, that would be false. Of
- 6:10:31course, it's NA. We can see it right
- 6:10:33there. It's a member of that list.
- 6:10:36Um, but if we check the different value
- 6:10:38that's completely not inside of a, 30,
- 6:10:40not NA, that would return true because
- 6:10:4230 is not a member of that list. And by
- 6:10:46the way, this in operator works for all
- 6:10:49kinds of collections. So it would work
- 6:10:50for a tupole, it would work for a set,
- 6:10:53it would work for a dictionary.
- 6:10:56Um, it would work it would work for all
- 6:10:59kinds of collections.
- 6:11:03Yeah, Roberto. Um, it's the fact that
- 6:11:06there are different variables. So we may
- 6:11:09sometimes it makes sense to have
- 6:11:11different variables that are that maybe
- 6:11:14they have the same value but they're
- 6:11:16different variables altogether different
- 6:11:18references
- 6:11:20and the reason that is is because maybe
- 6:11:21we have a copy of that data and then
- 6:11:25maybe we manipulate the this one. Maybe
- 6:11:28this is a copy of it and we manipulate
- 6:11:30this guy and we leave this guy the same
- 6:11:33to check the differences later.
- 6:11:38Yes, references are tied to the variable
- 6:11:41like A is a reference, B is a reference
- 6:11:45even though their data that they're
- 6:11:46pointing to is the same value.
- 6:11:50Maybe we just have a copy of it that and
- 6:11:53we manipulate one of those copies.
- 6:11:59Yeah, that's why
- 6:12:12What is the reason to check for what? If
- 6:12:14they're equal, like as references, A is
- 6:12:17B.
- 6:12:21Honestly, there's not many good reasons.
- 6:12:24Um, maybe if you want to know if
- 6:12:26something is a copy of something else,
- 6:12:27like you want to know that, like let's
- 6:12:29say you're checking later down the code
- 6:12:31and you have an A and a B and you want
- 6:12:32to know if one is a copy of the other,
- 6:12:34you can check and see if they're the
- 6:12:35same reference.
- 6:12:37That's the only reason I could think of
- 6:12:39why you would do that. It's rarely used.
- 6:12:42Rarely used, but it is an operator that
- 6:12:44I wanted to show you in case you do
- 6:12:46stumble across it um somewhere and
- 6:12:49you're reading about Python or something
- 6:12:50and you see the is
- 6:12:54Like that's the only reason I can really
- 6:12:55think of is to check if something is a
- 6:12:57copy of another meaning it's the same
- 6:12:59like maybe it's it's uh the same
- 6:13:02reference
- 6:13:04B equals to A then we could check A is B
- 6:13:06and we know that then they're the same
- 6:13:08reference.
- 6:13:11Yes, the is operator is comparing
- 6:13:12references not exact not the values.
- 6:13:14Yes, that's true.
- 6:13:23All right. How do we feel about this in
- 6:13:25operator? Like the the membership
- 6:13:28operator. Does that make sense? If
- 6:13:30you're checking if an if a value is a
- 6:13:33member of a collection,
- 6:13:36that's going to be highly useful later
- 6:13:38on. Highly useful. This one we will use
- 6:13:41quite a bit. The is we will probably
- 6:13:43rarely ever use, but but this one we
- 6:13:46will definitely use.
- 6:13:57Okay. So, I wanted to quiz you guys. Um,
- 6:14:01what is the main difference between the
- 6:14:02equals equals and the is operator? What
- 6:14:06is the main difference?
- 6:14:08We were we've just been discussing this,
- 6:14:10so hopefully this this one is easy. Been
- 6:14:14discussing it quite a bit.
- 6:14:33Yeah, Roberto, now now you know the
- 6:14:35answer. Perfect. Yeah, it is B. All you
- 6:14:37guys answering B. Perfect. It is B. Good
- 6:14:39job. You guys are right on top of that.
- 6:14:41Good job. So, just wanted to point that
- 6:14:43out. Like equals equals compares the
- 6:14:44values. We're doing a comparison. Um is
- 6:14:48checks if the references are the same,
- 6:14:51which is the variable reference like the
- 6:14:53memory location.
- 6:14:58Yes, it is. Identity is the same as Yep,
- 6:15:00that's what we mean. The references are
- 6:15:02the same.
- 6:15:06All right. So, wanted to do a short demo
- 6:15:08on the operators on comparison, etc.
- 6:15:11Wanted to just show off that demo um so
- 6:15:14you can see and practice with it a
- 6:15:16little bit. Uh so, let's hop over to
- 6:15:18demo six inside of the lesson one. I'm
- 6:15:22going to hop over to that.
- 6:15:30which will be our last thing we will do
- 6:15:32in lesson one and we'll move on to
- 6:15:33lesson two.
- 6:15:36Um, let me pull up demo six here.
- 6:15:44Give me a moment.
- 6:15:49All
- 6:16:01right,
- 6:16:05let me share my screen.
- 6:16:13All right, so demo six. Uh, hopefully
- 6:16:16you guys have access to this. This is
- 6:16:17the last one inside of lesson one. Um,
- 6:16:20now again, this one says try to use VS
- 6:16:23Code. You can if you want. Again, Collab
- 6:16:25works fine. You can use whatever you've
- 6:16:28been using, Jupiter, Collab, VS Code,
- 6:16:30whatever works for you. No big deal on
- 6:16:32which one you use.
- 6:16:37So, no worries on any of this.
- 6:16:48Oh, yeah. So, um, F is So, yeah, this
- 6:16:53that's a good question. What is F? So, F
- 6:16:55tells Python to format. F is short for
- 6:16:59format. It's basically format the the
- 6:17:02string which we're going to print by
- 6:17:04having some placeholders.
- 6:17:06And um
- 6:17:09we have variables called A and B. And
- 6:17:12this fills in the blank of these
- 6:17:14placeholders with whatever the values of
- 6:17:16A and B are. So F just allows us to
- 6:17:19format and fill in the blanks. Does that
- 6:17:22make sense? Like wherever these um
- 6:17:24braces are, we have a variable name
- 6:17:26inside of it A and B. And we are um
- 6:17:29going to fill in the blanks of those A
- 6:17:32and B um by uh you know by just
- 6:17:37replacing them whenever we do the print
- 6:17:39function.
- 6:17:42So there think of it as F is short for
- 6:17:45format.
- 6:17:47and we have a couple placeholders and
- 6:17:49those will be filled in by our variables
- 6:17:51A and B.
- 6:17:55Okay, so this demo uh does a bunch of
- 6:17:59operators. So it's going to do a bunch
- 6:18:01of comparisons where we input one number
- 6:18:04and we turn that into an integer. Input
- 6:18:06another number, turn that into an
- 6:18:08integer, and then do a bunch of
- 6:18:09comparisons. So I'm going to jump over
- 6:18:10to the notebook. I'll do that for us to
- 6:18:12to show that off. But that's all we're
- 6:18:15doing in in really in the beginning of
- 6:18:17this uh demo. Um so let me jump over to
- 6:18:23the notebook and show off that so we can
- 6:18:25see it
- 6:18:30but uh should be straightforward to
- 6:18:31follow because we've done a lot of that
- 6:18:32already.
- 6:18:38Okay. So hopping over to notebook. Again
- 6:18:40feel free to use whatever you want to
- 6:18:41use. You can use collab, you can use um
- 6:18:45uh Jupiter, you can use VS Code,
- 6:18:47whatever you use. Um let's store a
- 6:18:50variable as a and let's make it an
- 6:18:52integer
- 6:18:54and let's do enter your first number.
- 6:18:59So we will do that.
- 6:19:03Let's run that. So let's enter our first
- 6:19:05number. Let's put in 10 or whatever you
- 6:19:07want really, but I'm going to put in 10.
- 6:19:10So that gets stored as a. Now let's do a
- 6:19:13second number and let's do int
- 6:19:17input um enter your second number
- 6:19:24and let's input that.
- 6:19:28So now I'm going to put in a second
- 6:19:30number. Let's do 20.
- 6:19:35So now we have a and b. Now let's do
- 6:19:38some comparisons. So let's do print.
- 6:19:42Um then we can do uh let's do the f
- 6:19:46formatting like they had in there. Now
- 6:19:48what this is going to be is we are going
- 6:19:51to check a
- 6:19:54um greater than or let's do yeah let's
- 6:19:56do greater than b
- 6:20:00um is
- 6:20:02and then let's do uh comma a greater
- 6:20:06than b.
- 6:20:08Let's compare those two numbers. So now
- 6:20:10this is doing the comparison. A greater
- 6:20:13than b is going to compare those two
- 6:20:15integers. What should this return? What
- 6:20:18should a greater than b return? Should
- 6:20:21it be true or should it be false?
- 6:20:27Should be false. Right? So we should we
- 6:20:30should uh display false here.
- 6:20:36And that's what it is. 10 greater than
- 6:20:3820 is false.
- 6:20:41Okay, so that is false.
- 6:20:44Let's do another one.
- 6:20:47Let's do um let's try equals equals. So
- 6:20:50let's say um a
- 6:20:54equals equals to b is now what do we
- 6:20:58think this one's going to be?
- 6:21:03a equals equals to b.
- 6:21:06What should this one be?
- 6:21:10Yes, very good. This one should also be
- 6:21:14false.
- 6:21:18Let's run that.
- 6:21:20And that one will be false. Very good.
- 6:21:26Okay.
- 6:21:28So, I think we get that. Let me give you
- 6:21:29guys a Let me uh show you something
- 6:21:31interesting. Let me uh go a little off
- 6:21:34script from that demo document and let's
- 6:21:36introduce a third number called C. Let's
- 6:21:40do a third integer.
- 6:21:43So enter your third number.
- 6:21:47Let's do a third number.
- 6:21:52Let's enter um another value of 10.
- 6:22:03Okay. So, we have another number of 10.
- 6:22:07Now, what I want to do is let's just do
- 6:22:12uh multiple comparisons and do a logical
- 6:22:16operator between them. So let's do um a
- 6:22:21not equal to b
- 6:22:24and
- 6:22:27a less than c
- 6:22:31or sorry b
- 6:22:33less than c.
- 6:22:37What do we think this is going to
- 6:22:38return? This might be a little
- 6:22:39challenging. What do you think this is
- 6:22:41going to return?
- 6:23:01We should use F. We can. I'm just not
- 6:23:03printing. I'm just going to I'm just
- 6:23:05going to run the cell. I'm I'm kind of
- 6:23:07doing a shortcut and just print. I'm not
- 6:23:08going to print. I'm just going to run
- 6:23:10the cell and it should display what it
- 6:23:11is.
- 6:23:12Do we think it's going to be false?
- 6:23:14Yeah, you guys are right on top of it.
- 6:23:17Should be false. Now, let's break that
- 6:23:20down. Why is that false? So, A not equal
- 6:23:22to B is checking if A is not equal to B,
- 6:23:25which is true.
- 6:23:27A has a value of 10.
- 6:23:32So, A has a value of 10. So, 10 is not
- 6:23:35equal to 20. That makes sense. But 20 is
- 6:23:39not less than 10. So this part is false.
- 6:23:43So let's make a comment.
- 6:23:46So the reason reason this is false
- 6:23:50is because B is not less than C. So and
- 6:23:58returns false.
- 6:24:00Yes. Very good.
- 6:24:04Now,
- 6:24:06what if I take this code
- 6:24:10and do this?
- 6:24:16What's this?
- 6:24:26Perfect. Yeah, you guys are right on top
- 6:24:28of it. This should be true. And it is.
- 6:24:31Now that's because the moment we have at
- 6:24:34least one true which is going to be this
- 6:24:38that makes sense right that is going to
- 6:24:41be true.
- 6:24:54Okay, one more and then we can wrap up
- 6:24:57this demo. So I want to create a list.
- 6:25:00I'm going to call it X and I'm going to
- 6:25:02create a list of three numbers 10 20 30.
- 6:25:08Okay. Now, what do you think uh this
- 6:25:12result is?
- 6:25:14What what should this be?
- 6:25:23It should be true.
- 6:25:25Very good. Should be true.
- 6:25:32Now what should what should this be?
- 6:25:41This is also true. Very good. This
- 6:25:43should this should be true because this
- 6:25:45is not going to be inside of the list.
- 6:25:48So that makes sense. That is not true.
- 6:25:52Now what I want you to notice is nothing
- 6:25:54will change if I change this to a
- 6:25:55tupole.
- 6:25:57Nothing will change. We can still check.
- 6:26:00So we can still check if uh 10 is a
- 6:26:03member of this tupole and we can still
- 6:26:04check if 40 is not a member of this
- 6:26:06tupole. So nothing really changes,
- 6:26:08right? It's still in membership operator
- 6:26:11in checks is it a member of any
- 6:26:14collection?
- 6:26:19Okay,
- 6:26:21very good. Any questions about the these
- 6:26:23examples?
- 6:26:25Any questions? Do we feel comfortable
- 6:26:28with some of these operators? You guys
- 6:26:29were right on top of it. It was very
- 6:26:30impressive. You guys got those right
- 6:26:32away.
- 6:26:35Uh, you change print
- 6:26:39a= b and a is b. And I entered four both
- 6:26:44times
- 6:26:46and I got the same result. True.
- 6:26:53Uh so you had your A and your B. So you
- 6:26:56you had um
- 6:26:59you had A equals to 4
- 6:27:02and B equals to 4 and you checked
- 6:27:06uh you checked this.
- 6:27:16Yes.
- 6:27:20Yeah. So that is a little confusing and
- 6:27:23I can understand why. So the reason this
- 6:27:25ends up being true is because this is a
- 6:27:28scalar data. So for scalar data it's
- 6:27:32going to optimize in the memory to point
- 6:27:35because four the integer four occupies
- 6:27:39the same memory address always. But when
- 6:27:43we create an array, when we create a
- 6:27:45list, that is um a new object.
- 6:27:49So yeah, that one's a little confusing,
- 6:27:51but it's it's only because this is
- 6:27:54scalar data that um Python kind of
- 6:27:57knows, okay, four is the same like
- 6:28:01integer in the me in memory always
- 6:28:05regardless of if we're referencing it
- 6:28:07from this from two different variables.
- 6:28:09That's that's the reason I um yeah I
- 6:28:12forgot to mention that example but it's
- 6:28:15purely so the the reason this is true is
- 6:28:18because
- 6:28:20this is true because of scalar
- 6:28:25data optimization. Essentially it's it's
- 6:28:29not going to waste creating a new object
- 6:28:31when it's just a single integer four. it
- 6:28:33basically occupies the same memory
- 6:28:35address
- 6:28:37uh as as a as a reference.
- 6:28:42But when we build a list like that is a
- 6:28:44different object.
- 6:28:52Okay.
- 6:28:54Very good.
- 6:28:58All right. So,
- 6:29:00um, that will wrap up lesson one. What
- 6:29:05I'm going to do is go into lesson two.
- 6:29:07I'm going to pull up the lesson two
- 6:29:09notes, the slides for lesson two. So, if
- 6:29:11you have those, let's pull those up. Um,
- 6:29:14I do encourage you now, there is a
- 6:29:16guided practice at the end of lesson
- 6:29:17one, and that is for your yourself. I
- 6:29:21would encourage you as as kind of
- 6:29:22homework between now and and the next
- 6:29:24time we meet um to do the guided
- 6:29:27practice for lesson one. Try that out on
- 6:29:30your own. It is um you guys should have
- 6:29:32access to it from your LMS and the
- 6:29:34reference materials. There's a guided
- 6:29:36practice. Try that. Try those out. Okay?
- 6:29:40Try out the guided practice for lesson
- 6:29:41one. It's just it's just some additional
- 6:29:44practice of the things we just covered.
- 6:29:49Okay?
- 6:29:50Let me open up
- 6:29:53um
- 6:29:55lesson two.
- 6:30:04Give me a moment here.
- 6:30:10Okay.
- 6:30:12So, let me share my screen.
- 6:30:24Okay, so now we're going to move on to
- 6:30:27looking more closely at those things
- 6:30:29like lists, tupils, dictionaries, and
- 6:30:31then looking at control flow with
- 6:30:33conditional statements and loops. So
- 6:30:35we're going to get into the fun stuff, I
- 6:30:37would think, um that you guys may may
- 6:30:39have been waiting for.
- 6:30:43Okay, so we've talked about uh we just
- 6:30:45finished talking about lesson one where
- 6:30:46we have you know Python as a really
- 6:30:48important thing to learn and study
- 6:30:50because it's used all over the place
- 6:30:52with data science and a IML. So one of
- 6:30:54the things is we need to continue
- 6:30:56learning about it with things like loops
- 6:30:59things like if else statements to
- 6:31:01control the flow of our programs and
- 6:31:04these basic data structures.
- 6:31:08Let's continue forward. Um so this
- 6:31:11lesson we're going to talk about list
- 6:31:13tupils dictionary sets uh we're going to
- 6:31:15you know talk about the differences how
- 6:31:18we can access data within things like
- 6:31:20list how we can modify them um how we
- 6:31:22can access things from tupils what are
- 6:31:24the differences we'll review all that
- 6:31:26one of the big topics is going to be to
- 6:31:29um look at how we can control the flow
- 6:31:32meaning control the flow is using like
- 6:31:34decision logic like if this is true then
- 6:31:37do this else do this we'll talk about
- 6:31:39those kind of statements. We'll talk
- 6:31:41about iteration. So how we can do loops
- 6:31:44to repeat um statements of code that we
- 6:31:47want to do. Um we'll also talk about
- 6:31:49organizing our code a bit into
- 6:31:51functions. Um which is going to be
- 6:31:53really helpful to for our own
- 6:31:55organization and reuse and
- 6:31:57maintainability.
- 6:32:02Okay.
- 6:32:05So let's jump into it. Let's talk about
- 6:32:08some of those data structures.
- 6:32:11So, we've already talked about that
- 6:32:13these aggregate data structures exist um
- 6:32:17and they allow us to manipulate data
- 6:32:19inside of Python. Lists, tupil, sets,
- 6:32:22dictionaries are the main ones we're
- 6:32:23going to focus on.
- 6:32:26Let's start with lists. And I think
- 6:32:28lists are going to be something we're
- 6:32:30going to use quite a bit of throughout.
- 6:32:33So, they're going to be a really good
- 6:32:34place to start with. They are super
- 6:32:36popular in Python. A lot of people um
- 6:32:38use them to do to work with data. They
- 6:32:41are kind of the most basic um data most
- 6:32:45basic and useful data structure that
- 6:32:46there there is.
- 6:32:50So what is a list? A list is a an
- 6:32:53ordered mutable meaning it can be
- 6:32:56modified data structure that can hold
- 6:32:59elements of different data types. So
- 6:33:00it's a collection of data and there's no
- 6:33:04requirement that all the members of the
- 6:33:06list be the same type. In fact, you can
- 6:33:07have different types. You can have
- 6:33:09integers, you can have floats, you can
- 6:33:10have strings,
- 6:33:12um you can have even more abstract
- 6:33:14objects be members of a list. Um so any
- 6:33:18kind of data can live within a list. But
- 6:33:20the big thing is that it is modifiable.
- 6:33:23It's dynamic. You can change you can add
- 6:33:26things to it. You can remove you can
- 6:33:27change things. Um and it has an inherent
- 6:33:31ordering which is nice. So you can you
- 6:33:33can be reassured that there is some
- 6:33:36inherent position of items and it will
- 6:33:39maintain that order. So we can access
- 6:33:41things based on the order like we can
- 6:33:43access the first can access the last we
- 6:33:45can access anything in between.
- 6:33:48So that's nice. Um so what are some key
- 6:33:52characteristics? So uh lists support
- 6:33:55multiple data types. We talked about
- 6:33:57that. There's no requirement that
- 6:33:58they're all the same. They can have
- 6:33:59multiple. They allow for indexing, which
- 6:34:03we're going to talk about. This means
- 6:34:05that we can access things based on their
- 6:34:08index, which is their position within
- 6:34:10the list. So, we can always access the
- 6:34:12first thing, the last thing, anything in
- 6:34:14between, based on its position. That's
- 6:34:17another word for uh index because lists
- 6:34:21have natural ordering to them which is
- 6:34:24um really really powerful to ensure that
- 6:34:27there's um one you know there's a first
- 6:34:30position a second position third etc. So
- 6:34:34lists are really nice for that.
- 6:34:36Um
- 6:34:38uh they are modifiable which is really
- 6:34:41nice. So we can add things into the
- 6:34:43list. It's very dynamic. So once we
- 6:34:44create a list, we can throughout our
- 6:34:46program, we can add data to it, we can
- 6:34:48remove it, we can change. Um, lists also
- 6:34:51allow duplicates, which may be
- 6:34:53desirable. Like maybe we add something
- 6:34:55into our list that already exists.
- 6:34:57That's okay. List allow duplicates. This
- 6:35:00is going to be different than sets. Sets
- 6:35:03do not allow duplicates. Sets are just a
- 6:35:05bucket of unique things. So if we added
- 6:35:08a duplicate into a set, it would reject
- 6:35:10it. we wouldn't have any errors, but it
- 6:35:13just wouldn't um show up as a copy. We
- 6:35:16would just have a set is only going to
- 6:35:18maintain one copy of an item. It it only
- 6:35:21allows unique items. Lists allow you can
- 6:35:23have as many copies of data as you want
- 6:35:25inside of a list. Um so it does allow
- 6:35:28duplicates.
- 6:35:30Now, we already saw in terms of syntax,
- 6:35:33lists are um defined by brackets. So
- 6:35:37when you see those brackets um it
- 6:35:39defines a list and its items are
- 6:35:41separated by uh commas.
- 6:35:45So
- 6:35:47we can have a list that looks like this.
- 6:35:53Yeah, I was going to explain slice. We
- 6:35:54have a a couple slides about slicing
- 6:35:56coming up, but slicing just means that
- 6:35:59we can slicing means that we can grab a
- 6:36:02section of elements at a time from a
- 6:36:04list. So for instance we can uh actually
- 6:36:08let me use this example down here. A
- 6:36:10slice would mean we can grab like these
- 6:36:12first three slice of the of the list or
- 6:36:16we can grab the last five elements or
- 6:36:19whatever like this is a slice. It's just
- 6:36:22a a subset of the list that we can grab
- 6:36:25we can access.
- 6:36:27So slicing just means taking a subset,
- 6:36:31taking a smaller section of the list and
- 6:36:34we can grab all those elements at a
- 6:36:36time. And what that's called is within a
- 6:36:38slice.
- 6:36:42Yeah, like a slice of pizza. We're
- 6:36:44taking the whole thing and we're taking
- 6:36:45a small section of it.
- 6:36:50So that's and actually that's going to
- 6:36:53be possible because the list has a
- 6:36:56natural order to it.
- 6:37:01Does the data have to be sequential? No,
- 6:37:03it doesn't have to be. In fact, you it
- 6:37:05can be completely different types.
- 6:37:08Does that make sense? Like so look at
- 6:37:10this example down here. Like we have 10
- 6:37:122 5 hello. That's a valid list. You can
- 6:37:16have different types of data in there
- 6:37:17which isn't sequential at all
- 6:37:29in a slice. No, it doesn't have to be.
- 6:37:32So you can have like you can have a
- 6:37:34slice that picks every third element
- 6:37:38uh every other element. Um yeah, it
- 6:37:42doesn't have to be sequential. No, it
- 6:37:45can be customizable.
- 6:37:49What's also nice is you can slice from
- 6:37:51the beginning or you can slice from the
- 6:37:52end as well. So you can you can go from
- 6:37:54the end and slice backwards. Um or you
- 6:37:58can go from the beginning and slice
- 6:37:59forwards. So you can grab like every
- 6:38:01other element from the beginning. You
- 6:38:03can grab every other element from the
- 6:38:04back and work your way forward and stop
- 6:38:06at a certain point.
- 6:38:10Slicing is very nice. Yeah. So, I'm
- 6:38:12going to show us how to do that.
- 6:38:16Can you slice in the middle? Yeah, you
- 6:38:17can slice anywhere you want.
- 6:38:21Can you slice a pizza in the middle?
- 6:38:23Sure. Would you do that? Maybe not. But
- 6:38:26yeah, you can slice anywhere in the
- 6:38:28list. You can slice.
- 6:38:37Does slicing change the original list?
- 6:38:39No, it's just selection of a subset.
- 6:38:42No, it just it just extracts elements.
- 6:38:45It doesn't like permanently change it in
- 6:38:47any way. It just gives you a view. Think
- 6:38:51of it as like giving you a view of that
- 6:38:53subset.
- 6:39:00Okay, so you may be wondering when would
- 6:39:03I ever use list? So normally you use
- 6:39:06lists whenever you want an ordered
- 6:39:09collection that is dynamic meaning it
- 6:39:12may you may want to add things from it.
- 6:39:15Um we may want to modify things from it
- 6:39:19frequently. Um but we want something
- 6:39:22that is dynamic and has an ordering to
- 6:39:24it. Lists are perfect for that reason.
- 6:39:26So they can contain data that we can add
- 6:39:29to remove from change. So lists are very
- 6:39:32versatile. I think most use cases with
- 6:39:34manipulating
- 6:39:36um a collection of items would fall
- 6:39:39under a list. A list would be a very
- 6:39:41good choice. Um
- 6:39:46we can slice elements and use them.
- 6:39:48Yeah. Yeah. You can slice you can store
- 6:39:51a slice inside of a variable and use use
- 6:39:54the use that resulting variable. Yeah,
- 6:39:57for sure.
- 6:40:11order information like
- 6:40:16Yeah. Yeah, that's a good example. Yeah,
- 6:40:18order like the collection of orders
- 6:40:22would probably be in a list because it
- 6:40:23can change. Um, but the prices would be
- 6:40:26probably static. So, that would be more
- 6:40:29uh suitable for a tupole. Yeah, a tupole
- 6:40:32or maybe a dictionary. A dictionary is
- 6:40:34probably better because you can have
- 6:40:35like a product ID or a name that maps to
- 6:40:38a price. Probably a dictionary would be
- 6:40:40more appropriate for a price list. But
- 6:40:43but yeah, tupole maybe makes sense too.
- 6:40:51By the way, look at this list. Do you
- 6:40:52guys see how it has different types of
- 6:40:55data in it, right? It has so like the
- 6:40:57first element is an integer, the next is
- 6:40:59a string, the next is a float, the last
- 6:41:02is a boolean. That's totally valid in
- 6:41:04Python, which is kind of unique to
- 6:41:07Python. Like a lot of languages don't
- 6:41:09support that. A mixtyped array.
- 6:41:13Don't really they don't really have that
- 6:41:15notion of that.
- 6:41:26All right. So, what I wanted to get to
- 6:41:28was the positions. So, this is going to
- 6:41:31be this is really important to pay
- 6:41:33attention to because this is going to be
- 6:41:35something that we will be using
- 6:41:38throughout the program is how to access
- 6:41:41data by its position.
- 6:41:44Okay, which so there's another word for
- 6:41:46that. The position um sometimes you will
- 6:41:48hear called the index. So the index in
- 6:41:51the list of where these where data
- 6:41:54members live is their order like their
- 6:41:56position amongst amongst the list.
- 6:42:00So what's special about Python is that
- 6:42:04it has the first position
- 6:42:08is index zero which trips people up all
- 6:42:13the time. The very classic trip up of
- 6:42:17Python is that the first element of a
- 6:42:20list is at position zero. The next
- 6:42:23element is at position one. The next
- 6:42:26element is at position two. On and on
- 6:42:28and on. The last element is at position
- 6:42:32n minus one where n is the size of the
- 6:42:38um size of list. So, however many
- 6:42:41elements we have in our list.
- 6:42:43Um,
- 6:42:45so
- 6:42:47if you want to access the first element,
- 6:42:49you would be looking for the element
- 6:42:51that's at position zero. If you want to
- 6:42:54access the second element, that's this
- 6:42:56this name Bob, that is at position
- 6:42:59number one or index one. So that's
- 6:43:02something that trips up people is that
- 6:43:04it actually starts
- 6:43:08starts at zero
- 6:43:11which is really important to understand
- 6:43:12is that the positions start at zero.
- 6:43:18Why is that? Um that's a good question.
- 6:43:21I mean
- 6:43:23different so so different languages
- 6:43:26treat that differently. Um,
- 6:43:29it's more historical reasons that it was
- 6:43:32created that way. Uh, but I think it's
- 6:43:36based on your like
- 6:43:40think of it as like how far into the
- 6:43:42list you are. So, if you're in the
- 6:43:44beginning, that means you're basically
- 6:43:46at at um position zero because you
- 6:43:49haven't made any progress like
- 6:43:50traversing the list. I think that was
- 6:43:53the intuition.
- 6:43:57Oh, yeah. which is kind of what Tim says
- 6:43:58there is like yeah if you're at the very
- 6:44:01beginning of the list you're at position
- 6:44:02zero because you haven't made any you
- 6:44:04haven't made any forward progress in
- 6:44:06traversing so it's like you're at step
- 6:44:08zero you're at the beginning
- 6:44:17okay but okay so aside from why it was
- 6:44:20that way does it make sense that that
- 6:44:22the first like what I'm saying is the
- 6:44:26position zero is the first element,
- 6:44:27position one is the second element,
- 6:44:29position two is the third element, and
- 6:44:31on and on and on.
- 6:44:33That's that's how it works in Python.
- 6:44:37Okay.
- 6:44:39Now,
- 6:44:40what I'm telling you is the positions if
- 6:44:43you were to view it going left to right.
- 6:44:45So, in other words, in the forward
- 6:44:47direction, we start from zero and go up
- 6:44:49to the to the n minus one in terms of
- 6:44:52the position. Now, what's really
- 6:44:55convenient is the positions can also be
- 6:44:58indexed from back to front, meaning they
- 6:45:02can also be indexed going this way.
- 6:45:05And what's really nice is the very last
- 6:45:08element starts at position minus one.
- 6:45:14The very last element on the right
- 6:45:16starts at minus1 and then goes all the
- 6:45:19way up to minus n.
- 6:45:22Okay, now that's really convenient
- 6:45:24because if we want to access the very
- 6:45:26last element, I don't need to know how
- 6:45:28big the list is. I just need to access
- 6:45:30the position minus one. That guarantees
- 6:45:34to access the very last element. The
- 6:45:38minus one position is the very last
- 6:45:40element.
- 6:45:42And then it it so this is the very last
- 6:45:45element. This is the second to last,
- 6:45:48third to last, fourth to last, on and on
- 6:45:51and on. And then by the time you get to
- 6:45:53the front, it is minus n, which is like
- 6:45:57the
- 6:45:59uh number of elements in the list minus.
- 6:46:03So in this case, minus 6 is the very
- 6:46:05beginning.
- 6:46:06Now, why why in the world do we care
- 6:46:08about negative position? It's for that
- 6:46:10exact reason.
- 6:46:11we can um we can think about the
- 6:46:14positions from the end of the list going
- 6:46:17backward which is very powerful like if
- 6:46:19I want to access things from the very
- 6:46:21end I don't need to know exactly how
- 6:46:23many there are I just need to know minus
- 6:46:25one is the end minus two is the second
- 6:46:27to last minus3 is the third from last um
- 6:46:31which is pretty convenient
- 6:46:36okay
- 6:46:38so does that make sense on the negative
- 6:46:40index index. It's the negative index is
- 6:46:43going from right to left. It's from the
- 6:46:46back to the front. Minus one is the last
- 6:46:49element.
- 6:46:51And then you go second to last, third to
- 6:46:53last, right? And that is minus 2, -3,
- 6:46:57-4.
- 6:47:06So from right to left. Yeah. Right.
- 6:47:08Right to left. minus 6 is actually the
- 6:47:10first element
- 6:47:12because it's six back from the end which
- 6:47:14is the f which is the front. There's
- 6:47:16only six elements.
- 6:47:22Yeah.
- 6:47:28So are there any question about this is
- 6:47:30really important to understand really
- 6:47:33important to understand because this is
- 6:47:35how we are going to slice and access
- 6:47:37data is based on these positions.
- 6:47:45Uh is there syntax to get the number of
- 6:47:47elements in a list? Yes, it's the length
- 6:47:49function which is len. So length of list
- 6:47:55would give you uh the number of elements
- 6:47:58which in this case is six. So ln the
- 6:48:01length function gives you how many
- 6:48:03elements are in the list.
- 6:48:06len or length. So this guy this
- 6:48:11function.
- 6:48:17Uh when are you counting backwards?
- 6:48:19You're usually counting backwards when
- 6:48:22you want to know what's at the end of
- 6:48:24the list. So when you only care what has
- 6:48:26been added at the end and you want to
- 6:48:29maybe you want to get the last five
- 6:48:31elements.
- 6:48:32So you slice backwards from the from the
- 6:48:35end.
- 6:48:37That's that may be useful like maybe you
- 6:48:39want to know like it imagine a list is
- 6:48:42holding in your orders
- 6:48:44and so you want the last five orders. So
- 6:48:47you can just go from the back and go
- 6:48:49towards the front. Min -1, -2, -3, -4,
- 6:48:52-5
- 6:48:55would be the last five orders. There's
- 6:48:58going to be many scenarios where we want
- 6:48:59to count from the end.
- 6:49:02uh when we're manipulating data
- 6:49:05later on when we're when we're working
- 6:49:07with um bigger sets of data um in
- 6:49:10something like an like a matrix, it's
- 6:49:13going to be useful to grab like the last
- 6:49:15five rows, last 10 rows.
- 6:49:19So going from the end makes more sense.
- 6:49:21So you imagine if we have like a larger
- 6:49:23matrix of data um maybe we want to slice
- 6:49:27out these last 10 rows in which case we
- 6:49:30want to count from the minus like the
- 6:49:32last row is minus one and we want to go
- 6:49:35back towards the front.
- 6:49:43Can you sort? Yes, you can sort. Uh I
- 6:49:46I'll have an example of that in a
- 6:49:47second. Yeah, you can sort.
- 6:49:50There's a built-in sorting function in
- 6:49:53Python that that will allow you to sort
- 6:49:54data in a list. Yes,
- 6:49:57there's a built-in function for that.
- 6:50:07Okay. What I wanted to do is show you
- 6:50:11guys an example of accessing elements
- 6:50:13from the list. So assuming we know the
- 6:50:18position which is that index all we have
- 6:50:20to do is use brackets to access items of
- 6:50:24a list. So imagine we have this list
- 6:50:26called fruits which has some strings in
- 6:50:29it apple banana cherry mango and we want
- 6:50:33to access the first element. So that is
- 6:50:36just this code here fruits bracket zero.
- 6:50:40So the bracket tells the interpreter,
- 6:50:43hey, I want to access something within
- 6:50:45this list. And then all we have to do is
- 6:50:47give it the position that we want to
- 6:50:49access. So this is position zero, which
- 6:50:53is going to be the first element. This
- 6:50:54would retrieve the first element. Now,
- 6:50:58it's not removing it. It's just
- 6:51:02accessing it so we can view what that
- 6:51:05is. So, it's actually going to give us a
- 6:51:07copy of what that value is under the
- 6:51:10hood. Basically, a copy of it. It's not
- 6:51:12going to permanently. It's not going to
- 6:51:13delete it. It's not going to remove it.
- 6:51:16It's going to give us a
- 6:51:19copy of what is at position zero. In
- 6:51:22this case, apple.
- 6:51:24Now if we put in a position two in that
- 6:51:27bracket that should be remember the
- 6:51:29indexing is 0 1 2 3
- 6:51:34because there's four items. So the last
- 6:51:37element is n minus one which is three.
- 6:51:40So the item that's at position two is
- 6:51:43going to be cherry. So this should
- 6:51:45return the string cherry.
- 6:51:48Okay. So we use we always use this kind
- 6:51:52of syntax
- 6:51:55position
- 6:51:57to access the element that is at that
- 6:52:00position.
- 6:52:03Okay, pretty simple. And what's really
- 6:52:06nice by the way about this is this the
- 6:52:09same exact thing works for tupils
- 6:52:12because tupils also are ordered. So
- 6:52:15we're going to see that when we get into
- 6:52:16tupils. But the same exact things works
- 6:52:19where a tupole has a first item, a
- 6:52:21second item, a third which are index
- 6:52:22zero, 1, two, three. And we can access
- 6:52:25things in the exact same way. So we can
- 6:52:27access a tupole by its position as well.
- 6:52:31Same exact thing will happen position.
- 6:52:36So we can get the first item of a tupole
- 6:52:38with with uh by doing um zero and we can
- 6:52:42get the the last by doing minus one.
- 6:52:45and on and on.
- 6:52:55Any questions about the accessing?
- 6:53:01Can we get position based on value? Yes.
- 6:53:04So that is a special function called
- 6:53:06index.
- 6:53:08So if you do like list dot dot index
- 6:53:16and then you pass in a value like apple,
- 6:53:21this would return to you the index of
- 6:53:24the first occurrence. Not every
- 6:53:26occurrence, but just the first. So like
- 6:53:28because we could have multiple copies of
- 6:53:30Apple in our list, but the first time we
- 6:53:33stumble upon Apple, this would return
- 6:53:35this would return uh zero because apple
- 6:53:39occurs at index zero.
- 6:53:43So index function
- 6:53:46uh would return um the index
- 6:53:50which is which is the same word as
- 6:53:52position.
- 6:53:58Okay.
- 6:54:01All right. Any other questions? We can
- 6:54:03uh I think what we'll do is we can take
- 6:54:05if unless there's any other questions we
- 6:54:07can take a short break and then come
- 6:54:08back and continue.
- 6:54:11Uh, I want to talk about slicing.
- 6:54:22No, it would return none. It would
- 6:54:25return null. Basically none
- 6:54:36as opening closing braces are square on
- 6:54:40it if it's a list. No, it's actually the
- 6:54:42same as a tupole. Uh in terms of
- 6:54:44accessing it's the same
- 6:54:47uh
- 6:54:51it's it's in terms of accessing it's the
- 6:54:52same. Um, for creating a list, yes, it
- 6:54:56is square brackets. This creates a list.
- 6:54:59The brackets always create a list. For
- 6:55:02accessing elements, this is the same as
- 6:55:04a tupole. The the square brackets for
- 6:55:07accessing.
- 6:55:09Actually, in most things in Python, it's
- 6:55:12the same where we uh access data using
- 6:55:15the square brackets. A dictionary, same
- 6:55:18thing. We access keys by using the
- 6:55:21square brackets.
- 6:55:23So square brackets in Python is is
- 6:55:25basically like an access operator.
- 6:55:31Uh no, so list.index doesn't only work
- 6:55:34for the first item. What I'm saying is
- 6:55:36you can have lists that have multiple
- 6:55:38I'm saying it returns to you the first
- 6:55:41occurrence,
- 6:55:42the index of the first occurrence
- 6:55:44because we could have a copy of Apple
- 6:55:46later on in the list, right? So there's
- 6:55:49nothing that stops us from having a
- 6:55:51duplicate. So, we could have another
- 6:55:53apple down here. Like, let's say we had
- 6:55:55another apple at the end of the list. If
- 6:55:57I did list.index apple, it's only going
- 6:56:00to return to me this first one, even
- 6:56:03though there's another copy of it in the
- 6:56:04list.
- 6:56:07So, it only returns to you the first
- 6:56:09occurrence.
- 6:56:12But, you know, there's nothing stopping
- 6:56:13me from doing list.index of banana or
- 6:56:15cherry or whatever.
- 6:56:18Are there data types by this for all
- 6:56:20Unicode characters? Um, I think there's
- 6:56:23Yeah, I think there's special strings
- 6:56:25you can make that do Unicode,
- 6:56:28but I'm honestly not 100% sure.
- 6:56:32I would research that. I don't really
- 6:56:35know. I think you can do that with
- 6:56:37strings,
- 6:56:39special strings.
- 6:56:42Yeah, I don't think there's anything
- 6:56:43like I don't think there's anything
- 6:56:44inherently special about it that Python
- 6:56:48can't handle. It's just you sometimes
- 6:56:51have to like escape characters
- 6:56:54uh
- 6:56:56to distinguish them. But
- 6:57:00yeah.
- 6:57:02All right. Going back to our example,
- 6:57:07um we had the uh going back to that
- 6:57:12fruits list.
- 6:57:14We can access things from the end. So
- 6:57:16here's an example of us using the
- 6:57:18negative index, right? So minus one
- 6:57:21grabs the element that's at the last uh
- 6:57:24that's the last member of the list. So
- 6:57:26mango minus3
- 6:57:29um would be the uh third from last,
- 6:57:32which is going to be banana.
- 6:57:35So negative index really handy to access
- 6:57:39from the end of the list going backward.
- 6:57:42Um pretty useful there.
- 6:57:45There's an example of it.
- 6:57:51All right.
- 6:57:56All right. Let's talk about slicing.
- 6:57:58Okay. Let's talk about slicing. So,
- 6:58:00slicing allows us to extract a subset of
- 6:58:04items from a list using a specific range
- 6:58:08of indices. And so the syntax to do
- 6:58:11slicing
- 6:58:13is going to be using our brackets again,
- 6:58:16but um we use a colon to specify where
- 6:58:22we are starting and stopping our
- 6:58:24positions. And not only that, but we
- 6:58:27also use uh a colon to signal how many
- 6:58:30we want to step by. So do we want to do
- 6:58:33every other in which case we would step
- 6:58:35by two. Do you want to do every third
- 6:58:37element which would step by three? So
- 6:58:41most slicing is going to follow um this
- 6:58:44sort of syntax where we do a list and
- 6:58:48then we do like a start index and then
- 6:58:52we do colon
- 6:58:54and then we do stop index
- 6:58:57um and then we do colon uh step. Now
- 6:59:02what you will see is that the the step
- 6:59:06defaults to a step size of one meaning
- 6:59:09we grab every element in between
- 6:59:11starting and stop. Um so step size of
- 6:59:15one is the default. So we actually
- 6:59:18typically will not include the step size
- 6:59:20unless we specifically want to get every
- 6:59:23other which would be a step size of two
- 6:59:25or every third or every fourth or every
- 6:59:27fifth. Um so usually we leave off this
- 6:59:31step size and we just we just have a
- 6:59:33starting and a stop as part of our
- 6:59:35slice. Um and what what that does is it
- 6:59:39tells the interpreter to access
- 6:59:43everything between this start and stop.
- 6:59:47So um with one catch which is a very
- 6:59:51important catch that trips everyone up
- 6:59:54which is that Python um is very annoying
- 6:59:57and that it uh when you do slicing it
- 7:00:01allows you to include the starting
- 7:00:04index. So uh if we start somewhere we
- 7:00:09will guarantee that the slice will
- 7:00:10include that but it will not include the
- 7:00:14stopping index. it will do everything
- 7:00:17between there up to the stop index but
- 7:00:20not actually including what's what's at
- 7:00:23the stop position.
- 7:00:26So for example,
- 7:00:28this slice that you see um on the screen
- 7:00:33is a slice that would be starting at
- 7:00:36position two
- 7:00:39because remember um let me draw this
- 7:00:42out. This is position zero. This is
- 7:00:44position one, position two, position
- 7:00:47three, position four, five, and six. And
- 7:00:51so this is a slice that would start at
- 7:00:54position two
- 7:00:56is our start.
- 7:00:59And meaning we're guaranteed to get 34
- 7:01:02because we're starting there. But the
- 7:01:04stop for this slice would have to be
- 7:01:08here. This is our stop because we are
- 7:01:12going to include everything in between
- 7:01:16there. We're going to include all of
- 7:01:19this this uh slice as part of our uh
- 7:01:23what we can access. Um so that would be
- 7:01:27positions 2, three, and four. So this
- 7:01:29slice would be um basically like this
- 7:01:33list. Um it would start at two and go to
- 7:01:37five.
- 7:01:39And then it technically would be a step
- 7:01:41size of one, but remember we don't
- 7:01:43really need to include that. So this
- 7:01:45would really be list um two to five.
- 7:01:53That does feel annoying. Yes, it's it's
- 7:01:55because they you have to know where to
- 7:01:58start and stop. And so they cho Python
- 7:02:01chooses to to be um not inclusive of the
- 7:02:06stopping index, but it it includes the
- 7:02:09starting index. Um and and that's just a
- 7:02:14choice. That's a design choice of
- 7:02:15Python.
- 7:02:19Okay. So the colon gives us the slice.
- 7:02:22Um, so if you see if you see uh uh if
- 7:02:29you see
- 7:02:31a colon inside of a brackets, that
- 7:02:35signals you're grabbing a collection of
- 7:02:37items. So we So this this actually
- 7:02:40returns a smaller list. This returns a
- 7:02:43list of 34, 20, and 80. So it's a it's a
- 7:02:49slice meaning we get multiple items
- 7:02:51rather than just a single item from from
- 7:02:53the selection.
- 7:02:57No, 54 would not make the cut. 67
- 7:02:59doesn't make the cut either because
- 7:03:01remember we don't include the stopping
- 7:03:03index.
- 7:03:08Can you do two to four plus one? Yes,
- 7:03:11you can do arithmetic in there and
- 7:03:12Python will evaluate the arithmetic
- 7:03:14first. So it will do 4 + 1 first and
- 7:03:17determine that is five.
- 7:03:20Yes, you can do that.
- 7:03:31Okay. So before
- 7:03:34before I move on, uh does the slicing
- 7:03:39idea make sense? It is grabbing a
- 7:03:42collection of items from the larger list
- 7:03:45and we set it up with this syntax.
- 7:03:52Uh what do you think? What do you think
- 7:03:530 to six would return? What would be
- 7:03:55your guess?
- 7:04:09What's included in the slice though?
- 7:04:11It's not just 67. If you go 0 to six,
- 7:04:16how many? Like, you should be getting
- 7:04:17more than that. Yeah, you should get
- 7:04:20everything, right? Exactly. You should
- 7:04:22get all of those numbers up to
- 7:04:26uh 54. You would not include 54. So you
- 7:04:29would get 76, you get 12, you get 34,
- 7:04:3220, 80, and 67.
- 7:04:41Yes, that's true.
- 7:04:49Yeah. Yeah. So the the step relevance is
- 7:04:52that we can so from our slice we can
- 7:04:57choose um our step size of how many
- 7:05:00element like how much we want to skip
- 7:05:04positions within that slice. So step of
- 7:05:06one which is a default means we get
- 7:05:09every position we go we increment by one
- 7:05:13one position to the next to the next to
- 7:05:15the next. A step of two
- 7:05:19uh
- 7:05:21a step of two would be that uh I'm going
- 7:05:25to grab every other element from the
- 7:05:27start. So a step of two like if I sliced
- 7:05:31this
- 7:05:33and changed this to a step size of two,
- 7:05:36then that would only grab this and this
- 7:05:39because that would step over. It would
- 7:05:41take two steps to get to the next
- 7:05:43element of the slice.
- 7:05:47Whereas a step size of one is going to
- 7:05:49grab everything
- 7:05:53because it's going to go one index to
- 7:05:54the next. So think of the step size as
- 7:05:57how many positions are we incrementing?
- 7:06:05Yes. 2 to 7 would include 54. Yes,
- 7:06:08that's right.
- 7:06:102 to 7 would include 54.
- 7:06:20Okay.
- 7:06:23All right. Let's see some more examples.
- 7:06:27So if we have a list like this and we
- 7:06:30slice it from 1 to 4, the output is
- 7:06:34going to start at index one
- 7:06:38and go all the way up to index 4 but not
- 7:06:42include four. So index 4 is this guy and
- 7:06:46it's not going to include that. So it's
- 7:06:48going to be these three elements here
- 7:06:52would be our slice. Yep. 20 30 40. Very
- 7:06:55good. That's what it would be.
- 7:06:58So that's pretty useful
- 7:07:01to be able to slice. Let me show you
- 7:07:03another example.
- 7:07:08Oops, don't have another. Let me go
- 7:07:10back.
- 7:07:13Let me show you another example with
- 7:07:15this. Um, so what we can do is we can
- 7:07:20actually slice backwards as well. So,
- 7:07:25what do you Let me ask you guys this.
- 7:07:26What do you think this slice would be?
- 7:07:38Actually, let me erase this. Let me do
- 7:07:40minus
- 7:07:42or
- 7:07:47what do you think this would return?
- 7:07:51We can use negative index index uh
- 7:07:53indices in our slices.
- 7:07:56What do you think that would be?
- 7:08:04Very good. 30 40 50. So it's going to
- 7:08:07go. So remember -4 is the fourth from
- 7:08:11the last. So it's going to be this is
- 7:08:13minus4
- 7:08:15and then this is minus one. So we're not
- 7:08:19going to include minus one. So we should
- 7:08:21be doing this slice here.
- 7:08:24That should be the slice. 30 40 50 would
- 7:08:26be that.
- 7:08:29Okay. One other a couple other examples
- 7:08:32I wanted to give you is that you can act
- 7:08:36in certain special cases you can
- 7:08:38actually leave off the starting and
- 7:08:42stop. And what that would signal the
- 7:08:43interpreter is that you want to go all
- 7:08:45the way to the end or start all the way
- 7:08:47from the beginning. So if you do
- 7:08:50something like this,
- 7:08:52let me show you an example. If you do
- 7:08:54something like this and you do not
- 7:08:57include, you leave the the start blank.
- 7:09:01You leave the start blank and you go all
- 7:09:03the way up to minus one. What that would
- 7:09:05signal to the interpreter is by default
- 7:09:09um start at the beginning. So start at
- 7:09:11zero. Essentially start at zero. If you
- 7:09:14leave off a slice
- 7:09:17uh as your start that the interpreter
- 7:09:19assumes you want to start at the
- 7:09:21beginning.
- 7:09:23So what do you think this slice would be
- 7:09:25knowing that?
- 7:09:32Yes, exactly. You guys you guys are
- 7:09:33right on top of it. 10 to 50. Perfect.
- 7:09:35So it's going to be everything but the
- 7:09:37last
- 7:09:39everything but the last would be
- 7:09:40included in that slice.
- 7:09:43Perfect. And so the other thing is we
- 7:09:47can leave off the end which would signal
- 7:09:49that we want to go all the way to the
- 7:09:50end. So what do you guys think this is
- 7:09:53going to be?
- 7:10:00What would that be?
- 7:10:10Yes. So this is So this is actually
- 7:10:12going to include the end. So I know
- 7:10:15that's a little counterintuitive, but
- 7:10:16it's actually this this guarantees we
- 7:10:18include the end. So if it's blank, it's
- 7:10:22going to go all the way to the end,
- 7:10:23including the end. I know that's that's
- 7:10:26annoying. I don't blame you for thinking
- 7:10:28it should be 50,
- 7:10:31but it basically goes to n. Basically
- 7:10:34goes to n, which would mean that we
- 7:10:36remember the index is n minus one is the
- 7:10:38last index.
- 7:10:40So, so if it's blank, this means we
- 7:10:43should go we should end at n
- 7:10:47which is the length of the list meaning
- 7:10:49that um the last element is at n minus
- 7:10:52one. So we should include the n minus
- 7:10:54one. So yeah that would be 20 to 60
- 7:11:00actually sorry 30 to 60 because we uh
- 7:11:03index two is 30. So that would be So
- 7:11:06that slice would be this one all the way
- 7:11:09to the end.
- 7:11:13Good. I have one more example for you
- 7:11:16that's really going to throw you for a
- 7:11:17loop is this one. So what happens if we
- 7:11:20have
- 7:11:23this case?
- 7:11:31This probably won't for loop, but I'll
- 7:11:32have one more following that.
- 7:11:37All yeah, 10 to 60 everything. Perfect.
- 7:11:41You guys are right on top of that
- 7:11:42because we're leaving the starting blank
- 7:11:44meaning that we should start at the
- 7:11:46front. We're leaving the end blank
- 7:11:48meaning we should go all the way to the
- 7:11:49end. So that would be everything
- 7:11:50everything in between. So at that point
- 7:11:53we're not really slicing anything,
- 7:11:55right? We're not really slicing much.
- 7:11:56We're just taking the whole list.
- 7:11:59Now, one example I want to give you
- 7:12:03is
- 7:12:06what if we sliced
- 7:12:13and we had a step size
- 7:12:17of minus one.
- 7:12:22Any ideas what that would do?
- 7:12:27What does this do?
- 7:12:32What is a step size of minus one?
- 7:12:36So minus negative index goes from the
- 7:12:39back, right?
- 7:12:41So if we're step sizing minus one, what
- 7:12:44should we be doing effectively?
- 7:12:53So, so this is signaling we basically
- 7:12:54want to have the whole list but step by
- 7:12:57minus one.
- 7:13:04Yeah. So this this would be the reverse.
- 7:13:13So this would be the reverse list. So, I
- 7:13:17know that seems wacky, but that actually
- 7:13:18is a way to validly reverse a list in
- 7:13:21Python is to index to step size by minus
- 7:13:24one. Because what that means is you're
- 7:13:26slicing the entire array, but you're
- 7:13:29stepping by minus one. We know minus one
- 7:13:34um we know minus one goes
- 7:13:38backwards, right? Effectively, because
- 7:13:40it starts from the end. So, step size by
- 7:13:42minus one would be go back this way.
- 7:13:45each element going back this way. So
- 7:13:48that would that would effectively
- 7:13:49reverse the list.
- 7:13:54Yeah, there now there is a reverse
- 7:13:57function. A list has a reverse function.
- 7:14:00So that is in English. But this is like
- 7:14:04this is an alternative to to reversing
- 7:14:07minus one
- 7:14:09step size of minus one.
- 7:14:17Uh, no, Roberto, they're not quite the
- 7:14:19same because remember when you do two
- 7:14:22colon and then you leave out the blank,
- 7:14:25that means you're you're going all the
- 7:14:26way to the end, including the blank,
- 7:14:28including the end.
- 7:14:30So, -4 to minus one would be the same as
- 7:14:332 to
- 7:14:35uh 2 to six or 2 to 7. Two to six.
- 7:14:40Sorry. Two to six.
- 7:14:4560 to 10. Yep. It would be it. So this
- 7:14:47reverses it. Meaning this would be this
- 7:14:50would return to 60 then 50 then 40. It's
- 7:14:54the reverse of the list when you step
- 7:14:57size by minus one.
- 7:15:18I am sure.
- 7:15:21How about slice from position one?
- 7:15:26Slice from position one to second to
- 7:15:27last and reverse it.
- 7:15:31Uh what do you think that would be?
- 7:15:34You're starting at one going to second
- 7:15:36to last.
- 7:15:39And then rever like we already know
- 7:15:41what's reversing is step size minus one
- 7:15:43that will always reverse.
- 7:15:52What is colon colon?
- 7:15:57Colon is is the fact that we're leaving
- 7:16:01colon colon is not anything special.
- 7:16:03It's the fact that we're leaving the
- 7:16:04starting. It's just the syntax of
- 7:16:06slicing, right? colon colon is because
- 7:16:09we are we're leaving the starting and
- 7:16:11stopping blank
- 7:16:13which we can do. We're allowed to do in
- 7:16:15slicing. So that would signal that we're
- 7:16:16doing everything but we're going step
- 7:16:18size minus one.
- 7:16:25Yeah.
- 7:16:30Okay.
- 7:16:33Any
- 7:16:38other any other questions?
- 7:16:43How do we feel about slicing? Do you
- 7:16:45feel okay with it? Are we going to we're
- 7:16:48going to practice it more as we go
- 7:16:50along? Uh because we're going to use
- 7:16:52slicing quite a bit when we work with
- 7:16:54data, but does the concept of slicing
- 7:16:57make sense? Yeah, you need you need more
- 7:16:59p We'll do more of it. We're going to do
- 7:17:01slicing throughout the program.
- 7:17:07Yes, we'll do multi-dimensional uh in
- 7:17:10our next course.
- 7:17:12Not right now, but in our in our data
- 7:17:14science course, we'll do
- 7:17:15multi-dimensional.
- 7:17:32that help?
- 7:17:37Yeah. Step can be a very Yeah, sure.
- 7:17:40Sure. So, there's there's nothing that's
- 7:17:42stopping you from, you know, there's
- 7:17:44nothing that's stopping you from doing
- 7:17:45like let's say x is two and then we do
- 7:17:49um numbers
- 7:17:52numbers and then we have uh two to to
- 7:17:58six and then x
- 7:18:01Yeah, that's fine. There's nothing that
- 7:18:03would stop us from doing that.
- 7:18:07I do you mean that I think that's I
- 7:18:10think that's fine. There there would be
- 7:18:12nothing wrong with that.
- 7:18:16But you're right that it should be an
- 7:18:18integer. If it's if it's like a float,
- 7:18:19Python will complain. It needs to be an
- 7:18:22integer step size and it needs to be it
- 7:18:24needs to be uh
- 7:18:26um in order to get any meaningful data.
- 7:18:28We wouldn't want that step size to be
- 7:18:30too big or like it, you know, if we pick
- 7:18:32it to be like 20 and there's only five
- 7:18:34elements, that's not going to make any
- 7:18:36sense. The step size needs to be
- 7:18:38reasonable.
- 7:18:44All right. So, what I want to do is show
- 7:18:48you some functions that lists have. So
- 7:18:52probably one of the most useful
- 7:18:54functions a list has is the ability to
- 7:18:56append items to the list. Now this will
- 7:19:00add items to the list and particularly
- 7:19:03it will add it at the end. So this is
- 7:19:06this is something we will use quite a
- 7:19:08bit is the list.append
- 7:19:11function.
- 7:19:12So this will uh this will add this will
- 7:19:17modify the list and add a new element at
- 7:19:20the end. So append always appends to the
- 7:19:23end. Um
- 7:19:26and so this will uh allow us to um take
- 7:19:31this this string cherry and now the when
- 7:19:34we append it this list is now
- 7:19:36permanently been changed to have cherry
- 7:19:39at the very end. So there it is. It is
- 7:19:42now at the back of the list and it is it
- 7:19:45is now at the kind of end position when
- 7:19:49we do append. So here you see the list
- 7:19:52being really dynamic allowing us to add
- 7:19:55elements to it through this append
- 7:19:57function.
- 7:20:02So append really really useful allows us
- 7:20:05to to add we just pass in we pass in an
- 7:20:08element inside of the append uh function
- 7:20:11here that we want to add to the list and
- 7:20:14it will it will go to the back of the
- 7:20:16list.
- 7:20:23Um, lists also have a pop function
- 7:20:28um, which you pass in a position and it
- 7:20:33will remove that item that's at that
- 7:20:35position. And not only will it
- 7:20:38permanently remove the item that's at
- 7:20:40that position, but it will return it
- 7:20:42back to you. So, pop is really useful if
- 7:20:46you want to remove things um from the
- 7:20:50from the list. um if you don't provide a
- 7:20:54position. So if you don't provide any
- 7:20:56index that you want to remove from and
- 7:20:58you just do if you if you just do um pop
- 7:21:02without any uh thing in there that will
- 7:21:05always remove the last element by
- 7:21:08default. So always just so if we just
- 7:21:11did this it would remove the 40.
- 7:21:18Can we append in a specific position?
- 7:21:21Um, yes. You would use the insert
- 7:21:24function and then give it the index you
- 7:21:26want to insert into. So,
- 7:21:29would would uh be every list has ainsert
- 7:21:34function to to and then you put in a
- 7:21:36position you want to add it to.
- 7:21:40Is it common use for append and remove
- 7:21:42during? Yes. So append is really common
- 7:21:45to add new things to the list which
- 7:21:47maybe we're doing like data aggregation.
- 7:21:49We want to add things to a list and then
- 7:21:51take the average of the list. That's
- 7:21:53very common. Um remove. Yes. Maybe we're
- 7:21:57working our way through a collection of
- 7:21:59things and when we process it we want to
- 7:22:00remove it. So we can do pop to remove it
- 7:22:03from the list.
- 7:22:05Yes.
- 7:22:12Uh yes. So when so when we pop it
- 7:22:16permanently affects the list. So um
- 7:22:20everything gets shifted. Yes. All their
- 7:22:22positions get shifted according to what
- 7:22:24we removed. So like in this example um
- 7:22:2830 is uh 30 is index um two but when we
- 7:22:34pop it now becomes uh the last element.
- 7:22:38So it would now be eligible to be index
- 7:22:40minus one, right? Cuz when we remove 40,
- 7:22:4430 is now the end of the list. Um,
- 7:22:48for example,
- 7:22:54so yeah, everything shifts
- 7:22:57and you know this this example here um
- 7:23:01pops from index two. So we would go to
- 7:23:04index two, which is 30, and remove that.
- 7:23:07And so 40 now shifts up to be at index 2
- 7:23:11whereas previously it was at index 3.
- 7:23:19Insert as well. Yep. When you insert
- 7:23:21everything shifts. Yep.
- 7:23:29Okay. Here's a here's a really useful
- 7:23:31function as well. So we have the extend
- 7:23:34function
- 7:23:35um which would allow us to add in
- 7:23:38multiple elements. So this is the same
- 7:23:40as if we appended every individual item
- 7:23:43in this collection to the list. So
- 7:23:47extend takes a list and adds its
- 7:23:51elements to the other list. So notice
- 7:23:54that we have um
- 7:23:56we have a list of colors here, red and
- 7:23:59blue, and we're extending it with a list
- 7:24:02of green and yellow, which will result
- 7:24:05in the colors list now having all of
- 7:24:08those elements. So extend is really
- 7:24:10helpful if we want to add in multiple
- 7:24:12pieces of data to an existing list.
- 7:24:21a lot of questions. Um, can we get the
- 7:24:23index and values with a print command?
- 7:24:25Uh, yeah.
- 7:24:28Yeah. I mean, you can use a I'm not sure
- 7:24:30what example you have in mind, but yes.
- 7:24:37Yes. Append is one item. Extend is is
- 7:24:39taking an entire list and and adding all
- 7:24:42of those elements to the existing list.
- 7:24:45Yes. Append is for only one value at a
- 7:24:48time. Yes. Extend is when you're adding
- 7:24:51multiple values.
- 7:24:53Append is one value at a time. Yes.
- 7:25:10Uh okay. So let me ask you guys what do
- 7:25:12you think of this? Um,
- 7:25:15which of the following method adds a
- 7:25:17single element at the end of the list?
- 7:25:23So, adding a single element at the end
- 7:25:25of the list.
- 7:25:42Very good. It should be a Yeah, we
- 7:25:44append. Append adds and append always
- 7:25:47adds to the end.
- 7:25:53Very good.
- 7:25:59Okay. So, now we're going to have a
- 7:26:02demo. Um, and by the way, this is um
- 7:26:06this is a demo that uh is an existing
- 7:26:10notebook. So, um, what you would want to
- 7:26:14do is, especially if you're working in
- 7:26:15collab, is take the notebook. Now, this
- 7:26:18is within lesson two. So, we're going to
- 7:26:20do demo one and lesson two. You would
- 7:26:22want to take that notebook if you're
- 7:26:23working in Collab and upload it. I'll
- 7:26:26show you how to do that, but we're going
- 7:26:28to do we're going to do the demo that's
- 7:26:30inside of uh the first demo inside of
- 7:26:33lesson two.
- 7:26:36Um,
- 7:26:44so let me share my screen.
- 7:26:51Okay. So, if you're inside of Collab,
- 7:26:53what you're going to want to do is go to
- 7:26:55file and then upload notebook. So,
- 7:26:58you're going to want to go to upload
- 7:26:59notebook and then um pick the hopefully
- 7:27:03you've downloaded the demos in which
- 7:27:05case you have the notebook from from
- 7:27:08lesson two. There's a bunch of notebook
- 7:27:10files, the IP YMBs. You want to upload
- 7:27:14um those demo those demo notebooks.
- 7:27:18Okay. If you're working in collab, if
- 7:27:19you're working in Jupiter, um, or you're
- 7:27:22working in, uh, VS Code, you can just
- 7:27:24open that file, uh, within VS Code or
- 7:27:27Jupiter, um,
- 7:27:30and, uh, you should be good to go from
- 7:27:32there. So, I've I've already uh, done
- 7:27:34that. This is this is demo one inside of
- 7:27:38lesson two. Do you guys have access to
- 7:27:39that notebook?
- 7:27:42Demo one and lesson two.
- 7:27:47There should be lesson two has a bunch
- 7:27:50of notebooks that we're going to work
- 7:27:51through. Um,
- 7:27:57okay.
- 7:27:58Very good.
- 7:28:01So, if we run if we run this piece of
- 7:28:04code um that's in this first set or
- 7:28:08sorry first cell, it's going to um
- 7:28:11create this list which has different mix
- 7:28:13types. So this list has integers, it has
- 7:28:16strings, it has floats, but we can
- 7:28:20create this list. If we just hit run,
- 7:28:23um, we now have a list. And what I want
- 7:28:26to show you is if we were to check the
- 7:28:29type of this my list, um, of course,
- 7:28:32this should be a list, which it is.
- 7:28:38Okay. Uh, thank you for uploading that.
- 7:28:41Perfect.
- 7:28:44Okay. So let's go ahead and access a few
- 7:28:48elements. So we can access the first
- 7:28:50element here. We can access the element
- 7:28:53the fourth which would be at position
- 7:28:55three and then the seventh which would
- 7:28:57be at position six. We can access all of
- 7:29:00those
- 7:29:01and we put those into a new list here by
- 7:29:05putting them inside of the brackets. So
- 7:29:06that that means that we're accessing
- 7:29:08this first one. That's the first element
- 7:29:11of this list that we're creating. It's
- 7:29:1325, which is here.
- 7:29:17By the way, what do you guys think
- 7:29:18happens if we try to if we try to use an
- 7:29:20index that is too big for this list?
- 7:29:25What do you think would happen? Like if
- 7:29:27we if we tried to do if we tried to use
- 7:29:30code that would be like um my list and
- 7:29:34then we put in the index like 20. What
- 7:29:37do you think would happen?
- 7:29:40because there's definitely not 20 items
- 7:29:41in this list.
- 7:29:45Yeah, it'll be an error. So, let's try
- 7:29:46running that. This will give me an error
- 7:29:49that says it's out of range. Yeah, an
- 7:29:52exception, right? It would be an
- 7:29:54exception, which would say, uh, we have
- 7:29:56an index error. Um, we're trying to use
- 7:29:58an index that's too big for our list
- 7:30:01essentially.
- 7:30:05So, just pointing that out. Um, let me
- 7:30:07make a comment there.
- 7:30:10Um, this
- 7:30:13uh index is out of range for our list.
- 7:30:21So, we should get an error.
- 7:30:25See how I'm making a comment? Making a
- 7:30:28comment there to remind myself of why I
- 7:30:30got this error. So, remember, comments
- 7:30:33are useful.
- 7:30:38All right. So, we access things and we
- 7:30:42can uh put those inside of a list. Now,
- 7:30:44let's do negative index. So, we know
- 7:30:46negative -1 should give us the item
- 7:30:48that's at the very end of the list,
- 7:30:50which would be this 2.718.
- 7:30:54Um, so that should be there. And then
- 7:30:56minus 4 would be fourth from the back.
- 7:30:58Minus 7 would be seventh from the back.
- 7:31:02So, we can uh get those values. Not too
- 7:31:04bad.
- 7:31:06Then we have a slicing example.
- 7:31:10So 2 to 7 we know as a slice. This
- 7:31:13should um this is a slice that uh slice
- 7:31:18that starts at index 2 and goes to index
- 7:31:237 but doesn't include index 7.
- 7:31:31Right? So that should be the slice. Um,
- 7:31:34so if we run this code, um, this would
- 7:31:38extract everything starting at position
- 7:31:40two, which should be the third item of
- 7:31:42the list, all the way up to, uh,
- 7:31:45position 7.
- 7:31:48And one other thing I wanted to show you
- 7:31:50is we can extract how many elements are
- 7:31:53in the list.
- 7:31:55So I wanted to show you guys that this
- 7:31:58code tells us the length of the list
- 7:32:03which would be if we did length of my
- 7:32:07list.
- 7:32:08Yeah. Len. So we pass that in the the
- 7:32:11length function. We pass in my list um
- 7:32:14which should give us 10. So there's 10
- 7:32:16items in this list.
- 7:32:25So just wanted to call out that there is
- 7:32:28this length function that we can do with
- 7:32:30the list.
- 7:32:35Can you show printing index numbers for
- 7:32:38the list?
- 7:32:43Like do you mean an uh every number in
- 7:32:45its index?
- 7:32:50Do you mean that every number and its
- 7:32:52index? Yeah. So the code that does that
- 7:32:54is the enumerate function and we we
- 7:32:57would use a loop. Um so it would be
- 7:33:00something like for index
- 7:33:04um value in enumerate
- 7:33:08uh my list and then we could do um print
- 7:33:14uh index
- 7:33:16and value
- 7:33:24like that
- 7:33:27you know we haven't learned this We
- 7:33:29haven't learned any of this yet, but
- 7:33:30that's that's what it's doing.
- 7:33:33Yeah.
- 7:33:50Okay,
- 7:33:52cool. Uh let's see. So um finally what I
- 7:33:57wanted to show is that we can append. So
- 7:34:00if we uh take our list this is what it
- 7:34:03currently is
- 7:34:05and then we append a new element we can
- 7:34:08print out the list and you can see how
- 7:34:09it ends up at the end. So we take that
- 7:34:12original list and we just add a new
- 7:34:14element at the end. Um and then this by
- 7:34:18the way we didn't we didn't uh explain
- 7:34:20this but remove will find that value
- 7:34:24find the value 100 and remove it from
- 7:34:28the list. It will find the first
- 7:34:31occurrence.
- 7:34:33Yeah. So remove will find the first
- 7:34:36occurrence of this value. Pop is index
- 7:34:39based. Remove is value based. So remove
- 7:34:44will look for the 100 and and take that
- 7:34:47out of the list. But pop will um be
- 7:34:52index based. So if we do um
- 7:34:56if we do so pop is index based. So we
- 7:35:02could uh do my list.pop
- 7:35:05pop and we could pass in a zero which
- 7:35:09should remove um this will remove the
- 7:35:13first element
- 7:35:15and then we can uh print my list.
- 7:35:21So that removed the 25.
- 7:35:30Does remove all? No, I think it's just
- 7:35:33the first occurrence.
- 7:35:35This is the first occurrence. You'd have
- 7:35:36to do it multiple times if you have
- 7:35:38duplicates.
- 7:35:42I think I have to double check that, but
- 7:35:44I think it's just the first occurrence.
- 7:35:52Okay. All right. So, I know we're a
- 7:35:54minute over.
- 7:35:56Uh, thank you guys so much. What a great
- 7:35:59first couple of sessions. Um, I think
- 7:36:01we're picking up this really well. So
- 7:36:03very good job. A lot of great questions,
- 7:36:05a lot of good um back and forth. So I
- 7:36:08appreciate that. Hope you guys are
- 7:36:09learning and picking up this Python as
- 7:36:12we go along. Um we have a lot more to
- 7:36:14cover. So uh you know next time we meet
- 7:36:18um you know we will uh continue talking
- 7:36:21about the other data structures. So we
- 7:36:22have to talk about sets, tupils,
- 7:36:24dictionaries and then we have to get
- 7:36:26into uh loops and if else statements and
- 7:36:30then we'll eventually work our way to
- 7:36:32functions. So, a lot more to cover, but
- 7:36:34we'll get there. Um, and but hopefully
- 7:36:38you guys are learning a lot. Any
- 7:36:40homework to do? Not formally, but I
- 7:36:42would request that you guys work on the
- 7:36:45guided practice for lesson one. Work on
- 7:36:48the guided practice for lesson one if
- 7:36:51you can.
- 7:36:53Okay? So, go into your reference
- 7:36:55materials, find the guided practices,
- 7:36:58work on the lesson one guided practice
- 7:37:00between now and our next session.
- 7:37:02Thank you guys. Thank you so much. Uh
- 7:37:04thank you for I know these 4hour
- 7:37:06sessions are a lot. Appreciate your
- 7:37:08patience. Thank you so much.
- 7:37:11Have a great rest of your week.
- 7:37:13>> So you might be wondering what's changed
- 7:37:15in machine learning and why is it the
- 7:37:17best time to get into it now. Well,
- 7:37:19let's go back a few years. In the past,
- 7:37:22machine learning was more about building
- 7:37:23models based on historical data. It was
- 7:37:26about training algorithms to predict
- 7:37:28specific outcomes like classifying
- 7:37:30emails, spam or not spam, predicting
- 7:37:32house prices based on past data. But
- 7:37:34fast forward to 2026 and the landscape
- 7:37:37has changed dramatically. Today, machine
- 7:37:39learning isn't just about making
- 7:37:40predictions. It's about building systems
- 7:37:42that can learn, adapt, and improve over
- 7:37:44time. We're no longer creating
- 7:37:46algorithms to just run experiments
- 7:37:48offline. Now, machine learning systems
- 7:37:50are integrated into real world
- 7:37:52operations and are capable of making
- 7:37:54decisions that impact the business
- 7:37:56immediately. Here's an example to make
- 7:37:58it clearer. In the past, an e-commerce
- 7:38:00website might use machine learning to
- 7:38:02predict what products a customer might
- 7:38:04want based on their past purchases. Now,
- 7:38:07the systems can constantly learn from
- 7:38:09new customer data, continuously refining
- 7:38:11those predictions in real time as
- 7:38:13customers preferences are changing. The
- 7:38:15world of machine learning has evolved
- 7:38:17from theory to practice and this has
- 7:38:19created a huge demand for machine
- 7:38:21learning engineers who can build
- 7:38:23scalable systems and make them work in
- 7:38:25real world environments. The impact of
- 7:38:27machine learning is now directly tied to
- 7:38:29business outcomes and machine learning
- 7:38:31engineers are at the center of that
- 7:38:32transformation. You might be thinking
- 7:38:35okay I get it machine learning is
- 7:38:36impactful but what exactly does a
- 7:38:39machine learning engineer do compared to
- 7:38:40other roles in tech? That's a great
- 7:38:42question. In the world of machine
- 7:38:44learning, you'll hear about a few key
- 7:38:46roles such as data scientist, machine
- 7:38:48learning, and AI engineer. Let's break
- 7:38:50them down so you know exactly where you
- 7:38:52fit in. Data scientists are like the
- 7:38:54detectives of data. They spend their
- 7:38:57time analyzing large data sets, finding
- 7:38:59trends, and trying to extract meaningful
- 7:39:01insights. They build models, but their
- 7:39:04main focus is usually on data
- 7:39:05exploration, and experimenting with
- 7:39:07various algorithms. They don't typically
- 7:39:09focus on deploying those models into
- 7:39:12production environments. Machine
- 7:39:13learning engineers on the other hand
- 7:39:15these are architects. They take the
- 7:39:17models built by data scientists and
- 7:39:19build scalable deployable systems. They
- 7:39:22work on creating solutions that will not
- 7:39:24only work in the short term but can also
- 7:39:26scale to handle real world data in
- 7:39:28massive volumes. The machine learning
- 7:39:30engineer is responsible for ensuring
- 7:39:32that machine learning systems are
- 7:39:34integrated into businesses that can work
- 7:39:36seamlessly with existing technologies.
- 7:39:38AI engineers focus more on the
- 7:39:40application side of things. They build
- 7:39:42AI powered products like chatbots, voice
- 7:39:45assistants, and real-time systems. While
- 7:39:47their work often overlaps with ML
- 7:39:49engineers, they are typically more
- 7:39:51focused on the userfacing product and
- 7:39:53how machine learning fits into it. As an
- 7:39:55ML engineer, your primary focus is to
- 7:39:57take models and turn them into
- 7:39:59actionable solutions that are deployed
- 7:40:01in real world systems. We shall now move
- 7:40:03on to why 2026 is the right time to
- 7:40:05enter machine learning. Now that you
- 7:40:07know the role of an ML engineer, let's
- 7:40:09talk about why 2026 is the perfect time
- 7:40:11for you to jump into the field. You've
- 7:40:14probably heard that machine learning is
- 7:40:15a hot topic, but what does that mean for
- 7:40:17you as someone starting out in this
- 7:40:19field? So, I'll help you break that down
- 7:40:21for you. First, the demand for ML
- 7:40:23engineers has skyrocketed. The world is
- 7:40:26full of problems that needs solving, and
- 7:40:28machine learning has proven to be one of
- 7:40:30the most effective tools to solve them.
- 7:40:32From predicting customer preferences to
- 7:40:34automating critical business functions,
- 7:40:36machine learning is changing how
- 7:40:38businesses operate. Secondly, the tools
- 7:40:40used to build machine learning systems
- 7:40:42are more accessible than ever before. In
- 7:40:44the past, machine learning was viewed as
- 7:40:46something experimental, something that
- 7:40:48required a lot of effort just to set up.
- 7:40:50But today machine learning platforms and
- 7:40:52frameworks such as TensorFlow, PyTorch
- 7:40:54and Scikitlearn have matured
- 7:40:56significantly. These tools make it
- 7:40:58easier to build and deploy models that
- 7:41:00can scale to handle real world data. We
- 7:41:03shall now move on to why is this the
- 7:41:04right time for you to get started out as
- 7:41:06a machine learning engineer. So how do
- 7:41:08you get started? The first step is to
- 7:41:10build a strong foundation. You might be
- 7:41:12excited to start building models and
- 7:41:14diving into algorithms. But before that
- 7:41:16you need to understand the core concepts
- 7:41:18that drive all the machine learning
- 7:41:19systems. These include mathematics,
- 7:41:22programming and data handling. You don't
- 7:41:24need to be an expert in all of these
- 7:41:26areas, but you do need to understand the
- 7:41:28basics. Think of these as building
- 7:41:29blocks of everything that you will need
- 7:41:31to learn machine learning. Let's start
- 7:41:32with mathematics. You don't need to be a
- 7:41:34math genius, but you do need to
- 7:41:36understand the basics. There are three
- 7:41:38main areas of math that will help you
- 7:41:40get started out as an ML engineer, and
- 7:41:42those are linear algebra. This is the
- 7:41:45study of vectors and matrices which are
- 7:41:47used to manipulate and process data in
- 7:41:50machine learning models. If you have
- 7:41:52heard terms like feature vectors or
- 7:41:54matrix operations, that's linear algebra
- 7:41:56play. This is the core of how machine
- 7:41:59learning algorithms operate. This is the
- 7:42:01math behind optimizing models. You'll
- 7:42:03use calculus to adjust the parameters of
- 7:42:05machine learning models and minimize the
- 7:42:08error between predicted and the actual
- 7:42:10values. Specifically, derivatives are
- 7:42:12used to find the best way to fit a model
- 7:42:14into the data. Statistics. Machine
- 7:42:16learning is all about working with
- 7:42:18uncertaintity, and statistics will help
- 7:42:20you make sense of it. Whether you're
- 7:42:22dealing with probability distributions
- 7:42:23or hypothesis testing, statistics can
- 7:42:26help you understand patterns in data and
- 7:42:28make decisions with uncertain
- 7:42:30information. These mathematical concepts
- 7:42:32will help you build a more accurate
- 7:42:34model and optimize it effectively. Let's
- 7:42:37move on to programming stack. If you're
- 7:42:39new to programming, don't worry. Python
- 7:42:41is the language that you will want to
- 7:42:43learn. It's simple to get started with
- 7:42:45and has a huge ecosystem of libraries
- 7:42:47especially designed for machine
- 7:42:48learning. The key libraries that you
- 7:42:50will need to master include NumPy. This
- 7:42:53library is used to handle large arrays
- 7:42:55of matrices and data which is the
- 7:42:57backbone of most machine learning
- 7:42:59algorithms. If you plan on working with
- 7:43:01large data sets, you'll be using NumPy a
- 7:43:03lot. Pandas. This library is great for
- 7:43:06data manipulation and analysis. You'll
- 7:43:08use pandas to clean, organize, and
- 7:43:11transform data, making it ready for
- 7:43:12machine learning models. Scikitlearn.
- 7:43:15This library provides simple, easy to
- 7:43:17use tools for building machine learning
- 7:43:19models. It covers everything from data
- 7:43:21prep-processing models like regression
- 7:43:24and classification. Once you're
- 7:43:26comfortable with these tools, you'll be
- 7:43:27able to start building machine learning
- 7:43:29models and working with real world data.
- 7:43:31Next up is SQL, which stands for
- 7:43:33structured query language. As an ML
- 7:43:36engineer, you will be working on lots of
- 7:43:38data, and SQL will help you query
- 7:43:40databases to retrieve information that
- 7:43:42you need. You'll use SQL to extract
- 7:43:44data, filter it, and join tables
- 7:43:46together, making sure that you have the
- 7:43:48right data to your models. Once you have
- 7:43:50your data, the next step is data
- 7:43:52wrangling. Data wrangling is a huge part
- 7:43:55of your job, and if you master it, it
- 7:43:56will save you a lot of time and
- 7:43:58frustration while building models. We
- 7:44:00shall now move on to types of machine
- 7:44:02learning. Let's talk about the different
- 7:44:04types of machine learning. There are
- 7:44:05three main categories which are
- 7:44:07supervised learning, unsupervised
- 7:44:09learning and reinforcement learning.
- 7:44:11Speaking of supervised learning, this is
- 7:44:14when you have label data. You train your
- 7:44:16model on data where the answers are
- 7:44:18already known. The goal is for the model
- 7:44:20to learn the relationship between inputs
- 7:44:22and outputs so it can predict the future
- 7:44:24outcomes. Unsupervised learning. In this
- 7:44:27case, the model works with unlabelled
- 7:44:29data and it tries to find hidden
- 7:44:31patterns and groupings in the data. This
- 7:44:33is useful for tasks like clustering or
- 7:44:35anomaly detection. Reinforcement
- 7:44:37learning. This type of learning involves
- 7:44:39training an agent to make decisions by
- 7:44:42interacting with its environment and
- 7:44:43receiving rewards or penalties based on
- 7:44:46its actions. This is often used in game
- 7:44:48AI or robotics. We shall now move on to
- 7:44:51feature engineering and evaluation.
- 7:44:53After cleaning your data, the next step
- 7:44:55is feature engineering. This is the
- 7:44:57process of transforming raw data into
- 7:44:59meaningful features that help your model
- 7:45:01make better predictions. Once you've
- 7:45:03engineered your features, it's time to
- 7:45:05evaluate the model. You'll use metrics
- 7:45:07like accuracy, precision, and recall to
- 7:45:10see how well your model is performing.
- 7:45:12Proper evaluation ensures that your
- 7:45:14model is ready for real world
- 7:45:15applications and that it can handle data
- 7:45:18in production. We shall now move on to
- 7:45:20the six-month learning plan in order to
- 7:45:21become an ML engineer in 2026. So, how
- 7:45:24do you actually get started? Here's your
- 7:45:26six-month learning plan to guide your
- 7:45:28journey. Firstly, focus on learning
- 7:45:30Python, mathematics, and SQL. Dive into
- 7:45:33machine learning algorithms and hands-on
- 7:45:35projects. Then, learn deep learning and
- 7:45:37work with frameworks like TensorFlow or
- 7:45:39PyTorch. By the sixth month, you can
- 7:45:42work on a capstone project that covers
- 7:45:44the full pipeline from data collection
- 7:45:46to deployment. We shall now speak about
- 7:45:48the ML engineer tool stack. So let's
- 7:45:51talk about the tools that you will need
- 7:45:52as a machine learning engineer. You must
- 7:45:54be wondering what tools should I be
- 7:45:56learning to become a successful ML
- 7:45:58engineer in 2026. So let's simplify
- 7:46:00this. Git and GitHub are the first
- 7:46:03things on your list. At first glance,
- 7:46:05version control might seem like
- 7:46:06something that only coders need to worry
- 7:46:08about. However, it can be really
- 7:46:10crucial. Why? Because version control
- 7:46:13allows you to track every change that
- 7:46:15you make to your code, collaborate with
- 7:46:17teams, and roll back changes if
- 7:46:19something goes wrong. Imagine you're
- 7:46:21working on a huge project and you mess
- 7:46:22something up. Without version control,
- 7:46:25you could lose hours of work. But with
- 7:46:27Git and GitHub, you could go back and
- 7:46:29fix to any point and time. You might be
- 7:46:32thinking, okay, I get that Git is
- 7:46:34useful, but what about the tools that
- 7:46:35can actually help me build and deploy
- 7:46:37machine learning models? Now, that's a
- 7:46:39great question. So let's talk about the
- 7:46:41cloud platforms like AWS, Google Cloud,
- 7:46:44and Azure. In the past, machine learning
- 7:46:46models were often built and tested upon
- 7:46:48local machines, but we quickly realized
- 7:46:51that it wasn't scalable. These cloud
- 7:46:53platforms give you the ability to handle
- 7:46:55large data sets, run models on
- 7:46:57highowered servers, and scale them as
- 7:46:59your projects grow. Instead of relying
- 7:47:02on your personal laptop to do all of the
- 7:47:04heavy lifting, you can leverage the
- 7:47:05cloud to train models much faster and
- 7:47:08handle massive data sets without
- 7:47:09worrying about memory limitations. We
- 7:47:12shall now move on to the next topic
- 7:47:13which is on MLOps life cycle. You have
- 7:47:16now got all of your tools in place. But
- 7:47:18how do you go from all of these tools
- 7:47:20coming together in the MLOps life cycle.
- 7:47:22So you must be wondering what does all
- 7:47:24of the MLOps life cycle tools even look
- 7:47:26like and how do I manage the process
- 7:47:29from starting to finish? Well, here's
- 7:47:31the thing. MLOps is like a welloiled
- 7:47:33machine. It involves a series of stages
- 7:47:36that ensure that your models remain
- 7:47:37reliable, scalable, and adaptable to new
- 7:47:40data over time. Think of it like a car
- 7:47:42assembly line. First, you train the car
- 7:47:45and then you deploy it into the real
- 7:47:47world. After that, you need to monitor
- 7:47:49it and see if it runs smoothly. If it
- 7:47:51breaks down, you will have to retrain
- 7:47:52it. Let's break it down into key steps.
- 7:47:55Training your model. This is where you
- 7:47:57create the model using your training
- 7:47:59data. This is a very experimental phase
- 7:48:01where you try different algorithms,
- 7:48:03tweak parameters, and optimize the model
- 7:48:05for better performance. Deploying it to
- 7:48:08production. Now that your model is
- 7:48:10ready, it's time to put it into action.
- 7:48:12This means making the model available to
- 7:48:15users or clients. Whether it's in a
- 7:48:17mobile app or a web service, deployment
- 7:48:19is where the magic happens. Monitoring
- 7:48:22performance. Once your model is live,
- 7:48:24you can't just forget about it. It's
- 7:48:26like checking the car tire pressure
- 7:48:28after it's been on the road for a while.
- 7:48:30You need to continuously track how well
- 7:48:32your model is doing. If it starts to
- 7:48:34slip or underperform, it's time for
- 7:48:36tweaks and adjustments. Lastly, we have
- 7:48:39retraining. Over time, your model may
- 7:48:41need to be retrained with new data to
- 7:48:43stay accurate. This is especially true
- 7:48:45in industries where the environment is
- 7:48:47constantly changing like e-commerce or
- 7:48:50finance. Restraining ensures that the
- 7:48:52model stays relevant and continues to
- 7:48:54provide value. You may be wondering that
- 7:48:57this sounds like a lot of work and
- 7:48:59that's true. That's where MLOps tools
- 7:49:01like MLS help you streamline this
- 7:49:03process. By automating and managing this
- 7:49:06life cycle, you can spend less time
- 7:49:08dealing with the back end and more time
- 7:49:09focusing on building innovative
- 7:49:11solutions. We shall now move on to
- 7:49:13experiment tracking. As you dive deeper
- 7:49:15into machine learning, you'll quickly
- 7:49:17realize how important it is to track
- 7:49:19your experiments. At first, this might
- 7:49:21seem a little overwhelming. You might be
- 7:49:23thinking, I can't just run a model and
- 7:49:25hope for the best, right? But trust me,
- 7:49:28tracking your experiments is one of the
- 7:49:29best habits that you can develop early
- 7:49:31on. Think of it like logging your
- 7:49:33workout progress. Even if you don't
- 7:49:35track the results, how will you know
- 7:49:37whether you're improving or not? The
- 7:49:39same goes for machine learning. By using
- 7:49:41experiment tracking tools flow and
- 7:49:43weights and biases, you can keep a
- 7:49:45detailed record of every model and every
- 7:49:47hyperparameter and every evaluation
- 7:49:50metric. Now imagine you're building a
- 7:49:52recommendation system for an online
- 7:49:54store. You try different adjustments,
- 7:49:56algorithms, and get different results.
- 7:49:59With experiment tracking, you can easily
- 7:50:01compare which configurations work best.
- 7:50:03You can easily compare which
- 7:50:04configurations work best and learn from
- 7:50:06the past mistakes. You'll always know
- 7:50:09which experiment gives you the best
- 7:50:10results and which needs tweaking.
- 7:50:12Tracking experiments also helps in
- 7:50:14collaboration. If you're working with
- 7:50:16the team, being able to see everyone's
- 7:50:18experiments in one place makes it easier
- 7:50:20to understand their approach and build
- 7:50:22on each other's work. We shall now move
- 7:50:24on to projects and portfolio. Now that
- 7:50:26you've got your tools and workflow in
- 7:50:28place, it's time to focus on building a
- 7:50:30strong portfolio. You might be thinking,
- 7:50:32how do I make my portfolio stand out to
- 7:50:34potential employees? That's a great
- 7:50:37question. The answer is simple. Real
- 7:50:39world projects. Think about it.
- 7:50:42Employers want to see what you can do in
- 7:50:44practice. They don't just want to see
- 7:50:46theoretical knowledge. They want to see
- 7:50:47that you can solve problems and build
- 7:50:49working systems that can make a
- 7:50:51difference. So, for that reason, we have
- 7:50:53mini project examples. At this point,
- 7:50:56you're probably itching to start
- 7:50:57building something yourself. Well,
- 7:51:00hands-on projects are the best way to
- 7:51:02solidify what you've already learned.
- 7:51:04For example, you could create a customer
- 7:51:06churn prediction model or a fraud
- 7:51:09detection system. These are practical
- 7:51:11real world projects that demonstrate
- 7:51:13your ability to build solutions from
- 7:51:15start to finish. Make sure you document
- 7:51:18your projects clearly on GitHub and
- 7:51:20always include a detailed explanation of
- 7:51:23your approach, challenges, and results.
- 7:51:25Remember, the goal is to show that you
- 7:51:27understand the problem and build a
- 7:51:29working solution. We'll now move on to
- 7:51:31Kaggle and open source. As you continue
- 7:51:34to build your portfolio, I highly
- 7:51:36recommend diving into Kaggle
- 7:51:37competitions and contributing to open
- 7:51:39source projects. Kaggle is an amazing
- 7:51:42platform where you can work on real
- 7:51:43world data sets and solve problems that
- 7:51:46companies and research institutions are
- 7:51:48facing. Not only will you improve your
- 7:51:50skills, but you'll also have the
- 7:51:52opportunity to see how top data
- 7:51:53scientists and machine learning
- 7:51:55engineers approach similar problems.
- 7:51:57Contributing to open-source projects is
- 7:51:59another excellent way to showcase your
- 7:52:01skills. It shows that you can work well
- 7:52:03with others understanding existing
- 7:52:05systems and contribute to the community.
- 7:52:08Plus, it's a great way to gain
- 7:52:09visibility and make connections with
- 7:52:11other engineers. We shall now speak
- 7:52:13about the résumés which work for you.
- 7:52:15Speaking about your resume, when it
- 7:52:17comes to landing a job as a machine
- 7:52:19learning engineer, your resume needs to
- 7:52:21be focused on real world projects and
- 7:52:23practical experience. So, be sure to
- 7:52:26highlight the projects that you've
- 7:52:27worked on, the tools you've used, and
- 7:52:29most importantly, the impact that your
- 7:52:31work has had. If you build a
- 7:52:33recommendation engine that boosted
- 7:52:35product sales by 20%, make sure that you
- 7:52:37include that. Employers want to see how
- 7:52:39your work contributes to solving real
- 7:52:41business problems. Don't forget to
- 7:52:43include links to Kaggle profile or any
- 7:52:45other open source contributions that you
- 7:52:47have made. Employers love seeing code
- 7:52:50and showing them your projects is the
- 7:52:51best way to stand out. We shall now
- 7:52:53speak about what interviewers test in
- 7:52:552026. By now you must be wondering what
- 7:52:58do employers actually look for in an ML
- 7:53:00engineer. So in 2026 interviews are not
- 7:53:03just about technical knowledge.
- 7:53:05Employers just want to see how well you
- 7:53:07can communicate your thought process and
- 7:53:09real world problems. You'll likely face
- 7:53:11practical tests that challenge you to
- 7:53:13build a model or analyze data in real
- 7:53:16time. You'll also be tested on how well
- 7:53:18you explain your approach and justify
- 7:53:20the decisions that you made during the
- 7:53:22project. Prepare to discuss things like
- 7:53:25why you choose a specific algorithm for
- 7:53:27a problem or how you handle issues like
- 7:53:30data imbalance or overfitting along with
- 7:53:33what metrics you use to evaluate your
- 7:53:35model's performance. Being able to
- 7:53:38communicate clearly your process and how
- 7:53:40you arrived at your solution is a skill
- 7:53:42that will set you apart from other
- 7:53:43candidates. We shall now speak about the
- 7:53:46ML trends that you will need to follow.
- 7:53:48Machine learning is evolving fast and
- 7:53:50staying up to date is key to remaining
- 7:53:52relevant in this field. So here are a
- 7:53:54few trends to keep an eye on in 2026.
- 7:53:57Generative models. These models can
- 7:54:00generate new data based on patterns they
- 7:54:02learn and they're used for things like
- 7:54:04text generation, image creation, and
- 7:54:06even music composition. AutoML automated
- 7:54:10machine learning tools are making it
- 7:54:11easier for non-experts to build machine
- 7:54:13learning models. As a result, more
- 7:54:16people will be able to contribute to
- 7:54:18this field without needing to become
- 7:54:19experts. Privacy first models. As
- 7:54:22privacy concerns grow, machine learning
- 7:54:24models are being designed to work
- 7:54:26securely and ethically and these don't
- 7:54:28compromise on user privacy. Staying on
- 7:54:31top of these trends will help you remain
- 7:54:33competitive and innovative in the ever
- 7:54:35evolving field of machine learning. So
- 7:54:38in conclusion, I will say consistency is
- 7:54:40key in machine learning. You don't need
- 7:54:42to know everything right away, but you
- 7:54:44must stay committed to learning and
- 7:54:46building. Start small, keep
- 7:54:48experimenting, and keep improving. The
- 7:54:50road to become a machine learning
- 7:54:52engineer may seem long, but with the
- 7:54:53right tools and mindset, you can get
- 7:54:55there. Thank you for watching and I'll
- 7:54:57see you in the next one. Thanks for
- 7:54:59watching the machine learning engineer
- 7:55:00road map for 2026. We hope that this
- 7:55:02video gave you a clear path to becoming
- 7:55:04a successful ML engineer. Ready to take
- 7:55:07the next step? Explore the Simply Lance
- 7:55:09professional certificate in AI and
- 7:55:10machine learning in partnership with
- 7:55:12Purdue University. Gain hands-on
- 7:55:14experience and industry recognized
- 7:55:15credentials to boost your career.
- 7:55:17>> Machine learning is which is a subset of
- 7:55:20artificial intelligence, right? That's
- 7:55:22uh basically um machines learning from
- 7:55:26data
- 7:55:28in order to uh make decisions
- 7:55:30essentially. Um, so this was a big
- 7:55:34departure from the rules-based systems
- 7:55:37at the time, right, that were explicitly
- 7:55:39programmed to make decisions. So just
- 7:55:42think of an example like a really big
- 7:55:45kind of if this, then that, then that,
- 7:55:47then that, and and else if this, this,
- 7:55:50this, right? So bunch of rules that had
- 7:55:52to be pre-programmed in order to um come
- 7:55:56out with some final answer. Uh, with
- 7:55:58machine learning, it's the exact
- 7:56:00opposite of that. we're actually
- 7:56:01training something from examples from
- 7:56:03existing data um in order to predict
- 7:56:06something or um
- 7:56:10make some type of decision. Uh and so
- 7:56:12we're going to learn about the various
- 7:56:14ways we can do machine learning. But if
- 7:56:16you guys remember we
- 7:56:18um talked about some of this like the
- 7:56:20differences and the uh basically
- 7:56:23rules-based approaches to learning from
- 7:56:26data approach. Um and in included in
- 7:56:30that is going to be uh complex
- 7:56:32unstructured data. So things like
- 7:56:34images, text, audio. What handles those
- 7:56:37really well is uh deep learning which we
- 7:56:40will get to in the course after this.
- 7:56:43But uh those are certainly in there as
- 7:56:46learning from data even complex data.
- 7:56:51So we had this picture uh and I think
- 7:56:53this is kind of around where we left off
- 7:56:55last time was uh just distinguishing
- 7:56:59between those three terms. We see
- 7:57:01artificial intelligence, deep learning
- 7:57:03and machine learning kind of used
- 7:57:04interchangeably, but this is really how
- 7:57:05they fit in. Artificial intelligence is
- 7:57:08kind of a broad anything mimicking human
- 7:57:11intelligence. Um which doesn't have to
- 7:57:15be learning from data, but uh machine
- 7:57:17learning is part of that. And then um
- 7:57:20one way to accomplish machine learning
- 7:57:22is to use neural nets which is the focus
- 7:57:24of uh deep learning. Um and so deep
- 7:57:28learning has been has found a lot of
- 7:57:29success especially recently with uh
- 7:57:32those complex data types like images,
- 7:57:35speech, text, right? So deep learning
- 7:57:38used all over the place. Even in um
- 7:57:40modern like generative AI, we see deep
- 7:57:43learning used quite a bit. Um it really
- 7:57:46anything that's using neural nets is uh
- 7:57:49going to be deep learning.
- 7:57:52Um
- 7:57:53and again we'll focus on that later but
- 7:57:56we're going to be mainly focused on
- 7:57:57machine learning for this course
- 7:57:59primarily machine learning that does not
- 7:58:01use neural networks. Okay so just models
- 7:58:04that are not necessarily neural networks
- 7:58:07be our focus.
- 7:58:11So in machine learning we had an example
- 7:58:14of a game uh essentially um learning
- 7:58:19what decisions to make uh based on the
- 7:58:23uh kind of current um state of the
- 7:58:26board. This could be a um you know
- 7:58:29machine learning example that uh learns
- 7:58:32from many previous examples. So a lot of
- 7:58:35data around these games are used to
- 7:58:38train these um kind of robots that can
- 7:58:42play these games and play them at a very
- 7:58:43high level. Um so there's been a lot of
- 7:58:46successes actually in machine learning
- 7:58:48and deep learning um around
- 7:58:51uh playing games like chess or go
- 7:58:56um using machine learning algorithms. So
- 7:58:58pretty cool.
- 7:59:01All right. So I think this is where we
- 7:59:02ended. We last time we said there's a
- 7:59:03bunch of different use cases for machine
- 7:59:05learning. So um recommendation system is
- 7:59:08going to be a big one and we will
- 7:59:10actually study that uh in one of our
- 7:59:12final lessons of this course. Um chat
- 7:59:15bots like generative AI doing sentiment
- 7:59:18analysis chat bots we'll study later but
- 7:59:20those are certainly an application of
- 7:59:22learning from data in order to uh
- 7:59:25generate responses to text prompts
- 7:59:28right. Um spam filtering that's a good
- 7:59:30example like classifying an email as
- 7:59:33spam or not spam. Um that that gets
- 7:59:37trained from examples and uh learning
- 7:59:40from data such as previous emails. Um
- 7:59:44social media posts analysis is another
- 7:59:46kind of text data um use case but you uh
- 7:59:52can do a lot with that text like you can
- 7:59:54predict the sentiment um you can predict
- 7:59:58uh the category of what what the post is
- 8:00:01talking about um those kind of things
- 8:00:04all can be done with machine learning
- 8:00:07>> and many other use cases not on this
- 8:00:09list that we will uh cover
- 8:00:11>> you know as we as as we go further.
- 8:00:22Okay, so this is where we kind of left
- 8:00:24off. Um, so what's doing all the hard
- 8:00:28work here is
- 8:00:31>> uh machine learning algorithms. So these
- 8:00:32are things that will um these are things
- 8:00:36that will learn from the data. So they
- 8:00:39are uh they they are basically um
- 8:00:44algorithms or sets of rules that uh or
- 8:00:48mathematical rules I should say not
- 8:00:49formal rules like in the in the sense of
- 8:00:51a rule system but mathematical um
- 8:00:54formulas and mathematical uh rules
- 8:00:57essentially that help us learn from the
- 8:01:01data. So they correlate the data to some
- 8:01:03type of outcome. So some type of
- 8:01:06prediction uh whether that's going to be
- 8:01:09as we will see whether that could be
- 8:01:10like a number like we're predicting a
- 8:01:13price or demand or sales
- 8:01:16um or it could be a category like is
- 8:01:18this transaction fraud or not fraud or
- 8:01:21what's the probability that this is
- 8:01:23fraud um so we have different kinds of
- 8:01:27predictions we can make with machine
- 8:01:28learning.
- 8:01:30Um
- 8:01:31but uh we will study the kind of the
- 8:01:34differences of those coming up. Um but
- 8:01:37machine learning algorithms are really
- 8:01:39what power they're kind of the models,
- 8:01:41right? They're the models that help
- 8:01:42power uh machine learning to actually
- 8:01:46learn from data.
- 8:01:48So we're going to spend a lot of time in
- 8:01:50this course studying those algorithms
- 8:01:53like the different models that we can
- 8:01:54build and what their differences are,
- 8:01:56what their strengths are, what their
- 8:01:58weaknesses are. we'll we'll learn a lot
- 8:02:00about those.
- 8:02:04Okay. So, I guess you can imagine like
- 8:02:06everything is so data dependent, right?
- 8:02:09Um we're learning from data. So, uh it
- 8:02:12makes sense that the quality of data
- 8:02:14really really matters here in
- 8:02:16determining how strong the model can be.
- 8:02:19Um so you see this graph here charting
- 8:02:22kind of the um high quality data um
- 8:02:27versus just uh any old data but a decent
- 8:02:31enough quantity of it. Um you can see
- 8:02:34that performance and the performance is
- 8:02:36measured by some evaluation metric. Um,
- 8:02:40so think of it as uh something like an
- 8:02:42accuracy. Like if we were predicting
- 8:02:44fraud or not fraud, how accurate can our
- 8:02:47model get at actually detecting fraud
- 8:02:50um it gets better and better and better
- 8:02:53the graph shows that the higher quality
- 8:02:56of data that we have. So there's kind of
- 8:02:59that there's a there's a saying in
- 8:03:00machine learning um called garbage in
- 8:03:03garbage out. What that means is if you
- 8:03:06have poor data, even the best model in
- 8:03:08the world, poor data is not going to
- 8:03:11result in having a good model that can
- 8:03:13be accurate and perform well. Um, so it
- 8:03:16needs to be high quality, meaning um
- 8:03:20there needs to be a decent amount of it
- 8:03:21and it needs to be labeled appropriately
- 8:03:24as we will will talk about
- 8:03:27um and it needs to not have any, you
- 8:03:30know, significant outliers. it needs to
- 8:03:32be clean, not have those missing values,
- 8:03:35all of those things. Um, you can you
- 8:03:38have a good chance at deriving good
- 8:03:40predictions from higher quality data
- 8:03:44as this kind of shows.
- 8:03:49Okay.
- 8:03:52So, one thing we're going to learn um as
- 8:03:54we go along is
- 8:03:56quantity matters as well. So, not only
- 8:03:58quality, but a decent amount of it. And
- 8:04:01um we're going to learn those kind of
- 8:04:02rules of thumb like how much data do I
- 8:04:04need for certain algorithms. Um one
- 8:04:08thing that we will see is that uh the
- 8:04:11the basic machine learning models that
- 8:04:13we'll study don't need as much as a
- 8:04:16neural network would. It you know neural
- 8:04:19networks are going to require a lot more
- 8:04:22um than a basic machine learning model
- 8:04:25learning model. So uh that's something
- 8:04:28we will see as we go along. But uh this
- 8:04:30is something we'll talk about and
- 8:04:32discuss with each model that we study is
- 8:04:34kind of how much data do we actually
- 8:04:36need to produce a high quality model.
- 8:04:44Okay, any questions uh so far?
- 8:04:54Okay, let's talk about the different
- 8:04:57types of machine learning that we're
- 8:04:59going to discuss. Prim, there's going to
- 8:05:01be two primary ones that we will study
- 8:05:03in this course and then a couple others
- 8:05:06that'll be a little bit more advanced
- 8:05:07that we won't get to but worth knowing
- 8:05:10about. Um, so there's going to be four
- 8:05:11total that we'll study or talk about and
- 8:05:14they'll be on this list here, which is
- 8:05:18um supervised learning and unsupervised
- 8:05:20learning. Now, I'd say the majority of
- 8:05:23our focus will probably be on supervised
- 8:05:26learning, and we'll talk about what that
- 8:05:28means, but we'll also cover unsupervised
- 8:05:32learning as well. And so, we'll look at
- 8:05:34the most popular techniques in each of
- 8:05:36these types of machine learning.
- 8:05:39Um,
- 8:05:41and then we'll talk about these two, but
- 8:05:43not really study them because they're
- 8:05:45more advanced topics. Um, that that will
- 8:05:48be beyond the scope of what we'll do.
- 8:05:50But uh these are going to be um
- 8:05:54different styles of machine learning
- 8:05:58that are going to be characterized by um
- 8:06:02what kinds of predictions they make,
- 8:06:03what kind of data they need and require.
- 8:06:06Um and uh what kind of outcomes they're
- 8:06:10actually producing. Um, so let's let's
- 8:06:14get into each of these, but uh the the
- 8:06:17one that we'll probably spend the
- 8:06:18majority of our time on is going to be
- 8:06:20supervised learning, but we will study
- 8:06:23unsupervised learning as well. We'll
- 8:06:25study both and we're going to talk about
- 8:06:26we're going to define both of those um
- 8:06:28coming up. And again, these will be a
- 8:06:31little bit more advanced topics that we
- 8:06:32won't spend too much time on.
- 8:06:35Um, but but we'll discuss their
- 8:06:37relevancy in machine learning. Um, and
- 8:06:41give a good definition to it.
- 8:06:48Okay.
- 8:06:51All right. Let's start with supervised
- 8:06:54learning. Now this is going to be uh a
- 8:06:57term that really refers to
- 8:07:01using examples. So using labeled
- 8:07:06examples. So here we say labeled data to
- 8:07:10help our model train. In other words,
- 8:07:14help our model be able to predict guided
- 8:07:18by specific input output pairs. So
- 8:07:21supervised really refers to the fact
- 8:07:23that we have answers. We have examples.
- 8:07:29We have answers with those and we use
- 8:07:32that collection of data to build our
- 8:07:35model off of so that we can predict
- 8:07:39um those kinds of things like a price,
- 8:07:42like a category, like a spam, not spam.
- 8:07:47in this in this slide like we would be
- 8:07:49predicting if this shape is a square, a
- 8:07:52triangle or a circle.
- 8:07:54Um but but when we build a model for
- 8:07:58that, we have data that has an answer
- 8:08:02attached to it, right? We've talked
- 8:08:04about this before a little bit with
- 8:08:05labels. So there's a guide there that
- 8:08:10can guide us towards building our model.
- 8:08:12there's an actual every every example
- 8:08:15has an answer and that answer is really
- 8:08:18critical to help build our model off of.
- 8:08:21So, um that's it's almost like you have
- 8:08:26um a you have a bunch of exercises
- 8:08:31in let's say like a math textbook. You
- 8:08:33have a bunch of exercises and you have
- 8:08:35the answers and that way you can kind of
- 8:08:38check your work. you think about model
- 8:08:40training um that is the really a lot of
- 8:08:44that process of model training as we are
- 8:08:46going to discover is um basically
- 8:08:49checking our work against these answers
- 8:08:51in our data in our training data.
- 8:08:55Okay. So supervised learning is any type
- 8:08:58of machine learning that involves
- 8:09:01learning from labeled data in order to
- 8:09:03predict outcomes. Okay. Predict outcomes
- 8:09:06like now the the outcomes can be
- 8:09:09numerical. They can be like a price,
- 8:09:12temperature, demand, sales, revenue.
- 8:09:15They can be numerical, but they can also
- 8:09:17be categorical. So they can be like spam
- 8:09:20not spam, fraud, not fraud, cancer, not
- 8:09:22cancer. Um dog, cat, giraffe, those kind
- 8:09:27of categories. Um we could predict
- 8:09:30those. It's some type of outcome. Okay,
- 8:09:32some type of outcome. The key is we're
- 8:09:34using labeled examples to guide our
- 8:09:38model building. That's why it's called
- 8:09:40supervised learning.
- 8:09:42So we know in our data we know what the
- 8:09:46inputs are. Of course, those are going
- 8:09:48to be think of the inputs as like all of
- 8:09:49our columns and then we have a special
- 8:09:53label column that represents the output
- 8:09:55we're trying to predict. So if you think
- 8:09:57about that housing price data, the label
- 8:10:00could be the price. And that's something
- 8:10:02we would build a model to predict, but
- 8:10:04we have answers for all of our examples
- 8:10:07in our rows. We have answers to help
- 8:10:10guide our model building.
- 8:10:12They help tweak our model because we
- 8:10:14know the answer ahead of time. So
- 8:10:17they're they're really good examples to
- 8:10:19build our model off of.
- 8:10:23Okay. So that's that's supervised
- 8:10:25learning.
- 8:10:32Uh in this example, is circle not in the
- 8:10:34prediction because it's not part of the
- 8:10:36test data even though it's in the
- 8:10:38labeled data?
- 8:10:40Um no, it just not necessarily. It just
- 8:10:43means that like we learn against all of
- 8:10:47these examples that have these answers
- 8:10:50and then when we observe new examples um
- 8:10:53we can try to predict what those would
- 8:10:55be based on what we've seen before. So I
- 8:10:58if there was a you know it's just a
- 8:11:00coincidence we only have two two
- 8:11:02examples in our test data like we could
- 8:11:04have a circle here in which case we
- 8:11:06would predict circle
- 8:11:08that's fine or at least we would hope
- 8:11:10our model would predict circle right
- 8:11:12that's what we're hoping may or may not
- 8:11:14get it right
- 8:11:16um but it's it's only not there because
- 8:11:21we only like we're just assuming that we
- 8:11:23only have two examples we're testing
- 8:11:24against but in reality we would probably
- 8:11:26do a lot more than two.
- 8:11:29It's just it's just a coincidence
- 8:11:30really.
- 8:11:37In reality, we would test against a lot
- 8:11:39more data. And we're actually going to
- 8:11:41see why we would do that. Like why would
- 8:11:44we train our model and then kind of use
- 8:11:48additional data um to to evaluate it?
- 8:11:52It's actually really important that we
- 8:11:53do that step to get a sense of how good
- 8:11:56our model is before we take it out in
- 8:11:58the real world. So if we apply our model
- 8:12:01that we build on our label data to
- 8:12:05um this kind of set of test data that we
- 8:12:09haven't been exposed to before. It helps
- 8:12:12give us a sense of how good is our
- 8:12:14model. So it's test is usually used for
- 8:12:17evaluation.
- 8:12:20So that's something that's something
- 8:12:21we'll study.
- 8:12:24How do we train? Uh it depends on the
- 8:12:26model. Um so training will be a sense uh
- 8:12:31will be an algorithm that will um
- 8:12:34basically update the model according to
- 8:12:36the data. These labeled examples. Um
- 8:12:39every model is going to be different in
- 8:12:40exactly how it trains. So we're going to
- 8:12:44we're going to talk about that when we
- 8:12:45get to the individual models that we'll
- 8:12:47study.
- 8:12:49But uh loosely speaking, they're going
- 8:12:51to use the data to adjust itself. Like
- 8:12:54imagine adjust like tuning a bunch of
- 8:12:57knobs. Um, like the best example I can
- 8:13:00give you is we I think I did this one
- 8:13:03last week where you have kind of a
- 8:13:04function
- 8:13:06that predicts the price and let's say it
- 8:13:09has
- 8:13:11um weights like weight one with feature
- 8:13:14one, weight two with feature two,
- 8:13:18weight three with feature three. So
- 8:13:20imagine we had three input features and
- 8:13:22we we built an answer according to that.
- 8:13:25Essentially what we would do to train
- 8:13:27the model is adjust these
- 8:13:32um in order to get this correct based on
- 8:13:35our our labeled examples.
- 8:13:39Okay.
- 8:13:41So that's something we're going to learn
- 8:13:42about coming up shortly when we when we
- 8:13:44actually dive into model. Every model is
- 8:13:45going to be slightly different in how it
- 8:13:47trains, but at a high level it's going
- 8:13:49to use the training data with those
- 8:13:53examples, right? the labeled examples to
- 8:13:55help guide the formula essentially to
- 8:13:58adjust to generate the proper kind of
- 8:14:02model here. The these things are going
- 8:14:05to be adjusted according to the data
- 8:14:08in order to produce the correct output.
- 8:14:12So think about these as knobs that will
- 8:14:14turn.
- 8:14:19Okay.
- 8:14:22Uh which type of machine learning is
- 8:14:23used? Uh probably supervised um which is
- 8:14:26what we're talking about now. So
- 8:14:28probably supervised because most people
- 8:14:29want to
- 8:14:31um build some type of model to predict
- 8:14:34something. Uh so yeah, I'd say I'd say
- 8:14:38supervise.
- 8:14:44Yes, we're are we are definitely going
- 8:14:46to learn how to train. Yeah, we'll see.
- 8:14:48We'll do the code. Um, I'll tell you
- 8:14:51about how it's done. Yeah, we're
- 8:14:53definitely going to learn it. But what I
- 8:14:54was saying is it's kind of on a model by
- 8:14:56model basis.
- 8:14:58So, I want to wait till we get into the
- 8:15:00individual models, then we'll talk about
- 8:15:01how they're trained.
- 8:15:04But yeah, we'll we'll learn how to do
- 8:15:05that.
- 8:15:12But yeah, supervisor is used all over
- 8:15:13the place. Even even for uh generative
- 8:15:17models, they use supervised learning
- 8:15:19because um like an LLM
- 8:15:23is going to use labeled examples in
- 8:15:26order to train, right? In order to train
- 8:15:29how to generate responses according to
- 8:15:32prompts. Um it needs to learn against a
- 8:15:35lot of text examples.
- 8:15:38So that supervised learning is what um
- 8:15:41results in that model,
- 8:15:44right? Learning from those labeled
- 8:15:45examples.
- 8:15:56Okay.
- 8:16:02It is yeah image image uh a lot of um
- 8:16:07yeah a lot of image processing is
- 8:16:08supervised like object detection. So the
- 8:16:11yolo model is an object detection model.
- 8:16:14Yes. Um because it has to be trained
- 8:16:17right it has to be trained on uh it has
- 8:16:20to be trained on images
- 8:16:24with labels such as this is what object
- 8:16:26is in this image. This is the box around
- 8:16:29the object.
- 8:16:31Um yes. So if if it's if it ever uses
- 8:16:35label data to train and build the model,
- 8:16:38it is supervised. So YOLO is definitely
- 8:16:41supervised and we actually we will we
- 8:16:44will cover the YOLO model later on in in
- 8:16:46our deep learning course. We talk about
- 8:16:49object detection.
- 8:16:51So we'll we'll study that.
- 8:16:54But yeah, it's supervised
- 8:17:04Okay. So on the slide we have some
- 8:17:08common supervised learning algorithms
- 8:17:10that are we will study. So all of these
- 8:17:13we will study and understand what they
- 8:17:15do and how they work but just giving you
- 8:17:18some to name them. linear regression is
- 8:17:20kind of the one I just drew out which is
- 8:17:22the um this is the prototypical like
- 8:17:26easiest to understand model that is kind
- 8:17:28of the um exactly like this where we
- 8:17:31have a weight times a feature um a
- 8:17:34weight times a feature and then a weight
- 8:17:37times a feature
- 8:17:39and on and on and on. You can have as
- 8:17:41many as you want.
- 8:17:43um that is a linear regression. And so
- 8:17:46that is um that's a supervised model
- 8:17:49because we need this value here and we
- 8:17:52need all of our inputs in order to um
- 8:17:56actually train this model and generate
- 8:17:58all those weights
- 8:18:00um that that is uh that uses um labeled
- 8:18:05examples to help tune all those knobs.
- 8:18:07Um same with all these other models. So,
- 8:18:09we're going to talk about decision
- 8:18:10trees. We're going to talk about
- 8:18:11logistic regression and and SVMs, which
- 8:18:13are support vector machines. We'll talk
- 8:18:15about all of those, but they're all
- 8:18:17examples of supervised uh supervised
- 8:18:19learning.
- 8:18:21Okay, we'll talk about all of these.
- 8:18:25They're all supervised because they all
- 8:18:28require labeled examples in order to
- 8:18:30train them and and then subsequently use
- 8:18:33them. Okay.
- 8:18:45Okay. So what are some use case
- 8:18:46examples? So for for instance in uh
- 8:18:49supervised learning we may be predicting
- 8:18:50temperature based on yearly temperature
- 8:18:53trends. So we would have that yearly
- 8:18:56data as our um as our labeled examples
- 8:18:59and those would supervise the learning
- 8:19:02of a model that predicts temperature.
- 8:19:04Um, same thing with predicting crop
- 8:19:06yield based on um, seasonal crop quality
- 8:19:10changes. So maybe we have a bunch of
- 8:19:11features relating to crop quality. We
- 8:19:14could predict crop yield. Um, we would
- 8:19:18just need historical examples with those
- 8:19:21labels, right? What the crop yield is
- 8:19:23for each time period. Let's say we would
- 8:19:27just need those uh, supervised examples
- 8:19:29and we could easily build a model off of
- 8:19:32it.
- 8:19:33Um
- 8:19:35uh this this last one sorting waste
- 8:19:37based on known waste items and their
- 8:19:39corresponding waste types. Um that's
- 8:19:42kind of like spam. It's like filtering
- 8:19:44basically like a spam filtering. Um so
- 8:19:47think of it like the the shapes example.
- 8:19:49We sorting things into squares, circles,
- 8:19:52triangles. Um, same kind of idea here
- 8:19:54where we have a bunch of examples on
- 8:19:56what those um what those waste items
- 8:20:00should uh should belong to, like what
- 8:20:03waste bins they would go to, for
- 8:20:04example. Um, and those could be labeled
- 8:20:09and therefore then we could um
- 8:20:13understand what category of waste they
- 8:20:15belong to.
- 8:20:17Um, same thing with spam. something is
- 8:20:19fraud or not fraud, spam or not spam,
- 8:20:22cancer or not cancer. All of those are
- 8:20:23going to be supervised learning examples
- 8:20:25because they're going to require in
- 8:20:27order to train them, they're going to
- 8:20:29require data that has those labels.
- 8:20:32Okay? So, anything that has labels is
- 8:20:35going to be supervised learning.
- 8:20:39So, again, this is where we will spend
- 8:20:42probably the the majority of our time is
- 8:20:46doing supervised learning problems. ones
- 8:20:48that we have labeled data. We're
- 8:20:51building a model and we're going to
- 8:20:52predict those those uh labels
- 8:20:55essentially.
- 8:21:04Okay, before we go to unsupervised, any
- 8:21:07questions about uh supervised
- 8:21:24Okay.
- 8:21:27All right. So, supervised requires
- 8:21:29labels
- 8:21:31in order to have an example to go off of
- 8:21:34to build your model. And that's because
- 8:21:36you're predicting those kind of outcomes
- 8:21:39like spam or not spam, cancer not
- 8:21:41cancer. Now unsupervised learning is
- 8:21:45completely different. It's the opposite.
- 8:21:48So unsupervised learning is where we do
- 8:21:52not use labels whatsoever. So we're not
- 8:21:55using any labels at all. So it's it it
- 8:21:58can be completely unlabeled or even if
- 8:22:00it's labeled, we're not using labels in
- 8:22:02any way. But um we primarily would say
- 8:22:05it's unlabeled data. We have no guidance
- 8:22:08because we're not using the labels in
- 8:22:10any way. we have no guidance to um
- 8:22:14predict anything but that's because
- 8:22:16we're not really predicting anything in
- 8:22:17unsupervised learning. Generally what
- 8:22:19we're doing is looking for some
- 8:22:21structure or pattern.
- 8:22:23Okay, with unsupervised learning we're
- 8:22:25looking for some structure or pattern.
- 8:22:27So um one type of example that's very
- 8:22:32very popular is going to be this second
- 8:22:34one which is um identification
- 8:22:37identification of user groups based on
- 8:22:40similarities or commonalities. Now this
- 8:22:42is going to be a problem basically known
- 8:22:45as clustering
- 8:22:49and it's a problem we will study quite a
- 8:22:51bit. There's going to turn out to be
- 8:22:53lots of different algorithms that can
- 8:22:55accomplish clustering. So what
- 8:22:57clustering attempts to do is basically
- 8:22:59say um we have data that's like this and
- 8:23:02then data over here and then data over
- 8:23:06here. Let's just group these together.
- 8:23:08So like this should be one group. This
- 8:23:10should be one group and this should be
- 8:23:11one group. And we can find those
- 8:23:14structures and say okay this is group
- 8:23:17one this is group two and this is group
- 8:23:19three.
- 8:23:22One 2 3. And we can basically build what
- 8:23:26we would call clusters of data um based
- 8:23:29on how close together the points are
- 8:23:32kind of located in these kind of cluster
- 8:23:34zones like these boxes I've drawn.
- 8:23:38Okay. Now that doesn't require any label
- 8:23:40to do which is really fascinating. So
- 8:23:42unsupervised you don't need any label at
- 8:23:44all to accomplish the algorithm. Um so
- 8:23:47clustering is one good example. Um
- 8:23:50finding outliers or anomalies is
- 8:23:52another. So we don't necessarily have
- 8:23:54any label of what is an outlier or what
- 8:23:57is an anomaly. We are deriving that from
- 8:24:00the features alone. There's no guidance.
- 8:24:02There's no label um to doing like
- 8:24:05outlier detection or anomaly detection.
- 8:24:08Okay. So that's another good example.
- 8:24:10One that's not listed on here um but is
- 8:24:15also really important that we will study
- 8:24:16is something known as dimensionality
- 8:24:19reduction.
- 8:24:21So dim reduction and what that what this
- 8:24:25focuses on is basically compressing the
- 8:24:28data set a bit. So we take our data and
- 8:24:31basically compress it. Um so that but we
- 8:24:34do it in such a way that we retain as
- 8:24:37much information as we can. It's a very
- 8:24:40like smart compression and what it does
- 8:24:43is it lowers the dimension.
- 8:24:46um dimension. Think of the dimension as
- 8:24:48like number of columns.
- 8:24:52Number of columns.
- 8:24:54So imagine we had 100 columns in a data
- 8:24:57frame. What we could do is actually
- 8:24:58reduce that down to 10. So like 10% of
- 8:25:02that. So we reduce it down to 10. And um
- 8:25:06but those 10 are it's not like we
- 8:25:09chopped out um 90 other columns. we um
- 8:25:14smartly kind of compressed all that
- 8:25:15information into these 10 new columns um
- 8:25:19that are compressed versions of the
- 8:25:21hundred that we used to have. Um so
- 8:25:24dimensionality reduction is is another
- 8:25:26unsupervised technique. It requires no
- 8:25:28guidance, no label to do, but is um a
- 8:25:33really useful technique to reduce the
- 8:25:35size of your data if you're doing things
- 8:25:37with it. Um so this is another one that
- 8:25:40we will we'll study how to do it and
- 8:25:43basically more details behind it what
- 8:25:44the algorithms are.
- 8:25:46Um we'll so probably those two in
- 8:25:49unsupervised will spend the most amount
- 8:25:51of time on clustering and dimensionality
- 8:25:53reduction.
- 8:26:02Uh and supervised if some data is
- 8:26:03present but we didn't label it means in
- 8:26:06example we had circle triangle square in
- 8:26:09the training data we add pentagon
- 8:26:18but we didn't label that in that case.
- 8:26:23Uh yeah. So every um in supervised
- 8:26:27learning, every row, think about it as
- 8:26:30like every row in our data frame needs
- 8:26:32to have a label
- 8:26:34uh associated to it. It needs to have a
- 8:26:36a column that represents the label.
- 8:26:41So if we've never seen Pentagon before,
- 8:26:43I can't use that as a label.
- 8:26:48So it has to the pentagon has to exist
- 8:26:52in the data. if I'm going to be able to
- 8:26:54predict it,
- 8:27:01right? So, it can't predict, right? If
- 8:27:03we've never seen it before, we have no
- 8:27:05examples to go off. We have no guidance.
- 8:27:07So, how could we predict that?
- 8:27:09Right? We can't predict it
- 8:27:25if it's if it's in there. So if if we
- 8:27:28have labels of Pentagon, let's say, then
- 8:27:31yeah, we could predict Pentagon.
- 8:27:33We could
- 8:27:49remove. Remove what?
- 8:27:56We wouldn't if it was talking about the
- 8:27:57Pentagon, we wouldn't remove that. No,
- 8:27:59let me go back to that page. We wouldn't
- 8:28:02remove it. Um, it's just if it's not in
- 8:28:06our labels, we're not going to be able
- 8:28:07to predict it. So, Pentagon's a good
- 8:28:10example here. Uh, Pentagon is not one of
- 8:28:15our labels. So, it currently is not in
- 8:28:18our data set as one of the labels. We
- 8:28:20only have data that's either a triangle,
- 8:28:22circle, or
- 8:28:24square. We don't have pentagon. So, I
- 8:28:27would never be able to predict pentagon.
- 8:28:30I'll never be able to do that if I
- 8:28:31haven't seen examples of it before.
- 8:28:35Okay. But let's say we had that in
- 8:28:37there.
- 8:28:39So, we had Pentagon.
- 8:28:44So if we had Pentagon, um we could have
- 8:28:47an example of it in our labels
- 8:28:56and then Yeah, we it could be then we
- 8:28:58could predict it.
- 8:29:06Yeah. Yeah. The the don't get worried
- 8:29:08don't worry about the test data. So the
- 8:29:10test data is just saying here's a new
- 8:29:13here's a shape what is it okay that's a
- 8:29:15square here's a shape what is it okay
- 8:29:17that's a triangle and we could have as
- 8:29:19many of those examples as we want in our
- 8:29:21test data so we could have a circle and
- 8:29:24say okay what's this should be circle
- 8:29:29right the test data can be whatever it
- 8:29:31whatever it wants but yeah if if we've
- 8:29:33never seen pentagon before we're never
- 8:29:34going to be able to predict
- 8:29:45These are the the label data and labels
- 8:29:48are basically the talking about the same
- 8:29:51thing. The labels just mean what are the
- 8:29:54categories
- 8:29:56that are present in our data. So in this
- 8:30:00data we only have three labels that are
- 8:30:01present.
- 8:30:08So the labels is are relative to our
- 8:30:11label data, right? It's saying
- 8:30:14what labels,
- 8:30:18excuse me, what labels uh do we have
- 8:30:23in our data and we only have those three
- 8:30:25circle, triangle, square. So so Pentagon
- 8:30:28would not be part of those labels. We
- 8:30:30couldn't predict it.
- 8:30:44No. So unsupervised is not going to make
- 8:30:47a prediction. That's the big difference
- 8:30:49with unsupervised. They're not going to
- 8:30:51make a prediction like this. Um so
- 8:30:54unsupervised is not going to make a
- 8:30:56prediction. It's going to do something
- 8:30:57different like um basically say like
- 8:31:00these guys are similar, these are
- 8:31:02similar, these are similar, this is a
- 8:31:04cluster, this is a cluster, this is a
- 8:31:06cluster. It's not going to make a
- 8:31:09prediction. That's what supervised
- 8:31:12learning does.
- 8:31:16Clustering, yes, which is unsupervised.
- 8:31:19Yes, clustering does not require any
- 8:31:21labels. Unsupervised just means we don't
- 8:31:23have any labels. We don't require any
- 8:31:25labels.
- 8:31:43So the other thing unsupervised might do
- 8:31:46is it might say
- 8:31:49and again without the labels it might
- 8:31:51say that this is an outlier.
- 8:31:56it might say that this guy is an outlier
- 8:31:58because there's only there's only one of
- 8:32:00those and they're not like the other. So
- 8:32:02that that's something that um that's
- 8:32:05something that uh unsupervised could do.
- 8:32:14Um it it yeah and no. It kind of labels
- 8:32:19a cluster in the sense that um it would
- 8:32:23basically assign a number to it like
- 8:32:25this is cluster one, this is cluster
- 8:32:28two, this is cluster three.
- 8:32:33It'll assign a number to it, but it's
- 8:32:35not a very meaningful it doesn't assign
- 8:32:37like a prediction label in in the
- 8:32:40traditional sense of a label.
- 8:32:42It does provide like a numerical index
- 8:32:44for the cluster to because what we want
- 8:32:46to know is like okay this guy has the
- 8:32:49cluster of one. This guy belongs to
- 8:32:51cluster one. This guy belongs to cluster
- 8:32:53one. This guy belongs to cluster two.
- 8:32:55This guy belongs to cluster two. Does
- 8:32:57that make sense? So there needs to be
- 8:32:58some like index of what cluster you
- 8:33:00belong to.
- 8:33:03So it's kind of like a label but not in
- 8:33:06the traditional like prediction sense.
- 8:33:24Okay.
- 8:33:35Very good. So again, unsupervised, no
- 8:33:38labels. You're doing things like
- 8:33:42identifying clusters,
- 8:33:44um identifying outliers, doing
- 8:33:47dimensionality reduction. These are all
- 8:33:49like structure and pattern oriented
- 8:33:52things. They're not predictions of a
- 8:33:54label. Okay? They're not which is what
- 8:33:57we would see in supervised learning.
- 8:34:06Okay. So an example would be that we
- 8:34:09take we put in the data um we can group
- 8:34:13together uh data such as images into
- 8:34:17categories based on similarities um
- 8:34:20which would be like those clusters. So
- 8:34:21there's no these would be groups that we
- 8:34:24don't have any label on ahead of time
- 8:34:26like we don't have we don't say that
- 8:34:27this image should belong to this this
- 8:34:29image should belong to this we derive
- 8:34:32that from the characteristics of the
- 8:34:34data. Um so think like a good example is
- 8:34:38um customer groups. So we would identify
- 8:34:41customers based on like okay do they
- 8:34:44have similar spending levels? How many
- 8:34:46days do they go shopping in a week? How
- 8:34:49much money do they spend? And we can
- 8:34:51kind of group together customers based
- 8:34:53on similar qualities.
- 8:34:56Clustering will find those groups that
- 8:34:58should exist.
- 8:35:00um it will discover those groups based
- 8:35:03on um the similarities in the data, but
- 8:35:07there's no labels that that say like
- 8:35:09this person should be in this group,
- 8:35:12this person should be in this ahead of
- 8:35:14time. There's no labels of that. It gets
- 8:35:16derived during the algorithm. It's
- 8:35:18unsupervised,
- 8:35:22right? There's no unsupervised really
- 8:35:24literally means no guidance. There's no
- 8:35:27guidance to doing it. We just derive
- 8:35:29that from the structure of the data
- 8:35:31which is the similarities.
- 8:35:47Okay.
- 8:35:53All right. So,
- 8:35:56a couple more for you. So we had um
- 8:35:59supervised which uses the labels. We
- 8:36:04have unsupervised which uses no labels
- 8:36:07looking for structure. And then we have
- 8:36:09something that's kind of in between
- 8:36:12which is um what is known as
- 8:36:16semiupervised learning. And this is
- 8:36:19where you use a combination of a little
- 8:36:23bit of label data, but most of your data
- 8:36:25is actually unlabeled data. Um, and you
- 8:36:29try to get some use out of that label
- 8:36:32data in order to um build a model out of
- 8:36:37it. And so, uh, it uses the, um, it uses
- 8:36:42that label data to, um, generally
- 8:36:46provide some guidance on usually what
- 8:36:49happens with semi-supervised learning is
- 8:36:51you use your label data to kind of
- 8:36:54predict what the label should be for the
- 8:36:57unlabelled data and then you can go from
- 8:36:59there. So you can create artificial
- 8:37:01labels on this unlabeled data and then
- 8:37:05you can use all of it once it's all been
- 8:37:07labeled kind of like a supervised
- 8:37:09learning uh approach. So but but this is
- 8:37:12semi-supervised basically refers to the
- 8:37:14fact that you start out with most of
- 8:37:17your data not being labeled but you do
- 8:37:20have some labeled examples and what you
- 8:37:23can do is basically extrapolate those
- 8:37:25labels into the unlabeled data set and
- 8:37:28then provide some artificial labels and
- 8:37:32then now everything has a label you can
- 8:37:34do supervised learning.
- 8:37:36Okay. So, it falls kind of between um
- 8:37:40supervised and and unsupervised.
- 8:37:43Uh and there So, this is this is kind of
- 8:37:46rare. Most of the time you're not going
- 8:37:48to do that. You're actually just going
- 8:37:50to um prefer to just start with all
- 8:37:53label data. That's usually the preferred
- 8:37:55approach. Most of the time you'll
- 8:37:57actually just be doing supervised
- 8:37:58learning, not really semi-supervised
- 8:38:01learning. So, it's pretty rare, but um
- 8:38:09it it could like if Yeah, it could if
- 8:38:12the if we had a lot of examples of
- 8:38:14Pentagon and we wanted and so they were
- 8:38:16unlabeled and then we tried to guess
- 8:38:18what kind of shape they were um and
- 8:38:21provide an artificial label uh and then
- 8:38:25um then use that whole data set to build
- 8:38:27a model off of then then yeah, it could
- 8:38:29it could fall into this category. Okay.
- 8:38:38They Oh, going back to the question,
- 8:38:40they still use some kind of label data
- 8:38:41like age, gender. They use uh that's
- 8:38:44those aren't those aren't really labels.
- 8:38:46That's the features. So, yeah, they
- 8:38:48still use the core features of the data.
- 8:38:52They just don't have any like labels in
- 8:38:54the traditional sense of a label. Like
- 8:38:56you should think of a label as something
- 8:38:58we are trying to predict.
- 8:39:01So whether that's a price, whether
- 8:39:03that's like a category like spam, not
- 8:39:05spam, cancer, not cancer, it's something
- 8:39:07we'd be interested in kind of
- 8:39:09predicting. And so um in our data, we
- 8:39:12would have an answer for every row. We'd
- 8:39:15have one of our columns would be like
- 8:39:16the the result like the outcome answer
- 8:39:19that we're trying to predict. That's the
- 8:39:21label.
- 8:39:23So in unsupervised, we don't have any of
- 8:39:25the labels.
- 8:39:27We do have just the regular features
- 8:39:29like gender, age, income, square
- 8:39:33footage,
- 8:39:37bedrooms, bathrooms, all those things.
- 8:39:44Okay.
- 8:39:48So, we have semi-supervised that falls
- 8:39:50in between supervised. Now, the reason
- 8:39:52it falls between is be is because
- 8:39:54there's a decent amount of data that's
- 8:39:57unlabeled. In fact, a majority of it
- 8:39:59unlabeled. But what we can do is try to
- 8:40:03label it. We can try to take what we
- 8:40:05know from our existing labels and
- 8:40:08predict an artificial label and then use
- 8:40:11all that data together in kind of a
- 8:40:14supervised fashion for a model down the
- 8:40:16road.
- 8:40:22So that's kind of what this picture uh
- 8:40:24says is we can try to take um you know
- 8:40:28maybe we try to infer some labels based
- 8:40:30on we have some some labelled data here.
- 8:40:33We have most of our data is unlabeled
- 8:40:36and we try to supply some labels to it.
- 8:40:40Um like maybe we have a baby's category
- 8:40:42of teens, a tween, uh you know youth and
- 8:40:47um adults. Um and then we try so we we
- 8:40:51take our our labels and we try to
- 8:40:55extrapolate those into artificial labels
- 8:40:57for this unlabelled data so that we can
- 8:40:59use it now because then everything has a
- 8:41:02label at this point and then we can just
- 8:41:04go ahead and do supervised learning from
- 8:41:06there.
- 8:41:11So we can do supervised from there. What
- 8:41:13we would prefer to do and what we'll do
- 8:41:15in this course
- 8:41:17um is just start with supervised. We'll
- 8:41:20just start with the labels. We won't try
- 8:41:22to derive artificial labels usually.
- 8:41:24We'll just start with labels.
- 8:41:35So one example in the real world is
- 8:41:38something like Google photos which um
- 8:41:42whenever you take a picture it can
- 8:41:44provide uh uh labels based on previous
- 8:41:48uh images in your library. So it can it
- 8:41:52can produce tags or um labels on those.
- 8:41:56Uh generally when you take that picture
- 8:41:58it's kind of unlabeled unless you go in
- 8:42:01and specifically provide some tags and
- 8:42:03some labels. But um if you don't do that
- 8:42:06it can still it can still uh make it can
- 8:42:11artificially create one of those based
- 8:42:12on the other label data that you already
- 8:42:15have.
- 8:42:17So that's um
- 8:42:20that's an example.
- 8:42:28Okay.
- 8:42:32All right. Last one in terms of machine
- 8:42:34learning. So we have supervised, we have
- 8:42:38unsupervised.
- 8:42:39Uh then we had semi-supervised which is
- 8:42:42somewhere in between a mixture of having
- 8:42:43some unlabelled data and label data. Um
- 8:42:46now we're going to talk about
- 8:42:47reinforcement learning which is
- 8:42:49completely different. Um it's it's
- 8:42:52completely different than the other
- 8:42:53three. It's a type of machine learning
- 8:42:55where we uh basically learn from
- 8:42:58interaction with the environment. And
- 8:43:01you might ask what are we learning? We
- 8:43:04are learning what actions to take in the
- 8:43:08environment. Um and the way we do that
- 8:43:11is by reinforcing
- 8:43:13positive actions that lead to a a
- 8:43:16reward. Um, so that's where the word
- 8:43:20reinforcement comes from is we we
- 8:43:22basically uh imagine like a child
- 8:43:25that's, you know, learning from trial
- 8:43:27and error. Like they're trying to crawl,
- 8:43:28they're trying to walk and they keep
- 8:43:30falling down. um eventually they learn
- 8:43:33how to do it through trial and error and
- 8:43:35they might get a reward
- 8:43:38or they might um reinforce some of those
- 8:43:41positive movements that lead them to
- 8:43:43walk or crawl um or they might learn
- 8:43:48from the penalties, right? They might
- 8:43:49learn from uh some type of feedback. So
- 8:43:53they might learn from falling down like,
- 8:43:55"Oh, that hurts. I should uh support
- 8:43:57myself a little bit better, right?" Or
- 8:43:58be a little more coordinated. Um
- 8:44:02and so they they learn from those
- 8:44:04actions and their interaction with the
- 8:44:06environment. Um
- 8:44:09uh so this is a complex um algorithm
- 8:44:15essentially uh it's it deals a lot with
- 8:44:19um again taking actions. Usually when
- 8:44:22you take an action something changes in
- 8:44:24the environment um then you kind of
- 8:44:28observe some type of feedback. So, think
- 8:44:30about like a a board game where you're
- 8:44:33trying to figure out what move you
- 8:44:35should make. Or another good example is
- 8:44:37like with a robot um trying to navigate
- 8:44:40a maze. So, like what route should it
- 8:44:42take? Should it move forward? Should it
- 8:44:44move backward? Should it move left or
- 8:44:45right? Those are different actions it
- 8:44:47can take. Also, like a self-driving car,
- 8:44:50should it should it turn? Should it
- 8:44:52speed up? Should it slow down? Those are
- 8:44:54all good examples of things that have
- 8:44:56been trained from reinforcement
- 8:44:58learning.
- 8:45:05Uh yeah. So real world examples would be
- 8:45:08like in a board game, uh a a reward
- 8:45:10would be like if you win the game. Um or
- 8:45:14if you like capture a piece like in
- 8:45:17checkers or chess, that's a reward. A
- 8:45:20penalty would be like if you lose the
- 8:45:21game or lose one of your pieces, that
- 8:45:23could be a a penalty.
- 8:45:26um in a board game or sorry in like a a
- 8:45:31robot navigation task, it could get
- 8:45:33rewards for um moving in the right
- 8:45:36direction
- 8:45:38um towards the exit or like when it like
- 8:45:41let's say you wanted to train a robot on
- 8:45:42how to open the door and navigate a
- 8:45:45room. Um you would penalize it for
- 8:45:47bumping into the wall.
- 8:45:49Um you would give it a reward for moving
- 8:45:57usually oh like oh the algorithm
- 8:45:59themselves usually it's like a a step
- 8:46:02function um it's usually it's like a
- 8:46:04discrete function that kind of is based
- 8:46:07on the state so the reward it could be
- 8:46:10like um like depending on the let's
- 8:46:13let's go back to the board game example
- 8:46:15like the reward could be like or even
- 8:46:18the maze let's say like a navigating the
- 8:46:20maze like getting to this let's say this
- 8:46:22was the exit
- 8:46:25and this was the entrance.
- 8:46:29Then if they make it to here, they get a
- 8:46:31numerical like if they make it to the
- 8:46:33exit, they get a numerical reward of
- 8:46:34like plus 100, let's say. So it's just a
- 8:46:37number. And then if they uh like if they
- 8:46:41bump if they go into here, like let's
- 8:46:43say this is kind of like a death trap or
- 8:46:45like a pit, this this would be like a
- 8:46:47minus 100. So it could be like discrete
- 8:46:50numerical values could be the reward. If
- 8:46:53they're moving in the right direction
- 8:46:54like let's say we want to encourage
- 8:46:56going this way then we could give
- 8:46:57smaller intermediate rewards like this
- 8:46:59should be a plus like if you move
- 8:47:01forward this is a plus five this is a
- 8:47:04plus 10 this is a plus 15 if you're
- 8:47:07moving in the wrong direction away from
- 8:47:09the exit. Um that would be like a minus5
- 8:47:12or a minus 10. Does that make sense? So
- 8:47:15they're they're numerical in nature and
- 8:47:17what you're trying to do is collect the
- 8:47:19most reward. You're trying to get the
- 8:47:21largest reward you can through trial and
- 8:47:25error. So you you try this out many many
- 8:47:27many times. You basically simulate
- 8:47:30running through this maze many many
- 8:47:32times. And what dictates it what
- 8:47:36dictates like where I should go is based
- 8:47:39on what I've observed in the past. It's
- 8:47:41almost like you're a child remembering
- 8:47:42like, okay, what move should I make from
- 8:47:44this space? Like, if I'm here, if I'm
- 8:47:47here, which way should I go? Should I go
- 8:47:49down? Should I go right? Should I go
- 8:47:51left? You kind of know that from
- 8:47:53experience.
- 8:47:55Does that make sense? Based on the
- 8:47:56reward that I've seen in the past, like
- 8:47:58when I've moved down, I've gotten a
- 8:48:00higher reward than moving left or right.
- 8:48:04Does that make sense? So, yeah, it's
- 8:48:05it's a numerical value
- 8:48:08as a reward.
- 8:48:17Yeah, that's a great question. Um, how
- 8:48:20does it differentiate rewards based on
- 8:48:21gain and loss i.e. chess? So it's it's a
- 8:48:25very comp complicated uh answer but
- 8:48:27essentially every so in the chess board
- 8:48:32you can think of the board as like every
- 8:48:34every um
- 8:48:37space is a state.
- 8:48:41So I could be in this state I could be
- 8:48:43in this state and then it's not not only
- 8:48:46is every every uh space but where all
- 8:48:48the other pieces are. So there's lots of
- 8:48:50states that are possible.
- 8:48:53Um, so
- 8:48:55the way there's a way to quantify
- 8:48:59essentially what's the value of taking a
- 8:49:03certain action like moving my piece
- 8:49:04left, moving it right, moving it up or
- 8:49:07down um given the rest of the state. So
- 8:49:11you're you're right, it may be
- 8:49:12beneficial to sacrifice. Um, but we
- 8:49:16would learn that through experience that
- 8:49:18okay, the best move in this situation is
- 8:49:20to sacrifice.
- 8:49:22We would we would have to learn that
- 8:49:24through trial and error many many many
- 8:49:25times which is to say like okay if I'm
- 8:49:28in this current state of the world right
- 8:49:31all these pieces are distributed in this
- 8:49:33way the best move for me right now in
- 8:49:36the long run to get the most reward in
- 8:49:40the long run is to actually sacrifice my
- 8:49:42piece and move it right move it into
- 8:49:44like a bad position theoretically but we
- 8:49:47know from experience that's actually the
- 8:49:49most long-term reward is from that
- 8:49:51position
- 8:49:52like moving it right may be the best
- 8:49:54action for me. So what you learn is how
- 8:49:58to take actions
- 8:50:00and actions are usually like move right,
- 8:50:03move left, move up, move down. You think
- 8:50:05about like a self-driving car though,
- 8:50:07that's going to be like slow down, speed
- 8:50:09up, turn your wheel 10°,
- 8:50:13um those kind of actions.
- 8:50:19So the the short answer is it's there's
- 8:50:22a calculation there that you learn what
- 8:50:26the long-term value of every state is
- 8:50:30every unique state
- 8:50:32and then you're trying to basically say
- 8:50:35what action should I take from that
- 8:50:37state given that current state of the
- 8:50:40world.
- 8:50:49Okay.
- 8:50:51And I really I really like reinforcement
- 8:50:53learning. It's actually probably my
- 8:50:55favorite field of machine learning.
- 8:50:57Unfortunately, we won't be covering it
- 8:51:00um in our main uh course. We have
- 8:51:03offered uh electives around
- 8:51:05reinforcement learning in the past. So,
- 8:51:07um stay tuned. Maybe when we get to the
- 8:51:09end of this program, uh we'll offer an
- 8:51:11elective on it and if enough people sign
- 8:51:13up for it, we'll we'll run it. But, um
- 8:51:17we it's not part of our we don't really
- 8:51:19cover reinforcement learning as part of
- 8:51:20our main topics. It's it is an advanced
- 8:51:23uh more advanced topic than than what
- 8:51:25we'll cover. But, um I I really enjoy
- 8:51:28it. Find it very fascinating.
- 8:51:37Okay. So, all of this is kind of um
- 8:51:40illustrating what I was saying, which is
- 8:51:42um you think of like uh the thing that's
- 8:51:45interacting in the environment like the
- 8:51:46robot or the car or the human moving a
- 8:51:50chest piece is known as the agent. It's
- 8:51:54interacting with the environment by
- 8:51:55taking actions which updates the state
- 8:51:58um of of the environment. So that's
- 8:52:01that's why you see this word state here.
- 8:52:03This gets updated constantly every time
- 8:52:05you take an action. Um ultimately what
- 8:52:07reinforcement learning is trying to do
- 8:52:09is learn the best action like what would
- 8:52:12be the best action to take. Um
- 8:52:16and the best action is is the one that
- 8:52:18leads to the most long-term reward.
- 8:52:21That's the best action. Um, so you have
- 8:52:24to uh you have to learn what you know
- 8:52:29what leads to a good reward by kind of
- 8:52:31experiencing this over and over and over
- 8:52:33through trial and error. So there's a
- 8:52:36lot of um kind of simulation or letting
- 8:52:38the robot try something a lot um in
- 8:52:42order to kind of learn what's rewarding
- 8:52:44and what's not. Think about it again
- 8:52:46like I think a good example is like with
- 8:52:47children, right? you kind of have to let
- 8:52:49them try things until they learn on
- 8:52:52their own what's what can they do and
- 8:52:54what can they not do
- 8:52:56what's the best actions right
- 8:53:00so reinforcement learning has made its
- 8:53:03way into other places so I I said like a
- 8:53:05good example is self-driving cars or ro
- 8:53:08robotics a lot of reinforce
- 8:53:10reinforcement learning is used there one
- 8:53:11place it's found its way into recently
- 8:53:14is recommendation systems have kind of
- 8:53:17merged with reinforcement learning
- 8:53:18learning. Um, and this is because you
- 8:53:23you can imagine there's kind of a
- 8:53:25built-in reward for you clicking on a
- 8:53:28video and kind of watching it.
- 8:53:31Um, so that kind of reinforces that
- 8:53:33recommendation and then uh that's where
- 8:53:37um you can then kind of recommend a
- 8:53:40similar thing and see if that's
- 8:53:42rewarding and generates a click or
- 8:53:45generates some view time or watch time
- 8:53:47or whatever. Um so reinforcement
- 8:53:50learning has found its way into a lot of
- 8:53:52areas. Um recommendations being one of
- 8:53:55them because it's just natural for the
- 8:53:57idea of like what um should I recommend
- 8:54:00next to generate the most reward. In
- 8:54:03this case the reward is kind of
- 8:54:04correlated to did they click on it or
- 8:54:06not or did they how long did they watch
- 8:54:09for longer it's more rewarding.
- 8:54:12um those kind of things but uh place
- 8:54:15places where reinforcement learning have
- 8:54:17been used I said self-driving cars um
- 8:54:21games so uh one of the most famous
- 8:54:24examples if you want to look it up is
- 8:54:26the um Alph Go this was in 2016 um the
- 8:54:31Alph Go uh algorithm was a reinforcement
- 8:54:34learning bot that beat um some of the
- 8:54:38world's best Go players which go if
- 8:54:41you're not familiar Go is a um board
- 8:54:44game
- 8:54:45that is a little bit more uh complex
- 8:54:48than chess. It has more more uh it's a
- 8:54:52larger board um more pieces to it. Um
- 8:54:57but they there was a reinforcement
- 8:54:58learning powered bot that actually um
- 8:55:01learned how to play the game so
- 8:55:03effective it could beat um world kind of
- 8:55:06masters at the games was pretty amazing.
- 8:55:08Um that's the alpha go and that was by
- 8:55:11deep mind Google and deep mind in 2016.
- 8:55:15That was pretty that was only in 10
- 8:55:17years ago not that long.
- 8:55:19Um so certain uh we said recommendation
- 8:55:24uh even autocorrect um learning to
- 8:55:27predict like what is the best correction
- 8:55:30uh to generate a reward which would be
- 8:55:32like you accept that correction or you
- 8:55:34reject it would be a penalty. Um so
- 8:55:37reinforced learning has been adapted to
- 8:55:40these kind of problems very
- 8:55:41successfully. Let's take a look at the
- 8:55:44packages that we will use throughout. So
- 8:55:47um of course we will rely on these three
- 8:55:50which we've already relied on to do a
- 8:55:53lot of things like numpy to do numerical
- 8:55:56manipulations and calculations.
- 8:55:59Uh mapplot lib to do any plotting and
- 8:56:02not only mapp but maybe seabour as well.
- 8:56:05both of those to do plotting. Um, pandas
- 8:56:08is a big one because
- 8:56:11that's where all of our data is going to
- 8:56:12be manipulated and prepped before it
- 8:56:15goes into modeling.
- 8:56:17So, all of that stuff we learned from
- 8:56:19pandis is definitely going to be applied
- 8:56:21here in this course uh as we actually
- 8:56:24build models. Um, so of course like
- 8:56:27these old ones that we've been working
- 8:56:29with quite a bit um still going to be
- 8:56:31useful here in the modeling stage. Um,
- 8:56:35mainly for different reasons though,
- 8:56:37mostly to get our data prepared to do
- 8:56:40some type of modeling or maybe to
- 8:56:41visualize it before we do modeling to
- 8:56:43get a sense of what it looks like, those
- 8:56:46kind of things.
- 8:56:48Um,
- 8:56:50sci is sometimes useful for certain uh
- 8:56:55um processing like in unsupervised
- 8:56:58learning. We'll actually use scyp a
- 8:56:59little bit to do dimensionality
- 8:57:01reduction or help us do that. Um so
- 8:57:04scypi will be used here and there and
- 8:57:07we've seen it before with hypothesis
- 8:57:08testing we use scypi like the test and z
- 8:57:11test came from there. Um some of the
- 8:57:14unsupervised learning stuff will come
- 8:57:16out of there but the package we will use
- 8:57:19by far the most in this course is going
- 8:57:22to be scikitlearn
- 8:57:25which is here. Um and we've already seen
- 8:57:29a little bit about scikitlearn in terms
- 8:57:31of its pre-processing capability. So we
- 8:57:35use the uh minmax scaler and the
- 8:57:38standard scaler from there from the
- 8:57:40pre-processing module in scikitlearn.
- 8:57:43But it has um many different models
- 8:57:47built into it that we can use to help uh
- 8:57:50do our training and predictions. Um so
- 8:57:54it's a incredibly useful machine
- 8:57:56learning library. It is the industry
- 8:57:58standard machine learning library. Um if
- 8:58:01you're going to do anything in machine
- 8:58:03learning, it would be expected that you
- 8:58:05know how to use scikitlearn.
- 8:58:08Now what's really lucky about that is
- 8:58:10that scikitlearn is a really easy
- 8:58:13package to get used to. Nearly
- 8:58:15everything we do in scikitlearn will
- 8:58:17mostly follow the same pattern and so um
- 8:58:20the code will be extremely simple. They
- 8:58:22did a great job with that package of
- 8:58:24making things really user friendly,
- 8:58:26really simple. Um, it's a really
- 8:58:29fantastic package and we're going to get
- 8:58:30a lot of practice with it uh as we go
- 8:58:33along. Every model we build will
- 8:58:34essentially be from scikitlearn
- 8:58:37and not only like the models but um
- 8:58:40doing the training, doing the
- 8:58:41predictions and then doing the
- 8:58:43evaluation will all come from different
- 8:58:45uh scikitlearn u modules. So that'll be
- 8:58:49really nice and we'll get um good
- 8:58:52exposure to that package throughout the
- 8:58:53course. So if anything will come away
- 8:58:56from this course as um scikitlearn uh uh
- 8:59:01experts that'll be very nice. So this is
- 8:59:04this will be the new one for us learn
- 8:59:06but we'll get a lot of practice with it.
- 8:59:12Okay.
- 8:59:16All right. So just to recap that lesson
- 8:59:18before we move on to lesson three. Um we
- 8:59:21talked about machine learning as
- 8:59:22learning from data um which is included
- 8:59:25underneath the AI umbrella but deep
- 8:59:28learning is also included under machine
- 8:59:30learning because it's still learning
- 8:59:31from data but it's learning using neural
- 8:59:34networks.
- 8:59:35Um we talked about the four different
- 8:59:37types of machine learning. We had
- 8:59:38supervised, unsupervised,
- 8:59:41semi-supervised and reinforcement. So
- 8:59:44those are the the different types of
- 8:59:45machine learning that are out there. Um
- 8:59:48and then we talked about some of the pi
- 8:59:49Python packages uh that we will use. The
- 8:59:53main one being scikitlearn and of course
- 8:59:55we'll use our older like pandas to
- 8:59:57manipulate our data and get uh pass it
- 8:59:59into our model training etc. But
- 9:00:02scikitlearn will be uh our goto for
- 9:00:06anything machine learning.
- 9:00:10All right. So, I have some questions for
- 9:00:11you guys, some checks.
- 9:00:14So, let me know in the chat. What do you
- 9:00:16guys think? Uh, which of the following
- 9:00:19best describes machine learning?
- 9:00:26Which choice do you think makes the best
- 9:00:29is the best for this?
- 9:01:12Very good. Very good. I see I see a lot
- 9:01:14of choices for A and A would be the
- 9:01:16correct choice. So machine learning is
- 9:01:19definitely um a a subset of AI. It's
- 9:01:23underneath that AI umbrella, but of
- 9:01:25course we're learning from experience
- 9:01:27and of course that experience is
- 9:01:28recorded in the data um without being
- 9:01:32explicitly programmed. Uh so it's the
- 9:01:34exact opposite of BNC. We're definitely
- 9:01:36not learning from rules and it's
- 9:01:38definitely not just used for image and
- 9:01:40speech speech recognition. It can be
- 9:01:42used for many other things beyond those.
- 9:01:46So yeah, A is the best choice there.
- 9:01:49What we say here?
- 9:01:53Okay. What do you guys think about this?
- 9:01:55Which example illustrates the use of
- 9:01:57machine learning to enhance customer
- 9:01:58experience in an ecommerce company?
- 9:02:13In other words, what would be some what
- 9:02:14would be some uh typical use cases of
- 9:02:17machine learning?
- 9:02:45Good. So I think uh C is going to be the
- 9:02:48best answer here. Definitely C. So it's
- 9:02:51using machine learning to do uh fraud
- 9:02:54transactions. So so that would be a
- 9:02:56prediction probably a supervised
- 9:02:58learning, right? If if this is fraud or
- 9:03:00not fraud. Um, and then maybe some
- 9:03:03customer behavior uh that might be
- 9:03:06unsupervised. So maybe grouping together
- 9:03:08customers uh clustering them based on
- 9:03:10their data like their shopping behavior
- 9:03:13and characteristics. Um that that might
- 9:03:16be unsupervised but either way it's
- 9:03:18machine learning.
- 9:03:21Okay.
- 9:03:26Okay. Final one. What distinguishes deep
- 9:03:28learning from machine learning in
- 9:03:30artificial intelligence? So what's
- 9:03:32unique about deep learning?
- 9:04:03Oh, very good. Yep. So, deep learning
- 9:04:05uses neural networks as so you guys are
- 9:04:09right on top of that. Neural deep
- 9:04:10learning uses neural nets. That's what
- 9:04:12makes it unique. So, machine learning
- 9:04:14would be part A. Machine learning is
- 9:04:17focused on learning from data.
- 9:04:18Underneath of that is learning from data
- 9:04:20using neural networks which is what uh
- 9:04:23deep learning is.
- 9:04:27Very good.
- 9:04:30All right. Let's go to lesson three.
- 9:04:34And lesson 3 has two notebooks. We're
- 9:04:36going to be starting with 3.1.
- 9:04:40So, you'll want to open up that
- 9:04:41notebook. I'm going to go over to it
- 9:04:43now. Give you a moment to open that up.
- 9:04:52So, we're going to open the 3.1
- 9:04:54notebook. Um, there's two of them. We'll
- 9:04:57see how far if we can get into the
- 9:04:59second one today. probably will.
- 9:05:02Um, but we're going to do the uh we're
- 9:05:04going to start with 3.1 notebook. Do you
- 9:05:06guys have this notebook? Should be in
- 9:05:08your materials for for this course.
- 9:05:13Let me give you a moment to open that
- 9:05:14one.
- 9:05:30Do you guys have it?
- 9:05:44Okay.
- 9:05:46All right. So, we're going to start by
- 9:05:50talking about uh supervised learning.
- 9:05:54um in our machine learning journey. So
- 9:05:56remember we're going to talk about uh
- 9:05:58supervised and unsupervised after we do
- 9:06:00supervised
- 9:06:02um and there's going to be a lot to
- 9:06:03cover with supervised mainly because um
- 9:06:07there are uh two different types of
- 9:06:09problems we can tackle uh which will be
- 9:06:13uh we'll talk about in a moment
- 9:06:14predicting different kinds of values. Um
- 9:06:17but let's talk about the kind of what
- 9:06:18we're hoping to learn here which is um
- 9:06:21talk about the different kinds of
- 9:06:22problems that we'll study which are
- 9:06:24these these categories of supervised
- 9:06:26learning. Um those two categories are
- 9:06:28going to be called classification and
- 9:06:29regression. We'll talk about those and
- 9:06:31their differences and then talk about
- 9:06:33some applications and some uh example
- 9:06:36algorithms
- 9:06:38and that's just within this notebook. Um
- 9:06:403.2 two we'll get into uh regression in
- 9:06:44particular um which will be uh very very
- 9:06:48interesting. Okay. So that'll be our
- 9:06:50first models that we'll build will be
- 9:06:51over there in 3.2.
- 9:06:56Okay. So if you guys remember um
- 9:06:58supervised learning is where we learn
- 9:07:00from labeled data. So we have input and
- 9:07:03outputs in our in our data set. Um and
- 9:07:07so you train a model on this data that
- 9:07:10includes input features and
- 9:07:13corresponding outputs that are that are
- 9:07:16the labels, right? So um the goal is to
- 9:07:21learn a relationship between the input
- 9:07:24and the output. Of course, that's what
- 9:07:25any model is trying to do. Um, and what
- 9:07:28this allows us to do is then take that
- 9:07:32model and use it to make predictions on
- 9:07:34never-beforeseen
- 9:07:36uh data. Right? So then we have a
- 9:07:39predictive model out of that that we can
- 9:07:41use um going forward on new examples. Um
- 9:07:46so
- 9:07:47remember we will have in our data a
- 9:07:50bunch of features which are columns and
- 9:07:52then generally one of those columns will
- 9:07:54be the label that we're trying to
- 9:07:56predict.
- 9:07:58And our model is going to try to learn
- 9:08:00some type of relationship between those
- 9:08:02inputs and the output label. So the
- 9:08:05output label could be like fraud not
- 9:08:07fraud, cancer, not cancer, uh a price, a
- 9:08:11temperature, those kind of things.
- 9:08:15So let's talk about that. inside of um
- 9:08:18supervised learning there are two
- 9:08:19different types of learning that we can
- 9:08:23do and they're really based on the label
- 9:08:26or sometimes that label is known as the
- 9:08:29target that we're trying to predict. Um
- 9:08:32and depending on that type we get these
- 9:08:35two different categories of learning or
- 9:08:36two different types of learning. One is
- 9:08:39known as regression. So that's generally
- 9:08:42when we are predicting something that is
- 9:08:44continuous or something that is a
- 9:08:47numerical.
- 9:08:49So numerical
- 9:08:52numerical value. So think of price,
- 9:08:55think of temperature, think of revenue.
- 9:08:57We're trying to predict something like
- 9:08:59that. Um versus something that is
- 9:09:02categorical. So that the predicting
- 9:09:05something categorical would be like
- 9:09:06fraud, not fraud, spam, not spam. um
- 9:09:09those are discrete categories and the
- 9:09:13problem of predicting categories is is
- 9:09:15known as classification because we're
- 9:09:19trying to classify examples as belonging
- 9:09:22to one category or another.
- 9:09:26So we have these two main types of
- 9:09:29supervised learning problems. we have
- 9:09:31regression and we have classification
- 9:09:33and they're going to be handled slightly
- 9:09:36differently um for many reasons that
- 9:09:39we're going to uncover. Um one of the
- 9:09:42primary reasons is that of course we're
- 9:09:45predicting something that's continuous
- 9:09:47in the regression case versus something
- 9:09:48discreet. So the models have to be
- 9:09:50slightly different to account for that.
- 9:09:53Um but then a step beyond that is the
- 9:09:57evaluation has to be different too. Um I
- 9:10:00kind of alluded to this last week, but
- 9:10:02when you're predicting a regression,
- 9:10:03it's very very difficult to to get the
- 9:10:05exact numerical answer. So um generally
- 9:10:11we don't care about that. Um generally
- 9:10:15we don't care about getting it exactly
- 9:10:18uh we don't care about getting it
- 9:10:20exactly right.
- 9:10:22um we just care about getting it um
- 9:10:25we're just we care about getting it
- 9:10:26nearby, getting it close enough. Um
- 9:10:30whereas classification, we do care about
- 9:10:33getting it exactly right because it's a
- 9:10:34discrete category. So we're going to be
- 9:10:36able to evaluate that a little bit
- 9:10:38differently to say did we get the answer
- 9:10:40right or wrong. Regression is going to
- 9:10:42be did we get close? Um because it's we
- 9:10:45assume it's going to be nearly
- 9:10:46impossible to predict a continuous
- 9:10:49number. Um, that's very hard to do.
- 9:10:54Okay.
- 9:10:56So, any questions on
- 9:10:58uh that?
- 9:11:04Any questions on those two differences?
- 9:11:06Let me give you some examples. Maybe
- 9:11:07it'll it'll help too.
- 9:11:11So, again, the classification is going
- 9:11:12to be predicting uh something that's
- 9:11:14categorical. regression is going to be
- 9:11:17predicting something that is continuous.
- 9:11:24So think about trying to predict the
- 9:11:26price of a house based on those other
- 9:11:27features we talked about before like
- 9:11:29square footage, bedrooms, bathrooms, all
- 9:11:32of those things we predict the price.
- 9:11:34That would be a regression problem
- 9:11:35because the price is a continuous value.
- 9:11:39Let's take a look at an example here.
- 9:11:42Um, imagine we were trying to uh predict
- 9:11:46the temperature tomorrow. That's going
- 9:11:49to be a regression problem, a a
- 9:11:51supervised learning kind of regression
- 9:11:53problem because we're trying to predict
- 9:11:55a numerical temperature.
- 9:11:59Okay? And versus a category like a
- 9:12:03discrete category would be this would be
- 9:12:05a classification. So this is a
- 9:12:07regression on the left. This is a
- 9:12:09classification
- 9:12:11on the right. Classification
- 9:12:16um because we are um predicting one of
- 9:12:21two categories. Is it just hot or cold?
- 9:12:23Now, we're not saying exactly where that
- 9:12:25threshold is on what's hot or cold. That
- 9:12:27would be a decision on on what we want
- 9:12:29to what our discrete categories actually
- 9:12:32mean.
- 9:12:33But, um we only have two choices, hot or
- 9:12:37cold.
- 9:12:38versus predicting the entire temperature
- 9:12:41which would be um a numerical prediction
- 9:12:44of some exact number. Right? So that'd
- 9:12:48be a regression and then on the right
- 9:12:50would be a classification.
- 9:12:52Um now again why is this so different?
- 9:12:55You can see the types of predictions
- 9:12:56we're making are completely different.
- 9:12:57One's a number, one's a category. But
- 9:13:00again with the valuation it's like if
- 9:13:03the if the true answer in our labels was
- 9:13:0684
- 9:13:07and we predicted 83 that's a pretty good
- 9:13:10result. That's still pretty close.
- 9:13:13That's pretty close to this. So from an
- 9:13:14evaluation perspective that's pretty
- 9:13:17good. Um whereas like if I predicted
- 9:13:20cold and it's actually hot that's that's
- 9:13:22a wrong answer. So they're evaluated
- 9:13:25slightly different.
- 9:13:27Um, and that's something we're going to
- 9:13:30see as we talk about evaluation of our
- 9:13:33models once we build them is depending
- 9:13:35on if it's classification regression,
- 9:13:37there's going to be different ways of
- 9:13:38evaluating them.
- 9:13:41You can kind of see why it's very
- 9:13:43difficult to say, okay, we got exactly
- 9:13:4684 when it could be any number. Our
- 9:13:50model is going to be predicting a
- 9:13:51number. That's really hard to pin down
- 9:13:54an exact floatingoint number. So, the
- 9:13:56best we can do is kind of say, how close
- 9:13:58did I get? Like, this would be a worse
- 9:14:00answer. If I got something all the way
- 9:14:02down here, that's a really long distance
- 9:14:04to here. That's bad. That's a bad
- 9:14:06prediction. But if I get something
- 9:14:08really close, that's better, right?
- 9:14:11That's a decent prediction because it's
- 9:14:13pretty close,
- 9:14:15right?
- 9:14:17Of course, being perfect would be
- 9:14:18getting exactly right, but that would be
- 9:14:20nearly impossible to do.
- 9:14:33Okay.
- 9:14:38All right. Any questions on this?
- 9:14:41Does it make sense on regression versus
- 9:14:43classification? We're going to use those
- 9:14:44words quite a bit as we go along. So,
- 9:14:47regression predicting that continuous
- 9:14:48value. Classification predicting a
- 9:14:51category.
- 9:14:54And they're going to be um different
- 9:14:56models that do that
- 9:15:01different models being used for
- 9:15:02regression versus different models being
- 9:15:04used for classification.
- 9:15:12All right, let's talk about supervised
- 9:15:15learning uh applications here. So just
- 9:15:18to name a few, we have HR operations. is
- 9:15:22imaginary recruiter tasked with finding
- 9:15:23the best candidates. Um so supervised
- 9:15:26learning can help by um rejecting or
- 9:15:29accepting candidates. Now this is
- 9:15:30something that happens quite a bit even
- 9:15:32today. Um and that it's kind of like uh
- 9:15:37how recommendations happen like this
- 9:15:40this resume should be um recommended
- 9:15:42this should not um from a whole pool of
- 9:15:45applications. Um so there's those kind
- 9:15:49of use cases of of um predicting a
- 9:15:52category that would be like a
- 9:15:54classification. Should we should we
- 9:15:56accept or reject the the candidate?
- 9:15:59Um finance you see this all the time
- 9:16:01with things like risk and loan
- 9:16:04approvals.
- 9:16:05Um you can uh predict the the the
- 9:16:09category of like if the if the loan if
- 9:16:12we should accept or reject the loan
- 9:16:14application.
- 9:16:15um you know that would be a
- 9:16:17classification.
- 9:16:19Um what's interesting about
- 9:16:21classifications by the way so it says
- 9:16:23here like we can predict the likelihood
- 9:16:26of a of a loan being repaid
- 9:16:28um is a lot of classifications um we we
- 9:16:33say that they predict a category but
- 9:16:36under the hood they can actually predict
- 9:16:38a probability and we turn that
- 9:16:40probability into a category. So, um, you
- 9:16:45know, like we could say what's we could
- 9:16:47say the likelihood of her loan being
- 9:16:49repaid is very low. Let's say it's less
- 9:16:51than 50% probability. Um, then we could
- 9:16:54label this as reject,
- 9:16:57right? We could label that as a
- 9:16:59rejection. Um, if it's greater than 50%.
- 9:17:03Then we could label this as accept. So
- 9:17:06we can set a threshold there
- 9:17:09and say okay truly we're predicting a
- 9:17:12prob like our model spits out a
- 9:17:14probability but we turn that into a
- 9:17:16category by saying should we accept if
- 9:17:19it's less than 50% we should reject if
- 9:17:22it's greater than we should accept.
- 9:17:25Okay, so that's something we will see
- 9:17:26with some of our classification models
- 9:17:28is that they actually produce a
- 9:17:30probability and we turn that probability
- 9:17:32into a category label
- 9:17:36um by by doing something simple like
- 9:17:38this putting a threshold on it um for
- 9:17:41the for the category.
- 9:17:45So finances is used all over the place.
- 9:17:47Not only just loans like fraud, we
- 9:17:49talked about fraud, not fraud. That
- 9:17:50would be a classification.
- 9:17:52Um predicting sales revenue, that would
- 9:17:56be a regression, right? What is the
- 9:17:58revenue going to be in the next two
- 9:18:00quarters? That's going to be a
- 9:18:02regression problem.
- 9:18:05Uh emails like spam, not spam, that's
- 9:18:07going to be a classification.
- 9:18:10um that's going to operate on the that's
- 9:18:12going to take the text input and predict
- 9:18:14if this email is a spam or not spam.
- 9:18:18That's going to be a uh supervised
- 9:18:20learning problem, but it's going to be a
- 9:18:22classification problem,
- 9:18:25right? Uh manufacturing supervised
- 9:18:28learning is used to inspect and uh
- 9:18:31quality and classify products in
- 9:18:32different grades. For example, a factory
- 9:18:34might use a model to check for defects.
- 9:18:36So this actually something that happens
- 9:18:37is you look at images of products as
- 9:18:39they go through the assembly line and
- 9:18:42you can take a look at those images and
- 9:18:43predict if it's a high quality, low
- 9:18:45quality, medium quality. Um so they can
- 9:18:48be this is a classification, right?
- 9:18:50They're going into different categories
- 9:18:52of quality. Um so it's much much like a
- 9:18:55manual kind of intervention by some uh
- 9:18:58QA or quality control uh specialist.
- 9:19:04Okay. But that's a classification.
- 9:19:10So in the maritime industry, supervised
- 9:19:13learning can be used to predict current,
- 9:19:15so current level
- 9:19:17um and that can be used to forecast uh
- 9:19:20supply and demand. Um so those would be
- 9:19:23like regression models that are used to
- 9:19:26predict um kind of like temperature, but
- 9:19:28in this case like title levels.
- 9:19:34We talked about fraud already, so that's
- 9:19:36there. Um, that would be a
- 9:19:38classification.
- 9:19:43Okay,
- 9:19:46any questions on these uh examples?
- 9:19:50Of course, there's many more. Um
- 9:19:53recommendation is kind of like a
- 9:19:56supervised learning problem uh where you
- 9:19:59are
- 9:20:01taking examples of things that people
- 9:20:03have viewed in the past or or reviewed
- 9:20:06in the past and using that to predict
- 9:20:08what they would want to watch in the
- 9:20:10future. Um so recommendation is
- 9:20:14supervised learning. Um and it's like a
- 9:20:18classification, you know, trying to
- 9:20:20predict um uh certain number of
- 9:20:23categories of of uh shows or movies that
- 9:20:27you would want to watch. Um
- 9:20:31and that's something that we will study
- 9:20:33in the future. Recommend we'll we'll
- 9:20:35have a whole lesson dedicated to
- 9:20:36recommendation as well.
- 9:20:41All right.
- 9:20:43So when it comes down to the uh actual
- 9:20:48models themselves, so there's going to
- 9:20:50be lots of different models that we are
- 9:20:51going to cover. Um and they are um going
- 9:20:55to be different in their purpose and
- 9:20:58kind of their uh what kinds of problems
- 9:21:01they're used for. Um and uh their their
- 9:21:06how they actually train is going to be
- 9:21:08different. Um, but at a high level,
- 9:21:11they're all trying to do the same thing,
- 9:21:13which is learn some sort of relationship
- 9:21:15between the input data and the and the
- 9:21:17label, right? That's really what they're
- 9:21:19trying to do because they're all
- 9:21:20supervised. They're they have those
- 9:21:22labels, trying to build some
- 9:21:24relationship there. Um, they just do it
- 9:21:27differently.
- 9:21:29And what we're going to study is the
- 9:21:31pros and cons of a lot of these models,
- 9:21:33like when would I use one of them, when
- 9:21:34would I use another. Um, so we'll try to
- 9:21:37talk about that as we go along. Um, but
- 9:21:40they're all trying to learn some
- 9:21:42relationship between the input features
- 9:21:44and the output, right? So we have to
- 9:21:47keep that in mind. They're trying to
- 9:21:49model that relationship. They just do it
- 9:21:51in different ways. Okay? So as we go
- 9:21:53along and learn about new models, um, we
- 9:21:57will learn the details. will learn the
- 9:21:58ins and outs um and those pros and cons,
- 9:22:02but they're no matter what, they're all
- 9:22:04trying to
- 9:22:06uh learn that relationship, right? And
- 9:22:08be able to make predictions on new data.
- 9:22:13Okay,
- 9:22:15so here's a list of models that we will
- 9:22:18cover and work on throughout the uh the
- 9:22:22sessions that we have.
- 9:22:24um we're not going to do them all in one
- 9:22:26one sitting, but um the first one that
- 9:22:28we're going to start with and that we'll
- 9:22:30cover today is going to be linear
- 9:22:32regression.
- 9:22:33So we will cover linear regression and
- 9:22:35then we'll cover the rest of these guys
- 9:22:38mostly in the context of uh
- 9:22:41classification.
- 9:22:42So, um, what's interesting is some of
- 9:22:46these guys can actually be used for both
- 9:22:48regression and classification as long as
- 9:22:50you make, um, certain adjustments to
- 9:22:53them. They have variations that can be
- 9:22:56used to do classification and regression
- 9:22:58is very interesting. Um, but we're going
- 9:23:02to start with linear regression today
- 9:23:05and then work our way through the rest
- 9:23:07of these models when we do um, we're
- 9:23:09going to do a separate lesson four on
- 9:23:11classification. And so these all these
- 9:23:13guys will come from lesson four.
- 9:23:18Um and then uh we will do this guy in
- 9:23:23lesson three in the 3.2 notebook. We'll
- 9:23:26do all about linear regression.
- 9:23:30Yeah. I so logistic regression is a
- 9:23:32classification. Um which is kind of
- 9:23:35strange that its name is regression but
- 9:23:37it's doing a classification. But the the
- 9:23:40reason is that the logistic regression
- 9:23:42um computes a probability. So it does a
- 9:23:46regression to predict a number but that
- 9:23:49number is actually a probability. So it
- 9:23:51it produces a result that's between it
- 9:23:54produces a probability that's between um
- 9:23:58obviously uh zero and one.
- 9:24:03So it uh and then we take that
- 9:24:05probability and we turn it into a
- 9:24:07category
- 9:24:09like a spam not spam fraud not fraud.
- 9:24:12Um but so so logistic regression is kind
- 9:24:14of special. It's sort of like a
- 9:24:16regression but it's predicting a very
- 9:24:18specific type of value which is a
- 9:24:20probability. So for for that reason it's
- 9:24:23a classification uh algorithm primarily.
- 9:24:32So we'll study that one in lesson four.
- 9:24:35Uh but yeah, that's that's why it's
- 9:24:37under that kind of umbrella of
- 9:24:39classification is because it's it's
- 9:24:40producing a probability as its main
- 9:24:42output which we can then turn into a
- 9:24:46category as long as we interpret that
- 9:24:48probability as um in the right way. Uh
- 9:24:52like the probability of spam,
- 9:24:54probability of not spam.
- 9:25:00Okay.
- 9:25:03Okay. So, let me focus on um
- 9:25:07let me focus on linear regression. I'm
- 9:25:09not going to go through all of these
- 9:25:10other use cases because we haven't
- 9:25:12learned these models yet. Um so, I don't
- 9:25:16think they're good. Uh I don't think
- 9:25:19it's good to read about them yet until
- 9:25:21we've covered them. So, once we cover
- 9:25:23them in lesson four, I'll come back and
- 9:25:25describe these examples to you guys and
- 9:25:28we'll see why it makes sense. But I
- 9:25:30think for linear regression um which is
- 9:25:32what we'll cover next, let me talk about
- 9:25:34that example. So a prototypical example
- 9:25:36would be like predicting the house
- 9:25:38prices that we've seen in that house
- 9:25:40price data set.
- 9:25:42So um if we wanted to uh if we wanted to
- 9:25:47predict um if we wanted to estimate the
- 9:25:50market value of a house so the price
- 9:25:55um we could do that by using the
- 9:25:57features such as number of bedrooms,
- 9:25:59square footage, location, age of the
- 9:26:02property. Um and you know then when a
- 9:26:06new when a new house comes on the market
- 9:26:08we could estimate what the price should
- 9:26:10be based on those features. So linear
- 9:26:14regression is a good one to predict the
- 9:26:16price like a housing price. Um and we'll
- 9:26:19actually practice that in the next uh
- 9:26:22notebook.
- 9:26:25So we'll we'll uh and then all these
- 9:26:28other now there's descriptions of these
- 9:26:30other models but again we haven't
- 9:26:31covered these guys yet. So I don't want
- 9:26:32to really go through those until we get
- 9:26:35to those models. So we get to those I'll
- 9:26:37come back and mention the example.
- 9:26:40Uh can K andN be used for clustering?
- 9:26:43No. So um the clustering model is going
- 9:26:46to be different. It's going to be uh K
- 9:26:49means
- 9:26:51K means that's the primary clustering
- 9:26:53model. Not K nearest neighbors. K
- 9:26:56nearest neighbors is used for uh it can
- 9:26:59be used for regression. It can be used
- 9:27:00for classification.
- 9:27:04So we'll we'll talk about K andN which
- 9:27:06is the K nearest neighbors in lesson
- 9:27:08four.
- 9:27:11It sounds really similar. Yeah, it
- 9:27:13sounds really similar but K means is a
- 9:27:16clustering algorithm that's that's
- 9:27:17slightly different
- 9:27:19different uh there's no labels used at
- 9:27:22all. This K nearest neighbors is a is a
- 9:27:25supervised learning algorithm. It uses
- 9:27:28uh labels.
- 9:27:38Good. Any any other questions so far?
- 9:27:56Okay.
- 9:27:58So that being said, let's move on to the
- 9:28:013.2 notebook.
- 9:28:04Let's move on to that which will be our
- 9:28:07um first discussion around uh
- 9:28:11regression. So going into supervised
- 9:28:14learning and regression. Give you guys a
- 9:28:16moment to pull up this notebook.
- 9:28:19But yeah, you want to pull up the 3.2.
- 9:28:21We'll do this one next. So we'll focus
- 9:28:23in. So our plan is to do regression
- 9:28:26first and then we'll talk about
- 9:28:28classification in lesson four
- 9:28:35which we will cover all those other
- 9:28:37models which you you could use for
- 9:28:39classification uh on that list. But then
- 9:28:42we're going to talk about linear
- 9:28:43regression uh first.
- 9:28:53All right. So we have a a big agenda.
- 9:28:56This is a big notebook um to go through
- 9:28:58a lot of material here surrounding
- 9:29:02regression. So we're we're going to
- 9:29:04start with linear regression and see um
- 9:29:07how we actually perform it, what that
- 9:29:10model is doing. Um which we've kind of
- 9:29:13seen the idea of it a little bit
- 9:29:14already, so it should be somewhat
- 9:29:16familiar. Um and then we'll talk about
- 9:29:19how to adapt that linear regression idea
- 9:29:22to um nonlinear what's called nonlinear
- 9:29:26regression which is going to be using
- 9:29:27like polomial uh features. We'll talk
- 9:29:31about how to do that. Um and then a big
- 9:29:34big big topic for us is going to be
- 9:29:36evaluating the model. So it'll be it'll
- 9:29:39be quite easy to actually build it.
- 9:29:41building the model will be really easy
- 9:29:43but evaluating and interpreting that
- 9:29:46will be uh a lot of interesting work
- 9:29:50there um because we want to know what
- 9:29:53the performance of that model is once we
- 9:29:55have it built right we want to know how
- 9:29:57good of a model is it is it worth using
- 9:30:00or do we need to retrain it or get new
- 9:30:02data or change the model up to talk
- 9:30:05about that um how do you determine what
- 9:30:07to do based on that performance
- 9:30:10um and then we'll We'll talk about here
- 9:30:13um a couple things. We may not get to
- 9:30:14this today, but regularization
- 9:30:17which is used to boost the performance
- 9:30:19uh in certain situations um whenever the
- 9:30:23model is kind of performing um poorly
- 9:30:27against test data even though it
- 9:30:28performs pretty well on training data.
- 9:30:31In that scenario, you can use offshoots
- 9:30:34of linear regression that do some uh
- 9:30:36what's called regularization. We'll talk
- 9:30:38about that.
- 9:30:39Um, and then we'll talk about
- 9:30:41hyperparameter tuning, uh, generally as
- 9:30:44a strategy, which is something you
- 9:30:46generally do want to do when you're
- 9:30:47training machine learning models. Um, so
- 9:30:51again, these two we may not get to
- 9:30:53today, but um quite a quite a lot to get
- 9:30:57to be prior to that mainly centered
- 9:31:00around evaluation and building linear
- 9:31:03regression.
- 9:31:05Okay, so pretty cool. we'll get to our
- 9:31:07first kind of model here. This linear
- 9:31:09regression
- 9:31:10to start with.
- 9:31:14Okay,
- 9:31:16so let's start with uh linear regression
- 9:31:19here. Um, and really what linear
- 9:31:24regression is attempting to do and I
- 9:31:28want to show you this in this picture is
- 9:31:31draw this line sometimes what is known
- 9:31:34as the line of best fit. So this is our
- 9:31:37model that kind of goes through the data
- 9:31:40and it's generally a good predictor
- 9:31:44um because if you give me um features uh
- 9:31:50if you give me new features and let's
- 9:31:53say they are let's say you give me a
- 9:31:55feature that's right here.
- 9:31:57So you say, okay, I have a feature
- 9:31:59that's this value on the x- axis. Then I
- 9:32:02know all I have to do is plug that into
- 9:32:05my line equation, and I will generate a
- 9:32:09a value that's like right here.
- 9:32:12Okay, that's pretty that's on that line
- 9:32:14at that input. And that's going to be my
- 9:32:17prediction for what the output variable
- 9:32:19should be. It's just going to be
- 9:32:20something on that line. And what you can
- 9:32:23see is this line is a decent estimate
- 9:32:27for this data because it slices through
- 9:32:30this pretty evenly. So it's a good guess
- 9:32:32as to what the output should be given
- 9:32:36any one of these inputs. It's a it's a
- 9:32:38good estimator this line. And so our
- 9:32:41goal building a linear regression is to
- 9:32:43kind of build the equation of this line.
- 9:32:46So we want this equation.
- 9:32:50Equation of this line
- 9:32:54is going to be our model.
- 9:33:01Yes, it's going to look just like that.
- 9:33:03MX plus B or yeah, MX plus C. It's going
- 9:33:05to look exactly like that. uh except
- 9:33:08that it's going to be more than just MX
- 9:33:12because we have um generally more than
- 9:33:16one feature. So you think of X as a
- 9:33:17feature um it will be more than just MX.
- 9:33:20It will generally be like uh it'll
- 9:33:23generally look like this
- 9:33:30and then plus maybe some bias here plus
- 9:33:34an intercept. Yeah, it'll generally look
- 9:33:36like that. So, yeah, you're exactly
- 9:33:38right. MX plus B is the right idea.
- 9:33:41Exactly right.
- 9:33:44It'll generally look like that.
- 9:33:48Nonlinear, it can be adapted to
- 9:33:50nonlinear. Yeah. If we transform, we're
- 9:33:52going to talk about that. If we
- 9:33:54transform all of our features in a
- 9:33:56nonlinear way, um we can apply linear
- 9:33:59regression to it. Yes. And and that
- 9:34:01would be a nonlinear regression. So yes,
- 9:34:05we can do nonlinear things too.
- 9:34:08We'll talk about that.
- 9:34:18Okay. So linear regression again is the
- 9:34:21art or science I should say not really
- 9:34:24art but it is an exact science of
- 9:34:26finding the equation of this line that
- 9:34:29fits through this data. Um now why one
- 9:34:32thing you should be thinking about is
- 9:34:34why is this line a good predictor and
- 9:34:38the argument is that if you take a look
- 9:34:40at this distance from these blue points
- 9:34:42so let's say these blue points are
- 9:34:44actual data points this line is going to
- 9:34:48be found such that it minimizes this
- 9:34:52distance
- 9:34:54from the points to actually I should
- 9:34:57draw it this way from the points to the
- 9:34:59line. So, we want this distance to be um
- 9:35:04actually I should draw it that way. This
- 9:35:06way. We want this distance to be kind of
- 9:35:09at a minimum. So, it would be bad to
- 9:35:12draw a line all the way out here because
- 9:35:14then that's a lot of distance, right?
- 9:35:16So, and that would be a lot of error um
- 9:35:18contributed from not being able to
- 9:35:20predict those points in our data set
- 9:35:22very well. Um which is our training
- 9:35:25data. That's why we have labels, right?
- 9:35:27that that guide us in building this
- 9:35:29line. Um so our goal is to build that
- 9:35:32line especially so that this error or
- 9:35:36this distance can be as minimum as
- 9:35:39possible. Right? Which are all these
- 9:35:42distances from these points to the line.
- 9:35:45We want those to be as minimum as
- 9:35:48possible. So our goal is to find this
- 9:35:50equation.
- 9:35:52So we're going to build a model that's
- 9:35:54going to find this equation.
- 9:35:58of the line
- 9:36:01um such that our error
- 9:36:06is minimal.
- 9:36:09And what is the error? The error is the
- 9:36:12distance
- 9:36:16of our data points
- 9:36:22to to
- 9:36:26the line that we build. So essentially
- 9:36:28what we'll do in order to train this
- 9:36:30will be to adjust the parameters or the
- 9:36:33or in that like I think is really good
- 9:36:36you brought up the MX plus C. Basically
- 9:36:38the M and the C will adjust. So we
- 9:36:41adjust those accordingly to make this
- 9:36:44distance as small as possible.
- 9:36:48Okay? To minimize that distance as much
- 9:36:50as possible.
- 9:36:59Okay.
- 9:37:01So um where is regression used? We've
- 9:37:05already seen some examples. Here's some
- 9:37:07more uh advertising like predicting
- 9:37:10sales, predicting um oil and uh oil
- 9:37:15production and demand. Those are like
- 9:37:17forecast those are regression problems.
- 9:37:20Um retail like demand forecasting for
- 9:37:22inventory. Um healthcare predicting um
- 9:37:27uh the levels of certain um uh blood
- 9:37:32markers or you know something like that.
- 9:37:34um real estate predicting prices based
- 9:37:37on those uh talked about like square
- 9:37:40footage, bedrooms, bathrooms, those
- 9:37:42things. So regression is used again
- 9:37:44whenever we want to predict a number a
- 9:37:46numerical output um that's a regression
- 9:37:49problem.
- 9:37:55So this kind of regression we're talking
- 9:37:57about here is generally
- 9:38:01um known as uh a when that equation is
- 9:38:05linear that is known as a linear
- 9:38:08regression. So go back to that picture
- 9:38:10when we have a when that equation of the
- 9:38:13line that we find is a linear equation
- 9:38:16meaning that it is exactly the form I've
- 9:38:20been telling you. So it's it's something
- 9:38:22like um weight time feature
- 9:38:26plus weight time feature
- 9:38:29plus weight time feature
- 9:38:33and then maybe some intercept um term
- 9:38:37like some some bias term there.
- 9:38:40Um this is a linear equation because all
- 9:38:44of the features are to the single power.
- 9:38:47So it's a linear power and this is a
- 9:38:49linear combination of features with with
- 9:38:52those different weights. So this is a
- 9:38:54linear model
- 9:38:56because it is uh it's what in math we
- 9:39:00would call this a linear equation right
- 9:39:03everything is to the first power. It
- 9:39:05resembles mx plus b. It is a linear
- 9:39:07equation or linear model. Um so when we
- 9:39:13talk about linear regression that is a
- 9:39:16regression model so we're predicting
- 9:39:17some continuous target that assumes we
- 9:39:21are model our model is formed from this
- 9:39:25kind of equation a linear equation.
- 9:39:28So this is going to be our our model for
- 9:39:31a linear
- 9:39:33uh regression.
- 9:39:37Okay.
- 9:39:40And so when you when you train a linear
- 9:39:42regression, your goal is to learn these
- 9:39:45weights so that you can plug in um you
- 9:39:49can plug in any one of your uh input
- 9:39:52features and you um can generate a
- 9:39:56prediction. You can which is going to be
- 9:39:58something on that line, right? It's
- 9:40:00going to be a value that's sitting here
- 9:40:02on this line.
- 9:40:04We put in all of our features and we end
- 9:40:06up there somewhere on that line.
- 9:40:10This output.
- 9:40:14Okay.
- 9:40:25Okay. Let me pause there. Any questions
- 9:40:27on the linear model here or why it's
- 9:40:32called linear regression?
- 9:40:47Okay. And by the way in these notes um
- 9:40:51this bullet point here where it says it
- 9:40:52uses the least squares criterion to
- 9:40:54estimate the coefficients that is
- 9:40:56exactly what I said earlier with the
- 9:40:58distance. So the distance is based on
- 9:41:00the square
- 9:41:02of this this quantity like how far away
- 9:41:05you are from the line is based on this
- 9:41:07square distance here and here and here
- 9:41:11and here. So what we're trying to do is
- 9:41:14find the least distance or least squares
- 9:41:18which is that minimum distance. So
- 9:41:20that's how we find all of these weights
- 9:41:23is from minimize. We basically tune them
- 9:41:25enough using our labels. So here's our
- 9:41:29label which is the y. We basically plug
- 9:41:32in our data and tune those enough to
- 9:41:34minimize the error. It's it's a it's an
- 9:41:36optimization problem, right? We we're
- 9:41:39trying to find the minimum of this
- 9:41:44quantity which is that best fit line.
- 9:41:59Okay.
- 9:42:01So we have linear regression
- 9:42:04um and we can do a simple linear
- 9:42:08regression that only has one feature. So
- 9:42:11if it only has one feature that's
- 9:42:13exactly the so if there's only one input
- 9:42:16feature sometimes that is known as um
- 9:42:19simple regression or simple linear
- 9:42:21regression and there's only basically
- 9:42:23there's only one feature. So one
- 9:42:24independent variable is the feature.
- 9:42:28There's only one feature. And so this
- 9:42:31equation resembles the
- 9:42:36exact equation that you guys just put in
- 9:42:38there, which is um mx plus b,
- 9:42:42right? It resembles exactly that. um
- 9:42:45we're just using different symbols for
- 9:42:46those like beta beta 0 and beta 1 but um
- 9:42:50basically exactly that simple line
- 9:42:54there's only one feature. So and that's
- 9:42:57because that line is going to um that
- 9:43:01line is going to be generated uh
- 9:43:04according to that equation. So here's
- 9:43:05kind of what it looks like.
- 9:43:08This is the best fit line through all of
- 9:43:10these blue dots. This is something we're
- 9:43:12going to be able to build. we're going
- 9:43:14to be able to build that equation um
- 9:43:16pretty easily in scikitlearn.
- 9:43:20So we'll be able to find that um and it
- 9:43:22won't be too hard. So this line will be
- 9:43:26um y = beta 0 plus beta 1. So some
- 9:43:31weight beta 1 times the only feature we
- 9:43:35have x1.
- 9:43:37Okay. So in this case um we would be
- 9:43:41predicting sales. So sales would be the
- 9:43:43value basically the label that we're
- 9:43:45trying to predict and the feature that
- 9:43:48we're putting in is uh I think it's the
- 9:43:53number of TV expenses. Yep. TV expenses
- 9:43:58which is on the x- axis. So there's one
- 9:43:59feature which is um TV expense.
- 9:44:08So um on this graph this would be this
- 9:44:12would be our model.
- 9:44:14Okay that would be our model. We only
- 9:44:16have one feature and we have um these
- 9:44:20two weights. We have an intercept B 0
- 9:44:22and or beta 0 and then a one weight
- 9:44:25which gets applied to that one feature
- 9:44:28beta 1. And so our model would have
- 9:44:31certain value for beta 0 and a certain
- 9:44:33value for beta 1. That's what get that's
- 9:44:36these guys get learned
- 9:44:40learned during
- 9:44:44model
- 9:44:47training.
- 9:44:53Okay. So those are what get learned
- 9:44:56during our model training and they get
- 9:44:58learned by a a a least what's called a
- 9:45:01lease squares algorithm that is trying
- 9:45:03to minimize that distance. It tries to
- 9:45:05tweak beta 0 beta 1 to minimize this
- 9:45:08distance of this line
- 9:45:11um this line
- 9:45:13to all of these points
- 9:45:18trying to minimize this.
- 9:45:22So imagine taking a line and kind of
- 9:45:24moving it around and turning its its
- 9:45:28slope, its angle um to try to find that
- 9:45:31best fit,
- 9:45:33which reduces that error the most.
- 9:45:35Right? That's kind of what we're doing.
- 9:46:00Uh can I explain? Yeah. So uh sales is
- 9:46:04in dollars and and TV expense
- 9:46:08um
- 9:46:10uh
- 9:46:12TV actually I think it's the other way
- 9:46:13around. I think the sales is actually a
- 9:46:15quantity. So I this is number of sales
- 9:46:18that we have and TV expense is um I
- 9:46:22think I think it's in dollars. So how
- 9:46:25much money how much expense um did we
- 9:46:28put into the into the product and then
- 9:46:31this is how many sales did we have of
- 9:46:33that product.
- 9:46:37So I think it's the other way around.
- 9:46:44But what this what this graph is showing
- 9:46:46is the blue points are our actual data
- 9:46:51points. Okay. So so we have a collection
- 9:46:54like we have a data frame that has so
- 9:46:57imagine we had a data frame that has the
- 9:46:59uh true values.
- 9:47:02So it has the um TV expenses. Um, it has
- 9:47:06points that are like one. So, it has
- 9:47:10points that are like 120 and then the
- 9:47:13sale sales could be like 700
- 9:47:16700 units, let's say. And then it has um
- 9:47:20so this is just our data set, right?
- 9:47:21This would be like in a data frame that
- 9:47:23we have. And then we had ones that were
- 9:47:26um 50 and then this could be um this
- 9:47:30could be 400, let's say. And on and on
- 9:47:33and on, right? So this is our data and
- 9:47:36this data is plotted in the blue. So
- 9:47:38these are these blue points here,
- 9:47:41right? So these are the blue points here
- 9:47:43and the red points are is our model. So
- 9:47:47we built a linear regression model
- 9:47:50um where we are putting in some values.
- 9:47:53We're putting in some e fake x values
- 9:47:56here and generating some predictions
- 9:47:58which is this line
- 9:48:01this linear uh regression line.
- 9:48:04Right? And that line is derived from
- 9:48:07this data. Right? It gets learned from
- 9:48:10this supervised uh examples.
- 9:48:15Does that make sense?
- 9:48:17That line is derived from the data. It's
- 9:48:20actually um learned from like the line
- 9:48:23of best fit is learned from that data
- 9:48:30and the actual data is in the blue.
- 9:48:34So you can see we're trying to build
- 9:48:35this such that this distance is kind of
- 9:48:37a minimum.
- 9:48:39So it's an optimal fit
- 9:48:44to balance out these distances.
- 9:48:54So it's just plotting. So it's just
- 9:48:55building that relationship between the
- 9:48:57input and output. like when the when the
- 9:48:59expenses are higher, um we seem to have
- 9:49:01more sales.
- 9:49:30Uh what's perpendicular like the
- 9:49:32distance? This should be this should be
- 9:49:34perpendicular because it's a distance
- 9:49:35here.
- 9:49:39Is that what you mean? Like the distance
- 9:49:40from the real points to the line. Yeah,
- 9:49:42that should be perpendicular
- 9:49:44because it's it's a it's a distance
- 9:49:46formula.
- 9:50:06Okay.
- 9:50:10All right. So, more generally now do do
- 9:50:14we usually have one feature? No. So
- 9:50:17generally we expand this to the more
- 9:50:21general case where we have more than one
- 9:50:24feature like what we see in the housing
- 9:50:26data right where we could predict a
- 9:50:28price but we have many different inputs
- 9:50:30like bedrooms, bathrooms, square footage
- 9:50:33etc.
- 9:50:35So more broadly
- 9:50:39instead of simple linear regression we
- 9:50:42have what's known as multiple linear
- 9:50:44linear regression which means we have
- 9:50:46multiple variables or multiple features.
- 9:50:49Um so this is exactly the equation I've
- 9:50:52been talking about. Um so we just extend
- 9:50:56that that one into many features. So
- 9:50:59which is this case and then a intercept
- 9:51:02term which is uh um there as sometimes
- 9:51:06known as the bias. Um
- 9:51:10but this is the intercept term to kind
- 9:51:12of orient the line to start out in the
- 9:51:14right place. Um and uh but this is the
- 9:51:20um this is the equation that we would be
- 9:51:23building the model. This is our model
- 9:51:25essentially, right? This is the equation
- 9:51:26we would be learning.
- 9:51:29Intercept is like a constant. Yeah. So
- 9:51:31if if all of the features were zero, um
- 9:51:33this is what our our data would be. This
- 9:51:36is what our result would be. If
- 9:51:38basically if this was zero, this was
- 9:51:39zero, this was zero, it would reduce to
- 9:51:42this as the prediction. Yeah. It's like
- 9:51:44a constant. Yes.
- 9:51:53So in in geometry, the intercept is
- 9:51:56actually really important because it it
- 9:51:57orients where your line should start. So
- 9:51:59it orients like so so these values are
- 9:52:02kind of like the slope. They orient the
- 9:52:04tilt of it. Like should it be tilted
- 9:52:07like this or should it be more sloped?
- 9:52:09But the intercept orients where it
- 9:52:12should start like vertically like should
- 9:52:14it start all the way up here? Should it
- 9:52:16start more down here?
- 9:52:18Um, that's what the intercept kind of
- 9:52:21tells us.
- 9:52:31Okay, so this is the situation. This is
- 9:52:34going to be our linear regression model
- 9:52:35that we will be building most of the
- 9:52:37time because we will have again these
- 9:52:39are all going to be features.
- 9:52:41So this is some feature the X this is
- 9:52:45some feature this is some feature
- 9:52:49X1 etc. These are all features and what
- 9:52:53gets learned during the training are
- 9:52:55these coefficients. So all of these
- 9:52:56coefficients including the beta 0ero um
- 9:53:01will get learned. So these will get
- 9:53:03learned
- 9:53:05um from our data right they get learned
- 9:53:07they will be trained from our data um in
- 9:53:11order and and how do they get trained
- 9:53:13it's from reducing that distance we try
- 9:53:16to get that line of best fit by tweaking
- 9:53:18those betas enough to uh until we reach
- 9:53:22a minimum distance but there's there's
- 9:53:24an algorithm behind that um that that
- 9:53:26scikitlearn will run for us to find that
- 9:53:30best fit Um, so we don't need to do that
- 9:53:33manually, but that's that's the process
- 9:53:35is basically tweaking those weights to
- 9:53:38end up with that line of best fit. So in
- 9:53:41higher dimensions, instead of a line,
- 9:53:43you get more of what's called a plane
- 9:53:46here. Um, which kind of looks like this.
- 9:53:48So the best fit is actually this plane
- 9:53:51where all um, it kind of dissects all
- 9:53:54these points just like that um, in
- 9:53:57higher dimensions. So this is uh instead
- 9:53:59of a line you get this in in three
- 9:54:01dimensions you get this plane like this
- 9:54:04but it's still it's like a line of best
- 9:54:07it's just a more general line of best
- 9:54:09fit. It's still the same idea. Um we're
- 9:54:12still trying to um come up with the best
- 9:54:15coefficients to minimize that distance
- 9:54:17from our from our points to the line.
- 9:54:21Although in higher dimensions it's no
- 9:54:23longer a line. It's more like a plane
- 9:54:24like this. So you're trying to minimize
- 9:54:26this distance from here down to the
- 9:54:28plane
- 9:54:29here up to the plane
- 9:54:32in higher dimensions. So I want you to
- 9:54:34keep in mind what we're trying to do
- 9:54:37before we go into the code because the
- 9:54:39code's going to make it seem really
- 9:54:40really simple and that's because
- 9:54:42scikitlearn is great and that's what it
- 9:54:44does.
- 9:54:46But we should realize that there's
- 9:54:47something really complex going on which
- 9:54:49is again finding the best value of these
- 9:54:53weights
- 9:54:55that minimizes the distance of this line
- 9:54:58to the data points that we have. So
- 9:55:02there's an algorithm there that will
- 9:55:04keep trying to make adjustments to this
- 9:55:07based on those distances. So it's going
- 9:55:10to use those distances as a guide to
- 9:55:12kind of tweak them to find the one that
- 9:55:16results in the lowest amount of
- 9:55:18distance. So we keep making tweaks, keep
- 9:55:19making tweaks, keep making tweaks and
- 9:55:21eventually we try to find we converge to
- 9:55:24the set of weights that gives us that
- 9:55:26best fitting line. Um and and there's an
- 9:55:30algorithm there that occurs. Now luckily
- 9:55:33that gets abstracted for us a bit behind
- 9:55:36um scikitlearn
- 9:55:38um finding that best fit. So there'll be
- 9:55:41a function that we use in scikitlearn
- 9:55:44when we build the model that will go
- 9:55:46ahead and find the best weights for us
- 9:55:50and that's then we now have our optimal
- 9:55:53model right that then we can just plug
- 9:55:55in different values of these features
- 9:55:58and generate a prediction which is going
- 9:56:01to be this uh result right so so that's
- 9:56:04what we're ultimately trying to do is uh
- 9:56:07train the model which will uh find all
- 9:56:10those optimal weights and then uh we can
- 9:56:13predict with it which would be plugging
- 9:56:15in different feature values to to
- 9:56:18generate a prediction.
- 9:56:21Okay,
- 9:56:23so let's see how that happens. It's
- 9:56:24actually going to be super easy um with
- 9:56:27scikitlearn.
- 9:56:29So uh in this scenario we have um we're
- 9:56:33going to import our pandas because we're
- 9:56:35going to load our data from that. Um, so
- 9:56:38of course we need some data to work
- 9:56:40with. So we're going to load this uh
- 9:56:42CSV.
- 9:56:43Um, I
- 9:56:46uh so I was not actually able to find
- 9:56:49this CSV for this example, but I mean
- 9:56:52that's okay because we'll do some we'll
- 9:56:53do other examples where we'll work with
- 9:56:55the data. If you happen to have it, um,
- 9:56:58great. I didn't see it in in my files.
- 9:57:02So just have to take the word for it
- 9:57:03that these are the this is that TV and
- 9:57:06sales columns here um from this data
- 9:57:10set.
- 9:57:11Okay. Um as an example. So um just to
- 9:57:17see how it's fit um what we're going to
- 9:57:21do and this is going to be a very
- 9:57:24standard process for us for building a
- 9:57:27model. These steps are going to be very
- 9:57:29very standard for us which is going to
- 9:57:31be first of all splitting the features
- 9:57:35away from the label. That's the first
- 9:57:38step that we always will take. So if you
- 9:57:40take a look at this code, it's taking
- 9:57:43all rows but only the first column.
- 9:57:48Okay, so it's extracting all the
- 9:57:50features from the dataf frame um which
- 9:57:52happen to be which is just the first the
- 9:57:55first column uh which is the TV uh
- 9:57:59column right just that column there and
- 9:58:02our target variable which is our label.
- 9:58:06So our target variable aka the label um
- 9:58:10is the second column, right? It's that
- 9:58:14that sales column.
- 9:58:17Um and so our first step here, let me
- 9:58:21call that out here. First step is to
- 9:58:24always split apart
- 9:58:28features from labels.
- 9:58:31Okay, so we put all those features into
- 9:58:34a data frame called X and we have all of
- 9:58:37our labels into technically a series but
- 9:58:40uh sort of like a data frame, right? Um
- 9:58:44called Y, which is just the um which is
- 9:58:48just the uh uh labels. So that's just
- 9:58:52the TV values. Um now you're going to
- 9:58:55see why we do that. It's because we need
- 9:58:59um our our features and labels split
- 9:59:02apart to put them into the model
- 9:59:04building function. It expects our
- 9:59:08independent variables or our features to
- 9:59:10be separated from our answers or our
- 9:59:14labels that guide the model building.
- 9:59:17That's the first thing you got to do is
- 9:59:18separate those.
- 9:59:20Okay, so this code will separate those
- 9:59:22out into a capital X and a lowercase Y.
- 9:59:26And that's actually pretty industry
- 9:59:27standard notation. Whenever you split
- 9:59:30apart all your features, usually you put
- 9:59:33them into a data frame called capital X
- 9:59:35and then you have a lowercase Y to
- 9:59:38represent your labels. That's actually
- 9:59:40pretty standard.
- 9:59:42So it's pretty standard that um X
- 9:59:46represents
- 9:59:49features
- 9:59:50and
- 9:59:52Y represents labels
- 9:59:57label column
- 9:59:59whatever our label column is in this
- 10:00:01case it is the sales because we're going
- 10:00:04to be predicting sales
- 10:00:09using the TV column the TV quant uh
- 10:00:12expense quantity.
- 10:00:19Yeah. So what it so the assignment is
- 10:00:23that we are um the assignment is that we
- 10:00:27are
- 10:00:29uh we are um splitting apart our data.
- 10:00:34So that when we first read in the data
- 10:00:37um it is a data frame right that has two
- 10:00:40columns TV and sales.
- 10:00:44Oh perfect thank you Tim. I will I will
- 10:00:48go ahead and so if we look at this data
- 10:00:52it only has those two columns right it
- 10:00:55only has those two columns. Okay. So
- 10:00:59what we're doing with this is we are
- 10:01:02splitting apart
- 10:01:04our our independent variable our
- 10:01:07features. So this this X will contain
- 10:01:11our features
- 10:01:15and Y will contain
- 10:01:19our label.
- 10:01:23Does that make sense? We're splitting
- 10:01:24this data apart. So, we're only grabbing
- 10:01:26that first column here to be our
- 10:01:29features. And then we're we're grabbing
- 10:01:32the second column, which is the sales,
- 10:01:33because we're going to predict the
- 10:01:34sales. This is our label. We're going to
- 10:01:37we're going to build a model to predict
- 10:01:38the sales given the TV input, TV expense
- 10:01:43input. So, the first thing we have to do
- 10:01:46is split apart the features and the
- 10:01:48label.
- 10:01:51Okay, that's the first step we usually
- 10:01:52will take. And the reason we have to do
- 10:01:55that um just to reiterate the reason we
- 10:01:58have to do that is because our model
- 10:02:02will expect our our data features to be
- 10:02:05separate from the label. We will pass
- 10:02:07those in separately.
- 10:02:10X is TV. It's the first column
- 10:02:14because we're using right. It's the it's
- 10:02:16all rows but the first column
- 10:02:19which is TV.
- 10:02:24Why is sales? This is what we're
- 10:02:26predicting.
- 10:02:28We are predicting the sales given the TV
- 10:02:32expense value.
- 10:02:37Yeah. Which is why we split it into So
- 10:02:40this is the second column, right? The
- 10:02:42index one column.
- 10:02:47Uh you just put in read CSV and pass in
- 10:02:50the URL.
- 10:02:52So you could So exactly the code that
- 10:02:54was up earlier from temp
- 10:02:57um you just do this
- 10:03:01and then data equals ddread CSV URL.
- 10:03:09So we split our data into X and Y here.
- 10:03:13All right. Now, one other step that
- 10:03:15we're going to take that's a very very
- 10:03:18critical step and you're gonna we're
- 10:03:19going to see this step over and over and
- 10:03:23over and over again. So, splitting apart
- 10:03:25into X and Y will become we'll do that
- 10:03:27over and over and over and over again.
- 10:03:30Not only that, but doing this next step
- 10:03:34which is what's called a train test
- 10:03:37split. Now, let me show you what the
- 10:03:39train test split does. It takes our data
- 10:03:44And it's going to split apart our data
- 10:03:47that we have, our X and our Y data. It's
- 10:03:50going to split it apart into a
- 10:03:52percentage that will be used to train
- 10:03:54the data
- 10:03:57and then a percentage that will be used
- 10:03:59to test. Now, why would we want to do
- 10:04:02that? It's mainly so we can do
- 10:04:05evaluation. So we build the model over
- 10:04:08here and then we test it on data that
- 10:04:11has not seen before. So we reserve a
- 10:04:15percentage of the data to be used for
- 10:04:16test. Usually this this data is um
- 10:04:21somewhere between uh 20 to 30%.
- 10:04:26So somewhere between 20 to 30% of the
- 10:04:29original data. So that means the
- 10:04:31majority of it is used for training. So
- 10:04:34the majority of the of that X and Y over
- 10:04:37here is going to be between 70 to 80%.
- 10:04:42Will generally be used for for uh for
- 10:04:47training. Okay. So somewhere between 20
- 10:04:50to 30 the industry standard is some
- 10:04:52anywhere in between there. Um a lot of
- 10:04:54people like to use 30%, some people like
- 10:04:56to use 20%. Um anything in that range is
- 10:05:00acceptable. um we will I think we
- 10:05:03generally will favor like 30%.
- 10:05:06Um to be used for testing but um the the
- 10:05:11point is we don't we don't want to mix
- 10:05:13those together. We want those to be
- 10:05:15separated out so that we can have a fair
- 10:05:20evaluation, right? We want to train our
- 10:05:22data on this train our model on this
- 10:05:24data and then see how well it performs
- 10:05:28on this data that it has never seen
- 10:05:30before.
- 10:05:32Right? So in order to have data it's
- 10:05:34never seen before, we're going to take
- 10:05:35our X and our Y and we're going to split
- 10:05:37it using this function called train test
- 10:05:41split that will do this kind of
- 10:05:44splitting for us. Okay. So scikitlearn
- 10:05:47has a function called train test split
- 10:05:50that will go ahead and we're going to
- 10:05:51pass our x and our y and we'll pass in a
- 10:05:54percentage like 30% that we want to
- 10:05:58split out into a test set and then the
- 10:06:00remainder of that the 70% will be used
- 10:06:03for training the model.
- 10:06:07Okay.
- 10:06:10So what we're going to get let me redraw
- 10:06:13that. So, what we're going to get out of
- 10:06:15this for the train test split is we're
- 10:06:17going to we're going to have an X and a
- 10:06:19Y per
- 10:06:22training and test. So, we're going to
- 10:06:24get now we're going to get an X train
- 10:06:31and a Y train.
- 10:06:35So, we're going to get training features
- 10:06:36and training labels. And then we're
- 10:06:39going to get test features
- 10:06:44to plug into our model and and test
- 10:06:47answers or test labels
- 10:06:52to do evaluation because what we should
- 10:06:54be able to do is build the model over
- 10:06:56here and then apply the model on this
- 10:06:58data. Meaning we can take these features
- 10:07:01and plug it into our model and then see
- 10:07:04what answers we get and compare those
- 10:07:07answers to this testing data. Right? We
- 10:07:10should be able to do that to generate an
- 10:07:12evaluation.
- 10:07:16Okay? Now you may be wondering why do we
- 10:07:19do any of that? What's the purpose of
- 10:07:21that?
- 10:07:22Evaluating it on this test data gives us
- 10:07:26a good sense of will our model
- 10:07:30generalize to new examples. Right? If it
- 10:07:35performs pretty well on this data,
- 10:07:38that's a good signal like when it's
- 10:07:39performing pretty well on data it's
- 10:07:41never seen before, that's a good
- 10:07:44indicator that it's going to perform
- 10:07:45pretty well when we use it on brand new
- 10:07:48examples
- 10:07:50um in the future.
- 10:07:52Right. So that's a that's why we do this
- 10:07:57evaluation on this data that it has not
- 10:08:00seen before. It's going to see this
- 10:08:02training data, right? We're going to
- 10:08:04train the model on that data. But that
- 10:08:07model will never be exposed to this test
- 10:08:09data until we do the evaluation
- 10:08:12and and generate some metrics to see how
- 10:08:15good is this performing
- 10:08:18and does it have a good chance of
- 10:08:19generalizing to never before seen
- 10:08:22examples which is what we want right
- 10:08:24because we're going to use this model in
- 10:08:25the real world. It's going to be being
- 10:08:28used on new examples that it hasn't seen
- 10:08:30before. We want it to perform well. So,
- 10:08:33this is kind of our test, our
- 10:08:35evaluation.
- 10:08:39Okay. Any questions on the We're going
- 10:08:42to do this in a moment. I'll show you
- 10:08:44what it looks like in the code, but any
- 10:08:47conceptually, any questions on the train
- 10:08:49test split idea? It's a very very
- 10:08:52important idea that we um basically use
- 10:08:57part of the data to train it and then
- 10:08:58another part of it to evaluate. It's
- 10:09:01very important we do that. By the way,
- 10:09:03this has a term um this in machine
- 10:09:06learning this is called cross
- 10:09:10validation
- 10:09:14because we are using one data set to
- 10:09:18train the model and then we're cross
- 10:09:20over we're crossing that over into
- 10:09:23another data set to validate it which is
- 10:09:26the uh the the testing set.
- 10:09:33So this is called cross validation. Um
- 10:09:35there's actually many ways to do cross
- 10:09:37validation. That's something we'll
- 10:09:38study. This is a very simple way of
- 10:09:40doing cross validation. There's more
- 10:09:41complex ways. You can take your data and
- 10:09:44you can actually divide it into many
- 10:09:46sections
- 10:09:47and basically train it against most of
- 10:09:49these and evaluate it against one at a
- 10:09:51time and then rotate. So that's another
- 10:09:55way to do cross validation. We're going
- 10:09:56to study that. Um but this is the this
- 10:09:59is the simplest way to do it here.
- 10:10:08Okay.
- 10:10:11So let me show you what you get when you
- 10:10:12use train test split. So uh we're going
- 10:10:15to import from sklearn.
- 10:10:18We're uh from the model selection
- 10:10:20module. Now we haven't used this before.
- 10:10:23This is our first time using it. But
- 10:10:24here's our model selection. We're going
- 10:10:26to import this train test split function
- 10:10:30and we're going to use it on our X and Y
- 10:10:33and we're going to set a test size of
- 10:10:3730% which is which is.3. So our test
- 10:10:40size
- 10:10:42is 30%.
- 10:10:45Converted to decimal
- 10:10:48right converted to.3. So that means
- 10:10:51we're reserving 30% for that test set.
- 10:10:55Um you can set a random state. Now
- 10:10:57that's completely optional. Um the
- 10:11:00random state
- 10:11:02is for reproducibility
- 10:11:10because what the train test split is
- 10:11:11going to do is it's actually going to
- 10:11:13shuffle the data and then split it apart
- 10:11:16into the 7030.
- 10:11:18So um yes, the seed. Exactly. It's like
- 10:11:21a seed. So it's it's saying like when
- 10:11:24you do that shuffling every time I run
- 10:11:26this notebook I'm going to get the same
- 10:11:28result but it's going to be random the
- 10:11:30first it's going to be random but I'm
- 10:11:31gonna be able to reproduce that
- 10:11:33randomness with that random state. Yes,
- 10:11:38it is like a seed.
- 10:11:43Uh it's you can choose any number to be
- 10:11:46your your um your random state. It 42
- 10:11:49isn't important. You could choose zero.
- 10:11:51You could choose one. Um, you could
- 10:11:53choose any positive integer. Um, 42 is
- 10:11:57kind of like the uh industry standard.
- 10:12:01It's it's you'd have to look it up why
- 10:12:03it is. Um, apparently 42 is a special
- 10:12:07number. Um,
- 10:12:10in in kind of the history of development
- 10:12:12of this stuff, there's nothing really
- 10:12:14special about 42. You could choose a
- 10:12:16random you could choose a random seed to
- 10:12:18be uh zero. That's fine. It it doesn't
- 10:12:21really it doesn't really matter.
- 10:12:25Um you just want you could choose it to
- 10:12:27be uh one, two, three. Um you could
- 10:12:30choose it to be 15. You can choose it to
- 10:12:32be anything you want it to be. It's
- 10:12:34really so that your your shuffling is
- 10:12:37consistent. Every time you run this
- 10:12:39notebook, you get the same shuffle
- 10:12:41result. So I'm always going to get the
- 10:12:43same rows in these splits.
- 10:12:49Hitch. There it is. I knew it was from
- 10:12:51something.
- 10:12:57Yeah. So 42 is kind of like a
- 10:13:02it's it's just used ubiquitously
- 10:13:06uh you know as kind of a um paying
- 10:13:10tribute to the Hitchhiker's Guide to the
- 10:13:11Galaxy, but it's no it's there's nothing
- 10:13:14that special about 42. It doesn't it's
- 10:13:16not going to change our result or
- 10:13:18anything.
- 10:13:19It's just so that this train set split
- 10:13:22is going to shuffle our data and split
- 10:13:24it apart into 7030.
- 10:13:27You just want to set this to something
- 10:13:28so that you get a cons every time we run
- 10:13:30this notebook, we get a consistent
- 10:13:33shuffle.
- 10:13:34And so the data in these sets
- 10:13:38are uh consistent. That's all.
- 10:13:47Okay. But do you guys see how we pass in
- 10:13:50our X and our Y and we generate four we
- 10:13:53generate four different data uh
- 10:13:56quantities here which is we generate
- 10:13:58training features, test features,
- 10:14:01training labels and test labels because
- 10:14:03again we are generating these four
- 10:14:07different we're generating data on these
- 10:14:09two different sets. A training set and a
- 10:14:13test set. So we have training features,
- 10:14:17training label,
- 10:14:19and then test features, test label.
- 10:14:24Okay, that's why it's so important to
- 10:14:27split apart our data into the X and the
- 10:14:29Y. We need those split apart in order
- 10:14:32for this part to work.
- 10:14:36So by the way, these two steps we will
- 10:14:38always do for any model we build. We'll
- 10:14:41generally do X and Y and then train test
- 10:14:44split in order to generate the data that
- 10:14:48we will use for building our model.
- 10:14:55Okay. So this this data here is going to
- 10:14:58be what we actually use to guide the
- 10:15:00training of our model. So it's
- 10:15:01definitely supervised, right? Linear
- 10:15:03regression
- 10:15:05um we we will use that
- 10:15:15Okay, so we haven't built the model yet.
- 10:15:17We're just getting our data split apart
- 10:15:19and ready for the training. We haven't
- 10:15:21actually built our model yet, right?
- 10:15:24That'll be coming up uh in a moment.
- 10:15:27But this is getting our data ready. We
- 10:15:29started with our data frame. We split it
- 10:15:31apart into uh an x and a y. And we split
- 10:15:36that into a train test split. And um you
- 10:15:42know then we can uh then we can go ahead
- 10:15:45and um pass in to our model training
- 10:15:48which we'll do in a moment.
- 10:15:55Um you that's a good question. You could
- 10:15:57run so what you could do is you could
- 10:16:00run
- 10:16:01um should we import numpy? Let's see.
- 10:16:05We did. Okay. You could run the average
- 10:16:09on the um you could check the MP mean on
- 10:16:14the X train and see how it compares to
- 10:16:18um
- 10:16:20see how it compares to X.
- 10:16:25So you could you could do that and see
- 10:16:26what the average of this feature is um
- 10:16:29compared to the average of the original.
- 10:16:31They may not be perfect because we are
- 10:16:33taking a reduced data set size. So I
- 10:16:36don't think there's really any good
- 10:16:38there's not like a one-sizefits-all
- 10:16:40validation we can do because we're
- 10:16:41taking a random shuffle and taking a
- 10:16:44percent. We're taking 70% of the data
- 10:16:46out. So we're not guaranteed to maintain
- 10:16:48the same statistics. We can see if
- 10:16:50they're close.
- 10:16:52Um but does that make sense? Like we're
- 10:16:54not guaranteed to get the same stats
- 10:16:56because we're taking a slice of it.
- 10:16:57We're taking 70%.
- 10:17:00So it's not guaranteed to to to
- 10:17:03be the same distribution really.
- 10:17:10Delete that.
- 10:17:13Uh is it a good practice? Yes, it is.
- 10:17:17It is. Uh 30% is the industry standard.
- 10:17:20Anything between 20 to 30. So 0.2.25.3
- 10:17:25any of those are acceptable. It's really
- 10:17:27up to you. Um I mostly see 30%.
- 10:17:31Mo I think.3 is is a good good practice
- 10:17:34to use for sure.
- 10:17:37Um I did explain random state. Uh random
- 10:17:40state is so that you get consistent
- 10:17:43shuffling. Um you can set this to any
- 10:17:46integer that you want it to be. It it
- 10:17:48doesn't really matter. Um you can set it
- 10:17:51to uh 100, you can set it to 10, you can
- 10:17:54set it to 15. Um it just ensures because
- 10:17:57what this split will do is it will
- 10:18:00shuffle the data first. It'll shuffle
- 10:18:02the rows and then um split it apart into
- 10:18:05the into the train and test sets. So you
- 10:18:09set the random state so that the next
- 10:18:11time you run this you get the same
- 10:18:13consistent shuffling. That's the only
- 10:18:15that's the only thing it it helps you
- 10:18:17with because it is randomized but when
- 10:18:20you set a random state um it's so that
- 10:18:23like if you run it again you'll get the
- 10:18:25same shuffling.
- 10:18:27You'll get the same the shuffling
- 10:18:28matters because it it it uh dictates
- 10:18:31what ends up in in these sets.
- 10:18:39Okay.
- 10:18:41All right. So let's see let's do let's
- 10:18:44build the model
- 10:18:46um and let me show you how easy this is
- 10:18:49going to be to build the model and this
- 10:18:50is really how it's going to be for every
- 10:18:53single scikitlearn model will basically
- 10:18:55look the exact same for training it
- 10:18:58which is what's going to make it really
- 10:19:00really nice. So the first thing we have
- 10:19:02to do is import our model. So from
- 10:19:05scikitlearn we're going to be using a
- 10:19:07linear from the linear model package or
- 10:19:10the linear model module I should say
- 10:19:13within sklearn we're going to be
- 10:19:15importing the linear regression
- 10:19:18and we're going to create an instance of
- 10:19:20the linear regression here.
- 10:19:23Okay, so linear regression and look how
- 10:19:27easy this is going to be. Nearly all
- 10:19:31nearly all sklearn models use
- 10:19:37ffit function to train.
- 10:19:42So every one of them, no matter which
- 10:19:44one we use, like the decision tree, like
- 10:19:48the um logistic regression, any of those
- 10:19:52like we use for classification that are
- 10:19:53going to be coming up in lesson four,
- 10:19:55they're all going to look the same in
- 10:19:57terms of it's going to run.
- 10:20:00Which is um scikitlearn's
- 10:20:04uh generic function for training your
- 10:20:06model. So this will execute the training
- 10:20:10once we run this code. And what that
- 10:20:13again the linear regression training is
- 10:20:15going to do that least squares distance
- 10:20:18procedure or algorithm to try to find
- 10:20:22the right weights. It's trying to find
- 10:20:25those weights that minimize that squared
- 10:20:27distance uh from our line that it's
- 10:20:30trying to build to the data.
- 10:20:33And what I want you to notice is what we
- 10:20:36put into the ffit. See how we put in the
- 10:20:39training data where we put in the
- 10:20:40training features and we put in the
- 10:20:43training labels. Now this is supervised.
- 10:20:46So of course we put in the labels,
- 10:20:50right? Of course we put in these labels
- 10:20:53here and of course we put in our
- 10:20:55features here. So we're putting in all
- 10:20:58of our examples from our training split
- 10:21:03into this ffit which is going to train
- 10:21:06the model uh so that we can we can use
- 10:21:10it for prediction.
- 10:21:12Okay, it's really fast. If I run this,
- 10:21:15it's going to be pretty much instant.
- 10:21:18Pretty much instantly it gets trained.
- 10:21:20And you can see here we now have a
- 10:21:21linear regression. you can see in this
- 10:21:23little box. Um, and it and this
- 10:21:26information says that it has been
- 10:21:28fitted. So, it's now ready to be used,
- 10:21:30right? So, we now that's it. We've
- 10:21:33trained our model. We tr That's how easy
- 10:21:36that was. We did fit. Now, what we
- 10:21:38should realize is there's a lot of work
- 10:21:41going on behind the scenes of this ffit.
- 10:21:44Okay, there's a lot of work being done
- 10:21:46there to do the least squares algorithm
- 10:21:50and find those weights and and create
- 10:21:53that line of best fit. Right? So there
- 10:21:56there's a lot of work being going on
- 10:21:57there that's going on there behind the
- 10:21:59scenes, but scikitlearn is abstracting
- 10:22:02it away for us. Right? And all we have
- 10:22:04to do is fit when we're using this code.
- 10:22:08Really easy. Really easy. Fit. And there
- 10:22:12we go. We've trained our linear
- 10:22:14regression model.
- 10:22:18And by the way, if you want to see what
- 10:22:20the coefficients are, you can actually
- 10:22:23extract them if you do so if you take
- 10:22:25your lin regression and you do um
- 10:22:28coefficients like this
- 10:22:32coeff with a with an underscore. So this
- 10:22:36gives us the trained
- 10:22:39weights coefficients
- 10:22:43also known as the coefficients right.
- 10:22:46Um so if you run this you can see uh
- 10:22:48right now we have this coefficient here
- 10:22:53um which is the only coefficient we had
- 10:22:55on our feature. So we only had one
- 10:22:58feature coefficient there.
- 10:23:11And we can take a look at our intercept
- 10:23:17which is this.
- 10:23:20So this gives us the train weights
- 10:23:23and so we can look at the intercept we
- 10:23:25can look at the the the coefficient. Um
- 10:23:30so obviously if we have multiple
- 10:23:31features our model has many features
- 10:23:34it's going to have more values in that
- 10:23:36coefficient but the intercept is just
- 10:23:38the single value 7.23
- 10:23:41and then the coefficient
- 10:23:45is 0.046. So that's the weight that gets
- 10:23:48learned.
- 10:23:52Is there a size limit? No, not really.
- 10:23:54There's no size limit. Um,
- 10:23:57no, you can use as much data as you
- 10:23:59want.
- 10:24:01There's really no size limit other than
- 10:24:03what like what you can fit in memory.
- 10:24:08I'd say that's the only limit is
- 10:24:09basically what the amount of data that
- 10:24:11can fit in memory.
- 10:24:19Okay.
- 10:24:22All right. Were you guys able to run
- 10:24:23this? Were you guys able to run the
- 10:24:24linear regression ffit?
- 10:24:28Okay, perfect.
- 10:24:36Perfect. You are Okay, great. Great.
- 10:24:41So, we have a model and we can use it to
- 10:24:44predict. Um, and so that's actually what
- 10:24:46we're going to do next. If we go down
- 10:24:49here, um we're going to have a function
- 10:24:51that's going to um build a scatter plot
- 10:24:55of our original test data.
- 10:24:58Um so we're going to have our test data
- 10:25:01here.
- 10:25:03Um,
- 10:25:05and we're going to then take our uh
- 10:25:08we're going to take our training data
- 10:25:10and plot we're going to use the uh this
- 10:25:14data versus our sales predictions. So
- 10:25:18you can see we're going to you this is
- 10:25:20how by the way this is how you use the
- 10:25:22scikitlearn model to predict. You have a
- 10:25:25fit to train it and look at the function
- 10:25:28you use to predict. It's literally just
- 10:25:29called predict. That's how easy it is.
- 10:25:33and you pass in your data, all your
- 10:25:35features into this predict and it
- 10:25:37generates a prediction for every row. So
- 10:25:40every row in these features in this data
- 10:25:43frame um will end up with a prediction
- 10:25:47using our model. So what we're going to
- 10:25:49do is plot our training date uh features
- 10:25:54against the predicted sales to see how
- 10:25:58good of a fit that really was.
- 10:26:01Okay. to see to see the regression fit.
- 10:26:08Okay. And so there's the regression fit.
- 10:26:12We have all of our test data here
- 10:26:14plotted in the green. We have our blue,
- 10:26:16which is our um we have our our blue,
- 10:26:20which is our uh um training data line
- 10:26:24that we built our model on. So that's a
- 10:26:26pretty decent fit. Um, and then our test
- 10:26:29data is here. We just plotted in the
- 10:26:32green scatter. But the thing I want you
- 10:26:34to see is this prediction, right? We we
- 10:26:37were able to generate some predictions
- 10:26:39on that training um by running our
- 10:26:43predict function with our model. Now,
- 10:26:44this model has been trained. So, we've
- 10:26:47already fit it and now we're using it to
- 10:26:49predict, right? And so, we're predicting
- 10:26:51the sales and plotting that on the
- 10:26:54y-axis.
- 10:26:56So the sales are we're using the
- 10:26:57predicted sales there which is our blue
- 10:26:59line. So this is our line of best fit.
- 10:27:04So this is our model prediction.
- 10:27:12This is our model predictions. Right?
- 10:27:16You can see it's a pretty decent uh
- 10:27:17line, right? Pretty decent line of best
- 10:27:19fit. Of course, there's some error here
- 10:27:23like there, you know, it's not perfect,
- 10:27:25but it it does a decent job of being a
- 10:27:28best fit line.
- 10:27:43Okay, so look how easy that was to
- 10:27:47just to recap this to fit our model was
- 10:27:50a linear regression.fit. And of course,
- 10:27:52we're going to do more examples. So no
- 10:27:55worries uh on that. We're going to see
- 10:27:57this many many many times throughout
- 10:27:59this notebook. But we have linear
- 10:28:02regression.fit to train it. And then we
- 10:28:05have linear regression.predict
- 10:28:08to and we pass in our features and that
- 10:28:10generates a predicted output.
- 10:28:13Right. So what this is actually doing is
- 10:28:17is computing this quantity.
- 10:28:32We could do either.
- 10:28:34We could do either. Um, so we could do,
- 10:28:39so one thing we could do is plot uh, so
- 10:28:42we could swap it out. We, we could do
- 10:28:44either one. It doesn't, it's not a big
- 10:28:46deal to do the training set. We could
- 10:28:48do, so we could plot X test and then we
- 10:28:51could plot linear regression X test.
- 10:29:02So it's it's a similar line. Um it's
- 10:29:06just different input features, but the
- 10:29:08line is going to be the same. Just
- 10:29:11different inputs,
- 10:29:13but the coefficients are the same,
- 10:29:14right? It's the same line. It's just we
- 10:29:16generate different outputs.
- 10:29:22So yeah, you could do either one.
- 10:29:26This is this is honestly this is
- 10:29:28probably better. I see what you're
- 10:29:30saying. This is probably better because
- 10:29:31this is the line of best fit through
- 10:29:34this data. So that probably makes sense
- 10:29:36to do to do predict on the test set.
- 10:29:40Agreed on that. Probably makes about
- 10:29:43most sense.
- 10:29:50But you could do either one.
- 10:30:02Yeah, I think that would be the most I
- 10:30:04think that makes the most sense is for
- 10:30:05it to be on the same one just to
- 10:30:07validate. So like we could do we could
- 10:30:09do training here and then train and
- 10:30:12train just to see how that data lines
- 10:30:15up.
- 10:30:16Really, what we're trying to do is have
- 10:30:18our scattered data and then our line of
- 10:30:20best fit on the same plot. That's all
- 10:30:23we're trying to do, right? So, yeah, I
- 10:30:25think I think they should be the same.
- 10:30:30I think that makes sense.
- 10:30:34These values
- 10:30:37or which values do you want to see?
- 10:30:44Yeah, we could uh we could generate
- 10:30:46those if we just do um let's go down
- 10:30:49here. So the the line values
- 10:30:53um are going to be uh the prediction.
- 10:30:57So, um the the uh test
- 10:31:03predictions
- 10:31:05equals um
- 10:31:10test predictions equals linear
- 10:31:12regression.predict x test and then we
- 10:31:14could uh we could print out our test
- 10:31:17predictions.
- 10:31:21Yeah. So, we can see what those actual
- 10:31:23values are on our uh on the test set.
- 10:31:27Yeah.
- 10:31:37Um, we will do that. Yeah. So, you
- 10:31:39thought we were checking how well our
- 10:31:41data was trained. We will do that. Yes.
- 10:31:43We haven't learned how to evaluate this
- 10:31:44yet. We're going to talk about that
- 10:31:46coming up next. Yeah. We will do that.
- 10:31:49We just haven't learned how to do proper
- 10:31:51evaluation
- 10:31:53of a regression model.
- 10:31:55But yeah, it's something we're going to
- 10:31:57talk about for sure
- 10:31:59and see how to do in our code.
- 10:32:06Okay.
- 10:32:09All right. Any other uh questions on
- 10:32:12this example?
- 10:32:18Again big takeaways
- 10:32:21fit to train it and then predict to use
- 10:32:26it
- 10:32:28predict on the features to use the model
- 10:32:31and make predictions with it.
- 10:32:36So here is an example we we made all the
- 10:32:38predictions. This these are all the
- 10:32:39values that are on that line.
- 10:32:42These are all our predictions and notice
- 10:32:44they this is a truly regression right?
- 10:32:45These are all floatingoint values. Um,
- 10:32:48so this is definitely a regression,
- 10:32:50right?
- 10:33:02Okay.
- 10:33:10Uh, that's a good question. Um,
- 10:33:14I'm not sure if there is
- 10:33:18If there's like a verbose
- 10:33:22there's not really no there's not really
- 10:33:24a verbose you can I mean you can look at
- 10:33:26the source code if you really want to
- 10:33:28see you can view the source code to see
- 10:33:31um how it's done I can tell you I mean
- 10:33:34so generally linear regression is done
- 10:33:37in two ways either you use a formula um
- 10:33:40to to solve the optimization problem of
- 10:33:44minimizing like this this distance from
- 10:33:47the points to to the line. Um,
- 10:33:51or you use something called gradient
- 10:33:53descent, which is how a lot of these
- 10:33:55things do it is they iterate through a
- 10:33:59bunch of different iterations where they
- 10:34:00update these weights according to um a
- 10:34:04certain uh basically a gradient of the
- 10:34:08the error function. The error function
- 10:34:10in this case is the is the squared
- 10:34:13distance from the line to the uh to to
- 10:34:19the points.
- 10:34:20So uh we can compute the gradient of
- 10:34:23that and do um gradient descent. So if
- 10:34:26you really want to look into it, I would
- 10:34:28do some research on like linear
- 10:34:30regression gradient descent.
- 10:34:32Okay, linear regression gradient descent
- 10:34:35to see how that's uh how that's being
- 10:34:37done. Yeah, it it's it's a pretty simple
- 10:34:41procedure. Um, again, you have the the
- 10:34:45notion is that you want to minimize
- 10:34:48minimize the loss or the error. Uh, in
- 10:34:52this case, the loss is the square
- 10:34:54distance. So, it's like um there's like
- 10:34:58a it's a formula. It's like a sum of a
- 10:35:01square distance from your prediction
- 10:35:04um or your label sorry to your model
- 10:35:07which is the beta 0 um plus beta 1 x1
- 10:35:13plus beta 2 x2
- 10:35:16etc like your model and then squared. So
- 10:35:19this squared this is the squared
- 10:35:21distance here and you're minimizing this
- 10:35:24guy which is like a calculus problem.
- 10:35:27You you find you basically find the this
- 10:35:30is this is I'm getting so far into the
- 10:35:32weeds of this, but this is like a
- 10:35:34parabola and you work your way No, no,
- 10:35:37you're good. It's it's it's a good
- 10:35:39question. Um you work your way down to
- 10:35:42the minimum of it. Does that make sense?
- 10:35:44Like you're working your way down here
- 10:35:46and you do that through a descent
- 10:35:48process, like a descent iteration.
- 10:35:51Um
- 10:35:53so
- 10:35:54that's how these are found.
- 10:35:57Um, but you don't see that happening in
- 10:36:01the background. But if you look at the
- 10:36:02source code, it I guarantee you it would
- 10:36:04be it's either going to be this or
- 10:36:06they're going to use the they're going
- 10:36:07to use a a a matrix formula to basically
- 10:36:11solve an equation um that involves this
- 10:36:17basically the derivative of this set
- 10:36:19equal to zero and you find the minimum.
- 10:36:22Either way, you're finding the minimum
- 10:36:23of this.
- 10:36:28Okay. But yeah, I don't think Psycharn
- 10:36:31has like a uh maybe there's some type of
- 10:36:34verbose flag you can look for.
- 10:36:38I don't think they have that though. Not
- 10:36:40that I've seen.
- 10:36:51All right.
- 10:36:53So I have uh an important um concept to
- 10:36:57talk about next which is going to be uh
- 10:37:00called overfitting and underfitting
- 10:37:03um which is a really important concept
- 10:37:05that's related to the training and test
- 10:37:08data we just split apart to do
- 10:37:11evaluation.
- 10:37:13And um essentially the the issue with
- 10:37:16machine learning is that it's not
- 10:37:18perfect and it can struggle in different
- 10:37:20ways. And the two ways that it primarily
- 10:37:23struggles is going to be overfitting and
- 10:37:24underfitting. So overfitting is a
- 10:37:28situation where the model basically
- 10:37:33memorizes the training data so well that
- 10:37:36it's it fails to generalize to new
- 10:37:40examples. So what we see with
- 10:37:41overfitting is this exact sign here
- 10:37:45where we have really good performance on
- 10:37:47the training data. So when so when we do
- 10:37:49that train test split we see a really
- 10:37:51good accuracy or really low error on the
- 10:37:56training data but it does not perform
- 10:38:00anywhere near that on that test data
- 10:38:02split. So what that means is that the
- 10:38:05model is overfitting to the training
- 10:38:08data. it's basically memorizing it and
- 10:38:11it's not able to generalize very well.
- 10:38:15Now, why does that happen? It's usually
- 10:38:18because the model is way too complex.
- 10:38:21And that means generally you need to do
- 10:38:24something to reduce the complexity.
- 10:38:27Either you need to use a simpler model
- 10:38:30or you need to use some type of
- 10:38:32technique to mitigate overfitting. And
- 10:38:35we're going to we're going to study some
- 10:38:37of those techniques coming up in this
- 10:38:38notebook. Uh we might not get to it
- 10:38:40today, but we're going to study
- 10:38:42particularly what can we do to prevent
- 10:38:44overfitting because overfitting is the
- 10:38:46more common issue with machine learning
- 10:38:48models. They tend to do so well at
- 10:38:52learning from data that they pick up on
- 10:38:54small details and patterns in the
- 10:38:57training examples that they're exposed
- 10:38:58to. They don't do a great job at
- 10:39:01generalizing to new examples. they can
- 10:39:03struggle with that. So that's
- 10:39:06overfitting is struggling to generalize
- 10:39:09to new examples, but you do really well
- 10:39:11on your training data. So it appears
- 10:39:13like you have a good model, but it it's
- 10:39:16not able to go and make predictions on
- 10:39:17test data very well, which means we
- 10:39:20would not want to use that model in the
- 10:39:21real world, right? Because it's not able
- 10:39:24to generalize outside of what it's
- 10:39:26already seen. And that's not a good
- 10:39:28thing if we're trying to use it for real
- 10:39:29world examples, right?
- 10:39:32So overfitting is a real issue. Um you
- 10:39:35see it all the time. I've seen it many
- 10:39:37many times in the real world, real
- 10:39:39industry uh work that I've done.
- 10:39:42Overfitting is a is a challenge for a
- 10:39:44lot of machine learning models. And so
- 10:39:46we need some techniques to overcome
- 10:39:49overfitting and we're going to study
- 10:39:51some of those uh coming up shortly.
- 10:39:55Um, one of the things that we can do,
- 10:39:58one of the one of the things that we can
- 10:40:00do to detect overfitting is exactly what
- 10:40:03we just did, which is you split apart
- 10:40:06your data into training and testing so
- 10:40:08that you have a chance to do an
- 10:40:10evaluation to see if you're even
- 10:40:12overfitting in the first place. You want
- 10:40:14to see that performance be consistent
- 10:40:17from train to test, right? You want to
- 10:40:20see consistency. What you don't want to
- 10:40:22see is performance that drops off on the
- 10:40:25test data. It's much worse. You don't
- 10:40:28want to see that. That means that your
- 10:40:30model is overfit uh to your training
- 10:40:32data and it's not going to perform well
- 10:40:34in the real world.
- 10:40:37Okay. So, we're going to have a couple
- 10:40:38ways to uh overcome that. Talk about
- 10:40:42that. Um now, the opposite can actually
- 10:40:45happen as well, which is called
- 10:40:47underfitting.
- 10:40:48And underfitting
- 10:40:50refers to the fact that a model is too
- 10:40:53simple and it actually just performs
- 10:40:57poorly across the board. So if we see
- 10:40:59poor performance on the training and
- 10:41:04testing data, that's a good signal that
- 10:41:06the model's underfit and that means it's
- 10:41:10too simple usually and you should try
- 10:41:12using something more complex. Um, so the
- 10:41:15best way to combat underfitting is to
- 10:41:17use a more complex model. And as we go
- 10:41:21through and learn about the models,
- 10:41:23we're going to learn about which ones
- 10:41:24are simple and which ones are complex.
- 10:41:26So we're going to have a scale of kind
- 10:41:29of complexity. And if you're
- 10:41:31underfitting, you want to bump up to the
- 10:41:33to a more complex model. If you're if
- 10:41:36you're overfitting, one way of combating
- 10:41:38that is to actually go down to something
- 10:41:40more simple. Go the opposite way to
- 10:41:42something simpler. So we need to learn
- 10:41:44right now we've only learned linear
- 10:41:46regression
- 10:41:47but we will learn other models you know
- 10:41:49in the future and we'll we'll talk about
- 10:41:52uh their complexity and how they're
- 10:41:53related to each other.
- 10:41:56Okay, but these are two issues we see
- 10:41:58just to draw that out again is if we
- 10:42:01have a train test split where we have
- 10:42:037030 split let's say and we perform
- 10:42:06really well over here but we go to apply
- 10:42:08that model over here and it fails it's
- 10:42:11accuracy drops off significantly more
- 10:42:14error that's that's definitely
- 10:42:15overfitting which is not good
- 10:42:20right and then underfitting is just not
- 10:42:22performing well in either case so even
- 10:42:24on the training data itself self your
- 10:42:26your accuracy is not very good. So
- 10:42:28you're not really learning effectively.
- 10:42:31You're underfitting your model. So
- 10:42:34that's that's um underfitting case.
- 10:42:41Okay.
- 10:42:46All right. Now the issue is that it can
- 10:42:49be very difficult to balance these two
- 10:42:52and get it correct. That's what makes
- 10:42:53machine learning a little bit
- 10:42:54challenging is getting this balance
- 10:42:57correct of simplicity and complexity. So
- 10:43:01you don't want to be overly complex that
- 10:43:03you overfit, but you don't want to be
- 10:43:05overly simple that you underfit and
- 10:43:08you're not able to learn effectively. So
- 10:43:11there's a bit of a tradeoff there. And
- 10:43:12this trade-off is typically known in the
- 10:43:14community as bias variance trade-off. Um
- 10:43:18in which case, uh it's basically like a
- 10:43:20complexity simplicity trade-off. It's
- 10:43:22another word for that. Um,
- 10:43:25and so, uh, it's it's thought that, um,
- 10:43:30if you, uh, if you have very, um, if you
- 10:43:35have a situation where you're able to
- 10:43:36fit the training data very well, you
- 10:43:39risk not being able to generalize. In
- 10:43:42other words, you risk overfitting, and
- 10:43:44it's hard to um, it's hard to combat
- 10:43:48that in a way. Um, and um, on the
- 10:43:53reverse side, if you have something
- 10:43:54really simple, um, you risk not learning
- 10:43:58enough. Even if you're trying to combat
- 10:44:00that overfitting, you risk not learning
- 10:44:03enough and your model just doesn't
- 10:44:05perform as well as it could. So, there's
- 10:44:07a bit of a trade-off there of trying to
- 10:44:09find the right balance between something
- 10:44:11complex enough to learn, but something
- 10:44:14not overly complex that it's going to
- 10:44:17not generalize to new data. That's the
- 10:44:21challenge. Um, like I said, we are going
- 10:44:24to have techniques to overcome this. So
- 10:44:28luckily there are things to basically
- 10:44:30overcome this trade-off and um and help
- 10:44:34us along the way so that we don't
- 10:44:36overfit. They basically prevent
- 10:44:38overfitting
- 10:44:40um and allow us to use complex enough
- 10:44:42models um that that won't be overfit.
- 10:44:47This is in the um this was in our uh
- 10:44:51lesson 3.2 notebook. So you want to pull
- 10:44:54that one back up. We were working on
- 10:44:55Monday.
- 10:44:56Um, and just to recap this a little bit,
- 10:44:59remember we were building a linear
- 10:45:02regression, I wanted to recap some of
- 10:45:04the steps we took there, um, that we
- 10:45:08will be doing over and over again. And
- 10:45:10really the same kind of steps, uh, that
- 10:45:12we do here, we'll do in a lot of our
- 10:45:15model building. Pretty much all of our
- 10:45:17model building um, that we do, whether
- 10:45:19it's regression or classification,
- 10:45:21doesn't really matter. um we'll still be
- 10:45:23doing a lot of these steps which are um
- 10:45:27remember first we split apart our data
- 10:45:29into kind of a features and a label
- 10:45:33uh x and y and the reason that's
- 10:45:36important is because um the model
- 10:45:39training uses the features and the label
- 10:45:43um to help train the model right they
- 10:45:46use those separately um so we want to
- 10:45:48split those apart whenever we can and so
- 10:45:50we have usually Uh it's a good practice
- 10:45:53to call your features capital X and your
- 10:45:56labels lowercase Y. And what we do with
- 10:45:59that is remember we immediately split
- 10:46:02that into what we called a training and
- 10:46:05a test set. And the picture we had for
- 10:46:07that was something like this
- 10:46:11where we had about 70% of the data
- 10:46:15we used to train the model against and
- 10:46:18then the other 30% of the data we use to
- 10:46:21test the model against. Meaning that we
- 10:46:24build a model over here and we apply it
- 10:46:27to this set over here um to make
- 10:46:30predictions. And then the that's where
- 10:46:32the supervised learning really comes
- 10:46:33into play, right? is on this test set.
- 10:46:37We already have the answers. We already
- 10:46:39have the label. And so we can apply our
- 10:46:41model to this to the features over here.
- 10:46:44Predict uh what the the label should be
- 10:46:47and compare that. We can get a a metric,
- 10:46:50right, that compares how close we are in
- 10:46:53our prediction to the actual values. Um
- 10:46:56and that was some of our performance
- 10:46:58metrics. I'll recap some of those that
- 10:47:00kind of measure that distance away from
- 10:47:02our predictions to what the actual label
- 10:47:05is. Um, but remember we had this train
- 10:47:09test split function which helps us split
- 10:47:12apart our features and our labels into
- 10:47:15these uh four sets of data. So we have
- 10:47:18our training features, our testing
- 10:47:20features and then our training labels
- 10:47:22and our testing labels. So we have all
- 10:47:24of those and um really these two guys
- 10:47:27are going to be used to train the model.
- 10:47:30That's why they're called underscore
- 10:47:32train. They're going to be used to train
- 10:47:33that model and then the then we're going
- 10:47:35to predict on these set of features and
- 10:47:39then com use those predictions to
- 10:47:41compare to this set of labels right
- 10:47:44that's on the test test set. Um and you
- 10:47:47notice here our test size is set to 30%.
- 10:47:50Um, that's a pretty standard number.
- 10:47:52Anywhere between like 20 to 30% is
- 10:47:54pretty standard. Um, we'll typically
- 10:47:57use.3, but it could be 02. Anywhere in
- 10:48:00between is fine.
- 10:48:04Okay, so we had that. Hopefully that uh
- 10:48:06we remember that from Monday.
- 10:48:09So we had a train and a test set. And
- 10:48:11then building the model was actually
- 10:48:13really really easy. Once you have those
- 10:48:15train and test sets, um, we just import
- 10:48:17our model object. So from uh scikitlearn
- 10:48:20sklearn
- 10:48:22um linear model uh module from that
- 10:48:25package we import the linear regression
- 10:48:27model and then we do um linear
- 10:48:31regression.fit
- 10:48:32and we pass in our features and our
- 10:48:34labels and this is again this is where
- 10:48:37that supervised learning is really
- 10:48:38coming into play because we're passing
- 10:48:41in these labels.
- 10:48:43That's really what makes this work,
- 10:48:44right? We need those labels to help
- 10:48:46guide the model to make those updates.
- 10:48:48If you guys remember, the model is
- 10:48:51something that looks like this.
- 10:48:57So, this was a bunch of different
- 10:49:00coefficients
- 10:49:01um times the features,
- 10:49:04however many we have. Um, and so these
- 10:49:09labels are really taking the place of
- 10:49:11this and they're helping us um make the
- 10:49:15correct updates to these to these
- 10:49:17coefficients or sometimes we call them
- 10:49:20weights. Um, these B 0, B1, B2. Um, we
- 10:49:25find out what the optimal one is to get
- 10:49:27the best fit, right? To get the line of
- 10:49:29best fit. Um, that's what the model
- 10:49:33training when we call this fit. That's
- 10:49:35really what it's doing in the background
- 10:49:36is finding all those coefficients,
- 10:49:38right, to end up with the line of best
- 10:49:40fit that has the lowest amount of error.
- 10:49:46Okay, so hopefully that makes sense.
- 10:49:48That's just a fit um to train our
- 10:49:51models. And that's really going to be um
- 10:49:53the case for
- 10:49:56uh pretty much every single model that
- 10:49:59we uh train with scikitlearn. It's
- 10:50:01pretty much going to be a fit. we pass
- 10:50:03in our training uh features and our
- 10:50:06training labels.
- 10:50:09Okay, so we had that and this was the
- 10:50:12visualization of that where we had our
- 10:50:14test points kind of scattered and we see
- 10:50:17our line of best fit is the one that
- 10:50:19goes through there with that minimal
- 10:50:21error. That's that's the whole goal.
- 10:50:25Pretty decent predictor.
- 10:50:30Okay. And then we talked about
- 10:50:32overfitting, underfitting. So just to
- 10:50:34recap this, overfitting is the concept
- 10:50:37of our model basically memorizing our
- 10:50:39training data. It performs really well
- 10:50:41on that training set, but it is not able
- 10:50:44to generalize outside of that. So it
- 10:50:47performs poorly on the test set or data
- 10:50:49that it's never seen before. Um, and
- 10:50:52that's overfitting. So the reason that
- 10:50:56it overfits is generally the model is
- 10:50:58too complex and it needs to be um it
- 10:51:01needs to be simplified a bit. And one of
- 10:51:04the things we're going to do today is
- 10:51:06see a couple of ways we can alter the
- 10:51:08linear regression model um if we are
- 10:51:11overfitting to prevent overfitting. Um
- 10:51:15so there's going to be ways to handle
- 10:51:17this. Um and so we're going to explore
- 10:51:19some of those today.
- 10:51:22Uh underfitting is kind of the reverse
- 10:51:24of that. Remember it's where the model
- 10:51:26is not learning enough. So the
- 10:51:27performance is poor even on the training
- 10:51:29data. It's not good on the test data
- 10:51:32either. Um that is a sign that the model
- 10:51:35is probably too simple and maybe we
- 10:51:38should use something more complex like
- 10:51:40go from a linear regression maybe use a
- 10:51:42polomial regression. Um or maybe use an
- 10:51:45entirely different model altogether. Um,
- 10:51:48if we're underfitting, our performance
- 10:51:49is poor, it's a good signal we should
- 10:51:52try something else. Um,
- 10:51:55okay.
- 10:51:57So, we talked about those
- 10:52:01and one of the things we also talked
- 10:52:02about was evaluations. If you guys
- 10:52:05remember, we had different metrics that
- 10:52:07we could compute to get a gauge of how
- 10:52:10good our model is actually performing.
- 10:52:12Um, one of those was MSE, which is this
- 10:52:15mean squared error function. Um so we
- 10:52:17did this example during class last time
- 10:52:19on Monday um where we uh were able to
- 10:52:24generate the mean squared error. That's
- 10:52:27one of our metrics. And we can see what
- 10:52:29the mean squared error is on the
- 10:52:32training set and see what it is on the
- 10:52:33test set by um just passing in our um
- 10:52:37training predictions and our training
- 10:52:38labels, our test predictions and our
- 10:52:41test labels. pass those into this mean
- 10:52:43squared error function and it computes
- 10:52:45the MSE and that's that's a helpful
- 10:52:47function from the scikitlearn metrics
- 10:52:50um package um or module I should say and
- 10:52:55we'll be using that quite a bit to do
- 10:52:58you know evaluation of of especially of
- 10:53:00regression right mean squared error is
- 10:53:02pretty is probably the most common uh
- 10:53:06performance metric we can have and if
- 10:53:08you guys remember what it's really doing
- 10:53:10is measuring these distances So mean
- 10:53:12squared error is kind of like the
- 10:53:13average distance away from our our
- 10:53:16points to the actual um to the
- 10:53:19predictions which the predictions are
- 10:53:21all on this line. Um so it's like
- 10:53:25measuring on average how how much error
- 10:53:27do we have on average right? Um, and the
- 10:53:30idea is the closer to zero the better.
- 10:53:33Generally means that the distance away
- 10:53:35from our prediction to our points is
- 10:53:37pretty low. The closer to zero it is.
- 10:53:40Um, which is pretty desirable.
- 10:53:43So a low MSE is kind of what we're
- 10:53:45looking for. Um, closer to zero the
- 10:53:47better. And so um if one model has if
- 10:53:51one model has um a low lower MSE than
- 10:53:56another, it's it's a better performing
- 10:53:58model, right? It has less error.
- 10:54:02Okay. And then we also looked at the R R
- 10:54:05squared or sometimes known as R2 um
- 10:54:08score. Um this is another metric that we
- 10:54:12could use that measures the the
- 10:54:15variability
- 10:54:16um of uh the predictions and if our
- 10:54:21model is capturing that variability um
- 10:54:23well um and so R squ is has a range of 0
- 10:54:27to one one is better that means the
- 10:54:29model is capturing the the changes in in
- 10:54:32the um output it um our predictions
- 10:54:36follow along with those same changes um
- 10:54:38so they're pretty close um so closer to
- 10:54:41one would be a better score. So we have
- 10:54:45those kind of metrics. So like on this
- 10:54:47data um this would this would show that
- 10:54:50this model was underfitting remember
- 10:54:52because this
- 10:54:54mean this MSE was bad and this MSE was
- 10:54:58bad.
- 10:54:59Um and what we should think of these in
- 10:55:02the units of what our labels are. um
- 10:55:06especially if we take the square root of
- 10:55:08this the RMSSE that was another metric
- 10:55:11we had um the square root of this is
- 10:55:14actually in the exact units that we um
- 10:55:17have for our labels. So uh in this
- 10:55:20example this was the um this was the the
- 10:55:24units or the sales versus the TV
- 10:55:27products, right? Um and so this would
- 10:55:31indicate that on average if we take the
- 10:55:33square root of this um
- 10:55:35in the square root of this um we have uh
- 10:55:40um we're on average about 11 sales units
- 10:55:44off squared. So if we take the square
- 10:55:45roo of that um it's somewhere around 3
- 10:55:47to four um somewhere in between three
- 10:55:51and four units off. And this is as well.
- 10:55:54Um, and because both of these are still
- 10:55:57not close to zero, um, this would be
- 10:55:59under fit. And this shows that as well.
- 10:56:02This isn't that close to one. It's
- 10:56:04decent, but it's not, um, not that close
- 10:56:06to one. So, we would say, and
- 10:56:08performance is poor on both training and
- 10:56:11test sets. That's the key indicator of
- 10:56:13underfitting. It's poor on both.
- 10:56:20Yeah, exactly. High MSE correlates to
- 10:56:23underfitting. Yes. Yes. And it what's
- 10:56:25key is it's high MSE on both on both the
- 10:56:29training and the test sets.
- 10:56:33If you have a high MSE on your test set
- 10:56:35but a low MSE on your training set,
- 10:56:37that's overfitting, right? Where it's
- 10:56:40not generalizing from the training set
- 10:56:42to the test data that it hasn't seen
- 10:56:44before. That's overfitting. So the key
- 10:56:47is high MSE on both sets.
- 10:56:52All right. So we talked about that. Um
- 10:56:55we did polomial regression last time. So
- 10:56:58that was um doing
- 10:57:02that was uh making a curved graph um by
- 10:57:06transforming the features into polomial
- 10:57:08features and then doing linear
- 10:57:09regression with that. So you guys
- 10:57:11remember from Monday we did this where
- 10:57:14um we took our features and uh
- 10:57:17transformed them according to this
- 10:57:19polomial features from scikitlearn. So
- 10:57:22we can go all the way up to degree
- 10:57:23whatever degree we want. So we put in
- 10:57:25four here but there's nothing special
- 10:57:26about four really. This is just testing
- 10:57:28it out. um and we generate the the
- 10:57:31polomial features and we can fit a
- 10:57:34linear regression on those polomial
- 10:57:36features and we get a slightly better
- 10:57:39model, right? Um it fits the data a
- 10:57:42little bit better than just a straight
- 10:57:43line. This curved line with the polomial
- 10:57:48features um performs a little bit better
- 10:57:49and we could see that with the MSE,
- 10:57:51right? or we could evaluate the MSE of
- 10:57:53this um and it would be lower.
- 10:57:57It would be lower than the curve line.
- 10:57:58And so that's something we could do. Um
- 10:58:01we would just have to pass in these test
- 10:58:03predictions, the training predictions
- 10:58:05and then the the test labels and
- 10:58:07training labels and pass those into the
- 10:58:09mean squared error function and we could
- 10:58:10compute that, right? Wouldn't be hard to
- 10:58:12do.
- 10:58:16All right. And then finally where we
- 10:58:18left off um you know is on our
- 10:58:21performance metrics. So we talked about
- 10:58:23mean squared error. That's that average
- 10:58:25distance away from the labels to our
- 10:58:28predictions. Um and we take the square
- 10:58:31root of that. It's it's basically
- 10:58:33measuring the same thing but it's the
- 10:58:35square root of it is um more
- 10:58:37interpretable because it's in the same
- 10:58:38units as our label.
- 10:58:41um mean absolute error is is the average
- 10:58:45distance of the absolute value. So it's
- 10:58:47not the squared distance formula like a
- 10:58:49uklidian distance but it is a absolute
- 10:58:52value. So it's a little bit um less
- 10:58:54sensitive to outliers. They don't get
- 10:58:56magnified as much. Um but it's not
- 10:59:00typically used as much as a mean squared
- 10:59:03error would be with regression. um we
- 10:59:05talked about the last time because um
- 10:59:08the distance formula or that distance is
- 10:59:11actually what's used to train the model.
- 10:59:13So it's a more natural um fit for a
- 10:59:17performance metric for it.
- 10:59:22All right. And then we had R square. We
- 10:59:23just talked about that closer to zero
- 10:59:25would be um worse. Closer to one would
- 10:59:28be better. That means that the model
- 10:59:30explains um all the variability in the
- 10:59:33in the predictions. Uh it captures those
- 10:59:36predictions um closely to the labels
- 10:59:41um very well. So uh one would be better.
- 10:59:46Closer to one would be better.
- 10:59:49All right. So that's where we left off.
- 10:59:51Um we're gonna pick up from there with
- 10:59:53cross validation. um we've actually
- 10:59:56already seen one method of cross
- 10:59:58validation. So we're going to study um
- 11:00:00we're going to kind of recap that and
- 11:00:01and then um talk about cross validation
- 11:00:04in general um and look at some more
- 11:00:08sophisticated techniques of it um coming
- 11:00:10up next. But before I do that, any
- 11:00:13questions about anything we've covered
- 11:00:16um to this point in in the recap or
- 11:00:19anything from Monday? Any questions on
- 11:00:22that?
- 11:00:24All right. So let's talk about uh cross
- 11:00:27validation. Um now this term cross
- 11:00:32validation refers to a technique that
- 11:00:36evaluates performance. And what it does
- 11:00:39is it divides our data into essentially
- 11:00:43um training and test sets which we've
- 11:00:45kind of already seen. And then we are
- 11:00:47able to train a model on on the training
- 11:00:50set, evaluate it on the test set. And
- 11:00:52that's where that's where we get the
- 11:00:53name cross validation because we're
- 11:00:56crossing over our model from one batch
- 11:00:58of data used to train it over to another
- 11:01:01set of data used to validate those
- 11:01:03predictions. Um, and there's actually
- 11:01:06different ways to do cross validation.
- 11:01:08So cross validation is a bit of an
- 11:01:10umbrella term for multiple ways to do
- 11:01:12that. We've already seen one way of
- 11:01:14doing that um which I'm going to scroll
- 11:01:16down to is um known as a hold out cross
- 11:01:21validation. So that's um what we've been
- 11:01:23doing so far. So this is just um
- 11:01:26generating a train and a test set
- 11:01:30train um split.
- 11:01:33Um that's the that's what's known as the
- 11:01:36hold out cross validation method. Um and
- 11:01:39and this is exactly what we've been
- 11:01:41doing so far, which is you split your
- 11:01:43data into some type of split, usually
- 11:01:467030,
- 11:01:48um of a train and test
- 11:01:51and then you um train your model on this
- 11:01:54section of data and then apply it to
- 11:01:56this to evaluate performance. Right? So
- 11:01:59that's that's what's known as the hold
- 11:02:01out method. Um it is uh you know
- 11:02:06relatively simple. It's pretty fast to
- 11:02:08do. Um, but there are more robust ways
- 11:02:13to try to divide up our data a little
- 11:02:16bit uh more evenly. Instead of just
- 11:02:19having one split, we can actually do
- 11:02:21many splits, which is the idea of um the
- 11:02:24next kind of cross validation I'll
- 11:02:26cover. But hold out method is one that
- 11:02:29we've already studied. It's the most
- 11:02:30basic type of cross validation you can
- 11:02:33have. Um so hold out this is the most
- 11:02:36basic
- 11:02:38and we we've already been we've already
- 11:02:41been uh working with this type. Okay.
- 11:02:46So we've we've already seen hold out
- 11:02:48method. Let me uh explain to you a more
- 11:02:51sophisticated method a little bit more
- 11:02:53advanced of a cross validation um which
- 11:02:56is known as Kfold cross validation. So
- 11:02:59this is um going to be a little bit more
- 11:03:02advanced of a technique but this is the
- 11:03:04idea of kfold is that you take your data
- 11:03:07set
- 11:03:09and you split it into k number of what
- 11:03:14are called splits or folds. So you take
- 11:03:17your data and you let's say it was let's
- 11:03:19say k equals 5. So we have five splits
- 11:03:22here.
- 11:03:25Okay. So let's say k equals 5. We have
- 11:03:28five splits. So what we're going to do
- 11:03:33is we're going to we're going to train
- 11:03:35our model on K minus one of those folds.
- 11:03:39So if K was five, we had five splits.
- 11:03:42We're going to take our model and train
- 11:03:44it on four out of five of those uh
- 11:03:48splits. So let's say it's these four.
- 11:03:52We'll train it on these four.
- 11:03:56Okay. And then what we do is the one
- 11:03:59split that's left over, we will we will
- 11:04:03test our model against that split. So
- 11:04:05we'll test here.
- 11:04:10Okay. Now, this sounds very similar to
- 11:04:12the hold out method where we're doing a
- 11:04:14train test split, but it's a little bit
- 11:04:16this kful cross validation a little bit
- 11:04:18more sophisticated because we repeat
- 11:04:20this process that I just mentioned over
- 11:04:23and over for all combinations of the
- 11:04:25splits. So then what we'll do, this is
- 11:04:28just one trial that we'll do it again,
- 11:04:33but this time we will pick um four
- 11:04:36different splits. So, this time we might
- 11:04:39pick,
- 11:04:41let me do blue. This time we might pick
- 11:04:44this one, this one,
- 11:04:47um,
- 11:04:49this one,
- 11:04:51and this one.
- 11:04:55And then those four we will train our
- 11:04:57data on. And then we will test against
- 11:04:59this one. Okay. And we'll do we'll
- 11:05:02repeat this
- 11:05:05repeat for all combos of the folds.
- 11:05:15Okay. So we'll repeat that. So
- 11:05:17essentially what we're doing is rotating
- 11:05:19through. Every time we rotate through
- 11:05:22one of the folds is going to be left out
- 11:05:23as a test set. Now this is a little bit
- 11:05:27more robust than just a train test
- 11:05:29split, right? because we are exposing
- 11:05:32our model to more of the data in in
- 11:05:35doing this, right? Because we're going
- 11:05:37to split it evenly into five or 10
- 11:05:39splits. Those are pretty common um
- 11:05:42number of folds to use. 10 or five. Um
- 11:05:45those are the ones I've most commonly
- 11:05:47seen. Um but we're going to by rotating
- 11:05:52through which folds are being used for
- 11:05:53training, which ones being left out. um
- 11:05:56we are exposing our our model to more of
- 11:05:59the data this way than just doing a
- 11:06:00single train test split. Right? So now
- 11:06:04what do we do with with the results is
- 11:06:07every time we do this we we generate um
- 11:06:10an MSE let's say or some type of
- 11:06:12performance metric. So let's say we
- 11:06:14generate an MSE from this guy
- 11:06:17we generate an MSE from this version and
- 11:06:20we generate an MSE for all combos.
- 11:06:24each combo we generate MSE and then what
- 11:06:27we do is we average
- 11:06:30the metrics
- 11:06:33or the in this case uh if we use MSE we
- 11:06:36would average those together. So every
- 11:06:39time we do a fold combination and we
- 11:06:41keep four of them for training, one for
- 11:06:42test and we rotate through all those
- 11:06:45combinations, we are going to generate
- 11:06:47an MSE for every combination
- 11:06:50then we're just going to average those
- 11:06:52MSSE's to get a final. So the final MSE
- 11:06:56of cross val of this K-fold.
- 11:07:00So the final metric
- 11:07:03is just the average of the uh
- 11:07:06performance on all of the fold
- 11:07:08combinations. Okay. So our final MSE, we
- 11:07:12just average all those MSE from all of
- 11:07:14our combinations.
- 11:07:16Okay.
- 11:07:18Now, what's the advantage to doing this?
- 11:07:21It's way more robust of a estimate of
- 11:07:24the of the performance of the model
- 11:07:26because we're exposing it to all
- 11:07:29basically all of our data, right? We're
- 11:07:31getting a sense of how it performs
- 11:07:32across all those different folds. Um
- 11:07:35rather than just doing a single train
- 11:07:37test split, which is a bit it's basic,
- 11:07:40it works, but it's a bit basic. Um so
- 11:07:42this is more robust estimate of the
- 11:07:46performance.
- 11:07:47Now, what's the drawback to doing this
- 11:07:50is that it's more intensive. So, if you
- 11:07:52have a lot of data, this is going to be
- 11:07:54pretty expensive to do because you're
- 11:07:55going to have to especially you have a
- 11:07:57high number of folds, right? You're
- 11:07:58going to have to divide your data into k
- 11:08:01number of folds and you're going to have
- 11:08:03to do this over and over again. Um, and
- 11:08:05if it's a large data set, it might take
- 11:08:07your model a long time to train. It's
- 11:08:09going to be a little bit more uh
- 11:08:12computationally intense than if we just
- 11:08:15did a train test split.
- 11:08:17Okay, we just did a single like 7030
- 11:08:19split. We only do that once. We only
- 11:08:22train the model once, right? We train it
- 11:08:24on the 70, apply it to the 30% test data
- 11:08:28and evaluate performance that way. Um,
- 11:08:31so we're only really using the model and
- 11:08:33training the model once, but in this
- 11:08:36kfold, we're going to do it um, you
- 11:08:38know, k number of times essentially
- 11:08:42or I should say one for every
- 11:08:44combination that we have to work through
- 11:08:46of of all the folds.
- 11:08:52Okay.
- 11:08:55All right. Does that make sense? Any any
- 11:08:58questions on kf fold cross validation?
- 11:09:01So k K is an important uh number here.
- 11:09:05It it's how many folds how many splits
- 11:09:08do you have? A typical value for K is
- 11:09:10going to be somewhere like five or 10.
- 11:09:14So 10 folds or five folds. Those are
- 11:09:17pretty pretty standard
- 11:09:21from what from what I've seen.
- 11:09:25But does the does the concept make sense
- 11:09:27or is there any questions on it on in
- 11:09:29terms of um you're always going to leave
- 11:09:31one fold out. You're going to split it
- 11:09:33up into K number of folds. Always leave
- 11:09:35one out. Train on the rest of it.
- 11:09:38Evaluate on that one that gets left out
- 11:09:39and then rotate those through. And
- 11:09:41you're going to do that for every
- 11:09:42combination and average all those
- 11:09:44metrics.
- 11:09:52And by the way, there's going to be an
- 11:09:53easy function in scikitlearn that will
- 11:09:56do this for us. So managing all these
- 11:09:58combinations will be really easy. It's
- 11:10:01actually just built into scikitlearn. So
- 11:10:03we don't have to um we don't have to do
- 11:10:06this all by hand. Okay, this will be in
- 11:10:08scikitlearn. It'll handle doing all
- 11:10:10these combinations of folds for us and
- 11:10:13computing the average metric will be
- 11:10:15really easy. So um
- 11:10:19we don't have to worry about that. We're
- 11:10:20going to see an example of this coming
- 11:10:21up shortly.
- 11:10:23All right, of kfold cross validation,
- 11:10:28but this is a this is a really widely
- 11:10:30used technique. And again, like the
- 11:10:32purpose, you may be wondering like
- 11:10:33what's the purpose ultimately of doing
- 11:10:35this? It's to get a sense of if our
- 11:10:37model is going to perform well on new
- 11:10:39data. That's really what we want to
- 11:10:41know. Like is the model going to perform
- 11:10:43well when I start to use it on new data
- 11:10:45that it's never seen before? And this
- 11:10:48kffold is a decent indicator of that
- 11:10:52because we are varying which data it
- 11:10:55sees across many different folds. Right?
- 11:10:58So it's a it's kind of a good um proxy
- 11:11:02to exposing it to different kinds of
- 11:11:05data each time and seeing how it
- 11:11:07performs.
- 11:11:09All right? Because we're working our way
- 11:11:10through each one of the folds. There's
- 11:11:11always going to be one fold left out.
- 11:11:13We're going to change which fold gets
- 11:11:15left out each time. And um that's sort
- 11:11:18of mimicking the idea of we're going to
- 11:11:20apply our model to new data and see how
- 11:11:22it performs. And it's it's new data
- 11:11:26every fold.
- 11:11:45um how we know which model is best suits
- 11:11:49for which scenario because we have Yeah,
- 11:11:52that's a good question. Um,
- 11:11:54so my we're going to learn this as we go
- 11:11:57along because we haven't covered all the
- 11:11:59models yet, but generally the best
- 11:12:02advice I can give on that is
- 11:12:05you you generally want to start as
- 11:12:09simple as you can get and then if it's
- 11:12:11not performing well then work your way
- 11:12:13up to something more complex.
- 11:12:16So we are going to have models that are
- 11:12:18simpler. We're going to have models that
- 11:12:19are more complex. The rule of thumb is
- 11:12:22to start with the most simple model that
- 11:12:24works.
- 11:12:26So you're usually going to have the same
- 11:12:30ones that you're going to try in the
- 11:12:31beginning. And linear regression is a
- 11:12:34very simple model. It's usually the
- 11:12:36first one you want to try for regression
- 11:12:38because it's the simplest.
- 11:12:40Um, and for classification, we're going
- 11:12:42to have a similar like logistic
- 11:12:44regression is the simplest kind of
- 11:12:46classification model we could have. So
- 11:12:49usually want to start with that and then
- 11:12:51if it underfits like if we see it's
- 11:12:54producing a lot of error then we work
- 11:12:56our way up to a more sophisticated
- 11:12:59model.
- 11:13:01So um that's the way we that's the way
- 11:13:05it should usually go is simple to
- 11:13:07complex it based on their performance.
- 11:13:09So we evaluate it and then we can repeat
- 11:13:11the process. If it's not performing well
- 11:13:13we can try something different that's
- 11:13:15more complex if it's underfitting.
- 11:13:24Uh this is a good question. Does a model
- 11:13:26reset after training each K minus one
- 11:13:28fold? Um yeah, it's essentially like a
- 11:13:31blank model every time uh every fold. So
- 11:13:35um we imagine like you have a brand you
- 11:13:38have a fresh model every um k minus one
- 11:13:42combination. Yes.
- 11:13:50And the reason the reason it has to be
- 11:13:52that way is because you don't want the
- 11:13:55other folds influencing the model that
- 11:13:59like on on the next combination. You
- 11:14:02don't want the previous combination to
- 11:14:03influence the results on the next one,
- 11:14:05right? Um you want it to be a fresh
- 11:14:08evaluation on every combination of
- 11:14:11folds.
- 11:14:30Okay.
- 11:14:32All right. So, let me describe to you a
- 11:14:35variation on what we just um talked
- 11:14:38about with the K-fold. So, there's
- 11:14:40another cross validation known as
- 11:14:42stratified K-fold. And um this is the
- 11:14:46same exact procedure as k-fold except
- 11:14:49that when we this is used for
- 11:14:51classification.
- 11:14:52Um so when we do classification
- 11:14:55uh we want to make sure that the
- 11:14:57different categories are going to be um
- 11:15:00split amongst those folds in a
- 11:15:03proportional way. So we don't what we
- 11:15:05don't want to happen is um when we split
- 11:15:08apart the data. So, let's say we have
- 11:15:10let's say we're predicting um spam not
- 11:15:13spam. What we don't want to have happen
- 11:15:15when we do our splits is we don't want
- 11:15:18to have all of the spams end up in one
- 11:15:21fold and then every other fold has no
- 11:15:24spam, no spam, no spam, no spam, right?
- 11:15:28That's not very good. Um because if we
- 11:15:31if we train against all these guys, we
- 11:15:33have no shot at predicting spam when
- 11:15:35they've never seen spam before. So
- 11:15:38stratify kffold is is used in
- 11:15:40classification
- 11:15:42and it's to um it's to make our splits
- 11:15:46ensure that they have basically a
- 11:15:49balanced number of categories for each
- 11:15:51split. Um so that we don't end up with
- 11:15:54certain splits with way more spams than
- 11:15:57not spams. Um so we we do what's called
- 11:16:00stratifying where we make sure the
- 11:16:02proportions are balanced across each uh
- 11:16:05split. So this is only really useful in
- 11:16:07classification, not really necessary in
- 11:16:10regression because we're predicting a
- 11:16:11value. But if we were predicting a
- 11:16:13category,
- 11:16:15like in classification like fraud, not
- 11:16:18fraud, we don't want to do the split and
- 11:16:20have every single fraud example um by
- 11:16:23bad luck in our shuffling and split end
- 11:16:25up in one split and every other um every
- 11:16:29other split has no examples of fraud.
- 11:16:31Right? Right. So we want to stratify
- 11:16:32this to spread out those um frauds
- 11:16:35against all the other splits. Um so uh
- 11:16:40again um scikitlearn will take care of
- 11:16:42that for you. Um but if you're doing
- 11:16:44classification and you have an
- 11:16:46imbalanced data set um you you really
- 11:16:49want to make sure you stratify kfold. um
- 11:16:52imbalanced meaning that you have a a um
- 11:16:56different number. Like if you're doing
- 11:16:58fraud, not fraud, you have way more not
- 11:17:00frauds than frauds. Um when where that
- 11:17:02category is imbalanced,
- 11:17:05you want to make sure it's balanced
- 11:17:06across all your splits.
- 11:17:09Um so this is this is useful in
- 11:17:12classification only, not really
- 11:17:13regression, which is what we're talking
- 11:17:14about right now. Um but it's just a
- 11:17:17variation on this that ensures when we
- 11:17:19do those folds um the data is
- 11:17:22distributed evenly amongst those folds
- 11:17:24as much as we can. The labels are I
- 11:17:26should say.
- 11:17:28Okay. So that's stratified kfold. It's
- 11:17:32the same same procedure once we have our
- 11:17:34splits. It's the same where we do k
- 11:17:35minus one of them. We train test on that
- 11:17:38last fold um and then rotate through all
- 11:17:42the folds and and average all the
- 11:17:44metrics. the same exact procedure. It's
- 11:17:46just the splitting itself um is going to
- 11:17:49be balanced in a stratified kfold.
- 11:17:55Okay, so hold out we've already talked
- 11:17:56about um is just doing a single train
- 11:17:59test split. We've talked about that. One
- 11:18:02more variation that is a bit of an
- 11:18:03extreme version of Kfold. So it's
- 11:18:06actually the same process as Kfold, but
- 11:18:08it's an extreme version is if you set K
- 11:18:12equal to the number of data points. So
- 11:18:14you basically are um this is a really
- 11:18:17really extreme kfold where you um
- 11:18:20basically are training on all the data.
- 11:18:23Um so you're training on all the data
- 11:18:26except one point and then you test
- 11:18:29against that one point. Um now why would
- 11:18:33you ever do this? Um it's mainly so for
- 11:18:36this reason here. it's to um maximize
- 11:18:40the amount of training data that your
- 11:18:42model gets exposed to because instead of
- 11:18:44just doing instead of just doing five
- 11:18:46splits
- 11:18:48um which would be like
- 11:18:51you know these four folds are going to
- 11:18:53be used and then we um test against one
- 11:18:55fold um we're essentially going to use
- 11:18:5999% of the data right one point is going
- 11:19:02to be left out 99% of the data gets used
- 11:19:05to train um and then we're always going
- 11:19:08to leave out one point and and the issue
- 11:19:10is we're actually going to do that over
- 11:19:11and over and over again and rotate that
- 11:19:13one point to cover the whole data set.
- 11:19:16So we're going to train on 99% leave one
- 11:19:19that one point out
- 11:19:21and then rotate through every
- 11:19:23combination of points until we've left
- 11:19:25out every single point and then average
- 11:19:28all those together. Um so this is a this
- 11:19:30is an extreme kffold. Again the number
- 11:19:33of folds is actually equal to the number
- 11:19:35of data points in this case. So we have
- 11:19:37every point is its own fold and we train
- 11:19:40on everything but one test on that one.
- 11:19:43This gets you the maximum size of your
- 11:19:46training data because you're basically
- 11:19:48going to have every point but one used
- 11:19:50in the training.
- 11:19:52This gets you the maximum size. However,
- 11:19:54it gets you the maximum uh expense
- 11:19:58especially for large data sets. This is
- 11:20:00going to be usually you're not going to
- 11:20:01use this um especially for large data
- 11:20:04sets because it's just too extreme. It's
- 11:20:08going to take you a really long time to
- 11:20:09work through every single point being
- 11:20:12left out. Um it's just going to take a
- 11:20:15while to do.
- 11:20:17So for that reason, the leave one out um
- 11:20:20that that's why it's called leave one
- 11:20:21out because it's you're leaving one out
- 11:20:24every single time. Um is rarely used. I
- 11:20:27I don't really see it used that often,
- 11:20:30but it is an extreme version of K-fold
- 11:20:33cross validation.
- 11:20:36Okay. But rarely ever actually used. I
- 11:20:38think the the ones that get used the
- 11:20:40most are definitely the hold out method
- 11:20:42with just a regular train test split. Um
- 11:20:45and then uh the other one that gets used
- 11:20:47quite a bit is is Kfold
- 11:20:50or stratified K-fold if you're if you're
- 11:20:52doing classification, but certainly
- 11:20:54K-fold in the in a regression case.
- 11:20:58Okay.
- 11:21:02All right. Um, we're going to do an
- 11:21:05example with these guys. So, we'll do
- 11:21:07that next. Um, with with the different
- 11:21:09cross validation techniques. Um, but any
- 11:21:14questions on what they are doing
- 11:21:17conceptually before we actually do the
- 11:21:19code example?
- 11:21:36Okay,
- 11:21:38very good.
- 11:21:44All right, so let's see some examples.
- 11:21:47Um let's go into our code and build a
- 11:21:51model and do the different cross
- 11:21:54validation techniques on it. Um you're
- 11:21:57going to see it's actually going to be
- 11:21:58really easy to do and we it sounds
- 11:22:00complex like doing the kfold and leaving
- 11:22:03one out and testing it sounds kind of
- 11:22:05complex but I promise you scikitlearn
- 11:22:07makes it really easy to do. Um
- 11:22:11and so uh we won't need to do too much
- 11:22:14besides just use the right uh tools from
- 11:22:17scikitlearn. Uh so we're going to we're
- 11:22:19going to see that. Um so here we have
- 11:22:21some imports. The um primary uh thing
- 11:22:25that's a little bit new for us is going
- 11:22:27to be these um different kinds of cross
- 11:22:30validation techniques. So we have our
- 11:22:31kfold, we have our stratified kfold,
- 11:22:33leave one out um which are those
- 11:22:35different cross validation techniques.
- 11:22:38Um these are going to be used in
- 11:22:40combination with this cross val score
- 11:22:45which is going to keep track of the
- 11:22:47different um metrics and then average
- 11:22:49them
- 11:22:51uh while we do one of these um cross
- 11:22:55validation techniques. So this guy gets
- 11:22:58used in combination with one of these to
- 11:23:01um as as we're going to see in the code
- 11:23:04uh to average those metrics. um doing
- 11:23:07the different folds, right? Perform
- 11:23:09doing performance against the different
- 11:23:10folds. Okay. And then of course we need
- 11:23:13a model
- 11:23:15using linear regression. That's that's
- 11:23:16the one we've studied so far. Um and
- 11:23:20then we have just a regular metrics if
- 11:23:22we want to compute those. Um using maybe
- 11:23:25just hold out, right? And hold out um
- 11:23:28which which is just a regular train test
- 11:23:30split. Um we could use these guys to
- 11:23:32evaluate performance.
- 11:23:34But in a more sophisticated K-fold style
- 11:23:37of cross audition, we're going to use
- 11:23:39this to evaluate the the performance.
- 11:23:45Okay, let's see.
- 11:23:48So, we're going to be working with this
- 11:23:50housing with ocean proximity data. Um,
- 11:23:53you guys should have this one. Uh, so
- 11:23:58you guys should have this one. So, if
- 11:24:00you want to follow along and run it
- 11:24:01yourself, um, you can load that one in.
- 11:24:06Um, I want to make sure that I have it.
- 11:24:11Let me pull that one in. So, it should
- 11:24:12be this guy.
- 11:24:23I'm going to load that in so I can make
- 11:24:25sure I run it with you guys.
- 11:24:28Um,
- 11:24:37so let me run this.
- 11:24:41Do you guys have that data?
- 11:24:44The housing with ocean proximity?
- 11:24:48It's another it's another housing data
- 11:24:50set. Um,
- 11:24:53but it it's a little bit different than
- 11:24:55the ones we've seen before. It has a a
- 11:24:58special feature for how close it is to
- 11:25:00the ocean, the different locations.
- 11:25:10So, it looks kind of like this. If we
- 11:25:12load it in and do our head, which is
- 11:25:14usually what we do, right, we can see um
- 11:25:17we can see that it's got these features.
- 11:25:19So, it's got uh uh bedrooms, total
- 11:25:23rooms,
- 11:25:25um it's got uh median age. Now, this is
- 11:25:28this is looks a little strange for total
- 11:25:30rooms and um uh bedrooms and population,
- 11:25:35etc., but it's um
- 11:25:39it's it's got those uh it's got those
- 11:25:42because it's representing an entire
- 11:25:44neighborhood. So, it's an entire
- 11:25:46neighborhood and we're looking at this
- 11:25:49um this is actually going to be our
- 11:25:50label is this median house value for the
- 11:25:52entire neighborhood. So, what's that
- 11:25:54median value uh in the neighborhood? And
- 11:25:58this is the total number of bedrooms,
- 11:26:00total number of rooms, um population,
- 11:26:03households. So, how many houses are
- 11:26:06there? Um median income. And of course,
- 11:26:08these are scaled. So these are um likely
- 11:26:11times you know uh thousands
- 11:26:16um
- 11:26:18but um that's our data. We could
- 11:26:22describe it.
- 11:26:27So we can see the average age median age
- 11:26:31um which sounds a little um weird but
- 11:26:33that's it's because again this is the
- 11:26:36median of data within a neighborhood. Um
- 11:26:40so the average of those is about 28 or
- 11:26:4329. Um we have
- 11:26:47u
- 11:26:51total bedrooms. The we can look at the
- 11:26:53min. There's some data that only has
- 11:26:56one. So it's likely only one house in
- 11:26:58there. Um which is what this represents.
- 11:27:01There's only one house. So there there
- 11:27:03is some neighborhood that only has one
- 11:27:05house. Um, and we see the median, um, we
- 11:27:10see the minimum, uh, median house values
- 11:27:13there. And then the maximum down here,
- 11:27:16um, is a pretty big number.
- 11:27:216,000 households is the largest that we
- 11:27:23have in any any one of these
- 11:27:25neighborhoods.
- 11:27:27Okay. So, just a little bit of
- 11:27:29description of the data.
- 11:27:45Okay. So then we can run.info. So this
- 11:27:47is um let me ask you guys, were you able
- 11:27:50to load this? Were you able to run this?
- 11:27:56If you're following along, were you able
- 11:27:57to
- 11:27:59load it and take a look at dot head.
- 11:28:09Okay, great. Great.
- 11:28:12Okay, so we're able to load that and
- 11:28:15then look at dot head. Perfect. Um
- 11:28:20okay.
- 11:28:22Um and then we run describe which gives
- 11:28:25us that uh usual kind of statistical
- 11:28:27description. Uh so we can see some
- 11:28:30interesting stats about those.
- 11:28:37What do you guys notice about the info?
- 11:28:40Anything interesting that we see from
- 11:28:41there?
- 11:28:56Is there any missing data?
- 11:29:01Any features that have missing data? Can
- 11:29:04we see
- 11:29:07object? Yeah, object type usually is
- 11:29:09string. If it's an object type, that
- 11:29:11usually means string. Python when we
- 11:29:14read it into pandas it usually is just a
- 11:29:16string.
- 11:29:19So that that makes sense like we have
- 11:29:20mostly numerical features but then we
- 11:29:22have a this ocean proximity which is a
- 11:29:25string.
- 11:29:34Yeah. Total bedrooms has nles. That's
- 11:29:36right. Because you can see here this
- 11:29:38does not equal the number of uh rows
- 11:29:41that we have. So, this is the number of
- 11:29:42rows. There's about 20,000 rows. That's
- 11:29:44a good size data set, right? 20,000
- 11:29:47rows. That's decent. Um, we're
- 11:29:49definitely missing some data here for
- 11:29:52sure. Um, we could count how much we're
- 11:29:54missing exactly by running this is NATO
- 11:29:58sum. Um,
- 11:30:00and so we see that total bedrooms is
- 11:30:02missing about 200 uh 200 rows are
- 11:30:06missing total bedroom uh value.
- 11:30:12Okay. And then one thing I wanted to
- 11:30:14look at is yes, this is a string. So
- 11:30:16what remember what we can do with those?
- 11:30:18That's a categorical.
- 11:30:20So ocean proximity
- 11:30:26is a categorical
- 11:30:28string
- 11:30:30feature.
- 11:30:33So we can take a look at its value
- 11:30:35counts, which is usually a good idea to
- 11:30:37take a look and see what possible values
- 11:30:40that feature could be. So if we look at
- 11:30:43our
- 11:30:45um what are we calling this? Housing
- 11:30:47data.
- 11:30:51Housing data
- 11:30:54ocean
- 11:30:56proximity
- 11:31:01dot value counts.
- 11:31:09So here's the different types that that
- 11:31:11one can be. So there's some
- 11:31:12neighborhoods that are less than 1 hour
- 11:31:14from the ocean. There's some that are
- 11:31:16inland. There's some that are near the
- 11:31:18ocean. There's some that are near a bay.
- 11:31:21There's even five of them that are on an
- 11:31:22island. So, these are the different
- 11:31:25values of the ocean proximity. So,
- 11:31:28remember, you can always do that. If you
- 11:31:29see a string feature, you can always
- 11:31:32take a look at what its um categories
- 11:31:34are. And it looks like most things are
- 11:31:37less than 1 hour from the ocean, but
- 11:31:38it's kind of evenly distributed here. Um
- 11:31:42otherwise
- 11:31:44very few islands.
- 11:31:52But as you can imagine like this feature
- 11:31:54is probably going to be important for
- 11:31:56determining um what the value is, right?
- 11:32:00Probably going to be important.
- 11:32:07Okay. So, um, we need to deal with these
- 11:32:10NLES. If we're going to build a model,
- 11:32:12right? So, um, this is all of our
- 11:32:14typical data prep. If we want to build a
- 11:32:17model, we're going to have to deal with
- 11:32:18these NLES. What do you guys think we
- 11:32:20should do with the NLES? What would you
- 11:32:22what do you think for total bedrooms?
- 11:32:24What do you think is a good strategy to
- 11:32:26do? Keep in mind, we have 20,000 points,
- 11:32:3120,000 rows I should say, and about 200
- 11:32:35of them are null.
- 11:32:42Right. So about 200 are null. Um so what
- 11:32:47do you what do you guys think would be
- 11:32:48like a good strategy to deal with those
- 11:32:49nles in that case?
- 11:33:10average. We can't ignore it because we
- 11:33:14can't ignore that column.
- 11:33:17We can't ignore the whole column. So,
- 11:33:20something needs to go there.
- 11:33:29Probably don't want to make it zero.
- 11:33:31I think average is a decent average is a
- 11:33:34decent idea. Probably don't want to make
- 11:33:36it zero because um that would indicate
- 11:33:39that there's no bedrooms and yet we
- 11:33:41still have a bunch of total rooms. So,
- 11:33:43it probably doesn't make sense to do
- 11:33:45zero.
- 11:33:50Average, I think average could be a
- 11:33:52decent one.
- 11:33:54Now, in this example, what we're
- 11:33:56actually going to do is we're
- 11:34:01rows.
- 11:34:03We're actually going to drop the rows al
- 11:34:04together. Now, why are we doing that?
- 11:34:06It's because we have so much data and
- 11:34:10only 200 of them are null.
- 11:34:14Okay, only 200 of them are null. So,
- 11:34:16we're actually just going to drop the
- 11:34:17rows. Now, that's a choice.
- 11:34:22Um, that's a choice, right? Is that we
- 11:34:25could fill in with the average like you
- 11:34:26guys are suggesting. What we're actually
- 11:34:29going to do is just drop the rows. It It
- 11:34:31makes up less. It makes up about 1% of
- 11:34:35the whole data. So it's not that much of
- 11:34:38it is missing. We can drop those rows.
- 11:34:41So that's actually what we're going to
- 11:34:42do here is we remove all the roles with
- 11:34:46the NLES by doing drop NA. So this just
- 11:34:48drops them. So those rows are cut out.
- 11:34:51Um, it's arguable that we could replace
- 11:34:57it's arguable that we could just replace
- 11:34:58it with something and I think you guys
- 11:35:00have good thoughts which is the average
- 11:35:02a default
- 11:35:04um assume total bedrooms. We could we
- 11:35:07could try that. Yeah.
- 11:35:10Assign a value based on comparable home
- 11:35:13value. Yes, you could do that too.
- 11:35:14That's a good strategy is to look at the
- 11:35:17other rows that are similar to it and
- 11:35:19fill in a value. That's absolutely fair.
- 11:35:22Um, in this example, we're actually just
- 11:35:24going to drop those rows,
- 11:35:28but I think that's totally um totally
- 11:35:31valid.
- 11:35:33This is a choice.
- 11:35:36We could fill NA with different values
- 11:35:42such as the average,
- 11:35:45total bedrooms,
- 11:35:49um, derive a value, etc. So, we could
- 11:35:53derive something, which I think Brent,
- 11:35:55you have a good suggestion. That's a
- 11:35:57good suggestion. Um, we could derive
- 11:35:59something like that, uh, and fill in the
- 11:36:01blank, and that's I think that's totally
- 11:36:03valid. Um, we could take the average of
- 11:36:06the um bedrooms. Uh, I meant total rooms
- 11:36:10here. Sorry, total rooms. Um, we could
- 11:36:13fill in we could fill it in with the
- 11:36:15total rooms for that category um or for
- 11:36:18that row. Um, many options. In this
- 11:36:23case, we're actually just going to drop
- 11:36:24those rows because they make up such a
- 11:36:27small percentage relative to the 20,000
- 11:36:30rows that we have. It's about 1%, right?
- 11:36:33200 rows is about 1% of 20,000.
- 11:36:37So, we're just going to drop them. But
- 11:36:39that's a choice. We don't have to drop
- 11:36:41them. We could fill in with something.
- 11:36:43Um, and if we did that, we would use
- 11:36:46fill na rather than drop NA, right?
- 11:37:02Uh after dropping the rows, how many? So
- 11:37:05it's just so after we drop the rows, um
- 11:37:08after we drop the rows, it's just going
- 11:37:09to be we still have all our other rows
- 11:37:12are intact, right? So if we look at this
- 11:37:14now,
- 11:37:22we now have um slightly uh slightly less
- 11:37:26entries.
- 11:37:28So now we have this this many um rather
- 11:37:32than rather than this many,
- 11:37:36right? We dropped those 200
- 11:37:45But they're all filled in. Yeah, they're
- 11:37:47So all the other columns are still
- 11:37:48filled in. We're just we're we're
- 11:37:50cutting out the whole row. So if you
- 11:37:52think about our data set, um we have all
- 11:37:54these rows and all these columns. What
- 11:37:58we're doing is like if there's a null
- 11:38:00here, we're just we're just getting rid
- 11:38:02of that whole row, right? And so we
- 11:38:05still have all the other rows intact.
- 11:38:14Uh, we can drop them because we have a
- 11:38:16good sample size. Yes,
- 11:38:22that's exactly right, Ronald. Yep, we
- 11:38:24can drop them because we have we have
- 11:38:2520,000 rows and only 200 are missing
- 11:38:28values. So, that's totally fine.
- 11:38:35Uh, drop a removes all rows that has any
- 11:38:37null. Yes, that's true. It it will go
- 11:38:40ahead and just drop any row where
- 11:38:42there's any null, no matter what column
- 11:38:44it's in. Yes,
- 11:38:55index. Yeah, the index is not getting
- 11:38:57reset. Um that's true. So um what we
- 11:39:03what you can always do is you can reset
- 11:39:05the index. So, um, if you want to, it's
- 11:39:08optional. We we're not really going to
- 11:39:10use the index for anything that
- 11:39:12important, right? But what we could do
- 11:39:14is, uh, reset index.
- 11:39:19Uh,
- 11:39:21we could do that, right? Which will
- 11:39:22reset it.
- 11:39:36So now now it gets reset.
- 11:39:56But um let me actually I don't I don't
- 11:39:59really want to do that. I'm going to
- 11:40:01reset this.
- 11:40:08Um,
- 11:40:16yeah, we could do that.
- 11:40:38Okay.
- 11:40:40So now importantly there should be uh no
- 11:40:43missing data of this of this new one
- 11:40:45where we've dropped NAS right. So now
- 11:40:46this is good. If you now the reason we
- 11:40:49had to do this is because if we try to
- 11:40:51build a linear regression and we have
- 11:40:53nles in there um the the issue is like
- 11:40:57how do you build a model where you have
- 11:40:59something like this
- 11:41:08and these are null like what do how do
- 11:41:10you multiply a number by a null?
- 11:41:14Um, we can't really do that, right?
- 11:41:17We can't really do that. So, um,
- 11:41:27so therefore, uh, we need to get rid of
- 11:41:30NLES like the the null's not really
- 11:41:32going to work in there. So, uh, we need
- 11:41:35to get rid of them for linear regression
- 11:41:37to to really have a chance to work,
- 11:41:38right? To train it and be able to use
- 11:41:40it.
- 11:41:42You got to get rid of those nles.
- 11:41:50All right,
- 11:41:52any questions so far? So, we haven't
- 11:41:54done any modeling yet. We're doing some
- 11:41:55We're doing some data preparation before
- 11:41:57we get to the modeling. And we haven't
- 11:41:59done any cross validation yet. We
- 11:42:00haven't set that up. We're just doing
- 11:42:02our data preparation before we get to
- 11:42:04the modeling. Right. So, we've dropped
- 11:42:06some NAS. We've checked it. Um, we're
- 11:42:10going to do one more prep step, which is
- 11:42:12to um change that ocean proximity
- 11:42:16feature into something numerical because
- 11:42:18again, how do you build a model where
- 11:42:20you're inserting a string into those
- 11:42:23like beta 1, beta 2, beta 3 times of
- 11:42:25features? You can't really do that when
- 11:42:27it's a string. Um, so what we're going
- 11:42:30to do, and I'm going to get rid of this
- 11:42:32because I don't think we really need
- 11:42:33that. um is we are going to uh run this
- 11:42:38get dummies function which is our um our
- 11:42:42get dummies function is our usual one to
- 11:42:46uh our git dummies one is our usual one
- 11:42:48to um
- 11:42:51uh get our one hot encoding. So this is
- 11:42:55our uh one hot encoding here.
- 11:43:02We now are going to have data that's
- 11:43:05like this, right? So we have ocean. So
- 11:43:08So by the way, this prefix
- 11:43:10um this prefix is OP, which which is
- 11:43:14short for ocean proximity, right? So we
- 11:43:16have ocean proximity uh less than 1 hour
- 11:43:19from the ocean, ocean proximity inland,
- 11:43:22ocean proximity island, near bay, near
- 11:43:24ocean. So these first five rows are near
- 11:43:27the bay. Um so they have a one there and
- 11:43:30a zero in the other spots. So this is
- 11:43:32good. This one hot encodes that feature
- 11:43:35into these numerical uh values,
- 11:43:39right?
- 11:43:41Were you guys able to run that one? The
- 11:43:43get dummies
- 11:43:51So the reason that Yeah, that's a great
- 11:43:53question. How did it go ocean proximity?
- 11:43:54It's because um that is the only uh
- 11:43:58string feature we have. That's the only
- 11:44:01one we have. So it it's going to look
- 11:44:03for any non-numericals and one hot
- 11:44:05encode those however many however many
- 11:44:07there are. So whatever objects we have
- 11:44:10which are strings, it's going to
- 11:44:12automatically oneh hot encode those.
- 11:44:23Yeah, we could have right we could have
- 11:44:24went here and did Right. We could have
- 11:44:27done ocean
- 11:44:30proximity,
- 11:44:33but we only have one of those features.
- 11:44:36So, it's just going to do that to the
- 11:44:38whole data frame
- 11:44:40uh on that one feature. So, what we're
- 11:44:43going to do is um go ahead and split it
- 11:44:46into an X and a Y um which the X is
- 11:44:50always what includes our features. The Y
- 11:44:54is what we are trying to predict, which
- 11:44:55is the label. Now, um, in order to
- 11:44:59separate those out, what we're going to
- 11:45:01do is assign X to be the variable that
- 11:45:04is, um, our data frame minus this median
- 11:45:08house value column. So what this is
- 11:45:10doing is um uh it's not permanently
- 11:45:14dropping because we're not uh dropping
- 11:45:17it in place but it is returning us a
- 11:45:21copy of the data frame with the median
- 11:45:24house value column left out right it's
- 11:45:26dropped. So this is this is uh something
- 11:45:29we want to do because that will the rest
- 11:45:31of it will contain our features, right?
- 11:45:33So um this will temporarily or I should
- 11:45:38say return a copy of the DF with um
- 11:45:44median house value
- 11:45:48dropped,
- 11:45:50right? Median house value dropped. Um so
- 11:45:53we go ahead and drop that one. Uh now
- 11:45:57remember it's not permanent. It's just
- 11:45:58giving us uh the remainder of it which
- 11:46:01is this housing data dropping this and
- 11:46:04it's assigning that to x and then we're
- 11:46:06taking the actual median house value
- 11:46:08column from the original data and
- 11:46:11assigning that to y. So this is going to
- 11:46:13be our labels
- 11:46:17right. So this is what we are trying to
- 11:46:22predict.
- 11:46:25Okay, so that is our Y and that's always
- 11:46:27how it is. X is our features, Y is our
- 11:46:29labels. Um, hopefully that makes sense.
- 11:46:32What this is doing is this is going to
- 11:46:35get rid of that label column and
- 11:46:37everything else will be our features.
- 11:46:39And then this will get rid of this will
- 11:46:41just assign the label column to Y.
- 11:46:46All right. And then what we can do is
- 11:46:48pass X and Y into our train test split
- 11:46:51function. And this will generate the
- 11:46:54hold out set. So if we want to do the
- 11:46:56hold out cross validation, this is how
- 11:46:58we would do it is we would split the
- 11:47:01data into X-ray, X test, Y train, Y test
- 11:47:04um using train test split. So this is
- 11:47:08what we did last time. This would be
- 11:47:10this would be for hold out cross
- 11:47:15validation,
- 11:47:18right? where we are uh uh just have that
- 11:47:22one one set for testing and one set for
- 11:47:25uh one set for training one test one set
- 11:47:28for testing I should say right so this
- 11:47:31is pretty standard train test split um
- 11:47:33we pass in that x we pass in the y we
- 11:47:36use a 30% test size is pretty standard
- 11:47:39and random state so that we get the
- 11:47:41consistent shuffling if we were to run
- 11:47:43this multiple times um we we get that uh
- 11:47:46consistent randomization
- 11:47:50Okay,
- 11:47:51so we have that and so now our X train
- 11:47:55is a percentage um of the data frame of
- 11:47:59the 20,000 uh rows and the X test is uh
- 11:48:0430% of that. So it's only about 6,000
- 11:48:06rows which is what um the shape of that
- 11:48:08is.
- 11:48:12Yeah. X so X is our features. So we're
- 11:48:15we're putting all of our data in that is
- 11:48:17our features into X. And so the the um
- 11:48:21most efficient way of doing that is um
- 11:48:24the most efficient way of doing that is
- 11:48:26to
- 11:48:28uh just take our data and drop the
- 11:48:31median house value column because that's
- 11:48:33our label column. So we just remove
- 11:48:36that. The rest of the data is our
- 11:48:38features. So that's what that's what
- 11:48:40this X is, right? It's all of our
- 11:48:42feature data. All of our columns that is
- 11:48:44not the label column essentially is what
- 11:48:46that's doing. And then Y is our label
- 11:48:49column from our original data.
- 11:48:54Right? Y is our label column. And so
- 11:48:58this this will um contain all of our
- 11:49:00labels which is the median house value.
- 11:49:04X X X contains every column but the one
- 11:49:07we're gonna so we we ultimately decide
- 11:49:10that but X contains um X is everything
- 11:49:15that is not our dependent variable which
- 11:49:18is what we're predicting. So we're
- 11:49:20removing what we are trying to predict
- 11:49:22from X. X should be everything else.
- 11:49:25That's always how it's going to be. X is
- 11:49:27X is always going to be all of those
- 11:49:29independent variables that we're using
- 11:49:32to predict the median house value. So we
- 11:49:36are going to predict the median house
- 11:49:38value. We need to remove it from X.
- 11:49:42So we're we're taking everything but
- 11:49:44that column.
- 11:49:49So it's the whole data frame. It's the
- 11:49:52whole data frame minus this one column
- 11:49:54with just the dependent variable. Right.
- 11:49:59Exactly right. Removing the dependent
- 11:50:00variable and keeping all the
- 11:50:01independence. That's exactly right.
- 11:50:04Exactly right. So think about it in
- 11:50:07terms of the model. Let's go back to the
- 11:50:08features. Right. Think about it in terms
- 11:50:10of the model. We are trying to predict
- 11:50:13this this value. We're building a model
- 11:50:16to try to predict this. So we are going
- 11:50:19to make sure X is everything but this
- 11:50:24right. So this is actually just Y.
- 11:50:27That's our label. That's our dependent
- 11:50:29variable. Right? That's Y. Everything
- 11:50:32else is belongs to X. Everything else
- 11:50:35belongs to X including all of these.
- 11:50:42Right? We choose this one to be Y
- 11:50:44because we're building a model to
- 11:50:45predict that. That's our label.
- 11:50:49All right. So, we have our we use X and
- 11:50:52Y to do our train test split. So, we
- 11:50:54have our our training features and our
- 11:50:57test features and then our training
- 11:50:58label and test labels here. Um, pretty
- 11:51:01standard there.
- 11:51:04Um, okay. So, this is what's new is if
- 11:51:06we want to do K-fold uh validation, what
- 11:51:10we're going to do is create a kfold
- 11:51:12object. So, we have this kffold from
- 11:51:14scikitlearn that we already imported. we
- 11:51:17are going to create a kfold um where we
- 11:51:22are going to specify how many folds we
- 11:51:24want. So that is the in uh inslits
- 11:51:27parameter as this says um this is going
- 11:51:30to be uh uh in this case we're going to
- 11:51:34do 10 folds. That's pretty standard. So
- 11:51:36I think the typical number of folds that
- 11:51:38I've seen and I've worked with in my in
- 11:51:40my career is usually five or 10.
- 11:51:44Five or 10 folds is the standard.
- 11:51:49Okay. So, we're doing 10 folds in this
- 11:51:51case and we're setting a random state
- 11:51:53because we're going to do shuffling. So,
- 11:51:55in order to produce those folds, we're
- 11:51:57going to shuffle the data first and then
- 11:51:58split it into five folds, right? So,
- 11:52:01this this kfold object is going to
- 11:52:04manage creating these splits for us,
- 11:52:08right? These even splits. I know I I
- 11:52:10didn't draw it even, but um it's going
- 11:52:13to manage these five folds for us and
- 11:52:15it's going to shuffle the data and
- 11:52:17assign them to these different folds and
- 11:52:19we're and then what we're going to do is
- 11:52:21use those to do our training. We're
- 11:52:24going to execute the cross validation
- 11:52:26using this k-fold object.
- 11:52:29Okay, so we create the kffold
- 11:52:32um we initialize our model as well. So,
- 11:52:35of course, in order to train something
- 11:52:38uh in the Kfolds, we're going to need a
- 11:52:39model. In this case, we're using linear
- 11:52:42regression, right? Which is which is the
- 11:52:44model we've been studying so far. So,
- 11:52:46you have a linear regression. Um now,
- 11:52:49look how easy it's going to be in order
- 11:52:51to execute cross validation. All we need
- 11:52:54to do is um all we need to do is create
- 11:52:58a crossfile score function.
- 11:53:02um or I should say use the cross val
- 11:53:04score function from scikitlearn. So we
- 11:53:06use that with the model we want to
- 11:53:08train. So our model goes first. So
- 11:53:11that's the linear regression object.
- 11:53:14Then our data. So our extra our features
- 11:53:17and our label for our training.
- 11:53:20And then um let me skip over this for a
- 11:53:23second. I'll explain what this is in a
- 11:53:24second. Um but then we are using uh the
- 11:53:30cross validation technique is our kfold.
- 11:53:33So this is where our k-fold object goes
- 11:53:35in the CV parameter which is cross
- 11:53:37validation. So what cross validation
- 11:53:40strategy are you using? We're using
- 11:53:42kfold and the kfold we're using is this
- 11:53:44one we defined up here KF. So we're
- 11:53:47putting that right here for this. And
- 11:53:49then um in jobs um allows us to
- 11:53:54parallelize this. So if we set it to
- 11:53:55negative one, that's the that that's the
- 11:53:58default. Um it will do it will actually
- 11:54:01train across the different combinations
- 11:54:03in parallel. Um which speeds it up. So
- 11:54:06you want to you want to keep this to
- 11:54:07negative one if you can. So um now let
- 11:54:12me describe the scoring. So what this
- 11:54:15means is we put in our metric here. Um
- 11:54:20and so you can put mean absolute error,
- 11:54:23you can put in mean squared error. Um
- 11:54:25those are the two that we can use. And
- 11:54:29um the reason we it has a negative in
- 11:54:31front of it is because we want to find
- 11:54:34the one that has the lowest score.
- 11:54:38That's going to be our best model is the
- 11:54:41one that has the lowest score. So we
- 11:54:43take the absolute value.
- 11:54:46I'm sorry. we take the abs the the the
- 11:54:48metric and we take the negative of it um
- 11:54:52because the highest scoring one is going
- 11:54:56to be the closest to zero. Um so it's
- 11:54:59just a we use the we use the negative of
- 11:55:02the of the metric um because on the
- 11:55:05number line like the the highest um
- 11:55:08scoring one should be the least um or I
- 11:55:12should say the maximum negative that we
- 11:55:14can get. that's going to be closest to
- 11:55:16this to zero. So if here's zero, this
- 11:55:18would be like -1 is better than -10,
- 11:55:22right? So something that scores um the
- 11:55:25maximum negative uh absolute error would
- 11:55:29be closest to zero
- 11:55:32and something that has more is going to
- 11:55:34be on this side.
- 11:55:37So this is only the reason we need this
- 11:55:39is only just to keep track of the scores
- 11:55:42of each individual um fold. Okay.
- 11:55:47So the one so the reason we can do that
- 11:55:49is at the end we can kind of see which
- 11:55:51which combination performed the best. Um
- 11:55:55it's going to be the one that has the
- 11:55:57highest uh highest value of the negative
- 11:56:00which is closest to zero.
- 11:56:08That's just a convention.
- 11:56:11Yeah, it's just because um it's because
- 11:56:13the cross validation is looking to
- 11:56:15maximize the metric. So, whatever has
- 11:56:18the best score
- 11:56:20um whatever has the best score is
- 11:56:23considered the best uh performance. Um
- 11:56:25but we are using uh something where
- 11:56:28lower is better. So we we take the
- 11:56:31negative and like the the highest
- 11:56:33negative would be closest to zero,
- 11:56:37right? The highest negative is going to
- 11:56:38be closest to zero.
- 11:56:42So that so it's it's just because like
- 11:56:45we want the lower score to be the best.
- 11:56:49The lowest score should be the best.
- 11:56:52So we take the negative of it. Um and so
- 11:56:56something that is more negative is going
- 11:56:58to be worse. Yeah, that's the reason.
- 11:57:03So something that's down this way is
- 11:57:05going to be worse. Okay, so it runs this
- 11:57:11and what you can see is if we actually
- 11:57:13print this out, if we print out our
- 11:57:15k-fold scores, what we should get is 10
- 11:57:17different scores.
- 11:57:21And you can see um we have 10 different
- 11:57:24uh scores here which are all negative
- 11:57:27because we're taking the negative of the
- 11:57:28absolute of the mean absolute error. Um
- 11:57:32so what we would be looking for here is
- 11:57:37um we want to take the average of these
- 11:57:39scores but take the absolute value of
- 11:57:42them to get the best performance. So
- 11:57:44this is capturing like this is the score
- 11:57:46on the first fold combination. This is
- 11:57:49the score on the second fold
- 11:57:50combination. This is the score on the
- 11:57:52third fold combination and on and on and
- 11:57:55on. And these are the absolute errors.
- 11:57:58Okay, these are the absolute errors. Um
- 11:58:02so if we take a look at computing the uh
- 11:58:05average, which by the way, we don't need
- 11:58:07this import because we're using the
- 11:58:08numpy average. So that's fine. um we can
- 11:58:12take the absolute value of those um and
- 11:58:15take a look at the average MSE
- 11:58:20or sorry MAE. Now I want you to think
- 11:58:24about this this uh average performance.
- 11:58:27So this is our performance right here on
- 11:58:29the cross validation.
- 11:58:31This is our average
- 11:58:33MAE
- 11:58:35across all of our fold combinations. So
- 11:58:37that's a that's an indicator of our
- 11:58:38performance, right? um for the cross
- 11:58:41validation.
- 11:58:44Now, what are the units of our original
- 11:58:49uh the original median value? They're
- 11:58:53already in the thousands, right? So, if
- 11:58:56we go to that feature, they're already
- 11:58:59in these hundreds of thousands. So, this
- 11:59:03is not a very good error. It's it's kind
- 11:59:07of high, right? because it's in this is
- 11:59:1049,000.
- 11:59:12Um that's that's how far away we are in
- 11:59:15absolute value on average is $49,000 um
- 11:59:19dollars on the median value. That's not
- 11:59:22very good. So this score
- 11:59:26this score is
- 11:59:28um not very good. So this model is not
- 11:59:32performing that well and we can see that
- 11:59:34by comparing this error to our actual uh
- 11:59:37data. So this is right around 50,000
- 11:59:42and our median uh house values are in
- 11:59:45the hundreds of thousands. So on average
- 11:59:48we're 50,000 off when we make a
- 11:59:51prediction. That's a significant amount,
- 11:59:54right? It's a significant amount on
- 11:59:56average um when our when our data is in
- 11:59:59about the hundreds of thousands here.
- 12:00:03So we are um we have a significant
- 12:00:06amount of error 50,000 relative to the h
- 12:00:09to our units that our our data is in.
- 12:00:11Right? Um so this score is not very
- 12:00:14good. Um
- 12:00:17and so we see that from the cross
- 12:00:18validation. So look how easy the cross
- 12:00:20val is. Again we just do cross val
- 12:00:22score. We put in our model. We put in
- 12:00:24our data. We put in our cross validation
- 12:00:27uh strategy here which is k-fold and we
- 12:00:30can generate these metrics across all
- 12:00:32the fold combinations. So it's this
- 12:00:35function is taking care of rotating
- 12:00:37those and doing every combo with just
- 12:00:39the 10 different combinations here of
- 12:00:42the of the folds.
- 12:00:4410 different instances where you have
- 12:00:46you know 10 different folds are the ones
- 12:00:48that are left out for evaluation.
- 12:00:50Um so it's managing that for us using
- 12:00:53this data right using this training data
- 12:00:55here. Um and we uh we generate these um
- 12:01:02generate these scores.
- 12:01:06Okay. So that's kfold. It's not hard to
- 12:01:08do. All you have to do is um just use a
- 12:01:12cross file score. And we could change
- 12:01:14this to mean squared error. That's you
- 12:01:17know we could do that too. That'd be
- 12:01:18pretty easy. Um, so that'd be no issue.
- 12:01:23We just happen to be using the absolute
- 12:01:24error here. Of course, we could use
- 12:01:27squared error.
- 12:01:29Were you guys able to get this to run
- 12:01:31kfold scores?
- 12:01:34It produces an array of 10 10 different
- 12:01:36scores, which should make sense because
- 12:01:38those are these are the um we're
- 12:01:41splitting our data into 10 different
- 12:01:43folds,
- 12:01:45right?
- 12:01:4710 different folds and leaving one out
- 12:01:49to do our evaluation on. So the one that
- 12:01:51gets left out every time is what's
- 12:01:53producing these scores. So it's 10
- 12:01:55different ones get left out when we
- 12:01:57rotate through all the combinations.
- 12:02:02And so we average these scores
- 12:02:11and we get this amount of we get about
- 12:02:1350,000 in error on average.
- 12:02:22Um, what do you think would be what do
- 12:02:24you think would be acceptable? So, if
- 12:02:25our if we're predicting the price, like
- 12:02:27if we're a real estate agent and we're
- 12:02:29predicting these prices and they
- 12:02:32typically are
- 12:02:34Yeah, close to zero would be great.
- 12:02:36That'd be fantastic. Closer to zero
- 12:02:38would be better. The average is um
- 12:02:40206,000.
- 12:02:43So 50,000 is a decent percentage of
- 12:02:46that. Um so you know you can compute it
- 12:02:49as a percentage right? So 50,000 is a
- 12:02:52decent percentage of that. Um probably
- 12:02:55you want this to be less than 20,000
- 12:02:58would be about 10% error. 20,000
- 12:03:03right? So maybe like 30,000 somewhere in
- 12:03:07there.
- 12:03:12Yeah. 10% would be 5% error. 10,000
- 12:03:15would be 5% error. That's true. That's
- 12:03:16true. So that would be that would be
- 12:03:18much better. So being closer to zero,
- 12:03:20like the smaller the better, of course.
- 12:03:23Of course. Um but yeah, I would say an
- 12:03:26acceptable percentage of error is
- 12:03:28probably 20%.
- 12:03:31Probably 20%, which would be um like
- 12:03:3440,000 or less would probably be
- 12:03:36acceptable.
- 12:03:39Usually when we usually when you build
- 12:03:41models um 80% accuracy is usually uh
- 12:03:46considered decent.
- 12:03:48Usually considered decent
- 12:03:5180%. So I'd say 40,000 or less would be
- 12:03:55kind of ideal.
- 12:04:00Does that make sense to answer the
- 12:04:03question?
- 12:04:05That's a good question. What value is
- 12:04:06acceptable? I think probably less than
- 12:04:0940,000 would be ideal. That's right
- 12:04:12around 20% error.
- 12:04:16All right, so that's kfold. Um let's do
- 12:04:19just a regular hold out now. So this is
- 12:04:21just using our training and test data.
- 12:04:23Um doing model.fit and calculating an
- 12:04:26MSE on the test data. So this is this is
- 12:04:29just the um hold out strategy here where
- 12:04:32we just have um this is less robust but
- 12:04:35it's a lot quicker to do and easier to
- 12:04:37set up. Right? So um this is using the
- 12:04:41hold out strategy. So just a regular
- 12:04:46um train test split.
- 12:04:50Are we going to rebuild the model? No,
- 12:04:52not necessarily. There's some things we
- 12:04:54could do most likely. And like one thing
- 12:04:58we did not do was scale our features.
- 12:05:01Remember I said that's a pretty
- 12:05:02important thing to do is to scale our
- 12:05:05features. We did not do that. So that
- 12:05:07would be an enhancement to this that
- 12:05:08we're going to So I I actually do think
- 12:05:10we'll do that later. Yes. So I think we
- 12:05:13will actually do that now that I'm
- 12:05:14thinking about it. Yes. One of the
- 12:05:17things we can do is scale these features
- 12:05:19using like a minmax scaler, a standard
- 12:05:21scaler. that's actually going to help us
- 12:05:23um that's going to help us do better
- 12:05:26predictions.
- 12:05:30So that that's one thing we could do. Um
- 12:05:33but yeah, we will we'll try to see if we
- 12:05:36can get better.
- 12:05:38It should help it. Yeah, usually you
- 12:05:40want to scale you want to scale the
- 12:05:42data. That's something we didn't do in
- 12:05:43our preparation step. We did a lot of
- 12:05:45the things we should do. We removed nles
- 12:05:47and we did one hot encoding to the
- 12:05:49proximity feature like this one. Um
- 12:05:53those are good to do but we didn't scale
- 12:05:56any of these other we didn't scale any
- 12:05:58of the features right we didn't scale
- 12:06:00any of them. Um it you it will have an
- 12:06:04effect. It usually when we scale it
- 12:06:06it'll be a better model.
- 12:06:09It'll it'll learn a little bit better if
- 12:06:11we can scale the data. Um so that way
- 12:06:15like these
- 12:06:18um like ages aren't you know drastically
- 12:06:21different than like in scale then total
- 12:06:23bedrooms or income
- 12:06:26uh those kind of things. So we usually
- 12:06:28want these to be in a similar scale
- 12:06:30range.
- 12:06:33So we'll we will I think we'll scale
- 12:06:35them coming up in a bit and it should
- 12:06:38help the model.
- 12:06:41We've talked about that before, right?
- 12:06:42scaling usually is a good idea to do
- 12:06:44when you're prepping your data for
- 12:06:46modeling.
- 12:06:49No, you want to you want to scale your
- 12:06:51test data as well. You're going to do
- 12:06:53both. You're going to scale your
- 12:06:54training data, you're going to scale it.
- 12:06:56So, that's actually a good point you
- 12:06:58bring up is any transformations you do
- 12:07:00on your training to build your model,
- 12:07:02you should also do on your test set so
- 12:07:05you get an applesto apples comparison.
- 12:07:07You should always do the same
- 12:07:09transformations.
- 12:07:11Yes.
- 12:07:12Would scaling data impact K? Yeah, it
- 12:07:14could it could make it better. It could
- 12:07:16uh yeah, it should impact it. We should
- 12:07:18get a better model. So when we do the
- 12:07:20different folds, we'll get different
- 12:07:22we'll get better scores. Yeah, it it
- 12:07:24will impact
- 12:07:28uh yeah, if they're so that's a good
- 12:07:31point. If they're going to use our
- 12:07:32model, then yes, they have to scale the
- 12:07:34data as well. If they're gonna if we
- 12:07:36build the model on the assumption that
- 12:07:37the input is scaled, then yes, they have
- 12:07:40to also scale their data when they're
- 12:07:42using it with our model. That's true.
- 12:07:49I mean, not really. I'll show you why.
- 12:07:52There's something that's actually going
- 12:07:53to make it easier um that that will
- 12:07:56automate doing the scaling for them. So,
- 12:07:58they don't they don't have to do the
- 12:07:59scaling manually. it'll just it'll
- 12:08:02happen automatically when they use the
- 12:08:04model. I'm going to show you something
- 12:08:05that's going to automate that which is
- 12:08:07going to be called a pipeline.
- 12:08:09So that part will be automated and they
- 12:08:11won't have to do that. So it won't be
- 12:08:13heavy on the user. No, in theory it is,
- 12:08:17but
- 12:08:18has a really helpful tool to make it
- 12:08:20easy to do that. So I'm going to I'm
- 12:08:22going to show us that um later on in the
- 12:08:24notebook.
- 12:08:27No, the data data is not for a single
- 12:08:29house. It's for like a neighborhood. So
- 12:08:31there's a certain number of households
- 12:08:34in the neighborhood and this is the
- 12:08:36we're predicting the median house value
- 12:08:38of that neighborhood.
- 12:08:41Yeah. So there's a there's certain
- 12:08:42number of households. There's there's
- 12:08:43like an a median income, a population,
- 12:08:46certain number of people that live
- 12:08:47there. Um proximity generally of where
- 12:08:52that location is. It also has a latitude
- 12:08:54and longitude.
- 12:08:58So,
- 12:09:01and a median age in that neighborhood.
- 12:09:03So, yeah, it's not just a single house.
- 12:09:14Okay, let's go back to this was the hold
- 12:09:18out strategy. So, this is a lot simpler.
- 12:09:20This is just model.fit, right? This is
- 12:09:22just model.fit on the training uh data.
- 12:09:25And then we um can predict on the test
- 12:09:27features and generate test predictions.
- 12:09:30And then we can compute our error on
- 12:09:32those um we can compute our error
- 12:09:34amongst the test predictions and our
- 12:09:36test uh label. So that's our useful mean
- 12:09:40squared error function, right? To to
- 12:09:43compute the MSE. Um let's see what the
- 12:09:47MSE is. So MSE is right here.
- 12:09:53Um now what we can do is we can take the
- 12:09:56MSE
- 12:09:58and we can take the square root of it.
- 12:09:59So let's actually do that. Let's um do
- 12:10:02MP. square root of the
- 12:10:05um test
- 12:10:07MSE
- 12:10:11and we get um 67 we get 67,000.
- 12:10:16So that's pretty high on this. So when
- 12:10:19we just now look at the difference of
- 12:10:21that, right? When we just do a train
- 12:10:23test split, um
- 12:10:28when we just do a train test split, we
- 12:10:30get a worse score because it's not as
- 12:10:32it's not as robust, right? We're not
- 12:10:34showing that to many of the other uh
- 12:10:37folds. So we get a lot more error this
- 12:10:40way on the test data.
- 12:10:43So this is um actually worse performance
- 12:10:46just doing the train test split.
- 12:10:50This is a really higher.
- 12:11:02Yeah, we can. We can. I'm going to I'm
- 12:11:04going to show us how to how the scaling
- 12:11:06will be done automatically. Yes, we can.
- 12:11:10Um there's there's a really easy tool to
- 12:11:13do that will scale it automatically.
- 12:11:20It's going to be later in this notebook.
- 12:11:21I'll show us it.
- 12:11:30All right. So, just to recap this, this
- 12:11:32is fitting the model.
- 12:11:34This is fitting the model. This is
- 12:11:36making the predictions, right?
- 12:11:37Model.predict.
- 12:11:39So, this is making the predictions. And
- 12:11:41then this is calculating the error, the
- 12:11:43mean squared error, which is looking at
- 12:11:45our test labels versus our test
- 12:11:47predictions, right? And this is
- 12:11:49computing the distance, the average
- 12:11:51distance away from these values to these
- 12:11:54values,
- 12:11:56right?
- 12:11:57And then we can also compute the R squar
- 12:12:00R R 2 and we see that it's not a very
- 12:12:03good R squared. 65 uh is not a very
- 12:12:06great model
- 12:12:08um because it closer to one would be
- 12:12:10better. So this is still this is not
- 12:12:12very good.
- 12:12:14We know that we knew that from the cross
- 12:12:16file score. But this is just doing um
- 12:12:18this is just doing a hold out uh where
- 12:12:21we do a train and test split, right? So
- 12:12:23it's a little bit simpler, but it's not
- 12:12:25quite as robust. Um
- 12:12:28it's not quite as robust as the cross
- 12:12:30valve, but it works. Um it's, you know,
- 12:12:34we can do hold out. Um,
- 12:12:37we can do hold out uh to to quickly
- 12:12:39evaluate a model and see if we need to
- 12:12:42make any adjustments.
- 12:12:44It's a little bit quicker to run.
- 12:12:50Okay. And any questions on it? Does it
- 12:12:52make sense what we're doing here?
- 12:12:53Model.fit to train it predict to get our
- 12:12:57predictions. Um, this is pretty
- 12:13:00standard, right? To train is the
- 12:13:01model.fit it and then to use the model
- 12:13:03to predict. We predict on the test
- 12:13:05features.
- 12:13:07Um, so this is passing on on all of our
- 12:13:09features into this model to generate
- 12:13:13predictions for every row. That's
- 12:13:15something I also want to point out that
- 12:13:16may be a little bit confusing is this is
- 12:13:18a data frame. So we're passing in a
- 12:13:22bunch of rows of features with columns,
- 12:13:24right? So um, we're passing in a bunch
- 12:13:27of data that looks like this. And what
- 12:13:30we're doing is essentially making a
- 12:13:32prediction for every row. So this will
- 12:13:34generate a prediction. This row will
- 12:13:37generate a prediction. This row will
- 12:13:39generate a prediction. And on and on and
- 12:13:41on. So this this predict will predict
- 12:13:44for every row. And so we end up with
- 12:13:47this collection of predictions here for
- 12:13:49each row. And we're comparing those to
- 12:13:52the labels that we have for those rows
- 12:13:54from our from our supervised learning,
- 12:13:57right? From our data set.
- 12:13:59So that's truly supervised learning,
- 12:14:01right? We have the examples and we're
- 12:14:04comparing those to what our model is
- 12:14:06predicting to to get our performance.
- 12:14:16All right.
- 12:14:18So let's uh let's try the other just so
- 12:14:21you can see it. The leave one out. Now
- 12:14:23the leave one out cross validation is
- 12:14:25going to actually work the same way
- 12:14:26where we put in the leave one out um
- 12:14:30strategy inside of the cross file score.
- 12:14:32Now here we don't need to specify how
- 12:14:34many folds there are because we know how
- 12:14:37many there going to be. It's going to be
- 12:14:38the number of data points, right? So
- 12:14:41which is actually going to be quite
- 12:14:42large because there's 20,000 rows. So
- 12:14:45this is going to be extremely
- 12:14:48uh extremely um intensive because we are
- 12:14:53doing um you know 20,000 examples and
- 12:14:57leaving one example out to be our
- 12:15:00validation and then um doing that across
- 12:15:03every 20,000 uh examples.
- 12:15:06So we could do it though just to see how
- 12:15:08it works. Um we have this again leave
- 12:15:10one out. We generate our cross file
- 12:15:12score from our model our data and then
- 12:15:16same scoring that we had before and but
- 12:15:18this time we change our cross file to be
- 12:15:20instead of our k-fold object we have our
- 12:15:23leave one out object which is this
- 12:15:26um and then we could run this. We can
- 12:15:29compute our average uh across the all
- 12:15:32the folds. Now this is going to be a lot
- 12:15:34bigger of an array. It's going to be a
- 12:15:3620,000 size array and we're going to
- 12:15:39compute the average across it.
- 12:15:43So, let's do that. It's going to take a
- 12:15:45moment because there's lots. So, if you
- 12:15:47notice it when you run, it's going to
- 12:15:48take a little bit of time to run because
- 12:15:50it's running across all 20,000 examples
- 12:15:53and leaving one out. So, you have 20,000
- 12:15:57and then one left out to uh test
- 12:16:00against. So, it's quite intensive. You
- 12:16:03can see it's taking a lot more time.
- 12:16:11It's still running. It's taking a while.
- 12:16:23Okay, just let that run. Still running.
- 12:16:28So, if you guys try running this, it's
- 12:16:30going to take a little bit of time.
- 12:16:31Hopefully that makes sense why it's
- 12:16:33taking so long, right? It's because it's
- 12:16:35instead of doing 10 folds, it's it's
- 12:16:39putting every data point but one is the
- 12:16:41training set and then iterating through
- 12:16:43all 20,000 points.
- 12:16:47This takes a while to do.
- 12:17:04Let's see what our
- 12:17:08RAM our memory is a little increased.
- 12:17:20Okay,
- 12:17:22still running. That's okay. I'll let it
- 12:17:24run.
- 12:17:27Come back when it's finished.
- 12:17:34Yeah, exactly. This is a This is for
- 12:17:37This is giving us a performance
- 12:17:38evaluation. This is like the average
- 12:17:41error across all of our uh different
- 12:17:43folds. Um now this is the extreme case
- 12:17:46where we have the number of folds equals
- 12:17:48the number of points.
- 12:17:51Right? So it's an extreme case but yes
- 12:17:53it's just like kfold. It's giving us
- 12:17:55that performance estimate.
- 12:18:04Okay. It's about the same right. This is
- 12:18:07still around 50,000.
- 12:18:10Not much difference, right? Still right
- 12:18:12around there. But look how much longer
- 12:18:15it took. That took 2 minutes to run. The
- 12:18:17other one was pretty instant, right? So
- 12:18:19this this took about 2 minutes to run.
- 12:18:22So um definitely uh
- 12:18:28yeah, definitely don't want to run this
- 12:18:31uh too often. I think that it's
- 12:18:34generally preferred to do k-fold if
- 12:18:36you're going to do cross validation.
- 12:18:37Generally want to do k-fold or just the
- 12:18:39regular hold out train test split. Uh
- 12:18:42generally better than doing leave one
- 12:18:44out. It's just going to take too long
- 12:18:46and um it results in about the same kind
- 12:18:49of score as the kfold.
- 12:19:01Okay,
- 12:19:04any questions about um the cross
- 12:19:07validation that we just did.
- 12:19:28Okay,
- 12:19:29good. And as it says here that the
- 12:19:32stratified kfold is usually used for
- 12:19:34classification. Again, we're not doing
- 12:19:35classification yet. That's in going to
- 12:19:36be in lesson four. So, we don't need to
- 12:19:39worry too much about that. Just for
- 12:19:40regression, um regular kfold is
- 12:19:43preferred, right? Because we don't need
- 12:19:45to um worry about distributing
- 12:19:48categories amongst our folds uh in any
- 12:19:51regression problems.
- 12:19:53And as we see the error is kind of high.
- 12:19:56Um there's going to be some things we
- 12:19:57can do to improve that which will be uh
- 12:20:00later on we'll learn about some more
- 12:20:02advanced models. This signals that the
- 12:20:05performance is bad. We probably need a
- 12:20:07more complex model. Um one thing we
- 12:20:10could try before we try a complex model
- 12:20:12is to do scaling. We will try to do
- 12:20:15scaling. I'm going to show us how we can
- 12:20:17do that coming up um in a in a nice
- 12:20:20streamlined fashion. Um, but uh outside
- 12:20:24of that, if we still had bad
- 12:20:26performance, we would likely need to use
- 12:20:27a more advanced model. And we'll learn
- 12:20:30about more advanced models uh in the
- 12:20:33next lesson. And what's great is some of
- 12:20:35those advanced models can actually be
- 12:20:37used for regression. So they have
- 12:20:39variations that can be used for both
- 12:20:41classification and regression, which is
- 12:20:43pretty cool. So I'll point those out
- 12:20:45when we get to them. Um, okay.
- 12:20:50So what I want to talk about now is a
- 12:20:53way we can combat overfitting. So if we
- 12:20:56have overfitting which remember that is
- 12:20:58the case where the uh the we see good
- 12:21:03performance on the training data but
- 12:21:05then um it doesn't generalize over to
- 12:21:07the test data. We get poor performance
- 12:21:09on the test data. Um there's there's a
- 12:21:12drop off there. Um that would signal
- 12:21:15overfitting.
- 12:21:18overfitting
- 12:21:19and one way of um combating overfitting
- 12:21:23is to do something called regularization
- 12:21:26which we're going to talk about next. So
- 12:21:29the key idea in regularization
- 12:21:33is to
- 12:21:35change our uh the change the way we
- 12:21:39train. Essentially, what we're going to
- 12:21:41do is modify our training
- 12:21:46uh error function or sometimes called
- 12:21:50the objective function or loss function.
- 12:21:53We're going to change that to add a
- 12:21:56penalty to penalize excessive complex
- 12:22:00complexity. Essentially the the way that
- 12:22:03we're going to penalize is by making
- 12:22:05sure the size of the coefficients
- 12:22:08doesn't grow too much which should
- 12:22:11mitigate overfitting because remember in
- 12:22:14linear regression what we are learning
- 12:22:16are the coefficients right we're
- 12:22:18learning the beta 0 the beta 1 the beta
- 12:22:212 and on and on however many betas there
- 12:22:24are beta n we're learning all of those
- 12:22:26guys um through the regression error
- 12:22:29function we're trying to minimize that
- 12:22:31error function. That's how it trains. We
- 12:22:33talked about that on Monday.
- 12:22:36Um so what we're going to do is um
- 12:22:41basically penalize the these guys
- 12:22:44growing too big and making sure we kind
- 12:22:47of keep them small so that no one
- 12:22:51coefficient has a dominant uh effect on
- 12:22:54the model. And this should help with
- 12:22:56overfitting and complexity. should make
- 12:22:58the model simpler because all the
- 12:23:00coefficients are going to be encouraged
- 12:23:02to be smaller. They're not going to grow
- 12:23:04too big. Um and this this has the effect
- 12:23:07of making the model so basically make
- 12:23:10the model simpler.
- 12:23:14Make the model simpler is what these
- 12:23:17regularization techniques are
- 12:23:18essentially trying to achieve is is
- 12:23:20remove complexity, make them a little
- 12:23:22bit simpler, make these coefficients
- 12:23:24smaller so that you can generalize a bit
- 12:23:26better and and prevent overfitting. So
- 12:23:30we want to prevent
- 12:23:33uh overfitting,
- 12:23:35right, is what we want to do. Um so
- 12:23:38there's going to be a penalty and I'll
- 12:23:40show you where that penalty gets added
- 12:23:42and kind of what it looks like.
- 12:23:44Um but uh to control the level of that
- 12:23:48penalty we are actually going to
- 12:23:49introduce another parameter to our model
- 12:23:52um called alpha.
- 12:23:55Alpha is going to scale the penalty. So
- 12:23:58if alpha is really high that imposes a
- 12:24:02stronger penalty on the coefficients um
- 12:24:05which will make the model a lot simpler.
- 12:24:08So the higher the alpha the simpler the
- 12:24:10model we will get and we the the risk
- 12:24:14with that is we actually underfit. So if
- 12:24:17alpha is too big we may underfit the
- 12:24:20training data
- 12:24:22um a bit too much because it will make
- 12:24:24the model way too simple. Um and I again
- 12:24:27I'll show you what this means
- 12:24:28mathematically in a moment. Um but on
- 12:24:31the other hand if we have a lower alpha
- 12:24:34this will have a lower penalty. it's a
- 12:24:37weaker penalty term and that'll lead to
- 12:24:40a model that is um a bit more complex.
- 12:24:44Um which could um risk some level of
- 12:24:47overfitting. Um so there's so there's
- 12:24:51still the risk of overfitting if you
- 12:24:53have a low alpha. And of course if alpha
- 12:24:55goes all the way to zero, there's no
- 12:24:56penalty at all. So you're back to your
- 12:24:59original linear regression. Um which
- 12:25:02could risk a lot of overfitting,
- 12:25:04right? So you you generally want to pick
- 12:25:07an alpha um effectively and actually
- 12:25:10we're going to see h what's the best way
- 12:25:12to pick alpha. Um we're actually going
- 12:25:14to learn how to do that. I'm going to
- 12:25:15show us how doing some tuning techniques
- 12:25:18to pick what alpha should be. Um but um
- 12:25:23a a pretty industry standard alpha that
- 12:25:25most people default to is alpha equals
- 12:25:28to one. So just just one which signals
- 12:25:32that there should be some penalty. we
- 12:25:34just have alpha equal to one is a
- 12:25:35standard penalty. We don't want it to be
- 12:25:37too high. We don't want it to be too
- 12:25:39low. Like we don't want it to be a
- 12:25:40fraction. Um but a penalty of one is
- 12:25:43usually uh good enough.
- 12:25:47Okay, I'm going to show you where that
- 12:25:49comes into play in a moment.
- 12:25:52Um but the whole purpose of doing this
- 12:25:54is to mitigate overfitting, right? Um
- 12:25:57that's what and and doing this penalty
- 12:26:00is is called regularization. So adding
- 12:26:03so going beyond just regular linear
- 12:26:05regression adding this extra penalty to
- 12:26:07to the training process um to penalize
- 12:26:11large weights large coefficients
- 12:26:14um is known as regularization.
- 12:26:18Okay. Um and there's two common
- 12:26:21penalties that are added. Um so there's
- 12:26:23actually two different variations on the
- 12:26:25penalty. Um we're going to study both of
- 12:26:27them and um they're they're known as
- 12:26:30lasso. So if you take linear regression
- 12:26:32and add a particular type of penalty,
- 12:26:34it's known as lasso. If you add another
- 12:26:37type of penalty, it's known as ridge
- 12:26:39regression. We're going to study both of
- 12:26:41those and what their differences are.
- 12:26:43But these are the primary two
- 12:26:46uh regularization tech uh models that
- 12:26:49are used um to take a regular both of
- 12:26:52these take regular linear regression and
- 12:26:54just modify the training process a
- 12:26:57little bit in different ways. Two
- 12:26:59different ways. um using that alpha
- 12:27:03um to penalize the terms in slightly
- 12:27:06different mathematical ways. So we're
- 12:27:08going to learn about these two guys.
- 12:27:09Lasso regression there. Both of these
- 12:27:12are just offshoots of linear regression.
- 12:27:14So underlying model is still linear
- 12:27:16regression. It just adds different types
- 12:27:19of penalties to the training process.
- 12:27:22So both of these are still in the family
- 12:27:24of linear regression. In fact, in um in
- 12:27:29scikitlearn, they both come from they
- 12:27:31both are still from the linear model
- 12:27:33family in inside of the linear model
- 12:27:36module, which is where linear regression
- 12:27:37comes from. So there's still linear
- 12:27:39regression. They just have different
- 12:27:42styles of penalties added to them. Um
- 12:27:45which we're going to see.
- 12:27:48Okay, so just to recap that
- 12:27:51regularization is the process of adding
- 12:27:54a penalty to the training to discourage
- 12:27:58complexity. In this case, we're going to
- 12:28:00discourage large coefficients.
- 12:28:04And um this should help prevent
- 12:28:07overfitting.
- 12:28:09And so uh these are going to lead us to
- 12:28:12two different offshoots of linear
- 12:28:13regression that have two different
- 12:28:15penalties.
- 12:28:16lasso and ridge regression, which we're
- 12:28:18going to uh study next,
- 12:28:21but they they function the same way as
- 12:28:23linear regression. They will just have
- 12:28:26different penalty terms added onto their
- 12:28:28training process um to discourage
- 12:28:32uh discourage um again those large
- 12:28:35weights.
- 12:28:40Okay, any questions about regularization
- 12:28:42before we first look at our we're going
- 12:28:43to look at our first uh variation on on
- 12:28:47our first regularization technique which
- 12:28:48is going to be called lasso regression.
- 12:29:08Okay, let's look at lasso regression. So
- 12:29:10what is lasso regression? It's actually
- 12:29:13lasso is short for least absolute
- 12:29:16shrinkage and selection operator
- 12:29:18regression. Um and this will function by
- 12:29:23adding a particular penalty to the
- 12:29:27linear regression model. So again it's
- 12:29:29based on linear regression. That's the
- 12:29:30underlying model. It's just that during
- 12:29:33the training process we are going to um
- 12:29:36add a penalty which has the effect of
- 12:29:41shrinkage of the weights. That's why
- 12:29:43it's called shrinkage. It encourages
- 12:29:45smaller weights through that penalty and
- 12:29:48it also will shrink some of them so much
- 12:29:51that they'll become zero and so it has
- 12:29:53has an effect of kind of selection which
- 12:29:56means that some of them get wiped out to
- 12:29:58zero
- 12:30:00and this means that whatever is left
- 12:30:02over is kind of what's selected as our
- 12:30:05features because the other ones will
- 12:30:08have zero weight applied to them. So
- 12:30:10this penalty will really favor small
- 12:30:14weights um and penalize really large
- 12:30:18weights. In fact, it will favor small
- 12:30:20weights so much that some of them will
- 12:30:22actually um be shrunk to zero um during
- 12:30:26the training process. And the ones that
- 12:30:28are left over are the ones that um are
- 12:30:32the ones that are what we call selected
- 12:30:35because they are the ones that remain in
- 12:30:37in the training um after the other ones
- 12:30:40get uh coefficients of zero. Um now when
- 12:30:44you make some of the coefficients zero,
- 12:30:46you are inherently making the model
- 12:30:48simpler, right? There's less features
- 12:30:51involved in the prediction that or less
- 12:30:53features that have an effect on the
- 12:30:55prediction. So this definitely makes the
- 12:30:57model simpler. This lasso, this
- 12:31:00shrinkage and selection uh process makes
- 12:31:04makes the model simpler for sure. Um
- 12:31:08and this is supposed to reduce
- 12:31:10overfitting, right? If you make the
- 12:31:11model simpler, it's not as complex. It
- 12:31:14has less of a chance of memorizing
- 12:31:16training data and not generalizing over
- 12:31:19to test data. So our whole goal with uh
- 12:31:23regularization is to make our model
- 12:31:25better at generalization right over to
- 12:31:28test data from the original training
- 12:31:30data.
- 12:31:32Um so how does this happen? We have to
- 12:31:35go back to the
- 12:31:38uh training process. If you guys
- 12:31:40remember I I wrote out this equation a
- 12:31:42little bit earlier which is the
- 12:31:44distance. This is the sum of squared
- 12:31:46distance between our labels and our
- 12:31:48prediction.
- 12:31:50This is basically the mean squared error
- 12:31:52uh calculation that we're trying to
- 12:31:54reduce when we build our model using the
- 12:31:56training data. Um so this is just in
- 12:31:58standard linear regression. This is the
- 12:32:01um uh sum of squares uh distance right
- 12:32:05so this is this is what the model is
- 12:32:07trying to minimize when it learns these
- 12:32:09coefficients.
- 12:32:11So when it learns these coefficients,
- 12:32:13it's trying to minimize this guy.
- 12:32:18Minimize. It's trying to find the betas
- 12:32:21that minimize this quantity.
- 12:32:23Mathematically, that's what it's doing.
- 12:32:25Um, and there's there's a algorithm that
- 12:32:28will discover what the best betas are
- 12:32:30that actually minimize this. That gives
- 12:32:32us the line of best fit, right? That's
- 12:32:34what we've been talking about for
- 12:32:35regression.
- 12:32:37Now in regularization
- 12:32:41here's by the way here is that same
- 12:32:42thing but we've just inserted our model
- 12:32:44for the predictions. This is our model
- 12:32:47just a fancy way of writing down our
- 12:32:49model right it's the beta 0 plus all of
- 12:32:52these betas. So beta 1 x1 plus beta 2 x2
- 12:32:58plus on and on and on. Right? That's
- 12:33:01that's what this uh means if you're
- 12:33:03unfamiliar with the sigma notation. It
- 12:33:05just means sum. So it's the sum of all
- 12:33:07these guys or this term. Um so this is
- 12:33:12this here is just a regular linear
- 12:33:14regression
- 12:33:18uh training regular linear regression
- 12:33:21training. So we the training process
- 12:33:24solves for these parameters right it
- 12:33:27solves for these weights. We discover
- 12:33:29what those are by minimizing this
- 12:33:31quantity. That's the whole training
- 12:33:33process. Um but when we do lasso
- 12:33:37we add a penalty which is this
- 12:33:43here is our penalty.
- 12:33:47So basically um we take our linear
- 12:33:51regression training which is this and we
- 12:33:54add on a penalty which is this and you
- 12:33:57can see exactly what this penalty when
- 12:34:00when you minimize this penalty it's when
- 12:34:03these weights are small. So this
- 12:34:05encourages
- 12:34:06So minimizing this quantity encourages
- 12:34:10small weights
- 12:34:13encourages small betas
- 12:34:17beta I
- 12:34:19right you or in this case beta j sorry
- 12:34:23this encourages small beta js uh because
- 12:34:26we want this thing to be minimized
- 12:34:31minimize
- 12:34:33So um what's going to make this minimal
- 12:34:36is of course the line of best fit and
- 12:34:38small weights right are going to make
- 12:34:40are going to bring this error down the
- 12:34:43most.
- 12:34:46So um and here's our alpha right here's
- 12:34:48our alpha. So you can encourage a higher
- 12:34:51penalty with a larger alpha or a lower
- 12:34:53penalty. If alpha equals zero
- 12:34:56what happens to that term? It just goes
- 12:34:59away. So if alpha equals zero, there's
- 12:35:01no penalty and we're back to uh we're
- 12:35:05back to regular linear regression.
- 12:35:09We just have regular linear regression
- 12:35:11because we have no penalty at that point
- 12:35:12when alpha equals zero. So the smaller
- 12:35:15alpha is, the less penalty we're
- 12:35:19enforcing in in the regularization.
- 12:35:22Okay.
- 12:35:24Now what happens is in reality when you
- 12:35:27train with lasso. So this is lasso is
- 12:35:30this particular penalty. This is called
- 12:35:31the lasso penalty
- 12:35:34or sometimes um people call this the L1
- 12:35:37penalty.
- 12:35:39Um L1 just comes from the fact that this
- 12:35:42is the first power or absolute value. Um
- 12:35:46so it's not a squared penalty. It's a
- 12:35:48single uh single power penalty
- 12:35:51um there. But when you add this lasso
- 12:35:55penalty, what can happen is it c it does
- 12:35:58because the because you're minimizing
- 12:36:00this, it does encourage some of these
- 12:36:02weights to become zero.
- 12:36:05So some if you're really trying to get
- 12:36:08the lowest quantity of this,
- 12:36:11the lower the better.
- 12:36:14What makes this thing lower is of course
- 12:36:17if some of these go away. If some of
- 12:36:18these go to zero then that of course
- 12:36:21will lower this as much as we as much as
- 12:36:23possible. Right? So what happens during
- 12:36:25the training is some of these
- 12:36:27coefficients actually they're encouraged
- 12:36:29to be small because of this penalty. But
- 12:36:32some of them will actually becomes will
- 12:36:35actually become zero um in order to get
- 12:36:38the best model the best fit. Some of
- 12:36:40these will actually get so small that
- 12:36:42they'll basically become zero. And that
- 12:36:44means that that that feature basically
- 12:36:47has no effect anymore. It's it's been
- 12:36:51the model has been simplified, right?
- 12:36:53That feature no longer really has an
- 12:36:54effect.
- 12:37:00So just to call out the alpha again um
- 12:37:02if alpha zero some co uh basically you
- 12:37:06have your linear regression you're back
- 12:37:08to linear regression because alpha 0 is
- 12:37:10just wiping this out and you're back to
- 12:37:12linear regression. Um if alpha is
- 12:37:14infinity now if alpha is infinity that's
- 12:37:16an extreme. So if alpha is infinity the
- 12:37:19only way to make this minimize is if all
- 12:37:21your coefficients are zero. If every
- 12:37:23beta is zero, then this will lower the
- 12:37:25the error as as much as possible. So you
- 12:37:28basically have no model. So if all
- 12:37:31coefficients are zero, you have no model
- 12:37:32and that's useless. So you don't want
- 12:37:35your penalty, you don't want your alpha
- 12:37:37to be huge is what this is saying. You
- 12:37:40also don't want your alpha to be small.
- 12:37:41You're basically back to linear
- 12:37:42regression. So you want something in
- 12:37:44between. Um and the typical typical
- 12:37:48value is alpha equals 1.
- 12:37:51typical is alpha equals 1
- 12:37:55to have some level of penalty there. So
- 12:37:58just a regular kind of regular penalty
- 12:38:00term.
- 12:38:08But we are actually going to have a way
- 12:38:10to test and evaluate which alphas are
- 12:38:12the best.
- 12:38:27Um,
- 12:38:28basically you can, yeah, you can have a
- 12:38:31you can have a penalty that's close to
- 12:38:33zero. You can get rid of this if just a
- 12:38:36regular linear regression performs
- 12:38:38pretty well. You can basically have no
- 12:38:40penalty in that case.
- 12:38:42Yeah. So near zero or like it could be
- 12:38:46that adding a little bit of penalty
- 12:38:48actually helps the overfitting and it
- 12:38:50could be really small. One thing that
- 12:38:52we're basically going to do is have a
- 12:38:54strategy to try out different alphas,
- 12:39:00try different alphas
- 12:39:03and evaluate performance
- 12:39:07and then we can decide which. So that's
- 12:39:09what we're going to do is have a
- 12:39:11strategy to just plug in different
- 12:39:12alphas, generate the like train the
- 12:39:15model, and then see what its performance
- 12:39:17is and see if those alphas are good.
- 12:39:19What what which alpha is the best? We
- 12:39:22can evaluate that
- 12:39:24because we can train the model and see
- 12:39:26what it performance is,
- 12:39:31right?
- 12:39:34Yeah. Yeah. So we'll do that. We'll
- 12:39:35practice that.
- 12:39:44Okay, great. Any other questions about
- 12:39:46this lasso regression? So, remember this
- 12:39:48is linear regression here. This is the
- 12:39:50this is how you're training to find the
- 12:39:52betas in linear regression. So, this is
- 12:39:55just linear regression uh um training
- 12:40:00function there.
- 12:40:02We're adding a penalty which is this is
- 12:40:04the lasso penalty
- 12:40:07lasso penalty there right we're adding
- 12:40:09that this is known as regularization
- 12:40:12and the goal of regularization is to
- 12:40:15prevent overfitting so you add a penalty
- 12:40:18here this makes the model simpler which
- 12:40:21prevents overfitting it helps you
- 12:40:23generalize better when it's simpler Any
- 12:40:36questions conceptually on this? We're
- 12:40:38going to do a code example with it
- 12:40:39coming up, but any questions on this?
- 12:40:57Uh yeah, you you so that's the thing,
- 12:40:59Ronald, is you may be willing to
- 12:41:01sacrifice some accuracy in order to
- 12:41:04generalize to unseen data because
- 12:41:06remember that's what we're really trying
- 12:41:08to get after is we may be willing to
- 12:41:10sacrifice some accuracy on this training
- 12:41:12data in order to have it perform better
- 12:41:14on the test data, right? We may be
- 12:41:17willing to do that. That's a willing
- 12:41:19that's an okay sacrifice
- 12:41:22as long like if if it generalizes
- 12:41:24better. That's what we want. That's what
- 12:41:27we're trying to do here is add a
- 12:41:29penalty, make the model simpler and help
- 12:41:33it generalize better to new and unseen
- 12:41:36data. Right?
- 12:41:39That's that picture I've been using with
- 12:41:41the with the um train and test split.
- 12:41:45Where is the square?
- 12:41:47So in the model there's no square. So
- 12:41:50remember the model is the model is this
- 12:41:56um equation uh that has no squares in
- 12:41:58it. Right? It's beta 0 plus beta 1 x1
- 12:42:03plus beta 2 x2 plus beta n xn.
- 12:42:09That's the that's the linear regression
- 12:42:11model. This is the now this this is the
- 12:42:15model but this is the equation that
- 12:42:18helps us train and find the betas. This
- 12:42:21is how this is what we find the betas
- 12:42:23with. So we'll continue. Um we were
- 12:42:27talking about the lasso regression which
- 12:42:30uh adds it takes linear regression right
- 12:42:33which is this optimization and adds in a
- 12:42:36penalty um scaled by the alpha. Um, and
- 12:42:40what that does in order to minimize this
- 12:42:43whole thing, it encourages these to be
- 12:42:46small uh as possible. Um, which makes
- 12:42:49the model simpler, right? The weights
- 12:42:52don't get overly big and complex. Um,
- 12:42:55they they tend to stay small. In fact,
- 12:42:57some of them can even go all the way to
- 12:42:58zero. Um, which makes the model even
- 12:43:01more simpler,
- 12:43:03right? Um, so let's practice uh using it
- 12:43:07in code. It's actually really easy to
- 12:43:08use. It's going to be essentially the
- 12:43:10same uh style and and code as linear
- 12:43:14regression except we are um just going
- 12:43:17to have to uh put in our alpha parameter
- 12:43:21um when we use the lasso. So here we are
- 12:43:26um from the linear model family right
- 12:43:30which makes sense. It's a linear
- 12:43:31regression offshoot that has this
- 12:43:33penalty in it during the training. um we
- 12:43:35are grabbing our lasso regression. Um it
- 12:43:38also has a version of the lasso that
- 12:43:42we're going to take a look at that is
- 12:43:43used for cross validation which is
- 12:43:46really um convenient as well. So it has
- 12:43:49a cross validation lasso which is a
- 12:43:51really convenient um combination of
- 12:43:54basically cross val score and lasso um
- 12:43:57all in one. So it actually is really
- 12:43:59nice to use that way. Um so we'll take a
- 12:44:02look at that example. Um, but we are
- 12:44:05importing it. The main thing is going to
- 12:44:07be the lasso model here. Um, we're going
- 12:44:09to be using a different data set for
- 12:44:11this one. So, not the ocean uh data, but
- 12:44:14this hitters data, which is a baseball
- 12:44:16data set. Um, so it has 322 rows um with
- 12:44:2120 different columns and it looks like
- 12:44:23this. So, you want to download that one.
- 12:44:26Um, hopefully you guys have access to
- 12:44:28that one.
- 12:44:31Um,
- 12:44:33so I will upload it into
- 12:44:36this.
- 12:44:39So give me a moment.
- 12:44:45There's that. And then we can run this.
- 12:44:49Okay. So we are displaying the data and
- 12:44:52so it has um the the hitters names and
- 12:44:56then it has a bunch of different
- 12:44:57statistics. These are all baseball
- 12:44:58statistics.
- 12:45:00Um, if you're unfamiliar with with them,
- 12:45:01that's okay. It's not a big deal. Um,
- 12:45:04but just different baseball stats here.
- 12:45:09Okay. Were you guys able to load that?
- 12:45:11Um, if you're following along, were you
- 12:45:12able to load that? You should have
- 12:45:14access to this data. The hitters CSV.
- 12:45:18This is the one we're going to use for
- 12:45:19the lasso model
- 12:45:21to build a lasso model.
- 12:45:32Yeah.
- 12:45:39Okay. Able to load that one. Perfect.
- 12:45:42Okay. So, able to load that one. Um, and
- 12:45:45we take a look at the the head. Um, so
- 12:45:48we're actually going to uh drop this
- 12:45:51unnamed column because we don't care
- 12:45:53about their name. it's actually just the
- 12:45:55batter's name, which is not going to be
- 12:45:57useful in modeling. Um, so and remember
- 12:46:00that's generally true like an ID, a user
- 12:46:03ID, like a customer ID, a name, that's
- 12:46:06usually not going to be useful in any
- 12:46:08kind of modeling. So we're actually just
- 12:46:09going to drop that uh column and we're
- 12:46:12going to do it in place.
- 12:46:14And access equals 1 means we're dropping
- 12:46:16that column. Um, so we're going to drop
- 12:46:19that and we should no longer have that
- 12:46:21column. And we have all of these guys
- 12:46:23now. So you want to run that. This will
- 12:46:25drop that. Um this will drop drops the
- 12:46:30column in place.
- 12:46:33Um and now we can see we have uh all we
- 12:46:37have this data where um we have this
- 12:46:41data where it's now removed. So this
- 12:46:44that column is now gone and now we have
- 12:46:46these guys. Um, do you notice anything
- 12:46:50about this
- 12:46:52from the info?
- 12:46:58Looks like we have a couple categorical
- 12:46:59features, a few of them, league and
- 12:47:02division
- 12:47:04and new league. What do you notice about
- 12:47:07this?
- 12:47:16Nolles. Yep. So, there's definitely some
- 12:47:17missing data there um that we're going
- 12:47:20to have to deal with.
- 12:47:25So, it looks like there are uh there are
- 12:47:2959.
- 12:47:31Um there are 59. Now, we could we the
- 12:47:35alternative to doing that is we could uh
- 12:47:37we could just use our usual code where
- 12:47:39we do df.is is uh is null.
- 12:47:44Um and then we do uh dot sum to total
- 12:47:49those up across our different columns.
- 12:47:51And we can see that uh we have 59 of
- 12:47:54those in this salary column. That's this
- 12:47:57is the standard way of doing that,
- 12:47:58right?
- 12:48:02Standard way of doing that. And we have
- 12:48:03so we have 59 of those
- 12:48:1259 of those. So, we have to deal with
- 12:48:15it. Any ideas on how to deal with it?
- 12:48:22Any ideas on how to deal with it? This
- 12:48:24is now This is 59 out of 300.
- 12:48:28So,
- 12:48:29what do you guys think about that? It's
- 12:48:31a little bit different than 200 out of
- 12:48:3220,000. A little bit different. We have
- 12:48:36We have about 60 out of 300.
- 12:48:40There's a decent amount.
- 12:48:45Any ideas on how to handle this one?
- 12:48:50Replace. Yep, we should replace. What do
- 12:48:52you think we should replace with?
- 12:48:55It's a float. It's a floating point uh
- 12:48:58value.
- 12:49:09By the way, something unique about this
- 12:49:11that's a little different than usual,
- 12:49:12too, is that the uh this is actually the
- 12:49:16column we're going to use as our label.
- 12:49:18So, we're actually going to predict the
- 12:49:19salary based on the uh based on the um
- 12:49:24rest of the features. So, we definitely
- 12:49:27need to fill in these nles, right?
- 12:49:29because they're actually going to be the
- 12:49:30labels
- 12:49:32and we're missing some labels uh in our
- 12:49:34data. We we definitely need to fill them
- 12:49:37in. Yeah. So, we're going to replace
- 12:49:39them.
- 12:49:49All right. So, we'll we will replace
- 12:49:51them down below. That's going to be
- 12:49:52coming up. Uh we'll come back and
- 12:49:54replace them. um before we replace them,
- 12:49:56we're actually going to get our uh one
- 12:50:00hot encodings for those three different
- 12:50:02um features we have. Um so we do uh get
- 12:50:09dummies with this. Now um of course we
- 12:50:12don't need to do this if we just so this
- 12:50:15code we don't need to do if we just pass
- 12:50:17in the dtype here
- 12:50:21um which is uh then we don't need to do
- 12:50:24this. So we can comment this out.
- 12:50:28Um so now what I want you guys to notice
- 12:50:31is this is the alternative to what we
- 12:50:32did before where we are purposely just
- 12:50:35doing these columns not the whole data
- 12:50:37frame but just doing these columns and
- 12:50:40then we can um concatenate those these
- 12:50:46one hot encodings. We're going to
- 12:50:47concatenate back to the data frame.
- 12:50:50Right? So if we do our dummies and then
- 12:50:53do dummies.info info. Um, we can see
- 12:50:56that we end up with six new columns. And
- 12:50:59in fact, we can do dummies.head
- 12:51:03and take a look at what those are.
- 12:51:05Right? So, these are league A, league,
- 12:51:08uh, N, division E, W, division W, new
- 12:51:11league A, new league N.
- 12:51:14Okay.
- 12:51:16So, um these are uh these are our one
- 12:51:20hot encodings for these three different
- 12:51:22features which are strings, right? So,
- 12:51:24those those features were strings. If
- 12:51:25you go back up, those were our only
- 12:51:27string features we had. So, we've one
- 12:51:30hot encoded those so we can use them in
- 12:51:31our model. What we need to do is just
- 12:51:34concatenate this back to our data frame.
- 12:51:37Right? So, we just need to concatenate
- 12:51:39it back into our data.
- 12:51:48Okay. So, what we're going to do then is
- 12:51:51we're going to grab um we're going to
- 12:51:55grab Y as our salary. And of course,
- 12:51:57we're going to fill nles on that Y
- 12:51:59coming up shortly. But we're going to
- 12:52:01grab Y as our salary and X new. Now
- 12:52:05before building a full X, we're going to
- 12:52:08take a look at X numerical as our data
- 12:52:11frame minus these columns. The reason
- 12:52:14we're doing minus those is because we
- 12:52:17are going to concatenate our dummy
- 12:52:19variables back into this that are going
- 12:52:22to replace these guys. So we're going to
- 12:52:24replace these anyways with our one hot
- 12:52:26encodings. We don't want the strings. So
- 12:52:29we're going to get rid of those. And
- 12:52:32we're also going to get rid of the
- 12:52:33salary because that's going to be part
- 12:52:34of our that's just the label. So we
- 12:52:36don't want that in the X, the eventual
- 12:52:38X.
- 12:52:44Are you guys able to run this one?
- 12:52:48Hope I'm not going too fast. You guys
- 12:52:50able to run this? And does it make
- 12:52:52sense? What we're doing is we're putting
- 12:52:54our labels in Y, which is what we
- 12:52:56usually do. So we're going to predict
- 12:52:58the salary
- 12:53:00and we're getting ready to build the X.
- 12:53:02But before we first want to get rid of
- 12:53:04those one hot the the strings. This is
- 12:53:06getting rid of the strings
- 12:53:09and this is getting rid of the label and
- 12:53:11that's going to be part of our features.
- 12:53:12What we need to do is build our final X
- 12:53:14by concatenating our dummies with this.
- 12:53:17Do you guys see that? We're going to
- 12:53:19concatenate our dummies with this to
- 12:53:21build our final X.
- 12:53:23But but prior to doing that, we need to
- 12:53:26get rid of these string columns here. So
- 12:53:28we're dropping those
- 12:53:31dropping those from the uh data frame uh
- 12:53:35and getting a numerical uh x numerical
- 12:53:38here.
- 12:53:40You can see the columns of that are just
- 12:53:42these guys here. So the the results we
- 12:53:45need to concatenate our we need to
- 12:53:46concatenate this guy um into this and
- 12:53:50then that'll be our full x all of our
- 12:53:52features.
- 12:53:57Okay. So you can see x is going to be
- 12:54:00pd.con
- 12:54:02of this with our dummies.
- 12:54:06This with our dummies. And um
- 12:54:11uh instead of doing this, I'm actually
- 12:54:13going to do the full dummies. We don't
- 12:54:14need to um pick just a few columns.
- 12:54:18We're actually going to do our full
- 12:54:20dummies here and um do x equals 1. Now,
- 12:54:24the reason that's the case is because um
- 12:54:28this will get rid of one column per
- 12:54:32feature and basically assume that if you
- 12:54:35have a if you have a zero, the other one
- 12:54:37should be a one. If you have a one, the
- 12:54:39other one should be a zero. Um so it
- 12:54:42basically makes that assumption because
- 12:54:44we only have two of them. Um so whenever
- 12:54:47there's a one, the other should be zero.
- 12:54:50Um, so you can get away with just having
- 12:54:52these three, but um I think it makes
- 12:54:55more sense to just have to have the full
- 12:54:58dummies,
- 12:55:00but by process of elimination, you can
- 12:55:03get away with just using two of them
- 12:55:04because anytime you have a zero, the
- 12:55:06other one should be the other feature
- 12:55:08would have been would have been a one,
- 12:55:11right? And vice versa, when there's a
- 12:55:13one, the other feature would have been a
- 12:55:14zero.
- 12:55:20So we do that one.
- 12:55:31And you can see all of our uh all of our
- 12:55:34one hot encoding features end up back in
- 12:55:36there.
- 12:55:40So this is the code that I want you guys
- 12:55:42to run. I think it makes more sense. It
- 12:55:45follows along what we've been doing.
- 12:55:47um which will concatenate our dummies
- 12:55:49back to our features here to build out
- 12:55:52our full X. So now X is all of our
- 12:55:54features. Um remember X
- 12:55:58X X contains all of our features
- 12:56:04now.
- 12:56:05So X contains all of our features and so
- 12:56:09we have all of this now.
- 12:56:15Okay. Were you guys able to run this
- 12:56:17one?
- 12:56:19D. We have y, we have x. We still need
- 12:56:23to deal with the nles in y. So that
- 12:56:26something we still need to deal with.
- 12:56:33But hopefully you have this. Now
- 12:56:38all of these are numerical.
- 12:56:41So that should be good with the model.
- 12:56:44That's one thing about X is you should
- 12:56:47you our X should have all numerical
- 12:56:51features, right? Because it's going to
- 12:56:53go into a model to to learn those betas.
- 12:56:56So it needs to have all numerical
- 12:56:58features,
- 12:57:00right? These are going to be all
- 12:57:01numerical, which makes sense. We change
- 12:57:03we did one hot encoding to change all
- 12:57:05those guys to numerical.
- 12:57:15Sorry, I'm scrolling down.
- 12:57:20Okay, we do fill in the nator. Okay.
- 12:57:33Okay.
- 12:57:35Any questions so far? So, we're just
- 12:57:36getting our data ready. We haven't
- 12:57:37applied the lasso yet, but we're just
- 12:57:39doing some prep. Now, hopefully you guys
- 12:57:42recognize th these are some standard
- 12:57:45steps that we're taking when we do our
- 12:57:48modeling. We have to do these data prep
- 12:57:50steps. They're necessary. And so, if it
- 12:57:53seems like it's a lot of work, that's
- 12:57:55because it is. It is work that you do to
- 12:57:59prepare your data to get ready for
- 12:58:01modeling. You have to do that. Okay.
- 12:58:07So, we're doing that here. Um, now we're
- 12:58:10going to do our train test split because
- 12:58:12we're just going to do uh we're going to
- 12:58:14do hold out here. So, we're doing a
- 12:58:16train test split with about with a test
- 12:58:18size of about 0.25. So, again, anywhere
- 12:58:20between 02 to.3 would be okay.
- 12:58:24Um, so uh it's our choice. We could do
- 12:58:2902. We could do 3. We could do anywhere
- 12:58:31in between there. We're doing 0.25.
- 12:58:33That's fine. Um, that's okay. So, we we
- 12:58:37build our train test split right there.
- 12:58:40Um, so pretty pretty simple and we've
- 12:58:42seen that a bunch of times with our X
- 12:58:45and our Y data frames. There we have our
- 12:58:50train test split.
- 12:58:53Okay.
- 12:58:55Um, now what we're going to do is do our
- 12:58:58our scaling. So, we're we didn't do this
- 12:59:01last time, but we're going to do this
- 12:59:02now as uh because we should get in the
- 12:59:05habit of doing that. Um is um we're
- 12:59:10going to um go ahead and scale our
- 12:59:13features and we're going to use the
- 12:59:15standard scaler here uh to do that
- 12:59:19scaling. Okay. Now, we could use minmax
- 12:59:22scaler that's fine, too. We're just
- 12:59:24going to use the standard scaler here.
- 12:59:26Um and remember we are going to uh um
- 12:59:31use the standard scaler from sklearn and
- 12:59:34we're going to transform our features uh
- 12:59:38uh according to our um according to our
- 12:59:43training data. So we have our
- 12:59:47pre-processing standard scaler here. So
- 12:59:49we import that guy and then we um build
- 12:59:53our standard scaler and fit it on the
- 12:59:56training data only on the numerical
- 12:59:59features. Um so that's which is going to
- 13:00:04be uh all of these guys. So we're doing
- 13:00:08the scaling on all of these guys. Now,
- 13:00:10something to note is that we are not
- 13:00:13scaling all of these one hot encodings
- 13:00:16mainly because it doesn't make sense to
- 13:00:18scale those really. They're zero or one.
- 13:00:20They don't need to be scaled, right?
- 13:00:23They're already zero and one. So,
- 13:00:25they're they don't need to be even if we
- 13:00:27were doing minmax scaling, it's going to
- 13:00:29put them between zero and one. It
- 13:00:30wouldn't affect it really, right? So,
- 13:00:33these one hot encoding features, we're
- 13:00:35not going to scale because they're
- 13:00:36they're always going to be zero or one.
- 13:00:39There's no need to scale them really.
- 13:00:42Um, but we're going to scale all the
- 13:00:44other features here that are floats.
- 13:00:46So that's these guys here. These
- 13:00:50numerical features we're going to scale.
- 13:00:53Okay.
- 13:00:54Don't need to we don't really need to
- 13:00:56scale the one hot encoding. Uh, it's
- 13:00:58pretty much already scaled.
- 13:01:05Oh, you should change that. Um, go back
- 13:01:08and rerun go back and rerun this. But
- 13:01:10make sure you have your data type as int
- 13:01:13here.
- 13:01:15Make sure you add that in there to
- 13:01:17change that over to integer and rerun
- 13:01:19that and then rerun the rerun the
- 13:01:23concatenation.
- 13:01:24So make sure you run this
- 13:01:27and then u make sure you rerun this and
- 13:01:29rerun the concatenation part which is uh
- 13:01:34this
- 13:01:47Okay. So, we go ahead and fit the um
- 13:01:52scaler to this data and then we're going
- 13:01:54to transform our training features,
- 13:01:56those numerical features. Um we're and
- 13:02:00then we're going to uh transform these
- 13:02:02features uh uh the test features in the
- 13:02:05same way. So we're going to perform the
- 13:02:08same transformation from the scaler on
- 13:02:10the test data. So that's something
- 13:02:12really important I want to note here is
- 13:02:13that we always scale both the training
- 13:02:19and test data. We always scale both. Of
- 13:02:22course, we're going to train the model
- 13:02:23on the training data. Um, but we are
- 13:02:27going to also test it on the testing
- 13:02:30data and it also needs to be scaled
- 13:02:32because our model that we build is going
- 13:02:34to assume scaled features. The
- 13:02:37coefficients that it learns are going to
- 13:02:39be assuming scaled features.
- 13:02:42So we need to also scale our test data
- 13:02:46in the same way. So we're doing that as
- 13:02:49well.
- 13:02:53So, we scale that. And now we have our
- 13:02:56uh training and testing features have
- 13:02:58been scaled.
- 13:03:01No, we haven't replaced. We're going to
- 13:03:02do that. We have not yet. We're going to
- 13:03:05do that coming up in a minute. Yeah, we
- 13:03:08haven't done that. Um, it is it is the
- 13:03:11label. We definitely need to replace
- 13:03:13NLES. We just haven't done it yet
- 13:03:15because it's not in the features and
- 13:03:16we're doing all of our uh uh
- 13:03:18pre-processing to our pre-processing to
- 13:03:20our features.
- 13:03:30Yeah. So, we're definitely we need to
- 13:03:32we're going to in a minute.
- 13:03:36Okay. So, if you look at the data now,
- 13:03:38it's all been scaled. So, these are all
- 13:03:40um zcores. These are all on a much
- 13:03:43better scale now. Um, and these are we
- 13:03:48still have our one hot encoding features
- 13:03:49which are zero or one. So this scaling
- 13:03:53should lead to a better model than if we
- 13:03:56didn't scale. So scaling is really
- 13:03:59important. We can see that here.
- 13:04:06Okay.
- 13:04:08Now, um, let me ask you guys, were you
- 13:04:10able to run the scaling? Are you caught
- 13:04:12up to here? If you're following along,
- 13:04:15were you able to run the scaling?
- 13:04:28Okay, great. Great.
- 13:04:36Awesome.
- 13:04:39Okay. So, uh what we're going to do now
- 13:04:42is replace nulls in the uh replace nles
- 13:04:48by calculating the median of the data.
- 13:04:52So, what I want you to notice is that we
- 13:04:55are taking the NLES now this is um this
- 13:04:59is on purpose is we are purposely taking
- 13:05:02the NLES um out of the median
- 13:05:05calculation. So we're skipping the NLES
- 13:05:07when we compute the median because we
- 13:05:09don't want those NLES to affect the
- 13:05:11median calculation.
- 13:05:13Um so we compute a median salary here
- 13:05:16and then we fill our NLES with the
- 13:05:19median salary um from the training data.
- 13:05:23So this is our choice. This is a choice
- 13:05:26um to use the median and it's also a
- 13:05:30choice to use the training set median
- 13:05:34for both train and test. What we could
- 13:05:38have done, this is an alternative that
- 13:05:40we could have done is use the entire
- 13:05:43column and then um use the median of all
- 13:05:48of the data to replace. That's really up
- 13:05:50to us. Um this is one way of doing it.
- 13:05:53We could have done before we did the
- 13:05:56split. We could have um filled in with
- 13:06:00the median earlier. We chose to do it
- 13:06:03here mainly because it doesn't affect
- 13:06:05the features. So, we could have done
- 13:06:07this earlier and did it before we did
- 13:06:09the split and filled the NAS. Um really
- 13:06:12doesn't it's doesn't matter that much
- 13:06:14which way we do it. Um but we do need to
- 13:06:17fill in NLES. We cannot have those be
- 13:06:19null when we when we put it into our
- 13:06:21model. So some way we need to fill in
- 13:06:23NLES. Um and so in this strategy we're
- 13:06:27filling in our Y train um with the
- 13:06:30median salary from our training data.
- 13:06:32And same with this we're filling in with
- 13:06:34the median salary of the training data
- 13:06:36as well. But that's a choice f we could
- 13:06:39fill in with the mean with the average.
- 13:06:43Um we could fill in with the we could do
- 13:06:46it with all the data together before we
- 13:06:48split it. we could have filled in with
- 13:06:50all of the the median across the whole
- 13:06:52data set. Um either one works. You can
- 13:06:55do it either way, but we we did it um
- 13:06:59later here to show that it doesn't
- 13:07:01really affect the features. So, we can
- 13:07:03choose when we do it, right? It doesn't
- 13:07:05affect the features at all. So, we can
- 13:07:08do all of our pre-processing on the
- 13:07:09features and then do our label uh
- 13:07:12filling and nulls um if if we have them.
- 13:07:20uh x numerical. Um make sure you're
- 13:07:23running uh this
- 13:07:27uh x numerical was defined here
- 13:07:32when we split it apart um from
- 13:07:37uh when we dropped these columns here.
- 13:07:39So make sure you're running this. This
- 13:07:41is x numerical
- 13:07:43gets defined there.
- 13:07:45So, go back up to uh this cell
- 13:07:49where we split apart the y and we and we
- 13:07:51have the x here x numerical.
- 13:07:55Make sure you run this.
- 13:08:04Make sure you run this. And then you can
- 13:08:06run these. Then you run this to build x.
- 13:08:19All right.
- 13:08:21Are we up to here with this filling in
- 13:08:24the labels?
- 13:08:26Uh because then we can build our model
- 13:08:29once we're up to here. We've scaled
- 13:08:30everything. We filled in our NLES.
- 13:08:34We've gotten one hot encoding.
- 13:08:41Yeah, it is. That's why you know that's
- 13:08:42why we spend a lot of uh time on model
- 13:08:45on data preparation with pandas, right?
- 13:08:47That's why we did all that pandas work
- 13:08:49for sure. Yes, there is a lot of work
- 13:08:51before we can build a model.
- 13:08:54Yes, the mo do you guys notice that like
- 13:08:57the modeling is relatively easy. It's
- 13:08:58just a fit and predict. The modeling is
- 13:09:01actually really easy. It's all the other
- 13:09:03work that's that's more involved, right?
- 13:09:07more code.
- 13:09:10The modeling itself is really easy.
- 13:09:13It's just it's just one line of like
- 13:09:15ffit.
- 13:09:18Yeah, pretty easy to do.
- 13:09:23And then you do evaluation which is a
- 13:09:25couple lines.
- 13:09:47Yep. There's these are all the these are
- 13:09:50the common steps. All these steps we're
- 13:09:52doing are very very prototypical in
- 13:09:54model building is you let's just go back
- 13:09:57through this to see what we did right we
- 13:09:59imported our data
- 13:10:01um we analy we dropped this name column
- 13:10:04because it's not useful to us so we
- 13:10:06dropped that um we filled in the nles
- 13:10:10eventually um but you know if there were
- 13:10:13any nles in our features we would have
- 13:10:14to deal with those as well by replacing
- 13:10:16them or dropping the rows like we did
- 13:10:18earlier Um
- 13:10:21and then we do one hot encoding because
- 13:10:24of course we can't have any string
- 13:10:25columns going in our models. We got a
- 13:10:27one hot encode.
- 13:10:29Um we uh then build our X and Y by
- 13:10:33concatenating the one hot encoded back
- 13:10:36to the numerical features.
- 13:10:39Then we train test split. Right? That's
- 13:10:41pretty common. Or we could do cross
- 13:10:43validation either way. Um the K full
- 13:10:46cross validation. Then we scale. So, we
- 13:10:49didn't do this last time, but this is
- 13:10:50something we should get in the habit of
- 13:10:52is scaling um our features. So, we do
- 13:10:55that and now we're ready to model. So,
- 13:10:58now we're ready to model. Um so, that's
- 13:11:01this part.
- 13:11:04Okay. So, let's do the model. Um the
- 13:11:07model's actually uh pretty easy to do.
- 13:11:10So, we're going to use a lasso. So, we
- 13:11:12have a lasso model here. Notice what
- 13:11:14we're setting our alpha to. So the big
- 13:11:16parameter we really need ignore this
- 13:11:19iterations. We actually don't really
- 13:11:20need the we don't really need that
- 13:11:21parameter. Um so just ignore it for the
- 13:11:24moment. But the big one that we're
- 13:11:26setting here is the alpha. So when we
- 13:11:29did linear regression, we didn't need
- 13:11:31any parameters to go inside the linear
- 13:11:33regression object. We didn't need any
- 13:11:35parameters, right? Because there are
- 13:11:37really no parameters of it. But for
- 13:11:39lasso, the important one is the alpha.
- 13:11:42And so we need to know what to set alpha
- 13:11:44to. Um let's start with alpha equals 1.
- 13:11:49That's a good starting place. So a
- 13:11:51typical um starting point
- 13:11:55for alpha
- 13:11:58um is uh is one. So that's a typical
- 13:12:03starting point. And so we can set alpha
- 13:12:05equals to one. This max iterations is
- 13:12:08the the parameter that governs the
- 13:12:11training process because it is
- 13:12:13iterative. So if for some reason we we
- 13:12:16can't converge to the right betas and
- 13:12:18we've run it for 10,000 steps once we
- 13:12:21pass 10,000 steps, uh it will stop and
- 13:12:24just give us the betas at that point.
- 13:12:26But it will likely never hit this
- 13:12:29number. It'll converge before then. So
- 13:12:31um we don't really need to um specify
- 13:12:34it. So, I'm actually just going to get
- 13:12:36rid of it. Um, it's not really a big
- 13:12:38deal. It should converge before then.
- 13:12:41Um, but if if we want to set like a
- 13:12:44maximum step size in the optimization,
- 13:12:46we definitely could there. Uh, but not
- 13:12:49concerned about that too much. But
- 13:12:51here's our lasso. And then we're just
- 13:12:53going to do a fit on our data. So, look
- 13:12:56how easy that is. Just like a linear
- 13:12:58regression. Lasso.fit,
- 13:13:02right? So, we do fit. Um,
- 13:13:10oh, I didn't run this. I'm sorry. I got
- 13:13:12to run this. Okay. Actually, that's a
- 13:13:16good example of what happens when you
- 13:13:17don't when you have nulls, right? So, it
- 13:13:19says our our null contains nan. That's
- 13:13:21because I didn't run this. But now, that
- 13:13:24should be filled in. Now, we should be
- 13:13:26able to run this. Okay, perfect. So it
- 13:13:28runs.
- 13:13:36Okay. So you can see what the intercept
- 13:13:38is. Um this is one of our coefficients,
- 13:13:40right? The intercept is 457. And look
- 13:13:43now what's really interesting about the
- 13:13:44coefficients is look at what some of the
- 13:13:47coefficients are.
- 13:13:49Some of them are actually zero, which is
- 13:13:53really So some of them ended up being at
- 13:13:55zero, which is very very interesting.
- 13:13:58that means that those features get
- 13:14:02cancelceled out and they're basically
- 13:14:03not part of the model which is really
- 13:14:06interesting. Um so we have all these
- 13:14:09coefficients and some of them are zero.
- 13:14:16Yeah, negative0 is just because of the
- 13:14:19convergence like they started out
- 13:14:21negative and worked their way up to
- 13:14:23zero. it. Negative zero really just
- 13:14:26means zero, but they just were coming
- 13:14:29from they were like small negatives and
- 13:14:31ended up at zero
- 13:14:34during the training process. They were
- 13:14:36negative at one point and ended up zero.
- 13:14:39Um
- 13:14:40so yeah, negative 0 just obviously means
- 13:14:43zero. Um it's still still zero there.
- 13:14:54So what's interesting is some of these
- 13:14:56features ended up uh being zero which
- 13:14:59you don't usually see in a linear
- 13:15:01regression. So if we were to train this
- 13:15:03using a linear regression we typically
- 13:15:05wouldn't see that but some of these turn
- 13:15:07out to be zero because again we're
- 13:15:10encouraging those betas to be small.
- 13:15:13we're encouraging them to be uh small
- 13:15:16and so um you know what happens is some
- 13:15:20of them can be shrunk all the way down
- 13:15:22to zero meaning those features don't
- 13:15:24contribute that's a really simple model
- 13:15:26at that point right so we've taken
- 13:15:29something complex that includes all of
- 13:15:32these features and actually reduced it
- 13:15:34into something simple that only includes
- 13:15:36these features
- 13:15:38right
- 13:15:40so that's what it does um now we need to
- 13:15:44evaluate this to see how good of a model
- 13:15:46it is. But that's what this is saying
- 13:15:49here in this text is that um a positive
- 13:15:52uh coefficient indicates that as the
- 13:15:54independent variable increases the
- 13:15:56dependent variable also increases.
- 13:15:58Negative coefficient means as the
- 13:16:00independent variable increases dependent
- 13:16:02decreases because it's reducing the
- 13:16:04value. Um and lasso is known for feature
- 13:16:09selection by shrinking some of them to
- 13:16:10zero effectively removing those
- 13:16:13variables from the model from the
- 13:16:15equation right
- 13:16:17um
- 13:16:19so that's what happens
- 13:16:23some of them end up being zero
- 13:16:27were you guys able to run this this
- 13:16:30lasso uh fit which is the training of
- 13:16:33the lasso
- 13:16:52No, it doesn't ensure there's no
- 13:16:53overfit, but it helps with overfitting.
- 13:16:56It's supposed to help by making the
- 13:16:58model simpler. And this is definitely a
- 13:16:59simpler model because it's removing some
- 13:17:02of the features from the model
- 13:17:03essentially, right? Because some of the
- 13:17:05features aren't going to contribute.
- 13:17:07It's a simpler model.
- 13:17:09It doesn't it doesn't mean there's not
- 13:17:11going to be any overfitting, but it
- 13:17:13helps prevent it. That's what it's
- 13:17:15designed to do, help prevent it.
- 13:17:19Yeah. So, higher coefficient. Yes. The
- 13:17:22higher coefficient means it's a more
- 13:17:24important feature towards the
- 13:17:26prediction.
- 13:17:27Yes. That's what it means for sure. The
- 13:17:30higher the magnitude, the more of a
- 13:17:32contributor towards that prediction. Uh
- 13:17:35it is. Yes.
- 13:17:43And it's not just it's it could be
- 13:17:45higher positive or negative there. Like
- 13:17:47a higher negative is also a pretty big
- 13:17:50factor,
- 13:17:52right? So So you want to think about it
- 13:17:53in terms of absolute value.
- 13:18:01does not guarantee but helps. Yes, it
- 13:18:03doesn't guarantee it but it's designed
- 13:18:05to help overfitting, help prevent it.
- 13:18:07Yes, absolutely.
- 13:18:32Okay.
- 13:18:35So let's do some evaluation. Um so let's
- 13:18:38do in this case we are going to do our
- 13:18:42predict
- 13:19:02Oh, yeah. I'm not sure why that's the
- 13:19:05case.
- 13:19:07Interesting.
- 13:19:27We could try increasing the um max
- 13:19:31iterations.
- 13:19:43Okay, that's why. Yeah. So then you get
- 13:19:45that result with the with the higher max
- 13:19:47iterations.
- 13:19:49It doesn't get cut off there.
- 13:19:53I think that's why you probably left
- 13:19:54this in there,
- 13:19:57which is fine. You get about the same
- 13:19:58numbers.
- 13:20:06Yeah.
- 13:20:16All right. Let's evaluate this. So,
- 13:20:17we're going to to to do evaluation. I
- 13:20:19want you guys to see again, we should
- 13:20:21get in the habit of doing evaluation,
- 13:20:24which is taking our model and predicting
- 13:20:26on the training and predicting on the
- 13:20:29test sets, right? So we predict on the
- 13:20:31train set and calculate our MSE
- 13:20:35and we um calculate our R2 score um or R
- 13:20:41squar score I should say. Uh but again
- 13:20:44the MSE is the one we're really going to
- 13:20:46use mostly. Um but we calculate so we do
- 13:20:50our predictions and then we compare that
- 13:20:52into our mean squared error with our
- 13:20:54labels
- 13:20:56and we uh go ahead and do the same thing
- 13:21:00with the test. Right? So we do uh
- 13:21:02lasso.predict
- 13:21:04on our test features and we go ahead and
- 13:21:07compare that with the test labels. And
- 13:21:10so what we're doing there is generating
- 13:21:12our MSE.
- 13:21:15So we we take a look at our MSE and we
- 13:21:18get uh 84,000
- 13:21:21MSE. Um and so of course we could take
- 13:21:25the um what we could do with that is
- 13:21:29take a look at the um MSE on the uh we
- 13:21:34could do um MP. Square root
- 13:21:39and do the square root of the MSE test.
- 13:21:45and we get um 340. So this would be in
- 13:21:49the units of our label. So, if we go
- 13:21:51back and look at our label um for some
- 13:21:54of those um
- 13:22:08so uh we are in 300s and our data is
- 13:22:12like right around the 500. So, of
- 13:22:14course, if we describe this um we could
- 13:22:16see what the statistics are of it. So,
- 13:22:19we could do df.escribe describe and
- 13:22:21generate that. But that doesn't look
- 13:22:22like a very good error, right? If these
- 13:22:24are in the 400s, um that's that's not a
- 13:22:27very good error.
- 13:22:29So again, it's not a very great model.
- 13:22:32But one thing I want you to see is that
- 13:22:33it's it's not overfitting.
- 13:22:36Um if anything, it's actually
- 13:22:38underfitting, which is what this kind of
- 13:22:41um MSE suggests, right? because our
- 13:22:43error here is 84 uh excuse me 84,000.
- 13:22:48Um
- 13:22:54our our area here is 84,000,
- 13:22:58excuse me. And on the test set it's
- 13:23:01116,000.
- 13:23:02Um so these two errors are both bad. So
- 13:23:08it's not overfitting. This is actually
- 13:23:10underfitting. So it's not overfitting,
- 13:23:13it's actually underfitting. Um, and so
- 13:23:16that's the risk with something like
- 13:23:17lasso is that it's making the model a
- 13:23:20bit too simple and we actually risk
- 13:23:23underfitting, which is what happens. We
- 13:23:26have too much error across both the
- 13:23:29training and the test set. Overfitting
- 13:23:31is when we do we have really good
- 13:23:34performance on the training set, but bad
- 13:23:36performance on the test set. We're not
- 13:23:38overfitting.
- 13:23:40um we are uh underfitting because our
- 13:23:43performance is not good either way. Even
- 13:23:45this R squar is pretty low. It's not
- 13:23:48even at 50%.
- 13:23:56Okay, so that's so we we do the
- 13:23:59evaluation and again the evaluation just
- 13:24:00comes down to making predictions and
- 13:24:03computing our error amongst those
- 13:24:05predictions to our labels. That's always
- 13:24:07what the uh evaluation is going to be
- 13:24:14for MSE.
- 13:24:17What's the ideal MSE? What do you think
- 13:24:20it should be? What is So, think about it
- 13:24:23like this. The MSE represents the
- 13:24:25average distance between our predictions
- 13:24:30and the labels.
- 13:24:33So, if we're getting it right all the
- 13:24:36time, what's that distance going to be
- 13:24:38if we're always right? What's our
- 13:24:40distance from what's our distance from
- 13:24:43our predictions to our labels going to
- 13:24:46be if we're always getting it right?
- 13:24:49Zero. Yeah, there's not going to be any
- 13:24:51distance. It's going to be right. It's
- 13:24:53going to be perfectly aligned, right?
- 13:24:55There's going to be no distance there.
- 13:24:57So, yeah, an ideal MSE is zero.
- 13:25:00That's an ideal MSE.
- 13:25:03So, anything close to like the smaller
- 13:25:06the better for MSE. The smaller the
- 13:25:09better. Um, for this R squared, uh, it's
- 13:25:13it's a scale between 0 to one where one
- 13:25:16is the best. So, one would be perfectly
- 13:25:18aligned predictions. Um, so, and again,
- 13:25:22this this is we actually multiply by 100
- 13:25:25to get uh because it's it's a number
- 13:25:26between 0 and one. So we get about 47%
- 13:25:30which is not good.
- 13:25:41Okay.
- 13:25:56All right. Any questions on this
- 13:25:59evaluation?
- 13:26:13All right. I want to show you something
- 13:26:15which is
- 13:26:18Yeah, this that's true. the scale of it
- 13:26:21matters on the data because we should be
- 13:26:22you should always interpret your MSE in
- 13:26:25the scale of
- 13:26:27um your your labels because your labels
- 13:26:32like in this case our labels um you know
- 13:26:34we could take uh for example we could
- 13:26:37easily let's actually do that let's take
- 13:26:39the average
- 13:26:42let's take the average of our labels on
- 13:26:45the training data
- 13:26:49and and we can see what those are. Um,
- 13:26:52so the average is 500,
- 13:26:55right? The average is 500. And look at
- 13:26:58what our uh square root of our MSE is,
- 13:27:01which is in the same units as our
- 13:27:03original. Um, so we have uh quite a bit
- 13:27:07of error. 340 when our units are right
- 13:27:10around 500.
- 13:27:13So that's quite a bit of error.
- 13:27:23Yeah, MSSE of zero means our our uh our
- 13:27:26predictions are nearly identical to the
- 13:27:29test labels. Yes, that's what MSSE of
- 13:27:33zero means. There's zero distance.
- 13:27:38So closer to zero, the better.
- 13:27:47But we talked about it as you you really
- 13:27:49so the rule of thumb should be what is
- 13:27:53your RMSSE as a percentage of your
- 13:27:57typical value. So your typical value is
- 13:28:00in the 500s. Our our RMSSE is 340.
- 13:28:05That's just really high. That's over
- 13:28:07like 60% of that value.
- 13:28:11So that's just a lot. That's too much
- 13:28:14error. What we would love this RMSSE to
- 13:28:16be is under 20% of the typical value. So
- 13:28:20that means on average we are 20% or less
- 13:28:25off in our prediction. That would be
- 13:28:28good. That would be pretty good. That
- 13:28:30means we're like 80% accurate,
- 13:28:34right? That'd be pretty ideal. So you
- 13:28:36got to think about it in terms of this
- 13:28:37RMSSE which is in the same units as your
- 13:28:40labels.
- 13:28:43This is the
- 13:28:47RMSSE
- 13:28:50which is in the same units as the
- 13:28:54labels.
- 13:28:57So and then to interpret this we have
- 13:29:00340
- 13:29:02is compared to
- 13:29:05typical
- 13:29:07um salary unit of 500
- 13:29:11right so this is uh quite a bit when the
- 13:29:15typical value is 500 and we are off on
- 13:29:18average by 340 units
- 13:29:21that's so much relative to the typical
- 13:29:24value
- 13:29:26that's just too. That's a lot of error.
- 13:29:28That's not a very good model, right?
- 13:29:31It's underfitting. It's definitely
- 13:29:33underfitting.
- 13:29:45Yeah. So, that's a great question. What
- 13:29:46should we do from here? So, um because
- 13:29:49we're underfitting
- 13:29:57um we should use a more complex model.
- 13:30:02So uh we're going to learn about those
- 13:30:04in lesson four, but we should use
- 13:30:06something different. This linear
- 13:30:08regression is still too basic. Even with
- 13:30:10lasso, it's still too basic.
- 13:30:16Yeah, we're underfitting because we But
- 13:30:18it could also be we're underfitting with
- 13:30:20a regular linear regression. We should
- 13:30:22test that out. Um, and maybe it would be
- 13:30:24an exercise for you guys um to test that
- 13:30:28out yourself. It shouldn't be hard to
- 13:30:29do. Um, you already have all the data
- 13:30:32scaled. You So, do you see how you would
- 13:30:35do that? You would just come in here and
- 13:30:36build a linear regression rather than a
- 13:30:38lasso and dofit. And then you would
- 13:30:41evaluate it the same way with a predict.
- 13:30:44It's really easy to do that. And then we
- 13:30:46can compare that um to to this. It
- 13:30:51shouldn't be that hard to do that,
- 13:30:52right?
- 13:30:54And something you guys could do for
- 13:30:55sure. Um,
- 13:30:58is build the linear regression and
- 13:31:00actually compare it and see what kind of
- 13:31:04difference it makes. I mean, we honestly
- 13:31:06we could do it ourselves. We could do it
- 13:31:07right now. Maybe it's worth trying that.
- 13:31:12So, let's build a linear regression
- 13:31:16for comparison.
- 13:31:20So we have our linear regression
- 13:31:24uh is linear regression and then we do
- 13:31:28ffit linear regression.fit fit
- 13:31:32right so so this will train it um and
- 13:31:35then we can evaluate it so lin mse is
- 13:31:41mean squared error
- 13:31:44and then we can do our um let's do our
- 13:31:47training let's do the training and then
- 13:31:52um let's predict
- 13:31:54actually let me do that here
- 13:31:57uh y prediction
- 13:32:00train
- 13:32:03linear
- 13:32:05equals um linear regression.predict
- 13:32:10and then we're going to predict on our
- 13:32:12training features.
- 13:32:17Okay, do you guys see what I'm doing?
- 13:32:18I'm building a linear regression for
- 13:32:20comparison.
- 13:32:22I'm doing fit here to train it and then
- 13:32:25I'm making some predictions on the
- 13:32:26training set and we're going to evaluate
- 13:32:29those. I'm going to replace that here
- 13:32:30with y prred
- 13:32:34uh train
- 13:32:37linear. So these predictions
- 13:33:02Okay. So, if you guys want this code, I
- 13:33:03can paste it in.
- 13:33:15So, let's see what the RMSSE for just a
- 13:33:18linear model is.
- 13:33:21It's a little bit better. It's better
- 13:33:23for sure.
- 13:33:26So 289 is better than this 340. It's
- 13:33:30better. It's getting closer to zero.
- 13:33:33It's still underfitting though,
- 13:33:36right? And that's just on the training
- 13:33:38set. Let's look at the Let's do the same
- 13:33:41thing, but on
- 13:33:45Let's change this. Let's swap this out
- 13:33:47for um test
- 13:33:51And then let's do test.
- 13:33:55And then let's do test
- 13:34:01test.
- 13:34:04And then
- 13:34:06test test.
- 13:34:17Okay. Okay, so this is producing test
- 13:34:19predictions on the test set.
- 13:34:22We are generating an MSE test
- 13:34:27and then we're doing MSE test
- 13:34:30which is using the test labels and our
- 13:34:32test predictions and then we take the
- 13:34:35square root of that for RMSSE and then
- 13:34:36we're going to generate that. So it's
- 13:34:39still under fit. I mean this is still
- 13:34:41high. This is still high um on the test
- 13:34:44set and versus on the training set. So,
- 13:34:46it's still pretty high. Um, even the
- 13:34:49basic linear regression is under is
- 13:34:50still underfitting. Still underfitting,
- 13:34:53right? Even without the lasso,
- 13:34:56which is lasso is supposed to help with
- 13:34:58overfitting. It's definitely not
- 13:34:59overfitting. Um, it's definitely
- 13:35:02underfitting, but this is a signal that
- 13:35:05it's kind of overfitting because this is
- 13:35:07performing better on the training data
- 13:35:09and then it gets worse on the test data.
- 13:35:13Definitely gets worse, right?
- 13:35:23Did you guys follow?
- 13:35:27I'm just running this above I'm running
- 13:35:29this above this. It doesn't matter where
- 13:35:31you put it. We could uh we could move it
- 13:35:33down.
- 13:35:40We could move it down to I just ran I
- 13:35:43just picked a new cell right here and
- 13:35:46ran it. But we could move it actually
- 13:35:48let's do that. Let's move it down
- 13:35:53to
- 13:35:56after the lasso evaluation.
- 13:36:01Okay. So I just moved it there.
- 13:36:04And then let's move
- 13:36:07this down.
- 13:36:09So I just put it here after the um after
- 13:36:12this. So this is the um this is
- 13:36:16basically the objective function right
- 13:36:19of the training process. So during the
- 13:36:22algorithm that runs when we call ffit in
- 13:36:25scikitlearn it's going to find these
- 13:36:28betas right it's actually going to learn
- 13:36:30what these best betas are for our model.
- 13:36:34Um this is our model here, right? It's
- 13:36:36the combination of betas times our
- 13:36:37features um plus an intercept beta. Uh
- 13:36:42so that's our model. But um we penalize
- 13:36:45those large uh weights in absolute value
- 13:36:49by um adding a penalty term like this um
- 13:36:53where alpha is some level of penalty
- 13:36:56that we want to provide. Usually alpha
- 13:36:58equals 1 is okay. But um actually what
- 13:37:01we're going to learn uh to finish out
- 13:37:02this section is there's going to be a
- 13:37:04systematic way we can test out different
- 13:37:06alphas um that represent the level of
- 13:37:09penalty we want to uh apply to lasso or
- 13:37:12even ridge
- 13:37:14uh regression. So that was the lasso and
- 13:37:17um if you guys remember using it was
- 13:37:19super easy. Uh we worked through this
- 13:37:22problem with this um baseball data um
- 13:37:26and we had uh
- 13:37:29let's see scrolling down we um split out
- 13:37:31our numerical data and we did uh we one
- 13:37:36hot encoded our our categorical data
- 13:37:39combined it back together. Hopefully
- 13:37:41that um rings a bell there. Um and we
- 13:37:44actually scaled our data which is pretty
- 13:37:46standard to do is we do some type of
- 13:37:48scaling to our features especially our
- 13:37:50numerical features right want to scale
- 13:37:52those in some way whether it's minmax
- 13:37:54scale or standard scaler um want to do
- 13:37:57that and so we did that for this example
- 13:37:59and then we um ran the lasso regression
- 13:38:04which is pretty easy to use. You just
- 13:38:05use the lasso object and you pick an
- 13:38:07alpha here. Um, again, we are going to
- 13:38:10have a way to test out different alphas
- 13:38:14that could be candidates and we can see
- 13:38:16which one's the best. Um, so I'm going
- 13:38:19to show us that today coming up shortly.
- 13:38:23But that was that was the lasso. If you
- 13:38:25guys remember, we did that. Um, this it
- 13:38:28we compared that to a basic linear
- 13:38:30regression which is just this pretty
- 13:38:32straightforward just a fit and then
- 13:38:34predict and then we can generate mean
- 13:38:35squed error. Um, still not a very good
- 13:38:39mean squared error on this data, it's
- 13:38:41still fairly large. Um, so it's still
- 13:38:44not, no matter which model we use, it's
- 13:38:46still not very good, but at least we can
- 13:38:49practice doing that comparison. That's
- 13:38:50what we did last time. We did this on
- 13:38:53Wednesday.
- 13:38:54Um
- 13:38:56and then
- 13:38:58we saw that the effect of different
- 13:38:59alphas we had a lasso um
- 13:39:03we had a lasso uh cross validation
- 13:39:06example here. So beyond just using a
- 13:39:08regular lasso model that um scikitlearn
- 13:39:10has a lasso cv which allows you to try
- 13:39:13out different alphas uh with cross
- 13:39:16validation and um figure out what the
- 13:39:18best alpha is. Um, now we're actually
- 13:39:21going to have a different strategy
- 13:39:22that'll instead of just picking random
- 13:39:24ones, we can actually um supply multiple
- 13:39:28parameters that we may want to test. Um,
- 13:39:31as many as the models may support. And
- 13:39:33in some more complex models, we'll have
- 13:39:36more than one parameter like lasso only
- 13:39:38has the alpha. Um, technically it also
- 13:39:41has this max iterations, but really the
- 13:39:43only one that matters is this alpha.
- 13:39:45Other models have many more
- 13:39:47hyperparameters that we can um uh change
- 13:39:52and so we want a way to systematically
- 13:39:54test out those different combinations
- 13:39:57and to see which one leads to the best
- 13:39:59uh version of that model. Let's say the
- 13:40:01best results. So um we're going to
- 13:40:04explore that coming up. So we had lasso.
- 13:40:08Um now this is where we ended last time.
- 13:40:10We had ridge regression. If you guys
- 13:40:12remember, this one is just a slightly
- 13:40:15different penalty. Um,
- 13:40:18it takes the it I drew it out for us. It
- 13:40:21takes the same penalty we had before.
- 13:40:23So, it has that um residual sum of
- 13:40:26squares error, which is the main one we
- 13:40:29use for linear regression, but it has a
- 13:40:31penalty with an alpha. And then it has
- 13:40:33the sum of the beta squares
- 13:40:37beta i squares. So it penalizes it has a
- 13:40:42penalty but it penalizes slightly
- 13:40:44differently where it uses the square not
- 13:40:46the absolute value. That's the ridge
- 13:40:48regression. And this has the similar
- 13:40:50effect of you don't in order to minimize
- 13:40:52this right because our goal in training
- 13:40:54a model was to minimize this thing
- 13:40:57minimize this um quantity and find the
- 13:41:01best betas that minimize this. Um so
- 13:41:04generally yes you want to encourage
- 13:41:06lower values but the um once you get
- 13:41:10values that are a fraction if you square
- 13:41:12them they actually get smaller. Um, so,
- 13:41:16uh, it's it's not, um, it's not
- 13:41:20necessary to shrink them all the way to
- 13:41:22zero. They will get smaller as soon as
- 13:41:24they're kind of below one. Um, so they
- 13:41:27don't encourage it to completely go
- 13:41:29away, uh, like the absolute value does.
- 13:41:32It's just slightly different
- 13:41:33minimization. Um, so what we see with
- 13:41:36the ridge is we don't see the features
- 13:41:38kind of get wiped out completely like we
- 13:41:40do with a lasso. and lasso they get
- 13:41:42encouraged to be um to become zero
- 13:41:45because that's kind of the only way to
- 13:41:46minimize an absolute value. But with
- 13:41:48squares they can keep getting smaller
- 13:41:50and smaller and smaller um fractions and
- 13:41:54they don't have to become zero. It's not
- 13:41:56as harsh of a of a penalty.
- 13:41:59Um so uh the ridge was easy to use as
- 13:42:05well. Um and it also has an alpha that
- 13:42:09we can set. So, it's literally the same
- 13:42:12exact code, just a different model, just
- 13:42:15slightly different penalty, and it
- 13:42:17results in different coefficients. You
- 13:42:19notice that none of them are exactly
- 13:42:20zero. Like with the lasso, you can get
- 13:42:22ones that are exactly zero. We don't see
- 13:42:25that with the ridge. You remember that.
- 13:42:28Um, so we we s pointed out that last
- 13:42:30time. Notice the coefficients aren't
- 13:42:31zero. Um, and then we can evaluate it.
- 13:42:34So we did our MSE calculation which is a
- 13:42:37pretty standard thing where we use our
- 13:42:39model to predict on a training set,
- 13:42:41predict on a test set, evaluate those um
- 13:42:45by computing the metric like the mean
- 13:42:47squared error and we can see if we're
- 13:42:49overfitting underfitting. This is
- 13:42:51definitely the same kind of story we've
- 13:42:53seen with all these models is
- 13:42:54underfitting because the error is so big
- 13:42:56across both sets
- 13:42:58across training and test. So it's it's
- 13:43:00definitely underfitting.
- 13:43:02Um
- 13:43:04and same thing as lasso, it has a cross
- 13:43:06validation uh variation on it that
- 13:43:09allows you to try out different alphas
- 13:43:12and um do different folds. So 10 folds,
- 13:43:15five folds, whatever, and compute the um
- 13:43:19try to find the best alpha that way.
- 13:43:23Okay.
- 13:43:27All right.
- 13:43:29Any questions on this so far from last
- 13:43:32time from reviewing that a little bit?
- 13:43:35Hopefully that uh hopefully that is
- 13:43:38jogging your memory a little bit on
- 13:43:40ridge and lasso. Um you know where we're
- 13:43:43going to pick it up today is to finish
- 13:43:45out this lesson with one more model
- 13:43:49which is going to be a combination of
- 13:43:51ridge and lasso. So you can actually
- 13:43:54combine them together
- 13:43:56um in a linear fashion those penalties.
- 13:44:00So you can actually have both penalties,
- 13:44:02the absolute value and the square. And
- 13:44:04when you have both penalties um that's a
- 13:44:08special model called the elastic net uh
- 13:44:11regression or elastic net model. Um so
- 13:44:14this is a combination of lasso and ridge
- 13:44:18together. So you have lasso, you have
- 13:44:20ridge and then you have elastic net
- 13:44:21which combines both of those penalties.
- 13:44:24Um let me show you the equation.
- 13:44:28So here is the uh so here is the the
- 13:44:33model. This is the same that we've
- 13:44:35always had. This is our usual u model
- 13:44:39fitting for linear. This is a basic
- 13:44:42linear regression um loss function or
- 13:44:44objective function that we're trying to
- 13:44:46minimize to find the betas. Notice how
- 13:44:48we have both of our penalties though
- 13:44:50this time. So instead of just having one
- 13:44:52of the penalties, we actually have both.
- 13:44:54So we have the lasso penalty
- 13:44:57and then we have the ridge penalty here.
- 13:45:00So we actually use both of them and um
- 13:45:04try to find a balance of minimizing
- 13:45:06those two uh those two penalties.
- 13:45:11Okay. And notice how they instead of
- 13:45:13just a single alpha, we kind of have a
- 13:45:14balance on both of them.
- 13:45:18So, we can actually weight the lasso one
- 13:45:20more. We can weight the ridge one more.
- 13:45:23We can weight them the same. Uh we can
- 13:45:27um change that around as much as we
- 13:45:28want. So, they have two different
- 13:45:29weights there um that they could be.
- 13:45:34Um now what happens in reality is uh
- 13:45:39we're going to see this in the model is
- 13:45:41that um usually what happens is these
- 13:45:44get combined into a fraction. So there's
- 13:45:47usually a ratio of lambda 1 to lambda 2
- 13:45:51and this is known as the um this is
- 13:45:54sometimes known as the L1 ratio
- 13:45:58and this is a this is a a parameter
- 13:46:00inside the model that we'll be able to
- 13:46:02set um along with alpha. So we'll be
- 13:46:05able to set an alpha and then this
- 13:46:07ratio. Um the idea is is that um the
- 13:46:12ratio will uh allow us to control which
- 13:46:16one is more dominant. So if this number
- 13:46:19is bigger the um this lasso penalty will
- 13:46:23will be weighted more. If this ratio is
- 13:46:26smaller if it's less than one for
- 13:46:28example that means that the um ridge
- 13:46:31regression is more uh dominant. Um but
- 13:46:36the so we'll have this we'll have really
- 13:46:38this and this at our disposal and alpha
- 13:46:43is um
- 13:46:46alpha is kind of like a a you can think
- 13:46:48of it as a scale that is um so lambda 1
- 13:46:53kind of like lambda 1 plus lambda 2 um
- 13:46:57combined to equal alpha.
- 13:47:00So it's like our total level of penalty
- 13:47:03um our total level of penalty and we can
- 13:47:06set that equal to one. We can set it
- 13:47:08equal to whatever we want. Um and so
- 13:47:11these will be in this ratio and there'll
- 13:47:13be a total level of penalty that we can
- 13:47:15apply. So the model will actually use
- 13:47:18these two parameters when we when we do
- 13:47:20it. But that's how they're that's how
- 13:47:22they're all related.
- 13:47:25Okay. So ridge uses both penalties.
- 13:47:28That's the only difference between lasso
- 13:47:30or sorry elastic net uses both
- 13:47:32penalties. Um so one thing I want you to
- 13:47:35notice is that uh if we um if we want we
- 13:47:41could set this L1 ratio all the way to
- 13:47:43zero
- 13:47:45um which uh if we do that um the only
- 13:47:50way this L1 ratio could be zero would be
- 13:47:52if lambda 1 is zero. So it would just
- 13:47:54revert back to ridge regression. So it
- 13:47:56complet if if this is zero this will
- 13:47:59wipe out this term and we'll be back to
- 13:48:00ridge if the L1 ratio is zero.
- 13:48:05Okay.
- 13:48:10All right. So we have a elastic net
- 13:48:13model. Um now it's used the exact same
- 13:48:17way as we did the other models in the
- 13:48:19code. So we have elastic net um uh from
- 13:48:23the scikitlearn linear model family just
- 13:48:26exactly where we had linear regression
- 13:48:29lasso ridge all of those came from this
- 13:48:32linear model um elastic net also comes
- 13:48:35from there and then the cross validation
- 13:48:36version also comes from there um so
- 13:48:41let's see so when we build our model
- 13:48:43it's going to be um very very simple
- 13:48:46easy stuff because it's the same code
- 13:48:48that we always have um we just use the
- 13:48:51elastic net. We set an alpha alpha
- 13:48:54equals 1 is pretty standard um just like
- 13:48:56it is in in the last one ridge that's
- 13:48:58industry standard is one and then an L
- 13:49:02L1 ratio of.5
- 13:49:04that's pretty standard as well. What the
- 13:49:05L1 ratio.5 is is kind of a um
- 13:49:11uh kind of a that means that the lambda
- 13:49:141 to lambda 2 ratio is 1/2. Um, so
- 13:49:18that's that's a pretty standard uh ratio
- 13:49:20as well, but again, we could set this
- 13:49:23equal to one and they'd be kind of
- 13:49:25equally weighted. Um, L1 ratio of a half
- 13:49:28means that the uh ridge regard the the
- 13:49:32ridge penalty is a little bit more
- 13:49:34weighted uh in that in that situation.
- 13:49:39Okay.
- 13:49:40So uh once we have this model um we can
- 13:49:44do ffit and we can run that on our
- 13:49:46training data and we can um get we can
- 13:49:50figure out what our parameters are like
- 13:49:51our coefficients and our intercepts. Our
- 13:49:53model will have that but more
- 13:49:55importantly we can use our model to
- 13:49:56predict right so we can predict on the
- 13:49:58test set. Um let me go back and load our
- 13:50:02data and actually run this.
- 13:50:06So, we're going to be using the same
- 13:50:08data that we did for
- 13:50:10uh lasso,
- 13:50:14which is the I'm scrolling back up so I
- 13:50:16can load it. It's the baseball data
- 13:50:18here.
- 13:50:22Um,
- 13:50:25just run it from there.
- 13:50:27It's this hitters.csv. So, hopefully you
- 13:50:30have that one.
- 13:50:42Let me load this.
- 13:50:50Okay, so we loaded that and then that
- 13:50:52should load.
- 13:50:54Drop that unnamed column.
- 13:51:01We will get our dummies
- 13:51:10and then concatenate those split
- 13:51:15scale. I'm just rerunning things. I'm
- 13:51:18rerunning things so we can see our model
- 13:51:19one more time.
- 13:51:22So rerun that. Take a look at that. That
- 13:51:23looks good. and then
- 13:51:27fill in the NLES on the on those.
- 13:51:30Okay. So, we should be able to run our
- 13:51:34uh elastic net now.
- 13:51:43Okay. So, let's import that and then
- 13:51:46let's build our model. So, there we go.
- 13:51:47We build our model and the intercept is
- 13:51:51that. Now, of course, we can look at our
- 13:51:53coefficients. Let's look at that.
- 13:51:59Look at our coefficients. So remember
- 13:52:01the coefficients are the uh betas. These
- 13:52:03are our betas that are in our model. Um
- 13:52:06so we can take a look at those. Now um
- 13:52:08they're it's somewhere in between. It's
- 13:52:11not a full lasso where we're going to
- 13:52:12see some of these be zero. It's not a
- 13:52:14full ridge. Um so the coefficients we
- 13:52:17get are different. They're somewhere in
- 13:52:19between there. Those two models that
- 13:52:21we've already built. So not quite the
- 13:52:23same um somewhere in between there.
- 13:52:30Um and then we can use our model to make
- 13:52:32predictions and and compute the MSE
- 13:52:35uh or the RMSSE I should say as well. So
- 13:52:38we can take the mean squared error, pass
- 13:52:39that into the square root and compute
- 13:52:41the RMSSE. So still pretty bad. Um this
- 13:52:44is right around that 300 range of what
- 13:52:46we've gotten for our other RMSSE. So,
- 13:52:48it's not like elastic net is any better
- 13:52:50than those other like linear or lasso or
- 13:52:53ridge. And that's not surprising because
- 13:52:56it's just adding those extra penalties.
- 13:52:58We don't expect it to magically get
- 13:53:00better. It's actually a more complex
- 13:53:02um when we add when we add those in,
- 13:53:05we're actually reducing it and making it
- 13:53:07simpler. And we need something more
- 13:53:09complex, I should say. So, we're making
- 13:53:11it simpler um by by making penalizing
- 13:53:16our weights a little bit more. And so,
- 13:53:18it's still not a good fit. That's not
- 13:53:21really surprising, right? It's still not
- 13:53:23really a great fit.
- 13:53:25And we can we can even double check
- 13:53:27that. We know our RMSSE is pretty bad.
- 13:53:30Um but we can double check it with this
- 13:53:31R2 score. And it's, you know, still not
- 13:53:34good. Remember, a one would be really
- 13:53:36good. Um that'd be like a perfect linear
- 13:53:38model. This is um still pretty bad.
- 13:53:46Okay, so as we said, the alpha controls
- 13:53:48the overall strength. Um so the higher
- 13:53:51the alpha, the more overall penalty
- 13:53:54we're supplying, which makes the model
- 13:53:57simpler. Um uh but the L1 um ratio
- 13:54:02determines the mix or that ratio of the
- 13:54:06lambdas, the lasso to the ridge. Um if
- 13:54:09you have it be um exactly zero, you you
- 13:54:14revert all the way back to um if you if
- 13:54:18you put it at zero, you revert all the
- 13:54:20way back to ridge. One would be all the
- 13:54:22way to pure lassos. Somewhere in
- 13:54:23between, like one half is is good.
- 13:54:35Okay,
- 13:54:37so this is another example of trying out
- 13:54:40different values of alpha in the CV to
- 13:54:42see which one works. Now again, I'm
- 13:54:45going to show us in a minute a
- 13:54:46systematic way to do this, but this is
- 13:54:49just trying out um different alphas that
- 13:54:51we set up in this uh in this um
- 13:54:56uh range. So we have different uh values
- 13:54:59between minus2 and two um
- 13:55:02logarithmically.
- 13:55:03Um so these are uh logarithm values that
- 13:55:07are between this between minus2 and two
- 13:55:09and we choose a 100 different alphas and
- 13:55:12then we choose a 100 different um L1
- 13:55:14ratios between 0.01 and one and we run
- 13:55:18that we run this um cross validation
- 13:55:20with 10folds. So this is quite a bit.
- 13:55:22So, we're doing 10 folds and we're
- 13:55:25trying out a hundred different um
- 13:55:27options. Uh every time we do an option,
- 13:55:30we're trying out 10 folds to evaluate
- 13:55:31it. So, it's going to take a minute to
- 13:55:34run.
- 13:55:46It's still running here. But again, what
- 13:55:49this is doing is trying out different
- 13:55:50alphas and it's it's going to do a cross
- 13:55:54validation. And you guys remember the
- 13:55:56t-fold cross validation is where we take
- 13:55:58our data and we divide it into 10 folds
- 13:56:03and then we um train on nine of those
- 13:56:06and then test on the remaining fold and
- 13:56:08then we rotate all the folds 10 times.
- 13:56:11and that we average those mean squared
- 13:56:14error metrics together um against those
- 13:56:1810 different uh fold options to generate
- 13:56:22a basically like an average performance
- 13:56:25for that value of alpha. And we're doing
- 13:56:27that a 100 times for all these different
- 13:56:29100 alphas that there are and 100
- 13:56:32different L1 ratios that we're trying
- 13:56:33with them.
- 13:56:39So that's quite a bit of processing but
- 13:56:42uh it did finish.
- 13:56:46So we can see what our best alpha is and
- 13:56:48our best one ratio. So we get the best
- 13:56:50alpha is this best one ratio is this. Um
- 13:56:54and therefore we can uh build a model
- 13:56:57with those with just these two guys as
- 13:56:59the alpha and the L1 and um see how that
- 13:57:04performs.
- 13:57:06We build that model and then we predict
- 13:57:08on the test set and we generate the
- 13:57:10RMSSE. It's just a little bit better.
- 13:57:12It's still not It's just a little bit
- 13:57:14better, but it's still not good, right?
- 13:57:16It's still 338. It is just way too big.
- 13:57:20Remember, this is RMSSE, so it's in the
- 13:57:23units of our uh target variable. So,
- 13:57:27it's in the units of if we go back to
- 13:57:30our data, actually, I could just print
- 13:57:32it out here.
- 13:57:34um this RMSSE.
- 13:57:38If I just do this, we could take a look
- 13:57:40at um DF
- 13:57:42or I could look at Y test
- 13:57:47and you can see some of these values.
- 13:57:48These are these salary values in the
- 13:57:50hundreds, right? Some of them are in the
- 13:57:51thousands. Um but an error of like 338
- 13:57:56is just too big. That's a really big
- 13:57:58error. That means we would be off by an
- 13:57:59average of 300 when our our values if we
- 13:58:02just do the mean
- 13:58:06um
- 13:58:08is only 550 as on average is 550 but we
- 13:58:12have this amount of error on average um
- 13:58:15so that's just a way too big of a
- 13:58:17proportion of error right it's not a
- 13:58:19very good model and again we can verify
- 13:58:21that by looking at this R2 for.
- 13:58:31So if we go down here,
- 13:58:39still not very good.
- 13:58:43Here's some of our coefficients. So
- 13:58:45remember, you can always take your
- 13:58:46coefficients and line them up to your
- 13:58:48your data columns. Uh so that you can
- 13:58:51get a sense of what coefficient belongs
- 13:58:53with what feature. So that's all we're
- 13:58:56doing here is just creating a series
- 13:58:57where those coefficients instead of just
- 13:58:59printing out the coefficients, we're
- 13:59:00actually lining them up to the columns.
- 13:59:02So this tells us um remember the larger
- 13:59:05it is the more influence it kind of has
- 13:59:07on the on the final result. Um either
- 13:59:10way, so like this has a big negative
- 13:59:12influence. um this has a large positive
- 13:59:15influence.
- 13:59:22Okay, let me pause there. Any questions
- 13:59:25about the
- 13:59:27elastic net model?
- 13:59:31This is a really this model is a really
- 13:59:33good one to use when you are building a
- 13:59:36linear regression and it's performing
- 13:59:37well but it's overfitting. This is a
- 13:59:40really good one to use because you can
- 13:59:41balance
- 13:59:43lasso and ridge you can get the best of
- 13:59:45both worlds. So the the main strategy is
- 13:59:48if you are using a linear regression and
- 13:59:51you see overfitting
- 13:59:53um meaning that it's performing decently
- 13:59:56so on the training set
- 13:59:59it's performing okay but then on the
- 14:00:01test set like you know it's it's not
- 14:00:04underfitting. it's performing pretty
- 14:00:05well on the training set, but then on
- 14:00:07the test set it's um performance is much
- 14:00:11worse. That's overfitting. If you're
- 14:00:14overfitting, then this is a great model
- 14:00:15to use because we can try basically by
- 14:00:18by rotating through different alphas and
- 14:00:20different L1 ratios, we can try out
- 14:00:23different strengths of penalty and
- 14:00:26different variations on lasso and ridge
- 14:00:28together. This is a really good model to
- 14:00:30to use for those overfitting cases where
- 14:00:33linear regression is doing decently. Um,
- 14:00:37but it's overfitting,
- 14:00:39right? So far, we haven't ran that case
- 14:00:42because so far, no matter what model
- 14:00:44we've used, it's always underfit. So,
- 14:00:48anytime we have those underfitting
- 14:00:50cases, it signals that we should likely
- 14:00:53just use a more complex model. And we
- 14:00:56haven't learned about those yet.
- 14:00:58um we will coming up in lesson four, but
- 14:01:02um that's for this data. That's ultim
- 14:01:05ultimately what we'd want to do is
- 14:01:06probably use a more advanced model
- 14:01:08because it's underfitting um just using
- 14:01:10a linear regression and and then using
- 14:01:12the the overfitting variations of linear
- 14:01:14regression like lasso ridge and elastic
- 14:01:16net.
- 14:01:22Okay.
- 14:01:26Any questions on this on elastic then
- 14:01:35the TV? Yeah, we Yeah, I think I have
- 14:01:37it. I can share it with you.
- 14:01:44I said that and now I can't find it. I
- 14:01:46thought I had it.
- 14:01:55I don't have it. I thought I had it, but
- 14:01:57I don't.
- 14:01:59If anyone does have that one.
- 14:02:05Yeah, I'll look one more time. I thought
- 14:02:07I had that one.
- 14:02:11Um,
- 14:02:17yeah, it's not in there. I had it. Let
- 14:02:19me see.
- 14:02:32Yeah, I don't have it either. I thought
- 14:02:33I had it in here.
- 14:02:42Yeah, I don't have that one.
- 14:02:46I don't have that one. I'll have to find
- 14:02:48it. Uh I have this marketing data. I
- 14:02:50don't think this is the same one.
- 14:02:54I have this marketing data. I don't
- 14:02:55think that's the right one, but you can
- 14:02:56take a look at it.
- 14:03:01No, we're using So, for this example,
- 14:03:03we're using the same hitters data set
- 14:03:05that we used earlier for lasso.
- 14:03:08No, that's an earlier one.
- 14:03:14That's from the uh very beginning of the
- 14:03:18notebook. So that's the that's from this
- 14:03:21one.
- 14:03:25Oh, this Oh, this is where it is. Sorry.
- 14:03:27This is where it is. You can find it
- 14:03:29here.
- 14:03:34That's right. It was from a URL.
- 14:03:40It was used in the very beginning of the
- 14:03:41notebook.
- 14:03:43And we did we did this.
- 14:03:47Okay.
- 14:03:49That's right. That's why I didn't have
- 14:03:50it downloaded.
- 14:03:55Okay.
- 14:03:59All right. Any other questions on the
- 14:04:01elastic before I move I'm going to move
- 14:04:03on to uh finding those a systematic way
- 14:04:07to find the best hyperparameters.
- 14:04:10Um, I'm going to show you a couple
- 14:04:11strategies to doing that. Um, so far
- 14:04:14we've just ran CV with some random
- 14:04:16choices. Um, I'm going to show you a
- 14:04:18better, more systematic approach. That's
- 14:04:20kind of the industry standard for doing
- 14:04:22tuning. Um, so I'm going to I'm going to
- 14:04:25show you that next, but any questions on
- 14:04:26the elastic net?
- 14:04:33Okay. And again like you know
- 14:04:35scikitlearn makes it really easy for you
- 14:04:37guys
- 14:04:39because it just behaves the same way as
- 14:04:41any other model. You use the object and
- 14:04:44then you do ffit and predict right? So
- 14:04:46the ffit is going to train it um and the
- 14:04:50predict is going to allow you to use
- 14:04:52that model to predict. It's it's super
- 14:04:54easy that way. Every scikitlearn model
- 14:04:56is like that dofitit and predict. So it
- 14:04:59provides a really simple way to use
- 14:05:01basically every model.
- 14:05:07Okay,
- 14:05:09let's talk about let's finish up this
- 14:05:11lesson with a couple things. Um, one of
- 14:05:14those things is going to be
- 14:05:15hyperparameter tuning. So what is this?
- 14:05:19The hyperparameter tuning is a
- 14:05:21systematic way to find the best
- 14:05:24parameters in a machine learning model.
- 14:05:28So a lot of machine learning models have
- 14:05:30what are called hyperparameters.
- 14:05:33These are not the betas that we learn
- 14:05:35during the training that's learned from
- 14:05:37the data. These are settings that we set
- 14:05:40ahead of time like the alpha. That's a
- 14:05:43perfect example like alpha L1 ratio in
- 14:05:45in the elastic net. We set those up
- 14:05:48ahead of time and depending on what we
- 14:05:50pick for those we get different
- 14:05:51performance, right? And so what we
- 14:05:54really need is a systematic way to find
- 14:05:57the best settings for those
- 14:06:00hyperparameters as we are training our
- 14:06:02models. Um the the the main like idea
- 14:06:08behind this process though is going to
- 14:06:10be to systematically try out different
- 14:06:14combinations as many as we want to try.
- 14:06:17And so we're we're basically going to
- 14:06:19have a strategy for tuning that is going
- 14:06:22to exhaust all the combinations of those
- 14:06:26hyperparameters that we want to try
- 14:06:28until we find the one that performs the
- 14:06:31best. Um and and that strategy is known
- 14:06:34as grid search. Um and essentially what
- 14:06:39it does is it sets up a grid um where
- 14:06:42which is basically like a matrix to say
- 14:06:45okay which parameters do you want to
- 14:06:47try? I want to try um alpha and I want
- 14:06:50to try L1 ratio
- 14:06:53um L1 ratio like let's say I want to try
- 14:06:57these two. So we set these up in a grid
- 14:06:59where we say, "Okay, I want to try this
- 14:07:01value. I want to try this value. I want
- 14:07:02to try this value. This one, this one,
- 14:07:04this one, and on and as many as we want
- 14:07:06to try." So we could set up set those up
- 14:07:09systematically like a linear um a
- 14:07:12linearly spaced like I want to try every
- 14:07:14alpha between between 0 and 10 spaced by
- 14:07:18one um whatever. You know, we can set up
- 14:07:21different ranges of those, but that's
- 14:07:23going to be in this grid. And then the
- 14:07:25L1 ratio, same thing. We can try out
- 14:07:27different values of these that we want
- 14:07:28to try. Let's say there's many of those.
- 14:07:32Um maybe every um tenth between 0 to one
- 14:07:36I want to try out. Um so you set up your
- 14:07:40parameters and you can set up as many as
- 14:07:41you want in the grid. And then
- 14:07:43essentially what you're going to do to
- 14:07:45do grid search is you're going to work
- 14:07:47your way through every combination of
- 14:07:49those. So you're going to try out this
- 14:07:50combo. You're going to try out this
- 14:07:53combo. You're going to try out this
- 14:07:54combo.
- 14:07:56this combo. So the first value of alpha
- 14:08:00with every possible L1 ratio, then go to
- 14:08:03the next, try out the next value of
- 14:08:05alpha with every L1 ratio, and on and on
- 14:08:08and on. So we're going to try
- 14:08:11all combos
- 14:08:14in the grid.
- 14:08:16We're going to try all combos and we're
- 14:08:18going to find the lowest MSE
- 14:08:22combination. find lowest
- 14:08:25MSE
- 14:08:27combo.
- 14:08:28So whatever leads to the best model is
- 14:08:31going to be the um parameters that are
- 14:08:34that are deemed to be the best. And the
- 14:08:36idea is once we have found those we know
- 14:08:40that we can use we can go ahead and
- 14:08:42train a model with those best alpha and
- 14:08:44len ratio and on and on and on.
- 14:08:54Yeah, when you get an So this goes back
- 14:08:56to the error. Remember that for a
- 14:08:59regression,
- 14:09:01the error is this measurement of how far
- 14:09:04off we are, right? So if we have a bunch
- 14:09:06of points and we draw we fit a line
- 14:09:08through there, the the MSE is measuring
- 14:09:12this distance, right? So what do you
- 14:09:14think is a good distance? Like if our
- 14:09:16model is perfect,
- 14:09:19what's the best distance from our
- 14:09:21predictions to the actual points? Zero.
- 14:09:25Yes. So the lower the better. The lower
- 14:09:29the better. Um so for an R RMSSE, the
- 14:09:32lower the closer to zero the better.
- 14:09:35However, the RMSSE can be it's its units
- 14:09:39are interpreted in the units of our
- 14:09:41target.
- 14:09:43So what is deemed to be good is relative
- 14:09:46to our target. Like let's say our target
- 14:09:48is in the thousands. Like it averages in
- 14:09:51the thousands. If we produce an MSE of
- 14:09:5450 or sorry an RMSSE of 50, that's
- 14:09:59pretty good, right? Because our units
- 14:10:01are in the thousands
- 14:10:03and we're only on average we are off by
- 14:10:0750 units,
- 14:10:09right? Our distance away is about 50
- 14:10:11units. That's pretty good. So the RMSSE
- 14:10:15is relative to your target variable.
- 14:10:18Does that make sense? Yeah. It depends
- 14:10:20on the target. It depends on what you're
- 14:10:22trying to predict.
- 14:10:25So that's why we got RMSSE that were in
- 14:10:27the 300s for those hitters, but the
- 14:10:29average was the average of the target
- 14:10:31was in the 500s. So that's a really bad
- 14:10:35proportion of error relative to the
- 14:10:37average target value. Right? If our
- 14:10:41RMSSE was 300,
- 14:10:43but the target was sitting in the 500s,
- 14:10:47that's just too much error. Way too much
- 14:10:50error, right? That's just too big of a
- 14:10:52value. Um, our predictions are just off
- 14:10:56way too much
- 14:10:58in terms of that distance. So, this
- 14:10:59would be this was a bad model. It was
- 14:11:02underfit.
- 14:11:04We know that from the the RMSSE. So
- 14:11:07yeah, the RMSSE closer to zero, no
- 14:11:09matter what is good,
- 14:11:12zero is being perfect. Um, but it to
- 14:11:16know what's good, you need to know what
- 14:11:18your target is on average and then think
- 14:11:20of this as kind of a ratio to that
- 14:11:23average target. I think that's the best
- 14:11:26way to think about it.
- 14:11:37Okay, so going back to this grid idea is
- 14:11:42so the grid is just basically laying out
- 14:11:44all possible parameter combinations and
- 14:11:48trying them all out by fitting and
- 14:11:50predicting until and generating an a
- 14:11:53metric like an MSE
- 14:11:55until we find the one with the lowest
- 14:11:58MSE. So find the lowest MSE combination
- 14:12:02and that will be the best
- 14:12:04that will be the best combo and then if
- 14:12:07we once we know that best combo we can
- 14:12:09use that we can use that alpha we can
- 14:12:11use that L1 ratio and use that model
- 14:12:14going forward we can we can use those
- 14:12:16parameters in our model so this strategy
- 14:12:20it has a name it's known as grid search
- 14:12:24so it is a hyperparameter tuning process
- 14:12:26that tries out all combinations S.
- 14:12:30So what's the what's the U benefit to
- 14:12:34this is that we get to test out a lot of
- 14:12:37different combo combos of those
- 14:12:38parameters like the alpha and L1. So we
- 14:12:41can be confident what the best model is,
- 14:12:43right? So we can pick the alpha and L1
- 14:12:46perfectly because we're trying out a
- 14:12:47bunch of different combinations on the
- 14:12:49data to see which one's the best. What's
- 14:12:52the downside?
- 14:12:54It's expensive, right? It's an
- 14:12:56exhaustive search. So if you have many
- 14:13:00different parameters and you're trying
- 14:13:03out many different combinations, it can
- 14:13:06get exponentially
- 14:13:08expensive
- 14:13:09to perform this search. Okay, so grid
- 14:13:12search is great except for the fact that
- 14:13:15it can be expensive if you have many
- 14:13:17parameters with with very wide ranges
- 14:13:20that you're searching over because that
- 14:13:21that's a lot of combinations you have to
- 14:13:23test, right? And especially if you have
- 14:13:26a lot of data, that's going to be
- 14:13:28expensive
- 14:13:30um to do.
- 14:13:33So, uh we're going to practice doing
- 14:13:35grid search, but that is that's the pro
- 14:13:37and con. The pro is that we get to try
- 14:13:38out all these combinations and see which
- 14:13:40one's the best. The downside is it can
- 14:13:43be expensive to do that if you have a
- 14:13:44lot of parameters um that you want to
- 14:13:47tune for your model um and you have very
- 14:13:52uh many different choices that you're
- 14:13:53trying to evaluate for those and it just
- 14:13:56creates a really big um collection of
- 14:13:59combinations that you have to try out,
- 14:14:02right? Um that's the only downside to
- 14:14:05grid search.
- 14:14:08Now on the opposite end of the spectrum
- 14:14:10of that is a randomized search or random
- 14:14:13search and this will basically just um
- 14:14:17do a sampling of those parameters from
- 14:14:23um kind of fixed uh specified
- 14:14:25distribution. So essentially what you do
- 14:14:28is similarly you define your range. So
- 14:14:31you say I want to look at alphas um
- 14:14:34between zero sorry between let's say
- 14:14:38yeah 0 to 10. I want to look at a bunch
- 14:14:40of different alphas. Um, and I want to
- 14:14:42look at a bunch of different L1 ratios
- 14:14:45that are between 0ero to one.
- 14:14:490 to one. And um, what we do is we say,
- 14:14:53okay, I'm going to restrict only testing
- 14:14:5720, 30, 40 times. I'm not going to do
- 14:14:59all possible combinations. I'm just
- 14:15:02going to randomly sample something in
- 14:15:04this range and randomly sample something
- 14:15:07in this range. And so and I'm going to
- 14:15:10perform that experiment a fixed number
- 14:15:12of times. So let's say I set the uh
- 14:15:16sampling where I'm only going to do um
- 14:15:1920 evaluations.
- 14:15:21And so 20 times we're going to pick a
- 14:15:24combo randomly. So I'm going to pick an
- 14:15:28alpha and I'm going to pick an L1 ratio.
- 14:15:33L1 ratio.
- 14:15:36And um we are we are just going to uh
- 14:15:40sample those randomly from this range.
- 14:15:44Um and we're going to use those and test
- 14:15:47those out and then it's but otherwise
- 14:15:49it's the same as grid search. Whatever
- 14:15:50is the lowest MSE
- 14:15:53um so whatever is the lowest MSE is the
- 14:15:56best.
- 14:15:57So we evaluate those. We sample we train
- 14:16:00the model. Evaluate it. Whatever is the
- 14:16:03lowest MSE
- 14:16:05is the best is the best combo. Now,
- 14:16:09what's the benefit to this is it's a
- 14:16:12much more controlled experiment in the
- 14:16:16sense that we um aren't going to iterate
- 14:16:18through every possible combination in
- 14:16:20the grid. We're we basically set up a
- 14:16:23fixed number of times we're going to try
- 14:16:24out stuff.
- 14:16:26The risk to doing this is that you're
- 14:16:28not you're not exploring all
- 14:16:31combinations, right? Because you're
- 14:16:33randomly sampling, you may get unlucky
- 14:16:36and you may not stumble into the best.
- 14:16:39You you can make um samples and figure
- 14:16:42out what's the best amongst your
- 14:16:43samples, but you may not be covering all
- 14:16:46the combinations. Does that make sense?
- 14:16:48The grid search is going to try every
- 14:16:50combo. The random search is going to
- 14:16:53randomly sample those combos.
- 14:16:56So, it's not going to try every single
- 14:16:58one. It's going to try a limited number,
- 14:17:00however many you set up. Now, if you set
- 14:17:03that number really, really, really high.
- 14:17:06Now, you're starting to approach a grid
- 14:17:07search because now you're sampling so
- 14:17:10many of those combos that you basically
- 14:17:12are trying them all at that point,
- 14:17:16right? Um,
- 14:17:20so, so that's the way the random search.
- 14:17:22So by the way, both of these use cross
- 14:17:24validation in the sense that when you
- 14:17:27evaluate accommodation, you're actually
- 14:17:29doing it with cross validation. So when
- 14:17:31you do an evaluation, you're going to do
- 14:17:34probably 10 or five folds where you
- 14:17:36split your data and then you test it on
- 14:17:39the rest of the folds and evaluate or
- 14:17:41train it on the rest of the folds,
- 14:17:42evaluate it on one of them and generate
- 14:17:44an average MSE to get your evaluation.
- 14:17:49So every evaluation is using cross
- 14:17:51validation.
- 14:17:52That's why that's and hopefully you can
- 14:17:54see why this would be so expensive for a
- 14:17:56really big grid, right? Because you're
- 14:17:58trying out many different combinations
- 14:18:02and every combination is going to do a
- 14:18:04cross validation procedure. So it's
- 14:18:07going to train 10 times and test against
- 14:18:1010 different folds and average those
- 14:18:12together. it's going to be a pretty
- 14:18:13expensive operation
- 14:18:15for a really big grid, right, of of
- 14:18:18parameters.
- 14:18:19Um, but these are the two kind of
- 14:18:21systematic approaches we have at trying
- 14:18:24out different hyperparameters. Remember
- 14:18:27those those things are called
- 14:18:28hyperparameters. These are those choices
- 14:18:31that we have before we train our model.
- 14:18:34Um, those choices we have that affect
- 14:18:36the performance of the model like the
- 14:18:38alphas, the L1 ratios, those kind of
- 14:18:40things. um we have control over what
- 14:18:43they're going to be. This is a
- 14:18:44systematic approach to find out what the
- 14:18:46best
- 14:18:48uh value of those parameters is going to
- 14:18:51be on our data.
- 14:18:53Right?
- 14:18:57Okay. So before we practice this, we're
- 14:18:59going to practice with grid search
- 14:19:01first. Um
- 14:19:04any questions?
- 14:19:14Uh, I don't know if it has a built-in
- 14:19:16That's a good question. By time limit, I
- 14:19:17don't know if it has a built-in way of
- 14:19:18doing it, but you could certainly set up
- 14:19:20like a a a loop um to like to wrap
- 14:19:25around. Do you know what I mean? Like
- 14:19:26you could set up a loop where you check
- 14:19:28the time if it's if if the time elapsed
- 14:19:31as you're doing a search if the time
- 14:19:32elapsed is greater than the the time
- 14:19:35limit then you can kind of break early.
- 14:19:38Um so it's not hard to implement that
- 14:19:40but I don't know if it has that built
- 14:19:42in. I don't think it does
- 14:19:44because I don't think it really cares
- 14:19:46how long every evaluation takes. It's
- 14:19:48just going to exhaust all those
- 14:19:50especially in a grid search.
- 14:19:53But um yeah, I there's probably a way to
- 14:19:56manually kind of set up a time time
- 14:19:58loop.
- 14:20:04So hyperparameters are um settings that
- 14:20:08we have on the model itself. And a
- 14:20:11really good example of this is like the
- 14:20:13alpha and L1 ratio in the in the elastic
- 14:20:15net. So they're not things that we um
- 14:20:20learn from the data directly like the
- 14:20:22betas in the model like those get
- 14:20:24trained directly by doing the um least
- 14:20:28squares process right um by doing that
- 14:20:31gradient descent and all that
- 14:20:32optimization.
- 14:20:34Um so these are not learned from that.
- 14:20:36They're actually set ahead of time. And
- 14:20:39so what we're saying is the best way to
- 14:20:42understand the effects of those is to
- 14:20:44try out different combinations of those
- 14:20:46until we land on the best one. Right? So
- 14:20:49hyperparameters are those options we
- 14:20:51have in the model like the alpha like
- 14:20:54the alpha and l1 ratio in the uh elastic
- 14:20:57net. Many models have hyperparameters.
- 14:21:01Um we're actually going to see that in
- 14:21:03in future models that we study. they
- 14:21:05have options that you can set that
- 14:21:07affect their performance.
- 14:21:09And so this this is just a strategy to
- 14:21:11evaluate those different options to see
- 14:21:13which one's the best.
- 14:21:25Yeah. So again, hyperparameters, those
- 14:21:28are settings on the model itself um that
- 14:21:32affect the performance of it.
- 14:21:36And basically we have the two two
- 14:21:38strategies here. We can set up an
- 14:21:40exhaustive grid and search through all
- 14:21:41of those until we find the lowest MSE uh
- 14:21:44option or we can randomly sample
- 14:21:48potential options, try them out and see
- 14:21:50which one's the lowest as well. And do
- 14:21:52that a fixed number of times. Um
- 14:21:56sort of like a fixed number of trials
- 14:21:58almost. um which has a risk of not
- 14:22:02trying out every option but but
- 14:22:04hopefully you try out enough that you've
- 14:22:06explored the space a bit and you get
- 14:22:09some quality choices there but no
- 14:22:12guarantees right no guarantees you try
- 14:22:14everything which is what a grid search
- 14:22:15will do it will try everything
- 14:22:20okay now luckily per usual scikitlearn
- 14:22:26has something to manage this process for
- 14:22:28us in terms of grid search. Um so in
- 14:22:33that way we will not need to manage this
- 14:22:36process ourselves. We can just rely on
- 14:22:37scikitlearn. And so if you're doing
- 14:22:40hyperparameter tuning um this is going
- 14:22:42to come from the model selection module
- 14:22:45inside of sklearn. So we're going to
- 14:22:48import from from sklearn the model
- 14:22:50selection module. We have our grid
- 14:22:52search cross validation.
- 14:22:55Okay, that's what the CV stands for.
- 14:22:57grid search cross validation. So, this
- 14:23:00is going to do that grid search
- 14:23:01strategy. Um, we're going to set it up
- 14:23:03with our dictionary essentially of
- 14:23:06choices. So, we're going to say, hey,
- 14:23:08here's the alphas I want to try. Here's
- 14:23:09the L1 ratios I want to try. Um, and
- 14:23:12here's my other settings like uh how
- 14:23:15many folds I want to use, what my random
- 14:23:17state is for the shuffling. So, we'll
- 14:23:19set all that up. Um,
- 14:23:22and then we'll just run the grid search.
- 14:23:24And then what should come out of that is
- 14:23:26the best options for our parameters from
- 14:23:29the grid and then we can use those going
- 14:23:31forward in the we can build a model with
- 14:23:34those best options right so we're really
- 14:23:37doing some evaluation here of what is
- 14:23:39going to be those best alphas those best
- 14:23:410 to1 ratios on our data set right and
- 14:23:44the only way to really know that is to
- 14:23:47evaluate them because they're not things
- 14:23:48that are learned during the training
- 14:23:51hopefully that makes sense right they're
- 14:23:53not things that we learn directly from
- 14:23:55training. There are things that we have
- 14:23:57to set and then kind of evaluate and see
- 14:23:59how they affect things.
- 14:24:03Okay, so we have grid search CV. That's
- 14:24:05going to be our primary um tool to do
- 14:24:08the evaluations of the different
- 14:24:10hyperparameter options.
- 14:24:13Grid search CV. Um we're going to set up
- 14:24:15our cross validation uh object here. Now
- 14:24:19I want you to pay attention to this is
- 14:24:21that um it's a slightly different
- 14:24:24version than the kfold we had earlier.
- 14:24:26So we've used k-fold before with a
- 14:24:27certain number of folds. This would be
- 14:24:2910 folds and we can set a random state
- 14:24:32for the shuffling um that happens in the
- 14:24:34folds.
- 14:24:36But this is actually a slight different
- 14:24:37variation on it where it is a repeated
- 14:24:39kfold where we do three repeated trials.
- 14:24:43Now why would we do that? It's to be
- 14:24:46extra extra extra careful with the
- 14:24:49shuffling.
- 14:24:51So this what this means is we do three
- 14:24:52different shuffles. So we do kfold, we
- 14:24:56actually repeat it three times with
- 14:24:58three different shufflings. That's all
- 14:24:59that means. So the repeated kfold is
- 14:25:03actually a bit beyond the just basic
- 14:25:06kfold. What basic kfold will do will
- 14:25:09we'll will shuffle and then do our
- 14:25:12splits into 10 splits and then train on
- 14:25:14nine of those. test on the other one and
- 14:25:16rotate through all the splits.
- 14:25:19We're actually going to do that process
- 14:25:21three different times with three
- 14:25:24different shuffles. So this and we're
- 14:25:26going to average 30 results instead of
- 14:25:29just 10. So repeated kfold is just going
- 14:25:33above and beyond to do extra to repeat
- 14:25:36the kfold three different times. In this
- 14:25:38case only three. We could do more,
- 14:25:41but um now is that necessary to do? You
- 14:25:44could argue not necessarily. Um but it
- 14:25:47just provides extra robustness
- 14:25:50uh beyond just our single shuffle and
- 14:25:53then split and then rotation of those
- 14:25:55folds, right? We're doing it actually
- 14:25:57three different shuffles. Um so we're
- 14:26:00repeating our kfold three times uh for
- 14:26:03every now is the thing is we're doing
- 14:26:05that for every evaluation. So it is
- 14:26:07going to be more expensive than just a
- 14:26:09basic K-fold.
- 14:26:25So we have three different K-fold trials
- 14:26:28that we're doing essentially.
- 14:26:31Okay, hopefully that makes sense. This
- 14:26:33is the repeated K-fold. We haven't
- 14:26:35really seen that before. we've only
- 14:26:36worked with the Kfold, which would get
- 14:26:38rid of this repeats option and only have
- 14:26:41uh 10 splits in a random state for the
- 14:26:44for the single shuffle that we do. So,
- 14:26:46we can um recreate that same shuffle
- 14:26:48every time. Um but now we're actually
- 14:26:51going to do three random shuffles, uh
- 14:26:54three different trials. So, one shuffle
- 14:26:56creates the and then create the 10
- 14:26:58splits, evaluate, then go back and do
- 14:27:00another shuffle, another new 10 splits.
- 14:27:03So, one thing that should be um clear is
- 14:27:06that we get different splits every time
- 14:27:09because we're going to shuffle once,
- 14:27:12right? We're going to shuffle once and
- 14:27:14generate our splits
- 14:27:16and then we're going to shuffle again,
- 14:27:18generate these splits which are going to
- 14:27:19be different and then shuffle one more
- 14:27:21time for for three different times,
- 14:27:23right? And then get get these splits and
- 14:27:26then we're going to get 10 metrics here,
- 14:27:2810 metrics here, 10 metrics here, and
- 14:27:30then average all of those together.
- 14:27:34So, it's a bit more just going up extra
- 14:27:37above and beyond for a K-fold. Okay.
- 14:27:42All right. So, here comes the fun of
- 14:27:44when you do grid search. Now, the grid
- 14:27:48is actually just a dictionary. It's a
- 14:27:50Python dictionary where you declare what
- 14:27:54your parameters are going to be inside
- 14:27:55the dictionary and you set up a range of
- 14:27:58values that you're go or a list. It can
- 14:28:02be a list. It can be a range
- 14:28:04but some declaration of what you are
- 14:28:07going to test and evaluate inside of
- 14:28:09your grid search. So the grid is
- 14:28:12initialized as an empty dictionary.
- 14:28:15And then what we do is we say okay in my
- 14:28:18grid I want to check different alphas.
- 14:28:20So we're going to add a collection of
- 14:28:22alphas in here that we're going to test.
- 14:28:26So let me make a comment there. We add a
- 14:28:30add a range of alphas to test. And this
- 14:28:36range is a this is just like the Python
- 14:28:40range. Um
- 14:28:43this is just like a Python range um uh
- 14:28:46operator here where this is going to be
- 14:28:49uh every so it's going to be um every
- 14:28:54uh value between
- 14:28:57zero and one um uh steps with a step
- 14:29:03size
- 14:29:05of 0.1. So, it's going to try a bunch of
- 14:29:09different alphas um between uh zero and
- 14:29:130.1
- 14:29:14sorry 0 and one stepping by 0.1. So,
- 14:29:16it's going to try zero.1
- 14:29:182.3 point 4.5 6 right all the way up to
- 14:29:22one.
- 14:29:24So, that's what this will do. And it's a
- 14:29:25numpy range. So, it's just all those
- 14:29:27decimals between 0 to one.
- 14:29:31It you can use either that's valid.
- 14:29:34Yeah, you can do you can do that to
- 14:29:36create a dictionary or you can use the
- 14:29:38keyword um dict. You can use either one.
- 14:29:41Either one works.
- 14:29:45Whatever whatever you want to use.
- 14:29:46They're the same.
- 14:29:49Yeah. The the reason people prefer
- 14:29:52dictionary is because um sets are
- 14:29:56created with the same braces.
- 14:29:59So it it makes it clear what you're
- 14:30:01creating as a dictionary. If you use if
- 14:30:03you use this, that's the only advantage
- 14:30:06is it's just plainly obvious what you're
- 14:30:08making. Uh because technically you can
- 14:30:10make a set with the curly braces as
- 14:30:13well.
- 14:30:18Yeah,
- 14:30:21no worries. Um okay, so we have our
- 14:30:24alphas here. So what I want you to
- 14:30:27notice is that we are going to try out
- 14:30:29different alphas and we are that's the
- 14:30:31only parameter we are going to try in
- 14:30:33our ridge regression. So we're going to
- 14:30:36we're going to try ridge but just try
- 14:30:38different alphas in the in this range um
- 14:30:41in our grid search. So the grid search
- 14:30:44CV takes in a model. It takes in our
- 14:30:47grid dictionary which is really
- 14:30:48critical. We need that dictionary to
- 14:30:50declare what we're going to try.
- 14:30:53um we need a scoring to say to find the
- 14:30:56best. Now remember it uses the negative
- 14:30:59to find the lowest which is going to be
- 14:31:02the the least negative option.
- 14:31:06Um otherwise it wouldn't um just based
- 14:31:09on the optimization it would look for
- 14:31:11the highest value. Um so the highest
- 14:31:14would be closest to zero in this
- 14:31:15situation. Um so we use negative and
- 14:31:19again we could use squared error. It's
- 14:31:21using absolute. We could use um squared
- 14:31:25uh either either one works.
- 14:31:29Um more typical would probably be
- 14:31:31squared error, but um absolute is fine.
- 14:31:35Here's where we have our repeated kfold.
- 14:31:37So we pass in our um how we're doing CV.
- 14:31:40That can be it can be a kfold object. It
- 14:31:42can actually just be an integer, which
- 14:31:44is say I just want to do 10 splits or
- 14:31:46five splits um to to do every
- 14:31:49evaluation. But these are the bare
- 14:31:51minimum that you need. Just really the
- 14:31:53model and the grid and your CV. Um what
- 14:31:57metric you're using to evaluate what's
- 14:31:59going to be the best. And then this end
- 14:32:01jobs is to parallelize. If you have it
- 14:32:03set to minus one, it's going to it's
- 14:32:04going to try out all the grid options in
- 14:32:06parallel. Um which is nice. It's going
- 14:32:09to help speed up the overall search.
- 14:32:12Okay. So let me mark that down as n
- 14:32:17jobs equals minus one.
- 14:32:21tries out the combos in parallel.
- 14:32:26So in this situation, we actually don't
- 14:32:28have more than one parameter. We only
- 14:32:30have the alpha. So we're really just
- 14:32:32going to be systematically working our
- 14:32:34way through every alpha and evaluating
- 14:32:36which one's the best right with this.
- 14:32:39And notice that in order to use this
- 14:32:41grid search, all we have to do is call
- 14:32:43search.fit. So it works kind of like
- 14:32:46every other model does, right? It's the
- 14:32:49grid search.fit.
- 14:32:52And we pass in our data.
- 14:32:54And we um once we're once this prints
- 14:32:58out the results, you get a results
- 14:33:00object um which has a best score and
- 14:33:04then a dictionary with your best
- 14:33:06parameters. So, whatever your best grid
- 14:33:08member was or grid members, um it prints
- 14:33:12that out and you can So, for from that,
- 14:33:14we can um grab our best alpha, which
- 14:33:18which let's confirm what that ends up
- 14:33:20being.
- 14:33:26Oops. We need to import repeated kfold.
- 14:33:37So we'll import that.
- 14:33:44Oh, I didn't. Let's do from
- 14:33:47sklearn.linear
- 14:33:52model import ridge.
- 14:33:56Okay.
- 14:34:04Okay. So, it completed the search and
- 14:34:06what we found is this is the best score
- 14:34:08is 238 for the mean absolute error and
- 14:34:12the best alpha that we got was 0.9. So,
- 14:34:16the best alpha that worked here, the one
- 14:34:19that gave us the best score was actually
- 14:34:210.9 as the alpha. So what it did is it
- 14:34:24tried out everything between this range
- 14:34:27and 0.9 was the best. So it did cross
- 14:34:30validation, tried out every single combo
- 14:34:33in our grid.
- 14:34:35So if we want we could actually print
- 14:34:37out
- 14:34:40print our grid so we can see
- 14:34:43what our combinations were.
- 14:34:49So, it tried out all of these guys and
- 14:34:51the best one that we had was 0.9.
- 14:35:00Okay, so pretty cool how that works. And
- 14:35:03if we had other parameters, like if we
- 14:35:06were doing a elastic net, we could add
- 14:35:08those into our dictionary and it would
- 14:35:10do all combinations of those. So if we
- 14:35:13did um so for instance to add to our
- 14:35:16grid we could do grid
- 14:35:18um L1 ratio
- 14:35:21this would be for like an elastic net
- 14:35:22right now the ridge regression by itself
- 14:35:24doesn't have an L1 ratio parameter but
- 14:35:26just as an example um we could try out
- 14:35:29different ranges um similar range
- 14:35:32different one um maybe an exact list
- 14:35:35whatever we want to do. So this is going
- 14:35:37to try out different ones between 0ero
- 14:35:38to one
- 14:35:40as well. And so it's going to try out
- 14:35:42every combination of these from this
- 14:35:45grid.
- 14:35:47Okay, if we did that. But again, this
- 14:35:49the ridge regression doesn't have an L1
- 14:35:52ratio. The elastic net does. So that the
- 14:35:55ridge regression only has an alpha to as
- 14:35:58a hyperparameter. So we're only testing
- 14:36:00out that one.
- 14:36:05Okay. So that's grid search CV.
- 14:36:09Pretty useful. This is pretty useful in
- 14:36:11doing parameter tuning again when you
- 14:36:13want to try out ranges of different
- 14:36:15values and you can evaluate those to see
- 14:36:18which one is your best and then we can
- 14:36:21use that best going forward. So we can
- 14:36:23for instance this is what this code does
- 14:36:26below it is it fetches the best. Um you
- 14:36:29can do it this way or you can do it um
- 14:36:32the alternative is to do results.b best
- 14:36:34params
- 14:36:38and then you can just grab it like this
- 14:36:41alpha.
- 14:36:42Either way you can do get or like this
- 14:36:46um and it this is just a dictionary,
- 14:36:48right? And you can grab your alpha. So
- 14:36:50that's the 0.9 um and we can pass that
- 14:36:53alpha into the ridge regression and go
- 14:36:56back and refit it to our data um and
- 14:36:59then use that model going forward. So
- 14:37:01the grid search really just evaluates
- 14:37:04those different options, tells you
- 14:37:06what's the best according to this score,
- 14:37:10right?
- 14:37:12And you should, by the way, you should
- 14:37:14interpret this score in the positive
- 14:37:16sense. It's only negative because we're
- 14:37:19purposely making it negative to find out
- 14:37:22what the lowest option is, right?
- 14:37:24Because the lower is the better. So we
- 14:37:26we purposely make it negative to make it
- 14:37:28whatever is the least negative is the
- 14:37:30winner. Um more negative is worse.
- 14:37:35So it's really positive version of it is
- 14:37:39the is the true result for the error. Um
- 14:37:42and they are a tool from scikitlearn to
- 14:37:45put together your model with your
- 14:37:47pre-processing steps. So they kind of
- 14:37:49get automated together. Um and they
- 14:37:52combine everything into kind of a
- 14:37:54streamline process. You're going to see
- 14:37:56what that looks like, but it's a really
- 14:37:58nice um feature of scikitlearn. Um why
- 14:38:02would we care about pipelines? They help
- 14:38:05organize our code um so that we ensure
- 14:38:08that we basically always run the
- 14:38:10pre-processing steps before we train and
- 14:38:12use a model to with the predictions. Um,
- 14:38:15so it bundles those steps together,
- 14:38:17minimizes the risk of forgetting a step
- 14:38:20because one of the things that can
- 14:38:21happen is when you do pre-processing, if
- 14:38:24you're doing it on the training set, you
- 14:38:25have to do it on new test data as well
- 14:38:27when you put it through your model
- 14:38:29because your model is training against
- 14:38:30that pre-processed data.
- 14:38:33So in order to make sure you never
- 14:38:35forget that, you can bundle it all
- 14:38:37together in a pipeline which is going to
- 14:38:39make things really really easy to use
- 14:38:42and and make sure that those steps
- 14:38:44happen in a repeatable way. Um and it
- 14:38:49makes things easier to uh deploy that
- 14:38:52model as well because everything is
- 14:38:54together in one pipeline. So in the in
- 14:38:57the industry, I've seen this a lot. Um
- 14:39:00you know, people will do their initial
- 14:39:03exploration steps and initial model
- 14:39:05building. They may not use pipelines
- 14:39:07right away, but as they found their
- 14:39:10model, um they'll generally move it into
- 14:39:13a pipeline and all their steps into a
- 14:39:14pipeline so that it's uh easier to work
- 14:39:16with um when you're when you're
- 14:39:18deploying it, actually using it uh in in
- 14:39:22the real world. Um, so here's what a
- 14:39:25pipeline generally looks like. It's from
- 14:39:27scikitlearn. It's this pipeline object.
- 14:39:30Um, and the pipeline is made up of steps
- 14:39:33that we're going to see that that are
- 14:39:35various um uh basically um kinds of
- 14:39:40pre-processing we've seen before like a
- 14:39:42scaler or um filling in missing values.
- 14:39:46Those kind of things we can put here in
- 14:39:48the steps which is basically a list. um
- 14:39:51steps is just going to be a list of
- 14:39:53scikitlearn functions that we can apply
- 14:39:54to data. One of those being a model. Um
- 14:39:58and then whenever we use the pipeline,
- 14:40:00it's basically um you know, it's going
- 14:40:02to be something like pipeline.fit
- 14:40:05or pipeline.predict.
- 14:40:07So the pipeline kind of behaves like a
- 14:40:10model. It's just going to contain many
- 14:40:12more steps than that like the
- 14:40:14pre-processing steps we've worked with
- 14:40:16before. Um, and it also has some
- 14:40:19capabilities for caching. So you can
- 14:40:21like uh cache some of the data in
- 14:40:24memory. Um, so that if you're reusing
- 14:40:26the predictions, it kind of goes faster.
- 14:40:29Um, so there's some options for that
- 14:40:31too. I'm not too concerned about that at
- 14:40:33this stage, but the main thing is going
- 14:40:35to be filling out our steps and then
- 14:40:37using the pipeline.
- 14:40:40Okay.
- 14:40:42Um, so some important bits of
- 14:40:45information about the pipeline is that
- 14:40:46it is going to be a sequence of data
- 14:40:48transformations that will have at the
- 14:40:51very end of the pipeline the model
- 14:40:53because of course we're going to do
- 14:40:55transformations and then train a model
- 14:40:58or predict with a model. So every
- 14:41:03Oh, can you guys hear me? Okay,
- 14:41:06not able to hear me. Thanks for letting
- 14:41:08me know. Can you guys were you able to
- 14:41:09hear me so far?
- 14:41:14Okay. Make sure. Yeah, it might be on
- 14:41:16your internet or your your uh Yeah, it
- 14:41:20seems like seems like it's good. So,
- 14:41:24no, you can't hear me. Check your
- 14:41:26volume. Check your headphones if you're
- 14:41:28wearing headphones. Oh, no issues. Okay,
- 14:41:31perfect.
- 14:41:32Okay. Yeah, local internet issue. Yeah.
- 14:41:37Okay.
- 14:41:39Always let me know. always let me know
- 14:41:40cuz it could be the case that it is me.
- 14:41:43So, um always always make sure to let me
- 14:41:46know. Um but sounds like yeah, you may
- 14:41:49want to check on that. Um
- 14:41:52so, okay. What I was saying is every
- 14:41:55pipeline's going to have a uh a sequence
- 14:41:57of steps that go first and then the
- 14:41:59model at the end. Um so, the order
- 14:42:02really matters. Um
- 14:42:05uh so the order matters in the sense
- 14:42:08that we want our transformations to go
- 14:42:09first. Things like scaling, things like
- 14:42:12filling in missing values, we want those
- 14:42:13to be first and then we want our uh
- 14:42:17model to be last because we want those
- 14:42:19transformations to happen prior to
- 14:42:21training or prior to prediction. So
- 14:42:24usually what you'll see in these
- 14:42:26pipelines is a model at the end, right?
- 14:42:28a model that's going to be at the end of
- 14:42:31the pipeline because we want basically
- 14:42:34our processing steps then our training
- 14:42:36or our processing steps then our
- 14:42:38predictions. Um so everything in the
- 14:42:42pipeline though is going to be from
- 14:42:43scikitlearn. Uh that's how it gets
- 14:42:46automated in the sense that all of those
- 14:42:48things are going to have fit and
- 14:42:49transform functions built into them so
- 14:42:51the pipeline can use them. Uh, and then
- 14:42:54the last step is going to be a model
- 14:42:56that has a fit and a predict. So it's
- 14:42:59pretty standard that the last part of
- 14:43:00the pipeline is just going to be a
- 14:43:01model. Um,
- 14:43:05uh,
- 14:43:06so we can um, as we do more modeling,
- 14:43:11we're going to play around with the
- 14:43:12pipelines quite a bit and see how we can
- 14:43:14change up some of the parameters. like
- 14:43:15if we want to change a model's parameter
- 14:43:18um we can actually adjust it to do
- 14:43:20things like uh grid search or cross
- 14:43:22validation. So um we're going to see
- 14:43:26some examples of some pipelines but for
- 14:43:28right now mostly what we're going to see
- 14:43:30is how to build one and then how to use
- 14:43:32one. And then as we get into lesson
- 14:43:35four, we'll get some more practice with
- 14:43:37pipelines cuz we're going to start using
- 14:43:38them quite a bit uh to build our models
- 14:43:41rather than do manual steps uh all the
- 14:43:45manual pre-processing
- 14:43:47um and then kind of building a model
- 14:43:49from there. We'll just include all of it
- 14:43:51together in a pipeline.
- 14:43:55Okay, so the example we're going to do
- 14:43:56is with this housing with ocean
- 14:43:59proximity. So we've actually looked at
- 14:44:00this data set before. Um so we have uh
- 14:44:05this ocean proximity data set that has
- 14:44:07the feature of like how close it is to
- 14:44:09the ocean like the bay or the less than
- 14:44:111 hour. Remember we had that and it had
- 14:44:14the median house value for different
- 14:44:16neighborhoods. Um so we're going to work
- 14:44:18with that one again. Let me make sure I
- 14:44:20have that one uploaded.
- 14:44:24You guys should have this one. It should
- 14:44:25be in your uh data sets.
- 14:44:29Um, I'll I can upload it here in case
- 14:44:31you don't have it though.
- 14:44:42Does this use multi-threading? I think
- 14:44:44it does. Yeah, I think in order to do it
- 14:44:46can do uh um I think it can do
- 14:44:49processing in parallel for some of the
- 14:44:51pipeline steps. Um, now does it use that
- 14:44:54all the time? Not necessarily because
- 14:44:57some of it is sequential in nature where
- 14:44:59you have to do one step and then you do
- 14:45:01the next step and then you do the next
- 14:45:03step. So it's not like you can do them
- 14:45:04in parallel.
- 14:45:06Um in terms of the like you need to know
- 14:45:08the output of one step to compute the
- 14:45:10the output of the next step. Um so it
- 14:45:15can but it it doesn't always lend itself
- 14:45:18well. The thing that will use
- 14:45:20multi-threading is is like the training
- 14:45:22process could be parallelized
- 14:45:25like the fit um can be for some models
- 14:45:29it can be parallelized not every model
- 14:45:35it so long answer is or the short answer
- 14:45:38is that it depends
- 14:45:40depends on what kind of transforms
- 14:45:41you're doing and what kind of model
- 14:45:42you're using if you can really take
- 14:45:44advantage of
- 14:45:50Okay. So, we load our data here and take
- 14:45:53a look at that. Um, do you guys have
- 14:45:56this data set? Are you able to load it
- 14:45:58in? If you're following along, are you
- 14:46:00able to load it?
- 14:46:08Okay.
- 14:46:10And and again, we've worked with this
- 14:46:11data before, so hopefully it's somewhat
- 14:46:13familiar. Remember, every row represents
- 14:46:16a neighborhood and it has a we're going
- 14:46:17to end up trying to predict this median
- 14:46:20house value as our target um variable,
- 14:46:24our dependent variable. Um and we're
- 14:46:26going to use the rest of these features.
- 14:46:28Remember that um this feature is in
- 14:46:31particular going to need to be one hot
- 14:46:33encoded,
- 14:46:35right? It's going to be one hot encoded
- 14:46:37because it is currently a string and we
- 14:46:39need to turn that into a numerical
- 14:46:42feature which is the one hot encoded
- 14:46:43feature. So we're going to have to do
- 14:46:46that but we're going to do that as part
- 14:46:48of our pipeline.
- 14:46:50Okay. So we'll be able to include that
- 14:46:52in our pipeline steps uh to to do one
- 14:46:55hot encoding which is nice.
- 14:46:59All right. So we're going to split apart
- 14:47:00our data um as we normally do. So we're
- 14:47:04going to uh create our feature uh data
- 14:47:08frame which is everything but this
- 14:47:10median house value. So we go ahead and
- 14:47:12drop that column and then our target is
- 14:47:14the median house value. So it is just
- 14:47:16that column here. Pretty standard. Um
- 14:47:20and then we're going to train test split
- 14:47:23and um split it into 30%
- 14:47:27uh test data. And again random state you
- 14:47:30can choose whatever you want to be. that
- 14:47:31just affects the shuffling. Um, so
- 14:47:34whatever doesn't really matter what it
- 14:47:36is. It's just so that when you rerun
- 14:47:37this, you get the same result in the in
- 14:47:39the shuffle.
- 14:47:42Okay, so we have our train and our test.
- 14:47:47So you want to make sure you run those.
- 14:47:50All right, so what we're going to do is
- 14:47:52take a look at our data
- 14:47:54and see if we have any null values. Um
- 14:47:58if you guys remember this data actually
- 14:48:01did have null values. You can see it
- 14:48:02here in this this guy and exactly how
- 14:48:05many there are is from this the sum. So
- 14:48:08we have um 162 nles in in this data. Uh
- 14:48:13and this is just a training data. So of
- 14:48:15course you know the test data could have
- 14:48:17that in there as well. Um so that's
- 14:48:19something we're going to want to make
- 14:48:21sure we fill in the blanks on any data
- 14:48:23set we use whether we're using the
- 14:48:25training or test set. Um, like if we're
- 14:48:27doing training, we want to make sure
- 14:48:29that gets filled in. If we're doing
- 14:48:30predictions with the test set, want to
- 14:48:32make sure that gets filled in. Um, so we
- 14:48:35we should be doing that. Um, now
- 14:48:41what we're going to do is use this data
- 14:48:44to help uh train our pipeline or or use
- 14:48:49with our pipeline. We need to construct
- 14:48:51our pipeline. So far, we've just split
- 14:48:53apart our data. We haven't done anything
- 14:48:55with our processing steps in our model
- 14:48:57yet. Um so roughly
- 14:49:02it this should be the flow of our
- 14:49:03pipeline. What should happen is we
- 14:49:05should be doing some type of feature
- 14:49:07scaling
- 14:49:08um some type of uh feature um
- 14:49:12manipulation. So that could be
- 14:49:13engineering, that could be um that could
- 14:49:17be uh doing the one hot encoding. Um so
- 14:49:21extracting new features like one hot
- 14:49:23encoding,
- 14:49:26one hot encoding. Um we are going to be
- 14:49:29doing that and and by the way, this is
- 14:49:31split up into this is when we use our
- 14:49:33pipeline for training.
- 14:49:36Um it's going to look like this where we
- 14:49:37do our scaling, we do one hot encoding,
- 14:49:40um we have our model here. Um, so that
- 14:49:43could be a linear regression, that could
- 14:49:44be a lasso, that could be a ridge, it
- 14:49:46could be elastic net. Whatever model we
- 14:49:48end up using is going to be last in the
- 14:49:50pipeline. And we're going to run this
- 14:49:53pipeline. Ultimately, we're going to run
- 14:49:55pipeline.fit,
- 14:50:00right? We're going to run a fit function
- 14:50:02and we get a fitted model as the result
- 14:50:04of this pipeline.
- 14:50:07Then when we use it when we use our
- 14:50:10model for prediction,
- 14:50:14we use our model for prediction in this
- 14:50:16lower part. It's the same pipeline, same
- 14:50:19exact pipeline, but it's this model has
- 14:50:22now been trained.
- 14:50:24So we now have a trained model here. So
- 14:50:26the great thing about the pipeline is
- 14:50:28it's the same this is the same pipeline
- 14:50:30that we're using here. So it's just
- 14:50:33going to it's going to repeat those same
- 14:50:35transformations. It's going to do our
- 14:50:37scaling. It's going to do our one hot
- 14:50:39encoding. It's going to use our model
- 14:50:41and it's going to generate predictions
- 14:50:43and generate uh we can we can do
- 14:50:45predictions. We can do evaluation like
- 14:50:47in a cross validation. Um we can use it
- 14:50:50however we want to use it. Uh but notice
- 14:50:54that the pipeline makes it consistent
- 14:50:57between training and test. We're using
- 14:50:58the exact same transformations
- 14:51:01and the model is last. It's it's either
- 14:51:04being trained or it's being used for
- 14:51:05prediction, but it's last. Our
- 14:51:07transformations are upfront, which are
- 14:51:10things like our scaling, things like our
- 14:51:11one hot encoding, right? Those happen
- 14:51:14first. No matter what data we put
- 14:51:17through there, we put our training data
- 14:51:19through there, we put our test data
- 14:51:20through there, they're going to go
- 14:51:21through the same steps,
- 14:51:25right?
- 14:51:28So that's that's the design of the
- 14:51:29pipeline. That's what it's supposed to
- 14:51:31do. So our job is to create those steps.
- 14:51:36So we need to create those relevant
- 14:51:38steps and then put them together into
- 14:51:40this pipeline. Okay. So that's going to
- 14:51:43be the code we're going to see coming up
- 14:51:44is we're going to build out these steps
- 14:51:47and then put them together into the
- 14:51:48pipeline.
- 14:51:58Um any questions on this diagram? Does
- 14:52:00it make sense what we're trying to do
- 14:52:02with this pipeline? We want to have
- 14:52:04repeatable steps during the training,
- 14:52:06during a prediction process.
- 14:52:09Okay.
- 14:52:15All right.
- 14:52:21All right. So, um, a couple of things
- 14:52:24we're going to need is, uh, to first of
- 14:52:27all, let's jot down what steps we're
- 14:52:29actually going to do. We're going to
- 14:52:30need to deal with missing values. So,
- 14:52:32we're going to fill in we're going to
- 14:52:33need a pre-processing pre-processing
- 14:52:36step that fills in any nulls. We always
- 14:52:39need that, right? So, if there's nles,
- 14:52:42we're going to fill them in somehow.
- 14:52:44We're going to define how we do that in
- 14:52:46our in our step. Um and we also need to
- 14:52:50one hot encode and we need to scale
- 14:52:53right those are pretty standard steps
- 14:52:56that we've dealt with whenever we're
- 14:52:57building these models right so pretty
- 14:52:59standard things fill in nles one hot
- 14:53:02encode any categorical data whatever
- 14:53:04however much we have and then go ahead
- 14:53:07and um standardize which is the scaling
- 14:53:10so this this just is the same word for
- 14:53:13scaling our numeric features so we're
- 14:53:15going to we're going to define Windows.
- 14:53:18Um, so that's why we're going to go
- 14:53:21ahead and import from pre-processing.
- 14:53:23We're going to import our scaler. Um,
- 14:53:25again, we could use minmax scaler here.
- 14:53:27We're going to use standard scaler. Um,
- 14:53:30but we could use minmax. Um, we have our
- 14:53:33one hot encoder here. Now, usually when
- 14:53:37we do oneh hot encoding, we use pd.get
- 14:53:41dummies. This does the same thing as
- 14:53:44that, but because we're going to be
- 14:53:46building a pipeline, we actually want
- 14:53:48the scikitlearn version of git dummies.
- 14:53:52So this is the scikitlearn version of
- 14:53:54git dummies here. And it and we have to
- 14:53:56use that version in the pipeline because
- 14:53:59everything in the pipeline needs to be
- 14:54:00an sklearn object. It needs to be an
- 14:54:03sklearn tool or object.
- 14:54:06So um instead of using pandis get
- 14:54:09dummies we're using one hot encoder
- 14:54:11which is does the same thing. Okay. In
- 14:54:15fact it just this basically just uses
- 14:54:18pd.get dummies um under the hood.
- 14:54:24Okay. So it just uses that uh anyways.
- 14:54:26It's just code that builds on builds on
- 14:54:28that.
- 14:54:30Now what's really nice here is we're
- 14:54:32also going to use from sklearn.impute
- 14:54:35impute. We're going to use a simple
- 14:54:36imper now what this is is an automated
- 14:54:40way to fill in missing values. So this
- 14:54:42is a fancy way of basically doing the
- 14:54:45the fill na on a data frame. So simple
- 14:54:48imputer um we are going to basically
- 14:54:52fill in the blanks. What we're going to
- 14:54:54do when we create this object is give it
- 14:54:56a strategy of how to fill in blanks.
- 14:54:58Should you use the average? Should you
- 14:55:00use the median? Should you use the max?
- 14:55:02Should you use the min? should use a
- 14:55:04default value. We're going to tell it
- 14:55:06what to do in this object.
- 14:55:09Okay. So, we're going to we're so we're
- 14:55:12going to use this as our automated tool
- 14:55:14for filling in missing values. So,
- 14:55:16that's really nice. It has so this is
- 14:55:18going to be a critical part of our
- 14:55:20pipeline an imputer that's going to fill
- 14:55:23in missing values.
- 14:55:26So, we have that.
- 14:55:31Yeah. Coding to reduce coding. Exactly.
- 14:55:34Uh we have our pipeline now. So we have
- 14:55:36our pipeline. So our pipeline is going
- 14:55:38to hold everything. So we need the
- 14:55:39pipeline object um to hold everything
- 14:55:42and that comes from sklearn.pipeline.
- 14:55:45Um so everything's going to actually go
- 14:55:47into a pipeline object. We're going to
- 14:55:49see how that looks. Um and finally we're
- 14:55:53going to from skarn.compose we're going
- 14:55:56to use a column transformer. The reason
- 14:55:58we're going to do this is because we are
- 14:56:01going to specify for some columns like
- 14:56:04the numerical features we should be
- 14:56:06scaling
- 14:56:07for some columns like the categorical
- 14:56:10features we should be one hot encoding.
- 14:56:13So the column transformer will allow us
- 14:56:15to map different transformations to
- 14:56:18different sections of columns which is
- 14:56:20really useful. So this is actually going
- 14:56:22to be a critical part of our pipeline to
- 14:56:25apply to make sure we only apply this to
- 14:56:27numerical features and only apply this
- 14:56:30to categorical features. Right? So this
- 14:56:33column transformer will help us um to to
- 14:56:38apply pre-processing to particular
- 14:56:40columns. Um like that ocean proximity is
- 14:56:44the only one that really needs this but
- 14:56:46every other column is going to need this
- 14:56:48all the numerical features.
- 14:56:50So, we're going to use this column
- 14:56:52transformer. And again, we're going to
- 14:56:53see how this looks, but just trying to
- 14:56:56give you an idea of why we're importing
- 14:56:57all these things.
- 14:57:03Okay. So, let's import those.
- 14:57:06Uh, this mentions about the column
- 14:57:08transformer. We just talked about it. It
- 14:57:10allows us to have a particular column or
- 14:57:13group of columns get the right
- 14:57:14transformation. So again, uh, looking
- 14:57:17ahead to our pipeline, the numerical
- 14:57:20features are the ones that are going to
- 14:57:21need scaling, but the categorical
- 14:57:24features are the ones that are going to
- 14:57:26need one hot encoding. However many
- 14:57:27categoricals there are. In this case,
- 14:57:29there's really only one, which is that
- 14:57:30ocean proximity. Go back to our data.
- 14:57:33Um, you can even see that in the info,
- 14:57:36there's just that one. Um, and we see
- 14:57:38that here, right? Just this one string
- 14:57:40column that should be one hot encoded.
- 14:57:42All these other guys should be scaled.
- 14:57:45Right? They should all be uh uh standard
- 14:57:47scaled.
- 14:57:49So this will allow us to specify those
- 14:57:52distinctions.
- 14:57:56All right. So let's get started building
- 14:58:01our pipeline. So this is going to be
- 14:58:02really cool. We're going to build out
- 14:58:03the pipeline. Um let's extract our
- 14:58:08numerical data and our categorical data.
- 14:58:10Now this is a really neat way of doing
- 14:58:12that that I'm not sure we've seen
- 14:58:13before.
- 14:58:14Um so what this does is we'll take our
- 14:58:18data frame particular our training data
- 14:58:21frame and select our data
- 14:58:26that's what this select dtypes does is
- 14:58:28select data from it um which includes
- 14:58:31only the object type columns so only the
- 14:58:35object types. Now what's that?
- 14:58:38The object type is the string right? So
- 14:58:41this should select only this column
- 14:58:44because it's in the include.
- 14:58:47We go here include only object types in
- 14:58:50the result. And so this should only have
- 14:58:53our one categorical column which is
- 14:58:56ocean proximity. So, housing cat is
- 14:58:59going to have a reference to our uh it's
- 14:59:03going to be a list that has a a
- 14:59:05basically just our ocean proximity
- 14:59:08feature because this select dtypes will
- 14:59:11make sure we only pick object types and
- 14:59:15um
- 14:59:16grab those columns. So this is a way to
- 14:59:20neatly grab um our categorical features
- 14:59:24here by including the object types. Now
- 14:59:28on the flip side we can exclude object
- 14:59:30types and get everything else. So this
- 14:59:32is going to be all other columns which
- 14:59:35is excluding the object. So this is
- 14:59:38excluding this meaning we should get all
- 14:59:41of our numerical features that way. So
- 14:59:44this will be all of our numericals
- 14:59:47by excluding the object type and this
- 14:59:51will be our housing num which is short
- 14:59:54for numerical. So this excludes
- 14:59:58the uh object type meaning all numerical
- 15:00:06features
- 15:00:08right all numerical features there.
- 15:00:12Okay.
- 15:00:15So, if we were to uh let's double check
- 15:00:18this. Let's sanity check this. If we
- 15:00:19were to print out the housing
- 15:00:23cat, um this should be just the ocean
- 15:00:27proximity feature, which it is. So, just
- 15:00:30that one. If we were to print out the
- 15:00:32housing num, this should be all the
- 15:00:34numerical features, which are all these
- 15:00:37guys. So it's just a reference to those
- 15:00:39columns so that we can uh use those
- 15:00:43later when we're mapping uh this
- 15:00:45transform needs to go to this column
- 15:00:47like the one hot encoding needs to go to
- 15:00:49this column and the scaling needs to go
- 15:00:52to these columns right so we have those
- 15:00:56uh names of those columns already at our
- 15:00:58disposal. So, we're just doing that.
- 15:01:04And this is just a
- 15:01:07simple check.
- 15:01:11Uh, are you guys able to run this?
- 15:01:16If you're following along, let me pause
- 15:01:18there. Make sure I'm not going too fast.
- 15:01:26Uh it so the the issue with a specific
- 15:01:29data type like that is none of these are
- 15:01:31ants. They're actually all floats. So we
- 15:01:34did float. I think that should work. But
- 15:01:36yes, that's the idea.
- 15:01:41Great. I'm glad to hear that right there
- 15:01:43with me. Great. Glad to hear that.
- 15:01:54Okay. So, we have our columns picked out
- 15:01:56here, which we're going to use later.
- 15:01:59Okay.
- 15:02:02All right. So, let's go ahead and build
- 15:02:06out our steps for each of these types.
- 15:02:10So, um for our numerical features, let's
- 15:02:15build out our pipeline steps. So what
- 15:02:17we're going to do is build out a
- 15:02:18numerical pipeline. And it's going to be
- 15:02:21a pipeline with a list
- 15:02:25of tupils. And the reason these are
- 15:02:28tupils is because every tupil has a
- 15:02:30name. So here this is a name that we can
- 15:02:33it can be whatever we want it to be. So
- 15:02:35we're calling it imputer. We could call
- 15:02:38it anything we want. We could call it
- 15:02:39fill in the blanks. We could call it
- 15:02:41null filling. Call it whatever you want.
- 15:02:45We're calling it imputer because that's
- 15:02:46that's a pretty um easy name for it. An
- 15:02:50accurate name to what it's doing. Um but
- 15:02:53the important thing is after the name
- 15:02:55you give it, you put in the scikitlearn
- 15:02:59object that you are going to use to
- 15:03:01operate on your data. So in this case,
- 15:03:04we're using a simple impery
- 15:03:08of median. Now that's a choice. We could
- 15:03:11use a strategy of mean, max. Um, we
- 15:03:16could provide it a constant default
- 15:03:18value. But what this means is we are
- 15:03:22going to fill any blanks we find in
- 15:03:24those columns with the median value of
- 15:03:27that column. That's the strategy for the
- 15:03:29computer. So that's pretty cool. This is
- 15:03:31kind of an automated way to fill in the
- 15:03:32blanks using for any column using its
- 15:03:37median,
- 15:03:39right? And so we could change that. We
- 15:03:40could put mean here or max or min or
- 15:03:43whatever. Um
- 15:03:46but we are filling in the blank on any
- 15:03:48column with its median. And the reason
- 15:03:51this works is because we are going to
- 15:03:53apply this pipeline only to these
- 15:03:55numerical features. So that is fine.
- 15:03:59We're we're not going to apply it to the
- 15:04:01categorical features. We're going to
- 15:04:02apply it to only those numerical. So it
- 15:04:05should have a median value, right? So
- 15:04:08that that's totally fine. So we're going
- 15:04:11to now look at how we're constructing
- 15:04:13the steps. We have a list of tupils.
- 15:04:16Here's one tupole
- 15:04:19which is the imputer with a simple imper
- 15:04:22of strategy median. And then we can have
- 15:04:25as many tupils as we want which
- 15:04:27represent processing steps. So every let
- 15:04:30me write that down. Every tupil
- 15:04:34represents
- 15:04:36a pre-processing
- 15:04:38step on our data.
- 15:04:42Okay, so we have an imputer step named
- 15:04:46imputer and the reason it has a name is
- 15:04:49just so you can reference it in the
- 15:04:51pipeline if you need to. So you so it
- 15:04:53has like a a reference name um that you
- 15:04:56give it. Um but this is the more
- 15:04:59important part is the actual scikitlearn
- 15:05:01object that's doing the processing. So
- 15:05:04in this case a simple computer but
- 15:05:06notice that we have a secondary step
- 15:05:08which is our scaling. Now this makes
- 15:05:09sense. This is something we should be
- 15:05:11doing to our features is we should be
- 15:05:14scaling them. So here we we say okay
- 15:05:17let's fill in any blanks first.
- 15:05:20By the way order
- 15:05:23matters.
- 15:05:26So, and what I mean by that is the
- 15:05:30simple imputer
- 15:05:33is before the scaler. Now, that's
- 15:05:37important because what that means is we
- 15:05:40should be filling in any blanks before
- 15:05:42we attempt scaling.
- 15:05:45So, that order actually matters. We're
- 15:05:47going to fill in blanks first in this
- 15:05:50list. That's first. We're going to fill
- 15:05:52in blanks. Then we are going to scale
- 15:05:58right then we scale which makes sense
- 15:06:01right so we we fill in blanks first then
- 15:06:03we apply the scaler to scale our
- 15:06:05features so those are our two steps
- 15:06:10so so pretty simple um we are building
- 15:06:14out our two steps now this is just one
- 15:06:17piece of the puzzle we are going to put
- 15:06:18this pipeline together with our one hot
- 15:06:21encoding that's going to be coming up
- 15:06:23next and build out our final pipeline.
- 15:06:27But this is um a a pipeline that has two
- 15:06:30steps that will actually be used with a
- 15:06:32larger pipeline coming up where we we do
- 15:06:35one hot encoding to our categoricals and
- 15:06:37then we put a model in there at the end
- 15:06:40to train and and use for prediction. So
- 15:06:44um pipelines can actually be composed is
- 15:06:48is uh something to realize there is that
- 15:06:50we can have a pipeline that contains a
- 15:06:53few steps. We can have another pipeline
- 15:06:54over here that contains a few steps and
- 15:06:56we can actually um kind of put them
- 15:06:58together into a final pipeline that has
- 15:07:00both pipelines uh kind of merged
- 15:07:02together. Okay. So we're going to see
- 15:07:05that coming up when we construct our
- 15:07:07final one. Our final one, as you can
- 15:07:09imagine, needs to handle this mapping of
- 15:07:12basically saying, let's do one hot
- 15:07:13encoding to these guys and then do this
- 15:07:17pipeline here to these numerical
- 15:07:20features. That's what our final pipeline
- 15:07:23needs to handle. And it will. We're
- 15:07:25going to build that out.
- 15:07:28But let me pause here. Um, were you guys
- 15:07:32able to run this? Are you with me on
- 15:07:35this this pipeline here?
- 15:07:38Does that make sense? Those two steps
- 15:07:40one is filling in blanks with a median
- 15:07:44whatever column. So where so this is
- 15:07:47this is what's so amazing about this is
- 15:07:50this is going to automatically search
- 15:07:52for nulls and if you come across a
- 15:07:56column with a null, it's going to use
- 15:07:58the median of that column
- 15:08:02to fill in the blank, right? To fill in
- 15:08:04those nles.
- 15:08:18Okay,
- 15:08:21great. Glad to hear. Glad to hear.
- 15:08:25Okay.
- 15:08:27All right. So we are going to now um put
- 15:08:32this together with a column transformer
- 15:08:37to basically say what steps are going to
- 15:08:40be mapped to what columns.
- 15:08:44Um so now you can see what we're doing
- 15:08:47here is using the column transformer
- 15:08:49which is going to be a list of tupils
- 15:08:51again. So this is another um list of
- 15:08:55tupils.
- 15:08:57But the important thing is um
- 15:09:01each tupil
- 15:09:04has a name
- 15:09:06followed by so it has a name uh which
- 15:09:10again is is generic. You can say
- 15:09:12whatever you want it to be. So here
- 15:09:14we're kind of shortening this to
- 15:09:15numerical. This is short for
- 15:09:16categorical. But the important thing is
- 15:09:18it's followed by a pipeline
- 15:09:23slashstep
- 15:09:26followed by a pipeline slashstep
- 15:09:29um followed by a uh followed by a list
- 15:09:34of columns that it applies to. So you
- 15:09:39can see that pattern here. What we're
- 15:09:41saying is we're going to apply that
- 15:09:44numerical pipeline we just defined. So
- 15:09:46this is saved in a numerical pipeline
- 15:09:48object here. We're going to apply that
- 15:09:51to those numerical features. So this is
- 15:09:54that list
- 15:09:56of numerical features here. So that's
- 15:09:59how we do the mapping. We have a tupil
- 15:10:01here that says okay apply these steps to
- 15:10:04these columns.
- 15:10:06Those go together in that tupil, right?
- 15:10:09Apply these steps to this uh these
- 15:10:12columns. And then apply this step. Now
- 15:10:15what is the step? This is a one hot
- 15:10:17encoder
- 15:10:19which is going to uh uh encode um those
- 15:10:24features and it's going to uh ignore um
- 15:10:28basically nulls for now. That's a choice
- 15:10:31but it's going to ignore um uh basically
- 15:10:36ignore nles and and uh skip over them
- 15:10:39for now. We now we know there's no NLES
- 15:10:42because we already did an is NA from
- 15:10:45before and we know there's not any NLES
- 15:10:48in that ocean proximity. So this isn't
- 15:10:49going to be an issue. But that's what
- 15:10:52that would do.
- 15:10:54But we have a one hot encoder here which
- 15:10:57we're going to apply to our categorical
- 15:11:00features. Now of course that's just the
- 15:11:03ocean proximity feature but that but
- 15:11:05again you see the pattern in the tupole
- 15:11:07is apply this transform which is the one
- 15:11:10hot encoding to this column apply these
- 15:11:13numerical transforms which is a whole
- 15:11:15pipeline. So it's two steps in a
- 15:11:18pipeline of um
- 15:11:22uh an imputer and a scaler are going to
- 15:11:25be applied to this
- 15:11:28really nice. So those are going to be
- 15:11:29all together in this column transformer
- 15:11:32and that is our way to signal that for
- 15:11:33these numerical features use these
- 15:11:36steps. For our categorical features use
- 15:11:38this step and and you know if we had
- 15:11:41more than one step we were applying to
- 15:11:42categorical we could build a pipeline
- 15:11:45for the categorical and it would and do
- 15:11:47the same thing. We have more than one
- 15:11:49step here and so it's good practice when
- 15:11:52you have more than one step to just put
- 15:11:53that in a pipeline because we have more
- 15:11:55than one step. We'll just put that in
- 15:11:57this list inside of the pipeline and we
- 15:12:00can map that pipeline to those features.
- 15:12:03Here we only have one step. So it's okay
- 15:12:05to just put that there um and apply that
- 15:12:09to the categorical features. But if we
- 15:12:11had more than one step um it would be
- 15:12:14good practice to put that in a pipeline
- 15:12:17which is what we do here. Right? This
- 15:12:18pipeline is being mapped to these
- 15:12:20features. This step is being applied to
- 15:12:23this feature.
- 15:12:28Okay,
- 15:12:30how about that? Are you guys able to run
- 15:12:33that one? Does that make sense what we
- 15:12:35have set up so far? So, we're almost
- 15:12:37there. We almost have our final
- 15:12:38pipeline. We have our pre-processing
- 15:12:40basically done to say our numerical
- 15:12:43features should be processed with that
- 15:12:44other pipeline and our categorical
- 15:12:47features should be one hot encoded.
- 15:12:49We're getting close. The only thing
- 15:12:50we're really missing here is a model.
- 15:12:54The only thing we're really missing is
- 15:12:56to have our final model training
- 15:12:59pipeline is to actually include a model
- 15:13:01which should come at the end.
- 15:13:04Right? So it should we should be doing
- 15:13:06these steps first
- 15:13:09then doing modeling which we know right
- 15:13:12we we've done that uh many times. We've
- 15:13:14done our pre-processing and then we do
- 15:13:15our modeling.
- 15:13:19Any questions on that?
- 15:13:37Okay.
- 15:13:38Fantastic.
- 15:13:42All right.
- 15:13:45So, if we wanted to uh see if we wanted
- 15:13:48to test this so far, um we could. So we
- 15:13:51could run the pre-processing and
- 15:13:53actually run a fit transform on our data
- 15:13:56and this will um basically apply that
- 15:13:59pipeline to the data. Now this would be
- 15:14:01a sanity check. This is a good this is a
- 15:14:04good kind of um this is a good sanity
- 15:14:07check that our pre-processing
- 15:14:12works. So it's doing what we expected to
- 15:14:15do. It's not our final pipeline because
- 15:14:17we don't have our model in there yet.
- 15:14:19But this is just to ensure that all of
- 15:14:21the features are kind of behaving as we
- 15:14:23expect. So we can uh we can do that and
- 15:14:27we can take a look at the um results.
- 15:14:30This looks pretty good. This all of our
- 15:14:32numerical features ended up scaled
- 15:14:35which is pretty good. And we have one
- 15:14:37hot encoded features for that ocean
- 15:14:39proximity over here.
- 15:14:42Okay. So this looks pretty this looks
- 15:14:44reasonable of those steps being applied
- 15:14:46to the right columns. But this is a good
- 15:14:49kind of sanity check to just run our fit
- 15:14:51transform on our data to ensure those
- 15:14:55steps are actually happening and they
- 15:14:57are. You can see here the result of the
- 15:15:00scaling and the uh the one hot encoding.
- 15:15:04So that that all looks pretty
- 15:15:05reasonable,
- 15:15:09right?
- 15:15:11And uh what we should also do is make
- 15:15:15sure there are no nulls in this which
- 15:15:16there shouldn't be because we did the
- 15:15:18imper. So we should be doing uh is na
- 15:15:22dot
- 15:15:25sum
- 15:15:29and there is no nulls anymore. So that
- 15:15:31looks pretty good right? Those got
- 15:15:33filled in uh by doing our steps. our
- 15:15:37pipeline steps executed really nicely on
- 15:15:39our training data um and and we were off
- 15:15:44and running. And there's nothing unique
- 15:15:45about the training data. We could do
- 15:15:47this to our test data as well
- 15:15:50and verify that those steps are running
- 15:15:52and they would, right? There's nothing
- 15:15:54really that special about running it on
- 15:15:55the training data. Um it should also
- 15:15:59work on the test features as well and it
- 15:16:00does. You can check that for yourself.
- 15:16:05Okay.
- 15:16:10All right. So, that's pretty cool. We
- 15:16:12can uh verify all that's working.
- 15:16:21Any questions on that?
- 15:16:24We're almost there with our full
- 15:16:25pipeline. This this is this is not the
- 15:16:28full pipeline, but this is something
- 15:16:30that will run during our full pipeline.
- 15:16:32Of course, our features are going to be
- 15:16:34transformed according to those steps and
- 15:16:36then it will be uh put into our model to
- 15:16:38either predict or train with. Um
- 15:16:42so let's do that. Let's actually build
- 15:16:44out our final uh model here. So it's
- 15:16:48actually going to be really easy to do.
- 15:16:50All we need to do is um put in our
- 15:16:53model. So here we're going to import the
- 15:16:56ridge model here. Now, we could use any
- 15:16:59we could use linear regression, we could
- 15:17:00use lasso, we could use elastic net. Um,
- 15:17:03we're just going to use ridge um uh um
- 15:17:07just to test it out. And um we are going
- 15:17:11to uh now put in a final pipeline. So,
- 15:17:15we're going to use our pipeline. And so,
- 15:17:18we're going to create a new one here and
- 15:17:21map our pre-processing to our
- 15:17:23pre-processing that we've already built.
- 15:17:25So this is a column transformer that
- 15:17:27already has all of our steps. And then
- 15:17:30notice what comes after it is just the
- 15:17:32model. Now that's pretty pretty basic,
- 15:17:34but it makes sense that it should come
- 15:17:36after that model. Um and of course this
- 15:17:39is a generic name. We could we can name
- 15:17:41it whatever we want to. Um model ridge
- 15:17:44is pretty reasonable um to because it is
- 15:17:48a ridge uh regression. But uh of course
- 15:17:51we could we could change that.
- 15:17:55Okay. So that builds out our uh final um
- 15:17:58pipeline. So now we have a pipeline. And
- 15:18:02what's great about that is this signals
- 15:18:05that all of these steps should be
- 15:18:07completed prior to doing anything with
- 15:18:09this model. So all of those processing
- 15:18:12steps are going to run and then we're
- 15:18:14going to do ffit or predict and that. So
- 15:18:17that's really great. It ensures that all
- 15:18:19those steps are running together every
- 15:18:21single time we call.predict with this
- 15:18:24with this model. So we're just going to
- 15:18:26use the pipeline in place of the model
- 15:18:30to ensure that all those steps are
- 15:18:32running together. And this is our this
- 15:18:35is kind of our final pipeline that we
- 15:18:37would use uh with like something like
- 15:18:39ffit or predict.
- 15:18:42So let me make that a note of that. Now
- 15:18:45we can use this final pipeline just like
- 15:18:50a regular model i.e. pipeline.fit
- 15:18:56or pipeline.predict.
- 15:19:00So we could use it in ei in either
- 15:19:01fashion uh to to train the pipeline
- 15:19:06would be this guy and then use the
- 15:19:08pipeline to predict would be this. And
- 15:19:10what we should realize is under the
- 15:19:11hood, these steps are running first and
- 15:19:13then we train it or these steps run
- 15:19:16first then we use it for prediction.
- 15:19:24Okay,
- 15:19:26questions on that. Does that make sense
- 15:19:29on this final pipeline here? It's just
- 15:19:32now it it's really cool because we have
- 15:19:34a pipeline
- 15:19:36made up of a of a pipeline really,
- 15:19:39right? a pipeline made up of a pipeline.
- 15:19:41But that's scikitlearn allows you to do
- 15:19:42that to compose pipelines in this way.
- 15:19:46That's is pretty uh pretty uh normal
- 15:19:49there.
- 15:19:59Okay.
- 15:20:01What I want to show you is we can
- 15:20:03actually use this pipeline in a grid
- 15:20:06search. So that's pretty amazing. We can
- 15:20:08use this pipeline in any way we can use
- 15:20:10a mo like a regular model. It's just
- 15:20:13that now our pre-processing steps have
- 15:20:16kind of been packaged together with our
- 15:20:18model to ensure that they always run
- 15:20:21anytime we do any processing with this
- 15:20:22model. Um so for instance we can do a
- 15:20:27grid search just like we did with a
- 15:20:29regular with with just a model right
- 15:20:31with just this. Um we can do the same
- 15:20:34thing with the whole pipeline. Um, so
- 15:20:37the only catch is that you want to make
- 15:20:40sure in your grid you name things in the
- 15:20:44appropriate way inside of your your uh
- 15:20:46keys in your dictionary. So uh for
- 15:20:49instance um inside of the grid uh we're
- 15:20:54going to set up the alpha that would be
- 15:20:56used with this ridge regression by
- 15:20:59referencing its name. So this is model
- 15:21:01ridge is this is the name of the model
- 15:21:04inside of the pipeline. So you want to
- 15:21:06make sure that goes first.
- 15:21:08And then what scikitlearn does is it
- 15:21:11recognizes parameters that belong with
- 15:21:13this model by using a double underscore.
- 15:21:17So the so you have underscore alpha um
- 15:21:22here. So the double
- 15:21:25uh underscore
- 15:21:27signals a parameter
- 15:21:31belonging to model ridge in the in the
- 15:21:37pipeline.
- 15:21:39Okay, so we have a model ridge is just a
- 15:21:42reference to the model in our pipeline.
- 15:21:44That's the one we're going to test out
- 15:21:46these parameters with. and
- 15:21:48underscore_pha is just a way to say this
- 15:21:51alpha belongs to this model. Okay, it
- 15:21:56belongs so it's going to be used with
- 15:21:58that model in our pipeline. Um otherwise
- 15:22:02it's going to work exactly the same way.
- 15:22:04It's just we need to line up this naming
- 15:22:06convention of of scikitlearn.
- 15:22:08You just have to reference this to
- 15:22:11whatever name you provided here and then
- 15:22:14underscore parameter. So L1 ratio alpha
- 15:22:18whatever right would go there.
- 15:22:21Okay. So there is a range from 0.1 to2
- 15:22:26uh step size of 0.1
- 15:22:29um and then we do our grid search CV. So
- 15:22:32this is exactly the same setup as we had
- 15:22:34before. It's just that our model is now
- 15:22:38the pipeline. So our pipeline is going
- 15:22:40in there. Um we have our grid going in
- 15:22:43there. we have our scoring is the same,
- 15:22:46you know, negative absolute error. Um,
- 15:22:48we're using five-fold cross validation
- 15:22:51and we're parallelizing that search. Um,
- 15:22:54so we're going to search through these
- 15:22:55alphas and uh basically fit this to our
- 15:23:00um data and find the best um find the
- 15:23:05best alpha.
- 15:23:07So, it's going to try out all those
- 15:23:08combinations and try to come up with the
- 15:23:11best alpha.
- 15:23:14So looks like the best alpha was 0.1 for
- 15:23:17the ridge.
- 15:23:19Okay. Is the best. So then um if we
- 15:23:24wanted to we could uh then predict using
- 15:23:28the model um which would be doing
- 15:23:31something like this. Um, and we could
- 15:23:34also go back and do something like so we
- 15:23:38could
- 15:23:42now use um this param. So we could do
- 15:23:47model
- 15:23:50um equals ridge
- 15:23:55and then we could put in our alpha.
- 15:23:58Um, alpha is our results, our best
- 15:24:01parameters, and then we get that model
- 15:24:03ridge alpha. And then we just rebuild
- 15:24:05our our pipeline
- 15:24:10equals um pipeline and then we uh put in
- 15:24:14this new model here. So we could do
- 15:24:16this. This would be going back and just
- 15:24:20um putting in our best alpha here for
- 15:24:24this model and then uh ensuring that's
- 15:24:27part of our our pipeline. So we're just
- 15:24:28overwriting that pipeline with the best
- 15:24:30model there
- 15:24:36to get the best model in our pipeline.
- 15:24:42Okay,
- 15:24:44so that's all this is doing is just
- 15:24:46initializing a new um let me actually I
- 15:24:49can put this code in here.
- 15:24:56This is actually just getting this is
- 15:24:58just getting a model with the best alpha
- 15:25:00and then reinserting that into our our
- 15:25:03uh we're just overwriting our final
- 15:25:05pipeline there with the best model that
- 15:25:07we have.
- 15:25:10So pretty cool that pipeline can be used
- 15:25:13basically exactly like a model, right?
- 15:25:15It's it's going right here in the grid
- 15:25:17search and being used uh entirely like a
- 15:25:20basic model. So we do ffit
- 15:25:23um and that allows us to use it. We
- 15:25:26could do predict we could even do
- 15:25:28pipeline.predict once we we could go
- 15:25:30back and do final pipeline.fit
- 15:25:33um with this and then final
- 15:25:34pipeline.predict with this and evaluate
- 15:25:41Okay,
- 15:25:43so pretty cool that pipeline can be used
- 15:25:45uh basically exactly like how a model
- 15:25:47would be any way we'd use a model.f
- 15:25:50model.pred predict we can use a
- 15:25:51pipeline.
- 15:25:54So grid search is for instance something
- 15:25:56that can use a model in there. Um but
- 15:25:59instead of just a model we're ensuring
- 15:26:00we have our pre-processing steps kind of
- 15:26:03bundled with that model in this
- 15:26:04pipeline.
- 15:26:06Any
- 15:26:10questions on
- 15:26:12uh this example so far?
- 15:26:18Were you guys able to run it up to here?
- 15:26:20Were you able to run the grid search?
- 15:26:36Okay, great.
- 15:26:45Okay.
- 15:26:47Okay. So, this is this is uh just
- 15:26:49showing you what's actually happening
- 15:26:51underneath the hood is uh you know,
- 15:26:54we're doing some scaling. We're doing
- 15:26:56some one hot encoding. Um
- 15:27:00and we're doing some uh we're doing a
- 15:27:03model here. And that's all part of our
- 15:27:06pipeline. Um, and then we can use the
- 15:27:10pipeline however we want. So for
- 15:27:11example, I know it's not here, but for
- 15:27:14an example, we could use um once we do
- 15:27:17once we have this final pipeline um we
- 15:27:20can can use the final um
- 15:27:25pipeline to predict. So we can do um
- 15:27:29predictions
- 15:27:31equals final
- 15:27:34pipeline.predict
- 15:27:36and then we can pass in our test data.
- 15:27:38Now what happens on this is once we have
- 15:27:41ran our our pipeline.fit we have a
- 15:27:44trained pipeline and then when we run
- 15:27:47this final pipeline.predict uh this data
- 15:27:50is going to be transformed.
- 15:27:52It's going to go through those
- 15:27:53transformation steps and then we would
- 15:27:55apply our model to it at the end uh to
- 15:27:59to make those predictions and then we
- 15:28:00can evaluate those predictions which is
- 15:28:02what we're doing kind of here
- 15:28:05right.
- 15:28:12Okay.
- 15:28:16All right. So in conclusion uh we have
- 15:28:19gone through a lot of stuff here. Um,
- 15:28:22we've gone through regression, we've
- 15:28:24done the regularization on regression.
- 15:28:27So hopefully we have a good foundation
- 15:28:29on regression. Um, what we're going to
- 15:28:31do in a little bit is actually do some
- 15:28:33additional practice with regression on a
- 15:28:35new problem. We're going to do a
- 15:28:36capstone problem and do some additional
- 15:28:39regression work with that. Um, so we'll
- 15:28:43do that next. Um but the other thing we
- 15:28:47learned is how to evaluate the
- 15:28:48regression using things like mean
- 15:28:49squared error, RMSSE which is square
- 15:28:52root of that. Um which is which is
- 15:28:55really cool. So we have a sense of that
- 15:28:58error which is our distance from our
- 15:29:00prediction to the actual value. That's
- 15:29:02always what these uh that's always what
- 15:29:05these things are doing like this, right?
- 15:29:08This mean absolute error metric from
- 15:29:10scikitlearn is computing the average
- 15:29:13distance from these predictions to these
- 15:29:15test labels that we have right those
- 15:29:18actual values. Um and that gives us a
- 15:29:20sense on average how far away are our
- 15:29:23predictions
- 15:29:24um to see how good of a model that we
- 15:29:27have, right? And we should be evaluating
- 15:29:30that error generally
- 15:29:32um against the scale of our targets to
- 15:29:36see, you know,
- 15:29:38uh how far off we typically are.
- 15:29:43Okay. Any questions at all on this
- 15:29:45lesson on regression? Uh anything we
- 15:29:47covered up to this point?
- 15:29:49We're going to do some more practice
- 15:29:50with the next. We'll do the capstone. So
- 15:29:54we get So we just do some more
- 15:29:56regression problems.
- 15:30:07Yeah, it's a that's another bad score.
- 15:30:09It's a little bit hard to interpret this
- 15:30:10though because it's m ae. Um so one
- 15:30:14thing we could do is is compute mean
- 15:30:17squared error and then take the square
- 15:30:19root of it to get the RMSSE which is a
- 15:30:22much better uh evaluation metric in
- 15:30:25terms of our target. Um so we could
- 15:30:28actually run that. Uh if we go back here
- 15:30:31and um we could generate for instance we
- 15:30:34could generate the MSE which is the mean
- 15:30:37squared
- 15:30:39error
- 15:30:41and it's it's the same exact function uh
- 15:30:44of using our predictions
- 15:30:49um
- 15:30:50and then we could just print that out
- 15:30:52mean squared error.
- 15:30:56So we have mean squared error and then
- 15:30:58what we can do is let's take the um MP.
- 15:31:03Square root of that.
- 15:31:06So that way we can generate the RMSSE.
- 15:31:08So yeah that I mean that's pretty bad.
- 15:31:10That's uh pretty bad. Uh now let's let's
- 15:31:15go back and look at our
- 15:31:18uh data though. So let's take a look at
- 15:31:20the average for our Y. Um remember one
- 15:31:23thing we should be doing is taking a
- 15:31:25look at um what our uh let's take a look
- 15:31:28at y test mean
- 15:31:31to get an average value. So the average
- 15:31:34value is in the 200,000s. So
- 15:31:38this isn't this isn't awful. This is
- 15:31:4170,000. It's still a decent amount of
- 15:31:43error. It's not as bad as the models we
- 15:31:45have before though, right? This is an
- 15:31:47average
- 15:31:49median price of the house is in the
- 15:31:5226,000 range and our error is off by
- 15:31:56like 70,000,
- 15:31:59right?
- 15:32:01So, it's not good. Um, but it's not
- 15:32:07hor like as bad as the it's not as
- 15:32:10horrible as we've seen so far. Right.
- 15:32:12This is a little bit better of a model.
- 15:32:14A little bit better. closer to zero
- 15:32:16would be better, right? Um but the
- 15:32:19smaller the better. But uh remember this
- 15:32:22is the um these even the mean absolute
- 15:32:26error is is technically in similar units
- 15:32:29as the as the uh um
- 15:32:33as the target. So 50,000 60,000 here
- 15:32:3770,000 it's still a decent amount of
- 15:32:39error in terms of 200,000.
- 15:32:44Uh so far we only come up with models
- 15:32:45and test their accuracy with available
- 15:32:47data. We haven't used a model to make
- 15:32:48completely new predictions on No, we
- 15:32:50haven't done that. Uh except we know how
- 15:32:53to do that. Um it would so to make
- 15:32:56predictions on new data would be exactly
- 15:32:58how we're making them on our available
- 15:33:00data because we actually do that all the
- 15:33:03time. If we go back down to our model
- 15:33:05building,
- 15:33:07um it's it looks just like this, right?
- 15:33:09where we take so for instance we do
- 15:33:13predictions all the time on test data
- 15:33:15that was never involved in the training.
- 15:33:18So it's it's as if this data mimics new
- 15:33:22data that we've never seen before. So if
- 15:33:25we had new raw data it would just it
- 15:33:28would be the same exact process. the new
- 15:33:31now with our pipeline it makes it a
- 15:33:33little bit easier because with the
- 15:33:35pipeline
- 15:33:36um the raw data will go through those
- 15:33:39transformations which it should right
- 15:33:41the raw data should because if it's
- 15:33:42missing data it needs to be filled in if
- 15:33:44it has categoricals it needs to be one
- 15:33:46hot encoded so that's the purpose of the
- 15:33:49pipeline actually is to make sure that
- 15:33:53if we're dealing with raw data um those
- 15:33:56steps can happen on the data before it
- 15:34:00goes into the model. Right? So we so
- 15:34:05that's kind of the purpose of the
- 15:34:07pipeline
- 15:34:08is to ensure that we run those steps
- 15:34:11ahead of using it using a model with it.
- 15:34:15But but ultimately that's how it uh any
- 15:34:18scikitlearn model is going to be doing
- 15:34:20the predict even if it's a pipeline
- 15:34:22right it's going to be uh we just go
- 15:34:25back down here it's going to be um
- 15:34:27predict it's always going to be that on
- 15:34:30new data
- 15:34:32yeah
- 15:34:36okay
- 15:34:39really good question uh what I wanted to
- 15:34:41do next was do some practice um I wanted
- 15:34:45to go over to the capstone session five.
- 15:34:50So, in this course, we have some more
- 15:34:52capstone sessions. So, if you have a
- 15:34:54moment, you want to pull up those
- 15:34:56capstone session materials, the
- 15:34:58incremental capstone session materials.
- 15:35:00Um, we're going to be doing session five
- 15:35:03today. So, this is just um remember it's
- 15:35:06just extra practice that we do after
- 15:35:08we've covered some concepts. So we are
- 15:35:11going to do some regression practice now
- 15:35:14that we've uh covered regression um
- 15:35:16pretty fully and then um this will this
- 15:35:19will be good practice before we head
- 15:35:21into lesson four on classification. So I
- 15:35:24just want to do this practice now while
- 15:35:26it's fresh while the material is kind of
- 15:35:28fresh in our in our minds. Um do you
- 15:35:31guys have the capstone materials? Do you
- 15:35:34know where to get it in your LMS? It's
- 15:35:36in your LMS and the resources the uh
- 15:35:39capstone materials you want to download
- 15:35:41that so you can get the the data um and
- 15:35:45the the slides for the instructions
- 15:35:48right the or PDF I think for you guys
- 15:35:52but uh let me ask you do you have those
- 15:35:55we want to pull up session five if you
- 15:35:57have it
- 15:36:01thank you I was just going to share that
- 15:36:03appreciate that yeah so this is going to
- 15:36:05be session Question five. Um, you're
- 15:36:07also going to want the data that we're
- 15:36:11going to use with this, which is going
- 15:36:12to be the, uh, bike rental data set.
- 15:36:15I'll share that with you guys now.
- 15:36:18So, we're going to be using this bike
- 15:36:19rentals data set for this uh, for this
- 15:36:22capstone. Um, it's the one we're going
- 15:36:24to build a regression model off of.
- 15:36:29Okay. So, you should have that one from
- 15:36:31the the capstones data sets as well.
- 15:36:35All right. So, let's go through this.
- 15:36:37Um, we are going to be doing uh machine
- 15:36:41learning here. So, we're going to be
- 15:36:42doing uh so we're talking about that
- 15:36:45kind of example product that Aura
- 15:36:48product that was in our original
- 15:36:49capstone. Um, and in order to do uh to
- 15:36:54to to um make decisions, it's going to
- 15:36:57have to build some models. Um, in this
- 15:37:00case, it's going to be doing some bike
- 15:37:02rental modeling. um which is the data
- 15:37:04set we have. So we're moving away from
- 15:37:06that healthcare data set going into this
- 15:37:08bike rental data set as an example of
- 15:37:10the capabilities here. Um so we're going
- 15:37:14to do this first capstone. Um after we
- 15:37:18do classification, we'll do this
- 15:37:20practice. Um after we do unsupervised
- 15:37:23learning, we'll do this practice on
- 15:37:24clustering. And then after we do
- 15:37:26recommendation uh which is the last
- 15:37:29lesson um we'll come back and do
- 15:37:31practice with building a recommendation
- 15:37:32engine. Okay, but we're going to do this
- 15:37:34one today. Um and then we will uh do
- 15:37:39these other capstones as we go along. So
- 15:37:41this will be session six, session seven,
- 15:37:44and session 8
- 15:37:46uh later on in the course. Okay.
- 15:37:51Um,
- 15:37:56okay. A little bit about the data. So,
- 15:37:58uh, in this capstone, we're going to be
- 15:38:00working with a, um, a shop, like a
- 15:38:03retail shop that rents out bikes. And
- 15:38:06they have data, um, on a per day basis
- 15:38:10with, um, actually on a per hour basis
- 15:38:13on the number of bikes that they rented
- 15:38:15in every hour. Um, so maybe one hour
- 15:38:18they rented out 20 bikes, another hour
- 15:38:20they rented out 30 bikes. Um, another
- 15:38:23hour they rented out 15. So they have
- 15:38:26that data here in in the CSV. Um, they
- 15:38:29have other kinds of data like the
- 15:38:31environmental data like what the
- 15:38:33temperature was at that hour, the
- 15:38:35humidity, if there was snowfall if it's
- 15:38:37a holiday, the wind, the visibility, the
- 15:38:41due point, um, solar radiation,
- 15:38:43rainfall, uh, what season it was. um
- 15:38:47what uh what day it was like
- 15:38:51um the functional is like if it's uh I
- 15:38:54believe it's like if it's a weekend or
- 15:38:55weekday um which would be uh
- 15:38:58nonfunctional
- 15:39:00um so
- 15:39:03based on those features we have a bunch
- 15:39:05of tasks okay so we have um based on the
- 15:39:11uh rented by count hour of the day um
- 15:39:16temperature, humidity, wind speed,
- 15:39:17rainfall, and whatever other features
- 15:39:19we're that are in the data set. We're
- 15:39:21actually going to build a model to
- 15:39:22predict the bike count required for
- 15:39:24every hour to have a stable supply of
- 15:39:27rented bikes. So, our goal is actually
- 15:39:28going to be to predict the bike rental
- 15:39:31count per hour. Um and uh we are going
- 15:39:36to do that using all the features we
- 15:39:38have at our disposal. um like mostly
- 15:39:41those environmental features and what
- 15:39:43day it is, those kind of things. Um
- 15:39:47so we're going to load our data. We're
- 15:39:49going to do our usual check. So check
- 15:39:51for any NLES, handle those missing nles.
- 15:39:54Um we're actually going to practice
- 15:39:56converting our date because we actually
- 15:39:58do have things based on a date here. So
- 15:40:00we can convert it over to a datetime
- 15:40:02object, extract different uh features
- 15:40:05from that um like the month or the day
- 15:40:09of the week. Um we're going to check uh
- 15:40:12correlation using the heat map. We're
- 15:40:14going to do some plots, some very basic
- 15:40:16plots. The focus is going to be on the
- 15:40:18modeling. So, I probably won't spend too
- 15:40:20much time on the plots today, but um
- 15:40:23there's some plot tasks in here like the
- 15:40:25uh the the histogram of the bike count,
- 15:40:27the histogram of the numerical features.
- 15:40:30Um box plot of the bikes against the
- 15:40:33categoricals.
- 15:40:35Um so we can do some plots like that
- 15:40:38from Seabor for instance. Uh Seabor
- 15:40:40category plot of rented bike count
- 15:40:42against features like hour, holiday,
- 15:40:44rainfall, snowfall. Um, so we can see
- 15:40:47how that stacks up against like
- 15:40:48different hours of the day, different
- 15:40:50holidays, rainfall, different weather
- 15:40:52events. Um, then we're going to do then
- 15:40:55we're going to start building our model,
- 15:40:56right? So encode our categorical
- 15:40:58features. Um, identify target variable
- 15:41:01and do the split and then do scaling and
- 15:41:04do three different models. So we're
- 15:41:06actually going to build a linear
- 15:41:07regression. We're going to build a lasso
- 15:41:09regression and build a ridge regression
- 15:41:11for the hourly bike count. and we're
- 15:41:15going to see which model performs the
- 15:41:16best. Now,
- 15:41:19we could and should um build this into a
- 15:41:24pipeline. So, that could be something we
- 15:41:27practice. Um but this initial
- 15:41:31instructions actually doesn't require
- 15:41:33doing that, but I think it's really good
- 15:41:34practice to um build our model. So, we
- 15:41:38could use git dummies. Like it says
- 15:41:40here, hint to use git dummies. We could
- 15:41:42do that and build our model that way,
- 15:41:45but more practical would be doing the
- 15:41:47steps we did towards the end of lesson
- 15:41:49three, which is um actually putting
- 15:41:51everything together into a pipeline. All
- 15:41:53right, that'd be more practical and then
- 15:41:55fitting the pipeline and predicting with
- 15:41:57it um for evaluation.
- 15:42:01So, uh we'll do that. I think we'll do
- 15:42:04that instead because I think that'll be
- 15:42:07more practical. The the pipelines are
- 15:42:09really uh useful. So the things we'll
- 15:42:12have to do when we build our pipeline
- 15:42:13will be um making sure we handle the
- 15:42:16missing values. So we'll want that imper
- 15:42:18in there for numerical features. We'll
- 15:42:20want our one hot encoding very similar
- 15:42:22pipeline to the one we built earlier. Um
- 15:42:26and then we'll want to basically map
- 15:42:27those to the right columns using the
- 15:42:28column transformer. And then um we'll
- 15:42:31have our pipeline ready to go for
- 15:42:33training and prediction.
- 15:42:36Right.
- 15:42:38Okay. So, those are going to be our uh
- 15:42:41steps. Any questions on this before we
- 15:42:45kind of get started on it.
- 15:42:54Okay. So, let me go over to the
- 15:42:58notebooks and let me actually do a new
- 15:43:02notebook.
- 15:43:08And I'm going to name this uh
- 15:43:12capstone
- 15:43:14session five.
- 15:43:18Okay.
- 15:43:20So, I'm going to come back here and
- 15:43:22reference the steps here. Okay. So the
- 15:43:24steps are to load our data set and
- 15:43:27basically check for nles.
- 15:43:30So we should be pretty adept at doing
- 15:43:32that. We're going to import pandas as
- 15:43:35pd.
- 15:43:38And I need to make sure the data sets
- 15:43:40available. So I need to
- 15:43:45uh load that here. So we're going to use
- 15:43:49our bike rental.
- 15:44:04Florida bike rentals.
- 15:44:22And then look at the first five rows.
- 15:44:34Uh, I got a decoder error.
- 15:44:39UTF8 codec can decode by in
- 15:44:58one moment.
- 15:45:27I think we have an error in the data
- 15:45:29set. Was anybody able to get this to
- 15:45:30run?
- 15:45:32Hopefully they
- 15:45:35or did you get the same error as me?
- 15:45:50giving me the same error.
- 15:45:58I think we need to set a
- 15:46:14encoding
- 15:46:37having other issues. What other issues?
- 15:46:41Sorry, I think this data is kind of
- 15:46:43corrupted.
- 15:46:50I may need to open it externally.
- 15:46:55Okay, let me open it. There may be just
- 15:46:58a bad character that needs to be
- 15:47:01removed.
- 15:47:14Okay,
- 15:47:22let me try a different Let me try
- 15:47:24something real quick.
- 15:47:31I think the data is needs to be updated.
- 15:47:37Oh, saved it as the wrong file.
- 15:47:48One moment.
- 15:48:04Okay, let me try uploading this.
- 15:48:16Okay, there. That worked better. So, let
- 15:48:18me give you the data. I think it was uh
- 15:48:25yeah, I think it was that the degree
- 15:48:28code was giving it some issues. So, I
- 15:48:32actually just removed it.
- 15:48:39You can do that or you could just work
- 15:48:40with this one.
- 15:48:43I removed the temperature had a strange
- 15:48:46like degree symbol that wasn't being
- 15:48:47parsed.
- 15:48:51So, I just removed that in this data
- 15:48:53set. This one should work. The one I
- 15:48:55just sent you guys should work. Or yeah,
- 15:48:57I guess you could try the CP1252
- 15:49:01with the original data. See if that
- 15:49:02works. Did that work for you?
- 15:49:14Okay, it works with that encoding. So
- 15:49:16yeah, you could use that encoding or
- 15:49:19uh
- 15:49:21remove that temperature degree which is
- 15:49:23what I did from that. So I use that
- 15:49:26other data set.
- 15:49:34Yeah, that opens the data but
- 15:49:38it's not it's not in a data frame.
- 15:49:43We just we want it in a dataf frame to
- 15:49:45work with our models and doing all of
- 15:49:46our Yeah. Like that opens the file. If
- 15:49:49it just if it was a text file that's
- 15:49:51fine, but it's not in a data frame. Want
- 15:49:54in a structured data frame so we could
- 15:49:56use it.
- 15:50:04Okay. So, one of those methodologies
- 15:50:07hopefully works. So you can either work
- 15:50:08with the file I sent and read it like
- 15:50:10this or you can use the encoding.
- 15:50:13Did you guys were other people able to
- 15:50:15open it once they change the encoding?
- 15:50:29Okay.
- 15:50:42Okay, very good.
- 15:50:46Let's see what we have. Let's do
- 15:50:47df.info.
- 15:50:48Let's see what we have, which is pretty
- 15:50:50standard step to do once we first load
- 15:50:52in some data. Um, so we have about 14
- 15:50:56columns here. We have a date and then we
- 15:51:00have um which is an object right now but
- 15:51:03we're actually going to convert that
- 15:51:04over to a datetime object in a minute.
- 15:51:07Um we have our bike count which is what
- 15:51:09we want to ultimately predict. This is
- 15:51:11the target variable in our data.
- 15:51:13Remember that's going to be the problem
- 15:51:14is to predict that hourly bike count. Um
- 15:51:18the hour of the day that we are
- 15:51:20producing that bike count um is a
- 15:51:22feature. the temperature which is in
- 15:51:25degrees Celsius. Uh I I removed that
- 15:51:28from there but that's it was in Celsius.
- 15:51:31Um
- 15:51:33and then we have a bunch of different
- 15:51:35features which are uh a kind of boolean
- 15:51:38like yes no holiday no holiday the
- 15:51:42season. So these guys are going to be
- 15:51:46good um candidates for
- 15:51:50uh these are going to be good candidates
- 15:51:52for um doing our uh one hot encoding.
- 15:51:56Right? These are probably the three that
- 15:51:58we should pick to do uh some
- 15:52:01transformations to for our one encoding.
- 15:52:08Okay.
- 15:52:15Everything else is pretty numerical
- 15:52:16though, so those should be fine to keep
- 15:52:18those. But they would they're just going
- 15:52:20to be good candidates for scaling,
- 15:52:22right? Really good candidates for
- 15:52:23scaling.
- 15:52:26All right, let's go back to here.
- 15:52:37And let's go back to the task.
- 15:52:40So we loaded the data set. Um we're we
- 15:52:43are going to check for nulls. It looks
- 15:52:45like there's actually not any nles. So
- 15:52:47there may not be any uh imputing that we
- 15:52:49really need to do for this. Um but we
- 15:52:53can check. So we can use is na or is
- 15:52:57null and then do the sum.
- 15:53:00Um looks like we don't have any. So that
- 15:53:02that's good. There's no nulls.
- 15:53:07That's pretty good. So we don't have to
- 15:53:08worry about filling in any blanks
- 15:53:10really.
- 15:53:12Um but you know that would normally um
- 15:53:16that would be an important part of our
- 15:53:18pipeline right is filling in any nles.
- 15:53:21Looks like we don't have to worry about
- 15:53:22that here.
- 15:53:24So that's good.
- 15:53:29So we can
- 15:53:32uh
- 15:53:34we can extract we can do the date uh
- 15:53:38extraction. And now one thing that we're
- 15:53:39going to do is uh create purposely for
- 15:53:43this we're going to create a weekend or
- 15:53:45weekday feature. Okay, weekday or
- 15:53:50weekend which should be really easy to
- 15:53:52do from the day of the week. um which we
- 15:53:55should be able to extract from this uh
- 15:53:58from the date.
- 15:54:00Let's go back to Were you by the way,
- 15:54:02were you guys able to run this?
- 15:54:04Just want to make sure everyone's with
- 15:54:06me. Checking if there's any nulls. We
- 15:54:09did that.
- 15:54:26Okay.
- 15:54:35All right. So, we're following along
- 15:54:37there. We checked if there's any nles.
- 15:54:39Let's do um our conversion. So, let's do
- 15:54:43um df
- 15:54:45uh date.
- 15:54:48And this is going to be um
- 15:54:51PD.2
- 15:54:53date time and then df
- 15:54:56date.
- 15:54:59All right. And then let's sanity check
- 15:55:01that that it got converted by doing
- 15:55:03info.
- 15:55:25Oh, we might need this format.
- 15:55:43Um, maybe we should do
- 15:55:51Okay, so that actually worked. Let's
- 15:55:54see. Let's double check the
- 15:55:57date.
- 15:55:59Okay, so that worked. It extracted it
- 15:56:01into
- 15:56:03uh it extracted it into the right dates.
- 15:56:09Some weird encoding with this
- 15:56:17So it goes all the way up to 2018
- 15:56:20from
- 15:56:22the very beginning data is in 2017 the
- 15:56:24beginning of the year right. So it go
- 15:56:27and then the tail is all the way in
- 15:56:292018.
- 15:56:32This this extracts the date.
- 15:56:41If you guys want to run that
- 15:56:53you guys able to run this
- 15:56:57to extract the date. Okay, perfect. So,
- 15:57:00this extracts the date. By the way, the
- 15:57:02reason we need this is because our date
- 15:57:04formats are not uniform. Uh they were
- 15:57:07actually in different encodings. So some
- 15:57:08of them had the year uh the string was
- 15:57:11in a slightly different format where the
- 15:57:13year was last or the year was first. So
- 15:57:16if we do format mixed, it kind of
- 15:57:17rearrang it kind of puts it in a uniform
- 15:57:20arrangement with the year. It's year,
- 15:57:22month, date, but it parses that out um
- 15:57:26correctly. Uh so we have year, month,
- 15:57:29date in there.
- 15:57:31Um but the this strings were in a mixed
- 15:57:34format. So we uh put that argument in
- 15:57:38there to handle that case.
- 15:57:43Okay. So the reason we're going to do
- 15:57:45that is that we should be able to create
- 15:57:48a day of the week uh feature. Um so
- 15:57:52let's actually do that and add it to our
- 15:57:54data frame. So we should be able to
- 15:57:55create um day of week
- 15:58:00And we should be able to extract um
- 15:58:05the DF
- 15:58:06date. And then we do our usual DT dot um
- 15:58:12day
- 15:58:14of week.
- 15:58:19Okay. So, and then let's see what that
- 15:58:21does. So, if we add that, let's actually
- 15:58:24see um let's see what our new features
- 15:58:27are.
- 15:58:29So we add that it should go onto the end
- 15:58:32of the data frame as the day of the week
- 15:58:35is a numerical day of the week. So this
- 15:58:37this is the uh um
- 15:58:41this looks like a th or no a Wednesday.
- 15:58:44I think that's the third day.
- 15:58:48Sunday Monday being uh Monday being
- 15:58:51actually this would be a Thursday. I
- 15:58:52think Monday would be zero.
- 15:58:56This would be Thursday.
- 15:59:00Okay. So, we extract that day of the
- 15:59:02week.
- 15:59:07So, it's just we're adding a new column
- 15:59:09called day of week, which extracts the
- 15:59:12day of the week from the date feature,
- 15:59:14the datetime feature, uh the day of the
- 15:59:18week.
- 15:59:23And we're verifying that here. It's now
- 15:59:25a new feature called day of the week.
- 15:59:27which is a number.
- 15:59:40Does that make sense what this is doing?
- 15:59:48Yeah, it's a Thursday. So, I think yeah,
- 15:59:50Monday is zero
- 15:59:53and Sunday is six. So,
- 15:59:57um, Monday is
- 15:59:59zero, Sunday is
- 16:00:03six.
- 16:00:05Yeah. So, this should be a Thursday.
- 16:00:14This first five rows is a Thursday. And,
- 16:00:16and by the way, the data um, this is all
- 16:00:18on the Thursday. This is a different
- 16:00:20hours of the day.
- 16:00:23Different hours of the day. Um, and we
- 16:00:27the the thing that we don't know about
- 16:00:28this is the necessarily the time zone.
- 16:00:30So, it may seem strange that like this
- 16:00:32is zero, which is kind of like midnight.
- 16:00:35Um, it could be it could be in a
- 16:00:37different time zone. So, we don't know
- 16:00:38that uh necessarily, but the there is um
- 16:00:43you know 250 bikes, 200 bikes, 173. It
- 16:00:47starts to decrease over these hours.
- 16:00:51Just kind of notice that.
- 16:01:11Okay. So, by the way, uh we should be
- 16:01:14able to create a new feature. So let's
- 16:01:16create
- 16:01:18the weekend feature
- 16:01:21which is um we can create using
- 16:01:25uh weekend
- 16:01:28and we can take our DF
- 16:01:32um day of week
- 16:01:37and then we can just um
- 16:01:40say is this um greater than or equal to
- 16:01:45uh greater than or equal to five
- 16:01:50cuz that would be five or six. And we're
- 16:01:52going to
- 16:01:54um put this as
- 16:02:11actually. Let's just leave it like that.
- 16:02:12That should be fine.
- 16:02:17No, let's change it as type
- 16:02:20uh int.
- 16:02:25So, let's see this feature.
- 16:02:35Let's see if this works.
- 16:02:44So this is not a weekend uh because it's
- 16:02:47a Thursday, right? So anything bigger
- 16:02:49than five would be bigger than or equal
- 16:02:51to five would be five or six which would
- 16:02:53be Saturday, Sunday. Um so we know it's
- 16:02:56not a weekend day. This is uh zero.
- 16:03:02Okay.
- 16:03:09So just creating those features there
- 16:03:14and I can paste these in. Does that make
- 16:03:16sense what what I just did?
- 16:03:19Any questions on that code? And this is
- 16:03:22important. If you don't have this um it
- 16:03:25this will be a true or a false, but we
- 16:03:28want it to be a zero or a one. So when
- 16:03:31it's actually true, this should be a
- 16:03:33one. When it's false, it'll be a zero.
- 16:03:36So, we want it to be an integer rather
- 16:03:38than a true false. So, that's why I have
- 16:03:40this part. That's why I did this here
- 16:03:44to make sure it's um make sure it's an
- 16:03:46integer.
- 16:04:01Good.
- 16:04:05Okay,
- 16:04:07so we're pretty much uh doing that and
- 16:04:09we convert it to day and extract day. Um
- 16:04:14let's
- 16:04:16check our heat map.
- 16:04:20Let's do that.
- 16:04:26So let's import Seabour
- 16:04:30as SNS.
- 16:04:33So, we're going to do our heat map next.
- 16:04:35Let me uh make a note of that. So, we're
- 16:04:38going to do
- 16:04:40heat mapap
- 16:04:42heat map. Um,
- 16:04:45so we're going to do uh SNS
- 16:04:50heat map
- 16:04:52and then let's do uh df.correlation.
- 16:04:56And let's make sure we do numeric
- 16:05:00um only
- 16:05:02equals to true
- 16:05:10and then let's do annotate
- 16:05:13equals to true
- 16:05:15so that we get those uh correlation
- 16:05:17values that are displayed on the heat
- 16:05:19map. So what this is going to do is
- 16:05:21create our heat map with our correlation
- 16:05:23matrix um where we're making sure we
- 16:05:26only do the numerical features of course
- 16:05:28when we do uh the heat map.
- 16:05:34Let's generate that. Um okay so we have
- 16:05:39some
- 16:05:41uh pretty mild um correlations. Now the
- 16:05:45ones that are so the ones that are
- 16:05:48correlated are the temperature and
- 16:05:50dupoint temperature. Those are pretty
- 16:05:52correlated.
- 16:05:54Um so I think we could argue that we
- 16:05:56should drop one of those. Probably just
- 16:05:59the dupoint temperature we could drop.
- 16:06:01Um that's a really high correlation
- 16:06:04right right here.
- 16:06:06Let me draw it in red. That's a this
- 16:06:08this one here
- 16:06:10which is the Dupoint temperature against
- 16:06:12the regular temperature. That's super
- 16:06:14high. 0.91
- 16:06:16is nearly a onetoone correlation.
- 16:06:20So, that's a good candidate to be
- 16:06:22dropped. Uh, one of those guys, I would
- 16:06:24argue probably the Dupoint temperature
- 16:06:26we could get rid of dropping. Um, and
- 16:06:30just keep the regular temperature
- 16:06:31because they're nearly identical. Um, if
- 16:06:36you go down here,
- 16:06:38day of the week and weekend are
- 16:06:40correlated. Um, and that makes sense.
- 16:06:42That's a pretty strong correlation
- 16:06:44because of course if it's depending on
- 16:06:47what day of the week it is, it is the
- 16:06:48weekend or not. So we could go and we
- 16:06:52could go ahead and drop the day of the
- 16:06:54week column if we wanted to because
- 16:06:55we've already derived the weekend uh
- 16:06:58feature which is a simpler feature. Is
- 16:07:00it the weekend or is it not the weekend?
- 16:07:02Um so we could probably drop day of week
- 16:07:07and be okay. That's a pretty strong
- 16:07:09correlation. Otherwise, it's all pretty
- 16:07:12weak. I don't see any other strong
- 16:07:14correlations
- 16:07:16uh necessarily.
- 16:07:18Um so
- 16:07:21that seems pretty reasonable is that we
- 16:07:23could get rid of we could get rid of
- 16:07:24this one and we could get rid of day of
- 16:07:27week and probably be okay,
- 16:07:30right? Those are pretty strong
- 16:07:32correlations. is 08 and 0.91.
- 16:07:34Pretty strong.
- 16:07:42Okay. Were you guys able to run that one
- 16:07:43and see the the uh heat map? And does
- 16:07:46that make sense based on what I'm
- 16:07:48saying?
- 16:07:52So, we want to run the heat map
- 16:07:55and pass in that correlation. And we
- 16:07:56want to make sure we turn this to true
- 16:07:58to only do the numerical features. And
- 16:08:00then this to true to show the value
- 16:08:04on the uh on the heat map.
- 16:08:25able to run that one.
- 16:08:32Great.
- 16:08:47Okay.
- 16:08:54Um, that's a good question. Any
- 16:08:56recommendation on how to choose between
- 16:08:57the two? Uh,
- 16:09:00not really. I think
- 16:09:04I don't think it really matters. If
- 16:09:05they're correlated to each other, then
- 16:09:08uh including one of them,
- 16:09:11um, including one of them should give
- 16:09:13you the same information as the other,
- 16:09:15especially if they're really correlated.
- 16:09:17So, it doesn't really matter too much.
- 16:09:19Um
- 16:09:21the way that I would choose is to think
- 16:09:24about like I'll give you an example in
- 16:09:26this day of week versus weekend. Um we
- 16:09:30could I would argue drop day of week
- 16:09:34because the weekend is a simpler
- 16:09:36feature. It's only is zero or one and
- 16:09:39that's directly derived from day of
- 16:09:40week.
- 16:09:42So it has less um complexity to it. it's
- 16:09:47a little bit simpler of a feature and I
- 16:09:48think that makes it easier to work with.
- 16:09:51Um,
- 16:09:53but the truth is that uh we could do
- 16:09:57both options and try them out and
- 16:09:59evaluate the results, right? So, well to
- 16:10:02be to be truly thorough, what we could
- 16:10:05do is build a model where we've dropped
- 16:10:07this one and do the evaluation and then
- 16:10:09go back and build a model where we've
- 16:10:11dropped this one and do the evaluation.
- 16:10:14Right? So that's the proper way to do it
- 16:10:17is to actually just build both models
- 16:10:19with each one dropped and see which
- 16:10:21performs better.
- 16:10:28Um otherwise I I tend to prefer to go
- 16:10:31for the simplicity whatever one has kind
- 16:10:33of a lower range.
- 16:10:39But the truth is if they're if they're
- 16:10:41really strongly correlated, it's not
- 16:10:42going to matter too much. Uh because
- 16:10:44they're going to give you the same
- 16:10:45information, right? Uh because they're
- 16:10:48so strongly correlated. Like if I in
- 16:10:51this data, like if I know the
- 16:10:53temperature, I pretty much know what the
- 16:10:54Dupoint temperature is going to be.
- 16:10:57They're so correlated.
- 16:10:59So it doesn't it doesn't really matter
- 16:11:01which one I drop
- 16:11:09but I yeah prefer to go for the
- 16:11:10simplicity.
- 16:11:16Okay.
- 16:11:18Um going back to let's let's do this
- 16:11:23plot now of the distribution of the
- 16:11:25rented bike count. So that should be
- 16:11:28taking a look at the
- 16:11:30um
- 16:11:32the rented
- 16:11:35bike count and then just doing a
- 16:11:38histogram
- 16:11:43and we can see what that distribution
- 16:11:45is. Most of it is it less than 250.
- 16:11:50Um but there are some values that are
- 16:11:53really high, right? There are some days
- 16:11:56that turn out to be over 3,000 into the
- 16:11:593500 range. Um, that's quite a bit, but
- 16:12:03most of the days are stacked over here
- 16:12:05in this like 250 bucket. So, by far
- 16:12:08that's the most. And then it kind of
- 16:12:10decreases from there. Most values are in
- 16:12:14that range and then uh it kind of
- 16:12:17declines.
- 16:12:22So that should just be this simple
- 16:12:24histogram here.
- 16:12:27So this is the
- 16:12:34histogram.
- 16:12:38That should be a simple one to build.
- 16:12:44Were
- 16:12:50you guys able to run that one?
- 16:13:00Sweet.
- 16:13:02Okay, pretty basic. Just showing us how
- 16:13:05it's distributed. And we kind of noticed
- 16:13:06that most of it is in the 250 bucket or
- 16:13:09below. But there's a good amount that's
- 16:13:11out there beyond like in the 500,
- 16:13:147500,000
- 16:13:16um
- 16:13:18all the way up to there's some days that
- 16:13:20register with a 3,000 and above, right?
- 16:13:233500.
- 16:13:24So,
- 16:13:31in fact, we could look, we didn't do
- 16:13:33this, but it might be worth doing is we
- 16:13:35could look at the describe because we
- 16:13:38never looked at the maximum.
- 16:13:42Um,
- 16:13:44so for the bike count, there are some
- 16:13:47days that are zero
- 16:13:50and the maximum is 3500. 3556
- 16:13:56and the median is around 500 bikes.
- 16:14:00Um
- 16:14:02the average around 700.
- 16:14:18So that's just our usual describe
- 16:14:30All right, let's see what else. Uh,
- 16:14:34plot the histogram of all numerical
- 16:14:36features. So, uh, let's do that.
- 16:14:48So uh luckily we have a shortcut to do
- 16:14:51this. If you guys remember we have our
- 16:14:54SNS um pair plot and we can pass in our
- 16:15:01uh our data
- 16:15:03is our df right so we can pass that in.
- 16:15:07Um so this is kind of a nice this is a
- 16:15:10nice thing to run that will um generate
- 16:15:14the uh the scatter plots of all the
- 16:15:17features against each other kind of like
- 16:15:19the correlation but at the same time
- 16:15:21produce the histograms on the diagonal.
- 16:15:24Right? So this should be a nice plot to
- 16:15:26to see all the histograms
- 16:15:29uh in one kind of uh grid.
- 16:16:00That's just this one.
- 16:16:15Okay, so pretty this is obviously quite
- 16:16:18a bit of data, but this is all the
- 16:16:20features against each other. Um, look at
- 16:16:23this. I mean, a couple things you see
- 16:16:25right away is look at the ones that are
- 16:16:26really highly correlated like the
- 16:16:27temperature against the Dupoint
- 16:16:29temperature. Do you guys see this strong
- 16:16:32correlation here?
- 16:16:34So that's that's a very indicative of a
- 16:16:37very strong correlation, right? That's
- 16:16:39the temperature against the Dupoint
- 16:16:40temperature. That kind of makes sense
- 16:16:42that it's uh that was the 0.91
- 16:16:45correlation. So of course it's like
- 16:16:47that.
- 16:16:51Of course it looks like that, right? Uh
- 16:16:53just a very strong correlation on the on
- 16:16:56the x-axis or sorry on the diagonal is
- 16:16:58all the histograms. They're a little bit
- 16:17:00zoomed out so difficult to see. Um but
- 16:17:04um we can get a sense of how some of
- 16:17:07these are distributed like the first one
- 16:17:10uh
- 16:17:12sorry this first one
- 16:17:15which is the bite count we already did
- 16:17:17um temperature we can see how that's
- 16:17:20distributed this third one it's kind of
- 16:17:23evenly distributed
- 16:17:26left of due versus temperature this one
- 16:17:29that one is the hours
- 16:17:33versus the due point. So, it it's kind
- 16:17:36of evenly spaced out. Um, which would
- 16:17:40probably be which would make sense
- 16:17:42because this data is across two years.
- 16:17:44So, you're going to get a lot of
- 16:17:45seasonal data in there, right? So, it's
- 16:17:48going to it's going to vary across like
- 16:17:50seasons. Yeah. So, it kind of looks it's
- 16:17:52very evenly spread out.
- 16:17:55Okay. Were you able to get Parpot to
- 16:17:57run?
- 16:17:58Takes a moment to run.
- 16:18:01takes a moment, but it produces all of
- 16:18:02these plots, including all the
- 16:18:03histograms.
- 16:18:05And again, if we wanted to zoom in on
- 16:18:07any one particular histogram, we could
- 16:18:09do that. We just have to basically copy
- 16:18:11and paste this code and swap out this
- 16:18:13feature. Just swap out this feature and
- 16:18:15we can get a zoomed in uh plot of any
- 16:18:18one of those uh features like the
- 16:18:21temperature
- 16:18:23um or visibility or whatever it is.
- 16:18:27So, what I want to do, uh, I think we'll
- 16:18:28take one more break here. Um, and then
- 16:18:31what we'll do is we'll come back and
- 16:18:33just finish up our practice. Um, I'm
- 16:18:35going to start building the model. I
- 16:18:37know it wants us to do some additional
- 16:18:38plotting. Um, but I want to get to
- 16:18:42mainly the plotting elements, including
- 16:18:44doing the pipeline one more time. Um, so
- 16:18:49we're going to do that. Um, we're going
- 16:18:52to practice building out the pipeline in
- 16:18:53a similar fashion to exactly how we
- 16:18:56built the pipeline earlier and then do
- 16:18:59we're going to build our models and
- 16:19:00we're going to practice that coming up.
- 16:19:02I'll probably leave the plotting to you
- 16:19:04guys to do as kind of homework if you
- 16:19:06want to do that. Um, just because I want
- 16:19:08to get to the to the modeling. Hello and
- 16:19:11welcome to machine learning tutorial
- 16:19:13part one. This is part one of a machine
- 16:19:15learning series put on by SimplyLearn.
- 16:19:18My name is Richard Kersner. I'm with the
- 16:19:20SimplyLearn team. That's
- 16:19:21www.simplearn.com.
- 16:19:23Get certified, get ahead. What's in it
- 16:19:26for you today? Well, we'll start off
- 16:19:28with a brief explanation of why machine
- 16:19:30learning and what is machine learning.
- 16:19:32And then we'll get into a few of the
- 16:19:34types of machine learning. machine
- 16:19:35learning algorithms, linear regression,
- 16:19:38decision trees, support vector machine,
- 16:19:41and finally, we'll do a use case where
- 16:19:43we're going to classify whether a recipe
- 16:19:44is of a cupcake or a muffin using the
- 16:19:47SPM or the support vector machine.
- 16:19:49Sounds like a delicious way to explore
- 16:19:51machine learning. So, why machine
- 16:19:53learning? Why do we even care about
- 16:19:55having these computers come up and be
- 16:19:57able to do all these new things for us?
- 16:19:59Well, because machines can now drive
- 16:20:01your car for you. still very in the
- 16:20:03infant stage but it's just exploding as
- 16:20:05we see with uh Google's Whimo and then
- 16:20:08Uber had their program which
- 16:20:10unfortunately crashed. They know that
- 16:20:12this is huge. This is going to be the
- 16:20:13huge industry to change our whole
- 16:20:15transportation infrastructure. Machine
- 16:20:18learning is now used to detect over 50
- 16:20:20eye diseases. Do you know how amazing
- 16:20:22that is to have a computer that doublech
- 16:20:24checkcks for the doctor for things they
- 16:20:25might miss? That's just huge in the
- 16:20:27health industry. pretty soon they
- 16:20:29actually do already have that with in
- 16:20:30some areas where maybe not for eyes but
- 16:20:33for other diseases where they're using
- 16:20:34the camera on your phone to help
- 16:20:36pre-diagnose before you go in and see
- 16:20:38the doctor. And because the machine can
- 16:20:40now unlock your phone with your face, I
- 16:20:43mean, that's just cool having it being
- 16:20:44able to identify your face or your voice
- 16:20:47and be able to turn stuff on and off for
- 16:20:49you depending on where you're at and
- 16:20:50what you need. Talk about an ultimate
- 16:20:52automation our world we live in. And as
- 16:20:55we dig in deeper, we have a nice example
- 16:20:57of Facebook. As you can see here, they
- 16:20:59have the Facebook post with Halloween.
- 16:21:01Comment yes if you want it order here.
- 16:21:04Nobody likes spam posts on Facebook that
- 16:21:07annoy them into interacting with likes,
- 16:21:09shares, comments, and other actions. I
- 16:21:12remember the original ones were all if
- 16:21:13you don't click on here, you will have
- 16:21:16bad luck or some kind of fear factor.
- 16:21:18Well, this is a huge thing in a social
- 16:21:20media when people are getting spammed.
- 16:21:22And so this tactic known as engagement
- 16:21:25bait takes advantage of Facebook's
- 16:21:27newsfeed algorithm by choosing
- 16:21:29engagement in order to get the greater
- 16:21:31reach. To eliminate engagement bait, the
- 16:21:34company reviewed and categorized
- 16:21:36hundreds of thousands of posts to train
- 16:21:38a machine learning model that detects
- 16:21:39different types of engagement bait. So
- 16:21:41in this case, we have we're using
- 16:21:42Facebook, but this is of course across
- 16:21:44all the different social media. they
- 16:21:46have different tools are building and
- 16:21:47the Facebook scroll gif will be replaced
- 16:21:50kind of like a virus coming in there and
- 16:21:52notices that there's a certain setup
- 16:21:54with Facebook and it's able to replace
- 16:21:56it and they have like vote baiting react
- 16:21:59baiting share baiting they have all
- 16:22:01these different these are kind of
- 16:22:03general titles but there certainly are a
- 16:22:05lot of way of baiting you to go in there
- 16:22:06and click on something so they fed all
- 16:22:08this this data was fed into the machine
- 16:22:10and then they have the new post the new
- 16:22:12post comes up that takes over part of
- 16:22:14the Facebook setup up and that's what
- 16:22:16you're looking at. You're looking at
- 16:22:17this new post that's replaced like a
- 16:22:19virus has replaced that. So what
- 16:22:20Facebook did to eliminate this is they
- 16:22:22start scanning for keywords and phrases
- 16:22:24like this and checks the click-through
- 16:22:26rate. So it starts looking for people
- 16:22:28who are clicking through it without even
- 16:22:30looking at it or clicking through it and
- 16:22:32it's not something that normally would
- 16:22:33be clicked through. Once Facebook has
- 16:22:35scanned for these keywords and phrases,
- 16:22:37it is now able to identify the spam
- 16:22:40coming in and this makes your life
- 16:22:42easier. So you're not getting spammed.
- 16:22:44It's not like walking through an airport
- 16:22:45and in a lot of countries you have like
- 16:22:47hundreds of people trying to sell you
- 16:22:48time share. Come join us. Sign up for
- 16:22:51this. Eliminates that annoyingness. So
- 16:22:52now you can just enjoy your Facebook and
- 16:22:54your cat pictures. Or maybe it's your
- 16:22:56family pictures. Mine is family.
- 16:22:58Certainly people like their cat pictures
- 16:23:00too. Another good example is Google's
- 16:23:02Deep Mind project Alph Go. A computer
- 16:23:05program that plays a board game Go has
- 16:23:07defeated the world's number one go
- 16:23:09player and I hope I say his name right.
- 16:23:12Kijiji the ultimate go challenge game a
- 16:23:14three of three was on May 27th 2017 so
- 16:23:18that was just last year that this
- 16:23:20happened and what makes this so
- 16:23:22important is that you know go is just is
- 16:23:25a game so it's not like you're driving a
- 16:23:26car or something in our real world but
- 16:23:29they are using games to learn how to get
- 16:23:32the machine learning program to learn
- 16:23:35they want it to learn how to learn and
- 16:23:36that is a huge step a lot of this is
- 16:23:39still in its infant stage as far as
- 16:23:40development
- 16:23:41as we saw what happened with the as I
- 16:23:44referred to earlier the Uber cars. They
- 16:23:46lost their whole division because they
- 16:23:48jumped ahead too fast. So still an
- 16:23:50infant stage, but boy is this like the
- 16:23:52beginning of just an amazing world that
- 16:23:55is automated in ways we can't even
- 16:23:57imagine what tomorrow's going to look
- 16:23:59like. We've looked at a lot of examples
- 16:24:01of machine learning. So let's see if we
- 16:24:03can give a little bit more of a concrete
- 16:24:05definition. What is machine learning?
- 16:24:08Machine learning is the science of
- 16:24:09making computers learn and act like
- 16:24:11humans by feeding data and information
- 16:24:13without being explicitly programmed. And
- 16:24:16we see here we have a nice little
- 16:24:17diagram where we have our ordinary
- 16:24:19system, your computer nowadays, you can
- 16:24:22even run a lot of the stuff on a cell
- 16:24:24phone because cell phones have advanced
- 16:24:25so much. And then with artificial
- 16:24:27intelligence and machine learning, it
- 16:24:29now takes the data and it learns from
- 16:24:32what happened before and then it
- 16:24:33predicts what's going to come next. And
- 16:24:36then really the biggest part right now
- 16:24:38in machine learning that's going on is
- 16:24:39it improves on that. How do we find a
- 16:24:42new solution? So we go from descriptive
- 16:24:45where it's learning about stuff and
- 16:24:46understanding how it fits together to
- 16:24:48predicting what it's going to do to
- 16:24:50postcripting coming up with a new
- 16:24:52solution. And when we're working on
- 16:24:55machine learning, there's a number of
- 16:24:56different diagrams that people have
- 16:24:58posted for what steps to go through. A
- 16:25:00lot of it might be very domain specific.
- 16:25:03So if you're working on photo
- 16:25:05identification versus language versus
- 16:25:08medical or physics, some of these are
- 16:25:11switched around a little bit or new
- 16:25:12things are put in. They're very specific
- 16:25:14to the domain. This is kind of a very
- 16:25:15general diagram. First, you want to
- 16:25:17define your objective. Very important to
- 16:25:20know what it is you're wanting to
- 16:25:21predict. Then you're going to be
- 16:25:22collecting the data. So once you've
- 16:25:24defined an objective, you need to
- 16:25:25collect the data that matches. You spend
- 16:25:28a lot of time in data science collecting
- 16:25:30data and the next step preparing the
- 16:25:32data. You got to make sure that your
- 16:25:34data is clean going in. There's the old
- 16:25:36saying, bad data in, bad answer out or
- 16:25:40bad data out. And then once you've gone
- 16:25:43through and we've cleaned all this stuff
- 16:25:44coming in, then you're going to select
- 16:25:47the algorithm. Which algorithm are you
- 16:25:49going to use? You're going to train that
- 16:25:50algorithm. In this case, I think we're
- 16:25:52going to be working with SVM, the
- 16:25:54support vector machine. Then you have to
- 16:25:56test the model. Does this model work? Is
- 16:25:58this a valid model for what we're doing?
- 16:26:00And then once you've tested it, you want
- 16:26:02to run your prediction. You want to run
- 16:26:04your prediction or your choice or
- 16:26:06whatever output it's going to come up
- 16:26:07with. And then once everything is set
- 16:26:09and you've done lots of testing, then
- 16:26:12you want to go ahead and deploy the
- 16:26:13model. And remember I said domain
- 16:26:15specific. This is very general as far as
- 16:26:17the scope of doing something. A lot of
- 16:26:19models you get halfway through and you
- 16:26:21realize that your data is missing
- 16:26:23something and you have to go collect new
- 16:26:24data because you've run a test in here
- 16:26:26someplace along the line. You're saying,
- 16:26:28"Hey, I'm not really getting the answers
- 16:26:29I need." So, there's a lot of things
- 16:26:30that are domain specific that become
- 16:26:32part of this model. This is a very
- 16:26:34general model, but it's a very good
- 16:26:35model to start with. And we do have some
- 16:26:38basic divisions of what machine learning
- 16:26:40does that's important to know. For
- 16:26:42instance, do you want to predict a
- 16:26:44category? Well, if you're categorizing
- 16:26:46thing, that's classification. For
- 16:26:48instance, whether the stock price will
- 16:26:50increase or decrease. So in other words,
- 16:26:52I'm looking for a yes no answer. Is it
- 16:26:54going up or is it going down? And in
- 16:26:56that case, we'd actually say, is it
- 16:26:57going up? True. If it's not going up,
- 16:26:59it's false, meaning it's going down.
- 16:27:01This way, it's a yes, no. 01. Do you
- 16:27:04want to predict a quantity? That's
- 16:27:06regression. So remember, we just did
- 16:27:08classification. Now we're looking at
- 16:27:10regression. These are the two major
- 16:27:12divisions in what data is doing. For
- 16:27:14instance, predicting the age of a person
- 16:27:16based on the height, weight, health, and
- 16:27:18other factors. So based on these
- 16:27:20different factors, you might guess how
- 16:27:21old a person is. And then there are a
- 16:27:23lot of domain specific things like do
- 16:27:26you want to detect an anomaly? That's
- 16:27:28anomaly detection. This is actually very
- 16:27:31popular right now. For instance, you
- 16:27:32want to detect money withdrawal
- 16:27:33anomalies. You want to know when
- 16:27:34someone's making a withdrawal that might
- 16:27:36not be their own account. We've actually
- 16:27:38brought this up because this is really
- 16:27:40big right now. If you're predicting the
- 16:27:42stock whether to buy stock or not, you
- 16:27:44want to be able to know if what's going
- 16:27:45on in the stock market is an anomaly,
- 16:27:48use a different prediction model because
- 16:27:49something else is going on. You got to
- 16:27:51pull out new information in there or is
- 16:27:53this just the norm? I'm going to get my
- 16:27:55normal return on my money invested. So
- 16:27:58being able to detect anomalies is very
- 16:27:59big in data science these days. Another
- 16:28:02question that comes up which is on what
- 16:28:04we call untrained data is do you want to
- 16:28:07discover structure in unexplored data
- 16:28:10and that's called clustering. For
- 16:28:12instance, finding groups of customers
- 16:28:14with similar behavior given a large
- 16:28:16database of customer data containing
- 16:28:18their demographics and past buying
- 16:28:20records. And in this case, we might
- 16:28:23notice that anybody who's wearing
- 16:28:25certain set of shoes goes shopping at
- 16:28:27certain stores or whatever it is. are
- 16:28:29going to make certain purchases. By
- 16:28:31having that information, it helps us to
- 16:28:33market or group people together. So then
- 16:28:35we can now explore that group and find
- 16:28:37out what it is we want to market to them
- 16:28:39if you're in the marketing world. And
- 16:28:40that might also work in just about any
- 16:28:42arena. You might want to group people
- 16:28:44together whether they're uh based on
- 16:28:47their different areas and investments
- 16:28:49and financial background, whether you're
- 16:28:52going to give them a loan or not. before
- 16:28:54you even start looking at whether
- 16:28:55they're a valid customer for the bank,
- 16:28:57you might want to look at all these
- 16:28:58different areas and group them together
- 16:29:00based on unknown data. So, you're not
- 16:29:02you don't know what the data is going to
- 16:29:03tell you, but you want to cluster people
- 16:29:04together that come together. Let's take
- 16:29:07a quick detour for quiz time. Oh, my
- 16:29:10favorite. So, we're going to have a
- 16:29:12couple questions here under quiz time
- 16:29:15and um we'll be posting the answers in
- 16:29:17these part two of this tutorial. So,
- 16:29:20let's go ahead and take a look at these
- 16:29:22quiz times questions and hopefully
- 16:29:23you'll get them all right and it'll get
- 16:29:25you thinking about how to process data
- 16:29:27and what's going on. Can you tell what's
- 16:29:29happening in the following cases? Of
- 16:29:31course, you're sitting there with your
- 16:29:33cup of coffee and you have your checkbox
- 16:29:34and your pen trying to figure out what's
- 16:29:36your next step in your data science
- 16:29:38analysis. So, the first one is grouping
- 16:29:41documents into different categories
- 16:29:44based on the topic and content of each
- 16:29:46document. Very big these days. you know,
- 16:29:49you have legal documents, you have uh
- 16:29:52maybe it's a sports group documents,
- 16:29:53maybe you're analyzing newspaper
- 16:29:55postings, but certainly having that
- 16:29:58automated is a huge thing in today's
- 16:30:00world. B, identifying handwritten digits
- 16:30:03in images correctly. So, we want to know
- 16:30:06whether uh they're writing an A or
- 16:30:07capital A, B, C, what are they writing
- 16:30:10out in their hand digit, their
- 16:30:11handwriting. C behavior of a website
- 16:30:14indicating that the site is not working
- 16:30:17as designed. D, predicting salary of an
- 16:30:21individual based on his or her years of
- 16:30:24experience with HR hiring uh setup
- 16:30:27there. So stay tuned for part two. We'll
- 16:30:29go ahead and answer these questions when
- 16:30:31we get to the part two of this tutorial
- 16:30:33or you can just simply write at the
- 16:30:35bottom and send a note to SimplyLearn
- 16:30:36and they'll follow up with you on it.
- 16:30:39Back to our regular content. Now these
- 16:30:41last few bring us into the next topic
- 16:30:44which is another way of dividing our
- 16:30:45types of machine learning and that is
- 16:30:47with supervised unsupervised
- 16:30:51and reinforcement learning. Supervised
- 16:30:54learning is a method used to enable
- 16:30:56machines to classify predict objects,
- 16:30:58problems or situations based on labeled
- 16:31:01data fed to the machine. And in here you
- 16:31:03see we have a jumble of data with
- 16:31:05circles, triangles and squares. And we
- 16:31:08label them. We have what's a circle,
- 16:31:09what's a triangle, what's a square and
- 16:31:11we have our model training and it trains
- 16:31:13it. So we know the answer. Very
- 16:31:15important when you're doing supervised
- 16:31:16learning, you already know the answer to
- 16:31:18a lot of your information coming in. So
- 16:31:21you have a huge group of data coming in
- 16:31:23and then you have new data coming in. So
- 16:31:25we've trained our model. The model now
- 16:31:27knows the difference between a circle, a
- 16:31:29square, a triangle. And now that we've
- 16:31:31trained it, we can send in in this case
- 16:31:33a square and a circle goes in and it
- 16:31:35predicts that the top one's a square and
- 16:31:37the next one's a circle. And you can see
- 16:31:39that this is uh being able to predict
- 16:31:41whether someone's going to default on a
- 16:31:42loan because I was talking about banks
- 16:31:44earlier. Supervised learning on stock
- 16:31:46market whether you're going to make
- 16:31:48money or not. That's always important.
- 16:31:50And if you are looking to make a fortune
- 16:31:52in the stock market, keep in mind it is
- 16:31:54very difficult to get all the data
- 16:31:56correct on the stock market. It is very
- 16:31:58uh it fluctuates in ways you really hard
- 16:32:00to predict. So it's quite a roller
- 16:32:03coaster ride. If you're running machine
- 16:32:04learning on the stock market, you start
- 16:32:06realizing you really have to dig for new
- 16:32:08data. So we have supervised learning.
- 16:32:10And if you have supervised, we need
- 16:32:12unsupervised learning. In unsupervised
- 16:32:15learning, machine learning model finds
- 16:32:17the hidden pattern in an unlabeled data.
- 16:32:20So in this case, instead of telling it
- 16:32:22what the circle is and what a triangle
- 16:32:24is and what a square is, it goes in
- 16:32:26there, looks at them, and says for
- 16:32:27whatever reason, it groups them
- 16:32:29together. Maybe it'll group it by the
- 16:32:30number of corners. And it notices that a
- 16:32:33number of them all have three corners, a
- 16:32:35number of them all have four corners,
- 16:32:36and a number of them all have no
- 16:32:38corners. And it's able to filter those
- 16:32:40through and group them together. We
- 16:32:41talked about that earlier with looking
- 16:32:43at a group of people who are out
- 16:32:44shopping. We want to group them together
- 16:32:46to find out what they have in common.
- 16:32:48And of course, once you understand what
- 16:32:50people have in common, maybe you have
- 16:32:52one of them who's a customer at your
- 16:32:54store, or you have five of them are
- 16:32:55customer at your store, and they have a
- 16:32:57lot in common with five others who are
- 16:32:59not customers at your store. How do you
- 16:33:01market to those five who aren't
- 16:33:02customers at your store yet? They fit
- 16:33:04the demographs of who's going to shop
- 16:33:05there, and you'd like them to shop at
- 16:33:07your store, not the one next door. Of
- 16:33:09course, this is a simplified version.
- 16:33:10You can see very easily the difference
- 16:33:12between a triangle and a circle, which
- 16:33:13is might not be so easy in marketing.
- 16:33:15Reinforcement learning. Reinforcement
- 16:33:17learning is an important type of machine
- 16:33:19learning where an agent learns how to
- 16:33:21behave in an environment by performing
- 16:33:23actions and seeing the result. And we
- 16:33:25have here where the in this case a baby.
- 16:33:28It's actually great that they used an
- 16:33:29infant for this slide because the
- 16:33:31reinforcement learning is very much in
- 16:33:33its infant stages. But it's also
- 16:33:35probably the biggest machine learning
- 16:33:38demand out there right now or in the
- 16:33:40future. It's going to be coming up over
- 16:33:41the next few years is reinforcement
- 16:33:43learning and how to make that work for
- 16:33:45us. And you can see here where we have
- 16:33:47our action. In the action in this one,
- 16:33:49it goes into the fire. Hopefully, the
- 16:33:51baby didn't it's just a little candle,
- 16:33:53not a giant fire pit like it looks like
- 16:33:54here. When the baby comes out and the
- 16:33:56new state is the baby is sad and crying
- 16:33:59because they got burned on the fire. And
- 16:34:00then maybe they take another action. The
- 16:34:02baby's called the agent because it's the
- 16:34:04one taking the actions. And in this
- 16:34:06case, they didn't go into the fire. They
- 16:34:07went a different direction. And now the
- 16:34:09baby's happy and laughing and playing.
- 16:34:11Reinforcement learning is very easy to
- 16:34:13understand because that's how as humans
- 16:34:15that's one of the ways we learn. We
- 16:34:17learn whether it is you burn yourself on
- 16:34:19the stove, don't do that anymore. Don't
- 16:34:21touch the stove. In the big picture,
- 16:34:23being able to have machine learning
- 16:34:25programming or an AI be able to do this
- 16:34:27is huge because now we're starting to
- 16:34:29learn how to learn. That's a big jump in
- 16:34:33the world of computer and machine
- 16:34:34learning. And we're going to go back and
- 16:34:36just kind of go back over supervised
- 16:34:38versus unsupervised learning.
- 16:34:40Understanding this is huge because this
- 16:34:42is going to come up in any project
- 16:34:44you're working on. We have in supervised
- 16:34:47learning, we have labeled data. We have
- 16:34:49direct feedback. So someone's already
- 16:34:51gone in there and said, "Yes, that's a
- 16:34:53triangle. No, that's not a triangle."
- 16:34:55And then you predict an outcome. So you
- 16:34:56have a nice prediction. This is this
- 16:34:58this new set of data is coming in and we
- 16:35:00know what it's going to be. And then
- 16:35:01with unsupervised trading, it's not
- 16:35:03labeled. So we really don't know what it
- 16:35:06is. There's no feedback. So, we're not
- 16:35:08telling it whether it's right or wrong.
- 16:35:10We're not telling it whether it's a
- 16:35:12triangle or a square. We're not telling
- 16:35:14it to go left or right. All we do is
- 16:35:16we're finding hidden structure in the
- 16:35:18data, grouping the data together to find
- 16:35:20out what connects to each other. And
- 16:35:23then you can use these together. So,
- 16:35:25imagine you have an image and you're not
- 16:35:27sure what you're looking for. So, you go
- 16:35:29in and you have the unstructured data.
- 16:35:32Find all these things that are connected
- 16:35:34together and then somebody looks at
- 16:35:35those and labels them. Now you can take
- 16:35:38that label data and program something to
- 16:35:40predict what's in the picture. So you
- 16:35:42can see how they go back and forth and
- 16:35:44you can start connecting all these
- 16:35:46different tools together to make a
- 16:35:47bigger picture. There are many
- 16:35:49interesting machine learning algorithms.
- 16:35:51Let's have a look at a few of them.
- 16:35:53Hopefully this gave you a little flavor
- 16:35:54of what's out there and these are some
- 16:35:56of the most important ones that are
- 16:35:57currently being used. We'll take a look
- 16:35:59at linear regression, decision tree and
- 16:36:02the support vector machine. Let's start
- 16:36:04with a closer look at linear regression.
- 16:36:07Linear regression is perhaps one of the
- 16:36:09most well-known and well understood
- 16:36:10algorithms in statistics and machine
- 16:36:12learning. Linear regression is a linear
- 16:36:15model. For example, a model that assumes
- 16:36:17a linear relationship between the input
- 16:36:19variables x and the single output
- 16:36:22variable y. And you'll see this if you
- 16:36:24remember from your algebra classes, y =
- 16:36:27mx + c. Imagine we are predicting
- 16:36:30distance traveled y from speed x. Our
- 16:36:33linear regression model representation
- 16:36:35for this problem would be y = m * x + c
- 16:36:38or distance = m * speed + c where m is
- 16:36:43the coefficient and c is the y
- 16:36:45intercept. And we're going to look at
- 16:36:47two different variations of this. First,
- 16:36:49we're going to start with time is
- 16:36:50constant. And you can see we have a
- 16:36:52bicyclist. He's got a safety gear on,
- 16:36:54thank goodness. Speed equals 10
- 16:36:56meters/s. And so over a certain amount
- 16:36:59of time, his distance equals 36 km. We
- 16:37:02have a second bicyclist who's going
- 16:37:04twice the speed or 20 m/s. And you can
- 16:37:08guess if he's going twice the speed and
- 16:37:09time is a constant, then he's going to
- 16:37:11go twice the distance. And that's easy
- 16:37:14to compute. 36 * 2, you get 72 km. And
- 16:37:18so if you had the question of how fast
- 16:37:21would somebody going three times that
- 16:37:22speed or 30 m/s is, you can easily
- 16:37:25compute the distance in our head. We can
- 16:37:27do that without needing a computer, but
- 16:37:29we want to do this for more complicated
- 16:37:31data. So, it's kind of nice to compare
- 16:37:32the two. But, let's just take a look at
- 16:37:34that and what that looks like in a
- 16:37:35graph. So, in a linear regression model,
- 16:37:38we have our distance to the speed and we
- 16:37:40have our m equals the ve slope of the
- 16:37:44line. And we'll notice that the line has
- 16:37:45a plus slope. And as the speed
- 16:37:47increases, distance also increases.
- 16:37:49Hence, the variables have a positive
- 16:37:52relationship. And so your speed of the
- 16:37:54person which equals y = mx plus c
- 16:37:56distance traveled in a fixed interval of
- 16:37:58time. And we could very easily compute
- 16:38:00either following the line or just
- 16:38:02knowing it's 3 * 10 m/s that this is
- 16:38:05roughly 102 km distance that this third
- 16:38:07bicycle has traveled. One of the key
- 16:38:10definitions on here is positive
- 16:38:13relationship. So the slope of the line
- 16:38:16is positive. As distance increase so
- 16:38:18does speed increase. Let's take a look
- 16:38:20at our second example where we put
- 16:38:21distance is a constant. So we have speed
- 16:38:24equals 10 m/s. They have a certain
- 16:38:26distance to go and it takes him 100
- 16:38:29seconds to travel that distance. And we
- 16:38:31have our second bicyclist who's still
- 16:38:32doing 20 m/s. Since he's going twice the
- 16:38:35speed, we can guess he'll cover the
- 16:38:37distance in about half the time, 50
- 16:38:39seconds. And of course, you could
- 16:38:40probably guess on the third one, 100
- 16:38:42divided by 30 since he's going three
- 16:38:44times the speed. You can easily guess
- 16:38:46that this is 33.3333
- 16:38:49seconds time. We put that into a linear
- 16:38:51regression model or a graph. If the
- 16:38:53distance is assumed to be constant,
- 16:38:55let's see the relationship between speed
- 16:38:57and time. And as time goes up, the
- 16:38:59amount of speed to go that same distance
- 16:39:01goes down. So now your m equals a minus
- 16:39:04v slope of the line. As the speed
- 16:39:06increases, time decreases. Hence, the
- 16:39:08variable has a negative relationship.
- 16:39:11Again, there's our definition. positive
- 16:39:13relationship and negative relationship
- 16:39:15dependent on the slope of the line and
- 16:39:17with a simple formula like this um and
- 16:39:20even a significant amount of data. Let's
- 16:39:23uh see what the mathematical
- 16:39:24implementation of linear regression and
- 16:39:26we'll take this data. So suppose we have
- 16:39:28this data set where we have xyx= 1 2 3 4
- 16:39:325 standard series and the y value is 3
- 16:39:3622 43. When we take that and we go ahead
- 16:39:39and plot these points on a graph, you
- 16:39:42can see there's kind of a nice
- 16:39:43scattering and you could probably
- 16:39:44eyeball a line through the middle of it.
- 16:39:47But we're going to calculate that exact
- 16:39:48line for linear regression. And the
- 16:39:50first thing we do is we come up here and
- 16:39:52we have the mean of Xi. And remember
- 16:39:55mean is basically the average. So we
- 16:39:57added five plus 4 plus 3 plus 2 plus 1
- 16:40:00and divide by five. And that simply
- 16:40:02comes out as three. And then we'll do
- 16:40:04the same for y. We'll go ahead and add
- 16:40:06up all those numbers and divide by five.
- 16:40:09And we end up with a mean value of y of
- 16:40:11i equals 2.8 where the x i references
- 16:40:15it's an average or means value. And the
- 16:40:16yi also equals a means value of y. And
- 16:40:19when we plot that, you'll see that we
- 16:40:21can put in the y= 2.8 and the x= 3 in
- 16:40:25there on our graph. We kind of gave it a
- 16:40:27little different color so you could sort
- 16:40:28it out with the dashed lines on it. And
- 16:40:30it's important to note that when we do
- 16:40:32the linear regression, the linear
- 16:40:34regression model should go through that
- 16:40:36dot. Now, let's find our regression
- 16:40:38equation to find the best fit line.
- 16:40:40Remember, we go ahead and take our y= mx
- 16:40:42plus c. So, we're looking for m and c.
- 16:40:44So, to find this equation for our data,
- 16:40:47we need to find our slope of m and our
- 16:40:50coefficient of c. And we have y = mx + c
- 16:40:55where m equals the sum of x - x average
- 16:40:59* y - y average or y means and x means
- 16:41:02over the sum of x - x means squared.
- 16:41:06That's how we get the slope of the value
- 16:41:07of the line. And we can easily do that
- 16:41:09by creating some columns here. We have
- 16:41:11xy. Computers are really good about
- 16:41:14iterating through data. And so we can
- 16:41:16easily compute this and fill in a graph
- 16:41:18of data. And in our graph you can easily
- 16:41:21see that if we have our x value of 1 and
- 16:41:24if you remember the x i or the means
- 16:41:26value is 3. 1 - 3 equals a -2 and 2 - 3
- 16:41:32= a -1 so on and so forth. And we can
- 16:41:35easily fill in the column of x - x i y -
- 16:41:38yi. And then from those we can compute x
- 16:41:42- x i^ 2 and x - x i * y - yi. And you
- 16:41:47can guess it that the next step is to go
- 16:41:49ahead and sum the different columns for
- 16:41:50the answers we need. So we get a total
- 16:41:52of 10 for our x - x i^2 and a total of 2
- 16:41:56for x - x i * y - yi. And we plug those
- 16:42:01in, we get 2/10, which equals2. So now
- 16:42:04we know the slope of our line equals2.
- 16:42:06So we can calculate the value of c.
- 16:42:09That'd be the next step is we need to
- 16:42:10know where it crosses the y ais. And if
- 16:42:13you remember, I mentioned earlier that
- 16:42:15the linear regression line has to pass
- 16:42:18through the means value, the one that we
- 16:42:20showed earlier. We can just flip back up
- 16:42:22there to that graph. And you can see
- 16:42:24right here, there's our means value,
- 16:42:26which is 3 x= 3 and y= 2.8. And since we
- 16:42:30know that value, we can simply plug that
- 16:42:33into our formula. Y =2x + c. So we plug
- 16:42:38that in, we get 2.8 8 =2 * 3 + C. And
- 16:42:42you can just solve for C. So now we know
- 16:42:44that our coefficient equals 2.2. And
- 16:42:47once we have all that, we can go ahead
- 16:42:50and plot our regression line. Y =2 * X +
- 16:42:542.2. And then from this equation, we can
- 16:42:57compute new values. So let's predict the
- 16:43:00values of Y using X= 1 2 3 4 5 and plot
- 16:43:04the points. Remember the 1 2 3 4 5 was
- 16:43:07our original x values. So now we're
- 16:43:09going to see what y thinks they are, not
- 16:43:11what they actually are. And we plug
- 16:43:13those in, we get y of designated with y
- 16:43:16of p. You can see that x= 1 = 2.4, x= 2=
- 16:43:202.6, and so on and so on. So we have our
- 16:43:23y predicted values of what we think it's
- 16:43:26going to be when we plug those numbers
- 16:43:27in. And when we plot the predicted
- 16:43:29values along with the actual values, we
- 16:43:31can see the difference. And this is one
- 16:43:33of the things that's very important with
- 16:43:34linear regression in any of these models
- 16:43:36is to understand the error. And so we
- 16:43:38can calculate the error on all of our
- 16:43:40different values. And you can see over
- 16:43:41here we plotted um x and y and y
- 16:43:45predict. And we draw a little line so
- 16:43:46you can sort of see what the error looks
- 16:43:48like there between the different points.
- 16:43:50So our goal is to reduce this error. We
- 16:43:52want to minimize that error value on our
- 16:43:54linear regression model. Minimizing the
- 16:43:57distance. There are lots of ways to
- 16:43:59minimize the distance between the line
- 16:44:01and the data points like sum of squared
- 16:44:03errors, sum of absolute errors, root
- 16:44:06mean square error, etc. We keep moving
- 16:44:08this line through the data points to
- 16:44:10make sure the best fit line has the
- 16:44:12least squared distance between the data
- 16:44:14points and the regression line. So to
- 16:44:16recap with a very simple linear
- 16:44:18regression model, we first figure out
- 16:44:20the formula of our line through the
- 16:44:22middle and then we slowly adjust the
- 16:44:24line to minimize the error. Keep in mind
- 16:44:27this is a very simple formula. The math
- 16:44:29gets even though the math is very much
- 16:44:31the same, it gets much more complex as
- 16:44:33we add in different dimensions. So this
- 16:44:35is only two dimensions. Y equals MX + C.
- 16:44:38But you can take that out to X ZQ all
- 16:44:42the different features in there and they
- 16:44:44can plot a linear regression model on
- 16:44:46all of those using the different
- 16:44:47formulas to minimize the error. Let's go
- 16:44:50ahead and take a look at decision trees.
- 16:44:52A very different way to solve problems
- 16:44:54in the linear regression model. Decision
- 16:44:56tree is a treeshaped algorithm used to
- 16:44:58determine a course of action. Each
- 16:45:00branch of a tree represents a possible
- 16:45:02decision, occurrence, or reaction. We
- 16:45:05have data which tells us if it is a good
- 16:45:07day to play golf. And if we were to open
- 16:45:10this data up in a general spreadsheet,
- 16:45:12you can see we have the outlook, whether
- 16:45:14it's rainy, overcast, sunny,
- 16:45:17temperature, hot, mild, cool, humidity,
- 16:45:20windy, and did I like to play golf that
- 16:45:23day? Yes or no. So, we're taking a
- 16:45:25census. And certainly, I wouldn't want a
- 16:45:27computer telling me when I should go
- 16:45:29play golf or not. But you could imagine
- 16:45:30if you got up in the night before,
- 16:45:32you're trying to plan your day and it
- 16:45:34comes up and says, "Tomorrow would be a
- 16:45:36good day for golf for you in the morning
- 16:45:38and not a good day in the afternoon or
- 16:45:40something like that." This becomes very
- 16:45:41beneficial and we see this in a lot of
- 16:45:43applications coming out now where it
- 16:45:44gives you suggestions and lets you know
- 16:45:47what what would uh fit the match for you
- 16:45:49for the next day or the next purchase or
- 16:45:51the next uh whatever you know next mail
- 16:45:53out in this case is tomorrow a good day
- 16:45:56for playing golf based on the weather
- 16:45:57coming in. And so we come up and let's
- 16:46:00uh determine if you should play golf
- 16:46:02when the day is sunny and windy. So we
- 16:46:04found out the forecast tomorrow is going
- 16:46:05to be sunny and windy. And suppose we
- 16:46:08draw our tree like this. We're going to
- 16:46:10have our humidity. And then we have our
- 16:46:12normal, which is uh if it's if you have
- 16:46:14a normal humidity, you're going to go
- 16:46:16play golf. And if the humidity is really
- 16:46:18high, then we look at the outlook. And
- 16:46:20if the outlook is sunny, overcast, or
- 16:46:22rainy, it's going to change what you
- 16:46:24choose to do. So if you know that it's a
- 16:46:26very high humidity and it's sunny,
- 16:46:29you're probably not going to play golf
- 16:46:30cuz you're going to be out there
- 16:46:31miserable, fighting off the mosquitoes
- 16:46:33that are out joining you to play golf
- 16:46:35with you. Maybe if it's rainy, you
- 16:46:36probably don't want to play in the rain.
- 16:46:38But if it's slightly overcast and you
- 16:46:39get just the right shadow, that's a good
- 16:46:42day to play golf and be outside out on
- 16:46:44the green. Now, in this example, you can
- 16:46:47probably make your own tree pretty
- 16:46:49easily cuz it's a very simple set of
- 16:46:51data going in. But the question is, how
- 16:46:53do you know what to split? Where do you
- 16:46:54split your data? What if this is much
- 16:46:56more complicated data where it's not
- 16:46:58something that you would particularly
- 16:47:00understand? like studying cancer, they
- 16:47:03take about 36 measurements of the
- 16:47:05cancerous cells and then each one of
- 16:47:07those measurements represents how
- 16:47:10bulbous it is, how extended it is, how
- 16:47:12sharp the edges are, something that as a
- 16:47:14human we would have no understanding of.
- 16:47:16So how do we decide how to split that
- 16:47:18data up and is that the right decision
- 16:47:20tree? But so that's a question that's
- 16:47:21going to come up. Is this the right
- 16:47:23decision tree? For that we should
- 16:47:25calculate entropy and information gain.
- 16:47:28Two important vocabulary words there are
- 16:47:31the entropy and the information gain.
- 16:47:33Entropy. Entropy is a measure of
- 16:47:35randomness or impurity in the data set.
- 16:47:38Entropy should be low. So we want the
- 16:47:40chaos to be as low as possible. We don't
- 16:47:43want to look at it and be confused by
- 16:47:45the images or what's going on there with
- 16:47:46mixed data. And the information gain, it
- 16:47:49is a measure of decrease in entropy
- 16:47:51after the data set is split. Also known
- 16:47:54as entropy reduction. information gain
- 16:47:57should be high. So we want our
- 16:47:59information that we get out of the split
- 16:48:01to be as high as possible. Let's take a
- 16:48:03look at entropy from the mathematical
- 16:48:06side. In this case, we're going to
- 16:48:08denote entropy as I of P of and N where
- 16:48:12P is the probability that you're going
- 16:48:14to play a game of golf and N is the
- 16:48:17probability where you're not going to
- 16:48:19play the game of golf. Now, you don't
- 16:48:20really have to memorize these formulas.
- 16:48:22There's a few of them out there
- 16:48:23depending on what you're working with.
- 16:48:25But it's important to note that this is
- 16:48:26where this formula is coming from. So
- 16:48:28when you see it, you're not lost when
- 16:48:29you're running your programming, unless
- 16:48:31you're building your own decision tree
- 16:48:32code in the back. And we simply have a
- 16:48:35log 2 of p + n minus n / p + n * the log
- 16:48:40squar of n of p plus n. But let's break
- 16:48:43that down and see what actually looks
- 16:48:44like when we're computing that from the
- 16:48:46computer script side. Entropy of a
- 16:48:49target class of the data set is the
- 16:48:51whole entropy. So we have entropy play
- 16:48:53golf. And we look at this. If we go back
- 16:48:56to the data, you can simply count how
- 16:48:58many yeses and no in our complete data
- 16:49:00set for playing golf days. In our
- 16:49:03complete set, we find we have five days
- 16:49:06we did play golf and nine days we did
- 16:49:08not play golf. And so our I equals, if
- 16:49:11you add those together, 9 + 5 is 14. And
- 16:49:13so our I equals 5 over 14 and 9 over 14.
- 16:49:17That's our PNN values that we plug into
- 16:49:19that formula. And you can go 5 over
- 16:49:2214=.36.
- 16:49:249 over4=64.
- 16:49:26And when you do the whole equation, you
- 16:49:28get the -.36
- 16:49:30log<unk>^ 2 of.36 minus.64 log<unk> of
- 16:49:3664. And we get a set value. We get 94.
- 16:49:40So we now have a full entropy value for
- 16:49:42the whole set of data that we're working
- 16:49:44with. And we want to make that entropy
- 16:49:47go down. And just like we calculated the
- 16:49:49entropy out for the whole set, we can
- 16:49:51also calculate entropy for playing golf
- 16:49:54and the outlook. Is it going to be
- 16:49:55overcast or rainy or sunny? And so we
- 16:49:58look at the entropy. We have P of sunny
- 16:50:00times E of three of two. And that just
- 16:50:03comes out how many sunny days yes and
- 16:50:06how many sunny days no over the total,
- 16:50:08which is five. Don't forget to put the
- 16:50:10we'll divide that five out later on.
- 16:50:11equals P overcast = 4 comma 0 plus rainy
- 16:50:16= 2a 3 and then when you do the whole
- 16:50:18setup we have 5 over4 remember I said
- 16:50:22there was a total of five 5 over 14 *
- 16:50:25the i of 3 of 2 + 4 over 14 * the 4 0
- 16:50:30and 514 over i of 23 and so we can now
- 16:50:34compute the entropy of just the part
- 16:50:37that has to do with the forecast and we
- 16:50:39get 693 similar We can calculate the
- 16:50:42entropy of other predictors like
- 16:50:44temperature, humidity and wind. And so
- 16:50:46we look at the gain outlook. How much
- 16:50:48are we going to gain from this entropy
- 16:50:50play golf minus entropy play golf
- 16:50:52outlook? And we can take the original
- 16:50:550.94 for the whole set minus the entropy
- 16:50:58of just the rainy day and temperature
- 16:51:01and we end up with a gain of.247.
- 16:51:04So this is our information gain.
- 16:51:06Remember we define entropy and we define
- 16:51:09information gain. The higher the
- 16:51:10information gain, the lower the entropy,
- 16:51:13the better. The information gain of the
- 16:51:15other three attributes can be calculated
- 16:51:16in the same way. So we have our gain for
- 16:51:19temperature equals 0.029.
- 16:51:22We have our gain for humidity
- 16:51:23equals.152.
- 16:51:25And our gain for a windy day equals
- 16:51:270048. And if you do a quick comparison,
- 16:51:30you'll see the 247 is the greatest gain
- 16:51:34of information. So that's the split we
- 16:51:36want. Now let's build the decision tree.
- 16:51:38So, we have the outlook. Is it going to
- 16:51:40be sunny, overcast, or rainy? That's our
- 16:51:42first split because that gives us the
- 16:51:44most information gain. And we can
- 16:51:45continue to go down the tree using the
- 16:51:47different information gains with the
- 16:51:49largest information. We can continue
- 16:51:51down the nodes of the tree where we
- 16:51:53choose the attribute with the largest
- 16:51:54information gain as the root node and
- 16:51:56then continue to split each subnode with
- 16:51:59the largest information gain that we can
- 16:52:01compute. And although it's a little bit
- 16:52:02of a tongue twister to say all that, you
- 16:52:05can see that it's a very easy to view
- 16:52:07visual model. We have our outlook. We
- 16:52:09split it three different directions. If
- 16:52:11the outlook is overcast, we're going to
- 16:52:13play. And then we can split those
- 16:52:15further down if we want. So if the over
- 16:52:17outlook is sunny, but then it's also
- 16:52:19windy. If it's uh windy, we're not going
- 16:52:22to play. If it's uh not windy, we'll
- 16:52:24play. So, we can easily build a nice
- 16:52:26decision tree to guess what we would
- 16:52:28like to do tomorrow and give us a nice
- 16:52:30recommendation for the day. So, we want
- 16:52:32to know if it's a good day to play golf
- 16:52:34when it's sunny and windy. Remember the
- 16:52:35original question that came out,
- 16:52:37tomorrow's weather report is sunny and
- 16:52:38windy. You can see by going down the
- 16:52:40tree, we go outlook sunny, outlook
- 16:52:42windy. We're not going to play golf
- 16:52:44tomorrow. So, our little smartwatch pops
- 16:52:45up and says, I'm sorry, tomorrow's not a
- 16:52:48good day for golf. It's going to be
- 16:52:50sunny and windy. And if you're a huge
- 16:52:52golf fan, you might go, "Uh oh, it's not
- 16:52:55a good day to play golf." We can go in
- 16:52:57and watch a golf game at home. So, we'll
- 16:52:59sit in front of the TV instead of being
- 16:53:00out playing golf in the wind. Now that
- 16:53:02we looked at our decision tree, let's
- 16:53:04look at the third one of our algorithms
- 16:53:06we're investigating. Support vector
- 16:53:08machine. Support vector machine is a
- 16:53:10widely used classification algorithm.
- 16:53:12The idea of support vector machine is
- 16:53:14simple. The algorithm creates a
- 16:53:16separation line which divides the
- 16:53:18classes in the best possible manner. For
- 16:53:20example, dog or cat, disease or no
- 16:53:22disease. Suppose we have a labeled
- 16:53:24sample data which tells height and
- 16:53:27weight of males and females. A new data
- 16:53:30point arrives and we want to know
- 16:53:31whether it's going to be a male or a
- 16:53:33female. So we start by drawing a line.
- 16:53:36We draw decision lines. But if we
- 16:53:37consider decision line one, then we will
- 16:53:39classify the individual as a male. And
- 16:53:42if we consider decision line two, then
- 16:53:44it'll be a female. So you can see this
- 16:53:46person kind of lies in the middle of the
- 16:53:48two groups. So it's a little confusing
- 16:53:49trying to figure out which line they
- 16:53:50should be under. We need to know which
- 16:53:52line divides the classes correctly. But
- 16:53:54how the goal is to choose a hyper plane
- 16:53:57and that is one of the key words they
- 16:53:59use when we talk about support vector
- 16:54:01machines. Choose a hyper plane with the
- 16:54:04greatest possible margin between the
- 16:54:06decision line and the nearest point
- 16:54:07within the training set. So you can see
- 16:54:10here we have our support vector. We have
- 16:54:12the two nearest points to it and we draw
- 16:54:14a line between those two points. And the
- 16:54:17distance margin is the distance between
- 16:54:19the hyper plane and the nearest data
- 16:54:21point from either set. So we actually
- 16:54:23have a value and it should be equal
- 16:54:26distant between the two points that
- 16:54:28we're comparing it to. When we draw the
- 16:54:30hyperplanes, we observe that line one
- 16:54:32has a maximum distance. So we observe
- 16:54:35that line one has a maximum distance
- 16:54:37margin. So we'll classify the new data
- 16:54:39point correctly. And our result on this
- 16:54:41one is going to be that the new data
- 16:54:43point is MEL. One of the reasons we call
- 16:54:45it a hyper plane versus a line is that a
- 16:54:49lot of times we're not looking at just
- 16:54:51weight and height. We might be looking
- 16:54:53at 36 different features or dimensions.
- 16:54:56And so when we cut it with a hyper
- 16:54:58plane, it's more of a three-dimensional
- 16:55:00cut in the data, multi-dimensional that
- 16:55:03cuts the data a certain way. And each
- 16:55:05plane continues to cut it down until we
- 16:55:07get the best fit or match. Let's
- 16:55:09understand this with the help of an
- 16:55:11example. Problem statement. You always
- 16:55:12start with a problem statement when
- 16:55:14you're going to put some code together.
- 16:55:15We're going to do some coding now.
- 16:55:16Classifying muffin and cupcake recipes
- 16:55:18using support vector machines. So the
- 16:55:21cupcake versus the muffin. Let's have a
- 16:55:24look at our data set. And we have the
- 16:55:26different recipes here. We have a muffin
- 16:55:28recipe that has so much flour. I'm not
- 16:55:30sure what measurement 55 is in, but it
- 16:55:33has 55, maybe it's ounces, but it has a
- 16:55:36certain amount of flour, certain amount
- 16:55:38of milk, sugar, butter, egg, baking
- 16:55:41powder, vanilla, and salt. And so based
- 16:55:43on these measurements, we want to guess
- 16:55:45whether we're making a muffin or a
- 16:55:47cupcake. And you can see in this one, we
- 16:55:49don't have just two features. We don't
- 16:55:51just have height and weight as we did
- 16:55:53before between the male and female. In
- 16:55:55here, we have a number of features. In
- 16:55:57fact, in this, we're looking at eight
- 16:55:59different features to guess whether it's
- 16:56:01a muffin or a cupcake. What's the
- 16:56:04difference between a muffin and a
- 16:56:05cupcake? Turns out muffins have more
- 16:56:08flour, while cupcakes have more butter
- 16:56:10and sugar. So, basically, the cupcakes a
- 16:56:12little bit more of a dessert, where the
- 16:56:14muffin's a little bit more of a fancy
- 16:56:15bread. But how do we do that in Python?
- 16:56:18How do we code that to go through
- 16:56:20recipes and figure out what the recipe
- 16:56:21is? And I really just want to say
- 16:56:24cupcakes versus muffins like some big
- 16:56:28professional wrestling thing. Before we
- 16:56:30start in our cupcakes versus muffins, we
- 16:56:32are going to be working in Python.
- 16:56:34There's many versions of Python, many
- 16:56:36different editors. That is one of the
- 16:56:38strengths and weaknesses of Python is it
- 16:56:41just has so much stuff attached to it.
- 16:56:43It's one of the more popular data
- 16:56:45science programming packages you can
- 16:56:47use. In this case, we're going to go
- 16:56:49ahead and use Anaconda in Jupyter
- 16:56:52Notebook. The Anaconda Navigator has all
- 16:56:55kinds of fun tools. Once you're into the
- 16:56:58Anaconda Navigator, you can change
- 16:57:00environments. I actually have a number
- 16:57:02of environments on here. We'll be using
- 16:57:04Python 36 environment. So, this is in
- 16:57:07Python version 36. Although, it doesn't
- 16:57:09matter too much which version you use. I
- 16:57:12usually try to stay with the 3x because
- 16:57:14they're current unless you have a
- 16:57:15project that's very specifically in
- 16:57:17version 2x 27 I think is usually what
- 16:57:19most people use in the version two. And
- 16:57:22then once we're in our um Jupiter
- 16:57:24notebook editor, I can go up and create
- 16:57:26a new file and we'll just jump in here.
- 16:57:30In this case, we're doing SPM muffin
- 16:57:33versus cupcake. And then let's start
- 16:57:35with our packages for data analysis.
- 16:57:40And we almost always use a couple
- 16:57:41there's a few very standard packages we
- 16:57:43use. We use import oops import
- 16:57:50numpy
- 16:57:52that's for number python. They usually
- 16:57:54denote it as np that's very comma that's
- 16:57:57very common. And then we're going to
- 16:57:59import pandas as pd. And numpy deals
- 16:58:03with number arrays. There's a lot of
- 16:58:05cool things you can do with the numpy uh
- 16:58:07setup as far as multiplying all the
- 16:58:10values in an array in a numpy array data
- 16:58:12array. Pandas I can't remember if we're
- 16:58:15using it actually in this data set. I
- 16:58:17think we do as an import it makes a nice
- 16:58:19data frame. And the difference between a
- 16:58:21data frame and a numpy array is that a
- 16:58:24data frame is more like your Excel
- 16:58:25spreadsheet. You have columns, you have
- 16:58:28indexes. So you have different ways of
- 16:58:30referencing it easily viewing it. And
- 16:58:32there's additional features you can run
- 16:58:34on a data frame. And pandas kind of sits
- 16:58:36on numpy. So they you need them both in
- 16:58:38there. And then finally, we're working
- 16:58:41with the support vector machine. So from
- 16:58:45sklearn, we're going to use the sklearn
- 16:58:47model. Import SVM support vector
- 16:58:51machine.
- 16:58:53And then as a data scientist, you should
- 16:58:56always try to visualize your data. Some
- 16:58:59data obviously is too complicated or
- 16:59:01doesn't make any sense to the human. But
- 16:59:03if it's possible, it's good to take a
- 16:59:05second look at it so that you can
- 16:59:06actually see what you're doing. Now, for
- 16:59:08that, we're going to use two packages.
- 16:59:10We're going to import mapplot
- 16:59:12library.pipplot as plt. Again, very
- 16:59:15common. And we're going to import seabor
- 16:59:18as sns. And we'll go ahead and set the
- 16:59:21font scale in the SNS right in our
- 16:59:23import line. That's what this U
- 16:59:25semicolon followed by a line of data.
- 16:59:28We're going to set the SNS. And these
- 16:59:30are great because the the seabour sits
- 16:59:32on top of map plot library just like
- 16:59:35pandas sits on numpy. So it adds a lot
- 16:59:37more features and uses and control.
- 16:59:40We're obviously not going to get into
- 16:59:41mattplot library and seabour. It' be its
- 16:59:43own tutorial. We're really just focusing
- 16:59:45on the SVM, the support vector machine
- 16:59:48from sklearn. And since we're in Jupyter
- 16:59:52notebook, uh we have to add a special
- 16:59:54line in here for our mattplot library.
- 16:59:57And that's your percentage sign or amber
- 17:00:00sign mattplot library in line. Now, if
- 17:00:04you're doing this in just a straight
- 17:00:06code project, a lot of times I use like
- 17:00:08Notepad++
- 17:00:09and I'll run it from there. You don't
- 17:00:11have to have that line in there because
- 17:00:13it'll just pop up as its own window on
- 17:00:14your computer depending on how your
- 17:00:16computer's set up because we're running
- 17:00:18this in the Jupyter notebook as a
- 17:00:20browser setup. This tells it to display
- 17:00:23all of our graphics right below on the
- 17:00:26page. So that's what that line is for.
- 17:00:29Remember the first time I ran this, I
- 17:00:30didn't know that and I had to go look
- 17:00:31that up years ago. It's quite a
- 17:00:33headache. So mattplot library inline is
- 17:00:36just because we're running this on the
- 17:00:38web setup and we can go ahead and run
- 17:00:40this. make sure all our modules are in.
- 17:00:42They're all imported, which is great. If
- 17:00:44you don't have them import, you'll need
- 17:00:46to go ahead and pip. Use the pip or
- 17:00:48however you do it. There's a lot of
- 17:00:49other install packages out there,
- 17:00:51although pip is the most common. And you
- 17:00:53have to make sure these are all
- 17:00:54installed on your Python setup. The next
- 17:00:57step, of course, is we got to look at
- 17:00:58the data. You can't run a model for
- 17:01:01predicting data if you don't have actual
- 17:01:02data. So, to do that, let me go ahead
- 17:01:04and open this up and take a look. And we
- 17:01:07have our uh cupcakes versus muffins. and
- 17:01:10it's a CSV file or CSV meaning that it's
- 17:01:13commaepparated variable
- 17:01:15and it's going to open it up in a nice
- 17:01:17uh spreadsheet for me. And you can see
- 17:01:19up here we have the type we have muffin
- 17:01:21muffin muffin cupcake cupcake cupcake
- 17:01:24and then it's broken up into flour,
- 17:01:25milk, sugar, butter, egg, baking powder,
- 17:01:28vanilla and salt. So we can do is we can
- 17:01:31go ahead and look at this data also in
- 17:01:33our Python.
- 17:01:36Let us create a variable recipes equals
- 17:01:40we're going to use our pandas module
- 17:01:43read CSV. Remember is a commaepparated
- 17:01:46variable
- 17:01:48and the file name happened to be
- 17:01:50cupcakes versus muffins. Oops, I got
- 17:01:52double brackets there.
- 17:01:57Do it this way.
- 17:02:01There we go. cupcakes versus muffins.
- 17:02:05Because the program I loaded or the the
- 17:02:08place I saved this particular Python
- 17:02:10program is in the same folder, we can
- 17:02:12get by with just the file name. But
- 17:02:14remember, if you're storing it in a
- 17:02:15different location, you have to also put
- 17:02:16down the full path on there.
- 17:02:20And then because we're in pandas, we're
- 17:02:22going to go ahead and you can actually
- 17:02:25in line you can do this, but let me do
- 17:02:27the full print. You can just type in
- 17:02:29recipes.head head in the Jupyter
- 17:02:32notebook. But if you're running in code
- 17:02:34in a different script, you'd need to go
- 17:02:36ahead and type out the whole print
- 17:02:37recipes.
- 17:02:39And Pandanda's knows that's going to do
- 17:02:41the first five lines of data. And if we
- 17:02:44flip back on over to the spreadsheet
- 17:02:47where we opened up our CSV file,
- 17:02:50uh you can see where it starts on line
- 17:02:52two. This one calls it zero. And then 2
- 17:02:553 4 5 6 is going to match. Go and close
- 17:02:58that out because we don't need that
- 17:02:59anymore. And it always starts at zero.
- 17:03:02And these are it automatically indexes
- 17:03:04it since we didn't tell it to use an
- 17:03:06index in here. So that's the index
- 17:03:08number for the left hand side. And it
- 17:03:10automatically took the top row as
- 17:03:13labels. So pandas using it to read a CSV
- 17:03:17is just really slick and fast. One of
- 17:03:20the reasons we love our pandas, not just
- 17:03:22because they're cute and cuddly teddy
- 17:03:23bears.
- 17:03:26And let's go ahead and plot our data.
- 17:03:29And I'm not going to plot all of it. I'm
- 17:03:31just going to plot the uh sugar and
- 17:03:34flour. Now, obviously, you can see where
- 17:03:37they get really complicated if we have
- 17:03:39tons of different features. And so,
- 17:03:42you'll break them up and maybe look at
- 17:03:43just two of them at a time to see how
- 17:03:45they connect.
- 17:03:48And to plot them, we're going to go
- 17:03:49ahead and use Seabor. So, that's our
- 17:03:51SNS. And the command for that is SNS.LM
- 17:03:56plot. And then the two different
- 17:03:58variables I'm going to plot is flour and
- 17:04:00sugar.
- 17:04:02Data equals recipes. The hue equals
- 17:04:05type. And this is a lot of fun because
- 17:04:07it knows that this is pandas coming in.
- 17:04:10So this is one of the powerful things
- 17:04:12about pandas mixed with seabor and doing
- 17:04:17graphing. And then we're going to use a
- 17:04:19pallet set one. There's a lot of
- 17:04:21different sets in there. You can go look
- 17:04:23them up for seabor. We do a regular fit
- 17:04:26regular equals false. So, we're not
- 17:04:27really trying to fit anything. And it's
- 17:04:30a scatter KWS.
- 17:04:32A lot of these settings you can look up
- 17:04:33in Seabor. Half of these you could
- 17:04:35probably leave off when you run them.
- 17:04:37Somebody played with this and found out
- 17:04:38that these were the best settings for
- 17:04:40doing a Seabor plot. And let's go ahead
- 17:04:43and run that. And because it does it in
- 17:04:45line, it just puts it right on the page.
- 17:04:49And you can see right here that just
- 17:04:52based on sugar and flour alone, there's
- 17:04:55a definite split. And we use these
- 17:04:58models because you can actually look at
- 17:04:59it and say, "Hey, if I drew a line right
- 17:05:01between the middle of the blue dots and
- 17:05:03the red dots, we'd be able to do an SVM
- 17:05:07and and a hyper plane right there in the
- 17:05:09middle.
- 17:05:12Then the next step is to format or
- 17:05:16pre-process
- 17:05:20our data.
- 17:05:22And we're going to break that up into
- 17:05:23two parts.
- 17:05:26We need a type label. And remember,
- 17:05:29we're going to decide whether it's a
- 17:05:31muffin or a cupcake. Well, a computer
- 17:05:32doesn't know muffin or cupcake. It knows
- 17:05:34zero and one. So, what we're going to do
- 17:05:37is we're going to create a type label.
- 17:05:39And from this we'll create a numpy array
- 17:05:42nump where and this is where we can do
- 17:05:46some logic. We take our recipes from our
- 17:05:49panda and wherever type equals muffin
- 17:05:52it's going to be zero. And then if it
- 17:05:55doesn't equal muffin which is cupcakes
- 17:05:57it's going to be one. So we create our
- 17:05:59type label. This is the answer. So when
- 17:06:02we're doing our training model remember
- 17:06:04we have to have a a training data. This
- 17:06:06is what we're going to train it with. Is
- 17:06:07that it's zero or one? it's a muffin or
- 17:06:09it's not.
- 17:06:12And then we're going to create our
- 17:06:13recipe features.
- 17:06:16And if you remember correctly from right
- 17:06:17up here, the first column is type.
- 17:06:21So we really don't need the type column
- 17:06:22because that's our muffin or cupcake.
- 17:06:24And in pandas, we can easily sort that
- 17:06:27out.
- 17:06:29We take our value recipes
- 17:06:33columns. That's a pandas function built
- 17:06:36into pandas.
- 17:06:39values converting them to values. So
- 17:06:41it's just the column titles going across
- 17:06:43the top and we don't want the first one.
- 17:06:46So what we do is since it always starts
- 17:06:47at zero, we want one
- 17:06:51colon till the end.
- 17:06:54And then we want to go ahead and make
- 17:06:56this a list. And this converts it to a
- 17:06:59list of strings.
- 17:07:02And then we can go ahead and just take a
- 17:07:04look and see what we're looking at for
- 17:07:05the features. Make sure it looks right.
- 17:07:08Me go ahead and run that.
- 17:07:12And I forgot the S on recipes. So, we'll
- 17:07:14go ahead and add the S in there and then
- 17:07:16run that. And we can see we have flour,
- 17:07:18milk, sugar, butter, egg, baking powder,
- 17:07:22vanilla, and salt. And that matches what
- 17:07:24we have up here, right? Where we printed
- 17:07:26out everything but the type. So, we have
- 17:07:28our features and we have our label.
- 17:07:32Now, the recipe features is just the
- 17:07:34titles of the columns. We actually need
- 17:07:37the ingredients.
- 17:07:41And at this point, we have a couple
- 17:07:43options. One, we could run it over all
- 17:07:46the ingredients.
- 17:07:48And when you're doing this, usually you
- 17:07:50do. But for our example, we want to
- 17:07:51limit it so you can easily see what's
- 17:07:53going on because if we did all the
- 17:07:55ingredients, we have, you know, that's
- 17:07:57what, um, seven, eight different
- 17:08:00hyperplanes that would be built into it.
- 17:08:02We only want to look at one. So you can
- 17:08:04see what the SVM is doing.
- 17:08:06And so we'll take our recipes and we'll
- 17:08:08do just flour and sugar. Again, you can
- 17:08:12replace that with your recipe features
- 17:08:14and do all of them, but we're going to
- 17:08:15do just flour and sugar. And we're going
- 17:08:17to convert that to values. We don't need
- 17:08:19to make a list out of it because it's
- 17:08:21not string values. These are actual
- 17:08:23values on there. And we can go ahead and
- 17:08:26just print
- 17:08:28ingredients. And you can see what that
- 17:08:30looks like.
- 17:08:32Uh, and so we have just the nanoflower
- 17:08:34and sugar, just the two sets of plots.
- 17:08:38And just for fun, let's go ahead and
- 17:08:40take this over here and take our recipe
- 17:08:43features.
- 17:08:46And so if we decided to use all the
- 17:08:48recipe features, you'll see that it
- 17:08:50makes a nice column of different data.
- 17:08:52So it just strips out all the labels and
- 17:08:54everything. We just have just the
- 17:08:55values. But because we want to be able
- 17:08:57to view this easily in a plot later on,
- 17:09:01we'll go ahead and take that and just do
- 17:09:02flour and sugar.
- 17:09:05And we'll run that. And you'll see it's
- 17:09:07just the two columns.
- 17:09:10So the next step is to go ahead and fit
- 17:09:12our model.
- 17:09:14We'll go ahead and just call it model.
- 17:09:16And it's a SVM. We're using a package
- 17:09:19called SVC.
- 17:09:23In this case, we're going to go ahead
- 17:09:24and set the kernel equals linear. So,
- 17:09:27it's using a specific setup on there.
- 17:09:29And if we go to the reference on their
- 17:09:31website for the SVM,
- 17:09:34you'll see that there's about there's
- 17:09:36eight of them here. Three of them are
- 17:09:38for regression.
- 17:09:39Three are for classification. The SVC,
- 17:09:43support vector classification, is
- 17:09:45probably one of the most commonly used.
- 17:09:47And then there's also one for detecting
- 17:09:49outliers and another one that has to do
- 17:09:51with something a little bit more
- 17:09:52specific on the model. But SVC and SVR
- 17:09:55are the two most commonly used standing
- 17:09:57for support vector classifier and
- 17:10:00support vector regression. Remember
- 17:10:02regression is an actual value, a float
- 17:10:05value or whatever you're trying to work
- 17:10:07on. And SBC is a classifier. So it's a
- 17:10:10yes, no, true, false.
- 17:10:13But for this we want to know 01 muffin
- 17:10:15cupcake. If we go ahead and create our
- 17:10:17model and once we have our model
- 17:10:19created, we're going to do model.fit.
- 17:10:22And this is very common, especially in
- 17:10:23the sklearn. All their models are
- 17:10:25followed with the fit command.
- 17:10:28And what we put into the fit, what we're
- 17:10:30training with it is we're putting in the
- 17:10:32ingredients, which in this case we
- 17:10:34limited to just flour and sugar, and the
- 17:10:36type label. Is it a muffin or cupcake?
- 17:10:40Now, in more complicated data science
- 17:10:43series, you'd want to split into, we
- 17:10:46won't get into that today, where you
- 17:10:47split it into training data and test
- 17:10:50data. And they even do something where
- 17:10:52they split it into thirds, where a third
- 17:10:54is used for where you switch between
- 17:10:55which one's training and test. There's
- 17:10:57all kinds of things go into that. It
- 17:10:59gets very complicated when you get to
- 17:11:00the higher end. Not overly complicated,
- 17:11:02just an extra step, which we're not
- 17:11:04going to do today because this is a very
- 17:11:06simple set of data.
- 17:11:08And let's go ahead and run this. And now
- 17:11:09we have our model fit. And uh I got an
- 17:11:12error here. So let me fix that real
- 17:11:14quick. It's capital SBC. It turns out
- 17:11:18I did it lowercase.
- 17:11:20Support vector
- 17:11:23classifier. There we go. Let's go ahead
- 17:11:25and run that. And you'll see it comes up
- 17:11:27with all this information that it prints
- 17:11:29out automatically. These are the
- 17:11:32defaults of the model. You notice that
- 17:11:34we changed the kernel to linear. And
- 17:11:36there's our kernel linear on the
- 17:11:37printout. And there's other different
- 17:11:39settings you can mess with.
- 17:11:42We're going to just to leave that alone
- 17:11:44for right now. For this, we don't really
- 17:11:45need to mess with any of those.
- 17:11:49So, next we're going to dig a little bit
- 17:11:51into our newly trained model. And we're
- 17:11:55going to do this so we can show you on a
- 17:11:56graph.
- 17:11:58And let's go ahead and get the
- 17:12:00separating.
- 17:12:04and we're going to say uh we're going to
- 17:12:05use a W for our variable on here and
- 17:12:08we're going to do model.coreeficient_0.
- 17:12:14So what the heck is that? Again, we're
- 17:12:15digging into the model. So we've already
- 17:12:18got a prediction and a train. This is a
- 17:12:21math behind it that we're looking at
- 17:12:23right now. And so the w is going to
- 17:12:27represent two different coefficients.
- 17:12:30And if you remember, we had y = mx + c.
- 17:12:33So these coefficients are connected to
- 17:12:35that but in two-dimensional it's a
- 17:12:38plane.
- 17:12:40We don't want to spend too much time on
- 17:12:42this because you can get lost in the
- 17:12:44confusion of the math. So if you're a
- 17:12:46math wiz this is great. You can go
- 17:12:48through here and you'll see that we have
- 17:12:50a= minus w of 0 over w of 1. Remember
- 17:12:54there's two different values there. And
- 17:12:56that's basically the slope that we're
- 17:12:59generating.
- 17:13:01And then we're going to build an xx.
- 17:13:03What is xx? We're going to set it up to
- 17:13:06a numpy array. There's our np line
- 17:13:09space. So we're creating a line
- 17:13:12of values between 30 and 60. So it just
- 17:13:15creates a set of numbers for x. And then
- 17:13:18if you remember correctly, we have our
- 17:13:21formula y equals the slope * x
- 17:13:26plus the intercept. Well, to make this
- 17:13:29work, we can do this as y
- 17:13:32equals the slope times each value in
- 17:13:36that array. That's the neat thing about
- 17:13:38numpy. So, when I do a * xx, which is a
- 17:13:41whole numpy array of values, it
- 17:13:43multiplies a across all of them. And
- 17:13:45then it takes those same values and we
- 17:13:47subtract the model intercept. That's
- 17:13:50your uh we had mx plus c. So, that'd be
- 17:13:53the c from the formula y mx plus c.
- 17:13:57And that's where all these numbers come
- 17:13:59from. A little bit confusing because
- 17:14:00it's digging out of these different
- 17:14:02arrays. And then what we want to do is
- 17:14:04we're going to take this and we're going
- 17:14:06to go ahead and plot it. So plot the
- 17:14:09parallels to separating hyper plane that
- 17:14:11pass through the support vectors. And so
- 17:14:14we're going to create B equals a model
- 17:14:17support vectors. Pulling our support
- 17:14:19vectors out there. Here's our y, which
- 17:14:22we now know is a set of data. And we
- 17:14:24have uh we're going to create y down = a
- 17:14:27* xx + b1 - a * b 0. And then model
- 17:14:33support vector b is going to be set that
- 17:14:35to a new value the minus1 setup. And y y
- 17:14:38up = a * xx + b1 - a * b 0. And we can
- 17:14:45go ahead and just run this to load these
- 17:14:46variables up. If you wanted to know
- 17:14:49understand a little bit more of what's
- 17:14:50going on, you can see if we print
- 17:14:55y, let me just run that. You can see
- 17:14:58it's an array. This is a line. It's
- 17:15:00going to have in this case between 30
- 17:15:02and 60. So there's going to be 30
- 17:15:04variables in here. And the same thing
- 17:15:06with y y up y y y y y y y y y y y y y y
- 17:15:08y y y y y y y y y y y y y y y y y y y y
- 17:15:08y y y y y y y y y y y y y y y y y y y y
- 17:15:08y y y y y y y y y y y y y y y y y y y y
- 17:15:08y y y y y y down and we'll we'll plot
- 17:15:10those in just a minute on a graph so you
- 17:15:12can see what those look like.
- 17:15:14Just go ahead and delete that out of
- 17:15:15here and run that. So, it loads up the
- 17:15:18variables. Nice clean slate. I'm just
- 17:15:21going to copy this from before. Remember
- 17:15:23this? Our SNS, our Seabor plot, LM plot,
- 17:15:27flower, sugar. And I'll just go and run
- 17:15:30that real quick so you can see what
- 17:15:31remember what that looks like. It's just
- 17:15:32a straight graph on there. And then one
- 17:15:35of the neat things is because Seabour
- 17:15:37sits on top of piplot,
- 17:15:40we can do the piplot for the line going
- 17:15:42through. And that is simply plt.plot
- 17:15:47And that's our xx and y are two
- 17:15:51corresponding values xy. And then
- 17:15:53somebody played with this to figure out
- 17:15:55that the line width equals 2 and the
- 17:15:57color black would look nice. So let's go
- 17:16:00ahead and run this whole thing with the
- 17:16:01pi plot on there. And you can see when
- 17:16:04we do this, it's just doing flour and
- 17:16:06sugar on here.
- 17:16:08Corresponding line between the sugar and
- 17:16:10the flour and the muffin versus cupcake.
- 17:16:16Um, and then we generated the support
- 17:16:18vectors, the y down and y up. So let's
- 17:16:21take a look and see what that looks
- 17:16:23like.
- 17:16:24So we'll do our plot.
- 17:16:27And again, this is all against xx,
- 17:16:30our x value, but this time we have y
- 17:16:34down.
- 17:16:36And let's do something a little fun with
- 17:16:37this. We can put in a k dash dash. That
- 17:16:42just tells it to make it a dotted line.
- 17:16:46And if we're going to do the down one,
- 17:16:49we also want to do the up one. So here's
- 17:16:52our y
- 17:16:55up. And when we run that, it adds both
- 17:16:58sets of line. And so here's our support.
- 17:17:01And this is what you expect. You expect
- 17:17:03these two lines to go through the
- 17:17:04nearest data point. So the dash lines go
- 17:17:07through the nearest muffin and the
- 17:17:08nearest cupcake when it's plotting it.
- 17:17:11And then your SVM goes right down the
- 17:17:13middle. So it gives it a nice split in
- 17:17:14our data. And you can see how easy it is
- 17:17:16to see based just on sugar and flour
- 17:17:19which one's a muffin or a cupcake.
- 17:17:23Let's go ahead and create a function
- 17:17:28to predict
- 17:17:31muffin or cupcake.
- 17:17:34I've got my uh recipes. I pulled off the
- 17:17:37um internet and I want to see the
- 17:17:40difference between a
- 17:17:42muffin or a cupcake. And so we need a
- 17:17:44function to push that through. And uh we
- 17:17:46create a function with deaf. And let's
- 17:17:49call it muffin or cupcake. And remember,
- 17:17:51we're just doing flour and sugar today.
- 17:17:53We're not doing all the ingredients. And
- 17:17:55that actually is a pretty good split.
- 17:17:56You really don't need all the
- 17:17:57ingredients to know it's flour and
- 17:17:59sugar. And let's go ahead and do an if
- 17:18:02else statement. So if model predict
- 17:18:07is of flower and sugar equals zero. So
- 17:18:11we take our model and we do run a
- 17:18:13predict. It's very common in sklearn
- 17:18:14where you have a predict. You put the
- 17:18:17data in and it's going to return a
- 17:18:19value. In this case if it equals zero
- 17:18:21then print you're looking at a muffin
- 17:18:23recipe. Else if it's not zero that means
- 17:18:26it's one and you're looking at a cupcake
- 17:18:29recipe. That's pretty straightforward
- 17:18:31for
- 17:18:32function or def for definition. Deaf is
- 17:18:35how you do that in Python. And of
- 17:18:38course, if you're going to create a
- 17:18:38function, you should run something in
- 17:18:40it. And so, let's run a cupcake. And
- 17:18:42we're going to send it values 50 and 20.
- 17:18:44A muffin or a cupcake. I don't know what
- 17:18:45it is. And let's run this and just see
- 17:18:48what it gives us. It says, "Oh, it's a
- 17:18:50muffin. You're looking at a muffin
- 17:18:52recipe." So, it very easily predicts
- 17:18:54whether we're looking at a muffin or a
- 17:18:55cupcake recipe. Let's plot this. There
- 17:19:00we go. Plot this on the graph so we can
- 17:19:02see what that actually looks like. And
- 17:19:04I'm just going to copy and paste it from
- 17:19:06below where we plotting all the points
- 17:19:08in there.
- 17:19:09So, this is nothing different than we
- 17:19:11did before. If I run it, you'll see it
- 17:19:13has all the points and the lines on
- 17:19:15there. And what we want to do is we want
- 17:19:17to add another point. And we'll do
- 17:19:20pltot.
- 17:19:23And if you remember correctly, we did
- 17:19:24for our test we did 50
- 17:19:27and 20. And then somebody went in here
- 17:19:30and decided we'll do yo for yellow or
- 17:19:33it's kind of a orangeish yellow color is
- 17:19:34going to come out. Marker size nine.
- 17:19:36Those are settings you can play with.
- 17:19:38Somebody else played with them to come
- 17:19:40up with the right setup so it looks
- 17:19:41good. And you can see there it is
- 17:19:43graphed clearly a muffin.
- 17:19:46In this case in cupcakes versus muffins,
- 17:19:50the muffin has won. And if you'd like to
- 17:19:53do your own muffin cupcake contender
- 17:19:56series, you certainly can send a note
- 17:19:59down below and the team at SimplyLearn
- 17:20:01will send you over the data they use for
- 17:20:04the muffin and cupcake. And that's true
- 17:20:05of any of the data. We didn't actually
- 17:20:07run a plot on it earlier. We had men
- 17:20:09versus women. You can also request that
- 17:20:12information to run it on your data
- 17:20:14setup. So you can test that out.
- 17:20:17So to go back over our setup, we went
- 17:20:19ahead for our support vector machine
- 17:20:21code. We did a predict 40 parts flour,
- 17:20:2420 parts sugar. I think it was different
- 17:20:26than the one we did whether it's a
- 17:20:28muffin or a cupcake. Hence, we have
- 17:20:30built a classifier using SVM which is
- 17:20:33able to classify if a recipe is of a
- 17:20:36cupcake or a muffin. Which wraps up our
- 17:20:39cupcake versus muffin. So the key
- 17:20:41takeaways, what is machine learning? We
- 17:20:44discussed that with some of the
- 17:20:46different aspects of machine learning on
- 17:20:48there. We went into types of machine
- 17:20:50learning. If you memorize we have
- 17:20:52supervised, unsupervised and
- 17:20:54reinforcement learning. We discussed
- 17:20:56regression line or best fit and we did
- 17:20:59the building a decision tree and what
- 17:21:02the logic is behind that. And finally we
- 17:21:04did classification using SVM support
- 17:21:07vector machine and we did the code in
- 17:21:10there. Today we are diving into machine
- 17:21:12learning, the technology behind things
- 17:21:14like Netflix recommendations, CD, and
- 17:21:17even the face unlock of your phone.
- 17:21:19Machine learning helps devices get
- 17:21:21smarter by learning from data and
- 17:21:23predicting what we might like or need.
- 17:21:25And here's why machine learning is huge
- 17:21:27for your career. Right now, machine
- 17:21:29learning jobs are among the fastest
- 17:21:31growing roles worldwide. Companies in
- 17:21:33every industry, tech, healthcare,
- 17:21:36finance, and more, are looking for
- 17:21:38people with machine learning skills to
- 17:21:40improve their products, automate tasks,
- 17:21:42and make smarter decisions. Machine
- 17:21:45learning engineers in the US earn around
- 17:21:47$112,000 on average with plenty of room
- 17:21:50for growth as you gain experience. So,
- 17:21:52if you want to jump into this exciting
- 17:21:54field, learning machine learning can
- 17:21:56open doors to highpaying in- demand
- 17:21:58jobs. So in this video I'll guide you
- 17:22:00through the ultimate road map to master
- 17:22:02machine learning in 2025 one step at a
- 17:22:05time. So let's get started. So in the
- 17:22:08first month start with the foundations
- 17:22:10of programming. So programming is a
- 17:22:12language you'll use to communicate with
- 17:22:14your computer and bring machine learning
- 17:22:16algorithms to life. So this month is all
- 17:22:19about Python, the language of choice for
- 17:22:21most machine learning practitioners. So
- 17:22:23here's what to focus on. First, learn
- 17:22:26Python basics. Begin with Python's
- 17:22:28fundamentals like variables, data types,
- 17:22:31loops and functions. So spend time
- 17:22:34writing small programs daily to get
- 17:22:36comfortable. After that explore the key
- 17:22:38libraries like numpy, pandas and
- 17:22:41scikitlearn. So numpy is for numerical
- 17:22:43operations. It makes handling large data
- 17:22:46sets faster and easier. And pandas is to
- 17:22:49manipulate and analyze data. So pandas
- 17:22:52allow you to filter, sort and reshape
- 17:22:54data in a breeze. And then scikitlearn
- 17:22:56is for implementing algorithms in just a
- 17:22:58few lines of code. So now you might have
- 17:23:00heard about R, another language used in
- 17:23:02machine learning. But don't stress about
- 17:23:04it now. Python will serve you well,
- 17:23:06especially as a beginner, because it's
- 17:23:08simpler and more flexible. So aim to
- 17:23:10spend an hour or two each day coding. By
- 17:23:12the end of this month, you'll have a
- 17:23:14solid base to build on. Now, in the
- 17:23:16second month, get organized with version
- 17:23:18control and data structures. So this
- 17:23:20month is about learning how to organize
- 17:23:22and manage your code effectively and
- 17:23:24sharpening your problem solving skills
- 17:23:26with data structures and algorithms. So
- 17:23:28first is version control with git. So
- 17:23:31think of git as your project history
- 17:23:32tracker. So imagine working on a big
- 17:23:35project and making changes then
- 17:23:37realizing something went wrong. You want
- 17:23:39to go back to an earlier version, right?
- 17:23:41So that's where git comes in. And here's
- 17:23:44what you should practice. Number one is
- 17:23:46committing changes. So save different
- 17:23:49versions of your work as you progress.
- 17:23:51And then branching which means work on
- 17:23:53separate features without affecting your
- 17:23:55main code. And then comes merging which
- 17:23:58means combining changes from different
- 17:24:00versions once they are ready. So you
- 17:24:02have to set up an account on GitHub or
- 17:24:04GitLab to store your projects online. So
- 17:24:06not only will this be super useful, but
- 17:24:09it'll also start building your
- 17:24:10portfolio. Now next is data structures
- 17:24:13and algorithm. So think of data
- 17:24:15structures like tools in a toolkit. So
- 17:24:17each one like arrays, stacks, cues, etc.
- 17:24:20serves a specific purpose. So here's how
- 17:24:22to approach them. Number one, arrays and
- 17:24:24lists. Now arrays and lists are for
- 17:24:26storing data in sequence. After that,
- 17:24:29you can get familiar with stacks and
- 17:24:30cues. So stacks and cues are for tasks
- 17:24:33that need ordered data access. And then
- 17:24:35you have sorting and searching
- 17:24:36algorithms. So these make your programs
- 17:24:39more efficient. And that's super
- 17:24:41important in machine learning where data
- 17:24:42can get massive. So the goal here is to
- 17:24:44build up your problem solving skills
- 17:24:46which are key to machine learning
- 17:24:48success. So take it slow, practice daily
- 17:24:50and you'll see progress. Now in the
- 17:24:53third month, learn to access data with
- 17:24:55SQL. So in machine learning, a lot of
- 17:24:57work involves accessing and organizing
- 17:24:59data from databases. So SQL, a
- 17:25:02structured query language, is your
- 17:25:04ticket to getting the data you need for
- 17:25:05training ML models. So here's what you
- 17:25:08should focus on. Select and where. So
- 17:25:10these commands help you pull specific
- 17:25:12pieces of data and then you can move on
- 17:25:14to joins. Joins usually combine data
- 17:25:17from different tables. So this is so
- 17:25:19powerful that you'll use it all the
- 17:25:20time. And then comes group by and
- 17:25:22aggregate functions. They are great for
- 17:25:24summarizing data to find patterns. So
- 17:25:27spend time working with sample databases
- 17:25:29you can find online and practice writing
- 17:25:32queries. Being comfortable with SQL will
- 17:25:34save you time when preparing data for
- 17:25:36your models. Now after completing the
- 17:25:38third month you can move on to
- 17:25:40mathematics which is building your
- 17:25:42analytical mind. So this month we are
- 17:25:44tackling the math behind machine
- 17:25:45learning. So don't worry you don't need
- 17:25:47to be a math genius but understanding
- 17:25:49certain concepts will make everything
- 17:25:51feel less mysterious. So in this month
- 17:25:53you have to focus on linear algebra. So
- 17:25:56this is the math behind how models see
- 17:25:58data. So you can study vectors, matrices
- 17:26:00and operations like multiplication. Next
- 17:26:03comes calculus. So you'll use calculus
- 17:26:05to help your models learn. So you have
- 17:26:07to focus on derivatives and gradients
- 17:26:09which help minimize errors in your
- 17:26:11model. And then you can move on to
- 17:26:13probability and statistics. So
- 17:26:15understanding probability helps you make
- 17:26:17sense of data. So learn about
- 17:26:19distributions like normal distribution,
- 17:26:20bormal distribution and then variance
- 17:26:23and standard deviation. So once you have
- 17:26:25learned maths, next you'll be moving on
- 17:26:27to data handling and visualization which
- 17:26:30is the heart of machine learning as you
- 17:26:31all know. So with Matt under your belt,
- 17:26:34it's time to dig into data handling and
- 17:26:36visualization. So data preparation is
- 17:26:38vital because your model is only as good
- 17:26:40as the data you feed it. So number one
- 17:26:43comes data manipulation. So using pandas
- 17:26:45and numpy, you'll clean and organize
- 17:26:47your data. You might be removing missing
- 17:26:49values like clean up messy data so it
- 17:26:51doesn't confuse your model. And then
- 17:26:53you'll learn transforming variables like
- 17:26:55converting data into formats that work
- 17:26:57for models. And then you will move on to
- 17:26:59encoding categorical data like changing
- 17:27:02text data like female or male into
- 17:27:04numbers. Now once you're done with data
- 17:27:06manipulation, next comes data
- 17:27:07visualization. So visualization is how
- 17:27:10you get to see your data before training
- 17:27:12a model. So here you have to learn
- 17:27:13mattplot lip and seabboard. So you can
- 17:27:16create line charts, histograms, scatter
- 17:27:18plots and heat maps. So this lets you
- 17:27:21explore patterns and spot outliers. So
- 17:27:23understanding these patterns in your
- 17:27:25data is crucial for building effective
- 17:27:27models. Now in the sixth month you'll be
- 17:27:29moving on to the machine learning
- 17:27:31fundamentals. So now it's time to start
- 17:27:33building your own models. So you will
- 17:27:35focus on two main types of machine
- 17:27:37learning this month. Number one comes
- 17:27:39the supervised learning. So this is when
- 17:27:41you train a model on label data where
- 17:27:43the outcome is already known. So you'll
- 17:27:46work with algorithms like linear
- 17:27:47regression which predicts a continuous
- 17:27:49outcome. Then you'll work with decision
- 17:27:51trees which breaks down decisions into a
- 17:27:53tree structure. And then you have
- 17:27:55support vector machines under supervised
- 17:27:56learning which updates data into
- 17:27:58classes. Now after supervised learning
- 17:28:00comes unsupervised learning. So here
- 17:28:03your model identifies patterns in data
- 17:28:05without labeled outcomes. So two popular
- 17:28:08techniques in unsupervised learning is
- 17:28:09number one clustering like K means
- 17:28:11clustering which means group similar
- 17:28:13data points and then you have
- 17:28:15dimensionality reduction. This reduces
- 17:28:17data complexity by focusing on key
- 17:28:19features. So you can use scikitle learn
- 17:28:21to try out these algorithms on sample
- 17:28:23data sets. So this will give you
- 17:28:25hands-on experience with model training
- 17:28:27and you will learn to fine-tune them to
- 17:28:29get better results. Now before moving
- 17:28:31on, if you are interested in advancing
- 17:28:33your career in the field of AI and
- 17:28:35machine learning, simple learns
- 17:28:36post-graduate program delivered in
- 17:28:38collaboration with Purdue University and
- 17:28:40IBM is a perfect opportunity. This
- 17:28:43highly ranked program offers a
- 17:28:45comprehensive curriculum covering
- 17:28:47essential topics like machine learning,
- 17:28:49deep learning, NLP, computer vision,
- 17:28:52reinforcement learning, generative AI,
- 17:28:54prompt engineering, and many more. With
- 17:28:56hands-on experience to 25 plus projects
- 17:28:58and access to 20 plus cutting edge
- 17:29:00tools, you will gain the skills needed
- 17:29:02to excel in today's competitive job
- 17:29:04market. So join now and elevate your
- 17:29:07expertise with the backing of Produce
- 17:29:08academic excellence and IBM's
- 17:29:10industry-leading insight. You can find
- 17:29:12the course link in the description box
- 17:29:14and pin comments. Now moving on to the
- 17:29:16seventh month, you'll be building and
- 17:29:18training models with advanced libraries.
- 17:29:20So by now you have experimented with
- 17:29:22some basic models. So let's step it up
- 17:29:24with advanced tools like TensorFlow and
- 17:29:26PyTorch. So these libraries offer more
- 17:29:29flexibility and power. So TensorFlow and
- 17:29:31PyTorch. So here you can start with
- 17:29:33simple models and work your way up. So
- 17:29:35these libraries allow for building
- 17:29:37neural networks which you'll be studying
- 17:29:39more on the next month. Now once you
- 17:29:41have become familiar with TensorFlow and
- 17:29:43PyTorch, you can move on to model
- 17:29:44training and evaluation. So you have to
- 17:29:46learn to split data into training and
- 17:29:48testing sets and evaluate models using
- 17:29:51metrics like accuracy and precision. So
- 17:29:53your goal this month should be to get
- 17:29:55comfortable with these libraries and
- 17:29:57understand how they handle data and
- 17:29:59model training behind the scenes. So
- 17:30:01once you are done with this, you'll be
- 17:30:02moving on to the eighth month where
- 17:30:04you'll be dealing with advanced machine
- 17:30:06learning. So this month's concept will
- 17:30:08be number one on n symbol learning which
- 17:30:10means combining multiple models to get
- 17:30:12better predictions. So here you'll be
- 17:30:14learning about bagging for example
- 17:30:17random forests here multiple decision
- 17:30:19trees make predictions and then you have
- 17:30:21boosting like ada boost xg boost so
- 17:30:24models learn from each other's mistakes
- 17:30:26over here and after ensemble learning
- 17:30:28comes deep learning. So here you explore
- 17:30:31neural networks which mimic the human
- 17:30:33brain. So you'll learn about neural
- 17:30:34network basics. So you can start with
- 17:30:36simple fully connected networks and then
- 17:30:38you can move on to back propagation and
- 17:30:40gradient descent. So these helps your
- 17:30:42model learn and improve. So you can use
- 17:30:45TensorFlow or PyTorch to practice
- 17:30:47building neural networks. So you can
- 17:30:48work on projects to reinforce these
- 17:30:50concepts. Now moving on, you have two
- 17:30:52specialize on topics like NLP and
- 17:30:55computer vision. So machine learning
- 17:30:57applications are so powerful and here
- 17:30:59you'll get a taste of two major fields
- 17:31:01which is NLP or natural language
- 17:31:02processing. So here they work with text
- 17:31:05data with tasks like sentiment analysis
- 17:31:07and text classification. So you can
- 17:31:09start with basic pre-processing like
- 17:31:12tokenization, stop word removal and move
- 17:31:14to building simple NLP models. After
- 17:31:16that you can try computer vision. So for
- 17:31:18image data you have to learn CNN
- 17:31:21convolutional neural networks. So these
- 17:31:23network analyze visual patterns making
- 17:31:26them ideal for image classification. So
- 17:31:28you practice with open data sets like
- 17:31:30text, documents or images and apply the
- 17:31:32concepts you will learn to see results
- 17:31:34in real world applications. Now in the
- 17:31:3610th month you'll be dealing with model
- 17:31:38deployment which is bringing your models
- 17:31:40to life. So here you'll be using Flask
- 17:31:43or Django. So you can use these
- 17:31:44frameworks to create a web API so users
- 17:31:47can interact with your model. For
- 17:31:49example, build a web app that lets
- 17:31:50people upload images for classification.
- 17:31:53And then you can also try out Docker. So
- 17:31:55package your model and its dependencies
- 17:31:57so it can run on any machine. So this is
- 17:31:59super helpful for deploying models
- 17:32:01without compatibility issues. So by the
- 17:32:03end of this month, you'll be able to
- 17:32:04share your models with the world. So
- 17:32:06moving on to the 11th month, you'll be
- 17:32:08starting with cloud and production. So
- 17:32:11this month, you'll learn how to deploy
- 17:32:12models on the cloud and ensure they
- 17:32:14perform well in real world environments.
- 17:32:16So you'll be dealing with cloud
- 17:32:17platforms like AWS, Google Cloud or
- 17:32:20Azure. So you have to learn to deploy
- 17:32:22models of the cloud provider
- 17:32:23accessibility and scalability. And then
- 17:32:26comes monitoring and maintenance. So
- 17:32:28understand how to track your models
- 17:32:29performance over time and update it as
- 17:32:32needed. So these skills are essential
- 17:32:34for maintaining models in production and
- 17:32:36ensuring they stay reliable. And finally
- 17:32:38you will be creating real world projects
- 17:32:40and portfolio building. So here you have
- 17:32:42to choose topics that interest you and
- 17:32:45showcase your skills. So first you can
- 17:32:47start with full projects. So complete
- 17:32:49projects that go from data cleaning and
- 17:32:51model building to deployment. So ideas
- 17:32:53could be a sentiment analysis tool or an
- 17:32:55image recognition app. And then you have
- 17:32:57to build your portfolio. So organize and
- 17:32:59document your projects, host them on
- 17:33:01GitHub and create an online portfolio to
- 17:33:03share with potential employers or
- 17:33:05collaborators. So by following this road
- 17:33:07map, you'll be well prepared to handle
- 17:33:09real world machine learning challenges
- 17:33:11and have an impressive portfolio to show
- 17:33:13for it.
- 17:33:13>> Welcome to machine learning tutorial
- 17:33:16part two. My name is Richard Kersner
- 17:33:18with the SimplyLearn team. That is
- 17:33:20www.simplearn.com.
- 17:33:23Get certified, get ahead. Today in our
- 17:33:26second tutorial, we're going to cover K
- 17:33:28means linear regression along with going
- 17:33:31over the quiz questions we had during
- 17:33:33our first tutorial. What's in it for
- 17:33:36you? We're going to cover clustering.
- 17:33:38What is clustering? K means clustering
- 17:33:41which is one of the most common used
- 17:33:43clustering tools out there including a
- 17:33:45flowchart to understand K means
- 17:33:47clustering and how it functions and then
- 17:33:48we'll do an actual Python live demo on
- 17:33:51clustering of cars based on brands. Then
- 17:33:54we're going to cover logistic
- 17:33:55regression. What is logistic regression?
- 17:33:58Logistic regression curve and sigmoid
- 17:34:00function. And then we'll do another
- 17:34:02Python code demo to classify a tumor as
- 17:34:05malignant or benign based on features.
- 17:34:08And let's start with clustering. Suppose
- 17:34:10we have a pile of books of different
- 17:34:12genres. Now we divide them into
- 17:34:14different groups like fiction, horror,
- 17:34:17education, and as we can see from this
- 17:34:20young lady, she definitely is into heavy
- 17:34:22horror. You can just tell by those eyes
- 17:34:23and the maple Canadian leaf on her
- 17:34:25shirt. But we have fiction, horror, and
- 17:34:27education. And we want to go ahead and
- 17:34:29divide our books up. Well, organizing
- 17:34:31objects into groups based on similarity
- 17:34:33is clustering. And in this case, as
- 17:34:36we're looking at the books, we're
- 17:34:37talking about clustering things with
- 17:34:39known categories. But you can also use
- 17:34:41it to explore data. So you might not
- 17:34:43know the categories. You just know that
- 17:34:45you need to divide it up in some way to
- 17:34:47conquer the data and to organize it
- 17:34:49better. But in this case, we're going to
- 17:34:51be looking at clustering in specific
- 17:34:52categories. And let's just take a deeper
- 17:34:54look at that. We're going to use K means
- 17:34:57clustering. K means clustering is
- 17:34:59probably the most commonly used
- 17:35:00clustering tool in the machine learning
- 17:35:02library. K means clustering is an
- 17:35:05example of unsupervised learning. If you
- 17:35:07remember from our previous thing, it is
- 17:35:10used when you have unlabeled data. So we
- 17:35:12don't know the answer yet. We have a
- 17:35:14bunch of data that we want to cluster to
- 17:35:16different groups. Define clusters in the
- 17:35:18data based on feature similarity. So
- 17:35:21we've introduced a couple terms here.
- 17:35:23We've already talked about unsupervised
- 17:35:25learning and unlabeled data. So we don't
- 17:35:28know the answer yet. We're just going to
- 17:35:30group stuff together and see if we can
- 17:35:31find an unanswer
- 17:35:33connect. We've also introduced feature
- 17:35:36similarity. Features being different
- 17:35:38features of the data. Now, with books,
- 17:35:40we can easily see fiction and horror and
- 17:35:44history books. But a lot of times with
- 17:35:46data, some of that information isn't so
- 17:35:48easy to see right when we first look at
- 17:35:50it. And so, K means is one of those
- 17:35:51tools where we can start finding things
- 17:35:53that connect that match with each other.
- 17:35:55Suppose we have these data points and
- 17:35:57want to assign them into a cluster. Now
- 17:36:00when I look at these data points, I
- 17:36:01would probably group them into two
- 17:36:03clusters just by looking at them. I'd
- 17:36:04say two of these group of data kind of
- 17:36:06come together. But in K means we pick K
- 17:36:09clusters and assign random centrids to
- 17:36:12clusters where the K clusters represents
- 17:36:15two different clusters. We pick K
- 17:36:17clusters and say random centroidids to
- 17:36:19the clusters. Then we compute distance
- 17:36:21from objects to the centrids. Now we
- 17:36:24form new clusters based on minimum
- 17:36:26distances and calculate the centrids. So
- 17:36:29we figure out what the best distance is
- 17:36:32for the centrid. Then we move the
- 17:36:33centrid and recalculate those distances.
- 17:36:35Repeat previous two steps iteratively
- 17:36:38till the cluster centroid stop changing
- 17:36:40their positions and become static.
- 17:36:42Repeat previous two steps iteratively
- 17:36:44till the cluster centroid stop changing
- 17:36:46and the positions become static. Once
- 17:36:48the clusters become static, then K means
- 17:36:50clustering algorithm is said to be
- 17:36:52converged. And there's another term we
- 17:36:54see throughout machine learning is
- 17:36:56converged. That means whatever math
- 17:36:58we're using to figure out the answer has
- 17:37:00come to a solution or it's converged on
- 17:37:02an answer. Shall we see the flowchart to
- 17:37:04understand make a little bit more sense
- 17:37:06by putting it into a nice easy step by
- 17:37:08step? So we start, we choose K. We'll
- 17:37:11look at the elbow method in just a
- 17:37:13moment. We assign random centrids to
- 17:37:16clusters and sometimes you pick the
- 17:37:18centrids because you might look at the
- 17:37:20data in a in a graph and say ah these
- 17:37:21are probably the central points. Then we
- 17:37:24compute the distance from the objects to
- 17:37:26the centrids. We take that and we form
- 17:37:29new clusters based on minimum distance
- 17:37:31and calculate their centrids. Then we
- 17:37:33compute the distance from objects to the
- 17:37:35new centrids. And then we go back and
- 17:37:37repeat those last two steps. We
- 17:37:39calculate the distances. So as we're
- 17:37:42doing it, it brings into the new centrid
- 17:37:44and then we move the centrid around and
- 17:37:46we figure out what the best which
- 17:37:48objects are closest to each centrid. So
- 17:37:50the objects can switch from one centroid
- 17:37:52to the other as the centroidids are
- 17:37:53moved around and we continue that until
- 17:37:55it is converged. Let's see an example of
- 17:37:58this. Suppose we have this data set of
- 17:38:01seven individuals and their score on two
- 17:38:03topics A and B. Uh so here's our subject
- 17:38:07in this case referring to the person
- 17:38:09taking the uh test and then we have
- 17:38:12subject A where we see what they've
- 17:38:14scored on their first subject and we
- 17:38:15have subject B and we can see what they
- 17:38:17score on the second subject. Now let's
- 17:38:19take two farthest apart points as
- 17:38:21initial cluster centroidids. Now
- 17:38:23remember we talked about selecting them
- 17:38:25randomly or we can also just put them in
- 17:38:27different points and pick the furthest
- 17:38:28one apart so they move together. Either
- 17:38:30one works okay depending on what kind of
- 17:38:32data you're working on and what you know
- 17:38:34about it. So we took the two furthest
- 17:38:36points one and one and five and seven.
- 17:38:40And now let's take the two farthest
- 17:38:41apart points as initial cluster
- 17:38:43centrids. Each point is then assigned to
- 17:38:46the closest cluster with respect to the
- 17:38:48distance from the centrids. So we take
- 17:38:51each one of these points in there. We
- 17:38:52measure that distance. And you can see
- 17:38:53that if we measured each of those
- 17:38:55distances and you use the the
- 17:38:57Pythagorean theorem for a triangle in
- 17:38:59this case because you know the x and the
- 17:39:01y and you can figure out the diagonal
- 17:39:04line from that or you can just take a
- 17:39:05ruler and put it on your monitor. That'd
- 17:39:07be kind of silly but it would work if
- 17:39:08you're just eyeballing it. You can see
- 17:39:10how they naturally come together in
- 17:39:12certain areas. Now we again calculate
- 17:39:15the centroidids of each cluster. So
- 17:39:17cluster one and then cluster two and we
- 17:39:19look at each individual dot. There's
- 17:39:22one, two, three. We're in one cluster.
- 17:39:24Uh the centrid then moves over. It
- 17:39:26becomes 1.8 comma 2.3. So remember it
- 17:39:30was at 1 and one. Well, the very center
- 17:39:32of the data we're looking at would put
- 17:39:33it at the one point roughly 22, but 1.8
- 17:39:36and 2.3. And the second one, if we
- 17:39:38wanted to make the overall mean vector,
- 17:39:41the average vector of all the different
- 17:39:42distances to that centrid, we come up
- 17:39:44with 4, 1, and 54. So we've now moved
- 17:39:48the centrids. We compare each
- 17:39:50individual's distance to its own cluster
- 17:39:52mean and to that of the opposite cluster
- 17:39:54and we find build a nice chart on here
- 17:39:57that the as we move that centrid around
- 17:39:59we now have a new different kind of
- 17:40:01clustering of groups and using uklidian
- 17:40:03distance between the points and the mean
- 17:40:05we get the same formula you see new
- 17:40:07formulas coming up. So we have our
- 17:40:09individual dots distance to the mean
- 17:40:11centrid of the cluster and distance to
- 17:40:12the mean centrid of the cluster. Only
- 17:40:14individual three is nearer to the mean
- 17:40:16of the opposite cluster cluster two than
- 17:40:19its own cluster one. And you can see
- 17:40:21here in the diagram where we've kind of
- 17:40:23circled that one in the middle. So when
- 17:40:25we've moved the clust the centroidids of
- 17:40:27the clusters over one of the points
- 17:40:29shifted to the other cluster because
- 17:40:30it's closer to that group of
- 17:40:32individuals. Thus, individual 3 is
- 17:40:34relocated to cluster two, resulting in a
- 17:40:37new partition. And we regenerate all
- 17:40:39those numbers of how close they are to
- 17:40:41the different clusters. For the new
- 17:40:43clusters, we will find the actual
- 17:40:44cluster centroidids. So now we move the
- 17:40:47centrids over. And you can see that
- 17:40:49we've now formed two very distinct
- 17:40:50clusters on here. On comparing the
- 17:40:52distance of each individual's distance
- 17:40:54to its own cluster mean and to that of
- 17:40:56the opposite cluster, we find that the
- 17:40:58data points are stable. Hence, we have
- 17:41:00our final clusters. Now if you remember
- 17:41:03I brought up a concept earlier K mean on
- 17:41:05the K means algorithm choosing the right
- 17:41:08value of K will help in less number of
- 17:41:10iterations and to find the appropriate
- 17:41:12number of clusters in a data set we use
- 17:41:14the elbow method and within sum of
- 17:41:18squares WSS is defined as the sum of the
- 17:41:20squared distance between each member of
- 17:41:23the cluster and its centrid and so you
- 17:41:25see we've done here is we have the
- 17:41:27number of clusters and as you do the
- 17:41:30same K means algorithm over the
- 17:41:33different clusters and you calculate
- 17:41:35what that centrid looks like and you
- 17:41:36find the optimal you can actually find
- 17:41:38the optimal number of clusters using the
- 17:41:40elbow the graph is called as the elbow
- 17:41:42method and on this we guessed at two
- 17:41:44just by looking at the data but as you
- 17:41:46can see the slope you actually just look
- 17:41:48for right there where the elbow is in
- 17:41:50the slope and you have a clear answer
- 17:41:51that we want two different to start with
- 17:41:54k means equals two a lot of times people
- 17:41:56end up computing k means equals 2 3 four
- 17:41:59five until they find the value which
- 17:42:02fits on the elbow joint. Sometimes you
- 17:42:04can just look at the data and if you're
- 17:42:06really good with that specific domain
- 17:42:08remember domain I mentioned that last
- 17:42:10time you'll know that that where to pick
- 17:42:12those numbers and where to start
- 17:42:13guessing at what that k value is. So
- 17:42:15let's take this and we're going to use a
- 17:42:17use case using k means clustering to
- 17:42:20cluster cars into brands using
- 17:42:22parameters such as horsepower, cubic
- 17:42:24inches, make, year, etc. So, we're going
- 17:42:27to use the data set cars data having
- 17:42:30information about three brands of cars,
- 17:42:32Toyota, Honda, and Nissan. We'll go back
- 17:42:35to my favorite tool, the Anaconda
- 17:42:37Navigator with the Jupiter notebook. And
- 17:42:41let's go ahead and flip over to our
- 17:42:42Jupyter notebook. And in our Jupyter
- 17:42:45Notebook, I'm going to go ahead and just
- 17:42:46paste the uh basic code that we usually
- 17:42:49start a lot of these off with. We're not
- 17:42:51going to go too much into this code
- 17:42:52because we've already discussed numpy.
- 17:42:54We've already discussed mapplot library
- 17:42:56and pandas. Numpy being the number
- 17:42:58array, pandas being the pandas data
- 17:43:00frame and mattplot for the graphing. And
- 17:43:03don't forget uh since if you're using
- 17:43:04the Jupyter notebook, you do need the
- 17:43:06mattplot library in line so that it
- 17:43:08plots everything on the screen. If
- 17:43:10you're using a different Python editor,
- 17:43:13then you probably don't need that
- 17:43:14because it'll have a popup window on
- 17:43:16your computer. And we'll go ahead and
- 17:43:18run this just to load our libraries and
- 17:43:20our setup into here. The next step is of
- 17:43:23course to look at our data which I've
- 17:43:25already opened up in a spreadsheet. And
- 17:43:28you can see here we have the miles per
- 17:43:29gallon, cylinders, cubic inches,
- 17:43:32horsepower, weight pounds, how you know
- 17:43:34how heavy it is, time it takes to get to
- 17:43:3660. My card is probably on this one at
- 17:43:39about 80 or 90. What year it is? So this
- 17:43:42is you can actually see this is kind of
- 17:43:43older cars and then the brand Toyota,
- 17:43:46Honda, Nissan. So the different cars are
- 17:43:48coming from all the way from 1971 if we
- 17:43:51scroll down to uh the 80s. We have
- 17:43:54between the 70s and 80s a number of cars
- 17:43:55that they've put out. And let's uh we
- 17:43:58come back here. We're going to do
- 17:43:59importing the data. So we'll go ahead
- 17:44:01and do data set equals and we'll use
- 17:44:04pandas to read this in. And it's uh from
- 17:44:06a CSV file. Remember, you can always
- 17:44:08post this in the comments and request
- 17:44:11the data files for these either in the
- 17:44:13comments here on the YouTube video or go
- 17:44:15to simplylearn.com and request that. The
- 17:44:18car CSV, I put it in the same folder as
- 17:44:21the code that I've stored. So, my Python
- 17:44:23code is stored in the same folder, so I
- 17:44:25don't have to put the full path. If you
- 17:44:27store them in different folders, you do
- 17:44:28have to change this and double check
- 17:44:30your name variables. And we'll go ahead
- 17:44:31and run this. And uh we've chosen data
- 17:44:34set arbitrarily because, you know, it's
- 17:44:35a data set we're importing. And we've
- 17:44:37now imported our car CSV into the data
- 17:44:39set. As you know, you have to prep the
- 17:44:42data. So, we're going to create the X
- 17:44:43data. This is the one that we're going
- 17:44:45to try to figure out what's going on
- 17:44:47with. And then there is a number of ways
- 17:44:49to do this, but we'll do it in a simple
- 17:44:51loop so you can actually see what's
- 17:44:53going on. So, we'll do for i and x.c
- 17:44:57columns. So, we're going to go through
- 17:44:58each of the columns. And a lot of times
- 17:45:01it's important I I'll make lists of the
- 17:45:04columns and do this because I might
- 17:45:06remove certain columns or there might be
- 17:45:08columns that I want to be processed
- 17:45:10differently. But for this we can go
- 17:45:12ahead and take x of i and we want to go
- 17:45:16fill na and that's a pandas command. But
- 17:45:19the question is what are we going to
- 17:45:20fill the missing data with? We
- 17:45:22definitely don't want to just put in a
- 17:45:24number that doesn't actually mean
- 17:45:25something. And so one of the tricks you
- 17:45:27can do with this is we can take x of i.
- 17:45:31And in addition to that, we want to go
- 17:45:33ahead and turn this into an integer
- 17:45:35because a lot of these are integers. So
- 17:45:37we'll go ahead and keep it integers. And
- 17:45:39me add the bracket here. And a lot of
- 17:45:41editors will do this. They'll think that
- 17:45:42you're closing one bracket. Make sure
- 17:45:44you get that second bracket in there if
- 17:45:45it's a double bracket. That's always
- 17:45:47something that happens regularly. So
- 17:45:49once we have our integer of x of yi,
- 17:45:51this is going to fill in any missing
- 17:45:53data with the average. And I was so busy
- 17:45:55closing one set of brackets, I forgot
- 17:45:57that the mean is also has brackets in
- 17:45:59there for the pandas. So we can see
- 17:46:01here, we're going to fill in all the
- 17:46:02data with the average value for that
- 17:46:04column. So if there's missing data is in
- 17:46:06the average of the data it does have.
- 17:46:08Then once we've done that, we'll go
- 17:46:10ahead and loop through it again
- 17:46:12and just check and see to make sure
- 17:46:15everything is filled in correctly. And
- 17:46:17we'll print and then we take x is null.
- 17:46:21And this returns a set of the null value
- 17:46:23or the how many lines are null. And
- 17:46:25we'll just sum that up to see what that
- 17:46:27looks like. And so when I run this and
- 17:46:29so with the X, what we want to do is we
- 17:46:31want to remove the last column because
- 17:46:33that had the models. That's what we're
- 17:46:35trying to see if we can cluster these
- 17:46:36things and figure out the models. There
- 17:46:38is so many different ways to sort the X
- 17:46:41out. For one, we could take the X and we
- 17:46:44could go data set, our variable we're
- 17:46:47using, and use the eyelocation, one of
- 17:46:50the features that's in pandas, and we
- 17:46:53could take that and then take all the
- 17:46:55rows and all but the last column of the
- 17:46:58data set. And at this time, we could do
- 17:47:01values. We just convert it to values.
- 17:47:03So, that's one way to do this. And if I
- 17:47:05let me just put this down here and print
- 17:47:08X, it's a capital X we chose. and I run
- 17:47:11this, you can see it's just the values.
- 17:47:13We could also take out the values and
- 17:47:16it's not going to return anything
- 17:47:17because there's no values connected to
- 17:47:19it. What I like to do with this is
- 17:47:21instead of doing the location which does
- 17:47:24integers more common is to come in here
- 17:47:26and we have our data set and we're going
- 17:47:29to do data set dot or data set columns.
- 17:47:34And remember that lists all the columns.
- 17:47:36So if I come in here, let me just mark
- 17:47:40that as red and I print data set.c
- 17:47:45columns.
- 17:47:48You can see that I have my index here. I
- 17:47:49have my MPG cylinders everything
- 17:47:52including the brand which we don't want.
- 17:47:54So the way to get rid of the brand would
- 17:47:56be to do data columns of everything but
- 17:47:59the last one minus one. So now if I
- 17:48:01print this, you'll see the brand
- 17:48:03disappears. And so I can actually just
- 17:48:05take data set columns minus one and I'll
- 17:48:10put it right in here for the columns
- 17:48:12we're going to look at.
- 17:48:14And let's unmark this.
- 17:48:17And unmark this.
- 17:48:20And now if I do an x.ad
- 17:48:23I now have a new data frame. And you can
- 17:48:26see right here we have all the different
- 17:48:28columns except for the brand at the end
- 17:48:29of the year. And it turns out when you
- 17:48:33start playing with the data set, you're
- 17:48:35going to get an error later on and it'll
- 17:48:36say cannot convert string to float
- 17:48:40value. And that's because it for some
- 17:48:42reason these things the way they
- 17:48:43recorded them must have been recorded as
- 17:48:44strings. So we have a neat feature in
- 17:48:47here on pandas to convert. And it is
- 17:48:50simply convert objects.
- 17:48:55And for this we're going to do convert
- 17:48:57oops convert underscore
- 17:49:01numeric numeric equals true. And yes, I
- 17:49:05did have to go look that up. I don't
- 17:49:07have it memorized the convert numeric in
- 17:49:09there. If I'm working with a lot of
- 17:49:10these things, I remember them, but um
- 17:49:13depending on where I'm at, what I'm
- 17:49:14doing, I usually have to look it up. And
- 17:49:16we run that. Oops, I must have missed
- 17:49:18something in here. Let me double check
- 17:49:19my spelling. And when I double check my
- 17:49:21spilling, you'll see I missed the first
- 17:49:23underscore in the convert objects. And
- 17:49:25when I run this, it now has everything
- 17:49:27converted into a numeric value because
- 17:49:31that's what we're going to be working
- 17:49:32with is numeric values down here.
- 17:49:35And the next part is that we need to go
- 17:49:38through the data and eliminate null
- 17:49:40values. Most people when they're doing
- 17:49:42small amounts, you working with small
- 17:49:44data pools discover afterwards that they
- 17:49:46have a null value and they have to go
- 17:49:47back and do this. So, you know, be aware
- 17:49:50whenever we're formatting this data,
- 17:49:52things are going to pop up and sometimes
- 17:49:54you go backwards to fix it. And that's
- 17:49:56fine. That's just part of exploring the
- 17:49:58data and understanding what you have.
- 17:50:01And I should have done this earlier, but
- 17:50:03let me go ahead and increase the size of
- 17:50:05my window one notch.
- 17:50:09There we go. Easier to see.
- 17:50:12So, we'll do 4 I in working with X dot
- 17:50:16columns. will page through all the
- 17:50:17columns. And we want to take X of I and
- 17:50:21we're going to change that. We're going
- 17:50:22to alter it. And so with this, we want
- 17:50:25to go ahead and fill in X of I. Pandas
- 17:50:29has the fill in a. And that just fills
- 17:50:32in any non-existent missing data. And
- 17:50:36we'll put my brackets up. And there's a
- 17:50:38lot of different ways to fill this data.
- 17:50:41If you have a really large data set,
- 17:50:43some people just void out that data
- 17:50:45because if and then look at it later in
- 17:50:46a separate exploration of data. One of
- 17:50:49the tricks we can do is we can take our
- 17:50:53column and we can find the means
- 17:50:56and the means is in there or quotation
- 17:50:59marks. So we take the columns, we're
- 17:51:01going to fill in the non-existing one
- 17:51:03with the means. The problem is that
- 17:51:05returns a decimal float. So some of
- 17:51:08these aren't decimals. Certainly, you
- 17:51:11may need to be a little careful of doing
- 17:51:12this, but for this example, we're just
- 17:51:14going to fill it in with the integer
- 17:51:16version of this. Keeps it on par with
- 17:51:18the other data that isn't a decimal
- 17:51:20point.
- 17:51:23And then what we also want to do is we
- 17:51:24want to double check. A lot of times you
- 17:51:27do this first part first to double
- 17:51:29check, then you do the fill, and then
- 17:51:30you do it again just to make sure you
- 17:51:31did it right. So, we're going to go
- 17:51:33through and test for missing data. And
- 17:51:37one of the re ways you can do that is
- 17:51:40simply go in here and take our X of I
- 17:51:44column. So it's going to go through the
- 17:51:46X of I column. It says is null. So it's
- 17:51:48going to return any any place there's a
- 17:51:50null value. It actually goes through all
- 17:51:52the rows of each column is null. And
- 17:51:55then we want to go ahead and sum that.
- 17:51:57So we take that, we add the sum value.
- 17:51:59And these are all pandas. So is null is
- 17:52:01a panda command and so is sum. And if we
- 17:52:04go through that and we go ahead and run
- 17:52:05it
- 17:52:09and we go ahead and take and run that,
- 17:52:10you'll see that all the columns have
- 17:52:12zero null values. So we've now tested
- 17:52:15and double checked and our data is nice
- 17:52:16and clean. We have no null values.
- 17:52:18Everything is now a number value. We
- 17:52:20turned it into numeric and we've removed
- 17:52:23the last column in our data. And at this
- 17:52:26point, we're actually going to start
- 17:52:28using the elbow method to find the
- 17:52:30optimal number of clusters. So, we're
- 17:52:32now actually getting into the sklearn
- 17:52:34part. Uh, the K means clustering on
- 17:52:37here. I guess we'll go ahead and zoom it
- 17:52:40up one more notch so you can see what
- 17:52:41I'm typing in here.
- 17:52:44And then from sklearn going to or
- 17:52:48sklearn
- 17:52:50cluster, we're going to import K means.
- 17:52:57I always forget to capitalize the K and
- 17:52:59the M when I do this. So it's capital K,
- 17:53:01capital M K means.
- 17:53:05And we'll go and create a um array WCSS
- 17:53:09equals we'll make it an empty array. If
- 17:53:11you remember from the elbow method from
- 17:53:14our slide
- 17:53:16within the sums of squares, WSS is
- 17:53:18defined as the sum of squared distance
- 17:53:21between each member of the cluster and
- 17:53:23it centrid. So we're looking at that
- 17:53:25change in differences as far as a
- 17:53:27squared distance. And we're going to run
- 17:53:29this over a number of K mean values.
- 17:53:33In fact, let's go for I in range. We'll
- 17:53:36do 11 of them.
- 17:53:39Range zero of 11.
- 17:53:41And the first thing we're going to do is
- 17:53:42we're going to create the actual we'll
- 17:53:45do it all lowercase.
- 17:53:51And so we're going to create this object
- 17:53:54from the K means that we just imported.
- 17:53:58And the variable that we want to put
- 17:54:00into this is in clusters. We're going to
- 17:54:05set that equals to I. That's the most
- 17:54:07important one because we're looking at
- 17:54:08how increasing the number of clusters
- 17:54:11changes our answer. There are a lot of
- 17:54:14settings to the K means. Our guys in the
- 17:54:17back did a great job just kind of
- 17:54:19playing with some of them. The most
- 17:54:21common ones that you see in a lot of
- 17:54:23stuff is how you enit your K means. So
- 17:54:26we have K means plus plus. This is just
- 17:54:30a tool to let the model itself be smart
- 17:54:33how it picks it centrids to start with
- 17:54:35its initial centroidids. We only want to
- 17:54:37iterate no more than 300 times. We have
- 17:54:39a max iteration we put in there. We have
- 17:54:42the infinite the random state equals
- 17:54:44zero. You really don't need to worry too
- 17:54:46much about these when you're first
- 17:54:48learning this. As you start digging in
- 17:54:50deeper, you start finding that these are
- 17:54:51shortcuts that will speed up the process
- 17:54:55as far as a setup. But the big one that
- 17:54:57we're working with is the inclusters
- 17:54:59equals I. So, we're going to literally
- 17:55:02train our K means 11 times. We're going
- 17:55:04to do this process 11 times. And if
- 17:55:08you're working with big data, you know,
- 17:55:10the first thing you do is you run a
- 17:55:12small sample of the data so you can test
- 17:55:13all your stuff on it. And you can
- 17:55:15already see the problem that if I'm
- 17:55:17going to iterate through a terabyte of
- 17:55:19data 11 times and then the K means
- 17:55:22itself is iterating through the data
- 17:55:23multiple times. That's a heck of a
- 17:55:25process. So you got to be a little
- 17:55:27careful with this. A lot of times though
- 17:55:29you can find your elbow using the elbow
- 17:55:32method. Find your optimal number on a
- 17:55:34sample of data especially if you're
- 17:55:35working with larger data sources. So we
- 17:55:38want to go ahead and take our K means
- 17:55:39and we're just going to fit it. If
- 17:55:41you're looking at any of the sklearn,
- 17:55:43very common that you fit your model. And
- 17:55:45if you remember correctly, our variable
- 17:55:46we're using is the capital X. And once
- 17:55:49we fit this value, we go back to the um
- 17:55:53array we made. And we want to go and
- 17:55:54just append that value on the end.
- 17:55:57And it's not the actual fit we're
- 17:55:59pinning in there. It's when it generates
- 17:56:01it, it generates the value you're
- 17:56:03looking for is inertia. So k
- 17:56:05means.inertia will pull that specific
- 17:56:07value out that we need.
- 17:56:10And let's get a visual on this. We'll do
- 17:56:13our PLT plot. And what we're plotting
- 17:56:16here
- 17:56:17is first the x axis, which is range 0
- 17:56:2111. So that will generate a nice little
- 17:56:23plot there. And the wcss for our y axis.
- 17:56:30It's always nice to give our uh plot a
- 17:56:32title.
- 17:56:34And let's see, we'll just give it the
- 17:56:36elbow method for the title. And let's
- 17:56:38get some labels. So let's go ahead and
- 17:56:40do PLT X label.
- 17:56:43And what we'll do, we'll do number of
- 17:56:45clusters for that. And PLT Y label. And
- 17:56:50for that, we can do oops, there we go.
- 17:56:52WCSS since that's what we're doing on
- 17:56:54the plot on there. And finally, we want
- 17:56:56to go ahead and display our graph, which
- 17:56:58is simply plt. Oops.
- 17:57:02Show. There we go. And because we have
- 17:57:04it set to inline, it'll appear inline.
- 17:57:07Hopefully I didn't make a type error on
- 17:57:09there.
- 17:57:13And you can see we get a very nice
- 17:57:14graph. You can see a very nice elbow
- 17:57:16joint there at uh two and again right
- 17:57:19around three and four. And then after
- 17:57:21that there's not very much. Now as a
- 17:57:24data scientist, if I was looking at
- 17:57:26this, I would do either three or four.
- 17:57:29And I'd actually try both of them to see
- 17:57:31what the u output look like. And they've
- 17:57:33already tried this in the back. So,
- 17:57:35we're just going to use three as a setup
- 17:57:36on here. And let's go ahead and see what
- 17:57:38that looks like when we actually use
- 17:57:39this to show the different kinds of
- 17:57:42cars.
- 17:57:45And so, let's go ahead and apply the K
- 17:57:47means to the cars data set. And
- 17:57:50basically, we're going to copy the code
- 17:57:52that we loop through up above where K
- 17:57:54means equals K means number of clusters.
- 17:57:56And we're just going to set the number
- 17:57:57of clusters to three since that's what
- 17:58:00we're going to look for. And you could
- 17:58:01do three and four on this and graph them
- 17:58:04just to see how they come up
- 17:58:05differently. It'd be kind of curious to
- 17:58:06look at that. But for this, we're just
- 17:58:08going to set it to three. Go ahead and
- 17:58:10create our own variable Y k means for
- 17:58:13our answers. And we're going to set that
- 17:58:16equal to Whoops, my double equal there
- 17:58:19to K means. But we're not going to do a
- 17:58:22fit. We're going to do a fit predict is
- 17:58:25the setup you want to use. And when
- 17:58:27you're using untrained models, you'll
- 17:58:29see um a slightly different because
- 17:58:31usually you see fit and then you see
- 17:58:32just the predict. But we want to both
- 17:58:34fit and predict the k means on this. And
- 17:58:38that's fit underscore predict. And then
- 17:58:40our capital x is the data we're working
- 17:58:42with.
- 17:58:43And before we plot this data, we're
- 17:58:45going to do a little pandas trick. We're
- 17:58:47going to take our x value and we're
- 17:58:49going to set x as matrix. So we're
- 17:58:51converting this into a nice rows and
- 17:58:54columns kind of setup. But we want the
- 17:58:56we're going to have columns equals none.
- 17:58:57So it's just going to be a matrix of
- 17:58:59data in here. And let's go ahead and run
- 17:59:02that.
- 17:59:04A little warning. You'll see this
- 17:59:05warnings pop up because things are
- 17:59:06always being updated. So there's like
- 17:59:08minor changes in the versions and future
- 17:59:11versions. Let's set a matrix. Now that
- 17:59:13it's more common to set it values
- 17:59:16instead of doing as matrix, but mass
- 17:59:18matrix works just fine for right now and
- 17:59:20you'll want to update that later on. But
- 17:59:22let's go ahead and dive in and plot this
- 17:59:23and see what that looks like. And before
- 17:59:27we dive into plotting this data, I
- 17:59:29always like to take a look and see what
- 17:59:30I am plotting. So let's take a look at
- 17:59:33why K means. I'm just going to print
- 17:59:35that out down here. And we see we have
- 17:59:38an array of answers. We have 2 1 0 2 1
- 17:59:412. So it's clustering these different
- 17:59:45rows of data based on the three
- 17:59:47different spaces it thinks it's going to
- 17:59:49be.
- 17:59:51And then let's go ahead and print X and
- 17:59:53see what we have for X. And we'll see
- 17:59:55that X is an array. It's a matrix. So we
- 17:59:59have our different values in the array.
- 18:00:01And what we're going to do, it's very
- 18:00:03hard to plot all the different values in
- 18:00:05the array. So we're only going to be
- 18:00:07looking at the first two or positions
- 18:00:10zero and one. And if you were doing a
- 18:00:13full presentation in front of the board
- 18:00:16meeting, you might actually do a little
- 18:00:18different and and dig a little deeper
- 18:00:20into the different aspects because this
- 18:00:22is all the different columns we looked
- 18:00:23at. But we'll only look at columns one
- 18:00:25and two for this to make it easy. So
- 18:00:28let's go ahead and clear this data out
- 18:00:29of here and let's bring up our plot. And
- 18:00:32we're going to do a scatter plot here.
- 18:00:34So pl scatter.
- 18:00:37And
- 18:00:38this looks a little complicated. So
- 18:00:40let's explain what's going on with this.
- 18:00:42We're going to take the x values
- 18:00:46and we're only interested in y of k
- 18:00:48means equals 0, the first cluster. Okay?
- 18:00:52And then we're going to take value zero
- 18:00:54for the x-axis. And then we're going to
- 18:00:56do the same thing here. We're only
- 18:00:58interested in k means equals 0, but
- 18:01:01we're going to take the second column.
- 18:01:02So we're only looking at the first two
- 18:01:04columns in our answer or in the data.
- 18:01:07And then the guys in the back played
- 18:01:09with this a little bit to make it
- 18:01:10pretty.
- 18:01:12And they discovered that it looks good
- 18:01:14with a size equals 100. That's the size
- 18:01:16of the dots. We're going to use red for
- 18:01:19this one. And when they were looking at
- 18:01:22the data and what came out, it was
- 18:01:24definitely the Toyota on this. We're
- 18:01:26just going to go ahead and label it
- 18:01:27Toyota. Again, that's something you
- 18:01:29really have to explore in here as far as
- 18:01:32playing with those numbers and see what
- 18:01:34looks good. We'll go ahead and hit enter
- 18:01:35in there. And I'm just going to paste in
- 18:01:37the next two lines, which is the next
- 18:01:40two cars. And this is our Nissa and
- 18:01:43Honda. And you'll see with our scatter
- 18:01:46plot, we're now looking at where Y_K
- 18:01:48means equals 1. And we want the zero
- 18:01:51column and YK means equals 2. Again,
- 18:01:53we're looking at just the first two
- 18:01:55columns, zero and one. And each of these
- 18:01:57rows then corresponds to Nissan and
- 18:02:00Honda.
- 18:02:02And I'll go ahead and hit enter on
- 18:02:03there. And uh finally, let's take a look
- 18:02:05and put the centrids on there. Again,
- 18:02:08we're going to do a scatter plot.
- 18:02:11And on the centrids, you can just pull
- 18:02:13that from our K means, the uh model we
- 18:02:16created cluster centers. And we're going
- 18:02:19to just do um
- 18:02:22all of them in the first number and all
- 18:02:25of them in the second number, which is
- 18:02:2601 because you always start with zero
- 18:02:28and one.
- 18:02:30And then they were playing with the size
- 18:02:32and everything to make it look good.
- 18:02:34We'll do a size of 300. We're going to
- 18:02:36make the color yellow. And we'll label
- 18:02:38them. It's always good to have some good
- 18:02:39labels. Centroidids.
- 18:02:42And then we do want to do a title. PLT
- 18:02:45title.
- 18:02:47And pop up there. PLT title. So you
- 18:02:50always make want to make your graphs
- 18:02:51look pretty. And we'll call it clusters
- 18:02:52of car make. And one of the features of
- 18:02:56the plot library is you can add a
- 18:03:01legend. It'll automatically bring in it
- 18:03:03since we've already labeled the
- 18:03:05different aspects of the legend with
- 18:03:06Toyota, Nissan, and Honda.
- 18:03:09And finally, we want to go ahead and
- 18:03:10show so we can actually see it. And
- 18:03:13remember, it's in line. Uh so if you're
- 18:03:15using a different editor that's not the
- 18:03:16Jupyter notebook, you'll get a popup of
- 18:03:19this. And you should have a nice set of
- 18:03:21clusters here. So we can look at this
- 18:03:22and we have a clusters of Honda in
- 18:03:25green, Toyota in red, Nissan in purple.
- 18:03:29And you can see where they put the
- 18:03:30centroidids to separate them.
- 18:03:32Now when we're looking at this, we can
- 18:03:34also plot a lot of other different data
- 18:03:37on here as far because we only looked at
- 18:03:39the first two columns. This is just
- 18:03:40column one and two or 01 as as you label
- 18:03:44them in computer scripting. But you can
- 18:03:46see here we have a nice clusters of car
- 18:03:47making. and we were able to pull out the
- 18:03:49data and you can see how just these two
- 18:03:51columns form very distinct clusters of
- 18:03:54data. So if you were exploring new data
- 18:03:57you might take a look and say well what
- 18:03:58makes these different almost going in
- 18:04:01reverse you start looking at the data
- 18:04:03and pulling apart the columns to find
- 18:04:04out why is the first group set up the
- 18:04:07way it is. Maybe you're doing loans and
- 18:04:09you want to go, well, why is this group
- 18:04:11not defaulting on their loans and why is
- 18:04:13the last group defaulting on their
- 18:04:14loans? And why is the middle group 50%
- 18:04:16defaulting on their bank loans? And you
- 18:04:19start finding ways to manipulate the
- 18:04:21data and pull out the answers you want.
- 18:04:26So now that you've seen how to use K
- 18:04:28mean for clustering, let's move on to
- 18:04:31the next topic. Now let's look into
- 18:04:34logistic regression. The logistic
- 18:04:36regression algorithm is the simplest
- 18:04:38classification algorithm used for binary
- 18:04:41or multiclassification problems. And we
- 18:04:43can see we have our little girl from
- 18:04:45Canada who's into horror books is back.
- 18:04:47That's actually really scary when you
- 18:04:49think about that with those big eyes. In
- 18:04:51the previous tutorial, we learned about
- 18:04:52linear regression, dependent and
- 18:04:55independent variables. So to brush up,
- 18:04:58y= mx + c. Very basic algebraic function
- 18:05:03of uh y and x. The dependent variable is
- 18:05:06the target class variable we are going
- 18:05:08to predict. The independent variables X1
- 18:05:12all the way up to XN are the features or
- 18:05:15attributes we're going to use to predict
- 18:05:17the target class. We know what a linear
- 18:05:19regression looks like. But using the
- 18:05:21graph, we cannot divide the outcome into
- 18:05:23categories. It's really hard to
- 18:05:25categorize 1.5, 3.6, 9.8. Uh for
- 18:05:30example, a linear regression graph can
- 18:05:32tell us that with increase in number of
- 18:05:35hours studied, the marks of a student
- 18:05:37will increase, but it will not tell us
- 18:05:39whether the student will pass or not. In
- 18:05:41such cases where we need the output as
- 18:05:44categorical value, we will use logistic
- 18:05:46regression. And for that, we're going to
- 18:05:48use the sigmoid function. So you can see
- 18:05:50here we have our marks 0 to 100, number
- 18:05:53of hours studied. That's going to be
- 18:05:54what they're comparing it to in this
- 18:05:56example. And we usually form a line that
- 18:05:58says y = mx + c. And when we use the
- 18:06:01sigmoid function, we have p = 1 / 1 + e
- 18:06:06the minus y, it generates a sigmoid
- 18:06:09curve. And so you can see right here
- 18:06:11when you take the ln, which is the
- 18:06:13natural logarithm. I always thought it
- 18:06:16should be nl, not ln. That's just the
- 18:06:18inverse of uh e your e to the minus y.
- 18:06:22And so we do this, we get ln of p 1 - p
- 18:06:26= m * x + c. That's the sigmoid curve
- 18:06:29function we're looking for. And we can
- 18:06:31zoom in on the function and you'll see
- 18:06:33that the function as it deres goes to
- 18:06:36one or to zero depending on what your x
- 18:06:39value is. And the probability if it's
- 18:06:41greater than 0.5, the value is
- 18:06:44automatically rounded off to one
- 18:06:46indicating that the student will pass.
- 18:06:47So if they're doing a certain amount of
- 18:06:49studying, they will probably pass. Then
- 18:06:51you have a threshold value at the 0.5.
- 18:06:54It automatically puts that right in the
- 18:06:55middle usually. And your probability if
- 18:06:57it's less than 0.5, the value run it off
- 18:06:59to zero indicating the student will
- 18:07:01fail. So if they're not studying very
- 18:07:02hard, they're probably going to fail.
- 18:07:04This, of course, is ignoring the
- 18:07:06outliers of that one student who's just
- 18:07:07a natural genius and doesn't need any
- 18:07:09studying to memorize everything. That's
- 18:07:11not me, unfortunately. Have to study
- 18:07:14hard to learn new stuff. problem
- 18:07:16statement to classify whether a tumor is
- 18:07:19malignant or B9. And this is actually
- 18:07:22one of my favorite data sets to play
- 18:07:24with because it has so many features and
- 18:07:27when you look at them, you really are
- 18:07:29hard to understand. You can't just look
- 18:07:31at them and know the answer. So it gives
- 18:07:33you a chance to kind of dive into what
- 18:07:34data looks like when you aren't able to
- 18:07:36understand the specific domain of the
- 18:07:38data. But I also want you to remind you
- 18:07:40that in the domain of medicine, if I
- 18:07:42told you that my probability was really
- 18:07:45good at classified things that say 90%
- 18:07:48or 95% and I'm classifying whether
- 18:07:51you're going to have a malignant or a B9
- 18:07:54tumor, I'm guessing that you're going to
- 18:07:56go get it tested anyways. So you got to
- 18:07:57remember the domain we're working with.
- 18:07:59So why would you want to do that if you
- 18:08:01know you're just going to go get a
- 18:08:02biopsy? Because you know it's that
- 18:08:04serious. This is like an all or nothing.
- 18:08:07just referencing the domain. It's
- 18:08:08important. It might help the doctor know
- 18:08:11where to look just by understanding what
- 18:08:14kind of tumor it is. So it might help
- 18:08:17them or aid them on something they
- 18:08:18missed from before. So let's go ahead
- 18:08:20and dive into the code and I'll come
- 18:08:22back to the domain part of it in just a
- 18:08:24minute. So use case and we're going to
- 18:08:26do our normal imports here where we're
- 18:08:28importing numpy, pandas, seabour, the
- 18:08:30mattplot library and we're going to do
- 18:08:33mattplot library in line since I'm going
- 18:08:34to switch over to Anaconda. So, let's go
- 18:08:36ahead and flip over there and get this
- 18:08:38started. So, I've opened up a new window
- 18:08:40in my Anaconda Jupyter Notebook. And by
- 18:08:44the way, Jupyter Notebook, uh, you don't
- 18:08:46have to use Anaconda for the Jupyter
- 18:08:48Notebook. I just love the interface and
- 18:08:49all the tools that Anaconda brings. So,
- 18:08:52we got our import numpy aspy
- 18:08:55number array. We have our pandas pd.
- 18:08:58We're going to bring in Seabor to help
- 18:09:00us with our graphs as SNS. So many
- 18:09:03really nice tools in both Seabour and
- 18:09:05Mattplot library. And we'll do our
- 18:09:06mapplot library.pipplot as plt. And then
- 18:09:10of course we want to let it know to do
- 18:09:11it in line. And let's go and just run
- 18:09:13that. So it's all set up. And we're just
- 18:09:16going to call our data data. Not
- 18:09:18creative today. Uh equals pd. And this
- 18:09:20happens to be in a CSV file. So we'll
- 18:09:25use a pdread_csv.
- 18:09:28And I happen to name the file. renamed
- 18:09:31it data forp2.csv.
- 18:09:33You can of course um write in the
- 18:09:35comments below the YouTube and request
- 18:09:37for the data set itself or go to the
- 18:09:38SimplyLearn website and we'll be happy
- 18:09:40to supply that for you. And let's just
- 18:09:43um open up the data before we go any
- 18:09:45further and let's just see what it looks
- 18:09:46like in a spreadsheet.
- 18:09:48So when I pop it open in a local
- 18:09:50spreadsheet, this is just a CSV file,
- 18:09:53comma separated variables. We have an
- 18:09:55ID. So I guess the U categorizes for
- 18:09:58reference or what ID which test was
- 18:10:00done. The diagnosis M for malignant, B
- 18:10:04for B9. So there's two different options
- 18:10:06on there. And that's what we're going to
- 18:10:07try to predict is the M and B and test
- 18:10:09it. And then we have like the radius
- 18:10:12mean or average the texture average,
- 18:10:14perimeter mean, area mean, smoothness. I
- 18:10:17don't know about you, but unless you're
- 18:10:20a doctor in the field, most of the
- 18:10:22stuff, I mean, you can guess what
- 18:10:23concave means just by the term concave,
- 18:10:26but I really wouldn't know what that
- 18:10:28means in the measurements they're
- 18:10:29taking. So, they have all kinds of stuff
- 18:10:30like how smooth it is, uh, the symmetry,
- 18:10:33and these are all float values. You just
- 18:10:35page through them real quick, and you'll
- 18:10:37see there's, I believe, 36, if I
- 18:10:39remember correctly, in this one.
- 18:10:42So there's a lot of different values
- 18:10:43they take and all these measurements
- 18:10:45they take when they go in there and they
- 18:10:46take a look at the different growth, the
- 18:10:48tumorous growth. So back in our data and
- 18:10:52I put this in the same folder as a code.
- 18:10:54So I saved this code in that folder.
- 18:10:57Obviously if you have it in a different
- 18:10:58location, you want to put the full path
- 18:11:00in there and we'll just do uh pandas
- 18:11:05first five lines of data with the data
- 18:11:07head. And we run that. We can see that
- 18:11:10we have pretty much what we just looked
- 18:11:12at. We have an ID. We have a diagnosis.
- 18:11:15If we go all the way across, you'll see
- 18:11:17all the different columns coming across
- 18:11:19displayed nicely for our data.
- 18:11:23And while we're exploring the data, our
- 18:11:26uh Seabor, which we referenced as SNS,
- 18:11:29makes it very easy to go in here and do
- 18:11:31a joint plot. You'll notice the very
- 18:11:34similar to because it is sitting on top
- 18:11:36of the U plot library. So, the joint
- 18:11:39plot does a lot of work for us. And
- 18:11:41we're just going to look at the first
- 18:11:42two columns that we're interested in,
- 18:11:44the radius mean and the texture mean.
- 18:11:46We'll just look at those two columns and
- 18:11:49data equals data. So that tells it which
- 18:11:52two columns we're plotting and that
- 18:11:53we're going to use the data that we
- 18:11:55pulled in. Let's just run that. And it
- 18:11:58generates a really nice graph on here.
- 18:12:00And there's all kinds of cool things on
- 18:12:02this graph to look at. I mean, we have
- 18:12:03the texture mean and the radius mean
- 18:12:05obviously the axes. You can also see
- 18:12:10and uh one of the cool things on here is
- 18:12:12you can also see the histogram. They
- 18:12:13show that for the radius mean where is
- 18:12:15the most common radius mean come up and
- 18:12:17where the most common texture is. So
- 18:12:20we're looking at the tech the on each
- 18:12:22growth it's average texture and on each
- 18:12:25radius it's average uh radius on there
- 18:12:28gets a little confusing because we're
- 18:12:29talking about the individual objects
- 18:12:31average. And then we can also look over
- 18:12:33here and see the the histogram showing
- 18:12:36us the median or how common each
- 18:12:39measurement is. And that's only two
- 18:12:42columns. So let's dig a little deeper
- 18:12:44into Seabor. They also have a heat map.
- 18:12:47And if you're not familiar with heat
- 18:12:49maps, a heat map just means it's in
- 18:12:51color. That's all that means. Heat map.
- 18:12:53I guess the original ones were plotting
- 18:12:55heat density on something. And so ever
- 18:12:57since then it's just called a heat map.
- 18:12:58And we're going to take our data and get
- 18:13:00our corresponding numbers to put that
- 18:13:02into the heat map. And that's simply
- 18:13:04data.coR
- 18:13:06for that. That's a pandas expression.
- 18:13:09Let's remember we're working in a pandas
- 18:13:11data frame. So that's one of the cool
- 18:13:12tools in pandas for our data. And let's
- 18:13:15just pull that information into a heat
- 18:13:17map and see what that looks like. And
- 18:13:19you'll see that we're now looking at all
- 18:13:21the different features. We have our ID.
- 18:13:24We have our texture. We have our area,
- 18:13:26our compactness, concave points. And if
- 18:13:29you look down the middle of this chart
- 18:13:31diagonal going from the upper left to
- 18:13:32bottom right, it's all white. That's
- 18:13:35because when you compare texture to
- 18:13:38texture, they're identical. So they're
- 18:13:40100% or in this case perfect one in
- 18:13:43their correspondence.
- 18:13:45And you'll see that when you look at say
- 18:13:48area or right below it, it has almost a
- 18:13:51black on there. when you compare it to
- 18:13:53texture. So these have almost no
- 18:13:54corresponding data. They don't really
- 18:13:56form a linear graph or something that
- 18:13:58you can look at and say how connected
- 18:13:59they are. They're very scattered data.
- 18:14:02This is really just a really nice graph
- 18:14:04to get a quick look at your data.
- 18:14:06Doesn't so much change what you do, but
- 18:14:08it changes verifying. So when you get an
- 18:14:11answer or something like that or you
- 18:14:12start looking at some of these
- 18:14:13individual pieces, you might go, "Hey,
- 18:14:15that doesn't match. according to showing
- 18:14:18our heat map, this should not correlate
- 18:14:21with each other. And if it is, you're
- 18:14:22going to have to start asking, well,
- 18:14:23why? What's going on? What else is
- 18:14:25coming in there? But it does show some
- 18:14:27really cool information on here. I mean,
- 18:14:30we can see from the ID, there's no real
- 18:14:33one feature that just says if you go
- 18:14:36across the top line that lights up.
- 18:14:39There's no one feature that says, hey,
- 18:14:40if the area is a certain size, then it's
- 18:14:42going to be B9 or malignant. It says
- 18:14:44there's some that sort of add up and
- 18:14:46that's a big hint in the data that we're
- 18:14:49trying to ID this whether it's malignant
- 18:14:51or B9. That's a big hint to us as data
- 18:14:54scientists to go okay we can't solve
- 18:14:56this with any one feature. It's going to
- 18:14:59be something that includes all the
- 18:15:00features or many of the different
- 18:15:02features to come up with a solution for
- 18:15:03it. And while we're exploring the data
- 18:15:07let's explore one more area and let's
- 18:15:09look at data isnull. We want to check
- 18:15:12for null values in our data. If you
- 18:15:15remember from earlier in this tutorial,
- 18:15:18we did it a little differently where we
- 18:15:19added stuff up and sum them up. You can
- 18:15:22actually with pandas do it really
- 18:15:23quickly. Data.isnull and summit. And
- 18:15:25it's going to go across all the columns.
- 18:15:27So when I run this,
- 18:15:30you're going to see all the columns come
- 18:15:32up with no null data.
- 18:15:36So we've just just to rehash these last
- 18:15:39few steps. We've done a lot of
- 18:15:42exploration. We have looked at the first
- 18:15:44two columns and seen how they plot with
- 18:15:47the seabour with a joint plot which
- 18:15:49shows both the histogram and the data
- 18:15:52plotted on the XY coordinates. And
- 18:15:54obviously you can do that more in detail
- 18:15:58with different columns and see how they
- 18:15:59plot together. And then we took and did
- 18:16:02the Seabor heat map the SNS
- 18:16:05heat mapap of the data. And you can see
- 18:16:07right here where it did a nice job
- 18:16:09showing us some bright spots where stuff
- 18:16:11correlates with each other and forms a
- 18:16:13very nice combination or points of
- 18:16:16scattering points. And you can also see
- 18:16:18areas that don't.
- 18:16:20And then finally, we went ahead and
- 18:16:22checked the data. Is the data null
- 18:16:24value? Do we have any missing data in
- 18:16:26there? Very important step because it'll
- 18:16:28crash later on. If you forget to do this
- 18:16:31step, it will remind you when you get
- 18:16:33that nice error code that says null
- 18:16:35values. Okay. So, not a big deal if you
- 18:16:38miss it, but it it's no fun having to go
- 18:16:40back when you're when you're in a huge
- 18:16:42process and you've missed this step and
- 18:16:44now you're 10 steps later and you got to
- 18:16:45go remember where you were pulling the
- 18:16:47data in.
- 18:16:49So, we need to go ahead and pull out our
- 18:16:51X and our Y. So, we just put that down
- 18:16:54here and we'll set the X equal to. And
- 18:16:57there's a lot of different options here.
- 18:16:59Certainly we could do X equals all the
- 18:17:01columns except for the first two because
- 18:17:04if you remember the first two is the ID
- 18:17:05and the diagnosis. So that certainly
- 18:17:08would be an option. But what we're going
- 18:17:10to do is we're actually going to focus
- 18:17:12on the worst. The worst radius, the
- 18:17:14worst texture, parameter area,
- 18:17:17smoothness, compactness, and so on. One
- 18:17:20of the reasons to start dividing your
- 18:17:22data up when you're looking at this
- 18:17:24information is sometimes the data will
- 18:17:28be the same data coming in. So if I have
- 18:17:30two measurements coming into my model,
- 18:17:33it might overweigh them. It might
- 18:17:35overpower the other measurements because
- 18:17:37it's measuring it's basically taking
- 18:17:38that information in twice. That's a
- 18:17:40little bit past the scope of this
- 18:17:41tutorial. I want you to take away from
- 18:17:43this though is that we are dividing the
- 18:17:45data up into pieces and our team in the
- 18:17:47back went ahead and said hey let's just
- 18:17:49look at the worst. So I'm going to
- 18:17:51create a an array and you'll see this
- 18:17:54array radius worst texture worst
- 18:17:56perimeter worst. We've just taken the
- 18:17:58worst of the worst and I'm just going to
- 18:18:00put that in my X. So this X is still a
- 18:18:02pandas data frame but it's just those
- 18:18:05columns. And our Y, if you remember
- 18:18:08correctly, is going to be Oops, hold on
- 18:18:10one second. It's not X. is data. There
- 18:18:13we go. So, x equals data and then it's a
- 18:18:16list of the different columns, the worst
- 18:18:17of the worst. And if we're going to take
- 18:18:19that, then we have to have our answer
- 18:18:21for our y for the stuff we know. And if
- 18:18:24you remember correctly, we're just going
- 18:18:25to be looking at
- 18:18:28the diagnosis. That's all we care about
- 18:18:30is what is it diagnosed? Is it B9 or
- 18:18:32malignant? And since it's a single
- 18:18:35column, we can just do diagnosis. Oh, I
- 18:18:37forgot to put the brackets. There we go.
- 18:18:39Okay. So, it's just diagnosis on there.
- 18:18:42And we can also real quickly do like an
- 18:18:44X do. If you want to see what that looks
- 18:18:46like and Y head
- 18:18:50and run this and you'll see um it only
- 18:18:53does the last one. I forgot about that.
- 18:18:55If you don't do print, you can see that
- 18:18:57the the Y.D is just mm because the first
- 18:19:00ones are all malignant. And if I run
- 18:19:02this, the X do head is just the first
- 18:19:04five values of radius worst, texture
- 18:19:07worst, parameter worst, area worst, and
- 18:19:09so on. I'll go ahead and take that out.
- 18:19:13So, moving down to the next step, we've
- 18:19:17built our two data sets, our answer and
- 18:19:20then the features we want to look at.
- 18:19:23In data science, it's very important to
- 18:19:26test your model. So we do that by
- 18:19:29splitting the data
- 18:19:32and from sklearn model selection we're
- 18:19:34going to import train test split. So
- 18:19:37we're going to split it into two groups.
- 18:19:39There are so many ways to do this. I
- 18:19:41noticed in one of the more modern ways
- 18:19:43they actually split it into three groups
- 18:19:45and then you model each group and test
- 18:19:48it against the other groups. So you have
- 18:19:50all kinds and there's reasons for that
- 18:19:51which is past the scope of this and for
- 18:19:53this particular example isn't necessary
- 18:19:56for this. We're just going to split it
- 18:19:57into two groups. one to train our data
- 18:19:59and one to test our data. And the
- 18:20:02sklearn uh.mmodel selection we have
- 18:20:05train tests split. You could write your
- 18:20:07own quick code to do this where you just
- 18:20:09randomly divide the data up into two
- 18:20:11groups but they do it for us nicely
- 18:20:14and we actually can almost we can
- 18:20:16actually do it in one statement with
- 18:20:17this where we're going to generate four
- 18:20:19variables capital X train capital X
- 18:20:23test. So we have our training data we're
- 18:20:25going to use to fit the model and then
- 18:20:27we need something to test it and then we
- 18:20:29have our y train. So we're going to
- 18:20:30train the answer and then we have our
- 18:20:32test. So this is the stuff we want to
- 18:20:33see how good it did on our model. And
- 18:20:36we'll go ahead and take our train test
- 18:20:38split that we just imported.
- 18:20:41And we're going to do X and our Y, our
- 18:20:43two different data that's going in for
- 18:20:45our split. And then the guys in the back
- 18:20:48came up and wanted us to go ahead and
- 18:20:49use a test size equals.3.
- 18:20:52That's test size. Random state. It's
- 18:20:55always nice to kind of switch a random
- 18:20:57state around, but not that important.
- 18:20:59What this means is that the test size is
- 18:21:01we're going to take 30% of the data and
- 18:21:04we're going to put that into our test
- 18:21:06variables, our Y test and our X test.
- 18:21:09And we're going to do 70% into the X
- 18:21:11train and the Y train. So, we're going
- 18:21:13to use 70% of the data to train our
- 18:21:15model and 30% to test it. Let's go ahead
- 18:21:18and run that and load those up. So now
- 18:21:21we have all our stuff split up and all
- 18:21:23our data ready to go. And now we get to
- 18:21:25the actual logistics part. We're
- 18:21:27actually going to do our create our
- 18:21:28model. So let's go ahead and bring that
- 18:21:30in from sklearn. We're going to bring in
- 18:21:33our linear model and we're going to
- 18:21:34import logistic regression. That's the
- 18:21:37actual model we're using. And let's
- 18:21:39we'll call it log model.
- 18:21:42Oops, there we go. Model. And let's just
- 18:21:44set this equal to our logistic
- 18:21:46regression that we just imported. So now
- 18:21:49we have a variable log model set to that
- 18:21:51class for us to use. And with most the
- 18:21:55uh models in the sklearn, we just need
- 18:21:58to go ahead and fix it. Fit do a fit on
- 18:22:01there. And we use our x train that we
- 18:22:04separated out with our y train. And
- 18:22:07let's go ahead and run this. So once
- 18:22:08we've run this, we'll have a model that
- 18:22:10fits this data that 70% of our training
- 18:22:13data.
- 18:22:15Uh, and of course it prints this out
- 18:22:17that tells us all the different
- 18:22:18variables that you can set on there.
- 18:22:20There's a lot of different choices you
- 18:22:21can make, but for Word do, we're just
- 18:22:23going to let all the defaults set. We
- 18:22:25don't really need to mess with those on
- 18:22:26this particular example. And there's
- 18:22:28nothing in here that really stands out
- 18:22:29as super important until you start
- 18:22:32fine-tuning it. But for what we're
- 18:22:34doing, the basics will work just fine.
- 18:22:36And then let's we need to go ahead and
- 18:22:38test out our model. Is it working? So
- 18:22:41let's create a variable Y predict. And
- 18:22:43this is going to be equal to our log
- 18:22:46model. And we want to do a predict.
- 18:22:49Again, very standard format for the
- 18:22:52sklearn library is taking your model and
- 18:22:54doing a predict on it. And we're going
- 18:22:56to test y predict against the y test. So
- 18:22:59we want to know what the model thinks
- 18:23:01it's going to be. That's what our y
- 18:23:02predict is. And with that, we want the
- 18:23:05capital xx test. So we have our train
- 18:23:08set and our test set. And now we're
- 18:23:10going to do our y predict. And let's go
- 18:23:12ahead and run that.
- 18:23:15And if we uh print
- 18:23:18y predict, let me go ahead and run that.
- 18:23:22You'll see it comes up and it predents a
- 18:23:25prints a nice array of uh B and M for B9
- 18:23:28and malignant
- 18:23:30for all the different test data we put
- 18:23:32in there. So, it does pretty good. We're
- 18:23:34not sure exactly how good it does, but
- 18:23:36we can see that it actually works and is
- 18:23:38functional. Was very easy to create.
- 18:23:40You'll always discover with our data
- 18:23:42science that as you explore this, you
- 18:23:45spend a significant amount of time
- 18:23:47prepping your data and making sure your
- 18:23:50data coming in is good. Uh there's a
- 18:23:52saying, good data in, good answers out.
- 18:23:56Bad data in, bad answers out. That's
- 18:23:59only half the thing. That's only half of
- 18:24:02it. Selecting your models becomes the
- 18:24:04next part as far as how good your models
- 18:24:06are. and then of course fine-tuning it
- 18:24:08depending on what model you're using. So
- 18:24:11we come in here, we want to know how
- 18:24:12good this came out. So we have our Y
- 18:24:14predict here, log model.predict X test.
- 18:24:20So for deciding how good our model is,
- 18:24:23we're going to go from the
- 18:24:24sklearn.metrics,
- 18:24:26we're going to import classification
- 18:24:28report. And that just reports how good
- 18:24:30our model is doing. And then we're going
- 18:24:31to feed it the model data. And let's
- 18:24:33just print this out. and we'll take our
- 18:24:36uh classification report
- 18:24:39and we're going to put into there
- 18:24:43our test our actual data. So this is
- 18:24:46what we actually know is true and our
- 18:24:49prediction what our model predicted for
- 18:24:51that data on the test side. And let's
- 18:24:54run that and see what that does.
- 18:24:57So we pull that up. You'll see that we
- 18:24:59have um a precision for B9 and malignant
- 18:25:03B and M. And we have a precision of 93
- 18:25:06and 91, a total of 92. So it's kind of
- 18:25:10the average between these two of 92.
- 18:25:12There's all kinds of different
- 18:25:13information on here. Your F1 score,
- 18:25:16your recall, your support coming through
- 18:25:19on this. And for this, I'll go ahead and
- 18:25:22just flip back to our slides that they
- 18:25:23put together for describing it. And so
- 18:25:26here we're going to look at the
- 18:25:26precision using the classification
- 18:25:28report. And you see this is the same
- 18:25:30print out I had up above. Some of the
- 18:25:32numbers might be different because it
- 18:25:34does randomly pick out which data we're
- 18:25:36using. So this model is able to predict
- 18:25:39the type of tumor with 91% accuracy. So
- 18:25:43we look back here that's you will see
- 18:25:45where we have uh B9 and malignant. It
- 18:25:47actually has 92 coming up here. We're
- 18:25:49looking about a 92 91% precision. And
- 18:25:52remember I reminded you about domain.
- 18:25:54So, when we're talking about the domain
- 18:25:55of a medical domain with a very
- 18:25:58catastrophic outcome, you know, at 91 or
- 18:26:0092% precision, you're still going to go
- 18:26:03in there and have somebody do a biopsy
- 18:26:06on it. Very different than if you're
- 18:26:08investing money and there's a 92% chance
- 18:26:10you're going to earn 10% and 8% chance
- 18:26:14you're going to lose 8%, you're probably
- 18:26:16going to bet the money because at that
- 18:26:17odds, it's pretty good that you'll make
- 18:26:19some money. And in the long run, you do
- 18:26:20that enough, you definitely will make
- 18:26:22money. And also with this domain, I've
- 18:26:24actually seen them use this to identify
- 18:26:26different forms of cancer. That's one of
- 18:26:29the things that they're starting to use
- 18:26:30these models for because then it helps a
- 18:26:32doctor know what to investigate. So that
- 18:26:34wraps up this section. We're finally
- 18:26:37we're going to go in there and let's
- 18:26:38discuss the answers to the quiz asked in
- 18:26:40machine learning tutorial part one. Can
- 18:26:43you tell what's happening in the
- 18:26:44following cases? Grouping documents into
- 18:26:47different categories based on the topic
- 18:26:50and content of each document. This is an
- 18:26:52example of clustering where K means
- 18:26:54clustering can be used to group the
- 18:26:55documents by topics using bag of words
- 18:26:58approach. So if you gotten in there that
- 18:27:00you're looking for clustering and
- 18:27:02hopefully you had at least one or two
- 18:27:04examples like K means that are used for
- 18:27:06clustering different things then give
- 18:27:08yourself a two thumbs up. B identifying
- 18:27:11handwritten digits in images correctly.
- 18:27:14This is an example of classification.
- 18:27:17The traditional approach to solving this
- 18:27:18would be to extract digit dependent
- 18:27:20features like curvature of different
- 18:27:22digits etc. and then use a classifier
- 18:27:24like SVM to distinguish between images.
- 18:27:27Again, if you got the fact that it's a
- 18:27:29classification example, give yourself a
- 18:27:31thumb up. And if you're able to go, hey,
- 18:27:33let's use SVM or another model for this,
- 18:27:36give yourself those two thumbs up on it.
- 18:27:38C. Behavior of a website indicating that
- 18:27:41the site is not working as designed.
- 18:27:44This is an example of anomaly detection.
- 18:27:47In this case, the algorithm learns what
- 18:27:49is normal and what is not normal,
- 18:27:51usually by observing the logs of the
- 18:27:53website. Give yourself a thumbs up if
- 18:27:55you got that one. And just for a bonus,
- 18:27:57can you think of another example of
- 18:27:59anomaly detection? One of the ones I use
- 18:28:01it for in my own business is detecting
- 18:28:03anomalies in stock markets. Stock
- 18:28:06markets are very fickled and they behave
- 18:28:08very erratic. So finding those erratic
- 18:28:10areas and then finding ways to track
- 18:28:12down why they're erratic. Was something
- 18:28:14released in social media? Was something
- 18:28:16released you can see where knowing where
- 18:28:18that anomaly is can help you to figure
- 18:28:21out what the answer is to it in another
- 18:28:23area. D predicting salary of an
- 18:28:25individual based on his or her years of
- 18:28:28experience. This is an example of
- 18:28:30regression. This problem can be
- 18:28:32mathematically defined as a function
- 18:28:33between independent years of experience
- 18:28:35and dependent variables salary of an
- 18:28:38individual. And if you guess that this
- 18:28:40was a regression model, give yourself a
- 18:28:42thumbs up. And if you were able to
- 18:28:43remember that it was between independent
- 18:28:46and dependent variables and that terms,
- 18:28:49give yourself two thumbs up. Summary. So
- 18:28:52to wrap it up, we went over what is K
- 18:28:55means and we went through also the chart
- 18:28:58of choosing your elbow method and
- 18:29:00assigning a random centrid to the
- 18:29:02clusters, computing the distance and
- 18:29:04then going in there and figuring out
- 18:29:06what the minimum centroidids is and
- 18:29:08computing the distance and going through
- 18:29:09that loop until it gets the perfect
- 18:29:11centrid. And we looked into the elbow
- 18:29:13method to choose K based on running our
- 18:29:16clusters across a number of variables
- 18:29:17and finding the best location for that.
- 18:29:19We did a nice example of clustering cars
- 18:29:22with K means even though we only looked
- 18:29:23at the first two columns to make it
- 18:29:25simple and easy to graph. You can easily
- 18:29:27extrapolate that and look at all the
- 18:29:29different columns and see how they all
- 18:29:31fit together. And we looked at what is
- 18:29:33logistic regression. We discussed the
- 18:29:35sigmoid function. What is logistic
- 18:29:38regression? And then we went into an
- 18:29:40example of classifying tumors with
- 18:29:42logistics. I hope you enjoyed part two
- 18:29:45of machine learning. So in today's
- 18:29:47session we will discuss what RNN model
- 18:29:49is. Moving ahead we will see why should
- 18:29:52we use RNN. After that we will see how
- 18:29:55does RNN work recurrent neural network.
- 18:29:58After covering these topics we will move
- 18:30:00forward and see types of RNN recurrent
- 18:30:03neural network and applications of RNN.
- 18:30:06At the end we will do a hands-off lab
- 18:30:09demo of sentiment analysis using RNN. So
- 18:30:11before starting let us have a simple
- 18:30:13question to brush our knowledge. So
- 18:30:15question is what are the application of
- 18:30:17RNN? Okay, NLP,
- 18:30:20time series, image captioning and all of
- 18:30:24the above. Please answer in the comment
- 18:30:25section below and we will update the
- 18:30:27correct answer in the pin comments or
- 18:30:29you can pause this video, give it a
- 18:30:31thought and answer in the comment
- 18:30:33section. Before we move on to the
- 18:30:34programming part, let's discuss what RNN
- 18:30:37is and proceed further for the same. So
- 18:30:39what is RNN? Recurrent neural network.
- 18:30:41So RNN work on the principle of saving
- 18:30:44output on a particular layer and feeding
- 18:30:46this back to the input in order to
- 18:30:47predict the output of the layer. This is
- 18:30:50how can convert a feed neural network
- 18:30:52into a recurrent neural network RN. The
- 18:30:54node in different layers of neural
- 18:30:56network are compressed to form a single
- 18:30:58layer of recurrent neural network. A B
- 18:31:00and C are the parameters of neural
- 18:31:03network. Now that you understand what
- 18:31:05RNN is, let's look at the way why RNN.
- 18:31:08Okay. So why RNN? RNN were created
- 18:31:11because there are few issues in the feed
- 18:31:13forward neural network cannot handle the
- 18:31:14sequential data considers only the
- 18:31:16current input cannot memorize previous
- 18:31:19input. Okay. So the solution of these
- 18:31:21issues is RNN and RNN can handle
- 18:31:24sequential data accepting the current
- 18:31:26input data and previously received input
- 18:31:28data. So RNN can memorize previous input
- 18:31:31due to their internal memory. So moving
- 18:31:33forward let's see how does RNN networks
- 18:31:37work. Okay. So the input layer X takes
- 18:31:40an input to the neural network and
- 18:31:42process it and the passes it into the
- 18:31:44middle layer. The middle layer edge can
- 18:31:46consist of multiple hidden layers each
- 18:31:49with its own activation function and
- 18:31:51weight and biases. If you have a neural
- 18:31:53network where the various parameters of
- 18:31:55different hidden layers are not affected
- 18:31:57by the previous layer that is the neural
- 18:32:00network does not have the memory then
- 18:32:02you can use RNN. So the RNN will
- 18:32:06standardize the different activation
- 18:32:07function and weights and biases so that
- 18:32:09each hidden layer has the same
- 18:32:11parameter. Then instead of creating
- 18:32:13multiple hidden layers, it will create
- 18:32:15one end loop over it as many time it has
- 18:32:18required. So moving forward let's see
- 18:32:20types of RNN. So there are four types of
- 18:32:23RNN
- 18:32:25one to one,
- 18:32:27one to many, many to many and many to
- 18:32:30one.
- 18:32:32So let's see one to one RNN. So this
- 18:32:35type of neural network is known as the
- 18:32:37vanilla neural network. It is used for
- 18:32:39general machine learning problem which
- 18:32:41has a single input and a single output.
- 18:32:43Now see
- 18:32:45one to many RNN. This type of neural
- 18:32:48network has a single input and multiple
- 18:32:50outputs. An example of this is a image
- 18:32:52captioning. Now let's see many to one
- 18:32:55RNN. This RNN take a sequence of input
- 18:32:58and generates a single output. Sentiment
- 18:33:00analysis is a good example of this kind
- 18:33:02of neural network where a given sentence
- 18:33:04can be classified as expressing positive
- 18:33:06or negative sentiment. And the last one
- 18:33:08is many to many RNN. This RNN takes a
- 18:33:12sequence of inputs and generates a
- 18:33:13sequence of output. Machine translation
- 18:33:15is the one of the example. So moving
- 18:33:17forward, let's see application of
- 18:33:19recurrent neural network. First one is
- 18:33:22image captioning. RNNs are used to
- 18:33:25caption an image by analyzing the
- 18:33:28activities present. The second one is
- 18:33:30time series prediction. Any time series
- 18:33:32problem like predicting the prices of
- 18:33:34stocks in a particular month can be
- 18:33:36solved using RNN. And the third one is
- 18:33:39natural language processing. Text mining
- 18:33:41and sentiment analysis can be carried
- 18:33:43out using RNN or NLP. Natural language
- 18:33:47processing. The fourth one is machine
- 18:33:49translation. Given an input in one
- 18:33:52language, RNNs can be used to translate
- 18:33:54the input into different language as
- 18:33:56output. So now let's move to the
- 18:33:58programming part. First we will import
- 18:34:00some libraries major libraries for the
- 18:34:02first we will import for the data frame.
- 18:34:05So I will write import
- 18:34:12pd.
- 18:34:13The second one is import numpy
- 18:34:17as np.
- 18:34:20So pandas is a software library written
- 18:34:22for the python programming language for
- 18:34:24data manipulation and analysis. In
- 18:34:26particular, it offers a data structure
- 18:34:27and operations for manipulating
- 18:34:29numerical tables and the time series.
- 18:34:32And this numpy numpy is a library for
- 18:34:34the Python programming language adding
- 18:34:36support to four large multi-dimensional
- 18:34:39array and matrices along with a large
- 18:34:41collection of highle mathematical
- 18:34:44function to operate on these arrays.
- 18:34:46Okay. So for plotting we will import
- 18:34:49some libraries like seabon
- 18:34:54as
- 18:34:56SNS. This is nothing just a short form
- 18:34:59of we don't have to write again and
- 18:35:01again CON c we can write SNS. So then
- 18:35:04another one is from
- 18:35:08wordcloud
- 18:35:12port
- 18:35:24mattplot lib
- 18:35:27dot
- 18:35:28pip plot
- 18:35:30s plt. library.
- 18:35:34Okay, so Seabone is a library that uses
- 18:35:36Matt plot lib underneath to plot graphs.
- 18:35:39It will be used to visualize zandom
- 18:35:41distribution and the word cloud is a
- 18:35:44visual representations
- 18:35:46of words. Cloud creators are used to
- 18:35:48highlight popular words and phrases
- 18:35:50based on frequency and relevance. They
- 18:35:52provide you with quick and simple visual
- 18:35:55insights that can lead to more in-depth
- 18:35:57analysis. And this mattplot lil mattplot
- 18:36:00lib is a plotting library for the python
- 18:36:01programming language and its numerical
- 18:36:03mathematic
- 18:36:04ext extension numpy. It provides an
- 18:36:07object- oriented API for embedding plots
- 18:36:09into application using general purpose
- 18:36:12UI.
- 18:36:13Okay. Like tinker wxython QT or gtk.
- 18:36:19So let's import some
- 18:36:22NLTK
- 18:36:25natural language toolkit.
- 18:36:28So
- 18:36:29I will write import
- 18:36:32NLTK.
- 18:36:36Okay. from
- 18:36:38NLTK
- 18:36:40dot stem
- 18:36:43importizer
- 18:36:54then from analytic dot corpus
- 18:36:58imports
- 18:37:06and
- 18:37:08from
- 18:37:10NL ticket dot tokenize
- 18:37:14port
- 18:37:21tokenize
- 18:37:25NLTK the natural language toolkit or
- 18:37:28more commonly NLTK is a suit of
- 18:37:30libraries and programs for symbolic and
- 18:37:33statical natural language processing for
- 18:37:35English written in Python programming
- 18:37:37language and this is stop words. Stop
- 18:37:39words are words that are so common they
- 18:37:41are basically ignored by typical
- 18:37:44tokenizers and this word tokenize is a
- 18:37:46function in Python that splits a given
- 18:37:48sentence into words using the analytical
- 18:37:51library. Okay. So let's import some
- 18:37:54scikitlearn
- 18:37:56library. So for that I will write from
- 18:37:59skarn
- 18:38:02dot model
- 18:38:05collection
- 18:38:08import
- 18:38:10train
- 18:38:12test.
- 18:38:17Okay. Then from skarn
- 18:38:25dot feature
- 18:38:31extraction
- 18:38:34dot text import
- 18:38:39vectorizer.
- 18:38:44And then from
- 18:38:47skarn dot matrices
- 18:38:51matrix
- 18:38:52import
- 18:38:54confusion metric
- 18:39:01classification.
- 18:39:08Okay.
- 18:39:10So, scikitlarn is a free source software
- 18:39:13machine learning library for Python
- 18:39:15programming language. It features
- 18:39:17various classification, regression and
- 18:39:18clustering algorithms including support
- 18:39:21vector machine learning, logistic
- 18:39:22regression and many others like random
- 18:39:25forest classifier. And this train test
- 18:39:28split method is used to split our data
- 18:39:30into train and test set. First, we need
- 18:39:32to divide our data into features like X
- 18:39:34and Y labels. And this TF ID vectorzer
- 18:39:38converts a collection of raw documents
- 18:39:40into a matrix of TF features. The fast
- 18:39:44text or what to vectorizer what
- 18:39:47embedding Python implementation and this
- 18:39:49confusion matrix. A confusion matrix is
- 18:39:51a table that is used to define the
- 18:39:53performance of a classification
- 18:39:54algorithm. Okay.
- 18:39:57Then we'll import some libraries like
- 18:40:00prom skarn
- 18:40:08do linear model
- 18:40:15port
- 18:40:17logistic
- 18:40:21regression.
- 18:40:23So then from
- 18:40:26colonm
- 18:40:32port
- 18:40:40and from
- 18:40:51import
- 18:40:54random
- 18:40:56forests classifier.
- 18:41:01Okay. Then from
- 18:41:04skarn dot name base
- 18:41:10portoli
- 18:41:15base.
- 18:41:19Okay.
- 18:41:22So everything is correct. You will see
- 18:41:26while running. So logistic regression
- 18:41:28estimate the probability of an event
- 18:41:31occurring such as voted or didn't vote
- 18:41:34based on a given data set of the
- 18:41:36independent variable.
- 18:41:40L SVC logistic regression estimate
- 18:41:44sorry linear support vector machine SVC
- 18:41:47is an algorithm that attempts to find a
- 18:41:49hyper plane to maximize the distance
- 18:41:52between classified samples and this
- 18:41:54random forest classifier creates a set
- 18:41:56of decision trees from a randomly
- 18:41:59selected subset of the training set
- 18:42:03and this Bernoli NBoli
- 18:42:05name base is a part of the name base
- 18:42:07family it is based on Bernoli
- 18:42:09distribution ution and accept only
- 18:42:11binary values that is zero or one.
- 18:42:15So let's import some tensorflow. So
- 18:42:17import
- 18:42:20tensorflow
- 18:42:22dot
- 18:42:24compad dot v2 and
- 18:42:29then import
- 18:42:32tensorflow
- 18:42:35data sets
- 18:42:37as tfds.
- 18:42:40So, TensorFlow is a free and open-source
- 18:42:42library for machine learning and
- 18:42:43artificial intelligence across a range
- 18:42:46of task but has a particular focus on
- 18:42:48training and inference of deep neural
- 18:42:50networks. Okay, let's import warnings.
- 18:42:57Nothing. Warning.
- 18:43:02The warnings
- 18:43:11import
- 18:43:18string
- 18:43:21import.
- 18:43:24So everything is basic. Just let's see
- 18:43:25the pickle. Typically is a Python is
- 18:43:28primarily used in serializing and
- 18:43:31deserializing a Python object structure.
- 18:43:33Okay, let's run it. Let's see how many
- 18:43:37error
- 18:43:42after that we will load the data set and
- 18:43:45uh we will go through data
- 18:43:46visualization. Okay. Word cloud cannot
- 18:43:49import name word cloud. Okay. C C will
- 18:43:52be capital here.
- 18:43:57random forest.
- 18:44:06Okay, it's still loading here. Let's
- 18:44:08see. Okay, so loading is done. So now
- 18:44:12let's load the data set. So we'll write
- 18:44:14data equals to PD dot
- 18:44:19read
- 18:44:21CSV
- 18:44:26name
- 18:44:37test.
- 18:44:48So you can find this data set on the
- 18:44:50description box below.
- 18:44:53According to question
- 18:45:04polarity
- 18:45:16ID,
- 18:45:18comma date,
- 18:45:20comma
- 18:45:22query
- 18:45:46per you forget comma
- 18:46:01polarity. Okay.
- 18:46:14Seems fine.
- 18:46:19Let me change this first
- 18:46:28is using RNN.
- 18:46:33Okay.
- 18:46:36So here I will write data plus data dot
- 18:46:42sample.
- 18:46:45Let's
- 18:46:49do one.
- 18:47:05Okay. So, let me like brief uh tell you
- 18:47:09that what we are going. Okay.
- 18:47:14Let me brief you like what we will do in
- 18:47:16this sentiment analysis using RNA. So in
- 18:47:20this demo like you will see uh text
- 18:47:22processing on Twitter data set and after
- 18:47:26that we will perform different machine
- 18:47:27learning algorithms on the data such as
- 18:47:29logistic regression random forest
- 18:47:32classifier SVC nas to classify positive
- 18:47:35and negative dudes. After that I will
- 18:47:38also build RNN recurrent neural network
- 18:47:40which is the best fit for such textual
- 18:47:43sentiment analysis. Okay. Since it's a
- 18:47:46sequential data set which is requirement
- 18:47:48for the RNN network. So let's dive into.
- 18:47:52So now
- 18:47:54we will see the data data visualization
- 18:47:56data set details target like the
- 18:47:58polarity of the tweets zero negative.
- 18:48:01Okay. then the date like date of the
- 18:48:04tweet and the polarity and the user that
- 18:48:07what tweeted then the text okay so I
- 18:48:10will write print
- 18:48:15data set
- 18:48:21data
- 18:48:23shape
- 18:48:25okay
- 18:48:30let me first do like this. Yeah.
- 18:48:35So there are 20
- 18:48:40or you can say two like rows and six
- 18:48:43number of columns. Okay. So it is a huge
- 18:48:46data. I will you can find this data set
- 18:48:49from the description box below. So here
- 18:48:52let's see the data
- 18:48:57and why I use head. Head is used for
- 18:49:01like
- 18:49:03for showing
- 18:49:07top 10 rows of the data set. If you will
- 18:49:10use tail instead of head, it will show
- 18:49:12the last 10 rows of the data set. Okay.
- 18:49:15Here polarity zero. Zero means negative
- 18:49:18and four means positive. Okay. Like you
- 18:49:21can consider 01.
- 18:49:25This is ID, date, then query. then user
- 18:49:29then the text.
- 18:49:35Okay. So
- 18:49:38here I will do data
- 18:49:41clarity.
- 18:49:51Okay. These are the 04. Okay.
- 18:49:54Uniqueness. Zero means negative and the
- 18:49:56four means positive. replacing the value
- 18:49:58four as one for the ease of
- 18:50:00understanding what I said to you you can
- 18:50:02consider as 01. So data
- 18:50:05polarity
- 18:50:09to data
- 18:50:12polarity
- 18:50:19to one
- 18:50:22and then data.
- 18:50:30So now you can see 0 1 0 1 0 1 1 0.
- 18:50:34Okay.
- 18:50:36So if you will write only head it will
- 18:50:38show the top five rows only. Okay.
- 18:50:43So now let's use one Python function
- 18:50:46describe
- 18:50:48data dotribe.
- 18:50:59So as you can see here count is two lakh
- 18:51:02and the mean of the particular row is
- 18:51:05this and the ID is this standard
- 18:51:08deviation minimum value the 25% the 50%
- 18:51:12and the 75% and the maximum
- 18:51:15okay let's see the number of positive
- 18:51:18versus negative tagged sentence okay
- 18:51:27so here I I will write positives
- 18:51:32to data
- 18:51:34polarity
- 18:51:41data dot polarity
- 18:51:44= 1.
- 18:51:47Then it is
- 18:51:50data
- 18:51:54polarity
- 18:51:58data dot polarity
- 18:52:03is equals to zero.
- 18:52:08Print
- 18:52:11total
- 18:52:14length of the data is
- 18:52:27dot format
- 18:52:30data
- 18:52:31dot shape. Yep.
- 18:52:50Now I will print
- 18:52:54the total length, the negative and the
- 18:52:56positive. Okay. So number of positive
- 18:53:09Okay.
- 18:53:14Format
- 18:53:18positives.
- 18:53:24So I will copy
- 18:53:29this and paste it here.
- 18:53:31And here I will do the changes for the
- 18:53:34negatives.
- 18:53:43Okay. Now let's see.
- 18:53:50So here polarity is not defined.
- 18:54:07So as you can see the total length of
- 18:54:09the data is two lakh and the number of
- 18:54:11positive sentences is like one lakh 46
- 18:54:18and number of negatives okay spelling
- 18:54:21this
- 18:54:25the number of negative text sentences
- 18:54:2899,954.
- 18:54:30Okay. So now we have a brief data.
- 18:54:53So now let's get a word count p of text.
- 18:54:56So for this I will write
- 18:55:05count
- 18:55:06words
- 18:55:13done
- 18:55:16length of
- 18:55:18start split.
- 18:55:23Okay.
- 18:55:29And now let's plot a word count
- 18:55:31distribution for both positive and
- 18:55:33negative. So I will create a bar plot.
- 18:55:36So for that I will write it
- 18:55:40word
- 18:55:42count.
- 18:55:45data
- 18:55:48text
- 18:55:50dot apply
- 18:55:54but count.
- 18:56:00Okay, then I will write P positive= data
- 18:56:05then
- 18:56:08count
- 18:56:15data dot polarity
- 18:56:19is equals to 1
- 18:56:23and
- 18:56:25let me copy this Here
- 18:56:35I will write zero
- 18:56:41and okay then
- 18:56:44plt dot figure
- 18:56:48and figure size
- 18:56:52equ= to
- 18:56:5312 Thanks.
- 18:57:07Okay. Then plt
- 18:57:12LT dota
- 18:57:1845
- 18:57:23then plt dot x label
- 18:57:29word count
- 18:57:33plt dot y label
- 18:57:38and frequency
- 18:57:44we'll write uh g
- 18:57:49dot
- 18:57:55comma n
- 18:57:58Uh
- 18:58:19alpha also 0.5 Five
- 18:58:30positive.
- 18:58:39Okay. Then let's make a legend also.
- 18:58:46Location should be
- 18:58:53Right.
- 18:59:00False.
- 18:59:03Data word count equals to
- 18:59:15Okay, my bad.
- 18:59:27So as you can see the positive and the
- 18:59:29negatives.
- 18:59:31Okay.
- 18:59:33So these are the like word count
- 18:59:35distribution for both positive and
- 18:59:37negative. Okay.
- 18:59:41Now let's uh what we can do we can do
- 18:59:45the get like get the common words in
- 18:59:47training data set for the training data
- 18:59:49set. So for that I will do
- 18:59:53from
- 18:59:55collections
- 18:59:57import
- 19:00:01counter
- 19:00:04or
- 19:00:07words
- 19:00:08to
- 19:00:12for
- 19:00:18test
- 19:00:22data
- 19:00:24text
- 19:00:30line
- 19:00:32dot split
- 19:00:37forward. Word
- 19:00:40and words.
- 19:00:48If length of
- 19:00:51word
- 19:00:53than two
- 19:00:58all
- 19:01:01dot
- 19:01:05one dot lower
- 19:01:13here I can write counter
- 19:01:16all words
- 19:01:19dot most
- 19:01:22common then I need 20.
- 19:01:31So as you can see these are the most
- 19:01:33common word used like in every sentence
- 19:01:36the and you for have that I am but just
- 19:01:40like this out over all.
- 19:01:43So these are the most common words like
- 19:01:46it used the is used like 64,000 times
- 19:01:49and like this UR is used for 8,000 times
- 19:01:56something like that. So now we will do
- 19:01:58some data pro data processing. Okay. Now
- 19:02:02let's do the data processing.
- 19:02:05So
- 19:02:08div
- 19:02:17and SNS dot current plot
- 19:02:24data
- 19:02:26polarity.
- 19:02:29Okay,
- 19:02:34these are the uh negatives and this
- 19:02:37positives.
- 19:02:41There is a slight change I guess that is
- 19:02:45why it's not looking
- 19:02:49so much of different like there's a
- 19:02:51slight
- 19:02:5346 different so that is why it's looking
- 19:02:56almost same. Okay.
- 19:03:01So now removing the unnecessary columns
- 19:03:04like query, user, word count, data dot
- 19:03:08drop,
- 19:03:13date
- 19:03:17query
- 19:03:25and word count.
- 19:03:35X = 1
- 19:03:38comma
- 19:03:40place= to true.
- 19:03:45Okay.
- 19:03:47Uh A will be true.
- 19:03:57So here I will write data
- 19:04:01what is this? No. Okay my bad.
- 19:04:15So here I will write data dot drop
- 19:04:19id
- 19:04:21comma
- 19:04:23one
- 19:04:36then data dot head
- 19:04:42the data see we have only the to the
- 19:04:44polarity and the text. Okay.
- 19:04:52So
- 19:04:55now uh let's see the null values.
- 19:05:00So data
- 19:05:02dot
- 19:05:09um
- 19:05:15data
- 19:05:20print. Okay.
- 19:05:43So there is no null values. So now
- 19:05:45converting pandas's object to a string
- 19:05:48type.
- 19:05:50For that we have to write
- 19:05:52text
- 19:05:55to data
- 19:05:59text.
- 19:06:07Yeah.
- 19:06:15Get as type.
- 19:06:22Yeah. So now download the stop words
- 19:06:26NLTK.
- 19:06:29Download
- 19:06:37words.
- 19:06:40words
- 19:06:43as you said
- 19:06:45stop words
- 19:06:51it's in English
- 19:07:00stop
- 19:07:16This
- 19:07:37These are some, you know, stop words.
- 19:07:42So moving forward, let's download
- 19:07:45NLTK dot download.net.
- 19:08:10So the pre-processing steps taken are
- 19:08:12like lower casting each text is
- 19:08:14converted to lower case then remover of
- 19:08:17URLs will do this we will do okay links
- 19:08:20starting with http or https or ww are
- 19:08:24replaced by like commas and removing
- 19:08:28usernames removing short words removing
- 19:08:30stop words like limitization is the
- 19:08:32process will do of for the converting a
- 19:08:35word to its base Okay. So for that
- 19:08:41what I will do
- 19:08:44we'll just copy the whole code for you.
- 19:08:49We'll explain you one by one what I've
- 19:08:51done.
- 19:08:53Okay.
- 19:08:58So this is a course for the URL pattern
- 19:09:00for removing all the WW, HTTPS and HTTP
- 19:09:05type of thing and removing
- 19:09:08them. Then I have used pattern for the
- 19:09:12lower casting removing all the URLs.
- 19:09:15Okay.
- 19:09:16Then removing all the usernames like at
- 19:09:19the red and removing punctuations
- 19:09:22and stop words.
- 19:09:25Okay. Like this.
- 19:09:30So now what we have to do data
- 19:09:40processed
- 19:09:43weights
- 19:09:48then data
- 19:09:51Next
- 19:10:04dot apply
- 19:10:08lambda
- 19:10:10x
- 19:10:12process
- 19:10:17then tweets.
- 19:10:24Okay.
- 19:10:35Then print
- 19:10:39next.
- 19:10:41reprocessing.
- 19:10:50It is taking time
- 19:11:17It will be completed. It will return
- 19:11:19here the text prep-processing is done.
- 19:11:22Okay.
- 19:11:24As you can see the text prep-processing
- 19:11:26is done. So now let's check
- 19:11:31data dot add
- 19:11:3410.
- 19:11:37As you can see see the at the rate and
- 19:11:40this slices are gone.
- 19:11:43Okay.
- 19:11:44So now the text is pre-processed.
- 19:11:48So now what we will do? We will analyze
- 19:11:50the data. So now we are going to analyze
- 19:11:52the pre-processed data to get an
- 19:11:54understanding of it. We will plot word
- 19:11:56clouds for positive and negative dudes
- 19:11:58from our data set and see which words
- 19:12:01occurs the most. Okay. First we will uh
- 19:12:05create for the negative words or
- 19:12:08negative tweets you can say. So I will
- 19:12:10write pl dot figure
- 19:12:13then figure size
- 19:12:1815.
- 19:12:23Okay. Then word cloud also
- 19:12:27word cloud
- 19:12:32x words
- 19:12:362,00 comma
- 19:12:39width = to 1,600
- 19:12:45comma
- 19:12:46height = to 800
- 19:12:52rate
- 19:12:58dot join the data dot polarity
- 19:13:05and I will write here polarity
- 19:13:10okay equals equals to zero
- 19:13:17again
- 19:13:24then
- 19:13:27processed tweets.
- 19:13:33Okay.
- 19:13:36Then here
- 19:13:38I have to write plt dot show
- 19:13:44me show
- 19:13:49wcolation
- 19:14:00linear.
- 19:14:04Perhaps you forget the comma here.
- 19:14:13So 2000
- 19:14:16then comma width
- 19:14:19dot generate
- 19:14:32here. what I can do.
- 19:14:44Let me run now. Let's see. Hope this
- 19:14:48time it will work.
- 19:14:55Guess is still loading.
- 19:15:16As you can see this is
- 19:15:22okay like today I am and work don't wish
- 19:15:27they need much. These are the most
- 19:15:30negative tweets. Okay, words from
- 19:15:33negative tweets you can say,
- 19:15:36right? So, let's see the positive
- 19:15:40tweets. Okay,
- 19:15:46so the thing will be same.
- 19:15:50Let me copy
- 19:15:53paste it here. So for this I will do one
- 19:16:01it will take a little bit of time
- 19:16:05to come loading like as you can see hit
- 19:16:09can't. Okay sorry
- 19:16:14these are the negative words.
- 19:16:18Okay still loading. So let's wait for
- 19:16:21like few seconds.
- 19:16:26Now you can see the positive words like
- 19:16:28love, okay, good, lol
- 19:16:33and awesome something like that. Okay,
- 19:16:37so these are some
- 19:16:39positive words. So now let's do the
- 19:16:42vectorzation and splitting the data like
- 19:16:44storing into input variable process to X
- 19:16:47and output variable polarity to Y. Okay,
- 19:16:50we'll do that.
- 19:16:57So x = to data
- 19:17:02possessed
- 19:17:07with
- 19:17:10values
- 19:17:13and pi= to data
- 19:17:18entity
- 19:17:21dot
- 19:17:22values.
- 19:17:24Okay.
- 19:17:26Now I will write here print
- 19:17:31dot shape
- 19:17:35print y dot.shape.
- 19:17:41Okay cool. So now what we will do we
- 19:17:43will convert text to word frequency
- 19:17:45vectors. Okay. TF to IDF. So this is an
- 19:17:48acronym that stand for term frequency to
- 19:17:52inverse document frequency which are the
- 19:17:54components of the resulting scores
- 19:17:56assigned to each word. Okay. So term
- 19:17:58frequency this summarize how often a
- 19:18:00given word appears within a document and
- 19:18:03inverse document frequency this
- 19:18:05downscales word that appear a lot across
- 19:18:07documents. Okay. So now here we will
- 19:18:11convert a collection of raw documents to
- 19:18:12a matrix of TF to IDF features. Okay.
- 19:18:16And then I will write
- 19:18:19enter
- 19:18:21kazut
- 19:18:23riser
- 19:18:30and sublinear
- 19:18:40x = to
- 19:18:45dot with
- 19:18:48transform
- 19:18:55printed.
- 19:19:06Okay.
- 19:19:11Print
- 19:19:15number of feature
- 19:19:24comma length
- 19:19:28vector
- 19:19:31do get
- 19:19:34their names.
- 19:19:41So number of feature words are like 1703
- 19:19:44to 1.
- 19:19:46Okay. Now we will do like
- 19:19:51now let's print the shape.
- 19:20:10So now we will do the split uh spread to
- 19:20:13train and test. So the pre-provised data
- 19:20:15is divided into two sets of data
- 19:20:17training data and the testing data. So
- 19:20:19data set upon which the model would be
- 19:20:21trained on contains 80% data and the
- 19:20:25test data is the data set upon which
- 19:20:27model would be tested again contains 20%
- 19:20:30of data. So for that I will write extra
- 19:20:37test
- 19:20:39comma
- 19:20:50Test
- 19:21:00test size
- 19:21:02= to 0.20 2
- 19:21:07random
- 19:21:09state
- 19:21:11101.
- 19:21:16Okay. Random state.
- 19:21:23So what I will do? I will do the you
- 19:21:25know print the shape of X train, Y
- 19:21:27train, X test, Y test like how many
- 19:21:29columns are there? Rows not column
- 19:21:32exactly the rows are there. Okay.
- 19:21:35So we'll paste there. So see
- 19:21:39extra train like this is a total was
- 19:21:41like two lakh.
- 19:21:44Okay. So 1 lakh 60,000 in training as we
- 19:21:49discussed earlier like 80% in training
- 19:21:51and 20% in testing. Okay.
- 19:21:56So now let's do the model building.
- 19:21:59Okay. Model evaluating functions. So now
- 19:22:03let's make a model.
- 19:22:07Okay.
- 19:22:08And first I will do I will write and
- 19:22:12then I will explain you the whole. Okay.
- 19:22:16So here what I did uh this will tell you
- 19:22:19the accuracy of the model of training
- 19:22:20data and the testing data. Okay. Then we
- 19:22:24will predict the values for test data
- 19:22:26set and the evaluation for the data set.
- 19:22:28Then we will compute and plot the
- 19:22:30confusion matrix.
- 19:22:32Okay, the both the categories negative
- 19:22:34positives. Okay, group name will be true
- 19:22:36negative and the false positive. Okay,
- 19:22:39so there's nothing that's let's run it.
- 19:22:46So now what we will do? We will do first
- 19:22:48for the logistic regression. So here I
- 19:22:51will write LG equals to
- 19:22:54logistic
- 19:22:58regression.
- 19:23:01Okay. Then history
- 19:23:04equals to LG do fit
- 19:23:09X train,
- 19:23:13Y train
- 19:23:16with model
- 19:23:19evaluate
- 19:23:23LG. Now let's see
- 19:23:30this is for the logistic regression.
- 19:23:34Okay,
- 19:23:38as you can see the accuracy of the
- 19:23:40training data is 83% the testing data is
- 19:23:4277%.
- 19:23:43Okay.
- 19:23:48So this is the confidence matrix the
- 19:23:51predictive value like these are the
- 19:23:53categories.
- 19:23:55Now let's see for the linear SPM. For
- 19:23:57that I will write SPM
- 19:24:02equals to
- 19:24:06SVC
- 19:24:11then SVM
- 19:24:14dot fit
- 19:24:18train.
- 19:24:23Then model
- 19:24:27evaluate
- 19:24:30of SVM.
- 19:24:37Okay.
- 19:24:39And after that we will do for random
- 19:24:40forest and the N base. Okay. Then we
- 19:24:43will start with the RNN.
- 19:24:45So as you can see the accuracy of
- 19:24:47training data is very pretty good 93%
- 19:24:50and logation is 83% and the testing is
- 19:24:55less than
- 19:24:57regression model. Let's see for the
- 19:25:00random forest. So I will write here RF
- 19:25:03equals to
- 19:25:06random forest
- 19:25:10fire
- 19:25:15m=
- 19:25:18to 20
- 19:25:23criterion = to
- 19:25:29tropy Okay.
- 19:25:32Then max
- 19:25:34depth equals to 50.
- 19:25:38Then RF dot fit
- 19:25:44X train,
- 19:25:48Y train
- 19:25:51and model
- 19:25:58evaluate.
- 19:26:05Okay,
- 19:26:14loading. Let's see the accuracy how it
- 19:26:16will come.
- 19:26:26After this we will do for the name base
- 19:26:28and after that we will move on to the
- 19:26:31our main model RNN recurrent neural
- 19:26:34networks.
- 19:26:35It's still loading.
- 19:26:41Guess it will take little bit of time.
- 19:26:51So as you can see the confusion matrix.
- 19:26:56Okay. So training data accuracy is 75%
- 19:27:01very less. So now let's see the last
- 19:27:04model name base. Okay. So, NB equals to
- 19:27:12NB
- 19:27:16NB dot fit
- 19:27:21SP,
- 19:27:26wide train.
- 19:27:33Okay. Then model
- 19:27:37evaluate.
- 19:27:45Oh, NAB base training 867.
- 19:27:50So, as for
- 19:27:52linear SEC has the best
- 19:27:55test training uh accuracy you can say
- 19:27:58and the best testing accuracy is 7670
- 19:28:0376.45 4 five
- 19:28:06see logistic regression. So now let's
- 19:28:08move to the our main model RNN. So what
- 19:28:12is RNN recurrent neural network at the
- 19:28:14start
- 19:28:16are the state-ofthe-art algorithm for
- 19:28:18sequential data and are used by Apple CD
- 19:28:21and Google search voice. It is the first
- 19:28:24algorithm that remembers its input due
- 19:28:27to an internal memory which make it
- 19:28:28perfectly suited for machine learning
- 19:28:30problem that involve sequential data.
- 19:28:33And there is one more thing embedding
- 19:28:35layer. Embedding layer is one of the
- 19:28:36available layers in KAS. This is mainly
- 19:28:39used in natural language processing
- 19:28:41related applications such as language
- 19:28:43modeling but it can also be used with
- 19:28:45other tasks that involve neural networks
- 19:28:47while dealing with NLP problems. We can
- 19:28:49use pre-trained word embedding such as
- 19:28:51glow.
- 19:28:53Alternately we can also train our own
- 19:28:56embeddings using kas emitting layer.
- 19:29:00LSTM layer long short-term memory
- 19:29:02networks usually called LSTMs I have
- 19:29:05made already many videos you can check
- 19:29:08it out were introduced by Skyer these
- 19:29:11have widely been used for speech
- 19:29:13recognition language processing
- 19:29:15sentiment analysis and text prediction
- 19:29:17before going deep into LSTM we should
- 19:29:20first understand the need of LSTM which
- 19:29:22can be explained by the drawback of
- 19:29:24practical use of RNN so let's start with
- 19:29:26RNA
- 19:29:28Okay.
- 19:29:31So here I will importing some libraries.
- 19:29:36Okay. So after that I will write import
- 19:29:41kas
- 19:29:48version
- 19:29:512.110. Okay fine.
- 19:29:57So now let's
- 19:30:03paint
- 19:30:19X test, comma,
- 19:30:23white train.
- 19:30:29Let's do train
- 19:30:33test
- 19:30:41weights, comma, data dot polarity
- 19:30:46dot values.
- 19:30:49Then test
- 19:30:52size equals to 0.2. Test size 0.2 means
- 19:30:59like 80 and 20%
- 19:31:02thing 80 to training and then 20% to
- 19:31:08testing.
- 19:31:13Okay.
- 19:31:15and let's
- 19:31:20the model evaluation. Okay.
- 19:31:24So I will these are relu sigmoid all the
- 19:31:28you know the layers.
- 19:31:32So now this epoch it will run till 5,000
- 19:31:35like count will go till 5,000. Okay see
- 19:31:40the 5,000 and it will go to 1 to 10. So
- 19:31:42it will take time. So I will get back to
- 19:31:44you after this completing this. Okay.
- 19:31:48Now as you can see uh the box
- 19:31:52ran successfully. Okay. So what should I
- 19:31:56do? But I will give some space here. So
- 19:31:59now we will see the positive and
- 19:32:01negative outcome. Okay. This is
- 19:32:04something like testing. Okay. We will
- 19:32:06test. We will predict. we will give one
- 19:32:10uh a sentence and then we will predict
- 19:32:13it is coming right or wrong. The
- 19:32:15accuracy is giving a right or wrong.
- 19:32:17Okay. So here I will write
- 19:32:20sequence
- 19:32:22equals to tokenizer
- 19:32:26dot text
- 19:32:31to
- 19:32:33sequences.
- 19:32:36Okay, then I'll write this
- 19:32:42data science
- 19:32:45article.
- 19:32:48This was
- 19:32:51okay.
- 19:32:53So here I will write test equals to P
- 19:32:58sequences
- 19:33:02and here I will write sequence
- 19:33:07Comma max length
- 19:33:13to
- 19:33:15max length.
- 19:33:24Then I will write here prediction equals
- 19:33:28to model.
- 19:33:37We write model
- 19:33:43then we'll write model 12
- 19:33:49dot predict
- 19:33:53then test.
- 19:33:55Okay.
- 19:33:57If diction
- 19:34:02is greater than 0.5 means 50%.
- 19:34:06Then
- 19:34:08it should print
- 19:34:11positive.
- 19:34:14Okay.
- 19:34:16Else
- 19:34:22negative.
- 19:34:26Okay. Let me run this.
- 19:34:30Okay. Sequential
- 19:34:32object has no okay spellic
- 19:34:40see the negative because here is the
- 19:34:42word worst it is showing correct. Now
- 19:34:45check from the RNN model. So model
- 19:34:48equals to kas dot models dot load
- 19:34:54models. Here we will load RNN model. RNN
- 19:34:58model
- 19:35:00SG file. It is pre-trained model. Okay.
- 19:35:04Pretend RNN model. So sequence
- 19:35:10tokenizer
- 19:35:14dot text
- 19:35:17to sequences.
- 19:35:22Then
- 19:35:24I will write here this this
- 19:35:31ML
- 19:35:34course
- 19:35:37is best.
- 19:35:40Okay. S equals to P sequences
- 19:35:47sequence
- 19:36:00X
- 19:36:04okay then prediction equals to model dot
- 19:36:09predict
- 19:36:11Then test
- 19:36:15if prediction is greater than 0.5
- 19:36:23in
- 19:36:24positive 0.5 means 50% more than 50%.
- 19:36:29Else
- 19:36:31print negative
- 19:36:42attribute load models
- 19:36:50positive because this ML course is best.
- 19:36:53So there is no negative word.
- 19:36:57Okay.
- 19:37:05So what we will do now we will do model
- 19:37:08saving loading and prediction. Okay. So
- 19:37:11for that uh I will write import pickle
- 19:37:17file = to open
- 19:37:23vectorzer
- 19:37:32then
- 19:37:40here I will Pickle
- 19:37:42dot dump
- 19:37:45dump
- 19:37:46vector file vector.
- 19:38:03Okay.
- 19:38:04So like this I have to write for name
- 19:38:07base logitation SVM and random forest.
- 19:38:12So
- 19:38:13what I will do
- 19:38:16right here.
- 19:38:17Okay. Let's run this.
- 19:38:20Okay.
- 19:38:35Now what we have to do? We have to
- 19:38:37predict using saved model. Okay.
- 19:38:47What we will do here? We will load model
- 19:38:49first and we will predict. Okay. So
- 19:38:55first I will write the function name
- 19:38:57load
- 19:38:59models
- 19:39:02and we will load the vectorzer. So file
- 19:39:05equals to
- 19:39:07open
- 19:39:09vectorzer
- 19:39:11dot pickle
- 19:39:16IBizer
- 19:39:31file.
- 19:39:33file dot close.
- 19:39:39Now I'm loading the logistic regression
- 19:39:41model. So for that we have to write open
- 19:39:57B
- 19:39:59LG to pick
- 19:40:04code
- 19:40:07file
- 19:40:09then file dot close
- 19:40:13then
- 19:40:18riser
- 19:40:20LG.
- 19:40:22Okay.
- 19:40:27Yeah. So now we will predict the
- 19:40:30sentiment. So for that I will write here
- 19:40:33predict
- 19:40:37riser
- 19:40:47text.
- 19:40:49Okay. Uh so here we will predict the
- 19:40:52sentiment. So for that
- 19:40:59text equals to
- 19:41:02process
- 19:41:11then demands for
- 19:41:15sentiment in
- 19:41:17text.
- 19:41:21Then text
- 19:41:24data
- 19:41:30dot transform.
- 19:41:37This is
- 19:41:42okay. Then sentiment
- 19:41:49model dot predict.
- 19:41:57So here I will make a list of text with
- 19:41:59sentiment. So for that I will write data
- 19:42:01equals to empty array. Then for text
- 19:42:07prediction
- 19:42:09and zip
- 19:42:12text x
- 19:42:14sentiment
- 19:42:18dt
- 19:42:20append
- 19:42:23text prediction.
- 19:42:26Okay.
- 19:42:29Then we will convert the list into pas
- 19:42:31data frames. So for that I will write df
- 19:42:33= to ad dot data plane
- 19:42:41comma columns
- 19:42:44person
- 19:42:46next
- 19:42:48comma
- 19:42:50sentiment
- 19:42:58then df equals to df dot
- 19:43:02Replace
- 19:43:07comma 1
- 19:43:18positive
- 19:43:24and here.
- 19:43:30Okay.
- 19:43:32So at last I will write here if
- 19:43:41to
- 19:43:47then here we will loading the model
- 19:43:50vectorzer
- 19:43:52comma ng plus load.
- 19:44:00Here we text to classify like what
- 19:44:03should be in the list. So like text
- 19:44:08here I will like I love machine
- 19:44:13name.
- 19:44:21So
- 19:44:33John
- 19:44:37be so
- 19:44:52So here df equals to date
- 19:45:02the command text
- 19:45:06then print
- 19:45:08df.
- 19:45:19I love machine learning. Positive. B is
- 19:45:21so active. Positive. J I feel so good.
- 19:45:24Negative. Okay. There is
- 19:45:39and
- 19:45:41yeah. See now it's coming. Okay.
- 19:45:47This is how you can do the sentiment
- 19:45:50analysis using uh RNN model. Here we
- 19:45:53have loaded RNN model. So it is showing
- 19:45:55right. So let's do them as poses
- 19:46:04right here.
- 19:46:09Add
- 19:46:11one.
- 19:46:25negative. Okay,
- 19:46:29RNN model is working. So right, today we
- 19:46:32are going to explore K nearest neighbors
- 19:46:34or KN&N which is one of the most popular
- 19:46:37algorithms in data science. Python is a
- 19:46:40powerful tool for data science and KNN
- 19:46:42is great for classifying data by
- 19:46:44predicting the category of sample base
- 19:46:46on its closest neighbor. This algorithms
- 19:46:49is used in many field like healthcare,
- 19:46:52finance and agriculture helping us make
- 19:46:55decision based on data. The best part it
- 19:46:58is really easy to use. You just need to
- 19:47:00pick a number for K and choose a
- 19:47:03distance function to compare data
- 19:47:05points. However, KN&N has it downsides.
- 19:47:08It doesn't work well with the large data
- 19:47:10set and it require proper scaling of the
- 19:47:12data to get accurate result. In this
- 19:47:15video, we will show you how KNN work
- 19:47:17with real data set, the Iris data set.
- 19:47:19We'll walk you through simple Python
- 19:47:21code and demonstrate how to find the
- 19:47:23best K value to maximize your model's
- 19:47:26accuracy. So stay tuned to seekn in
- 19:47:28action. So welcome to the demo part. So
- 19:47:30I'm here using Google Collab. So you can
- 19:47:33use any of your favorite ID like Jupyter
- 19:47:36notebook, Intelligi, Visual Code Studio,
- 19:47:39anything. Okay. So let me rename this
- 19:47:43file as KNN classification.
- 19:47:51Okay, cool. So let me tell you that KN&N
- 19:47:55can be used for the classification
- 19:47:56regression predictive problems. So KN
- 19:47:59falls in the supervised learning family
- 19:48:00of algorithms. Okay, so we will measure
- 19:48:03the distance between the K neighbors and
- 19:48:06the first step will be we will choose
- 19:48:08the number of K of neighbors. Then uh uh
- 19:48:11we'll take the k nearest neighbors of
- 19:48:14the new data point according to your
- 19:48:15distance metric. And the step three will
- 19:48:18be our among these case neighbors count
- 19:48:21the number of data points of each
- 19:48:23category. Okay. Step four we will assign
- 19:48:25the new data points to the category
- 19:48:27where you counted the most neighbors.
- 19:48:29Okay. So let's start. First let's import
- 19:48:32some library. import
- 19:48:35numpy
- 19:48:37as np and let's import
- 19:48:42pandas sp. So everyone knows what is
- 19:48:45numpy and the pandas. Okay, so numpy is
- 19:48:49a library for the python programming
- 19:48:51language adding support for uh you know
- 19:48:53large multi- dimensional arrays. Okay.
- 19:48:56Along with the large collection of uh
- 19:48:58what to say uh highlevel mathematical
- 19:49:01functions and various pandas is a
- 19:49:03software library written for the Python
- 19:49:04programming language for the data
- 19:49:06manipulation and analysis uh all the
- 19:49:09data frames and uh data structures it
- 19:49:12offers for the manipulating numerical
- 19:49:14tables. Okay. And the time series you
- 19:49:16can see. So moving forward uh we'll
- 19:49:18import our data set. So you can download
- 19:49:21the data set from the description box
- 19:49:23below. Okay. data set
- 19:49:26equals to pb dot read
- 19:49:30csv. The data name is iris dot csv.
- 19:49:34Okay. This is how you read uh your data
- 19:49:38in python. Okay. Yeah. Data set is
- 19:49:42loaded. So data set dot shape
- 19:49:47shape. Okay. Yeah. So data uh set dot
- 19:49:52shape is used for how many numbers of
- 19:49:57rows and columns present in your data
- 19:49:59set. Okay, 150 rows and six columns. So
- 19:50:01let me tell you brief about data set
- 19:50:04this data set. So this data set include
- 19:50:08three Iris species with uh 50 samples
- 19:50:11each as well as some properties about
- 19:50:14each flower. So one flower species is
- 19:50:17linearly separable from the other two
- 19:50:19you can say but the other two are not
- 19:50:22linearly separable from each other.
- 19:50:24Okay. And this shape I told you we can
- 19:50:27get a quick idea of how many instances
- 19:50:28of rows and columns are present in our
- 19:50:30data set. So let's see our data set.
- 19:50:34Data set dot
- 19:50:36head. So head is used for uh you know by
- 19:50:41you can see top five rows of your data
- 19:50:44set using head and if you will use tail
- 19:50:47you instead of head you can see the last
- 19:50:50five rows of your data set. Cool. Okay.
- 19:50:52So columns are ID sample length sample
- 19:50:55width petal length petal width and the
- 19:50:58species. Cool. Then moving forward let's
- 19:51:01describe our data sets. So these are the
- 19:51:03basics uh basic function. Okay.
- 19:51:08of Python you can say.
- 19:51:12So data set.escribed what describes do
- 19:51:16is it will give you count of all the
- 19:51:18rows mean value standard deviation value
- 19:51:21minimum value what is the 25% okay of
- 19:51:25all the values in the particular row.
- 19:51:28What is the 50%? What is the 75%? What
- 19:51:32is the maximum? Maximum is 150 you can
- 19:51:34say. Okay. and 25% of 150 is 38.25 25
- 19:51:37this okay of all the columns if it is
- 19:51:40normal uh you know character so it won't
- 19:51:42give you any data okay cool yeah so
- 19:51:46moving forward uh let's now take a look
- 19:51:49at the number of instances row belong to
- 19:51:52each classes okay so we will write data
- 19:51:56set dot
- 19:51:59group by
- 19:52:03species
- 19:52:04dot size. Okay. Uh spec S is capital
- 19:52:10that's why it's showing the error. Yeah.
- 19:52:14So you can see Iris Satossa are 50, Iris
- 19:52:17verical are 50 and virginica is 50.
- 19:52:20Okay. And the data type type is integer.
- 19:52:23Cool. So as you can see data set
- 19:52:25contains six columns like ID, sample
- 19:52:28length, sample width and petal length,
- 19:52:30petal width and spacing. The actual
- 19:52:32features are described by columns 1 to
- 19:52:34four. the last columns labels or
- 19:52:37samples. Okay. So firstly we need to
- 19:52:39split data into two arrays like X
- 19:52:41features and Y labels. So how we will do
- 19:52:44this? By writing code like feature
- 19:52:48columns equals to
- 19:52:51sample length sample width petal length
- 19:52:54petal width. Just remember you are
- 19:52:56writing correct name. Okay. Then x = to
- 19:53:01data set
- 19:53:03feature
- 19:53:06columns.
- 19:53:09Okay. Dot
- 19:53:12values.
- 19:53:13Then y = to
- 19:53:18data set
- 19:53:20species dot values. Cool. Then let me
- 19:53:25run it. So this is uh how we can split
- 19:53:29the data set okay into two arrays X and
- 19:53:31the Y. Okay. In X there are feature
- 19:53:34columns. These four columns are there
- 19:53:36and in Y species column is there. Okay.
- 19:53:39And what is the species column this
- 19:53:41Satossa venica and all this. Cool.
- 19:53:45Then now we will do label encoding. So
- 19:53:47as you can see labels are categorical
- 19:53:49Kverse classifier does not accept string
- 19:53:51labels. So we need to use label encoder
- 19:53:55to transform them into numbers. Okay.
- 19:53:57Then iris satossa correspond to zero.
- 19:54:00Iris vericy color correspond to one and
- 19:54:04I is virginica correspond to two. Okay.
- 19:54:07012. Cool. So how we can write from
- 19:54:13skarn dot pre-processing
- 19:54:18import
- 19:54:21label
- 19:54:23encoder okay so what we'll do label
- 19:54:26encoder transform them into numbers okay
- 19:54:29so I will write
- 19:54:31l equals to
- 19:54:34label encoder
- 19:54:36okay my bad label encoder. Okay. Then y
- 19:54:41= to ele alate transform y. Okay. So now
- 19:54:44I will run it. Yeah. Correct. So
- 19:54:47splitting data set into training set and
- 19:54:49the test set now. Okay. So now we'll
- 19:54:51split data set into training set and
- 19:54:53test set to check later on whether or
- 19:54:56not a classifier work correctly or not.
- 19:54:58Okay. So here I will write from skarn
- 19:55:02dot not cross. I will write it here.
- 19:55:06Skarn domodel selection
- 19:55:12import
- 19:55:14train test split. So what is train test
- 19:55:17split? Train test split is a model
- 19:55:19validation procedure that reveals how
- 19:55:21your model performs on your new data.
- 19:55:24Okay. And what is X train access Y train
- 19:55:28Y test in Python. Okay. Let me first
- 19:55:30write it and then I will let you know.
- 19:55:32Okay. Then I will write here
- 19:55:35x train comma x test comma y train comma
- 19:55:43y test. Okay
- 19:55:48train test
- 19:55:51split
- 19:55:52then x comma y
- 19:55:56comma test size
- 19:55:580.2 and the random is this. Okay. Okay.
- 19:56:02Some error came. Okay. Underscore model
- 19:56:05selection.
- 19:56:08Yeah. Cool. So what is X train X test? Y
- 19:56:12train Y test. Okay. So X train and Y
- 19:56:15train sets are used for training and
- 19:56:17fitting the model. Okay. So the X test
- 19:56:20and the Y test are the set used for
- 19:56:23testing the model and it's predicting
- 19:56:26the right outputs level. Okay. So here
- 19:56:30you can see test size is 0.2. into means
- 19:56:3380% is for testing or sorry 80% is for
- 19:56:37training and 20% is for testing for the
- 19:56:41new data. Cool. Yeah. So now we will see
- 19:56:44some uh let's do some data
- 19:56:46visualization. Okay. So here I will
- 19:56:48write import
- 19:56:50mattplot lib
- 19:56:53dotpipplot
- 19:56:56as plt
- 19:56:59then import
- 19:57:01cb
- 19:57:03as sns
- 19:57:05then here I will write person mattplot
- 19:57:08lib
- 19:57:10in line. So what is mattplot lib?
- 19:57:12Mattplot lib is a plotting library for
- 19:57:14the python programming language and it's
- 19:57:17numerical mathematic extension. Okay. So
- 19:57:19numpy it provides an object API for
- 19:57:22embedding plots into application using
- 19:57:25generating purpose GUI toolkits like
- 19:57:27kintter and python or gtk. Okay. Whereas
- 19:57:32seabon seabon is a library for making
- 19:57:34statical graph in python. It builds on
- 19:57:37the top of mattplot lip and integrates
- 19:57:39closely with pandas data structure.
- 19:57:42Okay, seborn helps us to explore and
- 19:57:45understand the data. Okay, so I will run
- 19:57:49it here. I will write from
- 19:57:54pandas dot plotting
- 19:57:58import
- 19:58:01parallel
- 19:58:04coordinates.
- 19:58:07Then plt dot figure
- 19:58:11size should be
- 19:58:1515, 10. Okay.
- 19:58:19Then parallel coordinates.
- 19:58:25Okay. Then data set dot drop. I don't
- 19:58:30need id,
- 19:58:33x is one.
- 19:58:36Okay. then comma spaces
- 19:58:41then plt dot title
- 19:58:44and let parallel
- 19:58:48coordinates
- 19:58:50plot okay and
- 19:58:53you can give some font size
- 19:58:56equals to 20
- 19:58:59then font
- 19:59:02weight
- 19:59:03equals to
- 19:59:05okay Let's add bold only bold then plt
- 19:59:10dot x label
- 19:59:14then
- 19:59:15features
- 19:59:17comma
- 19:59:20font size
- 19:59:22to 15 then plt
- 19:59:25dot y label then I'll write here
- 19:59:29features
- 19:59:31values
- 19:59:33comma font size
- 19:59:37equals to 15. Okay. then plt dot legend
- 19:59:45then locals to 1 comma
- 19:59:53I will write have frame on
- 19:59:56equals to true comma shadow
- 20:00:00equals to true comma face color
- 20:00:06equals to White
- 20:00:11T should be capital
- 20:00:15and comma edge color
- 20:00:19equals to
- 20:00:25okay plt dot show okay some error is
- 20:00:28there plig
- 20:00:30size okay spelling mistake Take.
- 20:00:37Okay. One more error. PLT. Legend. Okay.
- 20:00:41Face color. Okay. Some spelling stick.
- 20:00:45So yeah, let me make it output in full
- 20:00:48screen. Yeah. So parallel coordinates is
- 20:00:50a plotting technique for plotting you
- 20:00:52know multivariate data. So it allows one
- 20:00:56to see clusters in the data and to
- 20:00:59estimate other stat visually. So using
- 20:01:02parallel coordinate points uh you know
- 20:01:04are represented as the connected line
- 20:01:06segments as you can see. Okay. And each
- 20:01:10vertical line represent one attribute
- 20:01:12and one set of connected line segments
- 20:01:15represent one data point. Okay. And
- 20:01:18points that tend to cluster will appear
- 20:01:21closer together. Okay. So this uh this
- 20:01:25color is iris satossa and this is irisy
- 20:01:29color and this ba one is iris virginica
- 20:01:33as you can see in the legend. Okay,
- 20:01:35cool. So moving forward let's create
- 20:01:38another graph and curves. Okay so here I
- 20:01:42will write from
- 20:01:44p and dotplotting import scatter. Oh,
- 20:01:48this uh plotting import
- 20:01:53Andreo curves.
- 20:01:56Okay, then plt
- 20:01:58dot figure. It's AI based. So it's
- 20:02:02giving me suggestions. Suggestions.
- 20:02:04Suggestions. Okay. So sometimes
- 20:02:05suggestions are good but not always.
- 20:02:08Yeah. Let's carry on. Figure size is 15,
- 20:02:1210.
- 20:02:15Then Andrew
- 20:02:20curves
- 20:02:22then I will add data set dot drop id
- 20:02:25access spaces and plt and curves plot.
- 20:02:29Okay fine. Okay, let me add we don't
- 20:02:34need X label and all. Let me add legend
- 20:02:38plot
- 20:02:40legend.
- 20:02:41Then same LOC equals to 1. Then
- 20:02:47proposition size.
- 20:02:50Then I will write here size
- 20:02:54is 15
- 20:02:56of frame on equals to two comma shadow
- 20:03:03equals to true. So this is uh truly
- 20:03:07based upon you if you want to add legend
- 20:03:10or not or you can skip. If you want to
- 20:03:13skip you can skip. Okay. Face color
- 20:03:16equals to white. Then
- 20:03:20edge color equals to black. Okay. Then
- 20:03:26plt dot show.
- 20:03:29Yeah. So no error is there. So let me
- 20:03:32first view output in full screen. Okay.
- 20:03:36So Andrew curves these are the endoc
- 20:03:39curves. Okay. You can see the graph in
- 20:03:41the curve. So and curve allow one to
- 20:03:44plot multivariate data as a large number
- 20:03:47of curves. So that are created using the
- 20:03:50attributes of samples okay as
- 20:03:52coefficient for 4year series. Okay. So
- 20:03:54by coloring these curves differently for
- 20:03:57each class. It is possible to visualize
- 20:04:00data clustering. So curves belongs
- 20:04:03belonging to samples of the same class
- 20:04:06will usually be closer together and they
- 20:04:08form large structure. Okay. As you can
- 20:04:11see here and this is the legend why we
- 20:04:13are setting the four color is should be
- 20:04:15white and yeah and the edge color should
- 20:04:19be black okay and frame on shadow should
- 20:04:22be there you can see the shadow okay
- 20:04:24like this if you want to skip you can
- 20:04:26skip this part legend part legend part
- 20:04:28okay but like it's good to have it's
- 20:04:33good practice to have this cool so let's
- 20:04:36create one small pair plot okay so I
- 20:04:39will write here plt dot
- 20:04:42figure
- 20:04:46then I will write SNS dot pair plot
- 20:04:51then I will write data set
- 20:04:53dot drop
- 20:04:56we'll drop again id we don't need comma
- 20:05:00xis is one then I will write u equals to
- 20:05:07species
- 20:05:09size equals to three.
- 20:05:13Then markers equals to
- 20:05:17O SD. Yeah. Cool. Then plt dot show.
- 20:05:23It's running. Yeah. So first let me make
- 20:05:27it to full screen. Yeah. So here pair
- 20:05:30wise is useful when you want to
- 20:05:32visualize the distribution of the
- 20:05:34variable or the relationship between the
- 20:05:36multiple variable separately within
- 20:05:39subsets of your data set. Okay. So here
- 20:05:42this blue one is stoa verol and the
- 20:05:46virginica. Okay. So these are some uh
- 20:05:49graph you can say or here you can see
- 20:05:51sample length. Okay. So green one is
- 20:05:55virginica and versol are almost having
- 20:05:58same saple length. Okay. And here are
- 20:06:01the different types of graphs. Okay. If
- 20:06:04you don't want to uh let's say if you
- 20:06:07don't know how to read this graph you
- 20:06:10can use this graph or you can use this
- 20:06:12graph. Either you can use this graph.
- 20:06:14Okay. That's the power of pair pair wise
- 20:06:17plots you can say. And the sample width
- 20:06:20cm. Okay. Satossa. No. Okay, virginica
- 20:06:24and verticola again having almost same
- 20:06:28sample width. Okay, and here as well.
- 20:06:32Okay, this stoa is in different form.
- 20:06:35Okay. Yeah. So, we'll uh create one
- 20:06:39small uh pair of u one small graph then
- 20:06:42we'll move forward. Okay. Uh I will
- 20:06:45write here plt dot. So now we are
- 20:06:48creating box plot figure.
- 20:06:54Okay. Then data set dot drop.
- 20:06:58Then again ID commais should be one
- 20:07:05dot box plot. Okay. Then figure size I
- 20:07:09chose this. Yeah. Let's run it. Yeah.
- 20:07:12Again see this is the box plot. Petal
- 20:07:15length same petal width sap length sle
- 20:07:17width okay according to width it is
- 20:07:20showing like from this to this width we
- 20:07:23have the veric color and from this to
- 20:07:26this width this length sorry sample
- 20:07:28length we have uh that virginica lies
- 20:07:32here Iris virginica and from here to
- 20:07:35here iristosa lies okay the sample cool
- 20:07:39and like you can use 3D models you can
- 20:07:43use different types of charts. Okay, I
- 20:07:46did four and four are enough to read the
- 20:07:49data set. Okay, so now we will do uh
- 20:07:52some cannon classification. Okay, we
- 20:07:54will make prediction and we will see the
- 20:07:56accuracy of our data set. Okay, so now
- 20:07:59what I will do? I will write here from
- 20:08:04skarn
- 20:08:07dot
- 20:08:09neighbors
- 20:08:11import k neighbors. Okay.
- 20:08:15Write from skarn dot
- 20:08:21neighbors
- 20:08:23import
- 20:08:26k neighbor
- 20:08:29classifier.
- 20:08:31Okay.
- 20:08:33Then I will write here from
- 20:08:36skarn
- 20:08:37dot matrix
- 20:08:40import
- 20:08:43confusion
- 20:08:46matrix
- 20:08:48comma accuracy
- 20:08:50accuracy score. Okay, then I will write
- 20:08:53from
- 20:08:55skarn dot model selection
- 20:09:01import
- 20:09:04cross
- 20:09:05value score. Okay, so here what we I did
- 20:09:10uh we are fitting the classifier to the
- 20:09:12training set and loading the libraries
- 20:09:14basically. Okay, then we'll initiate a
- 20:09:18learning model K equals to three. Okay.
- 20:09:21So here I will write classifier.
- 20:09:24Let me give one space. Classifier equals
- 20:09:26to
- 20:09:30neighbors.
- 20:09:31Classifier
- 20:09:34and
- 20:09:37neighbors
- 20:09:39to three. Okay. Then fitting the model.
- 20:09:42I will write here classifier
- 20:09:45dot fit
- 20:09:50x train,
- 20:09:53y train.
- 20:09:56Okay. Then uh now we will predicting the
- 20:09:59test uh test set results. Okay. So y
- 20:10:02prediction equals to
- 20:10:06classifier dot predict
- 20:10:11x test. Cool. Okay. From Okay. Spelling
- 20:10:16mistake.
- 20:10:19Okay.
- 20:10:23Okay.
- 20:10:28Yeah. I guess it's fine now. So now uh
- 20:10:32let's evaluate the prediction. So here I
- 20:10:35will uh build the confusion matrix.
- 20:10:38Okay. So here I will write cm confusion
- 20:10:40matrix equals to confusion
- 20:10:44matrix
- 20:10:46uh y test
- 20:10:48y prediction. Okay. Then here I will
- 20:10:51write cm. Okay. So what is confusion
- 20:10:54matrix basically? So I confusion matrix
- 20:10:56uh you know is a table that shows how
- 20:10:59well a model performs by comparing its
- 20:11:02prediction to the actual values. Okay.
- 20:11:05So what it show is a confusion uh matric
- 20:11:08display the number of correct and
- 20:11:10incorrect prediction of each class in
- 20:11:13models whose either it can give you true
- 20:11:17positive true negative false positive
- 20:11:19false negative either it can give you
- 20:11:21zero or one. Okay. So yeah moving
- 20:11:24forward let's calculate the model
- 20:11:25accuracy. This is the main part. If the
- 20:11:27model accuracy is low it means uh your
- 20:11:31analysis of you know classification is
- 20:11:34not good. Okay. So accuracy it should be
- 20:11:38more than 80 at least. accuracy
- 20:11:41equals to
- 20:11:43accuracy
- 20:11:45score
- 20:11:48y test
- 20:11:50comma y prediction into 100 otherwise it
- 20:11:54will give me in points so I will write
- 20:11:56print
- 20:11:58model
- 20:12:01canon model accuracy
- 20:12:05is
- 20:12:07oh I will write plus STR I will round
- 20:12:10it. Okay. Accuracy
- 20:12:20I will add the person. So why I wrote
- 20:12:23this accuracy to round? So I don't want
- 20:12:26after points I need only two numbers.
- 20:12:28Okay. I don't want like 6 7 8 9 10 11 12
- 20:12:30like this. Okay. I need only like 80.20
- 20:12:34like this. Cool. Let's run this. see
- 20:12:38KN&N model accuracy is 96.67 67 that's
- 20:12:41why I wrote two here and the percent
- 20:12:43should be there so 96 which is very very
- 20:12:46very good okay so this is how you can
- 20:12:50find uh the accuracy so now let's find
- 20:12:53the optimal number of neighbors in K
- 20:12:56okay basically finding the best K so we
- 20:12:59will use uh using cross validation
- 20:13:02parameter okay tuning so first I will
- 20:13:05create the list of K for KN okay so here
- 20:13:08I will write K
- 20:13:10list
- 20:13:12equals to list
- 20:13:15range
- 20:13:181 comma 50A 2.
- 20:13:22Here I'm I will create the list of CV
- 20:13:25score. Okay. So here I will write CV
- 20:13:29scores.
- 20:13:32Okay. Okay. I have to give brackets.
- 20:13:34Yeah.
- 20:13:36So we'll here perform the 10fold cross
- 20:13:39validation. Okay, I will explain you
- 20:13:41what is cross validation. Don't worry.
- 20:13:43So first let me write for K N K listN
- 20:13:52equals to
- 20:13:55K neighbor classifier
- 20:13:58and here I will add N neighbor
- 20:14:04neighbors equals to K. Then scores
- 20:14:08equals to cross
- 20:14:12value score
- 20:14:15KN&N then X train
- 20:14:19Y train
- 20:14:21okay
- 20:14:23I will write here cross validation
- 20:14:25equals to 10
- 20:14:28comma scoring equals to I will write
- 20:14:31here accuracy
- 20:14:35Okay. And then here cv
- 20:14:40course dotappend
- 20:14:43to course dome.
- 20:14:46Okay. So now yeah let me run it. Okay.
- 20:14:52Comma some error came. Okay. The error
- 20:14:56is the scoring parameter.
- 20:15:00Yeah. Why? Because here accuracy you can
- 20:15:02see and I did the spelling mistake.
- 20:15:05Okay. So now what is cross validation?
- 20:15:07Of course cross validation uh you know
- 20:15:09determine the accuracy of your machine
- 20:15:11learning model by partitioning the data
- 20:15:14into two different groups. Okay called
- 20:15:17training set and testing set. You can
- 20:15:18see train test and the testing set. Okay
- 20:15:21X and Y. So the data is randomly
- 20:15:24separated into a certain number of
- 20:15:25groups or subsets called folds. Okay you
- 20:15:29can see the 10 folds we have wrote. Each
- 20:15:32fold contains about the same amount of
- 20:15:34okay and there is one more thing
- 20:15:36validation. So validation is a technique
- 20:15:39for assessing the accuracy of the model
- 20:15:41on data set. Okay. And this cross
- 20:15:43validation we did on new data set. Cool.
- 20:15:46So now let's find the best K. Okay. So
- 20:15:49here I will write best K equals to K
- 20:15:56list. Okay. Before that I will write one
- 20:15:59thing. I will write here MSE was
- 20:16:03changing to mclassification error. Okay.
- 20:16:06So equals to one
- 20:16:09that's X for X in for X in
- 20:16:16CV scores.
- 20:16:18Okay. Okay, list I will write MSE dot
- 20:16:24index
- 20:16:26minimum
- 20:16:28MSE.
- 20:16:30I will use this bracket square bracket.
- 20:16:33Okay. So then I will write print
- 20:16:37the best
- 20:16:38optimal
- 20:16:40number of
- 20:16:45neighbors
- 20:16:46is person
- 20:16:53best K. Okay, let's run it. See the best
- 20:16:58optimal number of neighbors. Okay. N
- 20:17:04neighbors is 9. Okay. So in the K
- 20:17:08nearest neighbor KN algorithms. Okay. So
- 20:17:11K represent the number of neighbors that
- 20:17:13are considered when classifying a query
- 20:17:15point. Okay. See the if we'll classify
- 20:17:19this particular point you will get the
- 20:17:22six. Okay. 1 2 3 4 5 6. Okay. And uh let
- 20:17:26me show you. If you will classify this
- 20:17:28portion only, so you will get 1 2 3 4 5
- 20:17:326 points like this. Okay. So the best
- 20:17:35optimum number of neighbor is nine.
- 20:17:37>> What is Python? Python is a high-level
- 20:17:41object-oriented programming language
- 20:17:43developed by Guido Van Roum in 1989 and
- 20:17:46was first released in 1991.
- 20:17:49Python is often called a batteries
- 20:17:51included language due to its
- 20:17:54comprehensive standard library. A fun
- 20:17:56fact about Python is that the name
- 20:17:59Python was actually taken from the
- 20:18:00popular BBC comedy show of that time
- 20:18:03Montipython's Flying Circus. Now let's
- 20:18:07look at the top features of Python
- 20:18:12first. So Python has a simple structure
- 20:18:14and a clearly defined syntax. This
- 20:18:16allows the learners to pick up the
- 20:18:18language quickly. So it is easy to learn
- 20:18:20and use.
- 20:18:22Python can run on different operating
- 20:18:24systems such as Windows, Linux, and Mac,
- 20:18:27making it a portable language. It
- 20:18:29enables programmers to develop the
- 20:18:30software for several competing platforms
- 20:18:32by writing a program only once.
- 20:18:36Third, Python is freely available at the
- 20:18:38official website since it is open
- 20:18:41source. This means that source code is
- 20:18:43also available to the public.
- 20:18:45Now, Python uses an object-oriented
- 20:18:47approach that encapsulates code within
- 20:18:49objects.
- 20:18:51Python provides a collection of
- 20:18:53libraries for various tasks such as
- 20:18:54machine learning, web development, and
- 20:18:56data analysis. And finally, in Python,
- 20:19:00you don't need to assign the data type
- 20:19:02of the variable. When you assign some
- 20:19:04value to the variable, it automatically
- 20:19:06allocates the memory to the variable at
- 20:19:08runtime.
- 20:19:10Now, with that, let's move on to the
- 20:19:12uses of Python programming.
- 20:19:15So, Python programming language is used
- 20:19:17to develop desktop applications and
- 20:19:19build web applications too. It is
- 20:19:22popularly used in the field of data
- 20:19:24science, machine learning and artificial
- 20:19:26intelligence to analyze data, build
- 20:19:28predictive models and make business
- 20:19:30decisions. Python is also widely used in
- 20:19:32game development. Now, let's see some of
- 20:19:35the popular Python frameworks and
- 20:19:37libraries.
- 20:19:38Python can be used for web development
- 20:19:40using frameworks like Zango, Flask,
- 20:19:43Pyramid and Churi.
- 20:19:45Now you can build graphical user
- 20:19:47interfaces using libraries and
- 20:19:48frameworks such as Tkinter or just KER.
- 20:19:52You can also use PI GTK, PIQT or PYJS or
- 20:19:57Python JavaScript.
- 20:20:00Now, Python is also used to perform
- 20:20:01machine learning tasks using libraries
- 20:20:03such as TensorFlow, PyTorch,
- 20:20:06Scikitlearn, Mattplot Lib, and Scypi.
- 20:20:09You can also perform mathematical
- 20:20:11computations using numpy and pandas.
- 20:20:15Now, let's look at the best ids that you
- 20:20:17can use to write programs in Python and
- 20:20:19perform specific tasks. So, we have
- 20:20:22Jupyter notebook, which is part of the
- 20:20:24Anaconda distribution that is widely
- 20:20:26used these days. Even for our demo in
- 20:20:29this video, we'll be using Jupyter
- 20:20:30Notebook. I'll show you in a while. Then
- 20:20:32we have the visual code editor from
- 20:20:34Microsoft. This is also one of the
- 20:20:36preferred IDEs by learners and
- 20:20:39companies. Then we also have the popular
- 20:20:42text editor called Sublime Text editor.
- 20:20:45Then we also have PyCharm followed by
- 20:20:48Python and Spider as our top idees. Now
- 20:20:52let's look at the top companies that are
- 20:20:54using Python in our day-to-day work.
- 20:20:58So we have Google, Kora, Facebook, even
- 20:21:02Netflix, Spotify, and Instagram. Now
- 20:21:04there are other top product- based,
- 20:21:06service- based and startups that also
- 20:21:08use Python programming. So what really
- 20:21:11is Python programming language?
- 20:21:14Python is an object-oriented highle
- 20:21:16programming language that supports
- 20:21:18built-in data structures and dynamic
- 20:21:20semantics.
- 20:21:21It supports multiple programming
- 20:21:23paradigms such as structured,
- 20:21:25object-oriented and functional
- 20:21:26programming.
- 20:21:28Python is often described as batteries
- 20:21:30included language because it has a
- 20:21:32comprehensive collection of standard
- 20:21:34libraries. Python supports different
- 20:21:36modules and packages which allows
- 20:21:38program modularity and code reuse.
- 20:21:43Python was developed by Guido Van Rosum
- 20:21:45and its implementation started in
- 20:21:47December 1989.
- 20:21:50Python 1.0 version was released in the
- 20:21:53year 1994. Python 2.0 came out in
- 20:21:56October 2000 while Python 3.0 was
- 20:21:59released in December 2008.
- 20:22:03Now that you have got an understanding
- 20:22:05of the Python programming language,
- 20:22:07let's now look at the top 10 reasons why
- 20:22:09you should learn Python.
- 20:22:12So at number 10, we have ease of use.
- 20:22:16One of the most common reasons to like
- 20:22:18Python is that it is quite easy to learn
- 20:22:20and code. It provides a simple syntax
- 20:22:23that improves readability and makes it
- 20:22:25easier to understand. So developers can
- 20:22:27create any desktop or machine based
- 20:22:29application using this language. Python
- 20:22:31is very versatile and is instrumental in
- 20:22:33artificial intelligence and machine
- 20:22:35learning. We will talk about this later
- 20:22:38in the session. Compared to Java or C++,
- 20:22:41it has fewer lines of codes.
- 20:22:45In the example here, we are printing a
- 20:22:47hello world program in Java. As you can
- 20:22:50see,
- 20:22:52if you have to write a program in Java,
- 20:22:54you first have to declare the class name
- 20:22:57along with its scope.
- 20:23:02Next, using curly braces, you need to
- 20:23:05pass the main method along with its
- 20:23:07arguments. And then using
- 20:23:09system.out.print print len method you
- 20:23:12can print hello world that's quite a
- 20:23:14tedious task isn't it
- 20:23:19the same task of printing hello world
- 20:23:20can be done using just one line of code
- 20:23:22in python as shown here you can write
- 20:23:25the print function and pass whatever you
- 20:23:28want to display inside the brackets and
- 20:23:30that will print the output it is so
- 20:23:32simple
- 20:23:36that is why Python is considered as a
- 20:23:37highle language and it's open source
- 20:23:41You can just download it from the
- 20:23:42website and start using it.
- 20:23:49At nine, we have active community.
- 20:23:53You need a community to learn new
- 20:23:55technology and friends are your best
- 20:23:57asset when it comes to learning a
- 20:23:59programming language. Python has large
- 20:24:02community support.
- 20:24:06It has an extensive and active community
- 20:24:08to assist engineers, developers,
- 20:24:10analysts, and data scientists with
- 20:24:12expert support in case of programming
- 20:24:14errors or issues with the software. You
- 20:24:17can just go ahead and put your queries
- 20:24:19in the community forum. The community
- 20:24:21members will address your queries in
- 20:24:22real quick time. Communities like Stack
- 20:24:25Overflow also brings many Python experts
- 20:24:27together to help learners.
- 20:24:30Python enhancement proposals or PEP is
- 20:24:33where the proposals and the improvements
- 20:24:35are announced. Also, there are a set of
- 20:24:38recommendations or core values called
- 20:24:40the Zen of Python written by Tim Peters
- 20:24:43that represents the guiding principles
- 20:24:45for Python development.
- 20:24:50Up next at 8, we have portable and
- 20:24:52extensible.
- 20:24:53Multiple cross- language operations can
- 20:24:55be performed effectively because Python
- 20:24:58is portable and extensible in nature.
- 20:25:01For example, if the users have a Python
- 20:25:03code written on Windows and they want to
- 20:25:06execute on a Mac operating system or
- 20:25:08Linux operating system or Solaris, they
- 20:25:10can easily do it without any amendment.
- 20:25:13They can also run this code on any
- 20:25:15platform flawlessly and without any
- 20:25:17interrupt.
- 20:25:20Due to its extensibility feature, you
- 20:25:22can integrate other programming
- 20:25:24languages such as Java,Net, C and C++
- 20:25:27codes with Python. The components of
- 20:25:29other programming languages can be used
- 20:25:31with Python and thus it can be used to
- 20:25:33make a crossplatform suitable
- 20:25:35application too. So it is a really good
- 20:25:38feature that Python provides.
- 20:25:45The next reason to learn Python is
- 20:25:47testing frameworks.
- 20:25:51Python supports several built-in
- 20:25:53flawless testing tools and frameworks
- 20:25:55that help in debugging and speeding of
- 20:25:57workflows.
- 20:26:00Some of the tools and frameworks
- 20:26:02supported by Python are Piest, Selenium
- 20:26:04and Splinter. This is the reason for
- 20:26:06which every tester tries to use Python
- 20:26:09based tools and frameworks to test any
- 20:26:11application or code or to validate it in
- 20:26:13an easier manner. Piest is the most
- 20:26:16recommended testing framework for
- 20:26:17functional, integrational, and unit
- 20:26:19testing. You can run Selenium test
- 20:26:22scripts using Python programming
- 20:26:24language to automate various tasks. And
- 20:26:26Splinter is an open-source tool for
- 20:26:28testing web applications using Python.
- 20:26:31It lets you automate browser actions
- 20:26:33such as visiting URLs and interacting
- 20:26:35with their items.
- 20:26:39At number six, we have libraries and
- 20:26:41packages.
- 20:26:44Another reason why Python has become so
- 20:26:46popular in the industry these days is
- 20:26:48that it has a massive collection of
- 20:26:49libraries and packages that make your
- 20:26:51task simple and easy. It has a range of
- 20:26:54libraries, packages, frameworks, and
- 20:26:56modules for data manipulation,
- 20:26:58statistical calculation, web
- 20:27:00development, machine learning, and data
- 20:27:03science.
- 20:27:08Python programmers have developed tons
- 20:27:10of free and open-source libraries that
- 20:27:11you can use. You can find many of them
- 20:27:14via Python package index, the repository
- 20:27:17of Python software. Python provides the
- 20:27:19default package called pip. Anaconda is
- 20:27:22a third party Python ecosystem. Other
- 20:27:24examples include numpy, sci and zango.
- 20:27:31Then we have scripting and automation.
- 20:27:37Python is not just a programming
- 20:27:38language. It can also be used for
- 20:27:40writing scripts for automating tasks and
- 20:27:42workflows without human intervention.
- 20:27:47The code can be written in the form of
- 20:27:49scripts and executed later. Further, it
- 20:27:52is interpreted by the machine and
- 20:27:54checked for errors at runtime. The
- 20:27:56machine is used to read and interpret
- 20:27:58the code. Once the developer checks the
- 20:28:01code, it can further run or be used
- 20:28:03several times without any interruption.
- 20:28:06This allows you to automate a set of
- 20:28:08certain tasks within a program or the
- 20:28:10same code can be used with other
- 20:28:12applications as well.
- 20:28:16At number four, we have web development.
- 20:28:22Another reason to learn Python is that
- 20:28:24it makes the web development process so
- 20:28:26much easier.
- 20:28:30It provides a wide collection of
- 20:28:31frameworks that make it easier for
- 20:28:33developers to develop web applications.
- 20:28:36Some of the examples are Zango, Flask,
- 20:28:39Pyramid, Turbo Gears, CherryPie, etc.
- 20:28:42These frameworks are written in Python
- 20:28:44which makes the code a lot faster and
- 20:28:46stable.
- 20:28:48The task which used to take hours in PHP
- 20:28:50can be finished in minutes using Python.
- 20:28:53Python is also used for web scraping.
- 20:28:56Django offers many elements of intricate
- 20:28:59programs such as template design,
- 20:29:01management panel, signing in, signing
- 20:29:04up, signing out, URL routing, etc.
- 20:29:10Once the user establishes the framework,
- 20:29:12all these features become ready to use.
- 20:29:16Flask is a microwave framework written
- 20:29:18in Python.
- 20:29:21of all the components that are part of
- 20:29:23this module, they are all ready to
- 20:29:25execute in the server context.
- 20:29:27Pinterest and LinkedIn use Flask.
- 20:29:30Pyramid offers more attributes than
- 20:29:32Flask. It will assist users with URL
- 20:29:35routing and authentication support.
- 20:29:38Turbo Gears is a highly recommended and
- 20:29:41scalable framework that supports
- 20:29:42features such as authentication,
- 20:29:44caching, identification, management of
- 20:29:47sessions, and pluggable applications.
- 20:29:54Up next at number three, we have machine
- 20:29:57learning.
- 20:29:59The growth of machine learning has been
- 20:30:00phenomenal in the last 5 years and it's
- 20:30:03rapidly changing the world around us.
- 20:30:05Python is one of the most preferred
- 20:30:07programming languages for machine
- 20:30:08learning because of its simple syntax
- 20:30:10and support for several machine learning
- 20:30:12libraries.
- 20:30:15Using different libraries and functions
- 20:30:16in Python, the system can learn and
- 20:30:19train itself from past data.
- 20:30:23Once the system is trained, it can then
- 20:30:25learn to adjust itself to new inputs.
- 20:30:29Finally, it can make predictions and
- 20:30:31perform humanlike tasks automatically.
- 20:30:37At number two, we have data science.
- 20:30:40Machine learning and data science go
- 20:30:42hand in hand. Python is robust, scalable
- 20:30:45and provides extensible visualization
- 20:30:47and graphics options. Hence, it is
- 20:30:49widely used in data science.
- 20:30:52Python has libraries such as numpy for
- 20:30:55numerical computation of data, pandas
- 20:30:57for operations to manipulate data on
- 20:30:59numerical tables and time series. It
- 20:31:01also provides simply for symbolic
- 20:31:04computation and sci for technical and
- 20:31:06scientific computations.
- 20:31:08It has another library called pyrain
- 20:31:10which is sought for python based
- 20:31:12reinforcement learning, artificial
- 20:31:14intelligence and neural network library.
- 20:31:16Scikitlearn is the machine learning
- 20:31:17library for creating classification,
- 20:31:19regression and clustering algorithms.
- 20:31:21And finally, it provides PyTorch and
- 20:31:24TensorFlow for deep learning.
- 20:31:28Finally coming to the most important and
- 20:31:30the top reason to learn Python which is
- 20:31:32career opportunities and salary.
- 20:31:36Python language provides a variety of
- 20:31:38job opportunities and promises a high
- 20:31:40growth graph with huge salary prospects.
- 20:31:43It is been used by most of the tech
- 20:31:45giants.
- 20:31:47Industry leaders using Python are
- 20:31:49Amazon, Google, Facebook, IBM, NASA,
- 20:31:53Netflix and YouTube.
- 20:31:57Next, you can see the Google trends
- 20:32:01but I have considered three programming
- 20:32:02languages Python, Java and C++. I have
- 20:32:06compared them for the past 12 months.
- 20:32:09You can see it clearly on your screens
- 20:32:10that Python has become a frontr runner
- 20:32:12in terms of popularity and web search
- 20:32:14volume. It means people are interested
- 20:32:17in Python. They want to learn it and use
- 20:32:19it in their work. You can also check for
- 20:32:22the YouTube search.
- 20:32:25There also you will find that Python
- 20:32:26programming language is the most
- 20:32:28searched language on YouTube.
- 20:32:32Now on your screens you can see the
- 20:32:35report of PPL which is popularity of
- 20:32:38programming language index. It is
- 20:32:41created by analyzing how often language
- 20:32:43tutorials are searched on Google. It is
- 20:32:45a leading indicator. The raw data comes
- 20:32:48from Google trends. The bar graph
- 20:32:50depicts that Python is the most popular
- 20:32:53and widely used programming language
- 20:32:54across the globe followed by Java then
- 20:32:57JavaScript and C.
- 20:33:00The popularity of programming language
- 20:33:02index can help you decide which language
- 20:33:04to study or which one to use in a new
- 20:33:07software project.
- 20:33:12The next graph shows the popularity of
- 20:33:14Python and Java over the years starting
- 20:33:17from 2004 till the current period which
- 20:33:20is 2020. Worldwide, Python is the most
- 20:33:23popular language. Python grew the most
- 20:33:26in the last 5 years by 19.4%. 4% and
- 20:33:29Java lost the most by minus 7.2%.
- 20:33:37Now let's talk about the different
- 20:33:38career opportunities and the job roles
- 20:33:40that you can get into if you learn
- 20:33:42Python language.
- 20:33:44First, you can become a Python developer
- 20:33:47where you will be asked to write and
- 20:33:49test codes, debug programs, and
- 20:33:51integrate applications with third party
- 20:33:53web services.
- 20:33:55Second, you can become a web developer.
- 20:33:58Here you will be responsible for writing
- 20:34:01serverside web application logic. Python
- 20:34:03web developers usually develop back-end
- 20:34:05components, connect the application with
- 20:34:08third party services and support the
- 20:34:10front-end developers by integrating
- 20:34:12their work with the Python application.
- 20:34:16You can also become a data analyst if
- 20:34:17you know Python. As a data analyst, you
- 20:34:20have to gather data from multiple
- 20:34:22sources using scripts. analyze that
- 20:34:24data, develop and implement databases
- 20:34:27and data collection systems.
- 20:34:31You can become a data scientist. As a
- 20:34:34data scientist, you need to understand
- 20:34:36the challenges in business and come up
- 20:34:38with the best solutions using modern
- 20:34:40tools and techniques to analyze,
- 20:34:41visualize, and build prediction models
- 20:34:43to make business decisions.
- 20:34:46Lastly, you can be a machine learning
- 20:34:48engineer where you can develop
- 20:34:50intelligent machines that can learn from
- 20:34:52vast volumes of data and apply knowledge
- 20:34:54without human intervention.
- 20:34:56So there's a lot of scopes if you learn
- 20:34:58Python. But before we move on, let's
- 20:35:00understand first what is Jupyter
- 20:35:02Notebook. So guys, as you can see all
- 20:35:03over here that Jupyter Notebook is a
- 20:35:06popular open-source tool that basically
- 20:35:08allows you to create and share documents
- 20:35:11which contains codes, equations, you can
- 20:35:14have visualizations also. Basically, it
- 20:35:17is used for data analysis, machine
- 20:35:19learning and scientific research which
- 20:35:21makes it a very essential tools for
- 20:35:23developers like data scientists and
- 20:35:25researchers alike. Now before installing
- 20:35:28Jupyter notebook I request you that you
- 20:35:30have Python installed in your system. So
- 20:35:33the requirement should be Python 3.6 or
- 20:35:36greater. So now let us officially
- 20:35:38navigate to the Python's website. So
- 20:35:40guys as you can see all over here. So on
- 20:35:42python.org if I click on download
- 20:35:45Python. So we're going to see that all
- 20:35:47over here download Python 3.125. So as I
- 20:35:51already told you that the requirement of
- 20:35:53Python should be greater than 3.6. So
- 20:35:55just you can click all over here and you
- 20:35:58can see the download has started.
- 20:36:03So guys as you can see all over here
- 20:36:05that we have installed the Python. Now
- 20:36:07let us open the file. So you can see the
- 20:36:10given software is going to installed on
- 20:36:11this directory. Okay. So just click all
- 20:36:14over here. So guys as you can see all
- 20:36:17over here the Python installation of
- 20:36:203.125 is in progress. Let's wait for
- 20:36:23some time till it gets installed.
- 20:36:27So as you can see guys all over here
- 20:36:29that we have successfully installed our
- 20:36:31Python. Now let us open our terminal and
- 20:36:33let us check whether Python is correctly
- 20:36:35installed. So we are going to type
- 20:36:38python
- 20:36:40/ version.
- 20:36:42So as you can see all over here we have
- 20:36:44successfully installed our Python. So
- 20:36:46guys that was our prerequisite. Now
- 20:36:49there are two ways to install Jupyter
- 20:36:52notebook. The first one can be pip.
- 20:36:54Okay, pip is a package manager or using
- 20:36:58Anocanda distribution. So let us see
- 20:37:00with pip first. So guys, pip is a
- 20:37:03package manager which is used to install
- 20:37:04and manage software packages libraries
- 20:37:07written in Python. So you can see all
- 20:37:09over here that the Python with version
- 20:37:11greater than 3.6 have default pip
- 20:37:14installed in them. Okay. So we can use
- 20:37:16pip command to install our Jupyter
- 20:37:19notebook. So guys as you can see all
- 20:37:21over here we have come to the official
- 20:37:23documentation of jupitter.org and it is
- 20:37:26saying that installing Jupyter lab with
- 20:37:28pip command. So what you can do guys you
- 20:37:30can just copy all over here. You can go
- 20:37:32right all over here and click on this.
- 20:37:36Now as you can see all over here it has
- 20:37:38started downloading the Jupyter lab.
- 20:37:57So guys, we are going to install our
- 20:37:59Jupyter lab with the pip command. So
- 20:38:02this is the official documentation of
- 20:38:04Jupyter notebook. Okay? And just all you
- 20:38:07have to do is copy this and type on your
- 20:38:10terminal. So as you can see all over
- 20:38:12here it has started downloading the
- 20:38:14packages which is required to download
- 20:38:16the Jupyter notebook. Let us wait for
- 20:38:18some time.
- 20:38:23Okay guys, so we have successfully
- 20:38:25completed this step. Now let us move on
- 20:38:27to our next step. So as you can see all
- 20:38:30over here. So we have installed. Okay.
- 20:38:34Then what we have to do then you can
- 20:38:36type this. We can launch the Jupyter lab
- 20:38:39with this command on the terminal. Now
- 20:38:41let us wait. So as you can see all over
- 20:38:44here guys, we have successfully
- 20:38:46installed our Jupyter notebook. So you
- 20:38:48can go all over here and just create a
- 20:38:51new notebook and you can also choose
- 20:38:53your kernel and you can start working on
- 20:38:56your Jupyter notebook. Suppose I'll show
- 20:38:58you one snippet. So 3 + 5. Let us try to
- 20:39:02run this notebook. So as you can see it
- 20:39:05is giving us the eight as answer. So it
- 20:39:07is following the Python syntax and in
- 20:39:09this way we have successfully installed
- 20:39:12our Jupyter notebook using the pip
- 20:39:14command. So now as you can also see all
- 20:39:18over here you can also install Jupyter
- 20:39:20notebook with this command pip install
- 20:39:22notebook and then you can just open it.
- 20:39:24This is also an another alternative.
- 20:39:27Similarly, you can install with VA also
- 20:39:29same command and just open the VA. Now,
- 20:39:34if you are using any other operating
- 20:39:35system like Mac OS or Linux, then you
- 20:39:38can install by brew install Jupyter Lab.
- 20:39:41So, home will be the package manager for
- 20:39:44Mac OS and Linux. So, I hope so you are
- 20:39:47pretty clear with how to install Jupyter
- 20:39:49notebook with the pep command. Now, I
- 20:39:51have downloaded Anacondas from this
- 20:39:54official website. So as you can see all
- 20:39:57over here this is the official website
- 20:39:59of Anaconda. Okay. Now just type your
- 20:40:02email and you can just download it. So
- 20:40:04similarly as you can see after
- 20:40:07installing I'm going to launch my
- 20:40:09installer and let us click next. Okay.
- 20:40:12Let us click agree. Okay. And let us
- 20:40:16install this on the given directory.
- 20:40:20Let us wait for some time till the
- 20:40:21installation gets complete.
- 20:40:26So guys as you can see all over here we
- 20:40:29have completed our installation of
- 20:40:30Anoconda. So just click on finish and
- 20:40:35you can say we have successfully
- 20:40:36installed our Anocanda. Now let us open
- 20:40:39our Anaconda navigator. So just click
- 20:40:43on.
- 20:40:51So as you can see all over here just
- 20:40:54right click on this and our Anocanda
- 20:40:56navigator will be opened. So as you can
- 20:40:58see all over here this is our Anocanda
- 20:41:00navigator and it is loading the packages
- 20:41:03and for us to install the Jupyter
- 20:41:06notebook. So as you can see all over
- 20:41:08here just click on launch. So guys if
- 20:41:10you click on launch it is going to open
- 20:41:12our Jupyter notebook. So as you can see
- 20:41:14all over here it is saying launching the
- 20:41:17Jupyter notebook and it is hosted on
- 20:41:19localhost 8889. So this is our hosted
- 20:41:23Jupyter notebook and in similarly you
- 20:41:25can create a new notebook all over here
- 20:41:27and in this way you can start working
- 20:41:29>> LLMs. If you ever wondered how machine
- 20:41:32learning can now understand and generate
- 20:41:34humanlike text, you are in the right
- 20:41:36place. From chatboards like Chat GPT to
- 20:41:38AI assistant that powers search engines,
- 20:41:40LLMs are transforming how we interact
- 20:41:43with technology. One of the most
- 20:41:44exciting advancement in this space is
- 20:41:46Google's Gemini or OpenAI Charging large
- 20:41:50language model designed to push the
- 20:41:52boundaries of what AI can achieve. In
- 20:41:54this video, we will explore what LLMs
- 20:41:56are, how they work, and why models like
- 20:41:58Geminy are critical for the future of
- 20:42:01AI. Google Gemini is part of a new wave
- 20:42:04of AI models that are smarter, faster,
- 20:42:06and more efficient. It is designed to
- 20:42:08understand context better, offer more
- 20:42:11accurate responses and integrate deeply
- 20:42:13into service like Google search and
- 20:42:16Google Assistant, providing more
- 20:42:17humanlike interactions. So we will break
- 20:42:19down the science behind LLMs including
- 20:42:21their massive training data set,
- 20:42:23transformer architecture and how models
- 20:42:25like Gemini use deep learning innovation
- 20:42:28to change industries. Plus we will
- 20:42:30compare Google Gemini to other popular
- 20:42:32LMS such as OpenAI Chity models showing
- 20:42:35how each of these technologies is used
- 20:42:36to power chat bots, virtual assistants
- 20:42:39and other AIdriven application. By end
- 20:42:41of this video, you will have a clear
- 20:42:42understanding of how large language
- 20:42:44models like Gemini work, their key
- 20:42:46features, and what they mean for their
- 20:42:48future AI. Don't forget to like,
- 20:42:50subscribe, and hit the bell icon to
- 20:42:52never miss any update from Simply Learn.
- 20:42:54So, what are the large language models?
- 20:42:56Large language models like Chargen
- 20:42:58pre-trained transformer 4 o and Google
- 20:43:02Gemini are sophisticated AI system
- 20:43:04designed to comprehend and generate
- 20:43:06humanlike text. These models are built
- 20:43:08using deep learning techniques and are
- 20:43:09trained on vast data set collected from
- 20:43:12the internet. They leverage self
- 20:43:13attention mechanism to analyze
- 20:43:15relationship between words or tokens
- 20:43:17allowing them to capture context and
- 20:43:19produce coherent relevant responses.
- 20:43:21LLMs have significant application
- 20:43:23including powering virtual assistant
- 20:43:25chatboards, content creation, language
- 20:43:27translation and supporting research and
- 20:43:29decision making. Their ability to
- 20:43:31generate fluent and contextually
- 20:43:32appropriate text has advanced natural
- 20:43:34language processing and improved human
- 20:43:36computer interaction. So now let's see
- 20:43:38what are large language model used for.
- 20:43:40Large language models are utilized in
- 20:43:42scenarios with limited or no domain
- 20:43:45specific data available for training.
- 20:43:46These scenarios include both few short
- 20:43:48and zero short training approaches which
- 20:43:50rely on the model's strong inductive
- 20:43:52bias and its capability to derive
- 20:43:54meaningful representation from a small
- 20:43:56amount of data or even no data at all.
- 20:43:59So now let's see how are large language
- 20:44:01models trained. Large language models
- 20:44:03typically undergo pre-training on a
- 20:44:06board. All encompassing data set that
- 20:44:08shares statical similarities with the
- 20:44:10data set specific to the target task.
- 20:44:12The objective of pre-training is to
- 20:44:14enable the model to require highlevel
- 20:44:16feature that can later be applied during
- 20:44:18the finetuning phase for specific task.
- 20:44:21So there are some training processes of
- 20:44:23LLM which involves several steps. The
- 20:44:25first one is text prep-processing. The
- 20:44:27textual data is transformed into a
- 20:44:28numerical representation that the LLM
- 20:44:30model can effectively process. This
- 20:44:32conversion may be involve techniques
- 20:44:34like tokenization encoding and creating
- 20:44:36input sequences. The second one is
- 20:44:38random parameter initialization. The
- 20:44:39model's parameter are initialized
- 20:44:41randomly before the training process
- 20:44:42begins. The third one is input numerical
- 20:44:44data. The numerical representation of
- 20:44:46the text data is fed into the model of
- 20:44:48processing. The model's architecture
- 20:44:50typically based on transformers allows
- 20:44:52it to capture the conceptual
- 20:44:53relationship between the words or tokens
- 20:44:56in the next. The fourth one is loss
- 20:44:58function calculation. A loss function
- 20:44:59calculation measures the discrepancy
- 20:45:01between the model's prediction and the
- 20:45:02actual next word or token in a syntax.
- 20:45:05The LLM model aims to minimize this loss
- 20:45:08during training. The fifth one is
- 20:45:09parameter optimization. The model's
- 20:45:11parameter are registered through
- 20:45:12optimization technique. This involves
- 20:45:14calculating gradient and updating the
- 20:45:16parameters accordingly gradually
- 20:45:18improving the model's performance. The
- 20:45:20last one is iterative training. The
- 20:45:21training process is repeated over
- 20:45:23multiple iteration or epox until the
- 20:45:26model's output achieve a satisfactory
- 20:45:28level of accuracy on that given task or
- 20:45:30data set. By following this training
- 20:45:32process, large language model learn to
- 20:45:34capture linguistic patterns, understand
- 20:45:35context and generate coherent responses
- 20:45:38enabling them to excel at various
- 20:45:39language related tasks. The next topic
- 20:45:42is how do large language models work. So
- 20:45:44large language models leverage deep
- 20:45:46neural network to generate output based
- 20:45:48on patterns learned from the training
- 20:45:50data. Typically a large language model
- 20:45:52adopts a transformer architecture which
- 20:45:54enables the model to identify
- 20:45:55relationship between words in a sentence
- 20:45:58irrespective of their position in the
- 20:45:59sequence. In contrast to RNAs that rely
- 20:46:02on recurrence to capture token
- 20:46:04relationship transformer neural network
- 20:46:07employ self attention as their primary
- 20:46:09mechanism. Self attention calculates
- 20:46:10attention scores that determine the
- 20:46:12importance of each token with respect to
- 20:46:14the other token in the text sequence
- 20:46:16facilitating the modeling of intricate
- 20:46:18relationship within the data. Next,
- 20:46:20let's see application of large language
- 20:46:22models. Large language models have a
- 20:46:24wide range of application across various
- 20:46:26domains. So here are some notable
- 20:46:27application. The first one is natural
- 20:46:29language processing NLP. Large language
- 20:46:31models are used to improve natural
- 20:46:33language understanding tasks such as
- 20:46:35sentiment analysis, named entity
- 20:46:37recognition, text classification, and
- 20:46:39language modeling. The second one is
- 20:46:41chatbot and virtual assistant. Large
- 20:46:43language models power conversational
- 20:46:45agents, chatbots, and virtual assistant
- 20:46:47providing more interactive and humanlike
- 20:46:49user interaction. The third one is
- 20:46:51machine translation. Large language
- 20:46:53models have been used for automatic
- 20:46:55language translation enabling text
- 20:46:57translation between different languages
- 20:46:59with improved accuracy. The fourth one
- 20:47:01is sentiment analysis. LLMs can analyze
- 20:47:03and classify the sentiment or emotion
- 20:47:05expressed in a piece of text which is
- 20:47:08valuable for market research, brand
- 20:47:10monitoring and social media analysis.
- 20:47:12The fifth one is content recommendation.
- 20:47:14These models can be employed to provide
- 20:47:15personalized content recommendations
- 20:47:17enhancing user experience and engagement
- 20:47:20on platforms such as news website or the
- 20:47:22streaming services. So these application
- 20:47:24highlight the potential impact of large
- 20:47:25language models in various domains for
- 20:47:27improving language understanding
- 20:47:29automation. So hello guys welcome to
- 20:47:31this demo part of this video. So here
- 20:47:35what I will do I will go to new then
- 20:47:37Python 3 file
- 20:47:40then here
- 20:47:42I will give it the name called
- 20:47:46exploratory
- 20:47:54data
- 20:47:56analysis. Basically we will so we have
- 20:48:00one data set file of
- 20:48:04roller coaster basically. So we will be
- 20:48:07using that and you can download that
- 20:48:09file from the description box below from
- 20:48:11the below link driving link. Okay. So we
- 20:48:14will be doing some small basic functions
- 20:48:17using Python and later on we will uh
- 20:48:20make some good charts. Okay. We will
- 20:48:22remove duplicates and all we will do all
- 20:48:24that
- 20:48:26thing. We'll do data preparation. We'll
- 20:48:28do feature engineering. Okay. And uh
- 20:48:32we'll remove the duplicates. We'll check
- 20:48:35for the duplicates. We'll make charts
- 20:48:38like histogram, KD blocks, box plot and
- 20:48:43like many more similar to that like heat
- 20:48:46map. We will make scatter plot group by
- 20:48:48comparison. Okay, we all do that, right?
- 20:48:52So just stick with me and you can write
- 20:48:56side along with me here this code. Okay.
- 20:49:00Okay. So let's start with importing
- 20:49:05pandas first.
- 20:49:07I guess everyone know what is pandas and
- 20:49:09numpy
- 20:49:11fine
- 20:49:13as np then I will import
- 20:49:17numpy
- 20:49:21as np why this is np and pd sorry my bad
- 20:49:26pd is here because I don't want to write
- 20:49:29again and again this pandas this numpy
- 20:49:32so basically I can write this small
- 20:49:34version okay Yeah. So if you guys don't
- 20:49:38know what is panda. So panda is very
- 20:49:40popular library for working with data.
- 20:49:43Okay. It's goal is to be the most
- 20:49:46powerful and flexible opensource tool.
- 20:49:48Okay. So it has reached that goal. So
- 20:49:51data frames are the center of pandas.
- 20:49:54And what is data frame? A data frame is
- 20:49:56structured like a table or a
- 20:49:57spreadsheet. Okay. The rows and the
- 20:49:59columns. Okay. Whereas numpy numpy is an
- 20:50:02open-source Python library again that
- 20:50:05facilitates uh you know efficient
- 20:50:07numerical operation on large quantities
- 20:50:09of data. Okay. So there are many
- 20:50:12functions in numpy as well those we can
- 20:50:17use in the pandas data frame. Got it. So
- 20:50:21we'll import one more
- 20:50:25dot piplot
- 20:50:28dotp. Okay. as
- 20:50:32plt. Okay, there is one more library
- 20:50:34mattplot lip for plotting the graph and
- 20:50:38and there is one more import
- 20:50:42seabbon
- 20:50:44as SNS. So what is se? Sebon is again
- 20:50:47the Python data visualization library
- 20:50:49based on Matt plot lip. Okay, it
- 20:50:51provides a highle interface for drawing
- 20:50:53attractive and informative statical you
- 20:50:56know graphs. Okay. Then I will write
- 20:50:59here plt dot style
- 20:51:03dot use.
- 20:51:06We'll write ggplot.
- 20:51:08Got it. ggplot.
- 20:51:12Fine.
- 20:51:15So yeah,
- 20:51:17let me run it. So what is ggplot? So
- 20:51:20ggplot is an again this is an opensource
- 20:51:23data visualization package for
- 20:51:25historical programming. Okay. So yeah uh
- 20:51:29you can say or a general scheme for data
- 20:51:32visualization which breaks up graph into
- 20:51:34you know semantic components such as
- 20:51:36scales and layers. Okay. ML lip py
- 20:51:42pyab
- 20:51:46p. Okay. Yeah. So now
- 20:51:50we will import our data set df. DF means
- 20:51:53data frame. You can write any word of
- 20:51:56your choice. Then pd again pandas dot
- 20:52:00read
- 20:52:01csv used for readings CSV file. Okay.
- 20:52:05Excel files. Got it? Then here I will
- 20:52:09write my
- 20:52:11this coaster
- 20:52:14dot CSV. Okay. I'm not writing any path
- 20:52:17because my this uh data set is here
- 20:52:22itself. Coaster Coaster CO this is
- 20:52:26coaster.csv CSV. Okay. If you have your
- 20:52:30data set in another location as in like
- 20:52:32in C drive, D drive or whatever. Okay.
- 20:52:34You can give that path.
- 20:52:37Okay. Let me run it. Yeah. So now we
- 20:52:41will do some data understanding. Okay.
- 20:52:44We'll see data frame shape head and tail
- 20:52:47data types and describe like small small
- 20:52:50function. We will use DF dot shape.
- 20:52:55Right? So we have 1087 rows and 56
- 20:52:59columns in our data set. Okay. Then we
- 20:53:04will see df.hat
- 20:53:06five.
- 20:53:10So df do.head means it will give me top
- 20:53:13five rows of my data set. Okay. You can
- 20:53:15see 1 2 3 4 5. Okay. Five rows and 56
- 20:53:19columns. Here 56 column but we have 1087
- 20:53:23rows. Okay. So we have coaster name,
- 20:53:26length, speed, location, status, opening
- 20:53:28date, type, this, this, this, this.
- 20:53:30Okay, we will do
- 20:53:34uh you know we'll make some graphs using
- 20:53:36this these columns. Okay, and there is
- 20:53:39one more df.tail
- 20:53:42again last five rows. So df.tail gives
- 20:53:45you the last last five rows. See 1086
- 20:53:481085. Okay. And if you want to see
- 20:53:53full data, it is here. Okay. 0 to 108.
- 20:53:59Fine. Yeah. So we have one more df doc
- 20:54:04columns to check
- 20:54:07all the columns.
- 20:54:09See coaster name, land, the speed,
- 20:54:11location, status, opening date, type and
- 20:54:14these all are my columns names. Fine. So
- 20:54:18we have one more data types. Actually we
- 20:54:22have like many data types but let me
- 20:54:24show you some important ones or you can
- 20:54:27say some the basics on one. Okay. So DF
- 20:54:31dot D types. Dypes means data types.
- 20:54:33Coaster name is object type length
- 20:54:35object and inversion is float. Okay.
- 20:54:39We'll check inversion.
- 20:54:43Where is inversion? Yeah float type.
- 20:54:45Okay,
- 20:54:47then everything is object and eer
- 20:54:50introduces int numeric one and latitude
- 20:54:54float. Okay, so basically we have three
- 20:54:58types float, int and object.
- 20:55:01Okay, then what we will do? Let's just
- 20:55:05quickly check this describe.
- 20:55:11Okay. So what is the count of this
- 20:55:14inversion 932?
- 20:55:17Okay. It won't include any you know what
- 20:55:22is it empty cells. Okay. It won't count
- 20:55:24empty cells. Right. Mean of this
- 20:55:27particular table then standard deviation
- 20:55:29then mean minimum value 25% 70% 75% and
- 20:55:33max. It will describe you all this.
- 20:55:36Okay. And there is one more df.info info
- 20:55:40to get the info. See coaster name 1087
- 20:55:45value non null object then length 953
- 20:55:49null. Okay. Like this.
- 20:55:52Yeah. So now moving forward what we will
- 20:55:55do? We will do data preparation. Okay.
- 20:55:58What comes in data preparation like
- 20:56:00dropping irrelevant columns and rows
- 20:56:02which we don't want. Okay. Then second
- 20:56:05thing is identifying duplicates columns.
- 20:56:07Then third is renaming columns.
- 20:56:10Then we'll do some feature creations,
- 20:56:13right?
- 20:56:15So if you want to drop a column, okay,
- 20:56:18how you can drop?
- 20:56:20So you have to write just df dot drop.
- 20:56:23First I will give here stag data
- 20:56:28repage.
- 20:56:30Okay. Yeah. So how you can drop a
- 20:56:33column? Okay. DF.
- 20:56:36Then here I will write
- 20:56:40opening
- 20:56:43date=
- 20:56:47to 1.
- 20:56:50Okay.
- 20:56:52Opening. Okay. X is wrong.
- 20:56:58Instead of this
- 20:57:01maybe what is the opening date? Okay. O
- 20:57:03is capital here.
- 20:57:08Instead of this you can use double
- 20:57:11equals to
- 20:57:14okay axis is not defined.
- 20:57:18My bad. So what you have to do? You have
- 20:57:21to give this here and yeah.
- 20:57:26Okay. Next is not defined.
- 20:57:29Yeah. Fine. Okay. So as you can see here
- 20:57:35first I will show you this question name
- 20:57:38length speed location status and opening
- 20:57:41date is there fine.
- 20:57:44So if you will go here question name
- 20:57:48length speed location status nothing
- 20:57:50opening date is there right so this is
- 20:57:53how you can drop a table fine for a
- 20:57:56while I'm making this as a comment maybe
- 20:58:00in future
- 20:58:02upcoming you know making graph I'll
- 20:58:04leave this okay so yeah so I will write
- 20:58:08here df equals to
- 20:58:11df
- 20:58:14poster name,
- 20:58:17comma, location
- 20:58:21then comma
- 20:58:24status.
- 20:58:27Then here I will write manufacturer.
- 20:58:30Fine. Then again comma.
- 20:58:34Then I will give here.
- 20:58:38Okay. Year
- 20:58:42introduced
- 20:58:45then
- 20:58:46comma
- 20:58:48latitude.
- 20:58:50Wait I will tell you why I'm doing this.
- 20:58:52Then longitude
- 20:58:55latitude longitude
- 20:58:58then
- 20:59:03type main.
- 20:59:07Okay. Then
- 20:59:10opening
- 20:59:13date
- 20:59:15clean.
- 20:59:19Fine.
- 20:59:21Then speed into
- 20:59:25m/ hour r
- 20:59:29speed
- 20:59:31and comma what else then
- 20:59:38height
- 20:59:40foot
- 20:59:42inversion
- 20:59:45clean
- 20:59:49geforce
- 20:59:52clean dot copy. Yeah.
- 20:59:58So here I will write okay
- 21:00:02type in
- 21:00:05okay one more mistake is here
- 21:00:10fine.
- 21:00:12So now what I will do
- 21:00:15I will write here opening
- 21:00:22date
- 21:00:24clean equals to PD2
- 21:00:29dot date time
- 21:00:32DF
- 21:00:36opening
- 21:00:38date
- 21:00:40Clean fine.
- 21:00:44What happened?
- 21:00:50So what does this pd do to date time do?
- 21:00:54Okay. So it converts argument to data
- 21:00:58time date time. Okay. So this function
- 21:01:01you know you can say converts a scalar
- 21:01:04array like or series or data frame
- 21:01:06dictionary like to a pandas datetime
- 21:01:08object. Okay.
- 21:01:12So now let's rename the columns.
- 21:01:15Okay.
- 21:01:17Then to you know for better things DF
- 21:01:23equals to DF dot rename
- 21:01:26columns equals to
- 21:01:32poster
- 21:01:34name. Then
- 21:01:37poster
- 21:01:39name,
- 21:01:45year
- 21:01:48introduced then
- 21:01:51here
- 21:01:54introduced.
- 21:01:55Okay.
- 21:01:58Opening date clean to
- 21:02:05open
- 21:02:08date.
- 21:02:10Okay. Then I will write speed
- 21:02:16I then write caps speed.
- 21:02:26Got it? Then
- 21:02:30height
- 21:02:31into foot
- 21:02:34I will write it as
- 21:02:41catch 50.
- 21:02:44So why I'm writing this because this is
- 21:02:46a very good practice as a professional
- 21:02:48way or as a data analyst or as a
- 21:02:52business analyst whatever you are
- 21:02:53working on machine learning projects or
- 21:02:55whatever this is a good practice.
- 21:02:59Okay. So inverions
- 21:03:02in versions
- 21:03:06clean
- 21:03:08then should be like
- 21:03:11invers
- 21:03:25fine.
- 21:03:26Let me run it.
- 21:03:29Okay. Now let's check the column.
- 21:03:33Yeah. So now you can see our column name
- 21:03:36is changed.
- 21:03:38Fine.
- 21:03:41So I will write this na
- 21:03:46dot sum.
- 21:03:51So what does this is na dot sum do? So
- 21:03:54is na function returns a boolean value
- 21:03:56of you know true if the value is n and
- 21:03:59the false otherwise and the sum function
- 21:04:01returns the sum of the true values which
- 21:04:04equals to the number of n values in the
- 21:04:06column. So here we have zero and n here
- 21:04:10same the status 213
- 21:04:13here 217 5
- 21:04:16height foot is 916 okay so now what we
- 21:04:19will do we will write here df dot
- 21:04:25location
- 21:04:27df dot duplicate
- 21:04:31I told you in starting
- 21:04:34we'll
- 21:04:35duplicate
- 21:04:37Okay.
- 21:04:39So, question name, location, status,
- 21:04:41man. Okay. It's duplicated now. Okay.
- 21:04:47So, now what we will do? I will check
- 21:04:50duplicates for the coaster name. So, you
- 21:04:52can write df. Loc
- 21:04:56then df
- 21:04:58df dot
- 21:05:01duplicated and subset equals to
- 21:05:06poster
- 21:05:08name then dot head. Okay. Head of five.
- 21:05:14Now everyone know right what do
- 21:05:18what this head does.
- 21:05:21Okay. Yeah. So here you can see so these
- 21:05:26are duplicate okay of the question name
- 21:05:30right.
- 21:05:32Why? because you can just check the
- 21:05:33typing
- 21:05:35and all. Fine. So now checking with an
- 21:05:39example duplicate. Let's check
- 21:05:42dot query
- 21:05:47question name
- 21:05:53crystal.
- 21:05:56Now you will get each equals cyclone.
- 21:06:00Okay. I took this crystal B cipher.
- 21:06:03Fine.
- 21:06:05Then run it. So now you can see 39 and
- 21:06:0943 are the same.
- 21:06:12Okay. Everything is same. So this one is
- 21:06:14duplicate. So now what I will do? I will
- 21:06:17write here. Just let me give some space.
- 21:06:19Yeah. DF dot columns.
- 21:06:24Then I will write here DF
- 21:06:29dot location.
- 21:06:32Then TF dot
- 21:06:35duplicated
- 21:06:39subset
- 21:06:41question
- 21:06:43name
- 21:06:46then
- 21:06:48location
- 21:06:52then
- 21:06:56opening
- 21:06:58date.
- 21:07:00Okay.
- 21:07:02dot
- 21:07:04d set
- 21:07:07index
- 21:07:09then drop that.
- 21:07:12Okay.
- 21:07:18So now what we will do we will do some
- 21:07:19feature understanding.
- 21:07:21Okay.
- 21:07:24So now we will do some feature
- 21:07:26understanding.
- 21:07:32Okay.
- 21:07:34So in this we will plot some feature
- 21:07:37distribution like histogram KD box plot.
- 21:07:39Okay. For that I will write here DF
- 21:07:45year introduced.
- 21:07:49Okay. Then I will write value
- 21:07:54count.
- 21:07:56So now what I will do? Let's create bar
- 21:08:01chart. Okay.
- 21:08:04Ax goals to der
- 21:08:10introduced dot
- 21:08:13value
- 21:08:16counts. Okay. Then dot
- 21:08:21add 10 max. I need
- 21:08:27then dot plot kind equals to I need bar
- 21:08:33then title is top
- 21:08:38I write what what what should we give
- 21:08:40top 10
- 21:08:43years
- 21:08:45coasters introduced.
- 21:08:49Okay, then
- 21:08:53I will write here ex dot set
- 21:08:59X label. Then I will write here
- 21:09:06introduced.
- 21:09:09Okay. Then I will write here ex dot
- 21:09:14set Y label.
- 21:09:18than
- 21:09:20count. Okay, let me run it. Okay,
- 21:09:24unexpected character after line
- 21:09:27continuation.
- 21:09:28Okay,
- 21:09:33we can remove this
- 21:09:36unexpected intended.
- 21:09:39Now let me run it. Yeah, so this is our
- 21:09:42bar plot. Okay. So this is how you can
- 21:09:46create bar plot using your data. Okay.
- 21:09:49So a bar plot is you know one of the
- 21:09:53most common types of graphics in this
- 21:09:55data visualization or EDA. It shows the
- 21:09:58relationship between a numerical and the
- 21:10:00categorical variable. Okay. So each
- 21:10:03entity of this categoric variable is
- 21:10:07represented as a bar.
- 21:10:10Got it? So now let's do some more. So
- 21:10:15what uh let's make a stogram class.
- 21:10:19Okay,
- 21:10:21histogram.
- 21:10:23So for that I will write
- 21:10:26a ex = to df
- 21:10:31speed
- 21:10:33r then plot
- 21:10:37kind equals to
- 21:10:44Then comma I will write bins here. Okay.
- 21:10:48Bins equals to 20
- 21:10:53then
- 21:10:55title
- 21:10:58coaster
- 21:10:59speed.
- 21:11:01Okay. meter per hour then AX
- 21:11:06dot set
- 21:11:09X level
- 21:11:12speed
- 21:11:15okay forgot to give this
- 21:11:18okay yeah so histogram displays
- 21:11:22numerical data by grouping data into
- 21:11:24bins of equal width so here each bin is
- 21:11:28plotted as a bar whose height correspond
- 21:11:30to how How many data points are in that
- 21:11:32bin? So bins you can say are also
- 21:11:35sometimes called the intervals or
- 21:11:36classes or buckets. Basically
- 21:11:39it's the same.
- 21:11:42So here for this this much is the bin.
- 21:11:44For this this much is a bin. This is
- 21:11:46like that. Okay.
- 21:11:50So now let's create KD plot. So for that
- 21:11:55it's very simple. Ax = to DF.
- 21:12:00Then I can write as speed in me of R
- 21:12:04then plot
- 21:12:08kinda
- 21:12:13then title
- 21:12:16I can write the coaster
- 21:12:21speed
- 21:12:23then ex dot set
- 21:12:27X
- 21:12:30speed
- 21:12:33can.
- 21:12:35Yeah. So, KD, what does KD means? A
- 21:12:39kernel density estimate. This plot is a
- 21:12:42method of, you know, visualization the
- 21:12:45distribution of observation in a data
- 21:12:47set analog to a stola. So, KD represent
- 21:12:51the data using a continuous probability
- 21:12:53density curve in one or more dimension.
- 21:12:57Okay. In one or more dimension that
- 21:12:59curve fine. So now let's do some feature
- 21:13:03relationship. So in this we will make
- 21:13:05scatter plot, heat map correlation and
- 21:13:08pair plot. Okay. Or we can also do some
- 21:13:10group by comparison. Fine. So now let's
- 21:13:13first make
- 21:13:16scatter plot. Okay. So, df dot plot
- 21:13:22kind equals to
- 21:13:25scatter
- 21:13:27comma
- 21:13:28x = to
- 21:13:31speed
- 21:13:32m/ hour
- 21:13:35comma y = to
- 21:13:39height in foot.
- 21:13:42Fine.
- 21:13:44Then comma let's give the title
- 21:13:48equals to coaster speed
- 21:13:52versus
- 21:13:54height.
- 21:13:56Fine then pl do
- 21:14:02okay some error is there
- 21:14:05speed meter per hour. Okay. Okay. S
- 21:14:08should be capital and H should be
- 21:14:12capital.
- 21:14:14Yeah.
- 21:14:16So this is scatter plot speed versus
- 21:14:19height. This is a speed versus height.
- 21:14:22Okay. If the I can see here height is
- 21:14:26directly proportional to speed somewhat
- 21:14:29because if you can see the 350 is the
- 21:14:31speed less than
- 21:14:33120 km/h. No it's not like that. Okay
- 21:14:36fine.
- 21:14:38So a scatter plot identifies the
- 21:14:41possible relationship between change
- 21:14:43observed in two different sets of
- 21:14:44variable. Here the variables are height
- 21:14:47and the speed. Okay. It provides a
- 21:14:49visual and aesthetical uh you know means
- 21:14:51to test the strength of relationship
- 21:14:53between two variable. Fine. Now let's
- 21:14:59okay now let's make one more to give you
- 21:15:02better idea. Okay with legend.
- 21:15:06Now I will wait SNS dot scatter
- 21:15:11plot then x = to
- 21:15:15speed
- 21:15:18r
- 21:15:22y = to
- 21:15:27height into foot
- 21:15:32whatever you can say then hue will be
- 21:15:34there the year
- 21:15:38introduced.
- 21:15:41Okay. Then data equals to DF.
- 21:15:47Then ex= to set title.
- 21:15:53Set title. Then
- 21:15:56coaster
- 21:15:58speed versus
- 21:16:01height.
- 21:16:03then plt dot show.
- 21:16:07Yeah.
- 21:16:08So now you can see
- 21:16:11so this color you know the light color I
- 21:16:14will let me zoom it. Okay. So this color
- 21:16:17we have some here which are introduced
- 21:16:20in '90s. Okay. And this color which are
- 21:16:23introduced in 1925 and these color which
- 21:16:26are introduced in 2000. Okay. So you can
- 21:16:30see like this as well.
- 21:16:32So now let's make pair plot. Okay, it's
- 21:16:37look amazing.
- 21:16:40So let me write SNS dot
- 21:16:44dot pair plot then TF comma
- 21:16:51variable is equals to I need
- 21:16:56introduce
- 21:16:57then speed
- 21:17:01meter per R
- 21:17:04comma S should be capital height
- 21:17:09foot. Just remember the spelling, okay?
- 21:17:13It's case sensitive.
- 21:17:15Inversion
- 21:17:18inversions, comma,
- 21:17:21inversions, comma, geforce.
- 21:17:26Okay. Comma. U I will put what? What?
- 21:17:30What? What? Okay. I will put type
- 21:17:36main.
- 21:17:37Fine then pl do
- 21:17:42okay let me run it okay some error is
- 21:17:46there
- 21:17:48year introduced
- 21:17:51y capital
- 21:17:55let's say here so that's why I'm saying
- 21:17:58just remember the proper spelling again
- 21:18:01some
- 21:18:04error in versions is version.
- 21:18:11Now let me run it.
- 21:18:14Okay. Again some
- 21:18:21what inversions?
- 21:18:23Okay. Let me check here while renaming
- 21:18:28inversions only.
- 21:18:31Okay. Fine.
- 21:18:33in versions.
- 21:18:37Let me paste in but nothing changes. Let
- 21:18:41me check again.
- 21:18:44Okay, some error is there.
- 21:18:48Wait. So yeah, you can see it's running
- 21:18:52fine.
- 21:18:54So after this what I will do?
- 21:18:58So what is the pair plot? So pair plot
- 21:19:00function allows the user to get you know
- 21:19:02then an axis grid via which each
- 21:19:05numerical variable is stored in the data
- 21:19:06is shared across the x and the y okay in
- 21:19:09the structured column.
- 21:19:12So this is how see you know type mean
- 21:19:14wood other and steel this red one is
- 21:19:17wood other or the blue one and the
- 21:19:20purple are the steel one okay the
- 21:19:24different different format okay now
- 21:19:26let's create the last graph which is
- 21:19:31heat map okay so let me write df
- 21:19:35correlation equals to df
- 21:19:39here introduced,
- 21:19:46comma,
- 21:19:48speed m/a
- 21:19:55height
- 21:19:57into foot
- 21:19:59canions,
- 21:20:09comma
- 21:20:10G force
- 21:20:15then
- 21:20:16drop then coalition okay DFO
- 21:20:29introduced I
- 21:20:36capital yeah
- 21:20:39so So this is correlation values of all
- 21:20:41the things right. So now let's write SNS
- 21:20:45dot
- 21:20:47heat map
- 21:20:50heat map TF C
- 21:20:55not will be true.
- 21:20:59Okay.
- 21:21:01Now let me run it. Yeah. So to create
- 21:21:05heat map in Python so you can use this.
- 21:21:07Okay. C bond library for this heat map.
- 21:21:11So this function takes a data frame you
- 21:21:13know as a input and generates a heat map
- 21:21:16type of things as the output. Okay. So
- 21:21:20this is how you can perform EDA using
- 21:21:22any data set or you can show your data
- 21:21:26or insights with a beautiful
- 21:21:28representation using graphs and all like
- 21:21:31this. Fine. Web scraping is a powerful
- 21:21:34technique that allows you to
- 21:21:36automatically extract data from website.
- 21:21:38Turning the vast amount of information
- 21:21:40available online into something you can
- 21:21:43easily analyze and use. Whether you are
- 21:21:45gathering data for research, building a
- 21:21:47data set for machine learning project,
- 21:21:48or just curious about how websites work
- 21:21:51behind the scenes. Web scraping is an
- 21:21:53essential skills to have in your
- 21:21:55toolkit. On the other hand, Python is
- 21:21:57one of the most popular programming
- 21:21:58languages for web scraping thanks to its
- 21:22:01simplicity and the wealth of libraries
- 21:22:03available. In this video, we will
- 21:22:05explore how to use Python to scrape data
- 21:22:07from website. And we will dive into
- 21:22:09practical examples using Python
- 21:22:10libraries like request and beautiful
- 21:22:13soup to fetch and parse web content. But
- 21:22:16it's not just about the code. Web
- 21:22:17scripping comes with its own set of
- 21:22:20challenges and ethical constitution. We
- 21:22:22will talk about how to scrape
- 21:22:24responsibly, respecting the rules set by
- 21:22:26websites and ensuring that your scraping
- 21:22:28activity don't negatively impact the
- 21:22:31sites you are collecting data from. So
- 21:22:33by the end of this video, you will have
- 21:22:35a solid understanding of how to start
- 21:22:37scraping data from the web using Python.
- 21:22:40Whether you are new to programming or
- 21:22:42looking to add web scraping to your
- 21:22:43skill set, this video will give you the
- 21:22:45knowledge and tools you need to get
- 21:22:47started. So let's jump in and see how
- 21:22:50Python can help you unlock the full
- 21:22:51potential of the web. So without any
- 21:22:53further ado, let's get started. So here
- 21:22:55I am using this Google Collab for the
- 21:22:57web scraping. Okay. You can use your own
- 21:23:01like Jupyter notebook, Visual Code
- 21:23:03Studio, any thing. Okay. So here I'll
- 21:23:06write scraping
- 21:23:09using Python.
- 21:23:12Okay. Then here first you have to
- 21:23:15install some libraries like
- 21:23:19you know uh request
- 21:23:22and you have to install beautiful. So
- 21:23:29you have to install that you know pandas
- 21:23:33because we will create one data frame
- 21:23:35and we will save it then we will check
- 21:23:38our data. Okay. And you can install some
- 21:23:42basic Python library like numpy and all
- 21:23:44that. Okay. So here first I will import
- 21:23:51request.
- 21:23:53Okay. Then I will write from PS4
- 21:23:59import
- 21:24:02beautiful
- 21:24:06soap.
- 21:24:09Okay.
- 21:24:11So
- 21:24:12then I will write import
- 21:24:15pandas
- 21:24:17as pay.
- 21:24:20Fine. Then now what I will do? Okay
- 21:24:24let's see what is this request and all
- 21:24:27the so request is an HTTP client library
- 21:24:30for the Python programming language. So
- 21:24:32request is one of the most you know
- 21:24:34downloaded Python libraries. Okay. It's
- 21:24:37like more like over 2,000 not exactly
- 21:24:412,000 sorry 200 or 300 million monthly
- 21:24:44download. Okay. So what it does it maps
- 21:24:46the HTTP protocol onto Python subject
- 21:24:49oriented semantics. And here beautiful
- 21:24:52soap. So
- 21:24:54SOAP is a Python package for you know
- 21:24:57parsing the HTML and XML documents
- 21:24:59including those with you know malphone
- 21:25:02markup. It creates a parse tree for
- 21:25:04documents that can be you know used to
- 21:25:06extract data HTML from HTML which is
- 21:25:10useful for web scraping. Let me run it
- 21:25:13here. Our second step will be
- 21:25:17define the URL and the headers. Okay,
- 21:25:21URL and the headers.
- 21:25:26So URL okay from where you want to
- 21:25:29extract your data. So I will write here
- 21:25:32simply learn
- 21:25:35then this okay let's open this PMP
- 21:25:37certification
- 21:25:39it yeah so here I will write this
- 21:25:44and headers
- 21:25:48equals to
- 21:25:53so these are the headers okay so like in
- 21:25:57which you know device or you're working
- 21:26:00on which browser you are working on. So
- 21:26:04these are for the headers. So now what
- 21:26:06what I will do I will send a get request
- 21:26:10to the simple page. Okay to this URL for
- 21:26:14that you know right let me write this
- 21:26:17sending
- 21:26:19get request.
- 21:26:23Okay, here write response
- 21:26:26equals to
- 21:26:28request
- 21:26:30dot get
- 21:26:33then URL
- 21:26:36comma headers
- 21:26:38equals to headers.
- 21:26:41Okay, let me run it. Okay, working fine.
- 21:26:45So now uh let's check if the request is
- 21:26:49successful or not. For that if
- 21:26:53response
- 21:26:55dot status code plus equals to = 200
- 21:27:03then
- 21:27:05so equals to
- 21:27:09false.
- 21:27:11Okay. Then response
- 21:27:17dot content
- 21:27:23do
- 21:27:26HTML
- 21:27:29dot parser. Okay.
- 21:27:34Here.
- 21:27:35So here I will write course
- 21:27:39titles. Why? because here I'm you know
- 21:27:43initializing the list to store our
- 21:27:46titles or whatever the things are okay
- 21:27:49basically the data okay why I'm writing
- 21:27:52course here because this is again course
- 21:27:54page that's why nothing else okay here I
- 21:27:57will give the empty list
- 21:28:00so now we will find all course titles
- 21:28:04based on the actual HTML structure what
- 21:28:06is HTML structure
- 21:28:08you have to go to the page right click
- 21:28:10Click then inspect.
- 21:28:14Okay. According to this HTML structure
- 21:28:17means what is the class name? What is
- 21:28:20the you know this PMP certification is
- 21:28:22which heading? H1 heading, H2 heading,
- 21:28:25which heading it is. Okay. So let's see
- 21:28:30which heading it is. Okay. Okay. This
- 21:28:34PMP certification training H1 heading.
- 21:28:37Okay. Just remember H1 heading.
- 21:28:41So here I will write for
- 21:28:45quotes in soap
- 21:28:48dot find
- 21:28:52or
- 21:28:55h1.
- 21:28:57Okay.
- 21:29:00Then title
- 21:29:05custom course text
- 21:29:09dot strip.
- 21:29:12Okay.
- 21:29:14Then here I'm write course
- 21:29:19titles
- 21:29:20dot append
- 21:29:24title. Okay.
- 21:29:27Now what I will do? I will check
- 21:29:31if any data was extracted before it not.
- 21:29:35Okay, I will write if course
- 21:29:40titles.
- 21:29:43Here I will create a data frame. DF
- 21:29:46equals to PD dot data frame.
- 21:29:53In this I will write course title. it
- 21:29:56will be our you know uh that column
- 21:29:59name. So for I will write it course
- 21:30:01data. Okay.
- 21:30:04Then
- 21:30:06course
- 21:30:09titles.
- 21:30:12Okay.
- 21:30:14Then here I will write df
- 21:30:18dot to
- 21:30:22csv.
- 21:30:24Then uh just create simply learn dot
- 21:30:28CSV. Okay. So our data will be saved in
- 21:30:31this simply learn dot csv. Okay. CSV.
- 21:30:35Fine. Here what I will do I will write
- 21:30:39index equals to false.
- 21:30:44Print
- 21:30:46data.
- 21:30:48Print data is saved in CSV.
- 21:30:53Simply learn dot CSV.
- 21:30:57Fine.
- 21:30:59So here I will write
- 21:31:03else
- 21:31:07print
- 21:31:09web page.
- 21:31:13Okay.
- 21:31:15write field
- 21:31:17to
- 21:31:21retrieve the web page. Okay, that's it.
- 21:31:27Let me run it.
- 21:31:30Okay, here if response storage to this
- 21:31:34group
- 21:31:36spawn object has no attribute
- 21:31:39status code. Okay.
- 21:31:46Okay. Title.
- 21:31:52It is text.
- 21:31:54Some minor spelling mistakes are there.
- 21:31:59Okay. Dear.
- 21:32:03Sorry. Sorry. My bad
- 21:32:09again.
- 21:32:12Okay. Sorry.
- 21:32:16D will be capital, F will be capital.
- 21:32:20Yeah. So now you can see data is saved
- 21:32:22in simply dot CSV. So what I will do? I
- 21:32:25will write data equals to everyone know
- 21:32:28how to read CSV in Python. Read
- 21:32:33CSV.
- 21:32:35What was our file name? Simply learn
- 21:32:39CSV.
- 21:32:41Okay. Let me copy paste. Run it. Link
- 21:32:45fine data.
- 21:32:49Okay.
- 21:32:51PMP certification training data is fed.
- 21:32:54Okay. Here you can see H1 we gave and
- 21:32:59that's why PMP certification training
- 21:33:00came. Let's take what's in H2.
- 21:33:06Okay. In H2 there is leading premier PMI
- 21:33:10partner something is there. Okay. I will
- 21:33:12change here
- 21:33:14H1 to H2 then I will run it. Okay
- 21:33:20fine.
- 21:33:22See leading premier P P P P P P P P P P
- 21:33:24P P P P P P P P P P PMI. Okay, let me
- 21:33:26make it bigger. Yeah. So now you can see
- 21:33:30course data we mentioned. So
- 21:33:33if you will see this H1 didn't come why
- 21:33:37because we have mentioned H2 that's why.
- 21:33:40So these all are the H2.
- 21:33:44Okay. So what is this? I don't know.
- 21:33:46Okay. This is a time series graph. This
- 21:33:49is Google collaping. Okay. So this is
- 21:33:52how you can literally retrieve your data
- 21:33:56from any website. Okay. Just remember
- 21:33:59that some websites don't give you access
- 21:34:02to their you know for their scrapping
- 21:34:05just like Amazon don't give so you have
- 21:34:06to use at that time API and this is not
- 21:34:10ethical also to use someone's data okay
- 21:34:15without asking or whatever you can
- 21:34:16>> welcome to the neural network tutorial
- 21:34:19my name is Richard Kersner I'm with the
- 21:34:21simply learn team what's in it for you
- 21:34:24well today we're going to cover what is
- 21:34:25a neural network what can neural neural
- 21:34:28networks do, how does a neural network
- 21:34:31work, types of neural networks, and then
- 21:34:33we're going to jump into a use case to
- 21:34:35classify between the photos of dogs and
- 21:34:38cats, and we'll do that on the KAS with
- 21:34:40the TensorFlow in the back, but it's a
- 21:34:42Python script. So, that's always my
- 21:34:44favorite part is when we dive into the
- 21:34:45actual script. So, what is a neural
- 21:34:48network? So, hi guys. I heard you want
- 21:34:51to know what a neural network is. Here
- 21:34:52we have uh looks like he just went
- 21:34:54shopping at a red tag sale. My robots's
- 21:34:56back. So as a matter of fact, you have
- 21:34:58been using neural network on a daily
- 21:35:01basis. In today's world, it's just
- 21:35:03amazing how much we use our new
- 21:35:04technology. We're not even aware of it.
- 21:35:06When you ask your mobile assistant to
- 21:35:08perform a search for you, you know, like
- 21:35:10saying you're Google or Siri or whoever
- 21:35:12you use, Amazon Web, self-driving cars.
- 21:35:15So that's the newest thing coming out.
- 21:35:17They're just now trying to make those
- 21:35:18legal in different states in the US and
- 21:35:20around the world. Even in the UK, they
- 21:35:22now have self-driving cars going up and
- 21:35:24down the street. It's pretty amazing.
- 21:35:25These are all neural network driven.
- 21:35:27Computer games use it. A lot of computer
- 21:35:29games are driven by neural networks in
- 21:35:31the back end as part of the game system
- 21:35:34and how it adjusts to the players. And
- 21:35:36it's also used in processing the map
- 21:35:38images on your phone. So every time you
- 21:35:39do a navigation someplace and it opens
- 21:35:42it up, they now use neural networks to
- 21:35:44help you find the quickest way to get
- 21:35:45there. Neural network. A neural network
- 21:35:48is a system or hardware that is designed
- 21:35:51to operate like a human brain. In
- 21:35:53today's development, this is so
- 21:35:55important to understand because we don't
- 21:35:57have anything else to compare it to. I'm
- 21:35:59sure someday in the future, the computer
- 21:36:01will redefine or the neural network or
- 21:36:03the AI artificial intelligence will
- 21:36:05redefine what these mean. But as far as
- 21:36:08we can today's world, in today's
- 21:36:10commercial development, we have to
- 21:36:11compare it to what humans do. So, it's
- 21:36:13we want to compare and how it operates
- 21:36:15to a human brain and how it solves
- 21:36:17problems like a human does. What can a
- 21:36:19neural network do? And really, we're
- 21:36:22just going to dive in deeper to we just
- 21:36:24covered and look at other examples. So,
- 21:36:26what can a neural network do? Well,
- 21:36:28let's list out the things neural
- 21:36:30networks can do for you. Translate text.
- 21:36:32Boy, we got Google Translate and
- 21:36:34Microsoft has their own translate. They
- 21:36:37have some really cool. They actually
- 21:36:38have an earpiece. It's supposed to start
- 21:36:39translating as you talk. What a cool
- 21:36:42technology. What a cool time to live.
- 21:36:44Identify faces. Can you imagine all the
- 21:36:46uses for facial identification? In the
- 21:36:48case of our uh sample or our code that
- 21:36:51we're going to look at later, we'll be
- 21:36:52identifying dogs and cats. So, not quite
- 21:36:54as detailed as uh understanding whose
- 21:36:56face belongs to who. I I'm waiting for
- 21:36:58the Google glasses to come out so I can
- 21:37:00see who's who and the identify faces as
- 21:37:02I'm walking around. Have a little name
- 21:37:03tag over them. Not out there yet, but
- 21:37:05boy, we are close. We can identify the
- 21:37:07faces and they have all kinds of
- 21:37:08technologies to bring that information
- 21:37:10back to us. Recognize speech goes along
- 21:37:13with the translate text. So now as
- 21:37:16you're talking into your assistant, it
- 21:37:18can use that to do commands, turn lights
- 21:37:20on, all kinds of things you can do with
- 21:37:22recognizing speech. Read handwritten
- 21:37:24text. They're starting to translate all
- 21:37:26these old text documents that they've
- 21:37:28had in storage instead of doing it
- 21:37:30individually where somebody's going
- 21:37:31through each text by themsel in a room.
- 21:37:34Picture like an old Raiders of the Lost
- 21:37:36Arc theme where he's in the back, you
- 21:37:37know, archaeologist studying the text.
- 21:37:39Now it's fed into a computer. They take
- 21:37:41a picture. They even use neural networks
- 21:37:43to take a scroll that is so messed up
- 21:37:46that they can't undo the scroll and they
- 21:37:48x-ray it and then they use that x-ray to
- 21:37:51translate the text off of it without
- 21:37:53ever opening the scroll. I mean just way
- 21:37:56cool stuff they're starting to do with
- 21:37:57all this. And of course control robots.
- 21:38:00What would be a neural network without
- 21:38:01bringing in the robots? And we have our
- 21:38:03own favorite robot in the middle who
- 21:38:05goes to our red tag cell and goes
- 21:38:06shopping for us. So, you know, these are
- 21:38:08just a few of the wonderful things that
- 21:38:10neural networks are being applied to.
- 21:38:12It's such an infant stage technology.
- 21:38:15What a wonderful time to jump in. And
- 21:38:17there are a lot of other things it goes
- 21:38:18into. I mean, we could spend just
- 21:38:20forever talking about all the different
- 21:38:21applications from business to whatever
- 21:38:23you can even imagine. They're now
- 21:38:25applying neural networks to help us
- 21:38:26understand. So, now we talked a little
- 21:38:29bit about all the cool things you can do
- 21:38:30with a neural network. Let's dive in and
- 21:38:32say, how does a neural network work? So
- 21:38:35now we've come far enough to understand
- 21:38:36how neural network works. Let's go ahead
- 21:38:39and walk through this in a nice
- 21:38:40graphical representation. They usually
- 21:38:42describe a neural network as having
- 21:38:44different layers. And you'll see that
- 21:38:46we've identified a green layer, an
- 21:38:48orange layer, and a red layer. The green
- 21:38:50layer is the input. So you have your
- 21:38:52data coming in. It picks up the input
- 21:38:54signals and passes them to the next
- 21:38:56layer. The next layer does all kinds of
- 21:38:58calculations and feature extraction.
- 21:39:00It's called the hidden layer. A lot of
- 21:39:02times there's more than one hidden
- 21:39:04layer. We're only showing one in this uh
- 21:39:06picture, but we'll show you how it looks
- 21:39:07like in a more detail in a little bit.
- 21:39:10And then finally, we have an output
- 21:39:11layer. This layer delivers the final
- 21:39:14result. So the only two things we see is
- 21:39:16the input layer and the output layer.
- 21:39:18Now let's make use of this neural
- 21:39:20network and see how it works. Wonder how
- 21:39:22traffic cameras identify vehicles
- 21:39:24registration plate on the road to detect
- 21:39:26speeding vehicles and those breaking the
- 21:39:29law? They got me going through a red
- 21:39:30light the other day. Well, last month.
- 21:39:32That's like the horrible thing. They
- 21:39:33send you this picture of you and all
- 21:39:35your information because they pulled it
- 21:39:36up off of your license plate and your
- 21:39:38picture. I shouldn't have gone through
- 21:39:39the red light. So, here we are and we
- 21:39:41have an image of a car and you can see
- 21:39:43the license plate on there. So, let's
- 21:39:45consider the image of this vehicle and
- 21:39:46find out what's on the number plate. The
- 21:39:49picture itself is 28x 28 pixels and the
- 21:39:52image is fed as an input to identify the
- 21:39:54registration plate. Each neuron has a
- 21:39:57number called activation that represents
- 21:39:59the grayscale value of the corresponding
- 21:40:01pixel range. And we range it from zero
- 21:40:04to one. One for a white pixel and zero
- 21:40:06for a black pixel. And you can see down
- 21:40:08here we have an example where one of the
- 21:40:09pixels is registered as like 082.
- 21:40:12Meaning it's probably pretty dark. Each
- 21:40:14neuron is lit up when its activation is
- 21:40:16close to one. So as we get closer to
- 21:40:18black on white, we can really start
- 21:40:20seeing the details in there. And you can
- 21:40:22see again the pixel shows us one up
- 21:40:24there. It's like part of the car and so
- 21:40:26it lights up. So pixels in the form of
- 21:40:28arrays are fed to the input layer. And
- 21:40:30so we see here the pixels of a car image
- 21:40:32fed as an input. And you're going to see
- 21:40:33that the input layer which is green is
- 21:40:36one dimension while our image is
- 21:40:38two-dimension. Now when we look at our
- 21:40:40setup that we're programming in Python,
- 21:40:41it has a cool feature that automatically
- 21:40:43does the work for us. If you're working
- 21:40:45with an older neural network pattern
- 21:40:47package, you then convert each one of
- 21:40:49those rows so it's all one array. So
- 21:40:51you'd have like row one and then just
- 21:40:53tack row two onto the end. You can
- 21:40:54almost feed the image directly into some
- 21:40:56of these neural networks. The key is
- 21:40:58though is that if you're using a 28x 28
- 21:41:01and you get a picture of this 30x30,
- 21:41:03shrink the 30x30 down to fit the 28x 28.
- 21:41:06So you can't increase the number of
- 21:41:09input in this case green dots. It's very
- 21:41:11important to remember when you work on
- 21:41:12neural networks. And let's name the
- 21:41:14inputs x1 x2 x3 respectively. So each
- 21:41:17one of those represents one of the
- 21:41:18pixels coming in. And the input layer
- 21:41:20passes it to the hidden layer. And you
- 21:41:22can see here we now have two hidden
- 21:41:23layers in this image in the orange. And
- 21:41:26each one of those pixels connects to
- 21:41:28each one of those hidden layers. And the
- 21:41:31interconnections are assigned weights at
- 21:41:33random. So they get these random weights
- 21:41:35that come through. If x1 lights up, then
- 21:41:38it's going to be x1 times this weight
- 21:41:40going into the hidden layer. And we sum
- 21:41:42those weights. The weights are
- 21:41:43multiplied with the input signal and a
- 21:41:45bias is added to all of them. So as you
- 21:41:47can see here we have X1 comes in and it
- 21:41:49actually goes to all the different
- 21:41:51hidden layer nodes or in this case uh
- 21:41:53whatever you want to call them network
- 21:41:55setup the orange dots and so you take
- 21:41:57the value of X1 you multiply it by the
- 21:42:00weight for the next hidden layer. So X1
- 21:42:03goes to hidden layer 1 X1 goes to hidden
- 21:42:06layer two X1 goes hidden layer 1 node
- 21:42:09two hidden layer one node three and so
- 21:42:11on. And the bias a lot of times they
- 21:42:14just put the bias in as like another
- 21:42:16green dot or another orange dot and they
- 21:42:18give the bias a value one and then all
- 21:42:21the weights go in from the bias into the
- 21:42:23next node. So the bias can change. We
- 21:42:26always just remember that you need to
- 21:42:27have that bias in there. There's things
- 21:42:29that can be done with it. Generally most
- 21:42:31of packages out there control that for
- 21:42:33you so you don't have to worry about
- 21:42:34figuring out what the bias is. But if
- 21:42:37you ever dive deep into neural networks,
- 21:42:38you got to remember there's a bias or
- 21:42:40the answer won't come out correctly. The
- 21:42:42weighted sum of the input is fed as an
- 21:42:44input to the activation function to
- 21:42:46decide which nodes to fire. And for
- 21:42:48feature extraction, as a signal flows
- 21:42:50within the hidden layers, the weighted
- 21:42:52sum of inputs is calculated and is fed
- 21:42:54to the activation function in each layer
- 21:42:56to decide which nodes to fire. So here's
- 21:42:58our feature extraction of the number
- 21:43:00plate. And you can see these are still
- 21:43:01hidden nodes in the middle. And this
- 21:43:03becomes important. We're going to take a
- 21:43:05little detour here and look at the
- 21:43:06activation function. So, we're going to
- 21:43:08dive just a little bit into the math so
- 21:43:10you can start to understand where some
- 21:43:12of the games go on when you're playing
- 21:43:13with neural networks in your
- 21:43:15programming. So, let's look at the
- 21:43:16different activation functions before we
- 21:43:18move ahead. Here's our friendly red tag
- 21:43:20shopping robot. And so, one is a sigmoid
- 21:43:23function. And the sigmoid function which
- 21:43:25is 1 over 1 + e to the minus x takes the
- 21:43:28x value and you can see where it
- 21:43:31generates almost a zero and almost a one
- 21:43:34with a very small area in the middle
- 21:43:35where it crosses over and we can use
- 21:43:37that value to feed into another
- 21:43:40function. So if it's really uncertain it
- 21:43:42might have a 0.1 or 2 or 3 but for the
- 21:43:45most part it's going to be really close
- 21:43:46to one and really close to this case
- 21:43:48zero zero to one the threshold function.
- 21:43:50So if you don't want to worry about the
- 21:43:52uncertainty in the middle, you just say,
- 21:43:54"Oh, if x is greater than or equal to
- 21:43:56zero, if not, then uh x is zero." So
- 21:43:58it's either zero or one. Really
- 21:44:00straightforward. There's no in between
- 21:44:02in the middle. And then you have the
- 21:44:04what they call the reel relu function.
- 21:44:06And you can see here where it puts out
- 21:44:08the value, but then it says, well, if
- 21:44:10it's over one, it's going to be one. And
- 21:44:13if it's uh less than zero, it's zero. So
- 21:44:15it kind of just deadends it on those two
- 21:44:17ends, but allows all the values in the
- 21:44:19middle. And again, this like the sigmoid
- 21:44:21function allows that information to go
- 21:44:23to the next level. So it might be
- 21:44:24important to know if it's a 0.1 or a
- 21:44:26minus.1. The next hidden layer might
- 21:44:29pick that up and say, "Oh, this piece of
- 21:44:31information is uncertain or this value
- 21:44:33has a very low certainty to it." And
- 21:44:35then the hyperbolic tangent function.
- 21:44:37And you can see here it's a 1 - e to the
- 21:44:40-2x over 1 + e - 2x. And it's very much
- 21:44:44along the same theme, a little bit
- 21:44:46different in here in that it goes
- 21:44:47between minus one and one. So you'll see
- 21:44:49some of these it goes 0ero to one, but
- 21:44:51this one goes minus one to one. And if
- 21:44:53it's less than zero, it's, you know, it
- 21:44:55doesn't fire and if it's over zero, it
- 21:44:57fires. And it also still puts out a
- 21:44:59value. So you still have a value you can
- 21:45:01get off of that just like you can with
- 21:45:03the sigmoid function and the relu
- 21:45:04function. Very similar in use. And I
- 21:45:07believe the originally used to be
- 21:45:08everything was done in the sigmoid
- 21:45:10function. That was the most uh commonly
- 21:45:11used. And now they just kind of use more
- 21:45:13the reloo function. The reason is one,
- 21:45:15it processes faster because you already
- 21:45:17have the value and you don't have to add
- 21:45:20another compute the 1 / 1 + e to the
- 21:45:22minus x for each hidden node and the
- 21:45:25data coming off works pretty good as far
- 21:45:27as putting it into the next level. If
- 21:45:29you want to know just how close it is to
- 21:45:30zero, how close is it not to
- 21:45:32functioning, you know, is it minus.1
- 21:45:34minus.2 usually they're float values.
- 21:45:36You get like minus point minus.00138
- 21:45:39or something. So, you know, important
- 21:45:41information, but the Reu is most
- 21:45:42commonly used these days as far as the
- 21:45:44setup we're using. But you'll also see
- 21:45:46the sigmoid function very commonly used
- 21:45:48also. Now that you know what an
- 21:45:50activation function is, let's get back
- 21:45:53to the neural network. So, finally, the
- 21:45:55model would predict the outcome of
- 21:45:56applying a suitable activation function
- 21:45:58to the output layer. So, we go in here,
- 21:46:00we look at this, and we have the optical
- 21:46:02character recognition OCR is used on the
- 21:46:04images to convert it into a text in
- 21:46:06order to identify what's written on the
- 21:46:08plate. And as it comes out, you'll see
- 21:46:10the red node. And the red node might
- 21:46:12actually represent just the letter A. So
- 21:46:14there's usually a lot of outputs when
- 21:46:16you're doing text identification. We're
- 21:46:18not going to show that on here, but you
- 21:46:19might have it even in the order. It
- 21:46:21might be what order the license plates
- 21:46:23in. So you might have ABCDE E FG, you
- 21:46:26know, all the alphabet plus the numbers.
- 21:46:28And you might have the 1 2 3 4 5 6 7 8 9
- 21:46:3210 places. So it's a very large array
- 21:46:34that comes out. It's not a small amount
- 21:46:36of uh, you know, we show three dots
- 21:46:38coming in, eight hidden layer nodes, you
- 21:46:40know, two sets of four. We just show one
- 21:46:42red coming out. A lot of times this is
- 21:46:44uh, you know, 28 * 28. If you did 30 *
- 21:46:4730, that's, you know, 900 nodes. So 28
- 21:46:50is a little bit less than that uh, just
- 21:46:52on the input. And so you can imagine the
- 21:46:54hidden layer is just as big. Each hidden
- 21:46:57layer is just as big if not bigger. Then
- 21:46:58the output is going to be there's so
- 21:47:00many digits. You know, it's a lot.
- 21:47:01There's it's a huge amount of input and
- 21:47:03output. But we're only showing you just,
- 21:47:04you know, it' be hard to show in one
- 21:47:06picture. And so it comes up and this is
- 21:47:07what it finally gets out in the output
- 21:47:09as it identifies a number on the plate.
- 21:47:11And in this case, we have 08-d3858.
- 21:47:16Error in the output is back propagated
- 21:47:18through the network and weights are
- 21:47:20adjusted to minimize the error rate.
- 21:47:22This is calculated by a cost function.
- 21:47:24When we're training our data, this is
- 21:47:27what's used and we'll look at that in
- 21:47:28the code when we do the data training.
- 21:47:30So, we have stuff we know the answer to
- 21:47:33and then we put the information through
- 21:47:35and it says yes, that was correct or no,
- 21:47:38cuz remember we randomly set all the
- 21:47:40weights to begin with. And if it's
- 21:47:41wrong, we take that error. How far off
- 21:47:43are you? You know, are you off by is it
- 21:47:45if it was like minus one, you're just a
- 21:47:47little bit off. If it's like minus 300
- 21:47:50was your output, remember when we're
- 21:47:51looking at those different options, you
- 21:47:53know, hyperbolic or whatever, and we're
- 21:47:54looking at the could doesn't have an
- 21:47:57limit on top or bottom. it actually just
- 21:48:00generates a number. So if it's way off,
- 21:48:02you have to adjust those weights a lot.
- 21:48:04But if it's pretty close, you might
- 21:48:05adjust the weights just a little bit.
- 21:48:07And you keep adjusting the weights until
- 21:48:09they fit all the different training
- 21:48:10models you put in. So you might have 500
- 21:48:13training models and those weights will
- 21:48:15adjust using the back propagation. It
- 21:48:17sends the error backward. The output is
- 21:48:19compared with the original result and
- 21:48:21multiple iterations are done to get the
- 21:48:23maximum accuracy. So, not only does it
- 21:48:25look at each one, but it goes through it
- 21:48:27and just keeps cycling through these the
- 21:48:29data making small changes in the network
- 21:48:31until it gets the right answers. With
- 21:48:33every iteration, the weights at every
- 21:48:35interconnection are adjusted based on
- 21:48:38the error. We're not going to dive into
- 21:48:40that math because it is a differential
- 21:48:41equation and it gets a little
- 21:48:43complicated, but I will talk a little
- 21:48:45bit about some of the different options
- 21:48:46they have when we look at the code. So,
- 21:48:49we've explored a neural network. Let's
- 21:48:51look at the different types of
- 21:48:52artificial neural networks. And this is
- 21:48:55like the biggest area growing is how
- 21:48:57these all come together. Let's see the
- 21:48:59different types of neural network. And
- 21:49:02again, we're comparing this to human
- 21:49:03learning. So here's a human brain. I
- 21:49:06feel sorry for that poor guy. So we have
- 21:49:08a feed for forward neural network.
- 21:49:11Simplest form of a they call it a ann a
- 21:49:14neural network. Data travels only in one
- 21:49:17direction input to output. This is what
- 21:49:20we just looked at. So as the data comes
- 21:49:22in, all the weights are added, it goes
- 21:49:24to the hidden layer, all the weights are
- 21:49:26added, it goes to the next hidden layer,
- 21:49:27all the weights are added, and it goes
- 21:49:29to the output. The only time you use the
- 21:49:31reverse propagation is to train it. So
- 21:49:34when you actually use it, it's very
- 21:49:35fast. When you're training it, it takes
- 21:49:37a while because it has to iterate
- 21:49:38through all your training data. And you
- 21:49:40start getting into big data because you
- 21:49:42can train these with a huge amount of
- 21:49:44data. The more data you put in, the
- 21:49:45better trained they get. The
- 21:49:47applications vision and speech
- 21:49:49recognition actually they're pretty much
- 21:49:51everything we talked about a lot of
- 21:49:52almost all of them use this form of
- 21:49:55neural network at some level radio basis
- 21:49:57function neural network this model
- 21:50:00classifies a data point based on its
- 21:50:02distance from a center point. What that
- 21:50:05means is that you might not have
- 21:50:07training data. So you want to group
- 21:50:09things together and you create central
- 21:50:11points and it looks for all the things
- 21:50:13you know some of these things are just
- 21:50:14like the other. If you've ever watched
- 21:50:16the Sesame Street as a kid, that dates
- 21:50:18me. So, it brings things together and
- 21:50:20this is a great way if you don't have
- 21:50:21the right training model, you can start
- 21:50:23finding things that are connected you
- 21:50:24might not have noticed before.
- 21:50:25Applications power restoration systems.
- 21:50:29They try to figure out what's connected
- 21:50:30and then based on that they can fix the
- 21:50:33problem if you have a huge power system.
- 21:50:35and self-organizing neural network
- 21:50:38vectors of random dimensions are input
- 21:50:40to discrete map comprised of neurons. So
- 21:50:43they basically find a way to draw they
- 21:50:46call them they say dimensions or vectors
- 21:50:48or planes because they actually chop the
- 21:50:51data in one dimension, two dimension,
- 21:50:53three dimension, four, five, six. They
- 21:50:55keep adding dimensions and finding ways
- 21:50:56to separate the data and connect
- 21:50:58different data pieces together.
- 21:51:00Applications used to recognize patterns
- 21:51:02in data like in medical analysis. The
- 21:51:04hidden layer saves its output to be used
- 21:51:07for future prediction. Recurrent neural
- 21:51:09networks. So the hidden layers remember
- 21:51:11its output from last time and that
- 21:51:13becomes part of its new input. Uh you
- 21:51:16might use that especially in robotics or
- 21:51:18flying a drone. You want to know what
- 21:51:19your last change was and how fast it was
- 21:51:22going to help predict what your next
- 21:51:24change you need to make is to get to
- 21:51:25where the drone wants to go.
- 21:51:27Applications text to speech conversation
- 21:51:30model. So, you know, I talked about
- 21:51:31drones, but you know, just identifying
- 21:51:33on Lexus or Google Assistant or any of
- 21:51:36these, they're starting to add in I'd
- 21:51:38like to play a song on my Pandora, and
- 21:51:41I'd like it to be at volume 90%. So, you
- 21:51:44now can add different things in there,
- 21:51:45and it connects them together. The input
- 21:51:47features are taken in batches like a
- 21:51:50filter. This allows a network to
- 21:51:51remember an image in parts. Convolution
- 21:51:54neural network. today's world in photo
- 21:51:57identification and taking apart photos
- 21:51:59and trying to you know have you ever
- 21:52:00seen that on Google where you have five
- 21:52:02people together this is the kind of
- 21:52:04thing separates all those people so then
- 21:52:06it can do a face recognition on each
- 21:52:08person applications used in signal and
- 21:52:10image processing in this case I use
- 21:52:12facial images or Google picture images
- 21:52:14as one of the options modular neural
- 21:52:17network it has a collection of different
- 21:52:20neural networks working together to get
- 21:52:22the output so wow we just went through
- 21:52:24all these different types of neural
- 21:52:26networks. And the final one is to put
- 21:52:28multiple neural networks together. I
- 21:52:30mentioned that a little bit when we
- 21:52:32separated people in a larger photo and
- 21:52:34individuals in the photo and then do the
- 21:52:36facial recognition on each person. So
- 21:52:38one network is used to separate them and
- 21:52:40the next network is then used to figure
- 21:52:43out who they are and do the facial
- 21:52:44recognition. Applications still
- 21:52:46undergoing research. This is a cutting
- 21:52:48edge. you hear the term pipeline and
- 21:52:51there's actual in Python code and in
- 21:52:53almost all the different neural network
- 21:52:55setups out there they now have a
- 21:52:57pipeline feature usually and it just
- 21:52:59means you take the data from one neural
- 21:53:02network and maybe another neural network
- 21:53:04or you put it into the next neural
- 21:53:05network and then you take three or four
- 21:53:07other neural networks and feed them into
- 21:53:09another one. So how we connect the
- 21:53:11neural networks is really just cutting
- 21:53:13edge and it's so experimental. I mean
- 21:53:15it's almost creative in its nature.
- 21:53:17There's not really a science to it
- 21:53:19because each specific domain has
- 21:53:22different things it's looking at. So if
- 21:53:23you're in the banking domain, it's going
- 21:53:25to be different than the medical domain
- 21:53:27than the automatic car domain. And
- 21:53:29suddenly figuring out how those all fit
- 21:53:31together is just a lot of fun and really
- 21:53:33cool. So we have our types of artificial
- 21:53:35neural network. We have our feed forward
- 21:53:37neural network. We have a radial basis
- 21:53:39function neural network. We have our
- 21:53:40Cohen self-organizing neural network,
- 21:53:43recurrent neural network, convolution
- 21:53:46neural network, and modular neural
- 21:53:48network where it brings them all
- 21:53:49together. And u no the colors on the
- 21:53:52brain do not match what your brain
- 21:53:53actually does, but they do bring it out
- 21:53:56that most of these were developed by
- 21:53:58understanding how humans learn. And as
- 21:54:00we understand more and more of how
- 21:54:02humans learn, we can build something in
- 21:54:05the computer industry to mimic that, to
- 21:54:07reflect that. And that's how these were
- 21:54:09developed. So exciting part, use case
- 21:54:12problem statement. So this is where we
- 21:54:13jump in. This is my favorite part. Let's
- 21:54:15use the system to identify between a cat
- 21:54:18and a dog. If you remember correctly, I
- 21:54:19said we're going to do some Python code.
- 21:54:21And you can see over here, my hair is
- 21:54:24kind of sticking up over the computer,
- 21:54:25cup of coffee on one side, and a little
- 21:54:27bit of old school. A pencil and a pen on
- 21:54:29the other side. Yeah, most people now
- 21:54:31take notes. I love the stickies on the
- 21:54:33computer. That's great. That's that is
- 21:54:35my computer. I have sticky notes on my
- 21:54:37computer in different colors. So, not
- 21:54:39too far from uh today's programmer. So,
- 21:54:41the problem is is we want to classify
- 21:54:43photos of cats and dogs using a neural
- 21:54:46network. And you can see over here we
- 21:54:48have quite a variety of dogs in the
- 21:54:50pictures and cats and you know just
- 21:54:53sorting out it is a cat is pretty
- 21:54:55amazing. And why would anybody want to
- 21:54:56even know the difference between a cat
- 21:54:58and a dog? Okay, you know why? Well, I
- 21:55:00have a cat door. It'd be kind of fun
- 21:55:02that instead of it identifying, instead
- 21:55:05of having like a little collar with a
- 21:55:06magnet on it, which is what my cat has,
- 21:55:08the door would be able to see, oh,
- 21:55:10that's the cat. That's our cat coming
- 21:55:11in. Oh, that's the dog. We have a dog,
- 21:55:13too. That's a dog I want to let in.
- 21:55:15Maybe I don't want to let this other
- 21:55:16animal in cuz it's a raccoon. So, you
- 21:55:18can see where you could take this one
- 21:55:19step further and actually apply this.
- 21:55:21You could actually start a little
- 21:55:23startup company idea, self-identifying
- 21:55:25door. So, this use case will be
- 21:55:28implemented on Python. I am actually in
- 21:55:30Python 3.6. It's always nice to tell
- 21:55:33people the version of Python because
- 21:55:35that does affect sometimes which modules
- 21:55:37you load and everything. And we're going
- 21:55:38to start by importing the required
- 21:55:40packages. I told you we're going to do
- 21:55:42this in Kass. So we're going to import
- 21:55:44from KAS models sequential from the Kass
- 21:55:48layers conversion 2D or COV2D max
- 21:55:51pooling 2D flatten and dense. And we'll
- 21:55:55talk about what each one of these do in
- 21:55:56just a second. But before we do that,
- 21:55:58let's talk a little bit about the
- 21:56:00environment we're going to work in. And
- 21:56:01uh you know, in fact, let me go ahead
- 21:56:03and open a uh the website, KASS's
- 21:56:06website, so we can learn a little bit
- 21:56:07more about KASS. So here we are on the
- 21:56:09Kurass website, and it's uh ke.io.
- 21:56:14That's the official website for Kurass.
- 21:56:16And the first thing you'll notice is
- 21:56:17that Kurass runs on top of either
- 21:56:20TensorFlow, CNTK, and I think it's
- 21:56:23pronounced Thano or Theo. What's
- 21:56:25important on here is that TensorFlow and
- 21:56:27the same is true for all these, but
- 21:56:28TensorFlow is probably one of the most
- 21:56:30widely used currently packages out there
- 21:56:32with the KAS. And of course, you know,
- 21:56:34tomorrow this is all going to change.
- 21:56:36It's all going to disappear and they'll
- 21:56:37have something new out there. So, make
- 21:56:38sure when you're learning this code that
- 21:56:40you understand what's going on and also
- 21:56:42know the code. I mean, look, when you
- 21:56:44look at the code, it's not as
- 21:56:45complicated once you understand what's
- 21:56:46going on. The code itself is pretty
- 21:56:48straightforward. And the reason we like
- 21:56:50KAS and the reason that people are
- 21:56:52jumping on it right now, it's such a big
- 21:56:54deal is if we come down here, let me
- 21:56:56just scroll down a little bit. They talk
- 21:56:58about user friendliness, modularity,
- 21:57:00easy extensibility, work with Python.
- 21:57:02Python's a big one because a lot of
- 21:57:04people in data science now use Python,
- 21:57:06although you can actually access Kass
- 21:57:08other ways. Is if we continue down here
- 21:57:10is layers. And this is where it gets
- 21:57:12really cool. When we're working with
- 21:57:14KASS, you just add layers on. Remember
- 21:57:16those hidden layers we were talking
- 21:57:18about? And we talked about the reelu
- 21:57:21activation. You can see right here. Let
- 21:57:22me just up that a little bit in size.
- 21:57:24There we go. That's big. I can add in an
- 21:57:27eelu layer. And then I can add in a
- 21:57:29softmax layer in the next instance. We
- 21:57:31didn't talk about softmax. So you can do
- 21:57:33each layer separate. Now if I'm working
- 21:57:35in some of the other kits I use, I take
- 21:57:38that and I have one setup and then I
- 21:57:40feed the output into the next one. This
- 21:57:42one I can just add hidden layer after
- 21:57:44hidden layer with the different
- 21:57:45information in it which makes it very
- 21:57:47powerful and very fast to spin up and
- 21:57:50try different setups and see how they
- 21:57:52work with the data you're working on.
- 21:57:53And we'll dig a little bit deeper in
- 21:57:54here. And a lot of this is very much the
- 21:57:57same. So when we get to that part, I'll
- 21:57:58point that out to you also. Now just a
- 21:58:01quick side note, I'm using Anaconda with
- 21:58:03Python in it. And I went ahead and
- 21:58:05created my own package and I called it
- 21:58:07the Kass Python 36 because I'm in Python
- 21:58:0936. Anaconda is cool that You can create
- 21:58:12different environments really easily. If
- 21:58:13you're doing a lot of different
- 21:58:14experimenting with these different
- 21:58:16packages, probably want to create your
- 21:58:17own environment in there. And the first
- 21:58:19thing, as you can see right here,
- 21:58:21there's a lot of dependencies. A lot of
- 21:58:22these you should recognize by now if
- 21:58:24you've done any of these videos. If not,
- 21:58:26kudos for you for jumping in today. PIP,
- 21:58:28install, numpy, sci, the scikitlearn,
- 21:58:32pillow, and h5py
- 21:58:34are both needed for the tensorflow and
- 21:58:37then putting the kass on there. And then
- 21:58:39you'll see here uh and pip is just a
- 21:58:41standard installer that you use with
- 21:58:42Python. You'll see here that we did pip
- 21:58:44install TensorFlow since we're going to
- 21:58:46do KAS on top of TensorFlow. And then
- 21:58:48pip install and I went ahead and used
- 21:58:50the GitHub. So git plusgit and you'll
- 21:58:52see here github.com. This is one of
- 21:58:55their releases, one of the most current
- 21:58:56release on there that goes on top of
- 21:58:58TensorFlow. And you can look up these
- 21:58:59instructions pretty much anywhere. This
- 21:59:01is for doing it on Anaconda. Certainly
- 21:59:04you'd want to install these if you're
- 21:59:05doing it in Iuntu server setup. you
- 21:59:07you'd want to get I don't think you need
- 21:59:09the H5 py and aru but you do need the
- 21:59:11rest in there because they are
- 21:59:12dependencies in there and it's pretty
- 21:59:14straightforward and that's actually in
- 21:59:15some of the instructions they have on
- 21:59:17their website so you don't have to
- 21:59:18necessarily go through this just
- 21:59:19remember their website on there and then
- 21:59:21when I'm under my uh Anaconda navigator
- 21:59:24which I like you'll see where I have
- 21:59:26environments and on the bottom I created
- 21:59:28a new environment and I called it KAS
- 21:59:30Python 36 just to separate everything
- 21:59:32you can say I have Python 3.5 and Python
- 21:59:3536 I used to have a bunch of other ones,
- 21:59:37but it kind of cleaned house recently.
- 21:59:39And of course, once I go in here, I can
- 21:59:41launch my Jupyter Notebook, making sure
- 21:59:43I'm using the right environment that I
- 21:59:45just set up. This, of course, opens up
- 21:59:47my um in this case, I'm using uh Google
- 21:59:50Chrome. And in here, I could go and just
- 21:59:52create a new document in here. And this
- 21:59:54is all in your um browser window when
- 21:59:56you use the Anaconda. Do you have to use
- 21:59:58Anaconda and Jupyter Notebook? No. You
- 22:00:01can use any kind of Python editor,
- 22:00:03whatever setup you're comfortable with
- 22:00:05and whatever you're doing in there. So,
- 22:00:07let's go ahead and go in here and paste
- 22:00:09the code in. And we're importing a
- 22:00:11number of different settings in here. We
- 22:00:14have import sequential. That's under the
- 22:00:16models because that's the model we're
- 22:00:17going to use as far as our neural
- 22:00:19network. And then we have layers and we
- 22:00:21have conversion 2D, max pooling 2D,
- 22:00:24flatten dense. And you can actually just
- 22:00:27kind of guess at what these do. We're
- 22:00:29talking we're working in a 2D
- 22:00:31photograph. And if you remember
- 22:00:32correctly, I talked about how the actual
- 22:00:35input layer is a single array. It's not
- 22:00:37in two dimensions. It's one dimension.
- 22:00:39All these do is these are tools to help
- 22:00:41flatten the image. So, it takes a
- 22:00:43two-dimensional image and then it
- 22:00:44creates its own proper setup. You don't
- 22:00:46have to worry about any of that. You
- 22:00:48don't have to do anything special with
- 22:00:49the photograph. You let the carass do
- 22:00:51it. And we're going to run this. And
- 22:00:52you'll see right here they have some
- 22:00:53stuff that is going to be depreciated
- 22:00:55and changed because that's what it does.
- 22:00:56Everything's being changed as we go. You
- 22:00:58don't have to worry about that too much.
- 22:01:00If you have warnings, if you run it a
- 22:01:01second time, the warning will disappear.
- 22:01:03And this has just imported these
- 22:01:04packages for us to use. Jupiter's nice
- 22:01:07about this that you can do each thing
- 22:01:08step by step. And I'll go ahead and also
- 22:01:10zoom in there. A little control plus.
- 22:01:13That's one of the nice things about
- 22:01:14being in a browser environment. So, here
- 22:01:17we are back. Another sip of coffee. If
- 22:01:20you're familiar with my other videos,
- 22:01:21you notice I'm always sipping coffee. I
- 22:01:23always have a in my case latte next to
- 22:01:24me, an espresso. So the next step is to
- 22:01:26go ahead and initialize. We're going to
- 22:01:28call it the CNN or classifier neural
- 22:01:31network. And the reason we call it a
- 22:01:33classifier is because it's going to
- 22:01:34classify it between two things. It's
- 22:01:36going to be cat or dog. So when you're
- 22:01:38doing classification, you're picking
- 22:01:40specific objects. You're specific. It's
- 22:01:43a true or false. Yes, no. It is
- 22:01:45something or it's not. So first thing
- 22:01:48we're going to create our classifier and
- 22:01:50it's going to equal sequential. So their
- 22:01:51sequential setup is the classifier.
- 22:01:54That's the actual model we're using.
- 22:01:56That's the neural network. So we call it
- 22:01:58a classifier. And uh the next step is to
- 22:02:01add in our convolution. And let me just
- 22:02:03do a uh let me shrink that down in size
- 22:02:06so you can see the whole line. And let's
- 22:02:07talk a little bit about what's going on
- 22:02:09here. I have my classifier and I add
- 22:02:11something. What am I adding? Well, I'm
- 22:02:13adding my first layer. This first layer
- 22:02:16we're adding in is probably the one that
- 22:02:18takes the most work to make sure you
- 22:02:20have it set correct. And the reason I
- 22:02:22say that is this is your actual input.
- 22:02:24And we're going to jump here to the part
- 22:02:25that says input shape equals 64x 64x3.
- 22:02:30What does that mean? Well, that means
- 22:02:32that our pictures coming in. And there's
- 22:02:35these pictures. Remember we had like the
- 22:02:36picture of the car was 128x 128 pixels.
- 22:02:40Well, this one is 64x 64 pixels. And
- 22:02:43each pixel has three values. That's
- 22:02:46where these numbers come from. And it is
- 22:02:48so important that this matches. I
- 22:02:50mentioned a little bit that if you have
- 22:02:52like a larger picture, you have to
- 22:02:53reformat it to fit this shape. If it
- 22:02:56comes in as something larger, there's no
- 22:02:58input notes. There's no input neural
- 22:03:00network there that will handle that
- 22:03:02extra space. So, you have to reshape
- 22:03:03your data to fit in here. Now, the first
- 22:03:06layer is the most important because
- 22:03:07after that, KAS knows what your shape is
- 22:03:10coming in here and it knows what's
- 22:03:12coming out and so that really sets the
- 22:03:15stage. Most important thing is that
- 22:03:17input shape matches your data coming in.
- 22:03:19And you'll get a lot of errors if it
- 22:03:20doesn't. You'll go through there and
- 22:03:22picture number 55 doesn't match it
- 22:03:24correctly. And guess what it does? It
- 22:03:26usually gives you an error. And then the
- 22:03:27activation, if you remember, we talked
- 22:03:29about the different activations on here.
- 22:03:31We're using the reelu model. Like I
- 22:03:34said, that is the most commonly used now
- 22:03:36because one, it's fast. Doesn't have the
- 22:03:39added calculations in it. It just says
- 22:03:41here's the value coming out based on the
- 22:03:43weights and the value going in. And um
- 22:03:47from there, you know, it's uh if it's
- 22:03:49over one, then it's good or over zero,
- 22:03:51it's good. If it's under zero, then it's
- 22:03:53considered not active. And then we have
- 22:03:55this conversion 2D. What the heck is
- 22:03:58conversion 2D? I'm not going to go into
- 22:04:01too much detail in this because this has
- 22:04:03a couple of things it's doing in here, a
- 22:04:05little bit more in-depth than we're
- 22:04:06ready to cover in this tutorial. But
- 22:04:08this is used to convert from the photo
- 22:04:11cuz we have 64x 64x3 and we're just
- 22:04:14converting it to two-dimensional kind of
- 22:04:16setup. So it's very aware that this is a
- 22:04:18photograph and that different pieces are
- 22:04:20next to each other. And then we're going
- 22:04:21to add in uh a second convolutional
- 22:04:24layer. That's what the cov stands for
- 22:04:272D. So it's these are hidden layers. So
- 22:04:29we have our input layer and our two
- 22:04:31hidden layers and they are
- 22:04:32two-dimensional because we're dealing
- 22:04:34with a two-dimensional photograph. And
- 22:04:36you'll see down here that on the last
- 22:04:37one, we add a max pooling 2D and we put
- 22:04:40a pool size equals 22. And so what this
- 22:04:43is is that as you get to the end of
- 22:04:45these layers, one of the things you
- 22:04:47always want to think of is what they
- 22:04:48call mapping and then reducing.
- 22:04:50Wonderful terminology from the big data.
- 22:04:52We're mapping this data through all
- 22:04:53these layers. And now we want to reduce
- 22:04:56it to only two sets. In this case, it's
- 22:04:59already in two sets because it's a 2D
- 22:05:01photograph. But we had, you know, two
- 22:05:02dimensions by we actually have 64x 64
- 22:05:05by3. So now we're just getting it down
- 22:05:07to a 2x two. Just the two dimension
- 22:05:10two-dimensional instead of having the
- 22:05:11third dimension of colors. And we'll go
- 22:05:13ahead and run these. We're not really
- 22:05:15seeing anything in our run script
- 22:05:17because we're just setting up. This is
- 22:05:18all set up. And this is where you start
- 22:05:20playing because maybe you'll add a
- 22:05:21different layer in here to do something
- 22:05:23else to see how it works and see what
- 22:05:25your output is. That's what makes KAS so
- 22:05:27nice is I can with just a couple flips
- 22:05:29of code put in a whole new layer that
- 22:05:31does a whole new processing and see
- 22:05:33whether that improves my run or makes it
- 22:05:35worse. And finally, we're going to do
- 22:05:38the final setup, which is to flatten
- 22:05:40classifier, add a flatten setup. And
- 22:05:43then we're going to also add a layer, a
- 22:05:44dense layer, and then we're going to add
- 22:05:46in another dense layer. And then we're
- 22:05:49going to build it. We're going to
- 22:05:50compile this whole thing together. So,
- 22:05:52let's flip over and see what that looks
- 22:05:53like. And we've even numbered them for
- 22:05:55you. So, we're going to do the
- 22:05:56flattening. And flatten is exactly what
- 22:05:58it sounds like. We've been working in a
- 22:06:00two-dimensional array of picture, which
- 22:06:03actually is in three dimensions because
- 22:06:04of the pixels. The pixels have a whole
- 22:06:06another dimension to it of three
- 22:06:08different values. And we've kind of
- 22:06:10resized those down to 2x two. But now
- 22:06:12we're just going to flatten it. I don't
- 22:06:14want to have multiple dimensions being
- 22:06:16worked on by tensor and by kas. I want
- 22:06:19just a single array. So, it's flattened
- 22:06:21out. And then step four, full
- 22:06:23connection. So we add in our final two
- 22:06:26layers. And you could actually do all
- 22:06:28kinds of things with this. You could
- 22:06:29actually leave out this some of these
- 22:06:31layers and play with them. You do need
- 22:06:33to flatten it. That's very important.
- 22:06:35Then we want to use the dents again.
- 22:06:37We're taking this and we're taking
- 22:06:40whatever came into it. So once we take
- 22:06:42all those different the two dimensions
- 22:06:43or three dimensions as they are and we
- 22:06:45flatten it to one dimension. We want to
- 22:06:47take that and we're going to pull it
- 22:06:49into units of 128. They got that. You're
- 22:06:52say where did they get 128 from? You
- 22:06:53could actually play with that number and
- 22:06:55get all kinds of weird results. But in
- 22:06:56this case we took the 64 + 64 is 128.
- 22:07:00You could probably even do this with 64
- 22:07:02or 32. Usually you want to keep it in
- 22:07:04the same multiple whatever the data
- 22:07:05shape you're already using is in. And
- 22:07:07we're using the activation the re lu
- 22:07:10just like we did before. And then we
- 22:07:11finally filter all that into a single
- 22:07:15output. And it has how many units? One.
- 22:07:17Why? Because we want to know whether
- 22:07:19true or false. It's either a dog or a
- 22:07:21cat. You could say one is dog, zero is
- 22:07:24cat. Or maybe you're a cat lover and
- 22:07:26it's one is cat and zero is dog. And if
- 22:07:28you love both dogs and cats, you're
- 22:07:31going to have to choose. And then we use
- 22:07:33the sigmoid activation. If you remember
- 22:07:35from before, we had the reel and there's
- 22:07:38also the sigmoid. The sigmoid just makes
- 22:07:40it clear it's yes or no. We don't want a
- 22:07:42any kind of in between number coming
- 22:07:44out. And we'll go ahead and run this.
- 22:07:46And you'll see it's still all in setup.
- 22:07:48And then finally, we want to go ahead
- 22:07:49and compile. And let's put the compiling
- 22:07:52our um classifier neural network. And
- 22:07:54we're going to use the optimizer atom.
- 22:07:56And I hinted at this just a little bit
- 22:07:59before. Where does atom come in? Where
- 22:08:01does an optimizer come in? Well, the
- 22:08:03optimizer is the reverse propagation.
- 22:08:06When we're training it, it goes all the
- 22:08:08way through and says error and then how
- 22:08:10does it readjust those weights. There
- 22:08:12are a number of them. Atom is the most
- 22:08:14commonly used and it works best on large
- 22:08:17data. Most people stick with the atom
- 22:08:20because when they're testing on smaller
- 22:08:21data, see if their model is going to go
- 22:08:23through and get all their errors out
- 22:08:24before they run it on larger data sets.
- 22:08:26They're going to run it on atom anyway,
- 22:08:28so they just leave it on atom most
- 22:08:30commonly used. But there are some other
- 22:08:32ones out there. You should be aware of
- 22:08:33that that you might try them if you're
- 22:08:35stuck in a bind or you might blur that
- 22:08:37in the future, but usually atom is just
- 22:08:39fine on there. And then you have two
- 22:08:40more settings. You have loss and
- 22:08:42metrics. We're not going to dig too much
- 22:08:44into loss or metrics. These are things
- 22:08:46you really have to explore KAS because
- 22:08:48there are so many choices. This is how
- 22:08:50it computes the error. There's so many
- 22:08:52different ways to on your back
- 22:08:53propagation and your training. So we're
- 22:08:55using the atom model, but you can
- 22:08:57compute the error by um standard
- 22:08:59deviation, standard deviation squared.
- 22:09:02They use binary cross entropy. I'd have
- 22:09:04to look that up to even know what that
- 22:09:05is. There's so many of these. A lot of
- 22:09:07times you just start with the ones that
- 22:09:08look correct that are most commonly used
- 22:09:11and then you have to go read the KAS
- 22:09:12site and actually see what these
- 22:09:14different losses and metrics and what
- 22:09:16different options they have. So, we're
- 22:09:18not going to get too much into them
- 22:09:20other than to reference you over to the
- 22:09:21KAS website to explore them deeper, but
- 22:09:23we are going to go ahead and run them.
- 22:09:25And now we've set up our classifier. So,
- 22:09:27we have an object classifier. And if you
- 22:09:29go back up here, you'll see that we've
- 22:09:31added in step one. We added in our layer
- 22:09:33for the input. We added a layer that
- 22:09:36comes in there and uses the reelu for
- 22:09:38activation. And then it pulls the data.
- 22:09:41So this is even though these are two
- 22:09:42layers, the actual neural network layer
- 22:09:44is up here. And then it uses this to
- 22:09:46pull the data into a 2x two. So into a
- 22:09:49two-dimensional array from a
- 22:09:50three-dimensional array with the colors.
- 22:09:52Then we flatten it. So there's our adder
- 22:09:54flatten. And then we add another dense
- 22:09:56what they call dense layer. this dense
- 22:09:58layer goes in there and it it downsizes
- 22:10:00it to 128. It reduces it. So you can
- 22:10:04look at this as uh we're mapping all
- 22:10:06this data down the two-dimensional setup
- 22:10:08and then we flatten it. So we map it to
- 22:10:10a flatten map and then we take it and
- 22:10:12reduce it down to 128 and we use the
- 22:10:15reel again. And then finally we reduce
- 22:10:18that down to just a single output and we
- 22:10:20use a sigmoid to do that to figure out
- 22:10:22whether it's yes, no, true, false, in
- 22:10:25this case cat or dog. And then finally
- 22:10:27once we put all these layers together we
- 22:10:29compile them. That's what we've done
- 22:10:30here and we've compiled them as far as
- 22:10:33how it trains to use these settings for
- 22:10:35the training back propagation. So if you
- 22:10:37remember we talked about training our
- 22:10:40setup and when we go into this you'll
- 22:10:42see that we have two data sets. We have
- 22:10:44one called the training set and the
- 22:10:46testing set. And that's very standard in
- 22:10:48any data processing is you need to have
- 22:10:51that's pretty common in any data
- 22:10:53processing is you need to have a certain
- 22:10:55amount of data to train it and then you
- 22:10:56got to know whether it works or not. Is
- 22:10:58it any good and that's why you have a
- 22:11:00separate set of data for testing it
- 22:11:02where you already know the answer but
- 22:11:03you don't want to use that as part of
- 22:11:04the training set. So in here we jump
- 22:11:07into part two fitting the classifier
- 22:11:10neuron network to the images and then
- 22:11:12from KAS let me just zoom in there. I
- 22:11:15always love that about working with
- 22:11:16Jupyter Notebooks. You can really see.
- 22:11:18We're going to come in here. We do the
- 22:11:19cross pre-processing an image. And we
- 22:11:21import image data generator. It's so
- 22:11:24nice of KAS. It's such a high-end
- 22:11:26product right now going out. And since
- 22:11:28images are so common, they already have
- 22:11:30all this stuff to help us process the
- 22:11:31data, which is great. And so, we come in
- 22:11:33here, we do train data gen, and we're
- 22:11:36going to create our object for helping
- 22:11:38us train for reshaping the data so that
- 22:11:40it's going to work with our setup. and
- 22:11:43we use an image data generator and we're
- 22:11:45going to rescale it. And you'll see here
- 22:11:47we have one point which tells us it's a
- 22:11:49float value on the rescale over 255.
- 22:11:53Where does 255 come from? Well, that's
- 22:11:55the scale in the colors of the pictures
- 22:11:57we're using. They're value from 0 to
- 22:11:59255. So, we want to divide it by 255 and
- 22:12:02it'll generate a number between 0 and 1.
- 22:12:05They have sheer range and zoom range.
- 22:12:08Horizontal flip equals true. And this,
- 22:12:10of course, has to do with if the photos
- 22:12:12are different shapes and sizes. Like I
- 22:12:14said, it's a wonderful package. You
- 22:12:15really need to dig in deep to see all
- 22:12:17the different options you have for
- 22:12:19setting up your images. For right now
- 22:12:21though, we're going to just stick with
- 22:12:22some basic stuff here. And let me go
- 22:12:23ahead and run this code. And again, it
- 22:12:25doesn't really do anything because we're
- 22:12:27still setting up the pre-processing.
- 22:12:29Let's take a look at this next set of
- 22:12:31code. And this one is just huge. We're
- 22:12:33creating the training set. So the
- 22:12:35training set is going to go in here and
- 22:12:37it's going to use our train data gen we
- 22:12:40just created flow from directory. It's
- 22:12:42going to access in this case the path
- 22:12:45data set training set. That's a folder.
- 22:12:48So it's going to pull all the images out
- 22:12:50of that folder. Now I'm actually running
- 22:12:52this in the folder that the data sets
- 22:12:54in. So if you're doing the same setup
- 22:12:57and you load your data in there and
- 22:12:59you're doing this, make sure wherever
- 22:13:01your Jupyter notebook is saving things
- 22:13:03to that you create this path or you can
- 22:13:05do the complete path if you need to, you
- 22:13:07know, C colon slash etc. And the target
- 22:13:10size, the batch size and class mode is
- 22:13:13binary. So the classes, we're switching
- 22:13:15everything to a binary value. Batch
- 22:13:17size. What the heck is batch size? Well,
- 22:13:18that's how many pictures we're going to
- 22:13:20batch through the training each time.
- 22:13:22And the target size 64x 64. A little
- 22:13:25confusing, but you can see right here
- 22:13:26that this is just a general training and
- 22:13:28you can go in there and look at all the
- 22:13:30different settings for your training
- 22:13:31set. And of course with different data,
- 22:13:33we're doing pictures. There's all kinds
- 22:13:35of different settings depending on what
- 22:13:36you're working with. Let's go ahead and
- 22:13:38run that and see what happens. And
- 22:13:39you'll see that it found 800 images
- 22:13:41belonging to one classes. So we have 800
- 22:13:44images in the training set. And if we're
- 22:13:47going to do this with uh the training
- 22:13:50set, we also have to format the pictures
- 22:13:52in the test set. Now, we're not actually
- 22:13:55doing any predictions. We're not
- 22:13:56actually programming the model yet. All
- 22:14:00we're doing is preparing the data. So,
- 22:14:01we're going to prepare a training set
- 22:14:03and the test set. So, any changes we
- 22:14:05make to the training set at this point
- 22:14:07also have to be made to the test set.
- 22:14:09So, we've done this thing. We've done a
- 22:14:11train data generator. We've done our
- 22:14:13training set. And then we also have
- 22:14:15remember our test set of data. So I'm
- 22:14:17going to do the same thing with that.
- 22:14:18I'm going to create a test data gen and
- 22:14:21we're going to do this image data
- 22:14:22generator. We're going to rescale one
- 22:14:24over 255. We don't need the other
- 22:14:26settings, just the single setting for
- 22:14:28the test data gen. And we're going to
- 22:14:30create our test set. We're going to do
- 22:14:31the same thing we did with the test set
- 22:14:33except that we're pulling it from the
- 22:14:34test set folder. And we'll run that. And
- 22:14:37you'll see in our test set we found
- 22:14:392,000 images. That's about right. We're
- 22:14:42using 20% of the images as test and 80%
- 22:14:45to train it. And then finally, we've set
- 22:14:47up all our data. We've set up all our
- 22:14:50layers, which is where all the work is
- 22:14:52is cleaning up that data, making sure
- 22:14:54it's going in there correctly. And we're
- 22:14:55actually going to fit it. We're going to
- 22:14:58train our data set. And let's see what
- 22:15:00that looks like. And here we go. Let's
- 22:15:02put the information in here. And let's
- 22:15:03just take a quick look at what we're
- 22:15:06looking at with our fit generator. We
- 22:15:08have our classifier.fit
- 22:15:09generator. That's our back propagation.
- 22:15:12So the information goes through forward
- 22:15:14with a picture and it says, "Oh, you're
- 22:15:16either right or you're wrong." And then
- 22:15:18the error goes backward and reprograms
- 22:15:21all those weights. So we're training our
- 22:15:23neural network. And of course, we're
- 22:15:25using the training set. Remember, we
- 22:15:27created the training set up here. And
- 22:15:28then we're going steps per epic. So it's
- 22:15:318,000 steps. Epic means that that's how
- 22:15:34many times we go through all the
- 22:15:36pictures. So we're going to rerun each
- 22:15:37of the pictures. and we're going to go
- 22:15:39through the whole data set 25 times, but
- 22:15:41we're going to look at each picture
- 22:15:43during each epic 8,000 times. So, we're
- 22:15:46really programming the heck out of this
- 22:15:47and going back over it. And then they
- 22:15:50have validation data equals test set.
- 22:15:52So, we have our training set and then
- 22:15:54we're going to have our test set to
- 22:15:56validate it. So, we're going to do this
- 22:15:57all in one shot and we're going to look
- 22:15:59at that and they're going to do 200
- 22:16:00steps for each validation and we'll see
- 22:16:02what that looks like in just a minute.
- 22:16:04Let's go ahead and run our training
- 22:16:05here. And we're going to fit our data.
- 22:16:07And as it goes, it says epic one of 25.
- 22:16:10You start realizing that this is going
- 22:16:12to take a while. On my older computer,
- 22:16:14it takes about 45 minutes. I have a dual
- 22:16:18processor. You know, we're processing uh
- 22:16:2110,000 photos. That's not a small amount
- 22:16:23of photographs to process. So, if you're
- 22:16:26on your laptop, you know, which I am,
- 22:16:27it's going to take a while. So, let's go
- 22:16:29ahead and uh go get our cup of coffee
- 22:16:31and a sip and come back and see what
- 22:16:33this looks like. So, I'm back. You
- 22:16:35didn't know I was gone. That was
- 22:16:36actually a lengthy pause there. I made a
- 22:16:39couple changes. Let's discuss those
- 22:16:41changes real quick and why I made them.
- 22:16:43So, the first thing I'm going to do is
- 22:16:44I'm going to go up here and insert a
- 22:16:46cell above and let's paste the original
- 22:16:48code back in there. And you'll see that
- 22:16:50the original thing was steps per epic
- 22:16:528,000, 25 epics, and validation steps
- 22:16:552,000. And I changed these to 4,000
- 22:16:58epics or 4,000 steps per epic, 10 epics,
- 22:17:02and just 10 validation steps. And this
- 22:17:05will cause problems if you're doing this
- 22:17:07as a commercial release. But for demo
- 22:17:09purposes, this should work. And if you
- 22:17:11remember our steps per epic, that's how
- 22:17:13many photos we're going to process. In
- 22:17:15fact, let me go ahead and get my drawing
- 22:17:16pen out. And uh let's just highlight
- 22:17:18that right here. We have 8,000 pictures
- 22:17:21we're going through. So for each epic,
- 22:17:23I'm going to change this to 4,000. I'm
- 22:17:24going to cut that in half. So, it's
- 22:17:26going to randomly pick 4,000 pictures
- 22:17:28each time it goes through an epic. And
- 22:17:29the epic is how many processes. So, this
- 22:17:32is 25. And I'm just going to cut that to
- 22:17:3410. So, instead of doing 25 runs through
- 22:17:378,000 photos each, which you can do the
- 22:17:39math of 25 * 8,000, I'm only going to do
- 22:17:4210 through 4,000. So, I'm going to run
- 22:17:44this 40,000 times through the processes.
- 22:17:47And the next thing I not you'll you'll
- 22:17:49want to notice is that I also changed
- 22:17:50the validation step. And this would
- 22:17:52cause some major problems in releasing
- 22:17:54cuz I dropped it all the way down to 10.
- 22:17:56What the validation step does is it says
- 22:17:58we have 2,000 photos in our training or
- 22:18:01in our testing set and we're going to
- 22:18:03use that for validation. Well, I'm only
- 22:18:05going to use a random 10 of those to
- 22:18:07validate. So, not really the best
- 22:18:09settings, but let me show you why we did
- 22:18:11that. Let's scroll down here just a
- 22:18:13little bit and let's look at the output
- 22:18:15here and see what that what's going on
- 22:18:17there. So, I've got my drawing tool back
- 22:18:19on, and you'll see here it lists a run.
- 22:18:22So, each time it goes through an epic,
- 22:18:24it's going to do 4,000 steps. And this
- 22:18:26is where the 4,000 comes in. So, that's
- 22:18:28where we have. We have epic one of 10,
- 22:18:294,000 steps. So, it's randomly picking
- 22:18:32half the pictures in the file and going
- 22:18:33through them. And then we're going to
- 22:18:34look at this number right here. That is
- 22:18:36for the whole epic, and that's 24, 411
- 22:18:40seconds. And if you remember correctly,
- 22:18:42you divide that by 60, you get minutes.
- 22:18:44If you divide that by 60, you get hours.
- 22:18:47Or you can just divide the whole thing
- 22:18:48by 60 * 60 which is 3600. If 3600 is an
- 22:18:52hour, this is roughly 45 minutes right
- 22:18:55here. And that's 45 minutes to process
- 22:18:57half the pictures. So if I was doing all
- 22:19:00the pictures, we're talking an hour and
- 22:19:02a half per epic times 36 or no 25. They
- 22:19:06had 25 up above 25. So that's roughly a
- 22:19:09couple days. A couple days of
- 22:19:11processing. Well, for this demo, we
- 22:19:12don't want to do that. I don't want to
- 22:19:14come back the next day. Plus, my
- 22:19:15computer did a reboot in the middle of
- 22:19:17the night. So, we look at this and we
- 22:19:19say, "Okay, let's we're just testing
- 22:19:20this out. My computer that I'm running
- 22:19:22this on is a dual core processor. Uh,
- 22:19:25runs 0.9 gigahertz per second. For a
- 22:19:28laptop, you know, it's good about 4
- 22:19:29years ago, but for running something
- 22:19:31like this, it's probably a little slow.
- 22:19:33So, we cut the times down. And the last
- 22:19:34one was validation. We're only
- 22:19:36validating it on a random 10 photos. And
- 22:19:38this comes into effect because you're
- 22:19:40going to see down here where we have
- 22:19:42accuracy, value loss, value accuracy,
- 22:19:46and loss. Those are very important
- 22:19:48numbers to look at. So the 10 means I'm
- 22:19:50only validating across 10 pictures. That
- 22:19:53is where here we have value. This is ACC
- 22:19:56is for accuracy. Value loss. We're not
- 22:19:58going to worry about that too much. And
- 22:19:59accuracy. Now accuracy is while it's
- 22:20:02running, it's putting these two numbers
- 22:20:04together. That's what accuracy is. And
- 22:20:06value accuracy is at the end of the
- 22:20:09epic. What's our accuracy into the epic?
- 22:20:11What is it looking at? In this tutorial,
- 22:20:12we're not going to go so deep, but these
- 22:20:14numbers are really important when you
- 22:20:16start talking about these two numbers
- 22:20:19reflect bias. That is really important.
- 22:20:22We just put that up there. And bias is a
- 22:20:24little bit beyond this tutorial, but the
- 22:20:26short of it is is if this accuracy,
- 22:20:28which is being our validation per step
- 22:20:30is going down and the value accuracy
- 22:20:34continues to go up, that means there's a
- 22:20:36bias. That means I'm memorizing the
- 22:20:38photos I'm looking at. I'm not actually
- 22:20:41looking for what makes a dog a dog, what
- 22:20:44makes a cat a cat. I'm just memorizing
- 22:20:46them. And so the more this discrepancy
- 22:20:48grows, the bigger the bias is. And that
- 22:20:50is really the beauty of the KAS neural
- 22:20:54network. It has a lot of built-in
- 22:20:55features like this that make that really
- 22:20:57easy to track. So let's go ahead and
- 22:20:59take a look at the next set of code. So
- 22:21:01here we are into part three. We're going
- 22:21:04to make a new prediction. And so we're
- 22:21:06going to bring in a couple tools for
- 22:21:07that. And then we have to process the
- 22:21:09image coming in and find out whether
- 22:21:11it's an actual dog or cat if we can
- 22:21:13actually use this to identify it. And of
- 22:21:15course the final step of part three is
- 22:21:17to print prediction. We'll go ahead and
- 22:21:19combine these. And of course you can see
- 22:21:20me there adding more sticky notes to my
- 22:21:22computer screen hidden behind the
- 22:21:24screen. And you know last one was don't
- 22:21:26forget to feed the cat and the dog.
- 22:21:29So let's go and take a look at that and
- 22:21:30see what that looks like in code and put
- 22:21:32that in our Jupyter notebook. All right.
- 22:21:34And let's paste that in here. And we'll
- 22:21:36start by importing numpy as np. Numpy is
- 22:21:40a very common package. I pretty much
- 22:21:42import it on any Python project I'm
- 22:21:44working on. Another one I use regularly
- 22:21:46is pandas. They're just ways of
- 22:21:47organizing the data. And then np is
- 22:21:49usually the standard in most machine
- 22:21:52learning tools as the return for the
- 22:21:54data array. Although you know you use a
- 22:21:56standard data array from Python. And we
- 22:21:58have cross pre-processing import image.
- 22:22:01This should all look familiar because
- 22:22:02we're going to take a test image and
- 22:22:04we're going to set that equal to in this
- 22:22:06case cat or dog one as you can see over
- 22:22:09here. And you know let me get my drawing
- 22:22:11tool back on. So let's take a look at
- 22:22:13this. We have our test image we're
- 22:22:15loading and in here we have test image
- 22:22:17one. And this one hasn't data hasn't
- 22:22:19seen this one at all. So this is all
- 22:22:20new. Oh, let me shrink the screen down.
- 22:22:22Let me start that over. So here we have
- 22:22:24my test image and we went ahead and the
- 22:22:26cross processing has this nice image
- 22:22:29setup. So we're going to load the image
- 22:22:31and we're going to alter it to a 64x 64
- 22:22:34print. So right off the bat, we're going
- 22:22:36to cross is nice that way. It
- 22:22:37automatically sets it up for us so we
- 22:22:39don't have to redo all our images and
- 22:22:41find a way to reset those. And then we
- 22:22:42use also to set the image to an array.
- 22:22:45So again, we're all in pre-processing
- 22:22:47the data just like we pre-processed
- 22:22:49before with our test information and our
- 22:22:51training data. And then we use the
- 22:22:53numpy. Here's our numpy that's uh from
- 22:22:56our um right up here. Import numpy as in
- 22:22:58p expand the dimensions test image axis
- 22:23:01equal zero. So it puts it into a single
- 22:23:03array. And then finally all that work
- 22:23:07all that pre-processing and all we do is
- 22:23:09we run the result. We click on here we
- 22:23:10go result equals classifier predict test
- 22:23:13image. And then we find out, well, what
- 22:23:15is the test image? And let's just take a
- 22:23:17quick look and just see what that is.
- 22:23:18And you can see when I ran it, it comes
- 22:23:20up dog. And if we look at those images,
- 22:23:23there it is. Cat or dog. Image number
- 22:23:25one. That looks like a nice floppy eared
- 22:23:27lab. Friendly with his tongue hanging
- 22:23:29out. It's either that or a very floppy
- 22:23:31eared cat. I'm not sure which. But
- 22:23:33according to our software, it says it's
- 22:23:34a dog. And uh we have a second picture
- 22:23:36over here. Let's just see what happens
- 22:23:37when we run the second picture. We can
- 22:23:39go up here and change this uh from dog
- 22:23:40image one to two. We'll run that. and it
- 22:23:43comes down here and says cat. You can
- 22:23:45see me highlighting it down there as
- 22:23:47cat. So, our process works. You're able
- 22:23:49to label a dog a dog and a cat a cat
- 22:23:51just from the pictures. There we go.
- 22:23:53Cleared my drawing tool. And the last
- 22:23:55thing I want you to notice when we come
- 22:23:56back up here to when I ran it, you'll
- 22:23:59see it has an accuracy of one and the
- 22:24:02value accuracy of one. Well, the value
- 22:24:05accuracy is the important one because
- 22:24:07the value accuracy is what it actually
- 22:24:09runs on the test data. Remember, I'm
- 22:24:11only testing it on. and I'm only
- 22:24:12validating it on a random 10 photos and
- 22:24:14those 10 photos just happened to come up
- 22:24:16one. Now, when they ran this on the
- 22:24:18server, it actually came up about 86%.
- 22:24:21This is why cutting these numbers down
- 22:24:23so far for a commercial release is bad.
- 22:24:26So, you want to make sure you're a
- 22:24:27little careful of that when you're
- 22:24:28testing your stuff that you change these
- 22:24:30numbers back when you run it on a more
- 22:24:31enterprise computer other than your old
- 22:24:33laptop that you're just practicing on or
- 22:24:35messing with. And we come down here and
- 22:24:37again, you know, we had the validation
- 22:24:38of cat. And so we have successfully
- 22:24:40built a neural network that could
- 22:24:42distinguish between photos of a cat and
- 22:24:44a dog. Imagine all the other things you
- 22:24:46could distinguish. Imagine all the
- 22:24:48different industries you could dive into
- 22:24:50with that. Just being able to understand
- 22:24:51those two difference of pictures. What
- 22:24:53about mosquitoes? Could you find the
- 22:24:55mosquitoes that bite versus the
- 22:24:56mosquitoes that are friendly? It turns
- 22:24:58out the mosquitoes that bite us are only
- 22:25:004% of the mosquito population, if even
- 22:25:02that, maybe 2%. There's all kinds of
- 22:25:04industries that use this and there's so
- 22:25:06many industries that are just now
- 22:25:08realizing how powerful these tools are.
- 22:25:11Just in the photos alone, there is a
- 22:25:13myriad of industries sprouting up. And I
- 22:25:15said it before, I'll say it again. What
- 22:25:17an exciting time to live in with these
- 22:25:19tools and that we get to play with. So
- 22:25:22key takeaways. Well, we covered what is
- 22:25:24a neural network. We use all kinds of
- 22:25:26processing the map images on your phone.
- 22:25:29We talked about things that a neural
- 22:25:30network can do. translate text, identify
- 22:25:33faces all the way to control robots, you
- 22:25:36know, lots of exciting things. How does
- 22:25:38a neural network work? So, we discussed
- 22:25:40that with the different layers going
- 22:25:42from the picture to the input layer to
- 22:25:44the hidden layers and their weights to
- 22:25:46the final output layer. We also talked
- 22:25:48about how it does the math and computing
- 22:25:50the output as yes or no, categorically
- 22:25:54true false. We discussed types of
- 22:25:55artificial neural networks. A lot of
- 22:25:57vocabulary there from the feed forward
- 22:25:59neural network which is the most
- 22:26:01commonly used. That's the one the neural
- 22:26:03network we used is a feed forward neural
- 22:26:04network that does backward propagation
- 22:26:06to train. And there's a lot of other
- 22:26:08ones out there. There's the radial
- 22:26:10biases, the cohen self-organizing
- 22:26:13recurrent neural network, convolution
- 22:26:15neural network, modular neural network.
- 22:26:17The big one was modular because it
- 22:26:19incorporates pieces of all the other
- 22:26:21ones. So that whatever you're working on
- 22:26:23now is a huge conglomerate of multiple
- 22:26:27networks. Just all cutting edge. All of
- 22:26:29it's new. People even working on it
- 22:26:31don't even know where it's going. Again,
- 22:26:32very exciting times. And finally, we dug
- 22:26:35through my favorite part. You can see
- 22:26:37with my uh latte on one side, my old
- 22:26:39school pens and pencil, and all my
- 22:26:41sticky notes working away. That's not
- 22:26:43actually me, by the way. You probably
- 22:26:45guessed that. And we walked through and
- 22:26:46actually did a cat and dog photo, a
- 22:26:48simple cat and dog photo. And you could
- 22:26:50see where some of the problems are in
- 22:26:51processing large amounts of photographs
- 22:26:53and data where that starts to become
- 22:26:55going from a single machine on my laptop
- 22:26:58with this, you know, lower amount of
- 22:27:00resources all the way to big data. How
- 22:27:02if you're processing hundreds and
- 22:27:04thousands of these photos, this now
- 22:27:06needs to be set up on an enterprise
- 22:27:07machine or even on a cluster of
- 22:27:09computers. Again, significantly past the
- 22:27:11scope of this. The neat part about it
- 22:27:13though is once you write this code, most
- 22:27:15of this code, they now have tools that
- 22:27:17you can almost take the same ideas, if
- 22:27:19not the actual code, and push it right
- 22:27:21onto a cluster computation. So really
- 22:27:23cool times for this Python. My name is
- 22:27:26Richard Kersner with the SimplyLearn
- 22:27:28team. That's www.simplearn.com.
- 22:27:30Get certified, get ahead. Although deep
- 22:27:33learning is uh been around for a while,
- 22:27:36it is just in its infant stages of
- 22:27:38development as far as exploding on the
- 22:27:40market. I mean it is right now they're
- 22:27:42building robots with it. Deep learning
- 22:27:44is used to train robots to perform human
- 22:27:47tasks. Music composition. Deep neural
- 22:27:49nets can be used to produce music by
- 22:27:51making computers learn the patterns
- 22:27:53involved in composing music. Image
- 22:27:55colorization. Neural network recognizes
- 22:27:57objects and uses information from the
- 22:27:59images to color them. Machine
- 22:28:01translation. Given a word, phrase or a
- 22:28:03sentence in one language, neural
- 22:28:04networks automatically translate them
- 22:28:06into another language. Google Translate
- 22:28:08is one such popular machine translator
- 22:28:10you may have come across. And you'll
- 22:28:13notice in here we didn't show any
- 22:28:14examples of straight numbers like uh
- 22:28:17projective cells in a business tracking
- 22:28:20your favorite stock. You can certainly
- 22:28:21do those with machine languages, but
- 22:28:23this is the next level. Uh save that for
- 22:28:25your regression models, your linear
- 22:28:27regression where you're actually
- 22:28:28processing and crunching just straight
- 22:28:30numbers. With machine learning and deep
- 22:28:32learning, we're going to a whole new
- 22:28:33level as far as what we can figure out
- 22:28:35on the computer. What's in it for you?
- 22:28:37We're going to cover what is deep
- 22:28:39learning. We're going to take a look at
- 22:28:40the biological versus artificial
- 22:28:42intelligence. What is neural network
- 22:28:44activation function in your neural
- 22:28:46network and the cost function and how do
- 22:28:48neural networks work. How do neural
- 22:28:50networks learn? So there's a little
- 22:28:52you'll see a switch right there. We just
- 22:28:54went from how are they working in the
- 22:28:55math in the background to exactly how
- 22:28:57are they learning. We'll be implementing
- 22:28:58the neural network. We'll do a gradient
- 22:29:00descent deep learning platforms and
- 22:29:02we'll give an introduction to TensorFlow
- 22:29:04and implementation in TensorFlow. That's
- 22:29:06Google's platform that they open sourced
- 22:29:08recently and it's probably one of the
- 22:29:09most cutting edges in deep learning and
- 22:29:12even it is still in the infant stage
- 22:29:13which is one of the reasons they
- 22:29:15released it to open source. What is deep
- 22:29:17learning? Deep learning is a sub field
- 22:29:19of machine learning that deals with
- 22:29:20algorithms inspired by the structure and
- 22:29:22function of the brain. And you can see
- 22:29:24we have a nice picture here. We have
- 22:29:25artificial intelligence which is kind of
- 22:29:27the big bubble that encompasses all
- 22:29:29these different things we're talking
- 22:29:30about. This is ability of machine to
- 22:29:32imitate intelligent human behavior. And
- 22:29:34in there we have machine learning
- 22:29:36application of AI that allows a system
- 22:29:38to automatically learn and improve from
- 22:29:41experience. And if you looked at any of
- 22:29:42our other videos, you'll know that
- 22:29:44machine learning covers a lot. So deep
- 22:29:46learning is a subcategory of that. But
- 22:29:48don't forget machine learning has all
- 22:29:50kinds of other tools that people use to
- 22:29:52do very basic uh descriptive and
- 22:29:54predictive and postcriptive uh
- 22:29:56analytics. And then you have deep
- 22:29:58learning application of machine learning
- 22:30:00that uses complex algorithms and deep
- 22:30:03neural nets to train a model. Let's take
- 22:30:05a look at the biological neuron versus
- 22:30:07the artificial neuron. Now remember in
- 22:30:09the human brain and and this is true for
- 22:30:11most animals there are a lot of
- 22:30:12different neurons going on. So this is
- 22:30:14the very basic one. I mean there's
- 22:30:16hundreds of different cells involved. So
- 22:30:19when we talk about neural networks and
- 22:30:20this is why I say it's in a very infant
- 22:30:22stage. They're really basing it on uh
- 22:30:24just the most basic thing that we're
- 22:30:26able to figure out going on in the
- 22:30:28neural networks. And you can see right
- 22:30:30here we have dendrites fetch information
- 22:30:32from an adjacent neurons and pass them
- 22:30:34on as inputs. So you have your data
- 22:30:36coming in and your data going out. Any
- 22:30:39computer model should be looking at that
- 22:30:40what's coming in what's going out. The
- 22:30:42data is fed as an input to the neuron.
- 22:30:44So we look at the artificial neuron. You
- 22:30:47can see we have our inputs. They come in
- 22:30:49each one is specially weighted into the
- 22:30:51neuron and then the neuron has an
- 22:30:52output. The cell nucleus processes the
- 22:30:55information received from the dendrites
- 22:30:57and the neuron processes the information
- 22:30:59provided as inputs. Axons are the cables
- 22:31:01over which the information is
- 22:31:03transmitted and the information is
- 22:31:05transferred over weighted channels. So
- 22:31:06you can look at that uh I mentioned
- 22:31:08weights briefly but you alter the data
- 22:31:10coming in. So those weights are what
- 22:31:12causes different information coming in
- 22:31:15to be weighted differently and processed
- 22:31:17differently. And the synapses receive
- 22:31:19the information from the axons and
- 22:31:20transmit it to the adjacent neurons.
- 22:31:22That's in your biological model. And
- 22:31:24then when we look at the artificial
- 22:31:25neuron, the output is a final value
- 22:31:27predicted by the artificial neuron. So
- 22:31:29as we dig deeper into looking at the
- 22:31:31theory behind the neural network and we
- 22:31:34kind of flip back and forth between
- 22:31:35these because there's two huge aspects
- 22:31:38of it. One is from the outside. What are
- 22:31:40you seeing and what's going on from the
- 22:31:42inside so you can find to do what you
- 22:31:44need to do and give the best results you
- 22:31:46can. And we start off with what do we
- 22:31:47feed? We feed an unlabeled image to a
- 22:31:50machine which identifies it without any
- 22:31:52human intervention. And so you can see
- 22:31:54here we have a circle that comes in at
- 22:31:56784 pixels and it comes in by 28x 28.
- 22:31:59And you can see how it colors in the um
- 22:32:01the circle on there. And we put a
- 22:32:03triangle in. The triangle in also comes
- 22:32:05in as 28x 28 and it has 784 pixels. So
- 22:32:08you'll see between these two both of
- 22:32:10them are 784 pixels. This machine is
- 22:32:13intelligent enough to differentiate
- 22:32:15between the various shapes. So that's
- 22:32:17what we want to use our neural network
- 22:32:18to do is to say hey this is a circle.
- 22:32:20This is a triangle. That's more of a
- 22:32:22categorical. You can also do a
- 22:32:23regression model where you're actually
- 22:32:25putting out float value or a numerical
- 22:32:27value. We'll be looking at the true
- 22:32:28false or the categorical model mostly
- 22:32:30because that's where you usually start
- 22:32:31at the different there is no real
- 22:32:33difference when you as far as the way
- 22:32:36the internal functioning goes when you
- 22:32:38start flipping between them other than
- 22:32:40well we'll talk about that in just a
- 22:32:41minute. So you can actually go between
- 22:32:42the two quite easily and the neural
- 22:32:44network provides this capability. So
- 22:32:46we're going to use this capability to
- 22:32:48look between those two. One of the
- 22:32:49things I want you to note in here is
- 22:32:51that we're looking at 784 pixels. We're
- 22:32:54looking at 784 inputs. That's very
- 22:32:57different than stock with a high low or
- 22:32:59last year's sales based on date or we're
- 22:33:02looking at just a couple of numbers and
- 22:33:04they're very clear. They're numbers.
- 22:33:06They're very clear what they are, which
- 22:33:07is something you'd put into a machine
- 22:33:09learning linear regression model. This
- 22:33:10is a step up from that in that we're
- 22:33:12looking at complex patterns and how do
- 22:33:14you figure those complex patterns out.
- 22:33:16So, a neural network is a system modeled
- 22:33:19on the human brain. And we looked at
- 22:33:21that comparing the two. Let's go ahead
- 22:33:22and look deeper into the neural network
- 22:33:24itself. We have our inputs coming in. So
- 22:33:26the inputs are fed to a neuron that
- 22:33:28processes a data and gives us an output.
- 22:33:30Input and output. This is the most basic
- 22:33:33structure of a neural network known as a
- 22:33:35perceptron. So if you see the term
- 22:33:37perceptron, that's what we're talking
- 22:33:38about. We're talking about this single
- 22:33:40node that has inputs and an output.
- 22:33:42However, neural networks are usually
- 22:33:44much more complex. Let's start with
- 22:33:47visualizing a neural network as a black
- 22:33:49box. And I always love that symbol. It's
- 22:33:51a black box. It's kind of magical. We
- 22:33:53have our inputs coming in and we want
- 22:33:55certain outputs. The box takes inputs,
- 22:33:57processes them, and gives an output.
- 22:34:00Let's have a look at what happens within
- 22:34:02this box. And you can see me there in my
- 22:34:04uh secret agent getup and I got my
- 22:34:06hidden hood and everything. I guess I'm
- 22:34:08part of the uh black skull or something
- 22:34:10like that group. Uh so let's take a look
- 22:34:12at what happens within this magic box.
- 22:34:14And remember, we're skipping back and
- 22:34:15forth between the theory of what's going
- 22:34:17on in the box, which you have to know
- 22:34:19how to fine-tune and how to build,
- 22:34:21versus looking at it from the outside.
- 22:34:23We're programming this box, and we have
- 22:34:25an input and an output to the box as a
- 22:34:27whole. Within the box exists a network
- 22:34:29that is a core of deep learning. And you
- 22:34:31can see here we're showing one layer and
- 22:34:33we have our grid coming in. The network
- 22:34:35consists of layers of neurons. Each
- 22:34:38neuron is associated with a number
- 22:34:40called the bias. And you can think of
- 22:34:42the bias uh if you overly simplify this
- 22:34:45and we're doing a linear regression
- 22:34:46model. This is your y intercept in your
- 22:34:49uklidian geometry. You have to have
- 22:34:50something that offsets it. And so you
- 22:34:52always have a bias in these cells.
- 22:34:54Neurons of each layer transmit
- 22:34:56information to neurons of the next layer
- 22:34:58over channels. And so you can see each
- 22:35:00of our layers going through from left to
- 22:35:02right. These channels are associated
- 22:35:04with numbers called weights. These
- 22:35:06weights along with the biases determine
- 22:35:08the information that is passed over from
- 22:35:10the neuron to neuron. So just like the
- 22:35:12bias is your y intercept in uklidian
- 22:35:15geometry. You could look at the an one
- 22:35:17weight. Remember this is very
- 22:35:19complicated. So we're not looking at
- 22:35:20just one weight. You could look at the
- 22:35:21weight as your slope of the line. Or if
- 22:35:23you're doing x= uh my y + c, it would be
- 22:35:27the m value. Neurons of each layer
- 22:35:29transmit information to neurons of the
- 22:35:31next layer. And you can see here as they
- 22:35:32light up going across into the final
- 22:35:34layer. and then to the output. And in
- 22:35:37this case, the output is going to be
- 22:35:38either uh a square in this one or it
- 22:35:41might light up the other one which is a
- 22:35:42circle. The output layer emits a
- 22:35:44predicted output. So in this case, we're
- 22:35:46looking at a classification uh true
- 22:35:49false. Is it a circle? Is it a triangle?
- 22:35:51Is it a square? Let's now go deeper.
- 22:35:53What happens within the neuron? So we're
- 22:35:55going to dig deeper and start getting a
- 22:35:57little bit closer to some of the math.
- 22:35:58Don't worry, you don't have to be a
- 22:36:00calculus expert and know your
- 22:36:01differential equations. Even though this
- 22:36:03is one giant differential equation, you
- 22:36:06don't need to understand those to
- 22:36:07understand what's going on. Within each
- 22:36:09neuron, the following operations are
- 22:36:11performed. The product of each input and
- 22:36:13the weight of the channel it's passed
- 22:36:15over is found. This is simply addition.
- 22:36:17We're going to sum up the weight times
- 22:36:19the output from the previous channel and
- 22:36:21plus the bias. Sum of the weighted
- 22:36:23products is computed. This is called the
- 22:36:25weighted sum. Bias unique to the neuron
- 22:36:27is added to the weighted sum. The final
- 22:36:29sum is then subjected to the particular
- 22:36:32function and we'll discuss those that
- 22:36:34particular function. That part is really
- 22:36:36important because those functions uh
- 22:36:38have a huge impact on how well your
- 22:36:40model performs under different
- 22:36:42conditions. The final sum is then
- 22:36:43subject to a particular function. This
- 22:36:45is the activation function. So if you
- 22:36:48ever hear the term activation function,
- 22:36:50that's what we're talking about. What
- 22:36:51activates this cell and what doesn't. As
- 22:36:53we dig deeper into activation function,
- 22:36:56an activation function takes the
- 22:36:57weighted sum of the input as its input
- 22:37:00adds a bias and provides an output. And
- 22:37:03a lot of times you'll actually see one
- 22:37:05formula for the sum of the weight the
- 22:37:07weighted sum and the bias. You'll just
- 22:37:09see that as a single line of everything
- 22:37:10added together. And here we've broken it
- 22:37:12apart because it makes it clear that
- 22:37:14this bias is not computed the same as
- 22:37:16the weighted sums. Here are the most
- 22:37:18popular types of activation function.
- 22:37:20And I always find these interesting
- 22:37:22because at one point I was sitting at a
- 22:37:24table with a gentleman who was finishing
- 22:37:26his PhD. He was in his last year and he
- 22:37:28said he went through all this stuff and
- 22:37:30he ended up just trying the four
- 22:37:32different activation functions on this
- 22:37:34particular problem he was working on. So
- 22:37:36knowing the math behind it doesn't
- 22:37:38necessarily mean you're going to know it
- 22:37:40right away. Uh so even somebody who
- 22:37:42might have a PhD and be doing the
- 22:37:43calculations on this comes back out of
- 22:37:46it and ends up just trying the different
- 22:37:48uh um activation functions to see what's
- 22:37:50going to make a difference. And a lot of
- 22:37:51times that's a final step. That's the
- 22:37:53kind of thing where you built your whole
- 22:37:54model. You've come back and you're like
- 22:37:56wait a minute can I do a better deal
- 22:37:58with a sigmoid function or the threshold
- 22:38:00or the rectifier. Knowing what they're
- 22:38:02doing is important so you can explain it
- 22:38:03to somebody else. And again you probably
- 22:38:05do this on a small set of data. If
- 22:38:07you're working with big data, uh you
- 22:38:08don't want to take down the full server
- 22:38:10farm just to test out your three
- 22:38:12different series. You take a small
- 22:38:14portion of that data, test it, and then
- 22:38:16you put it through to the big data. So
- 22:38:17let's take a look at this. We have the
- 22:38:19sigmoid function, and it's used for
- 22:38:20models where we have to predict the
- 22:38:22probability as an output. It exists
- 22:38:24between zero and one. And you'll see
- 22:38:26that's true of all of our activation
- 22:38:28functions we're working with. Either the
- 22:38:30cells on or off, it's true or false. And
- 22:38:33there might be a little variation in
- 22:38:34there which as an output could be used
- 22:38:37to compute uncertainty in your solution.
- 22:38:40So if you're getting a 7 with this
- 22:38:42activation function, it might be well
- 22:38:43I'm not sure if that's really a square
- 22:38:45or I'm not sure that's really a
- 22:38:46triangle. And that might be a flag for
- 22:38:49it to be looked at by a human observer
- 22:38:51at least in today's models where we're
- 22:38:53at right now. And you can see here we
- 22:38:54have the formula is simply equals 1 over
- 22:38:571 + e the minus x where x is your value
- 22:39:00coming in. and it's going to give you a
- 22:39:02result that looks very similar to the
- 22:39:03graph on there which is somewhere
- 22:39:04between zero and one. Um, and right in
- 22:39:07the middle you can see that there's a
- 22:39:09huge uh kind of you can go through all
- 22:39:11the different values and uncertainties
- 22:39:13involved. So the sigmoid function is
- 22:39:15probably the default on most of them. Uh
- 22:39:17the next one is the threshold function.
- 22:39:19It is a threshold-based activation
- 22:39:20function. If x value is greater than a
- 22:39:23certain value, the function is activated
- 22:39:24and fired. Else not. Pretty
- 22:39:26straightforward. Yes, no, true, false.
- 22:39:29um I don't want to test for
- 22:39:30improbabilities. I just want a straight
- 22:39:32answer. I don't want to know if there's
- 22:39:34a partial value on there. It either is
- 22:39:36true or it's false. And the rectifier
- 22:39:38function, it is the most widely used
- 22:39:40activation function. I would debate
- 22:39:42that. Um rectifier is pretty common one,
- 22:39:45although I see that the sigmoid function
- 22:39:46is used to be the basic one, but it's up
- 22:39:48there. The rectifier function is very
- 22:39:50commonly used. You get the output of X
- 22:39:52if X is positive and zero otherwise. And
- 22:39:54you can see here again just like um uh
- 22:39:57it's either you kind of get a value
- 22:39:59going up there. So max of x of zero. So
- 22:40:02it's it's again it's like the threshold
- 22:40:03function. Yes, no, true, false. Uh it's
- 22:40:06either zero or it's uh some kind of
- 22:40:08progressive value. And then we have the
- 22:40:10rectifier function. I would argue with
- 22:40:12this because the sigmoid function used
- 22:40:13to be the most common one. But with the
- 22:40:15rectifier function, it now says it is
- 22:40:17the most commonly used or widely used
- 22:40:19activation function and gives an output
- 22:40:21of X if X is positive and zero
- 22:40:24otherwise. This is kind of nice because
- 22:40:26it now says absolutely not or it gives
- 22:40:29you a value of probability. Now, when I
- 22:40:31say a value of probability, be very
- 22:40:33careful there. I'm not saying that it's
- 22:40:34going to tell you this is 75% chance of
- 22:40:36being a circle. I'm going to tell you
- 22:40:38that it says, hey, if this says 0.1, it
- 22:40:41probably needs to be looked at or 2 or
- 22:40:433. It's going to depend on your data as
- 22:40:45to what that value means. In general,
- 22:40:48that just means it's flagging it that if
- 22:40:49it's not a one, then chances are it
- 22:40:51needs to be looked at by a person and
- 22:40:53re-evaluated. And there's a hyperbolic
- 22:40:55tangent function. This function is
- 22:40:57similar to sigmoid function is bound to
- 22:40:59a range of minus1 to 1. So you can see
- 22:41:01there's our 1 - eus 2x and 1 plus over 1
- 22:41:05+ eus 2x. Again, it's very similar to
- 22:41:07the sigmoid function. The bonus of the
- 22:41:10hyperbolic function is you have that
- 22:41:12variable coming through the middle. So
- 22:41:14again, you can look at it and you have a
- 22:41:16little bit more weight as far as you can
- 22:41:18process that down the line. That's a
- 22:41:20little bit more advanced than than what
- 22:41:21we're looking at right now. And a lot of
- 22:41:22times it's not even necessary in a lot
- 22:41:24of our different uh uses for these
- 22:41:26activation functions. Now, we looked at
- 22:41:28activation functions and I kind of said
- 22:41:30those are a little bit like a black box
- 22:41:32because even if you know all the math, a
- 22:41:35lot of times you end up just playing
- 22:41:36with them to find out what works. And it
- 22:41:38also depends on what model you're
- 22:41:39working with, whether you need a flat
- 22:41:40yes, no, true, false, or you need to
- 22:41:42have something in the middle that says,
- 22:41:44hey, this isn't quite a one. You might
- 22:41:46need to process this with the human
- 22:41:47intervention. And you could look at
- 22:41:49that. Uh, one example would be
- 22:41:50self-driving cars. You don't want a car
- 22:41:52to be yes, no, I'm going to go through
- 22:41:54the the light. You want it to be like,
- 22:41:56okay, if it's uh almost yes, maybe we
- 22:41:58stop and have human intervention so we
- 22:42:00don't get an accident. Cost function is
- 22:42:02something you can really see and measure
- 22:42:04and is very important. The cost value is
- 22:42:07the difference between the neural net's
- 22:42:08predicted output and the actual output
- 22:42:11from a set of labeled training data. So
- 22:42:13we have our group of data that's a
- 22:42:15square circle and since we're looking at
- 22:42:17geometrical shapes, we've had somebody
- 22:42:19already labeled that data. They've
- 22:42:21already said this is a triangle, this is
- 22:42:23a square. And so if this is coming up
- 22:42:25and it's giving us and it's saying a
- 22:42:26square is a triangle and it's saying a
- 22:42:28triangle is a circle, the output is
- 22:42:30wrong. And so that output can then be
- 22:42:33measured in the versus the actual output
- 22:42:35and that's the cost. Uh you might also
- 22:42:37hear this as error because that's the
- 22:42:39error value being returned. How far off
- 22:42:41is it? And what we're looking for is the
- 22:42:43least cost or the least error value. And
- 22:42:45it's obtained by making adjustments to
- 22:42:47the weights and biases iteratively
- 22:42:49throughout the training process. And
- 22:42:51this this is called back propagation.
- 22:42:54And we're going to look in that a little
- 22:42:55deeper as we look into an example. It's
- 22:42:57really hard to see when you're just
- 22:42:58looking at arrows without actual numbers
- 22:43:01and where that flow is coming from. But
- 22:43:02you can look at this is here's our
- 22:43:04inputs. They put out a prediction. The
- 22:43:06prediction comes out and says, "Hey,
- 22:43:08we've already labeled this data cuz
- 22:43:09we're in training mode and the training
- 22:43:11data is off. This is the cost. Can we
- 22:43:13send that error or that cost back and
- 22:43:16adjust those weights?" And we do it in
- 22:43:18very small increments across large
- 22:43:20amounts of data so that those weights
- 22:43:23minimize that cost or that error. But
- 22:43:25what happens within these neurons? So
- 22:43:28let's look at a little example of this.
- 22:43:29Kind of helps if you have some kind of
- 22:43:31visual. Let's build a neural network
- 22:43:32that predict bike prices based on a few
- 22:43:35of its features. And we'll see here we
- 22:43:37have our CC, our mileage, and our ABS.
- 22:43:39And these are our three input layers.
- 22:43:41And then we have the bike price and the
- 22:43:43output layer. Now, it doesn't do us very
- 22:43:45good to just uh pump it in from the
- 22:43:47beginning and pump it out. And to be
- 22:43:48honest, I would use a machine learning
- 22:43:50linear regression model on this since
- 22:43:52these are just straight numbers. But
- 22:43:54because we want a simple example, we're
- 22:43:57going to put this through and show you
- 22:43:58as a neural network what that looks
- 22:43:59like. And we got to put a hidden layer
- 22:44:01in there. The hidden layer helps in
- 22:44:02improving the output accuracy. And you
- 22:44:05could look at this as a bunch of ores.
- 22:44:08So it might say, hey, when we compare
- 22:44:10these three values on the first hidden
- 22:44:13layer neuron, we're looking at one set
- 22:44:15of features and then we might weight
- 22:44:17them in the second one. So these are a
- 22:44:18bunch of different ores kind of how the
- 22:44:20math comes out in behind the scenes. And
- 22:44:22then they go out of course to the bike
- 22:44:23trace or the output layer. And each of
- 22:44:26the connections have a weight assigned
- 22:44:27with it. And you'll see here we have a
- 22:44:30mileage CC with the weight one and
- 22:44:31weight two going into our first neuron.
- 22:44:34And you'd also have your ABS going in
- 22:44:35there. And so X1 * weight 1 + X2 *
- 22:44:39weight 2 plus the bias of one. And step
- 22:44:42two is our activation. The activation
- 22:44:44function coming in there. When does this
- 22:44:45fire? And the neuron takes a subset of
- 22:44:48the inputs and processes it. And then we
- 22:44:50go through and we do that with the um
- 22:44:52second hidden layer neuron and the third
- 22:44:54one and so on. So you process each layer
- 22:44:56in order going forward. Now when I told
- 22:44:59you this is in its infant stage, they
- 22:45:01now have neurons that fire into the same
- 22:45:04layer or back a layer so that you now
- 22:45:06have a time series and there's all kinds
- 22:45:08of wild things that they're
- 22:45:09experimenting with on these layers. This
- 22:45:11basic setup has been around since the
- 22:45:13mid90s. It's only now because of our
- 22:45:16technology that it's open to almost
- 22:45:18everybody to play with it. And that's
- 22:45:19why I say this is in an infant stage in
- 22:45:21development is this basic math is here,
- 22:45:24but what we can do with it is amazing.
- 22:45:26And what they're actually doing with all
- 22:45:27these different things is amazing. And
- 22:45:28so we're just at the beginning of how to
- 22:45:30use all these different tools and our
- 22:45:32deep learning and our neural networks.
- 22:45:34Uh and so once we have our hidden layer
- 22:45:35computed, the information reaching the
- 22:45:37neurons in the hidden layer is subjected
- 22:45:39to the respective activation function.
- 22:45:41And so each one of these fires an
- 22:45:43activation output uh and then those are
- 22:45:45each weighted to the final output layer.
- 22:45:47So the processed information is now sent
- 22:45:50to the output layer once again over
- 22:45:52weighted channels. And you could look at
- 22:45:54this as each one of these is um I always
- 22:45:56look at this as like a group of people.
- 22:45:58They're all looking at the bulletin
- 22:45:59board and the first person says this is
- 22:46:00what I project sales for the company and
- 22:46:02the second person and the third and so
- 22:46:04on. And then their perspectives are
- 22:46:06weighted based on their expertise. So
- 22:46:08your accountant might have a very high
- 22:46:10weight where the um maybe your janitor
- 22:46:12has a very low weight because their
- 22:46:14expertise is not in accounting and then
- 22:46:16that goes into the output layer and once
- 22:46:18in the output layer it goes uh the
- 22:46:20output which is the predicted value is
- 22:46:22compared against the original value. So
- 22:46:24now we have our output layer and since
- 22:46:26we have like already a list of uh bikes
- 22:46:29with their the different setups and what
- 22:46:31their value is we can now generate an
- 22:46:33error from this. The cost function
- 22:46:35determines the error in prediction and
- 22:46:37reports it back to the neural network.
- 22:46:40So this is the cost. This is how far off
- 22:46:42it is. This is your error coming back.
- 22:46:44And as you can see, this is back
- 22:46:46propagation going on. So now our error
- 22:46:48is going in reverse because we know
- 22:46:50we're not completely correct on this
- 22:46:53particular channel. The weights are
- 22:46:55adjusted in order to reduce the error.
- 22:46:57So each time we go back, we are changing
- 22:46:59those weights to reduce that error. and
- 22:47:01we change them in small increments. You
- 22:47:04don't want to fit one input. Remember,
- 22:47:07you might have a data pool with a
- 22:47:09terabyte of data. You don't want to
- 22:47:11solve for the first set of data that
- 22:47:12comes in and that be the main solution
- 22:47:14because everything else will be off.
- 22:47:16This is going to confuse you. That's
- 22:47:17also called a bias. So, we have the bias
- 22:47:20in the cell where we're adding a value,
- 22:47:22the kind of like the y intercept, and we
- 22:47:24have a bias of the whole neural network,
- 22:47:27which means that it's weighted towards
- 22:47:28one set of answers. So we want to make
- 22:47:31small changes in these weights so we
- 22:47:32don't create a bias and the weights are
- 22:47:34adjusted in order to reduce the error or
- 22:47:36the cost. The network is now trained
- 22:47:38using the new weights. Once again the
- 22:47:40cost is determined and back propagation
- 22:47:42is continued until the cost cannot be
- 22:47:44reduced any further. So let's go ahead
- 22:47:46and plug in values and see how our
- 22:47:48neural network works. So here we come in
- 22:47:51here and initially our channels are
- 22:47:52assigned with random weights. This is
- 22:47:55important because if you assign them all
- 22:47:57with the same weight, you might be able
- 22:47:58to reproduce it. But it turns out that
- 22:48:01if I put all my weights as one or all my
- 22:48:03weights as zero, it takes longer to
- 22:48:05train where if you have random weights,
- 22:48:06they already have like a little bit of
- 22:48:08adjustment and ores built in and that
- 22:48:10will give us a better answer and train
- 22:48:12faster. Our first neuron takes a value
- 22:48:14of mileage and CC as inputs. So here
- 22:48:17comes our computation whatever those
- 22:48:19inputs are. And we do that again with
- 22:48:21the second neuron with those values
- 22:48:23coming in. You can see here we have
- 22:48:24weight three and so on and then our
- 22:48:26third neuron coming down and of course
- 22:48:28our fourth neuron. So we're adding all
- 22:48:30these different values coming in here in
- 22:48:32our hidden layer. The process value from
- 22:48:34each neuron is sent to the output layer
- 22:48:36over weighted channels. So again here's
- 22:48:38our weights coming in and we have N1,
- 22:48:40N2, N3 and N4. Once again the values are
- 22:48:43subjected to the activation function and
- 22:48:45a single value is emitted as the output.
- 22:48:47On comparing the predicted value to the
- 22:48:49actual value, we clearly see that our
- 22:48:51network requires training. So, here we
- 22:48:53have it that our bike price uh we put
- 22:48:55out, we thought it was worth 2,000 on
- 22:48:57our random weights and the bike actually
- 22:48:59was $4,000 on there. Guessing that's not
- 22:49:02US dollars cuz that'd be a very
- 22:49:04expensive bike. But maybe it is. There's
- 22:49:05some $2,000 $4,000 bikes out there. The
- 22:49:08cost function is calculated and back
- 22:49:10propagation takes place. And this is
- 22:49:12pretty simple. You can look at that as
- 22:49:14our um we're subtracting one value from
- 22:49:16the other. We square it and then we take
- 22:49:18half of that and that is propagated back
- 22:49:20up. And each layer generates its own
- 22:49:23errors. Let's go back one because you
- 22:49:24have your predicted Y and your actual Y.
- 22:49:27That goes back to the first layer. And
- 22:49:29then based on the value of the cost
- 22:49:31function, certain weights are changed.
- 22:49:33So when we look at the next layer, that
- 22:49:35error is not the original 4,000 - 2,000
- 22:49:39squar / 2. This error is based on the
- 22:49:42error of each cell generated. How far
- 22:49:45off is that cell as far as its weights.
- 22:49:47We're not going to show you. It's
- 22:49:48actually a very complicated differential
- 22:49:50equation. And you can probably write it
- 22:49:51out if you wanted to. You just write out
- 22:49:53each formula that goes into the next
- 22:49:55level and you add them all together and
- 22:49:56you can write it out all the way
- 22:49:57through. Computers make it so you don't
- 22:49:59have to. And our neural network is
- 22:50:01considered trained when the value for
- 22:50:02the cost function is minimum. So when we
- 22:50:05get our error way down as low as we can,
- 22:50:07that's when our neural network is
- 22:50:08trained. And there I mean just recently
- 22:50:10they've come up with all kinds of
- 22:50:12different means for measuring that
- 22:50:14particular value. a little bit beyond
- 22:50:16the scope of today's neural network, but
- 22:50:18you can actually you can actually see
- 22:50:19it. You know, how far do you do this
- 22:50:21until the neural network doesn't need to
- 22:50:23be trained anymore and you can overtrain
- 22:50:25a neural network. Now, the tools that
- 22:50:27we're looking at automatically let you
- 22:50:29know when to stop, which is really nice.
- 22:50:31And that is just like I said, we're at
- 22:50:33the beginning stages in neural networks
- 22:50:34and it's just really cool what they can
- 22:50:36do now and how much of it's automated
- 22:50:38and how much of it is experimental.
- 22:50:40Right now, let's take a look at gradient
- 22:50:41descent. But what approach do we take to
- 22:50:44minimize the cost function? So here we
- 22:50:46have nice error thing coming in. This is
- 22:50:49our cost or our error. Uh let's start
- 22:50:51with plotting the cost function against
- 22:50:53the predicted value. And so you can see
- 22:50:55they fed in multiple y's and these are
- 22:50:58the errors coming in and the cost of
- 22:51:00each of these inputs and changes going
- 22:51:02on. Note we start at a random point on
- 22:51:05the curve. So usually you put in you
- 22:51:08know you pick up your data and you
- 22:51:09randomly pick where to start in your
- 22:51:11data. A lot of times you just run it
- 22:51:12from the beginning because you're going
- 22:51:13through so much data, it's not that big
- 22:51:15of a deal. But you start with one point
- 22:51:17going in. So your forward propagation
- 22:51:19goes through. You're going to go ahead
- 22:51:20and find your cost or your error. It
- 22:51:22points that on the curve. And you can
- 22:51:24see how we're plotting it right here.
- 22:51:25Since the gradient at this point is
- 22:51:27positive, we may move right. So we're
- 22:51:29going to move a little bit to the right
- 22:51:30on here. And this time the gradient is
- 22:51:32negative. We move a little bit to the
- 22:51:33left. Eventually we try out the point
- 22:51:36where the gradient is zero. This is a
- 22:51:38least value of cost function. You have
- 22:51:41to be a little careful with this because
- 22:51:43this particular I mean they make it look
- 22:51:45nice and simple in this graph. Sometimes
- 22:51:48these curves look like stair steps and
- 22:51:50so there is global minimums and then
- 22:51:53there is local there might be a local
- 22:51:55point where the gradient is zero but
- 22:51:57it's not the global one. Uh so it might
- 22:51:59be way off to the left where it just
- 22:52:01happens to step down a little bit and
- 22:52:02you think you're in the right gradient.
- 22:52:04And with that we have all the right
- 22:52:05weights and we can say our network is
- 22:52:07trained. So here we have um just some
- 22:52:10major these are some of the big names
- 22:52:12out there right now in development for
- 22:52:14deep learning platforms. TensorFlow
- 22:52:16which we'll actually do an example in in
- 22:52:18a minute. Deep learning for J which is
- 22:52:20in the Java platform. Uh so if you're a
- 22:52:22Java programmer uh by the way is
- 22:52:24TensorFlow is accessed most people are
- 22:52:26using Python to access it but it is a
- 22:52:28system that's kind of separate from a
- 22:52:29lot of the programming languages which
- 22:52:31makes it a lot more um flexible as far
- 22:52:33as use. Deep learning forj is Java based
- 22:52:36and then cross is just exploding right
- 22:52:38now. And this is interesting. Cross is
- 22:52:40uh working with TensorFlow. It actually
- 22:52:42can sit on top of TensorFlow. And it can
- 22:52:44also do its own thing. Uh so if you're
- 22:52:47studying deep learning, you're getting
- 22:52:48into it, you want to know the basics of
- 22:52:50TensorFlow, but you also are going to
- 22:52:52want to know the upper level of KAS
- 22:52:53sitting on top of TensorFlow. We're just
- 22:52:55looking at TensorFlow today though in
- 22:52:57our example. And there's also Torch on
- 22:52:59there. There's a bunch more that we
- 22:53:00didn't list on here. Um even sklearn or
- 22:53:03the uh side package in Python has a
- 22:53:06neural network you can program a very
- 22:53:08basic one and it is the same basic one
- 22:53:10that you could do in TensorFlow if you
- 22:53:12stripped everything out of it and then
- 22:53:13TensorFlow has a lot of tools they've
- 22:53:15added in and so has KAS but we're going
- 22:53:17to be looking specifically at TensorFlow
- 22:53:18in our example and TensorFlow is an
- 22:53:21open- source tool used to define and run
- 22:53:23computations on what they call tensors
- 22:53:26very common language now so you more and
- 22:53:28more we see the term tensor as being a
- 22:53:30standard in the uh deep learning
- 22:53:32language and this was originally
- 22:53:34developed by Google. So let's dig a
- 22:53:36little bit big in there. What are
- 22:53:37tensors? Tensors are just another name
- 22:53:40for arrays. So a tensor of dimension
- 22:53:43five. You can see here we have ab kmq
- 22:53:46whatever. So it's an array coming in.
- 22:53:47And the tensor of dimension 54 more like
- 22:53:50a picture. Very common to see that in a
- 22:53:52picture. You can also see a tensor even
- 22:53:54more detailed than a picture as we go to
- 22:53:56the next one. Tensor of dimension 333.
- 22:53:58This is 3D space. You might have a
- 22:54:00picture that also has colors. That might
- 22:54:02be the third dimension. You might have
- 22:54:04four dimensions because you have both
- 22:54:06your grid and your different color
- 22:54:08channels and your zplot. You can see
- 22:54:11where you can now process a very
- 22:54:13highlevel set of data coming in whether
- 22:54:16as an image or features. They could be
- 22:54:18features that have nothing to do with
- 22:54:19images. So there's a lot of stuff you
- 22:54:21can do now with the tensors coming in.
- 22:54:23Thus this where the term tensorflow
- 22:54:25comes from. So we have um right now the
- 22:54:27TensorFlow is the most popular library
- 22:54:29in deep learning and I did mention KAS
- 22:54:31now works with TensorFlow. So there's a
- 22:54:33lot of stuff you can do between the two.
- 22:54:35Uh it's an open-source software library
- 22:54:37developed by Google. Uh so they hit a
- 22:54:40roadblock and they realized hey this is
- 22:54:43an infant stage technology. You know we
- 22:54:45thought it was going to be the next
- 22:54:46greatest thing and we were going to have
- 22:54:48a hold on it but it's really infant as
- 22:54:50far as how it's applied and what we can
- 22:54:52do with it. Let's open source it so
- 22:54:53everybody can work on it. uh let's take
- 22:54:55it to the next level. And that's really
- 22:54:56what open source does to a lot of these
- 22:54:58uh packages when they release them. And
- 22:55:00you can run on either a CPU or a GPU. So
- 22:55:03when we look at the details, if you have
- 22:55:05your graphic processing units, um what's
- 22:55:08nice about those is they run a lot
- 22:55:10faster. The downside is you have to play
- 22:55:12with them a little bit to get them up
- 22:55:13and running. And it's a hardware
- 22:55:14upgrade. When we run it, I'll be running
- 22:55:16it in the CPU mode. I have played with
- 22:55:18it in my GPU on my personal computer.
- 22:55:21you know, it does increase the
- 22:55:22processing. Uh, but I did run into some
- 22:55:24version problems with my Python and
- 22:55:26stuff like that. And when I did finally
- 22:55:28work it out, I went back to the CPU
- 22:55:29because it didn't increase my speed
- 22:55:31enough for what I was working on. But in
- 22:55:32a larger group, you might be able put
- 22:55:34that on. If you're working with a larger
- 22:55:36stack of computers, you might want to
- 22:55:37run it in the GPU. You can create a data
- 22:55:39flow graphs that have nodes and edges.
- 22:55:42So there's our edges coming in. We
- 22:55:44didn't talk about edges, but that's very
- 22:55:46up and cominging way of looking at your
- 22:55:48analytical data is how do different
- 22:55:50nodes connect? What do those edges look
- 22:55:52like in between them? And it's used for
- 22:55:54machine learning applications such as
- 22:55:55neural networks. It is mostly a neural
- 22:55:58network, but they have all kinds of
- 22:56:00tools which sit on top of our basic
- 22:56:02neural network. They have new stuff
- 22:56:03evolving into the TensorFlow library.
- 22:56:06So, it's very much uh just exploding.
- 22:56:09great time to jump into TensorFlow
- 22:56:10because there's all kinds of cool things
- 22:56:12we're doing with it and all kinds of
- 22:56:13cool applications you can now use uh
- 22:56:15TensorFlow for. So let's take a look at
- 22:56:17implementation in TensorFlow and we're
- 22:56:20going to build a neural network to
- 22:56:21identify handwritten digits using the uh
- 22:56:24Mnest database or the MNIST database and
- 22:56:28that stands for modified National
- 22:56:30Institute of Standards and Technology
- 22:56:32database. It is a collection of 70,000
- 22:56:34handwritten digits and the digit labels
- 22:56:37identify each of the digits from 0ero to
- 22:56:39nine. This is a cool example because
- 22:56:41it's simple enough that you could
- 22:56:43actually run this through some basic
- 22:56:45machine learning categorizing algorithms
- 22:56:48and train them and you'll get about the
- 22:56:50same answer because again it's it's
- 22:56:51simple grid. The digits on the grid
- 22:56:53don't have a huge amount of variation
- 22:56:55like you would say an automated driving
- 22:56:57car looking at the environment. So you
- 22:56:59can still do this with a lot of your um
- 22:57:02different linear models and stuff like
- 22:57:03that. You can solve this and you'll get
- 22:57:05about the same answer. When I ran a
- 22:57:06comparison between TensorFlow and
- 22:57:08between some basic uh regression models
- 22:57:11or category models uh in machine
- 22:57:13learning, they came up pretty even as
- 22:57:15far as their output. Uh so this is kind
- 22:57:17of where we start to see the complexity
- 22:57:20of something coming in this case a
- 22:57:22tensor you know or a grid of uh
- 22:57:24information where the deep learning
- 22:57:26model does as good as the regular models
- 22:57:29and when you get past this kind of
- 22:57:31complexity and features suddenly the
- 22:57:34neural networks come up with better
- 22:57:36answers better solutions and a better
- 22:57:38build and that's why there's such a move
- 22:57:40into neural networks is we live in a
- 22:57:42complicated world and it's just really
- 22:57:44cool we can do with this. So the
- 22:57:45handwritten digits from the um NIST
- 22:57:47database, they come in, the data set is
- 22:57:50used to train the machine, a new image
- 22:57:52of a digit is fed and the digit is
- 22:57:55identified. Um and if you've looked at
- 22:57:56any of our other machine learning tools
- 22:57:59where we're doing training, uh where we
- 22:58:00train our uh model to fit and then you
- 22:58:03test it out, this should look pretty
- 22:58:05familiar. Uh and there is some tools out
- 22:58:07there for say untrained categorizing uh
- 22:58:09where it's just looking for features
- 22:58:11that fit together. So there are tools
- 22:58:12that don't need that training. But this
- 22:58:14is where uh when we talk about neural
- 22:58:16networks, we do need to train them. And
- 22:58:17this is what we're looking at.
- 22:58:33So for this I'm going to use the
- 22:58:35Anaconda Navigator just because it's a
- 22:58:38very nice visual tool. You might be in
- 22:58:40PyCharm or one of your other IDEs for
- 22:58:43editing Python because we are looking at
- 22:58:45Python TensorFlow. And under Anaconda,
- 22:58:47we have the notebook, which is something
- 22:58:49we use pretty regularly. And they have
- 22:58:51the Jupyter Lab. The Jupyter Lab is the
- 22:58:53Jupyter notebook, but with tabs and a
- 22:58:55few new features. So, we'll be using the
- 22:58:57Jupyter Lab today. And under the
- 22:58:59environment, you'll want to go ahead and
- 22:59:01and uh if you haven't yet, uh you'll see
- 22:59:04that I have a number of different setups
- 22:59:05in here. Right now I have the Python
- 22:59:08version 36 and the TensorFlow. In this
- 22:59:12case I have TensorFlow 1.12. If we
- 22:59:14scroll down you can see that uh here we
- 22:59:16go. TensorFlow and it's version 1.12.
- 22:59:19And in here if you haven't yet you'll
- 22:59:21need to install those and go in and just
- 22:59:23open our terminal. And u if you've never
- 22:59:26used the Anaconda or if you're in your
- 22:59:29other thing you might have something
- 22:59:30simple like pip. Is what I use for my
- 22:59:33install. And you can simply do install
- 22:59:36TensorFlow. And that should bring in the
- 22:59:38most current version. Now, when I
- 22:59:40installed this a few months ago, Python
- 22:59:42version, I'm not going to run this
- 22:59:44because I already have installed on
- 22:59:45here. Python version 3.7, the newest one
- 22:59:48out, still had a couple glitches with
- 22:59:51the TensorFlow. I believe they've fixed
- 22:59:53it as of writing of this, but um I'm
- 22:59:55going to stick with 3.6 just so I don't
- 22:59:56get any surprises on there. So, this is
- 22:59:58Python version 36 with TensorFlow 1.12
- 23:00:02on here. And if you haven't installed it
- 23:00:04yet, you also want to install Numpy for
- 23:00:06this example. That's Numbers Python or
- 23:00:08uh nu py. You can just simply run an
- 23:00:12install on there. Keep in mind if you're
- 23:00:14in Anaconda uh and you've created one of
- 23:00:16these environments specific to this,
- 23:00:18keep withd.
- 23:00:21If you're going to use pip, keep with
- 23:00:22pip. Don't install one package with pip
- 23:00:24and one under because that's how they
- 23:00:26track those version numbers and how they
- 23:00:28fit together and you can end up with a
- 23:00:30problem. they don't pip doesn't see cond
- 23:00:32and vice versa. Uh so just keep that in
- 23:00:34mind when you're running your installs.
- 23:00:35We'll go ahead and open up Jupyter Lab
- 23:00:37and we're going to launch that. So
- 23:00:39here's my Jupyter Lab. One of the really
- 23:00:41cool features of Jupyter Lab is you have
- 23:00:42tabs now. So you can open up multiple uh
- 23:00:45notebooks. And this is nice cuz I have
- 23:00:46my notes I'm working on and then our
- 23:00:48actual window we're looking in. And
- 23:00:50we'll go ahead and zoom in a little bit
- 23:00:52here. There we go. So you have a nice u
- 23:00:54hopefully easy to see fonts. And then
- 23:00:56we'll go ahead and do a simple or get
- 23:00:58our imports out of the way. Um, and so
- 23:00:59we're going to import our TensorFlow as
- 23:01:01TF. Uh, that's pretty much a standard
- 23:01:03for TensorFlow, numpy, our numbers
- 23:01:06Python as py, and we'll import our matt
- 23:01:10plot library as plt. Again, these are
- 23:01:13very common. So if you see TF or py or
- 23:01:16plt, this is a standard that most people
- 23:01:18use. Do you have to? No, you could just
- 23:01:21do import numpy instead of doing as py.
- 23:01:23And then from
- 23:01:25tensorflow.acamples.tutorial
- 23:01:28tutorials. This is always nice because
- 23:01:30they actually include data set we're
- 23:01:32going to play with. So, we're going to
- 23:01:34import input data. So, there's our data
- 23:01:37coming in. That's all we're doing is
- 23:01:39telling it this is where it's coming
- 23:01:40from. And if we're going to tell where
- 23:01:41it's coming from, we need to go ahead
- 23:01:42and create a variable with that
- 23:01:44information in it. And we'll just call
- 23:01:45this uh mnist
- 23:01:48or minced. You know, I don't really know
- 23:01:49how they pronounce that. I should
- 23:01:50probably look that up. It's a very
- 23:01:52common data set to use. And there's our
- 23:01:54input data. And we're going to read data
- 23:01:57sets. And this is um if you look at
- 23:01:59this, we imported input data from our
- 23:02:01TensorFlow. And so this is a TensorFlow
- 23:02:05read statement for their tutorials. So
- 23:02:07this isn't like some special Python
- 23:02:09setup. This is just their setup. Makes
- 23:02:11it easy to pull it in. So once we get
- 23:02:13into their data sets, we need to go
- 23:02:14ahead and tell it what kind of data set.
- 23:02:16And again, this is what we brought in,
- 23:02:18but it's going to be the nint data. And
- 23:02:20this part is very important. one hot
- 23:02:24equals true. This means that instead of
- 23:02:27importing a value from 0 to 9, we
- 23:02:31evaluate the data set. It's going to
- 23:02:33bring it in as one hot. Whenever you see
- 23:02:35one hot encoder, we're flattening that
- 23:02:37out. And we have true false for zero,
- 23:02:40true false for one, true false for two.
- 23:02:42So our output, if you remember from our
- 23:02:44output uh from the slide we did earlier,
- 23:02:46uh in this case, I grabbed the one for
- 23:02:48bike price. Doesn't really matter which
- 23:02:50one we use. This has one output. So we
- 23:02:52have our bike price on this. We're going
- 23:02:54to have instead of one output, we're
- 23:02:56going to have 10 outputs representing
- 23:02:58each of the digits in there. And this
- 23:03:00code really isn't going to show us
- 23:03:02anything. It's good to see what we're
- 23:03:03actually looking at. So um let's go
- 23:03:05ahead and do a figure ax equals go into
- 23:03:08our plot library subplots 10, 10. And
- 23:03:13that is if you remember we talked about
- 23:03:15tensor. Tensor being data coming in.
- 23:03:17This is a 10 by 10 grid or 100 pixels on
- 23:03:20there. And if we're going to display it,
- 23:03:22uh let's go do K0
- 23:03:24for I and range 10. Just a simple loop
- 23:03:28through on the data. Let's do what is
- 23:03:30it? Uh for J and range 10. And I
- 23:03:33actually misqued that 10 uh 10 x 10 is
- 23:03:36not the actual size of the pixels. Uh
- 23:03:38the actual pixels are going to be um we
- 23:03:40look at the shapes and we'll get into
- 23:03:41that in just a second here. We'll take a
- 23:03:42quick look at shape on there. Uh turns
- 23:03:45out they're uh what are they? are, I
- 23:03:47believe, 28x 28. Uh, so let's take a
- 23:03:50look at that. And we're just going to
- 23:03:51plot these. What are we looking at? What
- 23:03:53are we working with? As a data
- 23:03:54scientist, you should always be looking
- 23:03:56back at your data and seeing what it
- 23:03:59looks like and get that human
- 23:04:00perspective because you just never know.
- 23:04:02You know, the the computer may put
- 23:04:04something out that looks makes no sense.
- 23:04:06And at that point, you want to go back
- 23:04:08and reevaluate what you did. Uh, so
- 23:04:10we're going to go ahead and plot. We're
- 23:04:11going to plot 10 digit, you know, 10 of
- 23:04:13the digits by 10 of the digits. And
- 23:04:14here's our ax. We'll create the J on our
- 23:04:17subplots and we're going to do an image
- 23:04:19show. We're going to look at the
- 23:04:20training image for images of K. And then
- 23:04:24we want to reshape this. We're going to
- 23:04:25reshape this. And we're going to reshape
- 23:04:27this 28x 28. That's how I knew I had it
- 23:04:30wrong is cuz I looked down my notes. I
- 23:04:31was like, oh no, that says 28. It's not
- 23:04:3310 x 10. And I should know that already
- 23:04:34cuz I've done enough messing with this
- 23:04:36data set that I should have remembered.
- 23:04:37Uh, but it's 28 x 28. And the aspect
- 23:04:40we're going to do is auto. And this is
- 23:04:41all, if you look at this, here's our
- 23:04:43variable NIST. the NIST is coming from
- 23:04:46data set. Uh so this is all part of the
- 23:04:48TF TensorFlow learning or examples
- 23:04:51tutorial in there. And then we'll go
- 23:04:52ahead and do K plus equals 1. So we just
- 23:04:56keep paging through our different um
- 23:04:58images. And let's see what that looks
- 23:04:59like. Let's go ahead and do a plot show.
- 23:05:02Uh and we'll go ahead and run this so we
- 23:05:03can take a look and see what we have
- 23:05:04here. And so we have a nice plot here.
- 23:05:06And you can just see that we have uh
- 23:05:08some random numbers showing up in each
- 23:05:09one of these little subplots. If you're
- 23:05:11wanting a copy of this code, put a note
- 23:05:14down in the YouTube video and let us
- 23:05:16know or come visit us at
- 23:05:17www.simplearn.com
- 23:05:19and we'll send you out a copy of what
- 23:05:20we're working on and get a copy of that
- 23:05:22for your own setup. Uh so now we've
- 23:05:24taken a look and we can just see we have
- 23:05:26here's our pictures that are coming on.
- 23:05:27We plotted them so we have an idea of
- 23:05:29what we're looking at. Let's go ahead
- 23:05:31and uh print. Let's look at the shape of
- 23:05:33the features. Uh so when we have this we
- 23:05:35have our nest train images and we'll do
- 23:05:39the shape on there. Let's take a look
- 23:05:41and just see what we're looking at uh as
- 23:05:43far as uh our count and everything. And
- 23:05:45so you can see here we have 55,000.
- 23:05:49That's basically how many images we have
- 23:05:50and this by 784. And in this data set
- 23:05:53there's also our labels. So let's take a
- 23:05:55look at that. We have our net train
- 23:05:57labels shape. Let's take a look and see
- 23:05:59what that looks like. Uh and there we
- 23:06:01have 10 because there's 10 digits. So we
- 23:06:03brought in that's our output we're
- 23:06:04looking at. And so we have there we go
- 23:06:0655,000. They match. They should match
- 23:06:08because you should have equal numbers in
- 23:06:10both of those. You know, here's our data
- 23:06:12in and here's our answer. If you
- 23:06:14remember, this is a bunch of zeros and
- 23:06:16with one each each one will be 0001
- 23:06:19would be what letter four or something
- 23:06:21like that. So, let's take a look at what
- 23:06:22our one hot encoding did for the first
- 23:06:24observation. And this is when we're
- 23:06:26exploring data, you really want to dig
- 23:06:28in there and just see what the heck am I
- 23:06:30looking at. So, we're going to look at
- 23:06:31the labels. And this would be the first
- 23:06:33label that comes up. And we'll go ahead
- 23:06:35and run this. And we look at that. You
- 23:06:37can see this is what I'm talking about.
- 23:06:380 0 or 1 is 0 2 is 0 3 is 0 four is 0 5
- 23:06:43is 0 6 is 0. 7 equals 1. So our very
- 23:06:47first label is a seven. But our very
- 23:06:49first label comes up that it's a seven.
- 23:06:51And so we don't have like 0 through 9.
- 23:06:54We have a bunch of zeros and just the
- 23:06:56one to mark it as a seven on here. So
- 23:06:58now we've kind of looked a quick look at
- 23:07:00the data. And in here you might ask some
- 23:07:02questions like what is 784? 24 * 24.
- 23:07:05Remember that's the size of our grid on
- 23:07:08there or our tensor coming in. So 784 is
- 23:07:10a setup on there. And we've gone through
- 23:07:12all this viewing the data. We'll go
- 23:07:14ahead and start looking at our
- 23:07:16tensorflow. So let's take our X
- 23:07:18variable. This is going to be our
- 23:07:19training set. We'll do a placeholder and
- 23:07:22then we're going to have these come in
- 23:07:24as float. Now if I remember correctly,
- 23:07:26they're actually, you know, zero or one
- 23:07:28for the values because they're either
- 23:07:30but we have them coming in as a float
- 23:07:31value. And we have a little bit of a
- 23:07:33shape coming in here. And there's our
- 23:07:34784. Uh so we let it know that this is
- 23:07:37what's what our input is for our
- 23:07:39TensorFlow. And this is our training
- 23:07:41set. So we'll just put a label on there
- 23:07:43to help us uh track that train set. And
- 23:07:46then W. And with W, we'll go ahead and
- 23:07:48do TF variables. And we'll do this as uh
- 23:07:51zeros variables TF zeros. And we'll set
- 23:07:54this as as 784 by 10. 10 being the
- 23:07:58output. 784 being our number of
- 23:08:01variables in and this is our weights.
- 23:08:04Remember we have a bias in there too.
- 23:08:06And I'll go back over this in just a
- 23:08:08second as we see how that fits together
- 23:08:11in our tensorflow. And we'll do this one
- 23:08:13um with our variables again. We have 10.
- 23:08:15So we're going to do the bias. We're
- 23:08:17going do it the same kind of format and
- 23:08:19setup on here. And so we'll do that as
- 23:08:22as TF zeros of 10. So we'll just create
- 23:08:25an array of 10 there. And this is our
- 23:08:26bias. So with these three lines um and
- 23:08:30there's actually they're coming out with
- 23:08:32the eager execution which would bypass
- 23:08:35some of what we're doing. But this is
- 23:08:36important to understand is the first
- 23:08:37thing you have to do with TensorFlow is
- 23:08:39we have to allocate a space for the
- 23:08:41variables and our TF placeholder and our
- 23:08:44TF variable with our weights and our
- 23:08:46biases. This actually hasn't done
- 23:08:48anything yet. So all it is is
- 23:08:49placeholders. That's why it's okay to
- 23:08:51use zeros. Um you could have just as
- 23:08:53easily used ones or anything else and it
- 23:08:55wouldn't matter. The next stage is to go
- 23:08:57ahead and set up some of the functions
- 23:09:00going on. But before we do that, just
- 23:09:02note that this hasn't done anything.
- 23:09:03Even if I execute it, all it's done is
- 23:09:05created placeholders until we do the
- 23:09:07final initialization. And so we need to
- 23:09:09go ahead and set up. We'll do y=
- 23:09:11tf.n.oftmax.
- 23:09:14And the code for this is tf.mmoxw.
- 23:09:18And this is our uh sum. Let's just put a
- 23:09:21note here so we can keep track of what's
- 23:09:22going on. We're finding weighted sum of
- 23:09:25inputs plus the bias. Uh so there's our
- 23:09:29plus b the bias and then we need to go
- 23:09:31ahead keep um let's do y underscore and
- 23:09:34again another placeholder and this one
- 23:09:36we'll set um it actually we'll put in as
- 23:09:38tf placeholder on here tf placeholder
- 23:09:41float none 10. There's our one hot
- 23:09:43encoder going on there. So our 10 values
- 23:09:46coming out and we'll do a cross entropy
- 23:09:50on here and this is going to be minus tf
- 23:09:52reduce sum and we'll do y here's our y
- 23:09:55underscore which is remember we have
- 23:09:57your y output and your actual output. Uh
- 23:10:00so this will be our y underscore time
- 23:10:02the tf log of y. And then finally um
- 23:10:06before we do the actual initialization
- 23:10:08of all our variables we'll set up our
- 23:10:10train step. This equals our gradient
- 23:10:12descent optimizer. Very important.
- 23:10:14Remember we looked at that chart on our
- 23:10:16um uh slides and so we've set up all
- 23:10:18these formulas and here's our gradient
- 23:10:20descent optimizer and as it keeps
- 23:10:22looking it keeps looking for that zero
- 23:10:24value. That's what we're doing with that
- 23:10:26particular formula. So let's take a look
- 23:10:28and see what we're doing here. We just
- 23:10:30put together all of our pieces for
- 23:10:32TensorFlow. And you know the devil's in
- 23:10:34the details. We have here our training
- 23:10:37set coming in. We have to put a
- 23:10:39placeholder on there. We have our uh
- 23:10:42variables with their weights. We have
- 23:10:44our biases coming out and then we put in
- 23:10:46our uh the weighted sum. So here's
- 23:10:49summizing our summation here. Then we
- 23:10:52have our y variable output. So there's
- 23:10:54our y um how it works and then of course
- 23:10:56the actual output on there. And then we
- 23:10:58have our cross entropy coming in and
- 23:11:01that's our minus tf.reduce sum the y *
- 23:11:04the tf log of y. And then the training
- 23:11:06step gradient descent optimizer and
- 23:11:08we're using a 0.01 in this and we're
- 23:11:10going to minimize cross entropy. So,
- 23:11:12we're going to let it do all the work.
- 23:11:14So, once we've set up all of these
- 23:11:16different layers, we've allocated for
- 23:11:18them, we need to go ahead and initialize
- 23:11:20them. So, we're going to do an init tf
- 23:11:22initialize, and it's going to be all
- 23:11:24variables. Uh, one of the cool things is
- 23:11:26they're in the process of doing away
- 23:11:28with this. So, all these steps would be
- 23:11:30bundled into one instead of having to
- 23:11:32have placeholders. You initialize them
- 23:11:33in the same process going on. And then
- 23:11:36finally, everything in TensorFlow is
- 23:11:38based on your session. Now, this is
- 23:11:40changing that there's other options to
- 23:11:42be able to run this, but we want to go
- 23:11:43ahead and do uh session. There's our TF.
- 23:11:46There's a TF session. And then we want
- 23:11:48to go ahead and do session run. And what
- 23:11:50are we going to run? Well, we did
- 23:11:51initialization of all our variables. Uh
- 23:11:54so, this is what we're running. And this
- 23:11:55is we're actually once we do this, we
- 23:11:57actually create our TensorFlow object.
- 23:12:00So, this whole piece of code right here
- 23:12:02is our TensorFlow object. We have our
- 23:12:05input coming in with our weighted
- 23:12:07variables coming in. Our soft max for
- 23:12:10our metal going out. How does it add it
- 23:12:12together for our y value? Uh and then we
- 23:12:14have the actual uh float value coming
- 23:12:17out. Checking on our all the way down.
- 23:12:19So you can see all the different stages
- 23:12:20going through that we're setting up. Um
- 23:12:22and this is one of the reasons that a
- 23:12:23lot of people like TensorFlow is because
- 23:12:26you can designate all these different
- 23:12:28pieces one step at a time. This is also
- 23:12:30one of the reasons people don't like
- 23:12:32TensorFlow is because you have to
- 23:12:34designate all the different layers
- 23:12:36coming down and there's a lot of steps
- 23:12:38being made right now to minimize this to
- 23:12:40make it either easier to automate it or
- 23:12:43to allow you to do more complicated
- 23:12:45things and all those steps are still at
- 23:12:47play. So it's worth looking into the
- 23:12:49more advanced version what's going on
- 23:12:51with KAS on top of TensorFlow. It's also
- 23:12:53important to understand what's going on
- 23:12:55in these individual levels if you're
- 23:12:56going to play with them. It's important
- 23:12:58to understand, hey, what's going on with
- 23:13:00the uh finding the weighted sum of the
- 23:13:02inputs plus the bias because there's
- 23:13:04other ways to do that. There's all kinds
- 23:13:06of other tools in there now, but this is
- 23:13:07the basic setup that you want to do on a
- 23:13:09TensorFlow coming in. And we want to go
- 23:13:11ahead and just run and admit our
- 23:13:12TensorFlow. So, let's go ahead and do
- 23:13:14that. Let's run this. We do get a
- 23:13:16warning here because uh there's a move
- 23:13:18to use global variables. This is one of
- 23:13:20the changes they're making, but it as
- 23:13:22far as this example, it's not going to
- 23:13:23make a difference because we're doing
- 23:13:25once we initialize it. This is
- 23:13:26initializing our variables. And again,
- 23:13:28these are only placeholders up here
- 23:13:29until we initialize them. And I would
- 23:13:31highly suggest put a note down there or
- 23:13:33or go over to simplylearn.com and let
- 23:13:35them know and have them email you a copy
- 23:13:38of the code. So, you can actually play
- 23:13:39with this code right here because this
- 23:13:41is the body of what's going on in
- 23:13:43TensorFlow. This is the build in neural
- 23:13:46networks. And then once we've done that,
- 23:13:48now comes kind of the fun part is we
- 23:13:50need to go ahead and train it. Uh so
- 23:13:52we've created our TensorFlow, we've
- 23:13:53created our uh network and now we need
- 23:13:56to go ahead and train it. Uh so let's
- 23:13:57put together that training code and
- 23:13:59let's just do uh for I in range u 0 to
- 23:14:031,000. So we're just going to look at uh
- 23:14:05the first 10,000 in our training. And
- 23:14:07the way we pull that data from our mints
- 23:14:10train next batch of 100. Uh so you look
- 23:14:13at this. We're going to be doing groups
- 23:14:14of 100 and then there's going to go
- 23:14:16through a thousand of them. This is very
- 23:14:18important that TensorFlow builds this
- 23:14:21in. This is one of the downsides of
- 23:14:23doing sklearn or one of the older
- 23:14:25packages is they don't let you batch
- 23:14:27groups in. Uh they wanted to have it all
- 23:14:29up front and then you have to build your
- 23:14:31own batch programs right now. Uh this
- 23:14:33lets us go ahead and do that. And you
- 23:14:34can see here we have batch x of s, batch
- 23:14:36y of s. So there's our x and our y. You
- 23:14:39could look at this as our training of X
- 23:14:41and our train of Y or the data N and the
- 23:14:45answer in. Uh and then we simply do our
- 23:14:47session run. Uh so here's our session
- 23:14:49that we've created. We're going to run
- 23:14:50it and we want to do the train step. We
- 23:14:53initialized our train step up here and
- 23:14:55our TF. And so there's our train step
- 23:14:58feed. It's a dictionary. Dictionary
- 23:15:00coming in which we're going to create
- 23:15:01right here. uh is x is our batch x of
- 23:15:06our sample comma and our y underscore is
- 23:15:09going to be our batch of our y sample.
- 23:15:12Uh and so this goes through and we've
- 23:15:14now hit the run button and we've trained
- 23:15:16our session. We've trained this setup on
- 23:15:19here. And once we've trained it, then we
- 23:15:21need to go ahead and find out how good
- 23:15:22our accuracy was and actually start
- 23:15:24running some predictions through there.
- 23:15:25Uh so we'll go ahead and create a a
- 23:15:27correct prediction. And this is where
- 23:15:29our tf.equal equal. We'll use our argmax
- 23:15:31y of one and tf argmax of y of
- 23:15:34underscore of one to help us get the
- 23:15:35correct predictions on there. And then
- 23:15:37we want to use that to feed into an
- 23:15:39accuracy. And so our accuracy is going
- 23:15:42to be tf reduce uh mean and we'll take
- 23:15:45that and we'll do um a cast and this is
- 23:15:48the correct prediction that we're
- 23:15:49sending in there. And it is a uh float
- 23:15:52value. Keep it simple. And let's go
- 23:15:54ahead and print this out so we can see
- 23:15:56what we're looking at. Uh so what are we
- 23:15:57printing out? Uh we need to do a session
- 23:15:59run. This session run is going to be on
- 23:16:01the accuracy. Where did accuracy comes
- 23:16:03from? This is our we're casting our TF
- 23:16:05on there with the correct predictions on
- 23:16:07that. So here's our accuracy feed in. So
- 23:16:10it needs a dictionary for the data
- 23:16:12coming in. We're going to create our
- 23:16:13dictionary and x is going to be our nest
- 23:16:16test images and y there is going to be
- 23:16:21our nest.est
- 23:16:23labels. Let me just double check and
- 23:16:25make sure I have that typed in there
- 23:16:26correctly. There we go. Oh, and let's go
- 23:16:28ahead and run that and see what comes
- 23:16:29up. And we end up with a N165
- 23:16:33for our accuracy, which means our
- 23:16:36trained neural network does a pretty
- 23:16:38good job letting us know what these
- 23:16:40different symbols are in guessing that a
- 23:16:42seven and a three and a four, uh,
- 23:16:44something that as humans we kind of take
- 23:16:46for granted. I even have trouble reading
- 23:16:48this. So, I don't know if I would be
- 23:16:49able like that first one, I would sit
- 23:16:51there for a long time figuring out
- 23:16:52that's a seven versus a two. That could
- 23:16:54have easily been a two to me. Sing seem
- 23:16:56to do a pretty good job analyzing this
- 23:16:58data. And this is used to analyze
- 23:17:00something very complicated on these
- 23:17:01images, very different than uh just a
- 23:17:04straight value of uh cost of sales and
- 23:17:07here's our return and our marketing. Uh
- 23:17:09we can now create this nice neural
- 23:17:10network that does all kinds of cool
- 23:17:12things. Do you know friends that
- 23:17:14according to the lending statistics the
- 23:17:16demand for AI and ML specialist is
- 23:17:18projected to surge by 40% between 2023
- 23:17:22to 2027.
- 23:17:24And on an average, an ML engineer is
- 23:17:27expected to earn around 133 and $336 per
- 23:17:31year. So if you are an aspiring ML
- 23:17:34engineer and thinking about what
- 23:17:36innovative projects you can show in your
- 23:17:38portfolio, then your wait is over cuz in
- 23:17:41this video I'll be covering eight
- 23:17:43amazing ML projects that you can
- 23:17:45showcase in your resume. So guys, let's
- 23:17:48start first with a beginner level
- 23:17:49project and the first project that we
- 23:17:51are going to encounter that is home
- 23:17:53value prediction. So guys, this project
- 23:17:56aims to develop a predictive model to
- 23:17:58estimate the value of residential
- 23:17:59properties. The model will analyze
- 23:18:02various features such as location,
- 23:18:04square, footage, number of bedrooms and
- 23:18:06bathrooms, age of the property and other
- 23:18:09relevant factors. By leveraging
- 23:18:11historical property data, the model will
- 23:18:14be able to provide accurate home value
- 23:18:15predictions which can be useful for real
- 23:18:18estate agents, buyers and sellers. So
- 23:18:21guys, the programming language that we
- 23:18:23are going to use all over here will be
- 23:18:25Python and machine learning libraries
- 23:18:27that we will be using will be
- 23:18:28scikitlearn, tensorflow, kas and for
- 23:18:31data handling libraries we have pandas,
- 23:18:34numpy and for visualization we have to
- 23:18:37use mattplot and seabon. Now what will
- 23:18:39be the approach for this one guys? So
- 23:18:42guys the first one that we have a data
- 23:18:44collection. So here what is going to
- 23:18:46happen guys? So first you have to
- 23:18:48collect the historical property data
- 23:18:49from the sources like Zillow
- 23:18:52retailer.com. You can also get database
- 23:18:54from the public real estate databases
- 23:18:56like Kaggle data sets where you have
- 23:18:58Zillow home value prediction. Ensure
- 23:19:00that the data set include features like
- 23:19:02location where you have latitude,
- 23:19:04longitude, square footage, number of
- 23:19:06rooms, year built, property type and
- 23:19:08previous sales. The next step that comes
- 23:19:11is data cleaning. You have to handle the
- 23:19:13missing values by using imputation
- 23:19:15techniques or removing incomplete
- 23:19:17records. Removing outliers that may skew
- 23:19:20the model's prediction, normalize or
- 23:19:22standardize the data to ensure
- 23:19:24consistency. The third one that we have
- 23:19:26is feature engineering. You have to
- 23:19:28create new features such as proximity to
- 23:19:30schools, crime rates and access to the
- 23:19:32public transportation. Encode categorial
- 23:19:35variables, example property type,
- 23:19:37location using techniques like one hot
- 23:19:38encoding. Generate interaction features
- 23:19:41that capture relationship between
- 23:19:42existing features. The fourth one that
- 23:19:44we have is model selection. Use
- 23:19:46regression models like linear
- 23:19:48regression, random forest, gradient
- 23:19:49boosting, neural networks. Experiment
- 23:19:52with different models to identify the
- 23:19:53best performing one. Now in the next
- 23:19:56phase all you have to do guys is model
- 23:19:58training and evaluation. Split the data
- 23:20:00set into training and test sets. Train
- 23:20:03the model on a training set and evaluate
- 23:20:05their performance on the testing set
- 23:20:07using metrics like RSM which means root
- 23:20:10mean squared error. You can use cross
- 23:20:12validation to ensure the model's
- 23:20:14robustness and avoid overfitting. The
- 23:20:17sixth one that we have all over here is
- 23:20:19hyperparameter tuning. You can optimize
- 23:20:21the model's hyperparameter using
- 23:20:23techniques such as grid search or random
- 23:20:25search to improve accuracy. And if
- 23:20:27you're looking forward to deploy your
- 23:20:29model, then you can develop a web
- 23:20:31interface using flask or Django to allow
- 23:20:33users to input property features and get
- 23:20:35predictions. You can deploy the model on
- 23:20:38the cloud platform like AWS for
- 23:20:40scalability.
- 23:20:41Now if we talk about the complexity
- 23:20:43level of this, we all know that it is a
- 23:20:45beginner level project. Now let us move
- 23:20:48on to the one more set that is music
- 23:20:50genre classification and generation. So
- 23:20:53guys this is also one of the most
- 23:20:55beginner level project. This project
- 23:20:57aims to develop a system that can
- 23:20:58classify music tracks into different
- 23:21:00genres and generate new music
- 23:21:02composition within specified genre. The
- 23:21:04goal is to build a model that analyzes
- 23:21:06audio features to categorize music and
- 23:21:08uses deep learning techniques to create
- 23:21:10new music. This project introduces
- 23:21:13advanced concept of audio processing,
- 23:21:15deep learning and generative models. So
- 23:21:17guys, what will be used in this? So
- 23:21:19we'll have programming language that
- 23:21:21will be Python. For audio processing,
- 23:21:23we'll be using librosa. For machine
- 23:21:25learning libraries, we'll be using
- 23:21:26tensorflow, kas, pytorch. For data
- 23:21:29handling libraries, we'll be using
- 23:21:30pandas, numpy. For visualization, we'll
- 23:21:32be using mattplot, seabon. For the data
- 23:21:35set guys, you can use gtzan music genre
- 23:21:38data set or you can get it from free
- 23:21:40music archive. So guys in the first
- 23:21:42phase we are going to have data
- 23:21:43collection. You can obtain data sets
- 23:21:45containing music tracks and their
- 23:21:47corresponding genre from the sources
- 23:21:49from GTN music genre data set and the
- 23:21:51free music archive. Ensure that a data
- 23:21:54set includes diverse genre and
- 23:21:55substantial number of tracks per genre.
- 23:21:58Next we'll go for data prep-processing.
- 23:22:00Use library librosa to load and
- 23:22:02pre-process audio files including
- 23:22:04feature extraction such as mil
- 23:22:06frequency, septal coefficients, chroma
- 23:22:08features and spectral contrast. You can
- 23:22:11normalize the extracted features to
- 23:22:13ensure consistent input for the given
- 23:22:15models. Now if you talk about feature
- 23:22:17engineering guys, you can extract
- 23:22:19additional features from the audio files
- 23:22:20such as tempo, beat, zero crossing rate
- 23:22:23etc. Create a feature matrix that
- 23:22:25represents the extracted audio features.
- 23:22:27Then go for the model selection. Use
- 23:22:30conventional neural network or recurrent
- 23:22:32neural networks for the music genre
- 23:22:33classification. Split the data set into
- 23:22:36training and testing data set. Now if we
- 23:22:38talk about model training and evaluation
- 23:22:40guys then you can train the selected
- 23:22:42classification model on the training
- 23:22:43set. Evaluate the model's performance on
- 23:22:46the testing set using metrics like
- 23:22:47accuracy, precision, recall and F1
- 23:22:50score. Use confusion matrices to
- 23:22:52understand the classification
- 23:22:53performance across different genres. Now
- 23:22:55if we talk about model selection and
- 23:22:57training for the music generation, what
- 23:22:59you will do guys? You can use the
- 23:23:01generative adversial networks or
- 23:23:03recurrent neural networks such as LSTM,
- 23:23:05long short-term memory for music
- 23:23:07generation. Train the generative model
- 23:23:09on the data set to create new music
- 23:23:11sequences. Next, we have model training
- 23:23:13and evaluation. You can train the
- 23:23:15generative model on sequences of audio
- 23:23:17features. You can evaluate the generated
- 23:23:19music by listening tests and by
- 23:23:21objective metrics like inception score
- 23:23:23or fche audio distance. If I talk about
- 23:23:26hyperparameter tuning guides, we can
- 23:23:27optimize the model. You can use the
- 23:23:29hyperparameters using techniques like
- 23:23:31grid search or random search to improve
- 23:23:33performance. If I talk about deployment
- 23:23:35guys, you can deploy these models on
- 23:23:37cloud platforms like AWS. Now let us
- 23:23:40move on to our next project. So guys,
- 23:23:43the complexity of this project is at the
- 23:23:44beginner level. Now let us move to the
- 23:23:47intermediate level projects. Next
- 23:23:49project that we have all over here is
- 23:23:50sentiment analysis of Twitter data. This
- 23:23:53project aims to develop a sentiment
- 23:23:55analysis model that can classify to its
- 23:23:57side positive, negative or neutral. The
- 23:24:00goal is to analyze public sentiment on
- 23:24:02various topics or events using natural
- 23:24:04language techniques. So guys, what will
- 23:24:06be used all over here? So in this we
- 23:24:09will have programming language like
- 23:24:11Python. Okay. NLP libraries, NLTK spacy.
- 23:24:15For machine learning libraries you can
- 23:24:16use scikitlearn, tensorflow, kas. For
- 23:24:19data handling libraries we have pandas,
- 23:24:21numpy. For visualization we have
- 23:24:23mattplot lilip seon and we can use the
- 23:24:25API Twitter for data collection. Now how
- 23:24:28you going to work on it guys? So guys if
- 23:24:31I talk about the data collection use a
- 23:24:33Twitter API to collect tweets based on
- 23:24:35specific hashtags like keywords or
- 23:24:37topics. Extract relevant fields like
- 23:24:39tweet text, user information, timestamp
- 23:24:42etc. Then if I talk about data
- 23:24:43prep-processing guys, clean the tweet
- 23:24:46text by removing special characters,
- 23:24:48links, mentions, hashtags and stop
- 23:24:50words. Tokenize the text and perform
- 23:24:53limization or stemming to reduce the
- 23:24:55words to their base form. Next, if you
- 23:24:57talk about feature engineering guys, you
- 23:24:59convert the clean text data into
- 23:25:00numerical representation using TF, back
- 23:25:03of words or word embedding. Now if I
- 23:25:06talk about model selection guys, you can
- 23:25:07choose a classification algorithm such
- 23:25:09as logistic regression, n bias or LSDM.
- 23:25:13Split the data set into training and
- 23:25:14testing data set. Now if I talk about
- 23:25:17model training and evaluation, then you
- 23:25:19can train the selected model on the
- 23:25:21training set. Evaluate the model's
- 23:25:23performance on the testing set using
- 23:25:24metrics like accuracy, precision,
- 23:25:26recall, and fn score. Use cross
- 23:25:29validation to ensure the model's
- 23:25:31robustness. If I talk about
- 23:25:32hyperparameter tuning guys, you can
- 23:25:34optimize the model's hyperparameters
- 23:25:36using grid search or random search to
- 23:25:38improve the performance for deployment
- 23:25:40which can be optional. You can deploy
- 23:25:42your model on AWS for real-time
- 23:25:44sentiment analysis. So if I talk about
- 23:25:46the complexity level guys, its
- 23:25:48complexity is intermediate. So guys, our
- 23:25:51next project is customer segmentation
- 23:25:53using K means clustering. This project
- 23:25:56aims to segment customers into distinct
- 23:25:58groups based on their purchasing
- 23:25:59behavior and demographic information.
- 23:26:01The objective is to understand customer
- 23:26:03segments and tailor marketing strategies
- 23:26:05accordingly. So guys, what programming
- 23:26:08languages we'll be using? So basically
- 23:26:10we'll be using Python. For machine
- 23:26:12learning libraries, we will have
- 23:26:13scikitlearn. For data handling
- 23:26:15libraries, we'll have pandas, numpy. For
- 23:26:17visualization libraries, we'll have
- 23:26:18mattplot lab, seabon. And the data set
- 23:26:21source will be e-commerce transaction
- 23:26:23data. how we are going to work on this
- 23:26:25one. For data collection, we can obtain
- 23:26:27a data set of e-commerce transactions
- 23:26:29that include customer demographics,
- 23:26:30purchase history, and product
- 23:26:32information. Next, we'll have data
- 23:26:33prep-processing. For data
- 23:26:35prep-processing, we are going to do the
- 23:26:37cleaning of the data by handling the
- 23:26:38missing values and outliers. Then, for
- 23:26:41feature engineering, we are going to
- 23:26:43create features like total purchase
- 23:26:44amount, purchase frequency, and recency
- 23:26:46of the purchases.
- 23:26:48Then, we are going to proceed for the
- 23:26:50model selection. You can use K means
- 23:26:52clustering to segment the customers into
- 23:26:54distinct groups. You can determine the
- 23:26:56optimal number of clusters using methods
- 23:26:58like album methods or silhou.
- 23:27:01Now if I talk about model training and
- 23:27:03evaluation, you can train the K means
- 23:27:05model on the process data set. You can
- 23:27:07evaluate the quality of clusters by
- 23:27:09analyzing intracluster and intercluster
- 23:27:11distances. Next we have the evaluation.
- 23:27:14You can visualize the clusters using
- 23:27:15techniques like PCA, principal component
- 23:27:17analysis, TSN etc. Next we have the
- 23:27:22hyperparameter tuning. Now now you can
- 23:27:25tune this model and interpret the
- 23:27:27characteristic of each segment. You can
- 23:27:29develop a target marketing strategies
- 23:27:30for each segment based on unique
- 23:27:32behavior and preferences. Now deployment
- 23:27:35is optional. You can develop a dashboard
- 23:27:37using flask or Django to visualize
- 23:27:39customer segments and track marketing
- 23:27:41campaigns. So guys if I talk about the
- 23:27:44complexity of this project. So this is
- 23:27:47an intermediate level project. So guys
- 23:27:49for data set you can use the Kaggle's
- 23:27:51customer segmentation data set which is
- 23:27:53available at the Kaggle's platform. Now
- 23:27:56the third intermediate level project
- 23:27:58that we have all over here is building a
- 23:28:00chatbot with Rasa. This project aims to
- 23:28:02build an intelligent chatbot using Rasa
- 23:28:05framework. The chatbot will be capable
- 23:28:06of understanding user queries and
- 23:28:08providing appropriate responses making
- 23:28:10it useful for customer support, personal
- 23:28:13assistance or information retrieval.
- 23:28:15What languages we are going to use? So
- 23:28:17it will be Python based. We'll have the
- 23:28:19NLP libraries like Rasa, NLTK, Spacey.
- 23:28:22For machine learning libraries, we are
- 23:28:24going to have scikitlearn, tensorflow,
- 23:28:26kas, etc. For data handling libraries,
- 23:28:28we are going to use pandas, numpy. So
- 23:28:31guys, this was what we are going to do
- 23:28:33it and how you can work on this one by
- 23:28:35collecting the data, collect the
- 23:28:37conversation data and FAQs from the
- 23:28:39target domain, annotate the data to
- 23:28:41create training examples for the
- 23:28:43chatbot. Next comes is data
- 23:28:45prep-processing. Clean the text data by
- 23:28:47removing special characters and
- 23:28:48normalizing the text. You can tokenize
- 23:28:50and limitize the text to prepare for a
- 23:28:53training. The third one we have the
- 23:28:55models training. You can use Rasa's
- 23:28:57NLU's component to train a model for
- 23:28:59intent recognition and entity
- 23:29:01extraction. You can define a dialog
- 23:29:03management policies to handle different
- 23:29:05conversation flows. Next guys, you can
- 23:29:07perform the feature engineering and
- 23:29:09integration. You can integrate the Ras
- 23:29:11NLU and core components to build
- 23:29:13complete chatbot. You can connect the
- 23:29:15chatbot to messaging platform like
- 23:29:17Facebook Messenger etc. For model
- 23:29:19selection and testing what you can do
- 23:29:21guys you can test this chatbot with
- 23:29:23various inputs to ensure that it handles
- 23:29:25the scenarios appropriately and you can
- 23:29:28also select the right model using this.
- 23:29:30Now if I talk about model training guys
- 23:29:32what you have to do you have to collect
- 23:29:34the user feedback and conversational
- 23:29:36logs to continuously work on training
- 23:29:38the model. Next, similarly you have to
- 23:29:41retrain the model periodically with the
- 23:29:43new data to see how it is working. So
- 23:29:46that will be your evaluation. Now for
- 23:29:48the hyperparameter tuning, what are you
- 23:29:50going to do guys? You have to check in
- 23:29:52those scenarios where it is able to tune
- 23:29:54up with those scenarios where it can
- 23:29:56handle the input appropriately. And next
- 23:29:59is deployment. So guys, for deploying
- 23:30:01it, you can use AWS. So guys, for a data
- 23:30:04set, you can use Ras open source. So
- 23:30:07that's a very good data set for you to
- 23:30:09proceed. So guys, if you talk about
- 23:30:11difficulty of this project, this is an
- 23:30:13intermediate level project. Now let us
- 23:30:15move on to the advanced level projects.
- 23:30:17For advanced level projects, the first
- 23:30:19one that comes up to my mind is movie
- 23:30:21similarity from plot summaries. Now this
- 23:30:24project aims to develop a system that
- 23:30:26can find out recommended movies similar
- 23:30:28to a given movie based on their plot
- 23:30:31summaries. By analyzing the textual
- 23:30:33content of the movie plot summaries, the
- 23:30:35model will identify similarities and
- 23:30:37suggests movies with similar themes,
- 23:30:39story lines or genres. This project
- 23:30:42introduces beginners to natural language
- 23:30:44processing text similarly measures and
- 23:30:46recommendation systems. What languages
- 23:30:48we are going to use guys? We'll be using
- 23:30:50Python NLP libraries like NLTK spacy
- 23:30:53machine learning libraries like
- 23:30:55scikitlearn data handling libraries like
- 23:30:57pandas numpy visualization we can use
- 23:31:00numpy and data set source will be IMDb
- 23:31:02or kegel so guys this process is also
- 23:31:05involving the data collection then you
- 23:31:07have to go for data cleaning then
- 23:31:09feature engineering next model selection
- 23:31:12so similar process as I have discussed
- 23:31:14in other projects so you have to also go
- 23:31:16through the same one next what you have
- 23:31:18to do device. Similarly, what you have
- 23:31:21to do, you have to train the model, then
- 23:31:23evaluate the model, then hypertune it
- 23:31:25and finally proceed for the deployment.
- 23:31:28So, this is overall process of this
- 23:31:30project. Try to research on the website
- 23:31:33a lot like how you can extract it. So,
- 23:31:36guys, you can use kegel or towards data
- 23:31:38science to research more about this
- 23:31:40project. Now guys, if I talk about the
- 23:31:42difficulty of this project and this is
- 23:31:44an advanced level project. Now let us
- 23:31:47move on to the next one that we have all
- 23:31:50over here that is image segmentation
- 23:31:52project for brain tumor prognosis. This
- 23:31:55is a very very amazing project and
- 23:31:57definitely you can put up on your
- 23:31:58portfolio. Basically guys this project
- 23:32:00aims to develop an image segmentation
- 23:32:02model to identify and delinate brain
- 23:32:05tumors from MRI scans. The goal is to
- 23:32:07accurately segment the tumor regions
- 23:32:09which can aid in prognosis treatment
- 23:32:11planning surgical interventions. This
- 23:32:13project introduces intermediate level
- 23:32:15concepts of computer visions, deep
- 23:32:17learning and medical image analysis. So
- 23:32:19guys, what we'll be using all over here
- 23:32:21for programming languages we can use
- 23:32:23python for deep learning libraries we
- 23:32:25can use tensorflow, kas, pytor. For
- 23:32:27image processing libraries, we can use
- 23:32:29opencv, scikit image. For data handling
- 23:32:31libraries, we can use pandas, numpy. For
- 23:32:34visualization, we can use mattplot,
- 23:32:36cbond. Now if I talk about what is the
- 23:32:39process of developing this project the
- 23:32:41first step will be same data collection.
- 23:32:44So next step you have to go for data
- 23:32:46prep-processing. Third step you have to
- 23:32:48do the model selection where you can use
- 23:32:50CNN models for image segmentation task.
- 23:32:53Then you go for model selection. Moving
- 23:32:56ahead you're going to have the model
- 23:32:58training and evaluation. You have to
- 23:32:59split the data set into training and
- 23:33:02validation and testing data sets. Next
- 23:33:05proceed for the evaluation phase. Okay,
- 23:33:07evaluate the model with certain metrics.
- 23:33:09So here I can give you certain idea like
- 23:33:11you can use dice coefficient,
- 23:33:13intersection over union or accuracy.
- 23:33:15Then go for hyperparameter tuning where
- 23:33:17you have to optimize the model's
- 23:33:19hyperparameters.
- 23:33:20You can use grid search or random search
- 23:33:22as we have discussed. And finally you
- 23:33:24can deploy this model on AWS. Now guys
- 23:33:27we have come to the final project. This
- 23:33:30is also a very amazing project guys. So
- 23:33:32guys the complexity level of this
- 23:33:34project is advanced level.
- 23:33:36Now let us move on to our final project
- 23:33:39that is the impact of climate change on
- 23:33:42birds. This is a very very amazing
- 23:33:44project and definitely you can add it on
- 23:33:46your resume. This project aims to
- 23:33:48analyze the impact of climate change on
- 23:33:50the bird population and migration
- 23:33:52patterns by examining various climatic
- 23:33:54factors and their correlation with bird
- 23:33:56species data. The project seeks to
- 23:33:58predict how climate change might affect
- 23:34:00bird behavior and distribution. This
- 23:34:02project will introduce you some advanced
- 23:34:04level concepts like time series
- 23:34:06analysis, environmental data modeling,
- 23:34:08etc. So guys, what programming languages
- 23:34:10we'll be using for data analysis? You
- 23:34:12can see we'll have pandas, numpy. For
- 23:34:14machine learning libraries, we are going
- 23:34:15to have scikitlearn, tensorflow. For
- 23:34:18visualization, we're going to have
- 23:34:19mattplot lab, plotly. For geospatial, we
- 23:34:22are going to have geopandas, folium. For
- 23:34:24data source, we're going to have public
- 23:34:26data sets on bird observation and
- 23:34:27climate data sources from eird. Now what
- 23:34:30will the process flow for this one guys?
- 23:34:32First you have to proceed for data
- 23:34:33collection. Gather bird observation from
- 23:34:35the data like EIRD which provides
- 23:34:37extensive record of bird sightings.
- 23:34:40Okay. And for climate data you can
- 23:34:42collect it from NOA including
- 23:34:44temperature, precipitation and other
- 23:34:46relevant climatic factors over the time.
- 23:34:48And similar next process will be the
- 23:34:50data prep-processing. Then you have to
- 23:34:52proceed for feature engineering. Then
- 23:34:54you have to go for model selection.
- 23:34:56Okay. Moving ahead you have to go for
- 23:34:58model training. then evaluation, then
- 23:35:01hyperparameter tuning and finally you
- 23:35:03have to deploy the model. So research
- 23:35:05about this project, see what models you
- 23:35:07are going to use. Suppose I can give you
- 23:35:09a hint about this. You can use time
- 23:35:10series analysis models like ARMA or ML
- 23:35:14models. You can also use random forest
- 23:35:16or gradient boosting for predicting
- 23:35:18impact on the bird population. So guys,
- 23:35:21use Google exhaustively to research
- 23:35:23about this project. This is also a very
- 23:35:25amazing project and it's going to give
- 23:35:26you a lot of idea. Now if I talk about
- 23:35:29the complexity level of this, it is an
- 23:35:31advanced level project.
- 23:35:33>> Welcome to deep learning interview
- 23:35:34questions. My name is Richard Kersner
- 23:35:37with the SimplyLearn team. That's
- 23:35:39www.simplearn.com.
- 23:35:41Get certified, get ahead. Today we're
- 23:35:44going to help you prepare for interview
- 23:35:45questions dealing with deep learning.
- 23:35:47And we're going to go from the very
- 23:35:48basics of neural networks and deep
- 23:35:51learning into some of the more commonly
- 23:35:54used models so you can have an
- 23:35:56understanding of what kind of questions
- 23:35:57are going to come up and what you need
- 23:35:59to know in interview questions. We'll
- 23:36:01start with a very general concept of
- 23:36:03what is deep learning. This is where we
- 23:36:05take large volumes of data in this case
- 23:36:07on cats and dogs or whatever. A lot of
- 23:36:09times you use um a training setup to
- 23:36:12train your model. Remember it's kind of
- 23:36:13like a magic black box going on there.
- 23:36:16And then we use that to extract features
- 23:36:17or extract information and in this case
- 23:36:19classify the image of a cat and a dog.
- 23:36:22So the primary takeaway we're talking
- 23:36:23about deep learning is it learns from
- 23:36:25large volumes of structured and even
- 23:36:28unstructured data and uses complex
- 23:36:30algorithms to train neural network. It
- 23:36:32also performs complex operations to
- 23:36:34extract hidden patterns and features.
- 23:36:36And if we're going to discuss deep
- 23:36:38learning in this very uh simplified
- 23:36:40overview and we also have to go over
- 23:36:42what is a neural network. This is a
- 23:36:44common image you'll see of a drawing of
- 23:36:46a forward propagation neural network and
- 23:36:49it's it's a human brain inspired system
- 23:36:51which replicate the way humans learn. So
- 23:36:53this has inspired how our own neurons
- 23:36:55and our brain fire but at a much
- 23:36:57simplified level. Obviously it's not
- 23:36:58ready to take over the human uh
- 23:37:00population and and be our leader yet.
- 23:37:02Not for many years. It's very much in
- 23:37:04its infant stage. But it's inspired by
- 23:37:06how our brains work. Um and they use a
- 23:37:08lot of other inspirations. You can study
- 23:37:10brains of moths and other animals that
- 23:37:13they've used to figure out how to
- 23:37:14improve these neural networks. The most
- 23:37:16common one consists of three layers of
- 23:37:18network and this is generally how you
- 23:37:20view these networks is you have an
- 23:37:21input, you have a hidden layer and an
- 23:37:23output. And the neural network is uh
- 23:37:26broken up into many pieces. But when we
- 23:37:28focus just on the neural network, it's
- 23:37:30always on the hidden layers that we're
- 23:37:31making all the adjustments and figuring
- 23:37:33out how to best set up those hidden
- 23:37:36layers for their functions to both train
- 23:37:39faster and to function better. When we
- 23:37:41look at this, of course, we have our
- 23:37:42input, hidden, and output. Each layer
- 23:37:44contains neurons called as nodes perform
- 23:37:47various operations. And you can see here
- 23:37:49we have the list of the nodes. We have
- 23:37:51both our input nodes and our output
- 23:37:52nodes and then our hidden layer nodes.
- 23:37:54And it's used in deep learning algorithm
- 23:37:56like CNN, RNN, GN, etc. We'll address
- 23:38:00some of these models a little closer, at
- 23:38:01least the most common models as we go
- 23:38:03down the list and we study the deep
- 23:38:05learning and the neural network
- 23:38:06framework. Let's start with what is a
- 23:38:09multi-layer perceptron or MLP a lot of
- 23:38:12time as they're referred to. And you'll
- 23:38:14see these abbreviations. I'll be honest,
- 23:38:16I have to write them down on a piece of
- 23:38:18paper and go through them because I
- 23:38:19never remember what they all mean even
- 23:38:21though I play with them all the time.
- 23:38:22What is a multi-layer perceptron? Well,
- 23:38:24if you look at the image on the right,
- 23:38:26it's very similar to what we just looked
- 23:38:27at. You have your input layer, your
- 23:38:29hidden layer, and your output layer. And
- 23:38:31that's exactly what this is. It has the
- 23:38:33same structure of a single layer
- 23:38:34perceptron with one or more hidden
- 23:38:36layers except the input layer, each node
- 23:38:38in the other layers uses a nonlinear
- 23:38:40activation function. What that means is
- 23:38:43your input layer is your data coming in
- 23:38:45and then your activation function is
- 23:38:47based upon all those nodes and weights
- 23:38:49being added together and then it has the
- 23:38:51output. MLP uses supervised learning
- 23:38:54method called back propagation for
- 23:38:56training the model. Very key word there
- 23:38:58is back propagation. Single layer
- 23:39:00perceptron can classify only linear
- 23:39:02separable classes with binary output 01.
- 23:39:06But the MLP can classify nonlinear
- 23:39:08classes. So let's break this down just a
- 23:39:10little bit. The multi-layer perceptron
- 23:39:12with an input layer and a hidden layer
- 23:39:14and an output layer. As you see that it
- 23:39:16comes in there, it has adds up all the
- 23:39:18numbers and weights depending on how
- 23:39:20your setup is. That then goes to the
- 23:39:22next layer. That then goes to the next
- 23:39:23hidden layer if you have multiple hidden
- 23:39:25layers. And finally to the output layer.
- 23:39:27The back propagation takes the error
- 23:39:30that it sees. So whatever the output is,
- 23:39:32it says, hey, this has an error to it.
- 23:39:33It's wrong. And then sends that error
- 23:39:35backwards from where it came from. And
- 23:39:37there's a lot of different functions
- 23:39:39used to uh train this based on that
- 23:39:42error and how that error goes backwards
- 23:39:44in the notes. Uh so forward is you get
- 23:39:46your answers. Backward is for training.
- 23:39:49You see this every day. Even my uh
- 23:39:51Google Pixel phone has this. It they
- 23:39:53train the neural network which takes a
- 23:39:55lot more data to train than it does to
- 23:39:57use. And then they load up that neural
- 23:39:59network into in this case I have a Pixel
- 23:40:012 which actually has a built-in neural
- 23:40:03network for processing pictures. And so
- 23:40:05it's just the forward propagation I use
- 23:40:07when it processes my photos, but when
- 23:40:10they were training it, you use the back
- 23:40:11propagation to train it with the errors
- 23:40:13they had. We'll be coming back to
- 23:40:15different models that are used. For
- 23:40:17right now though, multi-layer
- 23:40:18perceptron, MLP, put that down as your
- 23:40:21vocabulary word and of course back
- 23:40:24propagation. What is data normalization
- 23:40:26and why do we need it? This is so
- 23:40:29important. We spend so much time in
- 23:40:30normalizing our data and getting our
- 23:40:32data clean and setting it up. Uh so we
- 23:40:34talk about data there's a pre-processing
- 23:40:36step to standardize the data. So
- 23:40:38whatever we have coming in we don't want
- 23:40:40it to be a uh you know one gigabyte file
- 23:40:43here a 2 GBTE picture here and a 3
- 23:40:46kilobyte text there. Even as a human I
- 23:40:49can't process those all in the same
- 23:40:50group. I have to reformat them in some
- 23:40:52way that loops them together so they're
- 23:40:54a standardized format. We use this uh
- 23:40:56data normalization and and
- 23:40:58pre-processing to reduce and eliminate
- 23:41:01data redundancy. A lot of times the data
- 23:41:04comes in and you end up with two of the
- 23:41:06same images or uh uh the same
- 23:41:08information in different formats. Then
- 23:41:10we want to rescale values to fit into a
- 23:41:12particular range for achieving better
- 23:41:14convergence. What this means is with
- 23:41:17most neural networks they form a bias.
- 23:41:20We've seen this in recently in attacks
- 23:41:22on neural networks where they light up
- 23:41:24one pixel or one piece of the view and
- 23:41:26it skews the whole answer. So suddenly u
- 23:41:29because one pixel is really bright uh it
- 23:41:32doesn't know what to do. Well when we
- 23:41:34start rescaling it we put all the values
- 23:41:36between say minus one and one and we
- 23:41:38change them and refit them to those
- 23:41:40values. It helps get rid of that bias
- 23:41:42helps fix for some of those problems.
- 23:41:44And then finally we restructure the data
- 23:41:46and improve the integrity. We want to
- 23:41:47make sure that we're not missing values
- 23:41:49um or we don't have partial data coming
- 23:41:51in. One way to look at this is uh bad
- 23:41:54data in bad data out. And so you want
- 23:41:57clean data in and you want good answers
- 23:41:59coming out. One of the most basic models
- 23:42:01used is a Boltzman machine. So let's
- 23:42:04address what is a Boltzman machine. And
- 23:42:06if you know we just did the MLP
- 23:42:08multi-layer perceptron. So now we're
- 23:42:10going to come into almost a simplified
- 23:42:12version of that. And in this we have our
- 23:42:14visible input layer and we have our
- 23:42:16hidden layer. The Boltzman machines are
- 23:42:18almost always shallow. They're usually
- 23:42:19just two-layer neural nets that make
- 23:42:21stochastic decisions whether a neuron
- 23:42:24should be on or off. True or false? Yes.
- 23:42:26No. First layer is a visible layer and
- 23:42:28second layer is the hidden layer. Nodes
- 23:42:30are connected to each other across
- 23:42:32layers, but no two nodes of the same
- 23:42:34layer are connected. Hence, it is also
- 23:42:36known as restricted Boltzman machine.
- 23:42:39Now that we've covered a basic MLP or
- 23:42:41multi-layer perceptron, and we've gone
- 23:42:44over the Boltzman machine, also known as
- 23:42:46the restricted Boltzman machine, let's
- 23:42:47talk a little bit about activation
- 23:42:49formulas. And this is a huge topic that
- 23:42:53can get really complicated but it also
- 23:42:55is automated. So it's very simple. So
- 23:42:57you have both a complicated and a simple
- 23:42:59at the same time. So what is the role of
- 23:43:02activation functions in a neural
- 23:43:03network? Activation function decides
- 23:43:05whether a neuron should be fired or not.
- 23:43:08That's the most basic one and that
- 23:43:10actually changes a little bit because
- 23:43:11it's either whether fired or not in this
- 23:43:13case activation function or what value
- 23:43:16should come out when it's fired. But in
- 23:43:17these models, we're looking at just the
- 23:43:19boltsman restricted layers. So this is
- 23:43:22what causes them to fire. Either they
- 23:43:24don't or they do. It's a yes or no,
- 23:43:26true, false, all or nothing. It accepts
- 23:43:28the weighted sum of the inputs, the bias
- 23:43:30as input to any activation function. So
- 23:43:33whatever activation function is, it
- 23:43:35needs to have the sum of the weights
- 23:43:37times the input. So each input, if you
- 23:43:39remember on that model, and let's just
- 23:43:40go back to that model real quick. And
- 23:43:42then you always have to add a bias. And
- 23:43:44you can look at the bias if you remember
- 23:43:46from your uklitian geometry. You draw a
- 23:43:49straight line. Formula for that line has
- 23:43:51a y-coordinate at the end. It's always
- 23:43:54um cx plus m or something like that
- 23:43:56where m is where it crosses the
- 23:43:58ycoordinates. If you're doing a straight
- 23:44:00line with these weights, it's very
- 23:44:02similar, but a lot of times we just add
- 23:44:04it in as its own weight. We take it as a
- 23:44:07node of a one value coming in and then
- 23:44:09we compute its new weight. And that's
- 23:44:11how we compute that bias just like we
- 23:44:12compute all the other weights coming in.
- 23:44:14The node which gets fired depends on the
- 23:44:16y value. And then we have a step
- 23:44:18function. And the step function this is
- 23:44:20where remember I said it's going to get
- 23:44:21complicated and simple all at the same
- 23:44:23time. We have a lot of different step
- 23:44:25functions. We have the sigmoid function.
- 23:44:28We have just a standard step function.
- 23:44:30We have the ru is pronounced like ray
- 23:44:33the ray of from the sun and lu like a
- 23:44:35name. So ru function. And we have the
- 23:44:37tangent h function. And if you look at
- 23:44:39these, they all have something similar.
- 23:44:41They all either force it to be um one
- 23:44:44value or the other. They force it to be
- 23:44:45in the case of the first three a zero or
- 23:44:47one. And in the last one, it's either a
- 23:44:50minus one or one. And you can easily
- 23:44:51convert that to a 0, one, yes, no, true,
- 23:44:54false. And on this, one of the most
- 23:44:56common ones is the step function itself
- 23:44:58because there is no middle value. There
- 23:45:00is no um uh discrepancy that says, well,
- 23:45:03I'm not quite sure. But as you get into
- 23:45:05different models, probably the most
- 23:45:07commonly used used to be the sigmoid was
- 23:45:09most commonly used, but I see the relu
- 23:45:11used more often. Really, depending on
- 23:45:13what you're doing, you just have to play
- 23:45:15with these and find out which one works
- 23:45:16best depending on the data in your
- 23:45:19output. The reason to have a non01
- 23:45:22answer or something kind of in the
- 23:45:24middle is when you're looking at this
- 23:45:25and it's coming out, you can actually
- 23:45:27process that middle ground as part of
- 23:45:30the answer into another neural network.
- 23:45:32So it might be that the relu function
- 23:45:34says hey this is only a 6 not a one and
- 23:45:38uh even though the one is what's going
- 23:45:41into the next neural network or the next
- 23:45:43hidden layer as an input the 6 value
- 23:45:47might also be going in there to let you
- 23:45:49know hey this is not a straight up one
- 23:45:51or straight up zero it's someplace in
- 23:45:53the middle this is a little uncertain
- 23:45:54what's coming out here so it's a very
- 23:45:56powerful tool in the basic neural
- 23:45:58network you usually just use the step
- 23:45:59function it's yes or no let's take a um
- 23:46:02a big step back and take a kind of an
- 23:46:05overview. The next function is what is a
- 23:46:08cost function that we're going to cover.
- 23:46:10This is so important because this is
- 23:46:12your end result that you're going to do
- 23:46:14over and over again and use to decide
- 23:46:16whether the model is working or not,
- 23:46:18whether you need to try a different step
- 23:46:19function, whether you need to try a
- 23:46:21different activation, whether you need
- 23:46:22to try a fully different model used. Uh
- 23:46:24so what is the cost function? Cost
- 23:46:26function is a measure to evaluate how
- 23:46:29good your model's performance is. It is
- 23:46:31also referred as loss or error used to
- 23:46:34compute the error of the output layer
- 23:46:36during back propagation. There's our
- 23:46:38back propagation where we're training
- 23:46:39our model. That's one of our key words.
- 23:46:42Mean squared error is an example of a
- 23:46:44popular cost function. And so here we
- 23:46:46have the cost function C = half of Y - Y
- 23:46:50predicted. Um and then you square that.
- 23:46:52So the first thing is um you know real
- 23:46:54quick if you haven't done statistics
- 23:46:56this is not a percentage. It's not a
- 23:46:58percentage of how accurate it is. is
- 23:47:00just a measurement of the error and we
- 23:47:02take that error if we're training it and
- 23:47:04we push that error backwards through the
- 23:47:06neural network and we use that through
- 23:47:08the different training functions
- 23:47:10depending on what model you're using to
- 23:47:12train the neural network. So when you
- 23:47:14deploy the network you're usually done
- 23:47:15training it because it takes a lot of
- 23:47:17computational force to train it. Um this
- 23:47:19is a very simple model and so you deploy
- 23:47:21the train one. Uh but we want to know
- 23:47:22how your error is and so how do we do
- 23:47:24that? Well you split your data. part of
- 23:47:26your data is for trading and part of
- 23:47:28your data is for testing. And then we
- 23:47:30can also test the error on there. So
- 23:47:32it's very important. And then we're
- 23:47:33going to go one more step on this. We
- 23:47:36got to look at both the local and the
- 23:47:38global setup. It might work great to
- 23:47:40test your data on what you have on your
- 23:47:42computer, but that's different than in
- 23:47:44the field. So, when we're talking about
- 23:47:46all these different tests and the error
- 23:47:48test as far as your loss, you don't you
- 23:47:50want to make sure that you're in a
- 23:47:52closed environment when you do initial
- 23:47:53testing, but you also want to open that
- 23:47:55up and make sure you follow up with the
- 23:47:56testing on the larger scale of data
- 23:47:58because it will change. It might not fit
- 23:47:59the larger scale. There might be
- 23:48:01something in there in the way you
- 23:48:02brought the data in specifically or the
- 23:48:04data group you used or um any of those
- 23:48:06could cause an error. So, it's very
- 23:48:08important to remember that we're looking
- 23:48:09at both the local and the global context
- 23:48:11of our error. And just one other side
- 23:48:14note on a lot of the newer models of
- 23:48:16neural networks by comparing the error
- 23:48:19we get on the data our training data
- 23:48:21with a portion of the test data we can
- 23:48:24actually figure out how good the model
- 23:48:25is whether it's overfitted or not. We'll
- 23:48:27go into that a little bit more as we go
- 23:48:29into some of the different models. So we
- 23:48:31have our output. We're able to um figure
- 23:48:34out the error on it based on the square
- 23:48:35means usually although there's other uh
- 23:48:37functions used. So we want to talk about
- 23:48:39what is gradient descent? Another
- 23:48:41vocabulary word gradient descent is an
- 23:48:44optimation algorithm to minimize the
- 23:48:47cost function or to minimize the error.
- 23:48:49Aim is to find the local or global
- 23:48:51minima of a function. Determine the
- 23:48:53direction the model should take to
- 23:48:55reduce the error. So as we're looking at
- 23:48:57this, we have our uh squared error that
- 23:48:59we just figured out the co based on the
- 23:49:01cost function. It says how bad is my
- 23:49:03model fitting the data I just put
- 23:49:05through it. And then we want to reduce
- 23:49:07that error. So how do you figure out
- 23:49:08what direction to do that in? Well, it
- 23:49:10could be that you're looking at just
- 23:49:12that line of that line of data coming
- 23:49:14in. So that would be a local minima. We
- 23:49:16want to know the error of that
- 23:49:17particular setup coming in. And then you
- 23:49:19have your global your global minima. We
- 23:49:21want to minimize it based on the overall
- 23:49:24data we're putting through it. And with
- 23:49:25this we can figure out the global
- 23:49:28minimum cost. We want to take all those
- 23:49:30local minimum costs of each piece of
- 23:49:33data coming in and figure out the global
- 23:49:34one. How are we going to adjust this
- 23:49:36model to fit all the data? We don't want
- 23:49:38it to be biased just on three or four
- 23:49:40lines of data coming in. We want it to
- 23:49:42kind of extrapolate a general answer for
- 23:49:45all the data coming in. But this of
- 23:49:46course uh we mentioned it briefly about
- 23:49:49back propagation. This is where really
- 23:49:51comes in handy is training our model.
- 23:49:53Neural network technique to minimize the
- 23:49:56cost function helps to improve the
- 23:49:58performance of the network. Back
- 23:49:59propagates the error and updates the
- 23:50:01weights to reduce the error. So as you
- 23:50:03can see here is a very nice depiction of
- 23:50:05a back propagation. We have our
- 23:50:08predicted y coming out and then we have
- 23:50:10since it's a training set we already
- 23:50:12know the answer and the answer comes
- 23:50:13back and based on case of the square
- 23:50:16means was one of the functions we looked
- 23:50:18at uh one of the activation functions
- 23:50:20based on cost function that cost
- 23:50:22function then depending on what you
- 23:50:24choose for your back propagation method
- 23:50:26and there's a number of them will change
- 23:50:28the weights it will change the weight
- 23:50:29going to each of one of those nodes in
- 23:50:31the hidden layer and then based upon the
- 23:50:34error that's still being carried back
- 23:50:35it'll change the weights going to the
- 23:50:37next hidden layer and then it computes
- 23:50:39an error level on that and sends that
- 23:50:41back up. And you're going to say, well,
- 23:50:43if it computes the error into the first
- 23:50:44hidden layer and fixes it, why would it
- 23:50:46stop there? Well, remember, we don't
- 23:50:48want to create a biased neural network.
- 23:50:52So, we only make small adjustments on
- 23:50:54these weights. We don't make a big
- 23:50:56adjustment that changes everything right
- 23:50:57off the bat. So, no matter how far back
- 23:50:59you go, you're always going to have a
- 23:51:00small amount of error, and that's still
- 23:51:02going to continue to go all the way back
- 23:51:03up the hidden layers. For right now,
- 23:51:05focus on the back propagation is taking
- 23:51:08that error and moving it backwards on
- 23:51:11the neural network to change the weights
- 23:51:13and help program it so that it'll have
- 23:51:15the correct answers. So far, we've been
- 23:51:17talking about forward propagation neural
- 23:51:20networks. Everything goes forwards, goes
- 23:51:22left to right. Uh but let's let's take a
- 23:51:23little detour and let's see what is the
- 23:51:25difference between a feed forward neural
- 23:51:27network and a recurrent neural network.
- 23:51:30Now, this is in the function, not when
- 23:51:31we're training it using the back
- 23:51:33propagation. So, you've got new
- 23:51:34information coming in and you want to
- 23:51:36get the answer and there's a couple
- 23:51:37different networks out there and we want
- 23:51:39to know we have a feed forward neural
- 23:51:40network and we have a new uh vocabulary
- 23:51:42term recurrent neural network. A feed
- 23:51:45forward neural network signals travel in
- 23:51:47one direction from input to output. No
- 23:51:50feedback loops considers only the
- 23:51:52current input cannot memorize previous
- 23:51:54inputs. One example of one of these feed
- 23:51:58forward neural networks. And we've
- 23:51:59covered a number of them, but one of the
- 23:52:00ones that has a big highlight nowadays
- 23:52:02is the CNN, a convolutional neural
- 23:52:05network. TensorFlow, the one put out by
- 23:52:07Google is probably most known for their
- 23:52:10CNN, where the information goes forward.
- 23:52:12It uh first takes a picture, splits it
- 23:52:15apart, goes through the individual
- 23:52:16pixels on the picture, so it picks up a
- 23:52:18different reading, then calculates based
- 23:52:20on that, goes into a regular feed
- 23:52:22forward neural network, and then gives
- 23:52:24you a categorization on there. Now,
- 23:52:26we're not covering the CNN today, but we
- 23:52:28do have a video out that you can look up
- 23:52:30on YouTube put out by SimplyLearn, the
- 23:52:33convolutional neural network. wonderful
- 23:52:35tutorial. Check that out and learn a lot
- 23:52:37more about the convolutional neural
- 23:52:38network. But you do need to know that
- 23:52:40the CNN is a forward propagation neural
- 23:52:43network only. So it's only moving in one
- 23:52:45direction. So we want to look at a
- 23:52:46recurrent neural network. Signals travel
- 23:52:48in both directions making it a looped
- 23:52:51network. Considers the current input
- 23:52:53along with the previous received inputs
- 23:52:55for generating the output of a layer.
- 23:52:57Has the ability to memorize past data
- 23:52:59due to its internal memory. And you can
- 23:53:01see they have a nice uh image here. We
- 23:53:03have our um input and for some reason
- 23:53:06they always do the recurrent neural
- 23:53:07network um in reverse from bottom up in
- 23:53:10the images. It's kind of a standard
- 23:53:11although I'm not sure why. Your X goes
- 23:53:13into your hidden layer and your hidden
- 23:53:16layer the answer for part of the answer
- 23:53:17from that it generates feeds back into
- 23:53:20the hidden layer. So now you have an
- 23:53:22input of both X and part of the hidden
- 23:53:24layer and then that feeds into your
- 23:53:25output. Now if we go back to the forward
- 23:53:28let me just go back a slide and we're
- 23:53:30looking at uh our forward propagation
- 23:53:32network. One of the tricks you can do to
- 23:53:34use just a forward propagation network
- 23:53:37is if you're in a what they call a time
- 23:53:39sequence, that's a good uh term to
- 23:53:41remember or a time series meaning that
- 23:53:43it's sequential data. Each term comes
- 23:53:45after the other. You can trick this by
- 23:53:48creating your input nodes as with the
- 23:53:51history. So if you know that uh you have
- 23:53:53values one, five and seven going in and
- 23:53:55you know what the output is from one
- 23:53:57what those outputs are, you can expand
- 23:53:59the input to include the history input.
- 23:54:02That's one of the ways to trick a
- 23:54:03forward propagation network into looking
- 23:54:05at that. But when you do with a
- 23:54:06recurrent neural network, you let the
- 23:54:09hidden layer do that for you. It sends
- 23:54:11that data and reprocesses it back into
- 23:54:13itself. What are some of the
- 23:54:14applications of recurrent neural
- 23:54:16network? The RNN can be used for
- 23:54:19sentiment analysis and text mining.
- 23:54:21Getting up early in the morning is good
- 23:54:23for health and it's a positive
- 23:54:24sentiment. One of the catches you really
- 23:54:26want to look at this when you're looking
- 23:54:27at the language is that I could switch
- 23:54:29this around and totally negate the
- 23:54:32meaning of what I'm doing. So, it no
- 23:54:33longer be positive. So, when you're
- 23:54:35looking at a sentence, knowing the order
- 23:54:37of the words is as important as the
- 23:54:39meaning of the words. You can't just
- 23:54:41count how many good words there are
- 23:54:42versus bad words to get positive
- 23:54:45sentiment. You know, have to know what
- 23:54:46they're addressing. And there's lots of
- 23:54:48other different uses. Uh, kids are
- 23:54:50playing football or soccer as we call it
- 23:54:52in the US. RN can help you caption an
- 23:54:54image. So based on previous information
- 23:54:56coming in, it refeeds that back in and
- 23:54:58you have a image setter. And then time
- 23:55:01series problems like predicting the
- 23:55:03prices of stocks in a month or quarter
- 23:55:05or sell of product can be solved using
- 23:55:07an RNN. And this is a really good
- 23:55:10example. You have whatever your stocks
- 23:55:11were doing earlier this month will have
- 23:55:13a huge effect of what they're doing
- 23:55:15today if you're investing. So having an
- 23:55:17RNN model, a recurrent neural network
- 23:55:19feeding into itself what was happening
- 23:55:21previously allows it to take that model
- 23:55:23and program in that whole series without
- 23:55:25having to put in the whole a month at a
- 23:55:28time of data. You can only put in one
- 23:55:30day at a time. But if you keep them in
- 23:55:31order, it will look back and say, "Oh,
- 23:55:33this because of what happened yesterday,
- 23:55:34I need some information from that and
- 23:55:35I'm going to use that to help predict
- 23:55:37today's." And so on and so on. We're
- 23:55:39going to go back to our activation
- 23:55:40functions. Remember I told you uh ReLU
- 23:55:43was one of the most common functions
- 23:55:44used. Uh so let's talk a little bit more
- 23:55:46about ReLU and also softmax. Softmax is
- 23:55:49an activation function that generates
- 23:55:51the output between zero and one. It
- 23:55:54divides each output such that the total
- 23:55:56sum of the outputs is equal to one. It
- 23:55:58is often used in the output layers.
- 23:56:00Softmax L of the N equals E to L the N
- 23:56:03over the absolute value of E to the L.
- 23:56:05So what does this function mean? I mean
- 23:56:07what is actually going on here? So we
- 23:56:09have our uh output nodes and our output
- 23:56:12nodes are giving us uh let's say they
- 23:56:13gave us 1.2.9 and point4. As a human
- 23:56:17being I look at that and I say well the
- 23:56:19greatest value is 1.2. So whatever
- 23:56:21category that is if you have three
- 23:56:23different categories maybe you're not
- 23:56:24just doing if it's a cat or it's a dog
- 23:56:27or u oh let's say it's a cow. We had
- 23:56:29cats and dogs earlier. Why the cats and
- 23:56:31dogs are hanging out with a cow. I don't
- 23:56:33know. But we have a value and it might
- 23:56:34say 1.2 2 is a cat, 0.9 is the dog, and
- 23:56:38point4 is a cow. Uh, for some reason, it
- 23:56:40thinks that there's a chance of it being
- 23:56:41any one of these three items, and that's
- 23:56:43how it comes out of the output layer.
- 23:56:44Well, as a human, I can look at 1.2 and
- 23:56:47say this is definitely what it is. It's
- 23:56:48definitely a cat or whatever it is. Uh,
- 23:56:50maybe it's looking at different kinds of
- 23:56:51cars might be a better whether it's a
- 23:56:53car, truck, or a motorcycle. Maybe
- 23:56:55that'd be a better example. Well, from a
- 23:56:57computer standpoint, that might be a
- 23:56:59little confusing because they're just
- 23:57:00numbers waving at us. And so with the
- 23:57:02soft max, we want all those numbers to
- 23:57:06always add up to one. So when I add
- 23:57:08three numbers together, I want the final
- 23:57:10output to be one on there. And so it
- 23:57:12goes through this formula changes each
- 23:57:13of these numbers. In this case, it
- 23:57:15changes them to 46.34
- 23:57:18and 2. They all add up to one. And
- 23:57:20that's a lot easier to register because
- 23:57:22it's very set. It's a set output. It's
- 23:57:24never going to be more than one. It's
- 23:57:26never going to be less than zero. And so
- 23:57:27you can see here that there's probably a
- 23:57:29pretty high chance that it's the first
- 23:57:30one. So you're as a human being, we have
- 23:57:32no problem knowing that. But this output
- 23:57:34can then also go into say another input.
- 23:57:37So it might be an automated car that's
- 23:57:39picking up images and it says that image
- 23:57:41in front of us is probably a big truck.
- 23:57:43We should deal with it like it's a big
- 23:57:45truck. It's probably not a motorcycle.
- 23:57:46Um or whatever those categories are.
- 23:57:48That's the softmax part of it. But now
- 23:57:50we have the ru. Well, what where's the
- 23:57:52ru coming from? Well, the ru is what's
- 23:57:54generating the 1.2 and the 0.9 and the
- 23:57:57point4. And so if you remember our relu
- 23:57:59stands for rectified linear unit and is
- 23:58:03the most widely used activation
- 23:58:05function. We looked at a number of
- 23:58:06different activation functions including
- 23:58:08tangent h the step function. Remember I
- 23:58:10said the step function is really used if
- 23:58:12that's what your actual output is
- 23:58:13because then you know it's a zero or
- 23:58:15one. But the relu if you have that as
- 23:58:17your output you now have a discrepancy
- 23:58:19in there. And if that's going into
- 23:58:21another neural network or another
- 23:58:23process having that discrepancy is
- 23:58:25really important. and it gives an output
- 23:58:26of x if x is positive and zero
- 23:58:29otherwise. So it says my x value is
- 23:58:31going to be somewhere between zero or
- 23:58:33one and then the uh usually unless it's
- 23:58:36really uncertain the output's usually a
- 23:58:38one or zero and then you have that
- 23:58:40little piece of uncertainty there that
- 23:58:41you can send forward to another network
- 23:58:43or you can look at to know that there's
- 23:58:45uncertainty involved and is often used
- 23:58:47in the hidden layers. This is what's
- 23:58:49coming out of the hidden layers into the
- 23:58:50output layer usually or as we reference
- 23:58:53the uh convolution neural network the
- 23:58:56CNN you'd have to go to another video to
- 23:58:58review the RLU is the most common used
- 23:59:01for convolutional part of that network
- 23:59:05has a bunch of little pieces that are
- 23:59:06very simplified looking at all the
- 23:59:08different images or different sections
- 23:59:09of the map and the RLU works really good
- 23:59:12for that like I said there's other
- 23:59:13formulas used but that this is the most
- 23:59:15common one and you'll see that in the
- 23:59:17hidden layers going maybe between one
- 23:59:18layer and the next layer. So just a
- 23:59:20quick recap, we have our soft max, which
- 23:59:23means that if you have uh numerous
- 23:59:25categories, only one of them is going to
- 23:59:27be picked, but you also want to have
- 23:59:29some value attached to it, how well it
- 23:59:31picked it, and you put that between 01.
- 23:59:33So it's very uh standardized. So we have
- 23:59:35our soft max. We looked at that. Let's
- 23:59:37go back one. We looked at that here
- 23:59:38where it transforms the numbers. And
- 23:59:39then we have our ReLU function which
- 23:59:41takes the information in the summation
- 23:59:44and puts it between a zero and a one
- 23:59:46where it's either clearly a zero or
- 23:59:48depending on how confident our model is,
- 23:59:52it'll go between the zero and one value.
- 23:59:54What are hyperparameters? Oh, this is a
- 23:59:57great interview question.
- 23:59:58Hyperparameters. When you are doing
- 24:00:00neural networks, this is what you're
- 24:00:01playing with most of the time once
- 24:00:03you've gotten the data formatted
- 24:00:05correctly. A hyperparameter is a
- 24:00:07parameter whose value is set before the
- 24:00:09learning process begins. Determines how
- 24:00:12a network is trained and the structure
- 24:00:14of the network. This includes things
- 24:00:15like the number of hidden units, how
- 24:00:17many hidden layers are you going to have
- 24:00:18and how many nodes in each layer.
- 24:00:20Learning rate. Learning rate is usually
- 24:00:23multiplied once you figured out the
- 24:00:25error and how much you want to change
- 24:00:26the weights. We talked about or I
- 24:00:28mentioned it earlier just briefly. You
- 24:00:29don't want to just make a huge change
- 24:00:31otherwise you're going to have a biased
- 24:00:32model. So you only take little
- 24:00:34incremental changes and that's what the
- 24:00:35learning rate is is those small
- 24:00:37incremental changes. Epics, how many
- 24:00:39times are you going to go through all
- 24:00:41the data in your training set? So one
- 24:00:43epic is one trip through all the data.
- 24:00:45And there's a lot of other things
- 24:00:46depending on which model you're working
- 24:00:48with and which programming script you're
- 24:00:50working with. Like the Python sklearn
- 24:00:53package will have it slightly different
- 24:00:55than say Google's TensorFlow package
- 24:00:58which will be a little bit different
- 24:00:59than the Spark machine learning package.
- 24:01:01So these are just some examples of the
- 24:01:03hyperparameters. And so you see in here
- 24:01:05we have a nice image of our data coming
- 24:01:07in and we train our model. Then we do a
- 24:01:09comparison to see how good our model is.
- 24:01:11And then we go back and we say, "Hey,
- 24:01:13this this model's pretty good, but it's
- 24:01:14biased." So then we send it back and we
- 24:01:17change our hyperparameters to see if we
- 24:01:19can get an unbiased model or we can have
- 24:01:20a better prediction on it that matches
- 24:01:22our data closer. What will happen if
- 24:01:24learning rate is set too low or too
- 24:01:27high? We have a nice couple graphs here.
- 24:01:29We have one over here. It says a
- 24:01:30learning rate set too low. And you can
- 24:01:32see that it slowly works its way down
- 24:01:34the curve. And on the right you can see
- 24:01:36a learning rate set too high. It's just
- 24:01:38bouncing back and forth. When your
- 24:01:39learning rate is too low, that's what we
- 24:01:41studied two slides ago. That's what the
- 24:01:43learning rate was. Training of the model
- 24:01:45will progress very slowly as we are
- 24:01:47making very tiny updates to the weights.
- 24:01:50We'll take many updates before reaching
- 24:01:51the minimum point. So I just mentioned
- 24:01:54epic going through all the data. You
- 24:01:56might have to go through all the data a
- 24:01:57thousand times instead of 500 times for
- 24:01:59it to train. Learning rate too high
- 24:02:01causes undesirable divergent behavior to
- 24:02:04the loss function due to drastic updates
- 24:02:06and weights. At times it may fail to
- 24:02:08converge or even diverge. So if you have
- 24:02:10your learning rate set too high and it's
- 24:02:12training too quickly, maybe you'll get
- 24:02:14lucky and it trains after one epic run,
- 24:02:16but a lot of times it might never be
- 24:02:18able to train because the weights are
- 24:02:19changing too fast. They they flip back
- 24:02:21and forth too easy. And you see down
- 24:02:22here we've introduced uh two new terms
- 24:02:25converge and diverge. Converge means
- 24:02:29that our model has reached a point where
- 24:02:32it's able to give a fairly good answer
- 24:02:34for all the data we put in. All those
- 24:02:36weights have adjusted and it's minimized
- 24:02:38the error. Diverge means that the data
- 24:02:40is so chaotic that it can never manage
- 24:02:42to to train to that data. The data is
- 24:02:45just too chaotic for it to train. So we
- 24:02:46have two new words there. Converge and
- 24:02:48diverge are important to know. Also what
- 24:02:50is dropout and batch normalization?
- 24:02:53Dropout is a technique of dropping out
- 24:02:56hidden and visible units of a network
- 24:02:58randomly to prevent overfitting of data.
- 24:03:00It doubles the number of iterations
- 24:03:02needed to converge the network. So here
- 24:03:04we have our standard neural network and
- 24:03:06then after applying dropout. Now it
- 24:03:08doesn't mean we actually delete the
- 24:03:10node. The node is still there and we're
- 24:03:12still going to use that node. What it
- 24:03:13means is that we're only going to work
- 24:03:16with a few of the nodes. Um, a lot of
- 24:03:18times I think the most common one right
- 24:03:20now used is 20%. Uh, so you'll drop out
- 24:03:2320% of the nodes when you do your
- 24:03:25training, you reverse propagate your
- 24:03:27data and then you'll randomly pick
- 24:03:29another 20 nodes the next time you go
- 24:03:31through an epic data training. So each
- 24:03:32time you go through one epic, you will
- 24:03:34randomly pick 20 of those nodes not to
- 24:03:36not to mess with. And this allows for
- 24:03:39less overfitting of the data. So by
- 24:03:41randomly doing this you create some I
- 24:03:43guess it just kind of pulls some nodes
- 24:03:44off to the side and says we're going to
- 24:03:45handle the data later on so we don't
- 24:03:46overfit. Batch normalization is a
- 24:03:49technique to improve the performance and
- 24:03:51stability of neural network. The idea is
- 24:03:54to normalize the inputs in every layer
- 24:03:56so that they have mean output and
- 24:03:57activation of zero and standard
- 24:03:59deviation of one. This question covers a
- 24:04:02lot of different things which is great.
- 24:04:04It's a great uh interview question
- 24:04:06because it pulls in that you have to
- 24:04:07understand what the mean value is. So a
- 24:04:10mean output activation of zero that
- 24:04:12means our average activation is zero. So
- 24:04:15when you normalize it remember usually
- 24:04:16we're going between minus1 and one on a
- 24:04:19lot of these. It's a very standard
- 24:04:20setup. So you have to be very aware that
- 24:04:22this is your mean output activation of
- 24:04:24zero. And then we have our standard
- 24:04:25deviation of one. So we want to keep our
- 24:04:28error down to just a one value. The
- 24:04:30benefits of this doing a batch
- 24:04:32normalization is it provides
- 24:04:34regularization. It trains faster, higher
- 24:04:37learning rates and weights are easier to
- 24:04:40initialize. What is the difference
- 24:04:42between batch gradient descent and
- 24:04:44stochastic gradient descent? Batch
- 24:04:47gradient descent. Batch gradient
- 24:04:49computes the gradient using the entire
- 24:04:51data set. It takes time to converge
- 24:04:53because the volume of data is huge and
- 24:04:55weights update slowly. So you can look
- 24:04:57at the batches. A lot of times if you're
- 24:04:59using big data, batch the data in, but
- 24:05:01you still go through a full epic. You
- 24:05:03still go through all the data on there.
- 24:05:05So bash gradient descent means you're
- 24:05:07going to use it to fit all the data and
- 24:05:08look for a convergence there. Stochastic
- 24:05:11gradient descent. Stochastic gradient
- 24:05:13computes the gradient using a single
- 24:05:15sample. It converges much faster than
- 24:05:17batch gradient because it updates weight
- 24:05:19more frequently. Explain overfitting and
- 24:05:21underfitting and how to combat them.
- 24:05:23Overfitting happens when a model learns
- 24:05:25the details and noise in the training
- 24:05:27data to the degree that it adversely
- 24:05:29impacts the execution of the model on
- 24:05:31the new information. It is more likely
- 24:05:33to occur with nonlinear models that have
- 24:05:35more flexibility when learning a target
- 24:05:37function. An example of this would be um
- 24:05:39if you're looking at say cars and trucks
- 24:05:42and motorcycles, it might only recognize
- 24:05:45trucks that have a certain box-like
- 24:05:47shape. It might not be able to notice a
- 24:05:50flatbed truck unless it's only a
- 24:05:52specific kind of flatbed truck or only
- 24:05:54Ford trucks because that's what it saw
- 24:05:55on the training set. This means that
- 24:05:57your model performs great on your train
- 24:05:59data and great on maybe a small test
- 24:06:02amount of data, but when you go to use
- 24:06:04it in the real world, it leaves out a
- 24:06:06lot and start and is not very functional
- 24:06:08outside of your small area, your slow
- 24:06:11laboratory data coming in. Underfitting,
- 24:06:13doing the opposite when you underfit
- 24:06:14your data. Underfitting alludes to a
- 24:06:16model that is neither well-trained on
- 24:06:18training data nor can generalize to new
- 24:06:21information. Usually happens when there
- 24:06:22is less and improper data to train a
- 24:06:25model. has a bur performance and
- 24:06:27accuracy. So if you're using underfitted
- 24:06:29data and you generate a model and you
- 24:06:30distribute that in a commercial zone,
- 24:06:32you'll have a lot of people unhappy with
- 24:06:34you because it's not going to give them
- 24:06:35very good answers. So we've explained
- 24:06:37overfitting and underfitting. So now we
- 24:06:39want to ask how to combat them.
- 24:06:41Combating overfitting and underfitting,
- 24:06:44resampling the data to estimate the
- 24:06:46model accuracy, k-fold cross validation,
- 24:06:49having a validation data set to evaluate
- 24:06:51the model. So when we do the reampling,
- 24:06:53we're randomly going to be picking out
- 24:06:55data and we'll run it a few times to see
- 24:06:57how that works depending on our random
- 24:06:59data and how we sample the data to
- 24:07:01generate our model and then we want to
- 24:07:03go ahead and validate the data set by
- 24:07:05having our training data and then
- 24:07:07keeping some data on the side uh testing
- 24:07:10data to validate it. How are weights
- 24:07:12initialized in a network? Initializing
- 24:07:14all weights to zero. All the weights are
- 24:07:16set to zero. This makes your model
- 24:07:18similar to a linear model. So if you
- 24:07:20have linear data coming in, doing a
- 24:07:22basic setup like that might work. All
- 24:07:23the neurons in every layer perform the
- 24:07:25same operation given the same output and
- 24:07:28making the deep net useless. Right?
- 24:07:30There's a key word. It's going to be
- 24:07:31useless if you initialize everything to
- 24:07:33zero. At that point be looking into some
- 24:07:35other uh machine learning tools.
- 24:07:37Initializing all weights randomly. Here
- 24:07:39the weights are assigned randomly by
- 24:07:41initializing them very close to zero. It
- 24:07:43gives better accuracy to the model since
- 24:07:45every neuron performs different
- 24:07:46computations. And here we have the
- 24:07:49weights are set randomly. We have our
- 24:07:50input layer, the hidden layers and the
- 24:07:52output layer. And W equals NP random
- 24:07:54random N layer size L, layer size L
- 24:07:57minus one. This is the most commonly
- 24:07:59used is to randomly generate your
- 24:08:01weights. What are the different layers
- 24:08:03in CNN? Convolutional neural network.
- 24:08:06First is the convolutional layer that
- 24:08:08performs a convolutional operation. We
- 24:08:11have our other video out if you want to
- 24:08:13explore that more. and go into detail
- 24:08:14exactly how the C the convolutional
- 24:08:17layer works in the CNN as far as
- 24:08:19creating a number of smaller uh picture
- 24:08:21windows that go over the data. Uh the
- 24:08:23second step is as a relu layer relu
- 24:08:26brings nonlinearity to the network and
- 24:08:28converts all the negative pixels to
- 24:08:30zero. Output is rectified feature map.
- 24:08:32So it goes into a mapping feature there.
- 24:08:34Pooling layer pooling is a down sampling
- 24:08:37operation that reduces the
- 24:08:38dimensionality of the feature map. So we
- 24:08:40have all our relu layer which is pulling
- 24:08:42all these little maps out of our
- 24:08:44convolutional layer. It's taking that
- 24:08:46picture and little creating little tiny
- 24:08:48neural networks to look at different
- 24:08:49parts of the picture. Uh then we need to
- 24:08:51pull it together and then finally the
- 24:08:53fully connected layer. So we flatten our
- 24:08:55pooling layer out and we have a fully
- 24:08:57connected layer recognizes and
- 24:08:59classifies the objects in the image. And
- 24:09:01that's actually your forward propagation
- 24:09:03reverse propagation training model
- 24:09:05usually. I mean there's a number of
- 24:09:06different models out there of course.
- 24:09:08What is pooling in CNN and how does it
- 24:09:11work? Pooling used to reduce the spatial
- 24:09:13dimensions of a CNN performs down
- 24:09:16sampling operation to reduce the
- 24:09:18dimensionality. Creates a pulled feature
- 24:09:20map by sliding a filter matrix over the
- 24:09:22input matrix. I mentioned that briefly
- 24:09:25on the previous slide. Um it's important
- 24:09:27to know that you have if you see here
- 24:09:29they have a rectified feature map. And
- 24:09:31so each one of those colors like the
- 24:09:33yellow color that might be one of the a
- 24:09:35smaller little neural network using the
- 24:09:37ReLU. You'll look at it'll just kind of
- 24:09:39um go over the main picture and look at
- 24:09:41all the different areas on the main
- 24:09:42picture. So you might step one 2 3 four
- 24:09:45spaces. Um and then you have another one
- 24:09:47that's also looking at features and it
- 24:09:49has a 2785. Each one of those is a map.
- 24:09:52So it might be the first one might be a
- 24:09:54map looking for cat ears and the second
- 24:09:56one looking for human eyes. When it does
- 24:09:58this, you then have this rectified
- 24:10:00feature map looking at these different
- 24:10:02features and the max pooling with a 2x
- 24:10:04two filters and a stride of two. Stride
- 24:10:06means instead of skipping every pixel,
- 24:10:07you're going to go every two pixels. You
- 24:10:09take the maximum values and you can see
- 24:10:11over here we look at a pulled feature
- 24:10:13map. One of the features says, hey, I
- 24:10:15had a max value of eight. So somewhere
- 24:10:17in here we saw a human eye labeled as
- 24:10:20eight. Pretty high label. And maybe
- 24:10:21seven was a human hand and maybe four
- 24:10:23was cat whiskers or something that we
- 24:10:25thought might be cat whiskers. Four is
- 24:10:27kind of a low number in this particular
- 24:10:29case compared to the other ones. So you
- 24:10:30have your full pool feature map. You can
- 24:10:32see the process here is we have our
- 24:10:34stepping, we look for the max value and
- 24:10:36then we create a poolled feature map of
- 24:10:38the maxed values. How does a LSTM
- 24:10:42network work? That's long shortterm
- 24:10:44memory. So the first thing to know is
- 24:10:46that an LSTMs are a special kind of
- 24:10:48recurrent neural network capable of
- 24:10:50learning long-term dependencies.
- 24:10:52remembering information for long periods
- 24:10:54of time is their default behavior. We
- 24:10:57did look at the RNN briefly talked about
- 24:10:59how the hidden layer feeds back into
- 24:11:01itself. With the LSTM has a much more
- 24:11:05complicated feedback and you can see
- 24:11:06here we have the hidden layer of T minus
- 24:11:09one and the hidden layer that's what the
- 24:11:11H stands for hidden layer of T and the
- 24:11:13formulas going in. As we can see here we
- 24:11:15have the hidden layers we have T minus
- 24:11:18one and then H of T where T stands for
- 24:11:21time. So this is a series remember
- 24:11:22working with series and we want to
- 24:11:24remember the past and you can see you
- 24:11:26have your ex your input of t and that
- 24:11:28might be a frame in a video as a frame
- 24:11:31comes in they usually use in this one
- 24:11:33the tangent h activation formula but you
- 24:11:36also see that it goes through a couple
- 24:11:38other formulas the omega formula and so
- 24:11:40when it combines these that then goes
- 24:11:42into the next layer your next hidden
- 24:11:44layer that then goes into the data
- 24:11:46that's submitted to the next input so
- 24:11:49you have your x of t + one. So when you
- 24:11:51have that coming in, then you have your
- 24:11:53H value that's coming forward from the
- 24:11:55last process. And depending on how many
- 24:11:57of these omega structures you put in
- 24:11:59there depends on how long-term the
- 24:12:01memory gets. So it's important to
- 24:12:03remember this is more for your long-term
- 24:12:05recurrent neural networks. The three
- 24:12:07steps in an LSTM, step one decides what
- 24:12:11to forget and what to remember. Step
- 24:12:13two, selectively update cell state
- 24:12:15values. So based on what we want to
- 24:12:17remember and forget, we want to update
- 24:12:18those cell values and then decides what
- 24:12:20part of the current state make it to the
- 24:12:22output. So now we have to also have an
- 24:12:24output on there. What are vanishing and
- 24:12:27exploding gradients? This is a great
- 24:12:29question that affects all our neural
- 24:12:30networks. While training an RNN, your
- 24:12:33slope can become either too small or too
- 24:12:35large and this makes the training
- 24:12:37difficult. When the slope is too small,
- 24:12:39the problem is known as vanishing
- 24:12:41gradient. So our slope, we have our
- 24:12:43change in x and our change in y. When
- 24:12:45the slope decreases gradually to a very
- 24:12:47small value, sometimes negative, and
- 24:12:49makes training difficult. When the slope
- 24:12:51tends to grow exponentially instead of
- 24:12:53decaying, this problem is called
- 24:12:55exploding gradient. The slope grows
- 24:12:57exponentially. You can see a nice graph
- 24:12:59of that here. Issues in gradient
- 24:13:00problem, long training time, poor
- 24:13:03performance, and low accuracy. What is
- 24:13:05the difference between epic, batch, and
- 24:13:07iteration in deep learning? Epic. An
- 24:13:09epic represents one iteration over the
- 24:13:12entire data set. So that's everything
- 24:13:14you're going to go ahead and put into
- 24:13:15that training model. Batch. We cannot
- 24:13:17pass the entire data set into the neural
- 24:13:19network at once. So we divide the data
- 24:13:21set into a number of batches. And then
- 24:13:24iteration. If we have 10,000 images as
- 24:13:27data and a batch size of 200, then the
- 24:13:30epic should run 10,000 times over 200.
- 24:13:32So that means we have our total number
- 24:13:34over the 200 equals 50 iterations. So in
- 24:13:37each epic we're running over all the
- 24:13:39data set, we're going to have 50
- 24:13:41iterations. And each of those iterations
- 24:13:43includes a batch of 200 images in this
- 24:13:46case. Why TensorFlow is the most
- 24:13:48preferred library in deep learning? Uh
- 24:13:50well, first TensorFlow provides both C++
- 24:13:52and Python APIs that makes it easier to
- 24:13:55work on. Has a faster compilation time
- 24:13:57than other deep learning libraries like
- 24:13:59KAS and torch. TensorFlow supports both
- 24:14:02CPUs and GPUs computing devices. So
- 24:14:05right now TensorFlow is at the top of
- 24:14:08the market because it's so easy to use
- 24:14:10for both programmer side and for
- 24:14:11hardware side and for the speed of
- 24:14:13getting something up and running. What
- 24:14:15do you mean by tensor in TensorFlow?
- 24:14:17Tensor is a mathematical object
- 24:14:19represented as arrays of higher
- 24:14:20dimensions. These arrays of data with
- 24:14:22different dimensions and ranks that are
- 24:14:24fed as input to the neural network are
- 24:14:26called tensors. And you can see here we
- 24:14:28have a tensor of dimensions five, four.
- 24:14:31So it's a two-dimensional tensor coming
- 24:14:32in. Um, you can look at an image like
- 24:14:34this that each one of those pixels is a
- 24:14:37different value if it's a black and
- 24:14:38white. So, it might be zero and ones and
- 24:14:40then each one represents a black and
- 24:14:41white image. In a color photo, you might
- 24:14:43um either find a different value system
- 24:14:46or you might have a tensor value that
- 24:14:48has the xy coordinates as we see here
- 24:14:50plus the colors. So, you might have
- 24:14:51three more different dimensions for the
- 24:14:54three different images, the red, the
- 24:14:56blue, and the yellow coming in. And even
- 24:14:58as you go from one layer or one tensor
- 24:15:01to the next, these layers might change.
- 24:15:03We might flatten them, might bring in
- 24:15:04numerous. In the case of the convergence
- 24:15:07neural network, we have all those
- 24:15:08smaller different mappings of features
- 24:15:10that come in. So each one of those
- 24:15:12layers coming through is a tensor. If it
- 24:15:14has multiple dimensions coming in and
- 24:15:15weights attached to it, what are the
- 24:15:17programming elements in TensorFlow?
- 24:15:19Well, we have our constants. Constants
- 24:15:21are parameters whose value does not
- 24:15:22change. To define a constant, we use
- 24:15:25tf.constant command. Example, A equals
- 24:15:27TF.Constant 2.0 TF float 32. So it's a
- 24:15:31tensor float value of 32. B equals TF
- 24:15:34constant 3.0. Print AB. If we did a
- 24:15:37print of AB, we'd have um TF.stant and
- 24:15:40then of course uh B is that instance of
- 24:15:42it. Variables. Variables allow us to add
- 24:15:44new trainable parameters to graph. To
- 24:15:47define a variable, we use TF.variable
- 24:15:49command and initialize them before
- 24:15:51running the graph in session. Example W
- 24:15:53equals TF variable.3 DT type TF float 32
- 24:15:57or B equals a TF variable minus 3, D
- 24:15:59type float 32. Placeholders.
- 24:16:01Placeholders allow us to feed data to a
- 24:16:04TensorFlow model from outside a model.
- 24:16:06It permits a value to be assigned later.
- 24:16:08To define a placeholder, we use TF
- 24:16:10placeholder command. Example A equals TF
- 24:16:12placeholder b= a * 2 with the TF session
- 24:16:16as SESS result equals session run B,
- 24:16:20feed dictionary equals A3.0. 0 print
- 24:16:22result. Uh so we have a nice example
- 24:16:24there of a placeholder session. A
- 24:16:26session is run to evaluate the nodes.
- 24:16:28This is called as the tensorflow
- 24:16:30runtime. So for example, you have a= tf
- 24:16:33constant 2.0 b= tf constant 4.0 c= a
- 24:16:36plus b. And at this point you'd go ahead
- 24:16:38and create a session equals tf session.
- 24:16:40And then you could evaluate the tensor C
- 24:16:43print session run C. That would input C
- 24:16:46as an input into your session. What do
- 24:16:48you understand by a computational graph?
- 24:16:51Everything in TensorFlow is based on
- 24:16:52creating a computational graph. It has a
- 24:16:55network of nodes where each node
- 24:16:57performs an operation. Nodes represent
- 24:16:59mathematical operation and edges
- 24:17:01represent tensors. Since data flows in a
- 24:17:04form of a graph, it is also called a
- 24:17:06data flow graph. And we have a nice
- 24:17:09visual of this graph or graphic image of
- 24:17:11a computational graph. And you can see
- 24:17:13here we have our input nodes, our add
- 24:17:15multiply nodes and our multiply node at
- 24:17:17the end. And then we have the edges
- 24:17:19where the data flows. So we have from A
- 24:17:22going to C, A going to D. You can see we
- 24:17:24have a two flowing, a four flowing.
- 24:17:26Explain generative adversarial network
- 24:17:29along with an example. Suppose there is
- 24:17:31a wine shop that purchases wine from
- 24:17:33dealers which they will resell later. So
- 24:17:35we have our dealer going to the wine,
- 24:17:36our shop owner that then sells it for a
- 24:17:38profit. But there are some malfactor
- 24:17:40dealers who sell fake wine. In this
- 24:17:42case, the shop owner should be able to
- 24:17:44distinguish between fake and authentic
- 24:17:46wine. The forger will try to different
- 24:17:48techniques to sell fake wine and make
- 24:17:50sure certain techniques go past the shop
- 24:17:52owner's check. So, here's our forger
- 24:17:54fake wine shop owner. The shop owner
- 24:17:56would probably get some feedback from
- 24:17:57the wine experts that some of the wine
- 24:17:59is not original. The owner would have to
- 24:18:01improve how he determines whether a wine
- 24:18:03is fake or authentic. Goal of forger to
- 24:18:05create wines that are indistinguishable
- 24:18:07from the authentic ones. Goal of shop
- 24:18:09owner to accurately tell if the wine is
- 24:18:11real or not. There are two main
- 24:18:12components of generative adversarial
- 24:18:15network. And we refer to as a noise
- 24:18:17vector coming in where we have our
- 24:18:19forger who's going to generate fake wine
- 24:18:21and then we have our real authentic wine
- 24:18:24and of course our shop owner who has to
- 24:18:25figure out whether it's real or fake.
- 24:18:27The generator is a CNN that keeps
- 24:18:30producing images that are closer in
- 24:18:32appearance to the real images while the
- 24:18:34discriminator tries to determine the
- 24:18:35difference between real and fake images.
- 24:18:38The ultimate aim is to make the
- 24:18:39discriminator learn to identify real and
- 24:18:41fake images. What is an autoenccoder?
- 24:18:44The network is trained to reconstruct
- 24:18:46its inputs. It is a neural network that
- 24:18:48has three layers. Here the input neurons
- 24:18:51are equal to the output neuron. The
- 24:18:53network's target outside is same as the
- 24:18:55input. It uses dimensionality reduction
- 24:18:58to restructure the input. Input image
- 24:19:01comes in. We have our Latin space
- 24:19:03representation and then it goes back out
- 24:19:05reconstructing the image. It works by
- 24:19:07compressing the input to a Latin space
- 24:19:09representation and then reconstructing
- 24:19:11the output from this representation.
- 24:19:13What is bagging and boosting? Bagging
- 24:19:15and boosting are ensemble techniques
- 24:19:17where the idea is to train multiple
- 24:19:19models using the same learning algorithm
- 24:19:21and then take a call. So we have in here
- 24:19:23where we're bagging. We take a data set
- 24:19:25and we split it. We're going to have our
- 24:19:26training data and our test data. Very
- 24:19:28standard thing to do. Then we're going
- 24:19:29to randomly select data into the bags
- 24:19:32and train your model separately. So we
- 24:19:34might have bag one, model one, bag two,
- 24:19:36model two, bag three, model 3, and so
- 24:19:38on. In boosting the emphasis is to
- 24:19:40select the data points which give wrong
- 24:19:42output in order to improve the accuracy.
- 24:19:45So in boosting we have our data set
- 24:19:47again we split it to test data and train
- 24:19:49data and we'll take a bag one and we'll
- 24:19:51train the model. Data points with wrong
- 24:19:53predictions then go into bag two and we
- 24:19:55then train that model and repeat.
- 24:19:57machine learning is which is a subset of
- 24:20:00artificial intelligence, right? That's
- 24:20:02uh basically
- 24:20:04um machines learning from data
- 24:20:08in order to uh make decisions
- 24:20:11essentially. Um so this was a big
- 24:20:14departure from the rules-based systems
- 24:20:17at the time, right? That were explicitly
- 24:20:19programmed to make decisions. So just
- 24:20:22think of an example like a really big
- 24:20:25kind of if this then that then that then
- 24:20:27that and and else if this this this
- 24:20:30right so bunch of rules that had to be
- 24:20:33pre-programmed in order to um come out
- 24:20:36with some final answer. Uh with machine
- 24:20:39learning it's the exact opposite of
- 24:20:40that. we're actually training something
- 24:20:42from examples from existing data um in
- 24:20:46order to predict something or um
- 24:20:50make some type of decision. Uh and so
- 24:20:52we're going to learn about the various
- 24:20:54ways we can do machine learning. But if
- 24:20:56you guys remember we um talked about
- 24:20:59some of this like the differences and
- 24:21:01the uh basically rules-based approaches
- 24:21:05to learning from data approach. Um and
- 24:21:09in included in that is going to be uh
- 24:21:11complex unstructured data. So things
- 24:21:14like images, text, audio. What handles
- 24:21:16those really well is uh deep learning
- 24:21:20which we will get to in the course after
- 24:21:22this. But uh those are certainly in
- 24:21:26there as learning from data even complex
- 24:21:28data.
- 24:21:31So we had this picture uh and I think
- 24:21:33this is kind of around where we left off
- 24:21:35last time was uh just distinguishing
- 24:21:40between those three terms. We see
- 24:21:41artificial intelligence, deep learning
- 24:21:43and machine learning kind of used
- 24:21:44interchangeably, but this is really how
- 24:21:46they fit in. Artificial intelligence is
- 24:21:48kind of a broad anything mimicking human
- 24:21:51intelligence. Um which doesn't have to
- 24:21:55be learning from data, but uh machine
- 24:21:57learning is part of that. And then um
- 24:22:00one way to accomplish machine learning
- 24:22:02is to use neural nets which is the focus
- 24:22:04of uh deep learning. Um and so deep
- 24:22:08learning has been has found a lot of
- 24:22:09success especially recently with uh
- 24:22:12those complex data types like images,
- 24:22:15speech, text, right? So deep learning
- 24:22:18used all over the place. Even in um
- 24:22:20modern like generative AI, we see deep
- 24:22:23learning used quite a bit. Um it really
- 24:22:26anything that's using neural nets is uh
- 24:22:29going to be deep learning.
- 24:22:32Um and again we'll focus on that later
- 24:22:35but we're going to be mainly focused on
- 24:22:37machine learning for this course.
- 24:22:39Primarily machine learning that does not
- 24:22:41use neural networks. Okay. So just
- 24:22:43models that are not necessarily neural
- 24:22:45networks
- 24:22:48be our focus.
- 24:22:51So in machine learning we had an example
- 24:22:54of a game uh essentially um learning
- 24:22:59what decisions to make uh based on the
- 24:23:03uh kind of current um state of the
- 24:23:06board. This could be a um you know
- 24:23:09machine learning example that uh learns
- 24:23:12from many previous examples. So a lot of
- 24:23:15data around these games are used to
- 24:23:18train these um kind of robots that can
- 24:23:22play these games and play them at a very
- 24:23:24high level. Um so there's been a lot of
- 24:23:26successes actually in machine learning
- 24:23:28and deep learning um around
- 24:23:32uh playing games like chess or go
- 24:23:36um using machine learning algorithms. So
- 24:23:38pretty cool.
- 24:23:41All right, so I think this is where we
- 24:23:42ended. Last time we said there's a bunch
- 24:23:43of different use cases for machine
- 24:23:45learning. So um recommendation system is
- 24:23:48going to be a big one and we will
- 24:23:50actually study that uh in one of our
- 24:23:53final lessons of this course. Um chat
- 24:23:56bots like generative AI doing sentiment
- 24:23:58analysis chat bots we'll study later but
- 24:24:00those are certainly an application of
- 24:24:02learning from data in order to uh
- 24:24:06generate responses to text prompts
- 24:24:08right. Um spam filtering that's a good
- 24:24:11example like classifying an email as
- 24:24:13spam or not spam. Um that that gets
- 24:24:17trained from examples and uh learning
- 24:24:20from data such as previous emails. Um
- 24:24:24social media posts analysis is another
- 24:24:26kind of text data um use case but you uh
- 24:24:32can do a lot with that text like you can
- 24:24:35predict the sentiment um you can predict
- 24:24:38uh the category of what what the post is
- 24:24:41talking about um those kind of things
- 24:24:44all can be done with machine learning
- 24:24:47>> and many other use cases not on this
- 24:24:49list that we will uh cover
- 24:24:51>> you know as we as we go further.
- 24:25:02Okay, so this is where we kind of left
- 24:25:04off. Um, so what's doing all the hard
- 24:25:08work here is
- 24:25:11>> uh machine learning algorithms. So these
- 24:25:12are things that will um these are things
- 24:25:16that will learn from the data. So they
- 24:25:19are uh they they are basically um
- 24:25:24algorithms or sets of rules that uh or
- 24:25:28mathematical rules I should say not
- 24:25:29formal rules like in the in the sense of
- 24:25:31a rule system but mathematical um
- 24:25:34formulas and mathematical uh rules
- 24:25:37essentially that help us learn from the
- 24:25:41data. So they correlate the data to some
- 24:25:43type of outcome. So some type of
- 24:25:46prediction uh whether that's going to be
- 24:25:49as we will see whether that could be
- 24:25:50like a number like we're predicting a
- 24:25:53price or demand or sales
- 24:25:56um or it could be a category like is
- 24:25:58this transaction fraud or not fraud or
- 24:26:01what's the probability that this is
- 24:26:03fraud um so we have different kinds of
- 24:26:07predictions we can make with machine
- 24:26:08learning
- 24:26:10um but uh we will study the kind of the
- 24:26:14differences of those coming up. Um, but
- 24:26:17machine learning algorithms are really
- 24:26:19what power they're kind of the models,
- 24:26:21right? They're the models that help
- 24:26:23power uh machine learning to actually
- 24:26:26learn from data.
- 24:26:28So, we're going to spend a lot of time
- 24:26:30in this course studying those algorithms
- 24:26:33like the different models that we can
- 24:26:34build and what their differences are,
- 24:26:37what their strengths are, what their
- 24:26:38weaknesses are. We'll we'll learn a lot
- 24:26:40about those.
- 24:26:44Okay. So I guess you can imagine like
- 24:26:46everything is so data dependent, right?
- 24:26:49Um we're learning from data. So uh it
- 24:26:52makes sense that the quality of data
- 24:26:55really really matters here in
- 24:26:56determining how strong the model can be.
- 24:26:59Um so you see this graph here charting
- 24:27:03kind of the um high quality data um
- 24:27:07versus just uh any old data but a decent
- 24:27:11enough quantity of it. Um you can see
- 24:27:14that performance and the performance is
- 24:27:16measured by some evaluation metric. Um,
- 24:27:20so think of it as uh something like an
- 24:27:23accuracy. Like if we were predicting
- 24:27:24fraud or not fraud, how accurate can our
- 24:27:27model get at actually detecting fraud,
- 24:27:30um, it gets better and better and better
- 24:27:33the graph shows that the higher quality
- 24:27:36of data that we have. So there's kind of
- 24:27:39that there's a there's a saying in
- 24:27:40machine learning um, called garbage in
- 24:27:44garbage out. What that means is if you
- 24:27:46have poor data, even the best model in
- 24:27:48the world, poor data is not going to
- 24:27:51result in having a good model that can
- 24:27:53be accurate and perform well. Um, so it
- 24:27:56needs to be high quality, meaning um
- 24:28:00there needs to be a decent amount of it
- 24:28:01and it needs to be labeled appropriately
- 24:28:04as we will will talk about
- 24:28:07um and it needs to not have any, you
- 24:28:10know, significant outliers. it needs to
- 24:28:12be clean, not have those missing values,
- 24:28:15all of those things. Um, you can you
- 24:28:18have a good chance at deriving good
- 24:28:20predictions from higher quality data
- 24:28:24as this kind of shows.
- 24:28:30Okay.
- 24:28:32So, one thing we're going to learn um as
- 24:28:34we go along is
- 24:28:36quantity matters as well. So, not only
- 24:28:38quality, but a decent amount of it. And
- 24:28:41um we're going to learn those kind of
- 24:28:42rules of thumb like how much data do I
- 24:28:45need for certain algorithms. Um one
- 24:28:48thing that we will see is that uh the
- 24:28:51the basic machine learning models that
- 24:28:53we'll study don't need as much as a
- 24:28:56neural network would. It you know neural
- 24:28:59networks are going to require a lot more
- 24:29:02um than a basic machine learning model
- 24:29:05learning model. So uh that's something
- 24:29:08we will see as we go along. But uh this
- 24:29:10is something we'll talk about and
- 24:29:12discuss with each model that we study is
- 24:29:14kind of how much data do we actually
- 24:29:16need to produce a high quality model.
- 24:29:24Okay, any questions uh so far?
- 24:29:34Okay, let's talk about the different
- 24:29:37types of machine learning that we're
- 24:29:39going to discuss. Pime, there's going to
- 24:29:41be two primary ones that we will study
- 24:29:43in this course and then a couple others
- 24:29:46that'll be a little bit more advanced
- 24:29:47that we won't get to but worth knowing
- 24:29:50about. Um, so there's going to be four
- 24:29:52total that we'll study or talk about and
- 24:29:55they'll be on this list here, which is
- 24:29:58um supervised learning and unsupervised
- 24:30:00learning. Now, I'd say the majority of
- 24:30:03our focus will probably be on supervised
- 24:30:06learning, and we'll talk about what that
- 24:30:08means, but we'll also cover unsupervised
- 24:30:12learning as well. And so, we'll look at
- 24:30:14the most popular techniques in each of
- 24:30:16these types of machine learning.
- 24:30:20Um,
- 24:30:21and then we'll talk about these two, but
- 24:30:24not really study them because they're
- 24:30:25more advanced topics um that that will
- 24:30:28be beyond the scope of what we'll do.
- 24:30:30But uh these are going to be um
- 24:30:34different styles of machine learning
- 24:30:38that are going to be characterized by um
- 24:30:42what kinds of predictions they make,
- 24:30:43what kind of data they need and require.
- 24:30:46Um and uh what kind of outcomes they're
- 24:30:50actually producing. Um, so let's let's
- 24:30:55get into each of these, but uh the the
- 24:30:57one that we'll probably spend the
- 24:30:59majority of our time on is going to be
- 24:31:00supervised learning, but we will study
- 24:31:03unsupervised learning as well. We'll
- 24:31:05study both and we're going to talk about
- 24:31:06we're going to define both of those um
- 24:31:08coming up. And again, these will be a
- 24:31:11little bit more advanced topics that we
- 24:31:12won't spend too much time on.
- 24:31:15Um, but but we'll discuss their
- 24:31:17relevancy in machine learning um and
- 24:31:21give a good definition to it.
- 24:31:28Okay.
- 24:31:32All right. Let's start with supervised
- 24:31:34learning. Now this is going to be uh a
- 24:31:38term that really refers to
- 24:31:41using examples. So using labeled
- 24:31:46examples. So here we say labeled data to
- 24:31:50help our model train. In other words,
- 24:31:54help our model be able to predict guided
- 24:31:58by specific input output pairs. So
- 24:32:01supervised really refers to the fact
- 24:32:04that we have answers. We have examples,
- 24:32:10we have answers with those and we use
- 24:32:12that collection of data to build our
- 24:32:15model off of so that we can predict
- 24:32:19um those kinds of things like a price,
- 24:32:23like a category, like a spam not spam.
- 24:32:27in this in this slide like we would be
- 24:32:29predicting if this shape is a square, a
- 24:32:32triangle or a circle.
- 24:32:34Um but but when we build a model for
- 24:32:38that, we have data that has an answer
- 24:32:42attached to it. Right? We've talked
- 24:32:44about this before a little bit with
- 24:32:45labels. So there's a guide there that
- 24:32:50can guide us towards building our model.
- 24:32:52there's an actual every every example
- 24:32:55has an answer and that answer is really
- 24:32:58critical to help build our model off of.
- 24:33:01So, um that's it's almost like you have
- 24:33:06um a you have a bunch of exercises
- 24:33:11in let's say like a math textbook. You
- 24:33:13have a bunch of exercises and you have
- 24:33:15the answers and that way you can kind of
- 24:33:18check your work. You think about model
- 24:33:20training um that is the really a lot of
- 24:33:24that process of model training as we are
- 24:33:26going to discover is um basically
- 24:33:29checking our work against these answers
- 24:33:32in our data in our training data.
- 24:33:35Okay. So supervised learning is any type
- 24:33:38of machine learning that involves
- 24:33:41learning from labeled data in order to
- 24:33:44predict outcomes. Okay, predict outcomes
- 24:33:47like now the the outcomes can be
- 24:33:49numerical. They can be like a price,
- 24:33:52temperature, demand, sales, revenue.
- 24:33:56They can be numerical, but they can also
- 24:33:58be categorical. So they can be like
- 24:33:59spam, not spam, fraud, not fraud,
- 24:34:01cancer, not cancer. Um, dog, cat,
- 24:34:05giraffe, those kind of categories. Um,
- 24:34:09we could predict those. It's some type
- 24:34:11of outcome. Okay, some type of outcome.
- 24:34:13The key is we're using labeled examples
- 24:34:17to guide our model building. That's why
- 24:34:20it's called supervised learning.
- 24:34:23So we know in our data we know what the
- 24:34:26inputs are. Of course, those are going
- 24:34:28to be think of the inputs as like all of
- 24:34:30our columns and then we have a special
- 24:34:34label column that represents the output
- 24:34:36we're trying to predict. So if you think
- 24:34:37about that housing price data, the label
- 24:34:40could be the price. And that's something
- 24:34:42we would build a model to predict, but
- 24:34:44we have answers for all of our examples
- 24:34:47in our rows. We have answers to help
- 24:34:50guide our model building.
- 24:34:52They help tweak our model because we
- 24:34:54know the answer ahead of time. So
- 24:34:58they're they're really good examples to
- 24:34:59build our model off of.
- 24:35:03Okay. So that's that's supervised
- 24:35:05learning.
- 24:35:12Uh in this example is circle not in the
- 24:35:14prediction because it's not part of the
- 24:35:16test data even though it's in the
- 24:35:18labeled data.
- 24:35:20Um no it just not necessarily. It just
- 24:35:23means that like we learn against all of
- 24:35:27these examples that have these answers
- 24:35:30and then when we observe new examples um
- 24:35:33we can try to predict what those would
- 24:35:35be based on what we've seen before. So I
- 24:35:38if there was a you know it's just a
- 24:35:40coincidence we only have two two
- 24:35:42examples in our test data like we could
- 24:35:44have a circle here in which case we
- 24:35:46would predict circle
- 24:35:48that's fine or at least we would hope
- 24:35:50our model would predict circle right
- 24:35:52that's what we're hoping may or may not
- 24:35:54get it right
- 24:35:56um but it's it's only not there because
- 24:36:01we only like we're just assuming that we
- 24:36:03only have two examples we're testing
- 24:36:04against but in reality we would probably
- 24:36:07do a lot more than two.
- 24:36:09It's just it's just a coincidence
- 24:36:10really.
- 24:36:17In reality, we would test against a lot
- 24:36:19more data. And we're actually going to
- 24:36:21see why we would do that. Like why would
- 24:36:24we train our model and then kind of use
- 24:36:28additional data um to to evaluate it?
- 24:36:32It's actually really important that we
- 24:36:34do that step to get a sense of how good
- 24:36:36our model is before we take it out in
- 24:36:38the real world. So if we apply our model
- 24:36:41that we build on our label data to
- 24:36:45um this kind of set of test data that we
- 24:36:49haven't been exposed to before. It helps
- 24:36:52give us a sense of how good is our
- 24:36:54model. So it's tested is usually used
- 24:36:57for evaluation.
- 24:37:00So that's something that's something
- 24:37:01we'll study.
- 24:37:04How do we train? Uh it depends on the
- 24:37:07model. Um so training will be a sense uh
- 24:37:11will be an algorithm that will um
- 24:37:14basically update the model according to
- 24:37:16the data. These labeled examples. Um
- 24:37:19every model is going to be different in
- 24:37:21exactly how it trains. So we're going to
- 24:37:24we're going to talk about that when we
- 24:37:25get to the individual models that we'll
- 24:37:27study.
- 24:37:29But uh loosely speaking, they're going
- 24:37:31to use the data to adjust itself. Like
- 24:37:35imagine adjust like tuning a bunch of
- 24:37:37knobs. Um, like the best example I can
- 24:37:40give you is we I think I did this one
- 24:37:43last week where you have kind of a
- 24:37:45function
- 24:37:46that predicts the price and let's say it
- 24:37:50has
- 24:37:51um weights like weight one with feature
- 24:37:54one, weight two with feature two,
- 24:37:58weight three with feature three. So
- 24:38:01imagine we had three input features and
- 24:38:03we we built an answer according to that.
- 24:38:05Essentially what we would do to train
- 24:38:07the model is adjust these
- 24:38:12um in order to get this correct based on
- 24:38:15our our labeled examples.
- 24:38:19Okay.
- 24:38:21So that's something we're going to learn
- 24:38:22about coming up shortly when we when we
- 24:38:24actually dive into model. Every model is
- 24:38:26going to be slightly different in how it
- 24:38:27trains, but at a high level it's going
- 24:38:29to use the training data with those
- 24:38:33examples, right? the labeled examples to
- 24:38:35help guide the formula essentially to
- 24:38:38adjust to generate the proper kind of
- 24:38:42model here.
- 24:38:44The these things are going to be
- 24:38:45adjusted according to the data
- 24:38:48in order to produce the correct output.
- 24:38:52So think about these as knobs that will
- 24:38:54turn.
- 24:38:59Okay.
- 24:39:02Uh which type of machine learning is
- 24:39:04used? Uh probably supervised um which is
- 24:39:06what we're talking about now. So
- 24:39:08probably supervised because most people
- 24:39:10want to
- 24:39:12um build some type of model to predict
- 24:39:14something.
- 24:39:16Uh so yeah, I'd say I'd say supervise.
- 24:39:25Yes, we're are we are definitely going
- 24:39:26to learn how to train. Yeah, we'll see.
- 24:39:28We'll do the code. Um, I'll tell you
- 24:39:31about how it's done. Yeah, we're
- 24:39:33definitely going to learn it. But what I
- 24:39:34was saying is it's kind of on a model
- 24:39:36bymodel basis.
- 24:39:38So, I want to wait till we get into the
- 24:39:40individual models, then we'll talk about
- 24:39:42how they're trained.
- 24:39:44But yeah, we'll we'll learn how to do
- 24:39:45that.
- 24:39:52But yeah, supervisor is used all over
- 24:39:54the place. Even even for uh generative
- 24:39:57models, they use supervised learning
- 24:40:00because um like an LLM
- 24:40:03is going to use labeled examples in
- 24:40:06order to train, right? In order to train
- 24:40:09how to generate responses according to
- 24:40:12prompts. Um it needs to learn against a
- 24:40:15lot of text examples.
- 24:40:18So that supervised learning is what um
- 24:40:21results in that model,
- 24:40:24right? Learning from those labeled
- 24:40:26examples.
- 24:40:36Okay.
- 24:40:42It is yeah image image uh a lot of um
- 24:40:47yeah a lot of image processing is
- 24:40:48supervised like object detection. So the
- 24:40:52YOLO model is an object detection model.
- 24:40:54Yes. Um because it has to be trained
- 24:40:57right. It has to be trained on uh it has
- 24:41:00to be trained on images
- 24:41:04with labels such as this is what object
- 24:41:06is in this image. This is the box around
- 24:41:09the object.
- 24:41:11Um yes. So if if it's if it ever uses
- 24:41:15label data to train and build the model,
- 24:41:18it is supervised. So YOLO is definitely
- 24:41:21supervised and we actually we will we
- 24:41:24will cover the YOLO model later on in
- 24:41:26our deep learning course. We talk about
- 24:41:29object detection.
- 24:41:31So we'll we'll study that.
- 24:41:34But yeah, it's supervised
- 24:41:45Okay. So on the slide we have some
- 24:41:48common supervised learning algorithms
- 24:41:50that are we will study. So all of these
- 24:41:53we will study and understand what they
- 24:41:55do and how they work but just giving you
- 24:41:58some to name them. linear regression is
- 24:42:00kind of the one I just drew out which is
- 24:42:02the um this is the prototypical like
- 24:42:06easiest to understand model that is kind
- 24:42:08of the um exactly like this where we
- 24:42:11have a weight times a feature
- 24:42:14um a weight times a feature and then a
- 24:42:17weight times a feature
- 24:42:19and on and on and on. You can have as
- 24:42:21many as you want.
- 24:42:23um that is a linear regression. And so
- 24:42:26that is um that's a supervised model
- 24:42:29because we need this value here and we
- 24:42:33need all of our inputs in order to um
- 24:42:36actually train this model and generate
- 24:42:38all those weights
- 24:42:40um that that is uh that uses um labeled
- 24:42:45examples to help tune all those knobs.
- 24:42:48Um same with all these other models. So,
- 24:42:49we're going to talk about decision
- 24:42:50trees. We're going to talk about
- 24:42:51logistic regression and and SVMs, which
- 24:42:53are support vector machines. We'll talk
- 24:42:56about all of those, but they're all
- 24:42:57examples of supervised uh supervised
- 24:42:59learning.
- 24:43:01Okay, we'll talk about all of these.
- 24:43:05They're all supervised because they all
- 24:43:08require labeled examples in order to
- 24:43:10train them and and then subsequently use
- 24:43:13them. Okay.
- 24:43:25Okay. So what are some use case
- 24:43:26examples? So for for instance in uh
- 24:43:29supervised learning we may be predicting
- 24:43:30temperature based on yearly temperature
- 24:43:33trends. So we would have that yearly
- 24:43:36data as our um as our labeled examples
- 24:43:40and those would supervise the learning
- 24:43:42of a model that predicts temperature.
- 24:43:44Um, same thing with predicting crop
- 24:43:46yield based on um, seasonal crop quality
- 24:43:50changes. So maybe we have a bunch of
- 24:43:51features relating to crop quality. We
- 24:43:54could predict crop yield. Um, we would
- 24:43:58just need historical examples with those
- 24:44:01labels, right? What the crop yield is
- 24:44:03for each time period. Let's say we would
- 24:44:07just need those uh, supervised examples
- 24:44:09and we could easily build a model off of
- 24:44:12it.
- 24:44:13Um
- 24:44:15uh this this last one sorting waste
- 24:44:17based on known waste items and their
- 24:44:19corresponding waste types. Um that's
- 24:44:22kind of like spam. It's like filtering
- 24:44:24basically like a spam filtering. Um so
- 24:44:27think of it like the the shapes example.
- 24:44:29We sorting things into squares, circles,
- 24:44:32triangles. Um, same kind of idea here
- 24:44:34where we have a bunch of examples on
- 24:44:36what those um what those waste items
- 24:44:40should uh should belong to, like what
- 24:44:43wastist bins they would go to, for
- 24:44:45example. Um, and those could be labeled
- 24:44:50and therefore then we could um
- 24:44:53understand what category of waste they
- 24:44:55belong to.
- 24:44:57Um, same thing with spam. Something is
- 24:44:59fraud or not fraud. spam or not spam,
- 24:45:02cancer or not cancer. All of those are
- 24:45:03going to be supervised learning examples
- 24:45:05because they're going to require in
- 24:45:07order to train them, they're going to
- 24:45:09require data that has those labels.
- 24:45:12Okay? So, anything that has labels is
- 24:45:15going to be supervised learning. So
- 24:45:20again, this is where we will spend
- 24:45:23probably the the majority of our time is
- 24:45:26doing supervised learning problems, ones
- 24:45:29that we have labeled data. We're
- 24:45:31building a model and we're going to
- 24:45:32predict those those uh labels
- 24:45:35essentially.
- 24:45:44Okay, before we go to unsupervised, any
- 24:45:47questions about uh supervised.
- 24:46:04Okay.
- 24:46:07All right. So supervised requires labels
- 24:46:11in order to have an example to go off of
- 24:46:14to build your model. And that's because
- 24:46:17you're predicting those kind of outcomes
- 24:46:19like spam or not spam, cancer or not
- 24:46:21cancer. Now unsupervised learning is
- 24:46:25completely different. It's the opposite.
- 24:46:28So unsupervised learning is where we do
- 24:46:32not use labels whatsoever. So we're not
- 24:46:35using any labels at all. So it's it it
- 24:46:38can be completely unlabeled or even if
- 24:46:40it's labeled, we're not using labels in
- 24:46:42any way. But um we primarily would say
- 24:46:45it's unlabeled data. We have no guidance
- 24:46:48because we're not using the labels in
- 24:46:50any way. We have no guidance to um
- 24:46:54predict anything, but that's because
- 24:46:56we're not really predicting anything in
- 24:46:57unsupervised learning. Generally, what
- 24:46:59we're doing is looking for some
- 24:47:01structure or pattern.
- 24:47:03Okay, with unsupervised learning, we're
- 24:47:05looking for some structure or pattern.
- 24:47:07So, um, one type of example that's very
- 24:47:12very popular is going to be this second
- 24:47:14one, which is, um, identification
- 24:47:17identification of user groups based on
- 24:47:20similarities or commonalities. Now, this
- 24:47:22is going to be a problem basically known
- 24:47:25as clustering
- 24:47:29and it's a problem we will study quite a
- 24:47:31bit. there's going to turn out to be
- 24:47:33lots of different algorithms that can
- 24:47:35accomplish clustering. So what
- 24:47:37clustering attempts to do is basically
- 24:47:39say um we have data that's like this and
- 24:47:43then data over here and then data over
- 24:47:46here. Let's just group these together.
- 24:47:48So like this should be one group, this
- 24:47:50should be one group and this should be
- 24:47:51one group. And we can find those
- 24:47:54structures and say okay this is group
- 24:47:57one, this is group two and this is group
- 24:48:00three.
- 24:48:02one, two, three. And we can basically
- 24:48:05build what we would call clusters of
- 24:48:07data um based on how close together the
- 24:48:11points are kind of located in these kind
- 24:48:13of cluster zones like these boxes I've
- 24:48:16drawn.
- 24:48:18Okay. Now, that doesn't require any
- 24:48:20label to do which is really fascinating.
- 24:48:22So, unsupervised, you don't need any
- 24:48:24label at all to accomplish the
- 24:48:26algorithm. Um so, clustering is one good
- 24:48:29example.
- 24:48:30um finding outliers or anomalies is
- 24:48:33another. So we don't necessarily have
- 24:48:34any label of what is an outlier or what
- 24:48:37is an anomaly. We are deriving that from
- 24:48:40the features alone. There's no guidance.
- 24:48:43There's no label um to doing like
- 24:48:45outlier detection or anomaly detection.
- 24:48:49Okay, so that's another good example.
- 24:48:50One that's not listed on here um but is
- 24:48:55also really important that we will study
- 24:48:57is something known as dimensionality
- 24:48:59reduction.
- 24:49:01So dim reduction and what that what this
- 24:49:05focuses on is basically compressing the
- 24:49:08data set a bit. So we take our data and
- 24:49:11basically compress it um so that but we
- 24:49:15do it in such a way that we retain as
- 24:49:17much information as we can. This is a
- 24:49:20very like smart compression and what it
- 24:49:23does is it lowers the dimension. Um
- 24:49:26dimension think of the dimension as like
- 24:49:29number of columns.
- 24:49:32Number of columns.
- 24:49:34So imagine we had 100 columns in a data
- 24:49:37frame. What we could do is actually
- 24:49:38reduce that down to 10. So like 10% of
- 24:49:42that. So we reduce it down to 10. And um
- 24:49:47but those 10 are it's not like we
- 24:49:49chopped out um 90 other columns. We um
- 24:49:54smartly kind of compressed all that
- 24:49:56information into these 10 new columns um
- 24:49:59that are compressed versions of the
- 24:50:01hundred that we used to have. Um so
- 24:50:04dimensionality reduction is is another
- 24:50:07unsupervised technique. It requires no
- 24:50:09guidance, no label to do, but is um a
- 24:50:13really useful technique to reduce the
- 24:50:15size of your data if you're doing things
- 24:50:17with it. Um so this is another one that
- 24:50:20we will we'll study how to do it and
- 24:50:23basically more details behind it, what
- 24:50:24the algorithms are.
- 24:50:26Um we'll so probably those two in
- 24:50:30unsupervised will spend the most amount
- 24:50:31of time on clustering and dimensionality
- 24:50:33reduction.
- 24:50:42uh and supervised if some data is
- 24:50:43present but we didn't label it means in
- 24:50:46example we had circle triangle square in
- 24:50:49the training data we add pentagon
- 24:50:58but we didn't label that in that case
- 24:51:04uh yeah so every um in supervised
- 24:51:07learning, every row, think about it as
- 24:51:10like every row in our data frame needs
- 24:51:12to have a label
- 24:51:14uh associated to it. It needs to have a
- 24:51:16a column that represents the label.
- 24:51:21So if we've never seen Pentagon before,
- 24:51:23I can't use that as a label.
- 24:51:29So, it has to the Pentagon has to exist
- 24:51:32in the data if I'm going to be able to
- 24:51:35predict it,
- 24:51:41right? So, I can't predict, right? If
- 24:51:43we've never seen it before, we have no
- 24:51:45examples to go off. We have no guidance.
- 24:51:47So, how could we predict that?
- 24:51:50Right? We can't predict it.
- 24:52:05if it's if it's in there. So if if we
- 24:52:08have labels of Pentagon, let's say, then
- 24:52:11yeah, we could predict Pentagon. We
- 24:52:14could
- 24:52:29remove. Remove what?
- 24:52:36We wouldn't if it was talking about the
- 24:52:37Pentagon, we wouldn't remove that. No,
- 24:52:39let me go back to that page. We wouldn't
- 24:52:42remove it. Um, it's just if it's not in
- 24:52:46our labels, we're not going to be able
- 24:52:47to predict it. So, Pentagon's a good
- 24:52:50example here. Uh, Pentagon is not one of
- 24:52:55our labels. So, it currently is not in
- 24:52:58our data set as one of the labels. We
- 24:53:00only have data that's either a triangle,
- 24:53:02circle, or a square. We don't have
- 24:53:05pentagon. So, I would never be able to
- 24:53:08predict pentagon. I'll never be able to
- 24:53:11do that if I haven't seen examples of it
- 24:53:13before.
- 24:53:15Okay. But let's say we had that in
- 24:53:17there.
- 24:53:19So we had Pentagon.
- 24:53:24So if we had Pentagon, um we could have
- 24:53:27an example of it in our labels
- 24:53:37and then yeah, we it could be then we
- 24:53:38could predict it.
- 24:53:46Yeah. Yeah. The the don't get worried.
- 24:53:48Don't worry about the test data. So the
- 24:53:51test data is just saying here's a new
- 24:53:53here's a shape. What is it? Okay, that's
- 24:53:55a square. Here's a shape. What is it?
- 24:53:57Okay, that's a triangle. And we could
- 24:53:59have as many of those examples as we
- 24:54:01want in our test data. So we could have
- 24:54:03a circle and say, okay, what's this?
- 24:54:06Should be circle,
- 24:54:09right? The test data can be whatever it
- 24:54:11whatever it wants. But yeah, if if we've
- 24:54:13never seen Pentagon before, we're never
- 24:54:15going to be able to predict it.
- 24:54:25These are the the label data and labels
- 24:54:29are basically the talking about the same
- 24:54:31thing. The labels just mean what are the
- 24:54:34categories that are present in our data.
- 24:54:39So in this data we only have three
- 24:54:41labels that are present.
- 24:54:48So the labels is are relative to our
- 24:54:51label data, right? It's saying
- 24:54:54what labels,
- 24:54:58excuse me, what labels uh do we have
- 24:55:03in our data and we only have those three
- 24:55:05circle, triangle, square. So so Pentagon
- 24:55:08would not be part of those labels. We
- 24:55:10couldn't predict it.
- 24:55:24No. So unsupervised is not going to make
- 24:55:27a prediction. That's the big difference
- 24:55:29with unsupervised. They're not going to
- 24:55:31make a prediction like this. Um so
- 24:55:34unsupervised is not going to make a
- 24:55:36prediction. It's going to do something
- 24:55:37different like um basically say like
- 24:55:40these guys are similar, these are
- 24:55:42similar, these are similar, this is a
- 24:55:45cluster, this is a cluster, this is a
- 24:55:46cluster. It's not going to make a
- 24:55:50prediction. That's what supervised
- 24:55:52learning does.
- 24:55:56Clustering, yes, which is unsupervised,
- 24:55:59yes, clustering does not require any
- 24:56:01labels. Unsupervised just means we don't
- 24:56:03have any labels. We don't require any
- 24:56:05labels.
- 24:56:24So the other thing unsupervised might do
- 24:56:26is it might say
- 24:56:29and again without the labels it might
- 24:56:31say that this is an outlier.
- 24:56:36it might say that this guy is an outlier
- 24:56:38because there's only there's only one of
- 24:56:40those and they're not like the other. So
- 24:56:42that that's something that um that's
- 24:56:45something that uh unsupervised could do.
- 24:56:54Um it it yeah and no. It kind of labels
- 24:56:59a cluster in the sense that um it would
- 24:57:03basically assign a number to it like
- 24:57:05this is cluster one, this is cluster
- 24:57:08two, this is cluster three.
- 24:57:13It'll assign a number to it, but it's
- 24:57:15not a very meaningful it doesn't assign
- 24:57:17like a prediction label in the in the
- 24:57:20traditional sense of a label.
- 24:57:22It does provide like a numerical index
- 24:57:24for the cluster to because what we want
- 24:57:26to know is like okay this guy has the
- 24:57:29cluster of one. This guy belongs to
- 24:57:31cluster one. This guy belongs to cluster
- 24:57:33one. This guy belongs to cluster two.
- 24:57:35This guy belongs to cluster two. Does
- 24:57:37that make sense? So there needs to be
- 24:57:38some like index of what cluster you
- 24:57:40belong to.
- 24:57:43So it's kind of like a label but not in
- 24:57:46the traditional like prediction sense.
- 24:58:00Very
- 24:58:15good. So again, unsupervised, no labels.
- 24:58:19You're doing things like
- 24:58:22identifying clusters,
- 24:58:24um identifying outliers, doing
- 24:58:27dimensionality reduction. These are all
- 24:58:29like structure and pattern oriented
- 24:58:32things. They're not predictions of a
- 24:58:34label. Okay? They're not which is what
- 24:58:37we would see in supervised learning.
- 24:58:46Okay. So an example would be that we
- 24:58:49take we put in the data um we can group
- 24:58:53together uh data such as images into
- 24:58:57categories based on similarities um
- 24:59:00which would be like those clusters. So
- 24:59:01there's no these would be groups that we
- 24:59:04don't have any label on ahead of time
- 24:59:06like we don't have we don't say that
- 24:59:07this image should belong to this this
- 24:59:09image should belong to this we derive
- 24:59:12that from the characteristics of the
- 24:59:14data. Um so think like a good example is
- 24:59:18um customer groups. So we would identify
- 24:59:21customers based on like okay do they
- 24:59:24have similar spending levels? How many
- 24:59:27days do they go shopping in a week? How
- 24:59:29much money do they spend? And we can
- 24:59:31kind of group together customers based
- 24:59:33on similar qualities.
- 24:59:36Clustering will find those groups that
- 24:59:38should exist.
- 24:59:41um it will discover those groups based
- 24:59:43on um the similarities in the data, but
- 24:59:47there's no labels that that say like
- 24:59:49this person should be in this group,
- 24:59:52this person should be in this ahead of
- 24:59:54time. There's no labels of that. It gets
- 24:59:57derived during the algorithm. It's
- 24:59:59unsupervised,
- 25:00:02right? There's no unsupervised really
- 25:00:04literally means no guidance. There's no
- 25:00:07guidance to doing it. We just derive
- 25:00:09that from the structure of the data
- 25:00:11which is the similarities.
- 25:00:27Okay.
- 25:00:33All right. So,
- 25:00:36a couple more for you. So we had um
- 25:00:39supervised which uses the labels. We
- 25:00:44have unsupervised which uses no labels
- 25:00:47looking for structure. And then we have
- 25:00:49something that's kind of in between
- 25:00:52which is um what is known as
- 25:00:56semiupervised learning. And this is
- 25:00:59where you use a combination of a little
- 25:01:03bit of label data, but most of your data
- 25:01:05is actually unlabeled data. Um, and you
- 25:01:09try to get some use out of that label
- 25:01:12data in order to um build a model out of
- 25:01:17it. And so uh it uses the um it uses
- 25:01:22that label data to um generally provide
- 25:01:27some guidance on usually what happens
- 25:01:29with semi-supervised learning is you use
- 25:01:32your label data to kind of predict what
- 25:01:35the label should be for the unlabelled
- 25:01:38data and then you can go from there. So
- 25:01:40you can create artificial labels on this
- 25:01:42unlabeled data and then you can use all
- 25:01:46of it once it's all been labeled kind of
- 25:01:48like a supervised learning uh approach.
- 25:01:51So but but this is semi-supervised
- 25:01:53basically refers to the fact that you
- 25:01:55start out with most of your data not
- 25:01:58being labeled but you do have some
- 25:02:01labeled examples and what you can do is
- 25:02:03basically extrapolate those labels into
- 25:02:06the unlabeled data set and then provide
- 25:02:10some artificial labels and then now
- 25:02:13everything has a label you can do
- 25:02:14supervised learning.
- 25:02:16Okay. So, it falls kind of between um
- 25:02:20supervised and and unsupervised.
- 25:02:23Uh and there so this is this is kind of
- 25:02:26rare. Most of the time you're not going
- 25:02:28to do that. You're actually just going
- 25:02:30to um prefer to just start with all
- 25:02:33label data. That's usually the preferred
- 25:02:35approach. Most of the time you'll
- 25:02:37actually just be doing supervised
- 25:02:38learning, not really semi-supervised
- 25:02:41learning. So, it's pretty rare, but um
- 25:02:49it it could like if Yeah, it could if
- 25:02:52the if we had a lot of examples of
- 25:02:54Pentagon and we wanted and so they were
- 25:02:56unlabeled and then we tried to guess
- 25:02:58what kind of shape they were um and
- 25:03:02provide an artificial label uh and then
- 25:03:05um then use that whole data set to build
- 25:03:07a model off of then yeah it could it
- 25:03:10could fall into this category. Okay.
- 25:03:18They Oh, going back to the question,
- 25:03:20they still use some kind of label data
- 25:03:21like age, gender. They use uh that's
- 25:03:24those aren't those aren't really labels.
- 25:03:26That's the features. So, yeah, they
- 25:03:29still use the core features of the data.
- 25:03:32They just don't have any like labels in
- 25:03:34the traditional sense of a label. Like
- 25:03:36you should think of a label as something
- 25:03:39we are trying to predict.
- 25:03:41So whether that's a price, whether
- 25:03:43that's like a category like spam, not
- 25:03:46spam, cancer, not cancer, it's something
- 25:03:48we'd be interested in kind of
- 25:03:49predicting. And so um in our data, we
- 25:03:52would have an answer for every row. We'd
- 25:03:55have one of our columns would be like
- 25:03:56the the result like the outcome answer
- 25:03:59that we're trying to predict. That's the
- 25:04:01label.
- 25:04:03So in unsupervised, we don't have any of
- 25:04:05the labels.
- 25:04:07We do have just the regular features
- 25:04:09like gender, age, income, square
- 25:04:13footage,
- 25:04:17bedrooms, bathrooms, all those things.
- 25:04:24Okay.
- 25:04:28So, we have semi-supervised that falls
- 25:04:30in between supervised. Now, the reason
- 25:04:32it falls between is be is because
- 25:04:35there's a decent amount of data that's
- 25:04:37unlabeled. In fact, a majority of it
- 25:04:39unlabeled. But what we can do is try to
- 25:04:43label it. We can try to take what we
- 25:04:45know from our existing labels and
- 25:04:48predict an artificial label and then use
- 25:04:51all that data together in kind of a
- 25:04:54supervised fashion for a model down the
- 25:04:56road.
- 25:05:02So that's kind of what this picture uh
- 25:05:04says is we can try to take um you know
- 25:05:08maybe we try to infer some labels based
- 25:05:10on we have some some labelled data here.
- 25:05:14We have most of our data is unlabeled
- 25:05:16and we try to supply some labels to it.
- 25:05:20Um like maybe we have a babies category
- 25:05:22of teens, a tween, uh you know youth and
- 25:05:27um adults. Um and then we try so we we
- 25:05:31take our our labels and we try to
- 25:05:35extrapolate those into artificial labels
- 25:05:37for this unlabelled data so that we can
- 25:05:39use it now because then everything has a
- 25:05:42label at this point and then we can just
- 25:05:44go ahead and do supervised learning from
- 25:05:46there.
- 25:05:52So we can do supervised from there. What
- 25:05:53we would prefer to do and what we'll do
- 25:05:55in this course
- 25:05:57um is just start with supervised. We'll
- 25:06:00just start with the labels. We won't try
- 25:06:02to derive artificial labels usually.
- 25:06:04We'll just start with labels.
- 25:06:15So one example in the real world is
- 25:06:18something like Google photos which um
- 25:06:22whenever you take a picture it can
- 25:06:24provide uh uh labels based on previous
- 25:06:28uh images in your library. So it can it
- 25:06:32can produce tags or um labels on those.
- 25:06:37Uh generally when you take that picture
- 25:06:39it's kind of unlabeled unless you go in
- 25:06:41and specifically provide some tags and
- 25:06:44some labels. But um if you don't do that
- 25:06:46it can still it can still uh make it can
- 25:06:51artificially create one of those based
- 25:06:53on the other label data that you already
- 25:06:55have.
- 25:06:57So that's um
- 25:07:00that's an example.
- 25:07:08Okay.
- 25:07:12All right. Last one in terms of machine
- 25:07:14learning. So we have supervised, we have
- 25:07:18unsupervised.
- 25:07:20Uh then we had semi-supervised which is
- 25:07:22somewhere in between a mixture of having
- 25:07:23some unlabelled data and label data. Um
- 25:07:26now we're going to talk about
- 25:07:27reinforcement learning which is
- 25:07:29completely different. Um it's it's
- 25:07:32completely different than the other
- 25:07:33three. It's a type of machine learning
- 25:07:36where we uh basically learn from
- 25:07:39interaction with the environment. And
- 25:07:41you might ask what are we learning? We
- 25:07:44are learning what actions to take in the
- 25:07:48environment. Um and the way we do that
- 25:07:51is by reinforcing
- 25:07:53positive actions that lead to a a
- 25:07:56reward. Um, so that's where the word
- 25:08:00reinforcement comes from is we we
- 25:08:02basically uh imagine like a child
- 25:08:05that's, you know, learning from trial
- 25:08:07and error. Like they're trying to crawl,
- 25:08:09they're trying to walk and they keep
- 25:08:10falling down. um eventually they learn
- 25:08:13how to do it through trial and error and
- 25:08:15they might get a reward or they might um
- 25:08:19reinforce some of those positive
- 25:08:21movements that lead them to walk or
- 25:08:24crawl um or they might learn from the
- 25:08:28penalties, right? They might learn from
- 25:08:31uh some type of feedback. So they might
- 25:08:34learn from falling down like, "Oh, that
- 25:08:35hurts. I should uh support myself a
- 25:08:38little bit better, right?" Or be a
- 25:08:39little more coordinated. Um
- 25:08:42and so they they learn from those
- 25:08:44actions and their interaction with the
- 25:08:46environment. Um
- 25:08:49uh so this is a complex um algorithm
- 25:08:55essentially uh it's it deals a lot with
- 25:08:59um again taking actions. Usually when
- 25:09:02you take an action something changes in
- 25:09:04the environment um then you kind of
- 25:09:08observe some type of feedback. So, think
- 25:09:10about like a a board game where you're
- 25:09:13trying to figure out what move you
- 25:09:15should make or another good example is
- 25:09:17like with a robot um trying to navigate
- 25:09:20a maze. So, like what route should it
- 25:09:23take? Should it move forward? Should it
- 25:09:24move backward? Should it move left or
- 25:09:26right? Those are different actions it
- 25:09:28can take. Also, like a self-driving car,
- 25:09:30should it should it turn? Should it
- 25:09:32speed up? Should it slow down? Those are
- 25:09:34all good examples of things that have
- 25:09:36been trained from reinforcement
- 25:09:38learning.
- 25:09:45Uh yeah. So real world examples would be
- 25:09:48like in a board game uh a a reward would
- 25:09:51be like if you win the game. Um or if
- 25:09:55you like capture a piece like in
- 25:09:57checkers or chess, that's a reward. A
- 25:10:00penalty would be like if you lose the
- 25:10:01game or lose one of your pieces, that
- 25:10:03could be a a penalty.
- 25:10:06um in a board game or sorry in like a a
- 25:10:11robot navigation task, it could get
- 25:10:14rewards for um moving in the right
- 25:10:16direction
- 25:10:18um towards the exit or like when it like
- 25:10:21let's say you wanted to train a robot on
- 25:10:22how to open the door and navigate a
- 25:10:25room. Um you would penalize it for
- 25:10:27bumping into the wall.
- 25:10:29Um you would give it a reward for moving
- 25:10:37usually oh like oh the algorithms
- 25:10:40themselves usually it's like a a step
- 25:10:42function um it's usually it's like a
- 25:10:45discrete function that kind of is based
- 25:10:47on the state so the reward it could be
- 25:10:50like um like depending on the let's
- 25:10:54let's go back to the board game example
- 25:10:55like the reward could be like or even
- 25:10:58the maze let's say like a navigating the
- 25:11:01maze like getting to this let's say this
- 25:11:02was the exit
- 25:11:05And this was the entrance.
- 25:11:09Then if they make it to here, they get a
- 25:11:11numerical like if they make it to the
- 25:11:13exit, they get a numerical reward of
- 25:11:14like plus 100, let's say. So it's just a
- 25:11:17number. And then if they uh like if they
- 25:11:21bump if they go into here, like let's
- 25:11:23say this is kind of like a death trap or
- 25:11:25like a pit, this this would be like a
- 25:11:27minus 100. So it could be like discrete
- 25:11:30numerical values could be the reward if
- 25:11:33they're moving in the right direction.
- 25:11:34Like let's say we want to encourage
- 25:11:36going this way then we could give
- 25:11:37smaller intermediate rewards like this
- 25:11:39should be a plus like if you move
- 25:11:41forward this is a plus five this is a
- 25:11:44plus 10 this is a plus 15 if you're
- 25:11:47moving in the wrong direction away from
- 25:11:49the exit. Um that would be like a minus5
- 25:11:52or a minus 10. Does that make sense? So
- 25:11:55they're they're numerical in nature and
- 25:11:58what you're trying to do is collect the
- 25:11:59most reward. You're trying to get the
- 25:12:01largest reward you can through trial and
- 25:12:05error. So you you try this out many many
- 25:12:07many times. You basically simulate
- 25:12:10running through this maze many many
- 25:12:13times. And what dictates it what
- 25:12:16dictates like where I should go is based
- 25:12:19on what I've observed in the past. It's
- 25:12:21almost like you're a child remembering
- 25:12:22like, okay, what move should I make from
- 25:12:24this space? Like if I'm here, if I'm
- 25:12:27here, which way should I go? Should I go
- 25:12:29down? Should I go right? Should I go
- 25:12:31left? You kind of know that from
- 25:12:33experience.
- 25:12:35Does that make sense? Based on the
- 25:12:36reward that I've seen in the past, like
- 25:12:39when I've moved down, I've gotten a
- 25:12:40higher reward than moving left or right.
- 25:12:44Does that make sense? So, yeah, it's
- 25:12:45it's a numerical value
- 25:12:48as a reward.
- 25:12:57Yeah, that's a great question. Um, how
- 25:13:00does it differentiate rewards based on
- 25:13:02gain and loss? I chess. So it's it's a
- 25:13:05very comp complicated uh answer but
- 25:13:07essentially every so in the chess board
- 25:13:12you can think of the board as like every
- 25:13:14every um
- 25:13:17space is a state.
- 25:13:21So I could be in this state I could be
- 25:13:23in this state and then it's not not only
- 25:13:26is every every uh space but where all
- 25:13:29the other pieces are. So there's lots of
- 25:13:31states that are possible.
- 25:13:33Um, so
- 25:13:36the way there's a way to quantify
- 25:13:39essentially what's the value of taking a
- 25:13:43certain action like moving my piece
- 25:13:44left, moving it right, moving it up or
- 25:13:47down um given the rest of the state. So
- 25:13:51you're you're right, it may be
- 25:13:52beneficial to sacrifice. Um, but we
- 25:13:56would learn that through experience that
- 25:13:58okay, the best move in this situation is
- 25:14:00to sacrifice.
- 25:14:02We would we would have to learn that
- 25:14:04through trial and error many many many
- 25:14:06times which is to say like okay if I'm
- 25:14:08in this current state of the world right
- 25:14:12all these pieces are distributed in this
- 25:14:14way the best move for me right now in
- 25:14:16the long run
- 25:14:18to get the most reward in the long run
- 25:14:21is to actually sacrifice my piece and
- 25:14:22move it right move it into like a bad
- 25:14:25position theoretically but we know from
- 25:14:28experience that's actually the most
- 25:14:29long-term reward is from that position
- 25:14:32like moving it right may be the best
- 25:14:35for me. So what you learn is how to take
- 25:14:38actions
- 25:14:40and actions are usually like move right,
- 25:14:43move left, move up, move down. You think
- 25:14:45about like a self-driving car though,
- 25:14:47that's going to be like slow down, speed
- 25:14:50up, turn your wheel 10 degrees. Um those
- 25:14:55kind of actions.
- 25:14:59So the the short answer is it's there's
- 25:15:02a calculation there that you learn what
- 25:15:06the long-term value of every state is
- 25:15:10every unique state
- 25:15:13and then you're trying to basically say
- 25:15:15what action should I take from that
- 25:15:17state
- 25:15:19given that current state of the
- 25:15:29Okay.
- 25:15:31And I really I really like reinforcement
- 25:15:34learning. It's actually probably my
- 25:15:35favorite field of machine learning.
- 25:15:38Unfortunately, we won't be covering it
- 25:15:40um in our main uh course. We have
- 25:15:43offered uh electives around
- 25:15:45reinforcement learning in the past. So,
- 25:15:47um stay tuned. Maybe when we get to the
- 25:15:49end of this program, uh we'll offer an
- 25:15:51elective on it and if enough people sign
- 25:15:53up for it, we'll we'll run it. But, um
- 25:15:57we it's not part of our we don't really
- 25:15:59cover reinforcement learning as part of
- 25:16:00our main topics. It's it is an advanced
- 25:16:03uh more advanced topic than than what
- 25:16:05we'll cover, but um I I really enjoy it.
- 25:16:08I find it very fascinating.
- 25:16:17Okay. So, all of this is kind of um
- 25:16:20illustrating what I was saying, which is
- 25:16:22um you think of like uh the thing that's
- 25:16:25interacting in the environment like the
- 25:16:27robot or the car or the human moving a
- 25:16:30chest piece is known as the agent. It's
- 25:16:34interacting with the environment by
- 25:16:36taking actions which updates the state
- 25:16:39um of of the environment. So that's
- 25:16:41that's why you see this word state here.
- 25:16:43This gets updated constantly every time
- 25:16:45you take an action. Um ultimately what
- 25:16:48reinforcement learning is trying to do
- 25:16:49is learn the best action like what would
- 25:16:52be the best action to take. Um
- 25:16:56and the best action is is the one that
- 25:16:58leads to the most long-term reward.
- 25:17:01That's the best action. Um, so you have
- 25:17:04to uh you have to learn what you know
- 25:17:09what leads to a good reward by kind of
- 25:17:12experiencing this over and over and over
- 25:17:14through trial and error. So there's a
- 25:17:16lot of um kind of simulation or letting
- 25:17:19the robot try something a lot um in
- 25:17:22order to kind of learn what's rewarding
- 25:17:24and what's not. Think about it again
- 25:17:26like I think a good example is like with
- 25:17:27children, right? you kind of have to let
- 25:17:29them try things until they learn on
- 25:17:32their own what's what can they do and
- 25:17:34what can they not do
- 25:17:37what's the best actions right
- 25:17:41so reinforcement learning has made its
- 25:17:43way into other places so I I said like a
- 25:17:46good example is self-driving cars or ro
- 25:17:48robotics a lot of reinforce
- 25:17:50reinforcement learning is used there one
- 25:17:52place it's found its way into recently
- 25:17:54is recommendation systems have kind of
- 25:17:57merged with reinforcement learning
- 25:17:58learning. Um, and this is because you
- 25:18:03you can imagine there's kind of a
- 25:18:05built-in reward for you clicking on a
- 25:18:08video and kind of watching it.
- 25:18:11Um, so that kind of reinforces that
- 25:18:13recommendation and then uh that's where
- 25:18:17um you can then kind of recommend a
- 25:18:20similar thing and see if that's
- 25:18:22rewarding and generates a click or
- 25:18:25generates some view time or watch time
- 25:18:27or whatever. Um so reinforcement
- 25:18:30learning has found its way into a lot of
- 25:18:32areas. Um recommendations being one of
- 25:18:35them because it's just natural for the
- 25:18:38idea of like what um should I recommend
- 25:18:40next to generate the most reward. In
- 25:18:43this case, the reward is kind of
- 25:18:44correlated to did they click on it or
- 25:18:46not or did they how long did they watch
- 25:18:49watch for longer it's more rewarding
- 25:18:52um those kind of things but uh place
- 25:18:56places where reinforcement learning have
- 25:18:57been used I said self-driving cars um
- 25:19:01games so uh one of the most famous
- 25:19:05examples if you want to look it up is
- 25:19:06the um Alph Go this was in 2016 um the
- 25:19:11Alph Go uh algorithm was a reinforcement
- 25:19:15learning bot that beat um some of the
- 25:19:18world's best Go players, which if you're
- 25:19:21not familiar, Go is a um board game
- 25:19:25that is a little bit more uh complex
- 25:19:29than chess. It has more more uh it's a
- 25:19:32larger board um more pieces to it. Um
- 25:19:37but they there was a reinforcement
- 25:19:39learning powered bot that actually um
- 25:19:41learned how to play the game so
- 25:19:43effective it could beat um world kind of
- 25:19:46masters at the games was pretty amazing.
- 25:19:49Um that's the alpha go and that was by
- 25:19:51deep mind Google and deep mind in 2016.
- 25:19:55That was pretty that was only 10 years
- 25:19:57ago not that long.
- 25:19:59Um so certain uh we said recommendation
- 25:20:04uh even autocorrect um learning to
- 25:20:07predict like what is the best correction
- 25:20:10uh to generate a reward which would be
- 25:20:12like you accept that correction or you
- 25:20:14reject it would be a penalty. Um so
- 25:20:17reinforced learning has been adapted to
- 25:20:20these kind of problems very
- 25:20:21successfully. Let's take a look at the
- 25:20:24packages that we will use throughout. So
- 25:20:27um of course we will rely on these three
- 25:20:30which we've already relied on to do a
- 25:20:33lot of things like numpy to do numerical
- 25:20:36manipulations and calculations.
- 25:20:39Uh mapplot lib to do any plotting and
- 25:20:42not only map lib but maybe seabour as
- 25:20:44well both of those to do plotting. Um,
- 25:20:48pandas is a big one because
- 25:20:51that's where all of our data is going to
- 25:20:53be manipulated and prepped before it
- 25:20:55goes into modeling.
- 25:20:57So all of that stuff we learn from
- 25:20:59pandis is definitely going to be applied
- 25:21:01here in this course uh as we actually
- 25:21:04build models. Um so of course like these
- 25:21:07old ones that we've been working with
- 25:21:09quite a bit um still going to be useful
- 25:21:12here in the modeling stage. Um mainly
- 25:21:16for different reasons though mostly to
- 25:21:17get our data prepared to do some type of
- 25:21:20modeling or maybe to visualize it before
- 25:21:23we do modeling to get a sense of what it
- 25:21:24looks like those kind of things.
- 25:21:29Um,
- 25:21:31sci-fi is sometimes useful for certain
- 25:21:35uh um processing like in unsupervised
- 25:21:38learning. We'll actually use scyp a
- 25:21:39little bit to do dimensionality
- 25:21:41reduction or help us do that. Um so
- 25:21:44scypi will be used here and there and
- 25:21:47we've seen it before with hypothesis
- 25:21:49testing. We use scypi like the t test
- 25:21:51and z test came from there. Um, some of
- 25:21:54the unsupervised learning stuff will
- 25:21:56come out of there, but the package we
- 25:21:58will use by far the most in this course
- 25:22:02is going to be Scikitlearn,
- 25:22:05which is here. Um, and we've already
- 25:22:09seen a little bit about scikitlearn in
- 25:22:11terms of its pre-processing capability.
- 25:22:14So, we use the uh minmax scaler and the
- 25:22:18standard scaler from there from the
- 25:22:20pre-processing module in scikitlearn.
- 25:22:23but it has um many different models
- 25:22:27built into it that we can use to help uh
- 25:22:30do our training and predictions. Um so
- 25:22:34it's a incredibly useful machine
- 25:22:36learning library. It is the industry
- 25:22:38standard machine learning library. Um if
- 25:22:42you're going to do anything in machine
- 25:22:43learning, it would be expected that you
- 25:22:46know how to use scikitlearn.
- 25:22:48Now what's really lucky about that is
- 25:22:50that scikitlearn is a really easy
- 25:22:53package to get used to. Nearly
- 25:22:55everything we do in scikitlearn will
- 25:22:57mostly follow the same pattern and so um
- 25:23:00the code will be extremely simple. They
- 25:23:02did a great job with that package of
- 25:23:04making things really user friendly,
- 25:23:06really simple. Um it's a really
- 25:23:09fantastic package and we're going to get
- 25:23:10a lot of practice with it uh as we go
- 25:23:13along. Every model we build will
- 25:23:14essentially be from scikitlearn
- 25:23:17and not only like the models but um
- 25:23:20doing the training doing the predictions
- 25:23:22and then doing the evaluation will all
- 25:23:24come from different uh scikitlearn u
- 25:23:27modules. So that'll be really nice and
- 25:23:30we'll get um good exposure to that
- 25:23:33package throughout the course. So if
- 25:23:35anything will come away from this course
- 25:23:38as um psychit learn uh uh experts
- 25:23:42that'll be very nice. So this is this
- 25:23:45will be the new one for us psychitlearn
- 25:23:47but we'll get a lot of practice with it.
- 25:23:52Okay.
- 25:23:56All right. So just to recap that lesson
- 25:23:59before we move on to lesson three. Um we
- 25:24:01talked about machine learning as
- 25:24:02learning from data. um which is included
- 25:24:06underneath the AI umbrella. But deep
- 25:24:08learning is also included under machine
- 25:24:10learning because it's still learning
- 25:24:11from data but it's learning using neural
- 25:24:14networks.
- 25:24:15Um we talked about the four different
- 25:24:17types of machine learning. We had
- 25:24:18supervised, unsupervised,
- 25:24:21semi-supervised and reinforcement. So
- 25:24:24those are the the different types of
- 25:24:25machine learning that are out there. Um
- 25:24:28and then we talked about some of the pi
- 25:24:30python packages uh that we will use the
- 25:24:33main one being scikitlearn and of course
- 25:24:35we'll use our older like pandas to
- 25:24:37manipulate our data and get it uh pass
- 25:24:39it into our model training etc.
- 25:24:42But scikitlearn will be uh our go-to for
- 25:24:46anything machine learning.
- 25:24:50All right. So, some questions for you
- 25:24:52guys, some checks.
- 25:24:55So, let me know in the chat. What do you
- 25:24:56guys think? Uh, which of the following
- 25:24:59best describes machine learning?
- 25:25:07Which choice do you think makes the best
- 25:25:09is the best for this?
- 25:25:52Very good. Very good. I see I see a lot
- 25:25:54of choices for A and A would be the
- 25:25:56correct choice. So machine learning is
- 25:25:59definitely um a a subset of AI. that's
- 25:26:04underneath that AI umbrella, but of
- 25:26:05course we're learning from experience
- 25:26:07and of course that experience is
- 25:26:09recorded in the data um without being
- 25:26:12explicitly programmed. Uh so it's the
- 25:26:14exact opposite of BNC. We're definitely
- 25:26:16not learning from rules and it's
- 25:26:18definitely not just used for image and
- 25:26:21speech recognition. It can be used for
- 25:26:23many other things beyond those. So yeah,
- 25:26:26A is the best choice there.
- 25:26:29What do we say here?
- 25:26:33Okay. What do you guys think about this?
- 25:26:35Which example illustrates the use of
- 25:26:37machine learning to enhance customer
- 25:26:38experience in an ecommerce company?
- 25:26:53In other words, what would be some what
- 25:26:54would be some uh typical use cases of
- 25:26:57machine learning?
- 25:27:25Good. So I think uh C is going to be the
- 25:27:28best answer here. Definitely C. So it's
- 25:27:31using machine learning to do uh fraud
- 25:27:34transactions. So so that would be a
- 25:27:36prediction probably a supervised
- 25:27:38learning right if if this is fraud or
- 25:27:40not fraud. Um and then maybe some
- 25:27:43customer behavior uh that might be
- 25:27:46unsupervised. So maybe grouping together
- 25:27:48customers uh clustering them based on
- 25:27:50their data like their shopping behavior
- 25:27:53and characteristics. Um that that might
- 25:27:57be unsupervised but either way it's
- 25:27:58machine learning.
- 25:28:01Okay.
- 25:28:06Okay. Final one. What distinguishes deep
- 25:28:09learning from machine learning and
- 25:28:11artificial intelligence? So what's
- 25:28:12unique about deep learning?
- 25:28:43Oh, very good. Yep. So, deep learning
- 25:28:45uses neural networks as so you guys are
- 25:28:49right on top of that. Neural deep
- 25:28:50learning uses neural nets. That's what
- 25:28:52makes it unique. So, machine learning
- 25:28:55would be part A. Machine learning is
- 25:28:57focused on learning from data.
- 25:28:58underneath of that is learning from data
- 25:29:01using neural networks which is what uh
- 25:29:03deep learning is.
- 25:29:07Very good.
- 25:29:10All right, let's go to lesson three.
- 25:29:14And lesson three has two notebooks.
- 25:29:16We're going to be starting with 3.1.
- 25:29:20So, you'll want to open up that
- 25:29:21notebook. I'm going to go over to it
- 25:29:23now. Give you a moment to open that up.
- 25:29:33So, we're going to open the 3.1
- 25:29:34notebook. Um, there's two of them. We'll
- 25:29:37see how far if we can get into the
- 25:29:39second one today. Probably will.
- 25:29:42Um, but we're going to do the uh we're
- 25:29:44going to start with 3.1 notebook. Do you
- 25:29:46guys have this notebook? Should be in
- 25:29:48your materials for for this course.
- 25:29:53Let me give you a moment to open that
- 25:29:54one.
- 25:30:10Do you guys have it?
- 25:30:27All right. So, we're going to start by
- 25:30:30talking about uh supervised learning
- 25:30:34um in our machine learning journey. So
- 25:30:36remember, we're going to talk about uh
- 25:30:38supervised and unsupervised after we do
- 25:30:40supervised. Um and there's going to be a
- 25:30:43lot to cover with supervised mainly
- 25:30:45because um there are uh two different
- 25:30:49types of problems we can tackle uh which
- 25:30:52will be uh we'll talk about in a moment
- 25:30:54predicting different kinds of values. Um
- 25:30:57but let's talk about the kind of what
- 25:30:59we're hoping to learn here which is um
- 25:31:01talk about the different kinds of
- 25:31:02problems that we'll study which are
- 25:31:04these these categories of supervised
- 25:31:06learning. Um those two categories are
- 25:31:08going to be called classification and
- 25:31:09regression. We'll talk about those and
- 25:31:11their differences and then talk about
- 25:31:13some applications and some uh example
- 25:31:16algorithms
- 25:31:18and that's just within this notebook. Um
- 25:31:203.2 two we'll get into uh regression in
- 25:31:24particular
- 25:31:25um which will be uh very very
- 25:31:28interesting. Okay. So that'll be our
- 25:31:30first models that we'll build will be
- 25:31:32over there in 3.2.
- 25:31:36Okay. So if you guys remember um
- 25:31:38supervised learning is where we learn
- 25:31:40from labeled data. So we have input and
- 25:31:43outputs in our in our data set. Um and
- 25:31:47you so you train a model on this data
- 25:31:50that includes input features and
- 25:31:53corresponding outputs that are that are
- 25:31:57the labels. Right? So um the goal is to
- 25:32:01learn a relationship between the input
- 25:32:04and the output. Of course that's what
- 25:32:05any model is trying to do. Um, and what
- 25:32:09this allows us to do is then take that
- 25:32:12model and use it to make predictions on
- 25:32:14never-beforeseen
- 25:32:16uh data. Right? So then we have a
- 25:32:19predictive model out of that that we can
- 25:32:21use um going forward on new examples.
- 25:32:25Um so
- 25:32:27remember we will have in our data a
- 25:32:30bunch of features which are columns and
- 25:32:33then generally one of those columns will
- 25:32:34be the label that we're trying to
- 25:32:36predict.
- 25:32:38And our model is going to try to learn
- 25:32:40some type of relationship between those
- 25:32:42inputs and the output label. So the
- 25:32:46output label could be like fraud not
- 25:32:47fraud, cancer not cancer, uh a price, a
- 25:32:51temperature, those kind of things.
- 25:32:55So let's talk about that. inside of um
- 25:32:58supervised learning there are two
- 25:33:00different types of learning that we can
- 25:33:03do and they're really based on the label
- 25:33:06or sometimes that label is known as the
- 25:33:09target that we're trying to predict. Um
- 25:33:12and depending on that type we get these
- 25:33:15two different categories of learning or
- 25:33:16two different types of learning. One is
- 25:33:19known as regression. So that's generally
- 25:33:22when we are predicting something that is
- 25:33:25continuous or something that is a
- 25:33:27numerical.
- 25:33:30So numerical
- 25:33:32numerical value. So think of price,
- 25:33:35think of temperature, think of revenue.
- 25:33:37We're trying to predict something like
- 25:33:39that. Um versus something that is
- 25:33:42categorical. So that the predicting
- 25:33:45something categorical would be like
- 25:33:46fraud, not fraud, spam, not spam. um
- 25:33:49those are discrete categories and the
- 25:33:53problem of predicting categories is is
- 25:33:55known as classification because we're
- 25:33:59trying to classify examples as belonging
- 25:34:02to one category or another.
- 25:34:06So we have these two main types of
- 25:34:09supervised learning problems. we have
- 25:34:11regression and we have classification
- 25:34:13and they're going to be handled slightly
- 25:34:16differently
- 25:34:17um for many reasons that we're going to
- 25:34:20uncover. Um one of the primary reasons
- 25:34:23is that of course we're predicting
- 25:34:26something that's continuous in the
- 25:34:27regression case versus something
- 25:34:28discrete. So the models have to be
- 25:34:31slightly different to account for that.
- 25:34:33Um but then a step beyond that is the
- 25:34:37evaluation has to be different too. Um I
- 25:34:40kind of alluded to this last week, but
- 25:34:42when you're predicting a regression,
- 25:34:43it's very very difficult to to get the
- 25:34:46exact numerical answer. So um generally
- 25:34:51we don't care about that. Um generally
- 25:34:56we don't care about getting exactly uh
- 25:34:59we don't care about getting it exactly
- 25:35:00right.
- 25:35:02um we just care about getting it um
- 25:35:05we're just we care about getting it
- 25:35:07nearby, getting it close enough. Um
- 25:35:10whereas classification, we do care about
- 25:35:13getting exactly right because it's a
- 25:35:14discrete category. So we're going to be
- 25:35:17able to evaluate that a little bit
- 25:35:18differently to say did we get the answer
- 25:35:20right or wrong. Regression is going to
- 25:35:22be did we get close? Um because it's we
- 25:35:25assume it's going to be nearly
- 25:35:26impossible to predict a a continuous
- 25:35:29number. Um, that's very hard to do.
- 25:35:34Okay.
- 25:35:36So, any questions on
- 25:35:38uh that?
- 25:35:44Any questions on those two differences?
- 25:35:46Let me give you some examples. Maybe
- 25:35:47it'll it'll help too.
- 25:35:51So, again, the classification is going
- 25:35:52to be predicting uh something that's
- 25:35:55categorical. regression is going to be
- 25:35:57predicting something that is continuous.
- 25:36:04So think about trying to predict the
- 25:36:06price of a house based on those other
- 25:36:08features we talked about before like
- 25:36:09square footage, bedrooms, bathrooms, all
- 25:36:12those things we predict the price. That
- 25:36:14would be a regression problem because
- 25:36:16the price is a continuous value.
- 25:36:19Let's take a look at an example here.
- 25:36:22Um, imagine we were trying to uh predict
- 25:36:26the temperature tomorrow. That's going
- 25:36:29to be a regression problem, a a
- 25:36:31supervised learning kind of regression
- 25:36:33problem because we're trying to predict
- 25:36:35a numerical temperature.
- 25:36:39Okay? And versus a category like a
- 25:36:43discrete category would be this would be
- 25:36:45a classification. So this is a
- 25:36:47regression on the left. This is a
- 25:36:49classification
- 25:36:52on the right. Classification
- 25:36:56um because we are um predicting one of
- 25:37:01two categories. Is it just hot or cold?
- 25:37:03Now, we're not saying exactly where that
- 25:37:05threshold is on what's hot or cold. That
- 25:37:08would be a decision on on what we want
- 25:37:10to what our discrete categories actually
- 25:37:12mean.
- 25:37:13But, um we only have two choices, hot or
- 25:37:17cold.
- 25:37:18versus predicting the entire temperature
- 25:37:21which would be um a numerical prediction
- 25:37:24of some exact number. Right? So that'd
- 25:37:28be a regression and then on the right
- 25:37:30would be a classification. Um now again
- 25:37:34why is this so different? You can see
- 25:37:35the types of predictions we're making
- 25:37:37are completely different. One's a
- 25:37:38number, one's a category. But again with
- 25:37:40evaluation it's like if the if the true
- 25:37:44answer in our labels was 84
- 25:37:48and we predicted 83 that's a pretty good
- 25:37:51result. That's still pretty close.
- 25:37:53That's pretty close to this. So from an
- 25:37:55evaluation perspective that's pretty
- 25:37:57good. Um whereas like if I predicted
- 25:38:00cold and it's actually hot that's that's
- 25:38:02a wrong answer. So they're evaluated
- 25:38:05slightly different.
- 25:38:08Um, and that's something we're going to
- 25:38:10see as we talk about evaluation of our
- 25:38:13models once we build them is depending
- 25:38:16on if it's classification regression,
- 25:38:17there's going to be different ways of
- 25:38:18evaluating them.
- 25:38:22You can kind of see why it's very
- 25:38:23difficult to say, okay, we got exactly
- 25:38:2684 when it could be any number. Our
- 25:38:30model is going to be predicting a
- 25:38:31number. That's really hard to pin down
- 25:38:34an exact floatingoint number. So, the
- 25:38:36best we can do is kind of say, how close
- 25:38:39did I get? Like, this would be a worse
- 25:38:40answer. If I got something all the way
- 25:38:42down here, that's a really long distance
- 25:38:44to here. That's bad. That's a bad
- 25:38:47prediction. But if I get something
- 25:38:48really close, that's better, right?
- 25:38:51That's a decent prediction because it's
- 25:38:53pretty close,
- 25:38:55right?
- 25:38:57Of course, being perfect would be
- 25:38:58getting exactly right, but that would be
- 25:39:00nearly impossible to do.
- 25:39:13Okay.
- 25:39:18All right. Any questions on this?
- 25:39:21Does it make sense on regression versus
- 25:39:23classification? We're going to use those
- 25:39:24words quite a bit as we go along. So
- 25:39:27regression predicting that continuous
- 25:39:29value classification predicting a
- 25:39:31category
- 25:39:34and they're going to be um different
- 25:39:36models that do that
- 25:39:41different models being used for
- 25:39:42regression versus different models being
- 25:39:44used for classification.
- 25:39:52All right, let's talk about supervised
- 25:39:55learning. uh applications here. So just
- 25:39:58to name a few, we have HR operations.
- 25:40:02Imagine your recruiter tasked with
- 25:40:03finding the best candidates. Um so
- 25:40:06supervised learning can help by um
- 25:40:08rejecting or accepting candidates. Now
- 25:40:10this is something that happens quite a
- 25:40:12bit even today. Um and that it's kind of
- 25:40:17like uh how recommendations happen like
- 25:40:20this this resume should be um
- 25:40:22recommended this should not um from a
- 25:40:24whole pool of applications. Um so
- 25:40:28there's those kind of use cases of of um
- 25:40:32predicting a category that would be like
- 25:40:34a classification. Should we should we
- 25:40:36accept or reject the the candidate?
- 25:40:39um finance. You see this all the time
- 25:40:41with things like risk and loan
- 25:40:44approvals.
- 25:40:45Um you can uh predict the the the
- 25:40:49category of like if the if the loan if
- 25:40:52we should accept or reject the loan
- 25:40:54application. Um you know that would be a
- 25:40:57classification.
- 25:40:59Um what's interesting about
- 25:41:01classifications by the way so it says
- 25:41:03here like we can predict the likelihood
- 25:41:06of a of a loan being repaid.
- 25:41:09um is a lot of classifications um we we
- 25:41:13say that they predict a category but
- 25:41:16under the hood they can actually predict
- 25:41:18a probability and we turn that
- 25:41:20probability into a category. So um you
- 25:41:25know like we could say what's we could
- 25:41:27say the likelihood of her loan being
- 25:41:29repaid is very low. Let's say it's less
- 25:41:31than 50% probability. Um then we could
- 25:41:35label this as reject,
- 25:41:38right? Right? We could label that as a
- 25:41:39rejection. Um if it's greater than 50%.
- 25:41:43Then we could label this as accept. So
- 25:41:46we can set a threshold there
- 25:41:49and say okay truly we're predicting a
- 25:41:52prob like our model spits out a
- 25:41:54probability but we turn that into a
- 25:41:57category by saying should we accept if
- 25:41:59it's less than 50% we should reject if
- 25:42:02it's greater than we should accept.
- 25:42:05Okay. So that's something we will see
- 25:42:06with some of our classification models
- 25:42:08is that they actually produce a
- 25:42:10probability and we turn that probability
- 25:42:12into a category label
- 25:42:16um by by doing something simple like
- 25:42:18this putting a threshold on it um for
- 25:42:21the for the category.
- 25:42:25So finances is used all over the place.
- 25:42:27Not only just loans like fraud, we
- 25:42:29talked about fraud, not fraud. That
- 25:42:30would be a classification.
- 25:42:32Um predicting sales revenue, that would
- 25:42:36be a regression, right? What is the
- 25:42:38revenue going to be in the next two
- 25:42:40quarters? That's going to be a
- 25:42:42regression problem.
- 25:42:45Uh emails like spam, not spam, that's
- 25:42:48going to be a classification.
- 25:42:50um that's going to operate on the that's
- 25:42:52going to take the text input and predict
- 25:42:54if this email is a spam or a not spam.
- 25:42:58That's going to be a uh supervised
- 25:43:00learning problem, but it's going to be a
- 25:43:02classification problem,
- 25:43:05right? Uh manufacturing supervised
- 25:43:08learning is used to inspect and uh
- 25:43:11quality and classify products in
- 25:43:12different grades. For example, a factory
- 25:43:14might use a model to check for defects.
- 25:43:16So this is actually something that
- 25:43:17happens is you look at images of
- 25:43:19products as they go through the assembly
- 25:43:21line and you can take a look at those
- 25:43:23images and predict if it's a high
- 25:43:25quality, low quality, medium quality. Um
- 25:43:28so they can be this is a classification,
- 25:43:30right? They're going into different
- 25:43:31categories of quality. Um so it's much
- 25:43:35much like a manual kind of intervention
- 25:43:37by some uh QA or quality control uh
- 25:43:42specialist.
- 25:43:44Okay. But that's a classification.
- 25:43:51So in the maritime industry, supervised
- 25:43:53learning can be used to predict current.
- 25:43:55So current level
- 25:43:57um and that can be used to forecast uh
- 25:44:00supply and demand. Um so those would be
- 25:44:03like regression models that are used to
- 25:44:07predict um kind of like temperature but
- 25:44:09in this case like title levels.
- 25:44:14We talked about fraud already, so that's
- 25:44:16there. Um, that would be a
- 25:44:18classification.
- 25:44:23Okay,
- 25:44:26any questions on these uh examples?
- 25:44:30Of course, there's many more. Um
- 25:44:34recommendation is kind of like a
- 25:44:36supervised learning problem uh where you
- 25:44:39are
- 25:44:41taking examples of things that people
- 25:44:43have viewed in the past or or reviewed
- 25:44:46in the past and using that to predict
- 25:44:48what they would want to watch in the
- 25:44:50future. Um so recommendation is
- 25:44:54supervised learning. Um and it's like a
- 25:44:58classification, you know, trying to
- 25:45:00predict um uh certain number of
- 25:45:03categories of of uh shows or movies that
- 25:45:07you would want to watch. Um
- 25:45:11and that's something that we will study
- 25:45:13in the future. Recommend we'll we'll
- 25:45:15have a whole lesson dedicated to
- 25:45:16recommendation as well.
- 25:45:21All right.
- 25:45:23So when it comes down to the uh actual
- 25:45:28models themselves, so there's going to
- 25:45:30be lots of different models that we are
- 25:45:31going to cover. Um and they are um going
- 25:45:36to be different in their purpose and
- 25:45:38kind of their uh what kinds of problems
- 25:45:41they're used for. Um and uh their their
- 25:45:46how they actually train is going to be
- 25:45:48different. Um, but at a high level,
- 25:45:51they're all trying to do the same thing,
- 25:45:53which is learn some sort of relationship
- 25:45:55between the input data and the and the
- 25:45:57label, right? That's really what they're
- 25:45:59trying to do because they're all
- 25:46:00supervised. They're they have those
- 25:46:02labels, trying to build some
- 25:46:04relationship there. Um, they just do it
- 25:46:07differently.
- 25:46:09And what we're going to study is the
- 25:46:11pros and cons of a lot of these models,
- 25:46:13like when would I use one of them, when
- 25:46:14would I use another. Um, so we'll try to
- 25:46:17talk about that as we go along. Um, but
- 25:46:20they're all trying to learn some
- 25:46:23relationship between the input features
- 25:46:25and the output, right? So you have to
- 25:46:27keep that in mind. They're trying to
- 25:46:29model that relationship. They just do it
- 25:46:31in different ways. Okay? So as we go
- 25:46:34along and learn about new models, um, we
- 25:46:37will learn the details. will learn the
- 25:46:38ins and outs um and those pros and cons,
- 25:46:42but they're no matter what, they're all
- 25:46:44trying to uh learn that relationship,
- 25:46:48right? And be able to make predictions
- 25:46:50on new data.
- 25:46:53Okay,
- 25:46:55so here's a list of models that we will
- 25:46:58cover and work on throughout the uh the
- 25:47:02sessions that we have. um we're not
- 25:47:04going to do them all in one one sitting,
- 25:47:07but um the first one that we're going to
- 25:47:09start with and that we'll cover today is
- 25:47:11going to be linear regression.
- 25:47:14So we will cover linear regression and
- 25:47:16then we'll cover the rest of these guys
- 25:47:18mostly in the context of uh
- 25:47:21classification.
- 25:47:23So, um, what's interesting is some of
- 25:47:26these guys can actually be used for both
- 25:47:28regression and classification as long as
- 25:47:30you make, um, certain adjustments to
- 25:47:33them. They have variations that can be
- 25:47:36used to do classification and regression
- 25:47:38is very interesting. Um but we're going
- 25:47:42to start with linear regression today
- 25:47:45and then work our way through the rest
- 25:47:47of these models when we do um we're
- 25:47:49going to do a separate lesson four on
- 25:47:51classification. So these all these guys
- 25:47:53will come from lesson four.
- 25:47:59Um and then uh we will do this guy in
- 25:48:03lesson three in the 3.2 notebook. We'll
- 25:48:06do all about linear regression.
- 25:48:10Yeah, I so logistic regression is a
- 25:48:12classification um which is kind of
- 25:48:15strange that its name is regression but
- 25:48:18it's doing a classification but the the
- 25:48:20reason is that the logistic regression
- 25:48:23um computes a probability. So it does a
- 25:48:26regression to predict a number but that
- 25:48:29number is actually a probability. So it
- 25:48:31it produces a result that's between it
- 25:48:34produces a probability that's between um
- 25:48:38obviously uh zero and one.
- 25:48:43So it uh and then we take that
- 25:48:45probability and we turn it into a
- 25:48:47category
- 25:48:49like a spam not spam fraud not fraud.
- 25:48:52Um but so so logistic regression is kind
- 25:48:54of special. It's sort of like a
- 25:48:57regression but it's predicting a very
- 25:48:58specific type of value which is a
- 25:49:00probability. So for for that reason it's
- 25:49:03a classification uh algorithm primarily.
- 25:49:12So we'll study that one in lesson four.
- 25:49:15Uh but yeah, that's that's why it's
- 25:49:17under that kind of umbrella of
- 25:49:19classification is because it's it's
- 25:49:21producing a probability as its main
- 25:49:22output which we can then turn into a
- 25:49:26category as long as we interpret that
- 25:49:28probability as um in the right way uh
- 25:49:32like the probability of spam,
- 25:49:34probability of not spam.
- 25:49:40Okay.
- 25:49:44Okay. So, let me focus on um
- 25:49:47let me focus on linear regression. I'm
- 25:49:49not going to go through all of these
- 25:49:50other use cases because we haven't
- 25:49:52learned these models yet. Um so, I don't
- 25:49:56think they're good. Uh I don't think
- 25:49:59it's good to read about them yet until
- 25:50:01we've covered them. So, once we cover
- 25:50:04them in lesson four, I'll come back and
- 25:50:06describe these examples to you guys and
- 25:50:08we'll see why it makes sense. But I
- 25:50:10think for a linear regression um which
- 25:50:12is what we'll cover next, let me talk
- 25:50:14about that example. So a prototypical
- 25:50:16example would be like predicting the
- 25:50:18house prices that we've seen in that
- 25:50:20house price data set.
- 25:50:22So um if we wanted to uh if we wanted to
- 25:50:27predict um if we wanted to estimate the
- 25:50:30market value of a house so the price
- 25:50:35um we could do that by using the
- 25:50:38features such as number of bedrooms,
- 25:50:40square footage, location, age of the
- 25:50:42property. Um and you know then when a
- 25:50:46new when a new house comes on the market
- 25:50:48we could estimate what the price should
- 25:50:50be based on those features. So linear
- 25:50:54regression is a good one to predict the
- 25:50:56price like a housing price. Um and we'll
- 25:50:59actually practice that in the next uh
- 25:51:02notebook.
- 25:51:05So we'll we'll uh and then all these
- 25:51:08other now there's descriptions of these
- 25:51:10other models but again we haven't
- 25:51:11covered these guys yet. So I don't want
- 25:51:13to really go through those until we get
- 25:51:15to those models. So we get to those I'll
- 25:51:17come back and mention the example.
- 25:51:20Uh can K andN be used for clustering?
- 25:51:23No. So um the clustering model is going
- 25:51:27to be different. It's going to be uh K
- 25:51:29means
- 25:51:31K means that's the primary clustering
- 25:51:33model. Not K nearest neighbors. K
- 25:51:36nearest neighbors is used for uh it can
- 25:51:39be used for regression. It can be used
- 25:51:40for classification.
- 25:51:44So we'll we'll talk about K andN which
- 25:51:46is the K nearest neighbors in lesson
- 25:51:48four.
- 25:51:51It sounds really similar. Yeah, it
- 25:51:54sounds really similar but K means is a
- 25:51:56clustering algorithm that's that's
- 25:51:57slightly different
- 25:51:59different uh there's no labels used at
- 25:52:02all. This K nearest neighbors is a is a
- 25:52:06supervised learning algorithm. It uses
- 25:52:08uh labels.
- 25:52:18Good. Any any other questions so far?
- 25:52:36Okay.
- 25:52:38So that being said, let's move on to the
- 25:52:413.2 notebook.
- 25:52:44Let's move on to that which will be our
- 25:52:47um first discussion around uh
- 25:52:51regression. So going into supervised
- 25:52:54learning and regression. Give you guys a
- 25:52:56moment to pull up this notebook.
- 25:52:59But yeah, you want to pull up the 3.2.
- 25:53:01We'll do this one next. So we'll focus
- 25:53:04in. And so our plan is to do regression
- 25:53:06first and then we'll talk about
- 25:53:08classification in lesson four
- 25:53:15which we will cover all those other
- 25:53:17models which you you could use for
- 25:53:20classification uh on that list but then
- 25:53:22we're going to talk about linear
- 25:53:23regression uh first.
- 25:53:33All right. So, we have a a big agenda.
- 25:53:36This is a big notebook um to go through
- 25:53:39a lot of material here surrounding
- 25:53:42regression. So, we're we're going to
- 25:53:44start with linear regression and see um
- 25:53:47how we actually perform it, what that
- 25:53:50model is doing. Um which we've kind of
- 25:53:53seen the idea of it a little bit
- 25:53:55already, so it should be somewhat
- 25:53:56familiar. Um and then we'll talk about
- 25:53:59how to adapt that linear regression idea
- 25:54:02to um nonlinear what's called nonlinear
- 25:54:06regression which is going to be using
- 25:54:08like polomial
- 25:54:10uh features. We'll talk about how to do
- 25:54:11that. Um and then a big big big topic
- 25:54:15for us is going to be evaluating the
- 25:54:17model. So it'll be it'll be quite easy
- 25:54:19to actually build it. building the model
- 25:54:22will be really easy but evaluating and
- 25:54:25interpreting that will be uh a lot of
- 25:54:30interesting work there um because we
- 25:54:33want to know what the performance of
- 25:54:34that model is once we have it built
- 25:54:36right we want to know how good of a
- 25:54:38model is it is it worth using or do we
- 25:54:41need to retrain it or get new data or
- 25:54:43change the model up to talk about that
- 25:54:46um how do you determine what to do based
- 25:54:48on that performance
- 25:54:50um and then we'll talk about here um a
- 25:54:53couple things. We may not get to this
- 25:54:55today, but regularization
- 25:54:57which is used to boost the performance
- 25:54:59uh in certain situations um whenever the
- 25:55:03model is kind of uh performing um poorly
- 25:55:07against test data even though it
- 25:55:08performs pretty well on training data.
- 25:55:11In that scenario, you can use offshoots
- 25:55:14of linear regression that do some uh
- 25:55:16what's called regularization. We'll talk
- 25:55:18about that.
- 25:55:20Um and then we'll talk about
- 25:55:21hyperparameter tuning uh generally as a
- 25:55:24strategy which is something you
- 25:55:26generally do want to do when you're
- 25:55:28training machine learning models. Um so
- 25:55:31again these two we may not get to today
- 25:55:35but um quite a quite a lot to get to be
- 25:55:38prior to that mainly centered around
- 25:55:40evaluation and building linear
- 25:55:43regression.
- 25:55:45Okay. So pretty cool. we'll get to our
- 25:55:47first kind of model here. This linear
- 25:55:49regression
- 25:55:51to start with.
- 25:55:55Okay,
- 25:55:56so let's start with uh linear regression
- 25:55:59here. Um, and really what linear
- 25:56:04regression is attempting to do and I
- 25:56:08want to show you this in this picture is
- 25:56:11draw this line sometimes what is known
- 25:56:14as the line of best fit. So this is our
- 25:56:17model that kind of goes through the data
- 25:56:20and it's generally a good predictor
- 25:56:25um because if you give me um features uh
- 25:56:30if you give me new features and let's
- 25:56:33say they are let's say you give me a
- 25:56:35feature that's right here.
- 25:56:38So you say, okay, I have a feature
- 25:56:39that's this value on the x- axis. Then I
- 25:56:43know all I have to do is plug that into
- 25:56:45my line equation, and I will generate a
- 25:56:49a value that's like right here.
- 25:56:52Okay, that's pretty that's on that line
- 25:56:54at that input. And that's going to be my
- 25:56:57prediction for what the output variable
- 25:56:59should be. It's just going to be
- 25:57:00something on that line. And what you can
- 25:57:04see is this line is a decent estimate
- 25:57:07for this data because it slices through
- 25:57:10this pretty evenly. So it's a good guess
- 25:57:13as to what the output should be given
- 25:57:16any one of these inputs. It's a it's a
- 25:57:18good estimator this line. And so our
- 25:57:21goal building a linear regression is to
- 25:57:23kind of build the equation of this line.
- 25:57:26So we want this equation.
- 25:57:30Equation of this line
- 25:57:34is going to be our model.
- 25:57:41Yes, it's going to look just like that.
- 25:57:43MX plus B or yeah, MX plus C. It's going
- 25:57:45to look exactly like that. uh except
- 25:57:48that it's going to be more than just MX
- 25:57:52because we have um generally more than
- 25:57:56one feature. So you think of X as a
- 25:57:57feature um it will be more than just MX.
- 25:58:00It will generally be like uh it'll
- 25:58:03generally look like this
- 25:58:10and then plus maybe some bias here plus
- 25:58:14uh an intercept. Yeah, it'll generally
- 25:58:16look like that. So, yeah, you're exactly
- 25:58:18right. MX plusb is the right idea.
- 25:58:21Exactly right.
- 25:58:24It'll generally look like that.
- 25:58:28Nonlinear. It can be adapted to
- 25:58:30nonlinear. Yeah. If we transform, we're
- 25:58:33going to talk about that. If we
- 25:58:34transform all of our features in a
- 25:58:36nonlinear way, um we can apply linear
- 25:58:39regression to it. Yes. And and that
- 25:58:42would be a nonlinear regression. So yes,
- 25:58:45we can do nonlinear things too.
- 25:58:48We'll talk about that.
- 25:58:58Okay. So linear regression again is the
- 25:59:01art or science I should say not really
- 25:59:04art but it is an exact science of
- 25:59:07finding the equation of this line that
- 25:59:09fits through this data. Um now why one
- 25:59:12thing you should be thinking about is
- 25:59:14why is this line a good predictor and
- 25:59:18the argument is that if you take a look
- 25:59:20at this distance from these blue points
- 25:59:22so let's say these blue points are our
- 25:59:24actual data points this line is going to
- 25:59:28be found such that it minimizes this
- 25:59:32distance
- 25:59:34from the points to actually I should
- 25:59:37draw it this way from the points to the
- 25:59:39line.
- 25:59:41So, we want this distance to be um
- 25:59:44actually I should draw it that way, this
- 25:59:47way. We want this distance to be kind of
- 25:59:50at a minimum. So, it would be bad to
- 25:59:52draw a line all the way out here because
- 25:59:54then that's a lot of distance, right?
- 25:59:56So, and that would be a lot of error um
- 25:59:59contributed from not being able to
- 26:00:01predict those points in our data set
- 26:00:02very well. Um which is our training
- 26:00:05data. That's why we have labels, right?
- 26:00:07that that guide us in building this
- 26:00:09line. Um so our goal is to build that
- 26:00:12line especially so that this error or
- 26:00:16this distance can be as minimum as
- 26:00:20possible. Right? Which are all these
- 26:00:22distances from these points to the line.
- 26:00:25We want those to be as minimum as
- 26:00:28possible. So our goal is to find this
- 26:00:30equation.
- 26:00:32So we're going to build a model that's
- 26:00:34going to find this equation.
- 26:00:39of the line
- 26:00:41um such that our error
- 26:00:46is minimal.
- 26:00:50And what is the error? The error is the
- 26:00:52distance
- 26:00:56of our data points
- 26:01:02to
- 26:01:06to the line that we build. So
- 26:01:08essentially what we'll do in order to
- 26:01:10train this will be to adjust the
- 26:01:13parameters or the or in that like I
- 26:01:15think it's really good you brought up
- 26:01:16the MX plus C. Basically the M and the C
- 26:01:19will adjust. So we adjust those
- 26:01:22accordingly to make this distance as
- 26:01:25small as possible.
- 26:01:28Okay to minimize that distance as much
- 26:01:30as possible.
- 26:01:39Okay.
- 26:01:41So, um where is regression used? We've
- 26:01:45already seen some examples. Here's some
- 26:01:47more uh advertising like predicting
- 26:01:50sales, predicting um oil and uh oil
- 26:01:55production and demand. Those are like
- 26:01:57forecast those are regression problems.
- 26:02:00Um retail like demand forecasting for
- 26:02:02inventory. Um healthc care predicting um
- 26:02:07uh the levels of certain um uh blood
- 26:02:12markers or you know something like that.
- 26:02:14Um real estate predicting prices based
- 26:02:18on those uh talked about like square
- 26:02:20footage, bedrooms, bathrooms, those
- 26:02:22things. So regression is used again
- 26:02:24whenever we want to predict a number a
- 26:02:26numerical output um that's a regression
- 26:02:29problem.
- 26:02:35So this kind of regression we're talking
- 26:02:37about here is generally
- 26:02:41um known as uh a when that equation is
- 26:02:45linear that is known as a linear
- 26:02:48regression. And so go back to that
- 26:02:49picture when we have a when that
- 26:02:52equation of the line that we find is a
- 26:02:55linear equation meaning that it is
- 26:02:58exactly the form I've been telling you.
- 26:03:00So it's it's something like um weight
- 26:03:04time feature
- 26:03:06plus weight time feature
- 26:03:09plus weight time feature
- 26:03:13and then maybe some intercept um term
- 26:03:17like some some bias term there. Um this
- 26:03:21is a linear equation because all of the
- 26:03:24features are to the single power. So
- 26:03:27it's a linear power and this is a linear
- 26:03:29combination of features with with those
- 26:03:32different weights. So this is a linear
- 26:03:35model
- 26:03:36because it is uh it's what in math we
- 26:03:40would call this a linear equation right
- 26:03:43everything is to the first power. It
- 26:03:45resembles mx plus b. It is a linear
- 26:03:48equation or a linear model. Um so when
- 26:03:53we talk about linear regression that is
- 26:03:56a regression model so we're predicting
- 26:03:58some continuous target that assumes we
- 26:04:01are model our model is formed from this
- 26:04:05kind of equation a linear equation.
- 26:04:08So this is going to be our our model for
- 26:04:12a linear
- 26:04:14uh regression.
- 26:04:17Okay.
- 26:04:20And so when you when you train a linear
- 26:04:23regression, your goal is to learn these
- 26:04:25weights so that you can plug in um you
- 26:04:30can plug in any one of your uh input
- 26:04:32features and you um can generate a
- 26:04:36prediction. You can which is going to be
- 26:04:38something on that line, right? It's
- 26:04:40going to be a value that's sitting here
- 26:04:42on this line.
- 26:04:44We put in all of our features and we end
- 26:04:46up there somewhere on that line.
- 26:04:50This output.
- 26:04:54Okay.
- 26:05:05Okay. Let me pause there. Any questions
- 26:05:07on the linear model here or why it's
- 26:05:12called linear regression?
- 26:05:28Okay. And by the way in these notes um
- 26:05:31this bullet point here where it says it
- 26:05:32uses the least squares criterion to
- 26:05:34estimate the coefficients that is
- 26:05:36exactly what I said earlier with the
- 26:05:38distance. So the distance is based on
- 26:05:40the square
- 26:05:42of this this quantity like how far away
- 26:05:45you are from the line is based on this
- 26:05:47square distance here and here and here
- 26:05:51and here. So what we're trying to do is
- 26:05:54find the least distance or least squares
- 26:05:58which is that minimum distance. So
- 26:06:00that's how we find all of these weights
- 26:06:03is from minimize. We basically tune them
- 26:06:06enough using our labels. So here's our
- 26:06:09label which is the y. We basically plug
- 26:06:12in our data and tune those enough to
- 26:06:14minimize the error. It's it's a it's an
- 26:06:16optimization problem,
- 26:06:19right? We we're trying to find the
- 26:06:21minimum of this quantity which is that
- 26:06:26best fit line.
- 26:06:40Okay.
- 26:06:41So we have linear regression
- 26:06:44um and we can do a simple linear
- 26:06:49regression that only has one feature. So
- 26:06:51if it only has one feature that's
- 26:06:53exactly the so if there's only one input
- 26:06:56feature sometimes that is known as um
- 26:06:59simple regression or simple linear
- 26:07:01regression and there's basically there's
- 26:07:03only one feature. So one independent
- 26:07:05variable is the feature.
- 26:07:08There's only one feature. And so this
- 26:07:11equation resembles the
- 26:07:16exact equation that you guys just put in
- 26:07:18there, which is um mx plus b,
- 26:07:22right? It resembles exactly that. Um
- 26:07:25we're just using different symbols for
- 26:07:27those like beta beta 0 and beta 1. But
- 26:07:30um basically exactly that simple line is
- 26:07:34only one feature. So, and that's because
- 26:07:37that line is going to um that line is
- 26:07:42going to be generated uh according to
- 26:07:44that equation. So, here's kind of what
- 26:07:46it looks like.
- 26:07:48This is the best fit line through all of
- 26:07:50these blue dots. This is something we're
- 26:07:52going to be able to build. We're going
- 26:07:54to be able to build that equation um
- 26:07:57pretty easily in scikitlearn.
- 26:08:00So, we'll be able to find that um and it
- 26:08:03won't be too hard. So this line will be
- 26:08:06um y = beta 0 plus beta 1. So some
- 26:08:11weight beta 1 times the only feature we
- 26:08:15have x1.
- 26:08:17Okay. So in this case um we would be
- 26:08:21predicting sales. So sales would be the
- 26:08:24value basically the label that we're
- 26:08:26trying to predict and the feature that
- 26:08:28we're putting in is uh I think it's the
- 26:08:34number of TV expenses. Yep. TV expenses
- 26:08:38which is on the x- axis. So there's one
- 26:08:40feature which is um TV expense.
- 26:08:48So um on this graph this would be this
- 26:08:52would be our model.
- 26:08:54Okay that would be our model. We only
- 26:08:56have one feature and we have um these
- 26:09:00two weights. We have an intercept B 0
- 26:09:03and or beta 0 and then a one weight
- 26:09:06which gets applied to that one feature
- 26:09:08beta 1. And so our model would have
- 26:09:11certain value for beta 0 and a certain
- 26:09:13value for beta 1. That's what get that's
- 26:09:16these guys get learned
- 26:09:21learned during
- 26:09:24model
- 26:09:27training.
- 26:09:33Okay. So those are what get learned
- 26:09:36during our model training and they get
- 26:09:38learned by a a a least what's called a
- 26:09:41lease squares algorithm that is trying
- 26:09:43to minimize that distance. It tries to
- 26:09:45tweak beta 0 beta 1 to minimize this
- 26:09:48distance of this line
- 26:09:51um this line
- 26:09:53to all of these points
- 26:09:58trying to minimize this.
- 26:10:02So imagine taking a line and kind of
- 26:10:04moving it around and turning its its
- 26:10:08slope, its angle um to try to find that
- 26:10:11best fit,
- 26:10:13which reduces that error the most.
- 26:10:16Right? That's kind of what we're doing.
- 26:10:40Uh can I explain? Yeah. So uh sales is
- 26:10:44in dollars and and TV expense
- 26:10:48um
- 26:10:50uh
- 26:10:52TV actually I think it's the other way
- 26:10:54around. I think the sales is actually a
- 26:10:56quantity. So I this is number of sales
- 26:10:59that we have and TV expense is um I
- 26:11:03think I think it's in dollars. So how
- 26:11:05much money how much expense um did we
- 26:11:08put into the into the product and then
- 26:11:11this is how many sales did we have of
- 26:11:13that product.
- 26:11:17So I think it's the other way around
- 26:11:24but what this what this graph is showing
- 26:11:27is the blue points are our actual data
- 26:11:31points. Okay. So so we have a collection
- 26:11:34like we have a data frame that has so
- 26:11:37imagine we had a data frame that has the
- 26:11:39uh true values.
- 26:11:42So it has the um TV expenses.
- 26:11:46Um it has points that are like one. So
- 26:11:50it has points that are like 120 and then
- 26:11:53the sale sales could be like 700
- 26:11:57700 units, let's say. And then it has um
- 26:12:00so this is just our data set, right?
- 26:12:01This would be like in a data frame that
- 26:12:03we have. And then we had ones that were
- 26:12:06um 50 and then this could be um this
- 26:12:10could be 400 let's say and on and on and
- 26:12:13on right so this is our data and this
- 26:12:16data is plotted in the blue so these are
- 26:12:19these blue points here
- 26:12:21right so these are the blue points here
- 26:12:24and the red points are is our model so
- 26:12:27we built a linear regression model um
- 26:12:31where we are putting in some values
- 26:12:33we're putting in some fake x values
- 26:12:36here and generating some predictions
- 26:12:38which is this line,
- 26:12:41this linear uh regression line, right?
- 26:12:44And that line is derived from this data,
- 26:12:48right? It gets learned from this
- 26:12:50supervised uh examples.
- 26:12:55Does that make sense?
- 26:12:57That line is derived from the data. it's
- 26:13:00actually um learned from like the line
- 26:13:03of best fit is learned from that data
- 26:13:10and the actual data is in the blue.
- 26:13:14So you can see we're trying to build
- 26:13:15this such that this distance is kind of
- 26:13:17a minimum
- 26:13:20so it's an optimal fit
- 26:13:24to balance out these distances.
- 26:13:34So it's just plotting. So it's just
- 26:13:35building that relationship between the
- 26:13:37input and output. Like when the when the
- 26:13:39expenses are higher, um we seem to have
- 26:13:42more sales.
- 26:14:01Okay.
- 26:14:10Uh what's perpendicular like the
- 26:14:12distance? This should be this should be
- 26:14:14perpendicular because it's a distance
- 26:14:15here.
- 26:14:19Is that what you mean? Like the distance
- 26:14:20from the real points to the line? Yeah,
- 26:14:22that should be perpendicular
- 26:14:24because it's it's a it's a distance
- 26:14:26formula.
- 26:14:46Okay.
- 26:14:50All right. So more generally now do do
- 26:14:54we usually have one feature? No. So
- 26:14:58generally we expand this to the more
- 26:15:01general case where we have more than one
- 26:15:04feature like what we see in the housing
- 26:15:06data right where we could predict a
- 26:15:08price but we have many different inputs
- 26:15:10like bedrooms, bathrooms, square footage
- 26:15:13etc.
- 26:15:16So more broadly
- 26:15:19instead of simple linear regression we
- 26:15:22have what's known as multiple linear
- 26:15:24linear regression which means we have
- 26:15:26multiple variables or multiple features.
- 26:15:29Um so this is exactly the equation I've
- 26:15:32been talking about. Um so we just extend
- 26:15:36that that one into many features. So
- 26:15:39which is this case and then a intercept
- 26:15:42term which is uh um there as sometimes
- 26:15:47known as the bias. Um
- 26:15:50but this is the intercept term to kind
- 26:15:52of orient the line to start out in the
- 26:15:54right place. Um and uh but this is the
- 26:16:00um this is the equation that we would be
- 26:16:03building the model. This is our model
- 26:16:05essentially, right? This is the equation
- 26:16:06we would be learning.
- 26:16:09Intercept is like a constant. Yeah. So
- 26:16:11if if all of the features were zero, um
- 26:16:14this is what our our data would be. This
- 26:16:16is what our result would be. If
- 26:16:18basically if this was zero, this was
- 26:16:19zero, this was zero, it would reduce to
- 26:16:22this as the prediction. Yeah. It's like
- 26:16:25a constant. Yes.
- 26:16:33So in in geometry, the intercept's
- 26:16:36actually really important because it it
- 26:16:37orients where your line should start. So
- 26:16:39it orients like so so these values are
- 26:16:42kind of like the slope. They orient the
- 26:16:44tilt of it. Like should it be tilted
- 26:16:47like this or should it be more sloped?
- 26:16:49But the intercept orients where it
- 26:16:52should start like vertically like should
- 26:16:54it start all the way up here? Should it
- 26:16:56start more down here?
- 26:16:58Um, that's what the intercept kind of
- 26:17:01tells us.
- 26:17:11Okay, so this is the situation. This is
- 26:17:14going to be our linear regression model
- 26:17:15that we will be building most of the
- 26:17:17time because we will have again these
- 26:17:19are all going to be features.
- 26:17:22So this is some feature the X this is
- 26:17:25some feature this is some feature
- 26:17:29X1 etc. These are all features and what
- 26:17:33gets learned during the training are
- 26:17:35these coefficients. So all of these
- 26:17:37coefficients including the beta 0ero um
- 26:17:41will get learned. So these will get
- 26:17:43learned
- 26:17:45um from our data right they get learned
- 26:17:48they will be trained from our data um in
- 26:17:51order and and how do they get trained
- 26:17:53it's from reducing that distance we try
- 26:17:56to get that line of best fit by tweaking
- 26:17:59those betas enough to uh until we reach
- 26:18:02a minimum distance but there's there's
- 26:18:04an algorithm behind that um that that
- 26:18:07scikitlearn will run for us to find that
- 26:18:10best fit Um, so we don't need to do that
- 26:18:13manually, but that's that's the process
- 26:18:15is basically tweaking those weights to
- 26:18:18end up with that line of best fit. So in
- 26:18:21higher dimensions, instead of a line,
- 26:18:23you get more of what's called a plane
- 26:18:26here. Um, which kind of looks like this.
- 26:18:28So the best fit is actually this plane
- 26:18:32where all um, it kind of dissects all
- 26:18:34these points just like that um, in
- 26:18:37higher dimensions. So this is uh instead
- 26:18:40of a line you get this in in three
- 26:18:42dimensions you get this plane like this
- 26:18:45but it's still it's like a line of best
- 26:18:47it's just a more general line of best
- 26:18:49fit. It's still the same idea. Um we're
- 26:18:52still trying to um come up with the best
- 26:18:55coefficients to minimize that distance
- 26:18:57from our from our points to the line.
- 26:19:01Although in higher dimensions it's no
- 26:19:03longer a line. It's more like a plane
- 26:19:04like this. So you're trying to minimize
- 26:19:06this distance from here down to the
- 26:19:08plane
- 26:19:09here up to the plane
- 26:19:12in higher dimensions. So I want you to
- 26:19:14keep in mind what we're trying to do
- 26:19:17before we go into the code because the
- 26:19:19code's going to make it seem really
- 26:19:20really simple and that's because
- 26:19:22scikitlearn is great and that's what it
- 26:19:24does.
- 26:19:26But we should realize that there's
- 26:19:27something really complex going on which
- 26:19:29is again finding the best value of these
- 26:19:33weights
- 26:19:35that minimizes the distance of this line
- 26:19:39to the data points that we have. So
- 26:19:42there's an algorithm there that will
- 26:19:44keep trying to make adjustments to this
- 26:19:47based on those distances. So it's going
- 26:19:50to use those distances as a guide to
- 26:19:52kind of tweak them to find the one that
- 26:19:56results in the lowest amount of
- 26:19:58distance. So we keep making tweaks, keep
- 26:19:59making tweaks, keep making tweaks and
- 26:20:02eventually we try to find we converge to
- 26:20:05the set of weights that gives us that
- 26:20:06best fitting line. Um and and there's an
- 26:20:10algorithm there that occurs. Now luckily
- 26:20:13that gets abstracted for us a bit behind
- 26:20:16um scikitlearn
- 26:20:18um finding that best fit. So there'll be
- 26:20:21a function that we use in scikitlearn
- 26:20:25when we build the model that will go
- 26:20:26ahead and find the best weights for us
- 26:20:30and that's then we now have our optimal
- 26:20:33model right that then we can just plug
- 26:20:35in different values of these features
- 26:20:38and generate a prediction which is going
- 26:20:41to be this uh result right so so that's
- 26:20:44what we're ultimately trying to do is uh
- 26:20:48train the model which will uh find all
- 26:20:50those optimal weights and then uh we can
- 26:20:54predict with it which would be plugging
- 26:20:56in different feature values to to
- 26:20:58generate a prediction.
- 26:21:02Okay,
- 26:21:03so let's see how that happens. It's
- 26:21:05actually going to be super easy um with
- 26:21:07scikitlearn.
- 26:21:09So uh in this scenario we have um we're
- 26:21:13going to import our pandas because we're
- 26:21:15going to load our data from that. Um, so
- 26:21:19of course we need some data to work
- 26:21:20with. So we're going to load this uh
- 26:21:22CSV.
- 26:21:23Um, I
- 26:21:26uh so I was not actually able to find
- 26:21:30this CSV for this example, but I mean
- 26:21:32that's okay because we'll do some we'll
- 26:21:33do other examples where we'll work with
- 26:21:35the data. If you happen to have it, um,
- 26:21:38great. I didn't see it in in my files.
- 26:21:42So just have to take the word for it
- 26:21:43that these are the this is that TV and
- 26:21:46sales columns here um from this data
- 26:21:50set.
- 26:21:52Okay. Um as an example. So um just to
- 26:21:57see how it's fit um what we're going to
- 26:22:01do and this is going to be a very
- 26:22:04standard process for us for building a
- 26:22:07model. These steps are going to be very
- 26:22:09very standard for us which is going to
- 26:22:11be first of all splitting the features
- 26:22:15away from the label. That's the first
- 26:22:18step that we always will take. So if you
- 26:22:20take a look at this code, it's taking
- 26:22:23all rows but only the first column.
- 26:22:28Okay, so it's extracting all the
- 26:22:30features from the data frame um which
- 26:22:33happen to be which is just the first the
- 26:22:35first column uh which is the TV uh
- 26:22:39column right just that column there and
- 26:22:42our target variable which is our label.
- 26:22:46So our target variable aka the label um
- 26:22:50is the second column, right? It's that
- 26:22:54that sales column.
- 26:22:57Um and so our first step here, let me
- 26:23:01call that out here. First step is to
- 26:23:05always split apart
- 26:23:08features from labels.
- 26:23:12Okay, so we put all those features into
- 26:23:14a data frame called X and we have all of
- 26:23:17our labels into technically a series but
- 26:23:20uh sort of like a data frame, right? Um
- 26:23:24called Y, which is just the um which is
- 26:23:28just the uh uh labels. So that's just
- 26:23:32the TV values. Um now you're going to
- 26:23:35see why we do that. It's because we need
- 26:23:39um our our features and labels split
- 26:23:42apart to put them into the model
- 26:23:44building function. It expects our
- 26:23:48independent variables or our features to
- 26:23:50be separated from our answers or our
- 26:23:54labels that guide the model building.
- 26:23:57That's the first thing you got to do is
- 26:23:58separate those.
- 26:24:01Okay, so this code will separate those
- 26:24:03out into a capital X and a lowercase Y.
- 26:24:06And that's actually pretty industry
- 26:24:08standard notation. Whenever you split
- 26:24:10apart all your features, usually you put
- 26:24:13them into a data frame called capital X
- 26:24:15and then you have a lowercase Y to
- 26:24:18represent your labels. That's actually
- 26:24:20pretty standard.
- 26:24:23So it's pretty standard that um X
- 26:24:26represents
- 26:24:29features
- 26:24:31and
- 26:24:32Y represents labels
- 26:24:37label column
- 26:24:39whatever our label column is in this
- 26:24:41case it is the sales because we're going
- 26:24:44to be predicting sales
- 26:24:49using the TV column the TV quant expense
- 26:24:53quantity.
- 26:25:00Yeah. So what it so the assignment is
- 26:25:03that we are um the assignment is that we
- 26:25:08are
- 26:25:09uh we are um splitting apart our data.
- 26:25:14So that when we first read in the data
- 26:25:17um it is a data frame right that has two
- 26:25:20columns TV and sales.
- 26:25:24Oh perfect thank you Tim. I will I will
- 26:25:28go ahead and so if we look at this data
- 26:25:33it only has those two columns right it
- 26:25:36only has those two columns. Okay. So
- 26:25:39what we're doing with this is we are
- 26:25:42splitting apart
- 26:25:44our our independent variable our
- 26:25:47features. So this this x will contain
- 26:25:52our features
- 26:25:56and y will contain
- 26:25:59our label.
- 26:26:03Does that make sense? We're splitting
- 26:26:04this data apart. So, we're only grabbing
- 26:26:06that first column here to be our
- 26:26:09features. And then we're we're grabbing
- 26:26:12the second column, which is the sales,
- 26:26:13because we're going to predict the
- 26:26:15sales. This is our label. We're going to
- 26:26:17we're going to build a model to predict
- 26:26:19the sales given the TV input, TV expense
- 26:26:23input. So, the first thing we have to do
- 26:26:26is split apart the features and the
- 26:26:28label.
- 26:26:31Okay, that's the first step we usually
- 26:26:33will take. And the reason we have to do
- 26:26:35that um just to reiterate, the reason we
- 26:26:39have to do that is because our model
- 26:26:42will expect our our data features to be
- 26:26:45separate from the label. We will pass
- 26:26:47those in separately.
- 26:26:50X is TV. It's the first column
- 26:26:54because we're using eyeling.
- 26:27:09We are predicting the sales given the TV
- 26:27:12expense value.
- 26:27:17Yeah. which is why we split it into so
- 26:27:20this is the second column right the
- 26:27:22index one column
- 26:27:28uh you just put in read CSV and pass in
- 26:27:30the URL so you could so exactly the code
- 26:27:34that was up earlier from Tim
- 26:27:37um you just do this
- 26:27:41and then data equals ed read CSV URL
- 26:27:49So we split our data into X and Y here.
- 26:27:53All right. Now, one other step that
- 26:27:56we're going to take that's a very very
- 26:27:58critical step and you're going to we're
- 26:27:59going to see this step over and over and
- 26:28:03over and over again. So splitting apart
- 26:28:05into X and Y will become we'll do that
- 26:28:08over and over and over and over again.
- 26:28:10Not only that, but doing this next step,
- 26:28:14which is what's called a train test
- 26:28:17split. Now, let me show you what the
- 26:28:19train test split does. It takes our data
- 26:28:24and it's going to split apart our data
- 26:28:27that we have, our X and our Y data. It's
- 26:28:30going to split it apart into a
- 26:28:32percentage that will be used to train
- 26:28:34the data
- 26:28:37and then a percentage that will be used
- 26:28:39to test. Now, why would we want to do
- 26:28:42that? It's mainly so we can do
- 26:28:45evaluation. So, we build the model over
- 26:28:48here and then we test it on data that
- 26:28:51has not seen before. So, we reserve a
- 26:28:55percentage of the data to be used for
- 26:28:57test. Usually this this data is um
- 26:29:01somewhere between uh 20 to 30%.
- 26:29:06So somewhere between 20 to 30% of the
- 26:29:09original data. So that means the
- 26:29:12majority of it is used for training. So
- 26:29:14the majority of the of that X and Y over
- 26:29:17here is going to be between 70 to 80%.
- 26:29:23will generally be used for for uh for
- 26:29:27training. Okay. So somewhere between 20
- 26:29:30to 30 the industry standard is some
- 26:29:32anywhere in between there. Um a lot of
- 26:29:35people like to use 30%, some people like
- 26:29:37to use 20%. Um anything in that range is
- 26:29:40acceptable. Um we will I think we
- 26:29:43generally will favor like 30%.
- 26:29:46um to be used for testing. But um the
- 26:29:50the point is we don't we don't want to
- 26:29:53mix those together. We want those to be
- 26:29:55separated out so that we can have a fair
- 26:30:00evaluation, right? We want to train our
- 26:30:02data on this train our model on this
- 26:30:05data and then see how well it performs
- 26:30:08on this data that it has never seen
- 26:30:10before.
- 26:30:12Right? So in order to have data it's
- 26:30:14never seen before, we're going to take
- 26:30:16our x and our y and we're going to split
- 26:30:17it using this function called train test
- 26:30:21split that will do this kind of
- 26:30:24splitting for us. Okay, so scikitlearn
- 26:30:27has a function called train test split
- 26:30:30that will go ahead and we're going to
- 26:30:32pass our x and our y and we'll pass in a
- 26:30:34percentage like 30% that we want to
- 26:30:38split out into a test set and then the
- 26:30:40remainder of that the 70% will be used
- 26:30:44for training the model.
- 26:30:47Okay.
- 26:30:50So what we're going to get let me redraw
- 26:30:53that. So what we're going to get out of
- 26:30:55this for the train test split is we're
- 26:30:57going to we're going to have an X and a
- 26:30:59Y per
- 26:31:02training and test. So we're going to get
- 26:31:05now we're going to get an X train
- 26:31:11and a Y train.
- 26:31:15So we're going to get training features
- 26:31:17and training labels. And then we're
- 26:31:19going to get test features
- 26:31:24to plug into our model and and test
- 26:31:28answers or test labels
- 26:31:32to do evaluation because what we should
- 26:31:34be able to do is build the model over
- 26:31:36here and then apply the model on this
- 26:31:38data. Meaning we can take these features
- 26:31:41and plug it into our model and then see
- 26:31:44what answers we get and compare those
- 26:31:47answers to this testing data. Right? We
- 26:31:50should be able to do that to generate an
- 26:31:52evaluation.
- 26:31:56Okay. Now you may be wondering why do we
- 26:31:59do any of that? What's the purpose of
- 26:32:01that?
- 26:32:03Evaluating it on this test data gives us
- 26:32:06a good sense of will our model
- 26:32:11generalize to new examples. Right? If it
- 26:32:15performs pretty well on this data,
- 26:32:18that's a good signal like when it's
- 26:32:19performing pretty well on data it's
- 26:32:21never seen before, that's a good
- 26:32:24indicator that it's going to perform
- 26:32:26pretty well when we use it on brand new
- 26:32:28examples
- 26:32:30um in the future.
- 26:32:33Right. So that's a that's why we do this
- 26:32:37evaluation on this data that it has not
- 26:32:40seen before. It's going to see this
- 26:32:43training data, right? We're going to
- 26:32:44train the model on that data. But that
- 26:32:47model will never be exposed to this test
- 26:32:49data until we do the evaluation
- 26:32:53and and generate some metrics to see how
- 26:32:56good is this performing
- 26:32:58and does it have a good chance of
- 26:32:59generalizing to never before seen
- 26:33:02examples which is what we want right
- 26:33:04because we're going to use this model in
- 26:33:06the real world. It's going to be being
- 26:33:08used on new examples that it hasn't seen
- 26:33:10before. We want it to perform well. So,
- 26:33:13this is kind of our test, our
- 26:33:15evaluation.
- 26:33:20Okay. Any questions on the We're going
- 26:33:23to do this in a moment. I'll show you
- 26:33:24what it looks like in the code, but any
- 26:33:27conceptually any questions on the train
- 26:33:29test split idea. It's a very very
- 26:33:32important idea that we um basically use
- 26:33:37part of the data to train it and then
- 26:33:38another part of it to evaluate. It's
- 26:33:41very important we do that. By the way,
- 26:33:43this has a term um this in machine
- 26:33:46learning this is called cross
- 26:33:50validation
- 26:33:55because we are using one data set to
- 26:33:58train the model and then we're cross
- 26:34:00over we're crossing that over into
- 26:34:03another data set to validate it which is
- 26:34:06the uh the the testing that.
- 26:34:13So this is called cross validation. Um
- 26:34:15there's actually many ways to do cross
- 26:34:17validation. That's something we'll
- 26:34:18study. This is a very simple way of
- 26:34:20doing cross validation. There's more
- 26:34:21complex ways. You can take your data and
- 26:34:24you can actually divide it into many
- 26:34:26sections
- 26:34:27and basically train it against most of
- 26:34:29these and evaluate it against one at a
- 26:34:32time and then rotate. So that's another
- 26:34:35way to do cross validation. We're going
- 26:34:36to study that. Um but this is the this
- 26:34:40is the simplest way to do it here.
- 26:34:48Okay.
- 26:34:51So let me show you what you get when you
- 26:34:52use train test split. So uh we're going
- 26:34:55to import from sklearn.
- 26:34:58We're uh from the model selection
- 26:35:00module. Now we haven't used this before.
- 26:35:03This is our first time using it. But
- 26:35:04here's our model selection. We're going
- 26:35:07to import this train test split function
- 26:35:10and we're going to use it on our X and Y
- 26:35:13and we're going to set a test size of
- 26:35:1730% which is which is.3. So our test
- 26:35:20size
- 26:35:23is 30%.
- 26:35:25Converted to decimal
- 26:35:29right converted to.3 so that means we're
- 26:35:32reserving 30% for that test set. Um you
- 26:35:36can set a random state. Now that's
- 26:35:37completely optional. Um the random state
- 26:35:42is for reproducibility
- 26:35:50because what the train test split is
- 26:35:51going to do is it's actually going to
- 26:35:53shuffle the data and then split it apart
- 26:35:56into the 7030.
- 26:35:58So um yes, the seed. Exactly. It's like
- 26:36:02a seed. So it's it's saying like when
- 26:36:04you do that shuffling every time I run
- 26:36:06this notebook I'm going to get the same
- 26:36:08result but it's going to be random the
- 26:36:10first it's going to be random but I'm
- 26:36:11going to be able to reproduce that
- 26:36:13randomness with that random state. Yes,
- 26:36:18it is like a seed.
- 26:36:23Uh it's you can choose any number to be
- 26:36:26your your um your random state. It 42
- 26:36:29isn't important. You could choose zero.
- 26:36:31You could choose one. Um, you could
- 26:36:33choose any positive integer. Um, 42 is
- 26:36:37kind of like the uh industry standard.
- 26:36:41It's it's you'd have to look it up why
- 26:36:43it is. Um, apparently 42 is a special
- 26:36:47number. Um,
- 26:36:50in kind of the history of development of
- 26:36:52this stuff, there's nothing really
- 26:36:54special about 42. You could choose a
- 26:36:56random You could choose a random seed to
- 26:36:58be uh zero. That's fine. It it doesn't
- 26:37:01really it doesn't really matter.
- 26:37:05Um you just want you can choose it to be
- 26:37:08uh one, two, three. Um you can choose it
- 26:37:11to be 15. You can choose it to be
- 26:37:13anything you want it to be. It's really
- 26:37:15so that your your shuffling is
- 26:37:17consistent. Every time you run this
- 26:37:19notebook, you get the same shuffle
- 26:37:21result. So I'm always going to get the
- 26:37:23same rows in these splits.
- 26:37:29Hitch. There it is. I knew it was from
- 26:37:31something.
- 26:37:38Yeah. So 42 is kind of like a
- 26:37:42it's it's just used ubiquitously
- 26:37:46uh you know as kind of a um paying
- 26:37:50tribute to the Hitchhiker's Guide to the
- 26:37:52Galaxy, but it's no it's there's nothing
- 26:37:54that special about 42. It doesn't it's
- 26:37:56not going to change our result or
- 26:37:58anything.
- 26:37:59It's just so that this train set split
- 26:38:02is going to shuffle our data and split
- 26:38:04it apart into 7030.
- 26:38:07You just want to set this to something
- 26:38:08so that you get a cons every time we run
- 26:38:11this notebook, we get a consistent
- 26:38:13shuffle.
- 26:38:15And so the data in these sets
- 26:38:18are uh consistent. That's all.
- 26:38:27Okay. But do you guys see how we pass in
- 26:38:30our X and our Y and we generate four we
- 26:38:33generate four different data uh
- 26:38:36quantities here which is we generate
- 26:38:38training features, test features,
- 26:38:41training labels and test labels because
- 26:38:43again we are generating these four
- 26:38:47different we're generating data on these
- 26:38:50two different sets a training set
- 26:38:53and a test set. So we have training
- 26:38:56features, training label,
- 26:39:00and then test features, test label.
- 26:39:04Okay, that's why it's so important to
- 26:39:07split apart our data into the X and the
- 26:39:09Y. We need those split apart in order
- 26:39:12for this part to work.
- 26:39:16So by the way, these two steps we will
- 26:39:18always do for any model we build. We'll
- 26:39:21generally do X and Y and then train test
- 26:39:24split in order to generate the data that
- 26:39:28we will use for building our model.
- 26:39:35Okay. So this this data here is going to
- 26:39:38be what we actually use to guide the
- 26:39:40training of our model. So it's
- 26:39:41definitely supervised, right? Linear
- 26:39:43regression
- 26:39:45um we we will use that
- 26:39:55Okay, so we haven't built the model yet.
- 26:39:57We're just getting our data split apart
- 26:39:59and ready for the training. We haven't
- 26:40:02actually built our model yet, right?
- 26:40:04That'll be coming up uh in a moment. But
- 26:40:07this is getting our data ready. We
- 26:40:09started with our data frame. We split it
- 26:40:11apart into uh an x and a y. And we split
- 26:40:16that into a train test split. And um you
- 26:40:22know then we can uh then we can go ahead
- 26:40:25and um pass in to our model training
- 26:40:28which we'll do in a moment.
- 26:40:36Um you that's a good question. You could
- 26:40:38run so what you could do is you could
- 26:40:40run
- 26:40:41um should we import numpy? Let's see.
- 26:40:45We did. Okay. You could run the average
- 26:40:49on the um you could check the MP mean on
- 26:40:54the X train and see how it compares to
- 26:40:58um
- 26:41:00see how it compares to X.
- 26:41:05So you could you could do that and see
- 26:41:07what the average of this feature is um
- 26:41:09compared to the average of the original.
- 26:41:11They may not be perfect because we are
- 26:41:13taking a reduced data set size. So I
- 26:41:16don't think there's really any good
- 26:41:18there's not like a one-sizefits-all
- 26:41:20validation we can do because we're
- 26:41:21taking a random shuffle and taking a
- 26:41:24percent. We're taking 70% of the data
- 26:41:26out. So we're not guaranteed to maintain
- 26:41:28the same statistics. We can see if
- 26:41:30they're close.
- 26:41:32Um but does that make sense? Like we're
- 26:41:34not guaranteed to get the same stats
- 26:41:36because we're taking a slice of it.
- 26:41:37We're taking 70%.
- 26:41:40So it's not guaranteed to to to
- 26:41:43be the same distribution really.
- 26:41:50Delete that.
- 26:41:53Uh is it good practice? Yes, it is.
- 26:41:57It is. Uh 30% is the industry standard.
- 26:42:00Anything between 20 to 30, so 0.2,
- 26:42:030.25.3,
- 26:42:05any of those are acceptable. It's really
- 26:42:07up to you. Um I mostly see 30%.
- 26:42:12Mo I think.3 is is a good good practice
- 26:42:15to use for sure.
- 26:42:17Um I did explain random state. Uh random
- 26:42:20state is so that you get consistent
- 26:42:23shuffling. Um you can set this to any
- 26:42:26integer that you want it to be. It it
- 26:42:28doesn't really matter. Um you can set it
- 26:42:31to uh 100, you can set it to 10, you can
- 26:42:34set it to 15. Um it just ensures because
- 26:42:38what this split will do is it will
- 26:42:40shuffle the data first. It'll shuffle
- 26:42:42the rows and then um split it apart into
- 26:42:45the into the train and test sets. So you
- 26:42:49set the random state so that the next
- 26:42:51time you run this you get the same
- 26:42:53consistent shuffling. That's the only
- 26:42:55that's the only thing it it helps you
- 26:42:57with because it is randomized but when
- 26:43:00you set a random state um it's so that
- 26:43:03like if you run it again you'll get the
- 26:43:05same shuffling.
- 26:43:07You'll get the same the shuffling
- 26:43:08matters because it it it uh dictates
- 26:43:11what ends up in in these sets.
- 26:43:19Okay.
- 26:43:21All right. So let's see let's do let's
- 26:43:24build the model. Um and let me show you
- 26:43:28how easy this is going to be to build
- 26:43:30the model. And this is really how it's
- 26:43:31going to be for every single scikitlearn
- 26:43:34model will basically look the exact same
- 26:43:37for training it which is what's going to
- 26:43:39make it really really nice. So the first
- 26:43:41thing we have to do is import our model.
- 26:43:45So from scikitlearn we're going to be
- 26:43:46using a linear from the linear model
- 26:43:49package or the linear model module I
- 26:43:53should say within sklearn we're going to
- 26:43:55be importing the linear regression
- 26:43:58and we're going to create an instance of
- 26:44:00the linear regression here.
- 26:44:03Okay, so linear regression and look how
- 26:44:07easy this is going to be. Nearly all
- 26:44:12nearly all sklearn models use
- 26:44:17ffit function to train.
- 26:44:22So every one of them, no matter which
- 26:44:25one we use, like the decision tree, like
- 26:44:28the um logistic regression, any of those
- 26:44:32like we use for classification that are
- 26:44:33going to be coming up in lesson four,
- 26:44:35they're all going to look the same in
- 26:44:37terms of it's going to run.fit,
- 26:44:40which is um scikitlearn's
- 26:44:44uh generic function for training your
- 26:44:46model. So this will execute the training
- 26:44:50once we run this code. And what that
- 26:44:53again the linear regression training is
- 26:44:55going to do that least squares distance
- 26:44:59procedure or algorithm to try to find
- 26:45:02the right weights. It's trying to find
- 26:45:05those weights that minimize that squared
- 26:45:07distance uh from our line that it's
- 26:45:10trying to build to the data.
- 26:45:13And what I want you to notice is what we
- 26:45:16put into the ffit. See how we put in the
- 26:45:19training data where we put in the
- 26:45:20training features and we put in the
- 26:45:23training labels. Now this is supervised.
- 26:45:27So of course we put in the labels,
- 26:45:30right? Of course we put in these labels
- 26:45:33here and of course we put in our
- 26:45:35features here. So we're putting in all
- 26:45:38of our examples from our training split
- 26:45:43into this ffit which is going to train
- 26:45:46the model uh so that we can we can use
- 26:45:50it for prediction.
- 26:45:52Okay, it's really fast. If I run this,
- 26:45:55it's going to be pretty much instant.
- 26:45:58Pretty much instantly it gets trained.
- 26:46:00And you can see here we now have a
- 26:46:02linear regression. you can see in this
- 26:46:03little box. Um, and it and this
- 26:46:06information says that it has been
- 26:46:08fitted. So, it's now ready to be used.
- 26:46:11Right? So, we now that's it. We've
- 26:46:13trained our model. We try that's how
- 26:46:15easy that was. We did ffit. Now, what we
- 26:46:18should realize is there's a lot of work
- 26:46:21going on behind the scenes of this ffit.
- 26:46:24Okay. There's a lot of work being done
- 26:46:26there to do the least squares algorithm
- 26:46:30and find those weights and and create
- 26:46:33that line of best fit. Right? So there
- 26:46:36there's a lot of work being going on
- 26:46:38there that's going on there behind the
- 26:46:39scenes, but scikitlearn is abstracting
- 26:46:42it away for us, right? And all we have
- 26:46:44to do is fit when we're using this code.
- 26:46:48Really easy. Really easy. Fit. And there
- 26:46:52we go. We've trained our linear
- 26:46:54regression model.
- 26:46:58And by the way, if you want to see what
- 26:47:01the coefficients are, you can actually
- 26:47:03extract them if you do so if you take
- 26:47:05your lin regression and you do um
- 26:47:09coefficients like this.
- 26:47:12COF with a with an underscore. So this
- 26:47:16gives us the trained
- 26:47:19weights
- 26:47:21coefficients
- 26:47:23also known as the coefficients right.
- 26:47:26Um so if you run this you can see uh
- 26:47:29right now we have this coefficient here
- 26:47:33um which is the only coefficient we had
- 26:47:35on our feature. So we only had one
- 26:47:38feature coefficient there.
- 26:47:51And we can take a look at our intercept
- 26:47:57which is this.
- 26:48:00So this gives us the train weights
- 26:48:03and so we can look at the intercept we
- 26:48:06can look at the the the coefficient. Um
- 26:48:10so obviously if we have multiple
- 26:48:12features our model has many features
- 26:48:14it's going to have more values in that
- 26:48:16coefficient but the intercept is just
- 26:48:18the single value 7.23
- 26:48:21and then the coefficient
- 26:48:25is 0.046. So that's the weight that gets
- 26:48:28learned.
- 26:48:32Is there a size limit? No, not really.
- 26:48:34There's no size limit. Um,
- 26:48:37no. You can use as much data as you
- 26:48:39want.
- 26:48:41There's really no size limit other than
- 26:48:43what like what you can fit in memory.
- 26:48:48I'd say that's the only limit is
- 26:48:49basically what the amount of data that
- 26:48:51can fit in memory.
- 26:49:00Okay.
- 26:49:02All right. Were you guys able to run
- 26:49:03this? Were you guys able to run the
- 26:49:05linear regression ffit?
- 26:49:09Okay, perfect.
- 26:49:16Perfect. Do you Okay, great. Great.
- 26:49:21So, we have a model and we can use it to
- 26:49:24predict. Um, and so that's actually what
- 26:49:27we're going to do next. If we go down
- 26:49:29here, um we're going to have a function
- 26:49:32that's going to um build a scatter plot
- 26:49:35of our original test data.
- 26:49:39Um so we're going to have our test data
- 26:49:41here.
- 26:49:43Um,
- 26:49:45and we're going to then take our uh
- 26:49:48we're going to take our training data
- 26:49:51and plot we're going to use the uh this
- 26:49:54data versus our sales predictions. So
- 26:49:58you can see we're going to you this is
- 26:50:00how by the way this is how you use the
- 26:50:02scikitlearn model to predict. You have a
- 26:50:05fit to train it and look at the function
- 26:50:08you use to predict. It's literally just
- 26:50:10called predict. That's how easy it is.
- 26:50:13and you pass in your data, all your
- 26:50:15features into this predict and it
- 26:50:17generates a prediction for every row. So
- 26:50:21every row in these features in this data
- 26:50:23frame um will end up with a prediction
- 26:50:27using our model. So what we're going to
- 26:50:30do is plot our training uh features
- 26:50:34against the predicted sales to see how
- 26:50:38good of a fit that really was.
- 26:50:41Okay. to see to see the regression fit.
- 26:50:49Okay. And so there's the regression fit.
- 26:50:52We have all of our test data here
- 26:50:54plotted in the green. We have our blue,
- 26:50:56which is our um we have our our blue,
- 26:51:00which is our uh um training data line
- 26:51:04that we built our model on. So that's a
- 26:51:06pretty decent fit. Um and then our test
- 26:51:10data is here. We just plotted in the
- 26:51:12green scatter. But the thing I want you
- 26:51:14to see is this prediction, right? We we
- 26:51:17were able to generate some predictions
- 26:51:19on that training um by running our
- 26:51:23predict function with our model. Now
- 26:51:24this model has been trained. So we've
- 26:51:27already fit it and now we're using it to
- 26:51:30predict, right? And so we're predicting
- 26:51:32the sales and plotting that on the y
- 26:51:34ais. So the sales are we're using the
- 26:51:38predicted sales there which is our blue
- 26:51:40line. So this is our line of best fit.
- 26:51:44So this is our model prediction.
- 26:51:52This is our model predictions. Right?
- 26:51:56You can see it's a pretty decent uh
- 26:51:58line, right? Pretty decent line of best
- 26:52:00fit.
- 26:52:02Of course, there's some error here. Like
- 26:52:04there, you know, it's not perfect, but
- 26:52:06it it does a decent job of being a best
- 26:52:09fit line.
- 26:52:23Okay.
- 26:52:24So look how easy that was to
- 26:52:27just to recap this to fit our model was
- 26:52:30a linear regression.fit and of course
- 26:52:32we're going to do more examples. So no
- 26:52:35worries uh on that we're going to see
- 26:52:37this many many many times throughout
- 26:52:39this notebook. But we have linear
- 26:52:42regression.fit to train it and then we
- 26:52:45have linear regression.predict
- 26:52:48to and we pass in our features and that
- 26:52:50generates a predicted output.
- 26:52:53Right? So what this is actually doing is
- 26:52:57is computing this quantity.
- 26:53:12We could do either.
- 26:53:14We could do either. Um, so we could do,
- 26:53:19so one thing we could do is plot uh, so
- 26:53:22we could swap it out. We, we could do
- 26:53:24either one. It doesn't, it's not a big
- 26:53:26deal to do the training set. We could
- 26:53:28do, so we could plot X test and then we
- 26:53:31could plot linear regression X test.
- 26:53:42So it's it's a similar line. Um it's
- 26:53:47just different input features, but the
- 26:53:48line is going to be the same. Just
- 26:53:51different inputs,
- 26:53:53but the coefficients are the same,
- 26:53:55right? It's the same line. It's just we
- 26:53:57generate different outputs.
- 26:54:02So yeah, you could do either one.
- 26:54:06This is This is honestly this is
- 26:54:08probably better. I see what you're
- 26:54:10saying. This is probably better because
- 26:54:11this is the line of best fit through
- 26:54:14this data. So that probably makes sense
- 26:54:16to do to do predict on the test set.
- 26:54:20Agreed on that. Probably makes about
- 26:54:23most sense.
- 26:54:30But you could do either one.
- 26:54:42Yeah, I think that would be the most I
- 26:54:44think that makes the most sense is for
- 26:54:46it to be on the same one just to
- 26:54:48validate. So like we could do we could
- 26:54:50do training here and then train and
- 26:54:52train just to see how that data lines
- 26:54:55up. Really, what we're trying to do is
- 26:54:58have our scattered data and then our
- 26:55:00line of best fit on the same plot.
- 26:55:03That's all we're trying to do, right?
- 26:55:05So, yeah, I think I think they should be
- 26:55:06the same.
- 26:55:10I think that makes sense.
- 26:55:14These values
- 26:55:17or which values do you want to see?
- 26:55:24Yeah, we could uh we could generate
- 26:55:26those if we just do um let's go down
- 26:55:29here. So the the line values
- 26:55:34um are going to be uh the prediction. So
- 26:55:38um the the
- 26:55:40uh test
- 26:55:43predictions
- 26:55:45equals um
- 26:55:50test predictions equals linear
- 26:55:52regression.predict predict x test and
- 26:55:54then we could uh we could print out our
- 26:55:56test predictions.
- 26:56:01Yeah. So we can see what those actual
- 26:56:03values are on our uh on the test set.
- 26:56:07Yeah.
- 26:56:18Um we will do that. Yeah. So you thought
- 26:56:20we were checking how well our data was
- 26:56:21trained. We will do that. Yes, we
- 26:56:23haven't learned how to evaluate this
- 26:56:24yet. We're going to talk about that
- 26:56:26coming up next. Yeah, we will do that.
- 26:56:29We just haven't learned how to do proper
- 26:56:31evaluation
- 26:56:33of a regression model.
- 26:56:36But yeah, it's something we're going to
- 26:56:37talk about for sure
- 26:56:39and see how to do in our code.
- 26:56:46Okay.
- 26:56:49All right. Any other uh questions on
- 26:56:52this example?
- 26:56:59Again, big takeaways
- 26:57:02fit to train it and then predict to use
- 26:57:06it.
- 26:57:08Predict on the features to use the model
- 26:57:11and make predictions with it.
- 26:57:16So here is example. We we made all the
- 26:57:18predictions. This these are all the
- 26:57:19values that are on that line.
- 26:57:22These are all our predictions. And
- 26:57:23notice they this is a truly regression,
- 26:57:25right? These are all floating point
- 26:57:27values. Um so this is definitely a
- 26:57:29regression, right?
- 26:57:43Okay.
- 26:57:50Uh, that's a good question. Um,
- 26:57:54I'm not sure if there is
- 26:57:58if there's like a verbose
- 26:58:03there's not really no there's not really
- 26:58:05a verbose. You can I mean you can look
- 26:58:06at the source code if you really want to
- 26:58:08see you can view the source code to see
- 26:58:11um how it's done. I can tell you I mean
- 26:58:14so generally linear regression is done
- 26:58:17in two ways. Either you use a formula um
- 26:58:21to to solve the optimization problem of
- 26:58:24minimizing like this this uh distance
- 26:58:27from the points to to the line. Um
- 26:58:32or you use something called gradient
- 26:58:33descent which is how a lot of these
- 26:58:35things do it is they iterate through a
- 26:58:39bunch of different iterations where they
- 26:58:40update these weights according to um a
- 26:58:44certain uh basically a gradient of the
- 26:58:48the error function. The error function
- 26:58:50in this case is the is the squared
- 26:58:53distance from the line to the uh to to
- 26:58:59the points.
- 26:59:01So uh we can compute the gradient of
- 26:59:03that and do um gradient descent. So if
- 26:59:06you really want to look into it, I would
- 26:59:08do some research on like linear
- 26:59:10regression gradient descent.
- 26:59:13Okay, linear regression gradient descent
- 26:59:15to see how that's uh how that's being
- 26:59:17done. Yeah, it it's it's a pretty simple
- 26:59:21procedure. Um, again, you have the the
- 26:59:25notion is that you want to minimize
- 26:59:28minimize the loss or the error. Uh, in
- 26:59:32this case, the loss is the square
- 26:59:34distance. So, it's like um there's like
- 26:59:38a it's a formula. It's like a sum of a
- 26:59:41square distance from your prediction
- 26:59:44um or your label sorry to your model
- 26:59:48which is the beta 0 um plus beta 1 x1
- 26:59:53plus beta 2 x2
- 26:59:56etc like your model and then squared. So
- 26:59:59this squared this is the squared
- 27:00:01distance here and you're minimizing this
- 27:00:04guy which is like a calculus problem.
- 27:00:07You you find you basically find the this
- 27:00:10is this is a I'm getting so far into the
- 27:00:12weeds of this, but this is like a
- 27:00:14parabola and you work your way No, no,
- 27:00:17you're good. It's it's it's a good
- 27:00:19question. Um you work your way down to
- 27:00:22the minimum of it. Does that make sense?
- 27:00:24Like you're working your way down here
- 27:00:26and you do that through a descent
- 27:00:28process, like a descent iteration.
- 27:00:31Um
- 27:00:33so
- 27:00:35that's how these are found.
- 27:00:37Um, but you don't see that happening in
- 27:00:41the background. But if you look at the
- 27:00:42source code, it I guarantee you it would
- 27:00:44be it's either going to be this or
- 27:00:46they're going to use the they're going
- 27:00:48to use a a a matrix formula to basically
- 27:00:51solve an equation um that involves this.
- 27:00:57Basically, the derivative of this set
- 27:00:59equal to zero and you find the minimum.
- 27:01:02Either way, you're finding the minimum
- 27:01:03of this.
- 27:01:09Okay. But yeah, I don't think Psycharn
- 27:01:12has like a uh maybe there's some type of
- 27:01:15verbose flag you can look for.
- 27:01:18I don't think they have that though. Not
- 27:01:21that I've seen.
- 27:01:31All right.
- 27:01:33So I have uh an important um concept to
- 27:01:37talk about next which is going to be uh
- 27:01:40called overfitting and underfitting
- 27:01:43um which is a really important concept
- 27:01:45that's related to the training and test
- 27:01:48data we just split apart to do
- 27:01:51evaluation.
- 27:01:53And um essentially the the issue with
- 27:01:56machine learning is that it's not
- 27:01:58perfect and it can struggle in different
- 27:02:00ways. And the two ways that it primarily
- 27:02:03struggles is going to be overfitting and
- 27:02:05underfitting.
- 27:02:06So overfitting is a situation where the
- 27:02:11model basically memorizes the training
- 27:02:14data so well that it's it fails to
- 27:02:18generalize to new examples. So what we
- 27:02:21see with overfitting is this exact sign
- 27:02:25here where we have really good
- 27:02:26performance on the training data. So
- 27:02:28when so when we do that train test split
- 27:02:30we see a really good accuracy or really
- 27:02:34low error on the training data but it
- 27:02:39does not perform anywhere near that on
- 27:02:41that test data split. So what that means
- 27:02:44is that the model is overfitting to the
- 27:02:48training data. It's basically memorizing
- 27:02:50it and it's not able to generalize very
- 27:02:54well.
- 27:02:56Now, why does that happen? It's usually
- 27:02:58because the model is way too complex.
- 27:03:01And that means generally you need to do
- 27:03:04something to reduce the complexity.
- 27:03:07Either you need to use a simpler model
- 27:03:10or you need to use some type of
- 27:03:12technique to mitigate overfitting. And
- 27:03:15we're going to we're going to study some
- 27:03:17of those techniques coming up in this
- 27:03:18notebook. uh we might not get to it
- 27:03:20today, but we're going to study
- 27:03:22particularly what can we do to prevent
- 27:03:24overfitting because overfitting is the
- 27:03:26more common issue with machine learning
- 27:03:28models. They tend to do so well at
- 27:03:32learning from data that they pick up on
- 27:03:34small details and patterns in the
- 27:03:37training examples that they're exposed
- 27:03:38to. They don't do a great job at
- 27:03:41generalizing to new examples. They can
- 27:03:43struggle with that.
- 27:03:45So that's overfitting is struggling to
- 27:03:48generalize to new examples, but you do
- 27:03:50really well on your training data. So it
- 27:03:53appears like you have a good model, but
- 27:03:55it it's not able to go and make
- 27:03:57predictions on test data very well,
- 27:03:59which means we would not want to use
- 27:04:01that model in the real world, right?
- 27:04:03Because it's not able to generalize
- 27:04:05outside of what it's already seen. And
- 27:04:07that's not a good thing if we're trying
- 27:04:08to use it for real world examples,
- 27:04:10right?
- 27:04:12So overfitting is a real issue. Um you
- 27:04:16see it all the time. I've seen it many
- 27:04:17many times in the real world, real
- 27:04:19industry uh work that I've done.
- 27:04:22Overfitting is a is a challenge for a
- 27:04:25lot of machine learning models. And so
- 27:04:26we need some techniques to overcome
- 27:04:29overfitting. And we're going to study
- 27:04:31some of those uh coming up shortly.
- 27:04:35Um, one of the things that we can do,
- 27:04:39one of the one of the things that we can
- 27:04:40do to detect overfitting is exactly what
- 27:04:43we just did, which is you split apart
- 27:04:46your data into training and testing so
- 27:04:48that you have a chance to do an
- 27:04:50evaluation to see if you're even
- 27:04:52overfitting in the first place. You want
- 27:04:54to see that performance be consistent
- 27:04:58from train to test, right? You want to
- 27:05:00see consistency. What you don't want to
- 27:05:02see is performance that drops off on the
- 27:05:05test data. It's much worse. You don't
- 27:05:08want to see that. That means that your
- 27:05:10model is overfit uh to your training
- 27:05:12data and it's not going to perform well
- 27:05:14in the real world.
- 27:05:17Okay. So, we're going to have a couple
- 27:05:18ways to uh overcome that. Talk about
- 27:05:22that. Um now, the opposite can actually
- 27:05:25happen as well, which is called
- 27:05:27underfitting. And underfitting
- 27:05:30refers to the fact that a model is too
- 27:05:33simple and it actually just performs
- 27:05:37poorly across the board. So if we see
- 27:05:40poor performance on the training and
- 27:05:44testing data, that's a good signal that
- 27:05:46the model's underfit and that means it's
- 27:05:50too simple usually and you should try
- 27:05:52using something more complex. Um, so the
- 27:05:55best way to combat underfitting is to
- 27:05:57use a more complex model. And as we go
- 27:06:01through and learn about the models,
- 27:06:03we're going to learn about which ones
- 27:06:05are simple and which ones are complex.
- 27:06:06So we're going to have a scale of kind
- 27:06:09of complexity. And if you're
- 27:06:11underfitting, you want to bump up to the
- 27:06:13to a more complex model. If you're if
- 27:06:16you're overfitting, one way of combating
- 27:06:18that is to actually go down to something
- 27:06:20more simple. Go the opposite way to
- 27:06:22something simpler. So we need to learn
- 27:06:24right now we've only learned linear
- 27:06:26regression
- 27:06:27but we will learn other models you know
- 27:06:30in the future and we'll we'll talk about
- 27:06:32uh their complexity and how they're
- 27:06:34related to each other.
- 27:06:36Okay, but these are two issues we see
- 27:06:38just to draw that out again is if we
- 27:06:41have a train test split where we have
- 27:06:437030 split let's say and we perform
- 27:06:46really well over here but we go to apply
- 27:06:49that model over here and it fails its
- 27:06:51accuracy drops off significantly more
- 27:06:54error that's that's definitely
- 27:06:56overfitting which is not good
- 27:07:00right and then underfitting is just not
- 27:07:02performing well in either case so even
- 27:07:04on the training data itself your your
- 27:07:06accuracy is not very good. So you're not
- 27:07:09really learning effectively. You're
- 27:07:11underfitting your model. So that's
- 27:07:15that's um underfitting case.
- 27:07:21Okay.
- 27:07:26All right. Now the issue is that it can
- 27:07:29be very difficult to balance these two
- 27:07:32and get it correct. That's what makes
- 27:07:33machine learning a little bit
- 27:07:34challenging is getting this balance
- 27:07:37correct of simplicity and complexity. So
- 27:07:41you don't want to be overly complex that
- 27:07:43you overfit, but you don't want to be
- 27:07:45overly simple that you underfit and
- 27:07:48you're not able to learn effectively. So
- 27:07:51there's a bit of a tradeoff there. And
- 27:07:53this trade-off is typically known in the
- 27:07:55community as bias variance trade-off. Um
- 27:07:58in which case uh it's basically like a
- 27:08:01complexity simplicity trade-off. It's
- 27:08:02another word for that. Um,
- 27:08:06and so, uh, it's it's thought that, um,
- 27:08:10if you, uh, if you have very, um, if you
- 27:08:15have a situation where you're able to
- 27:08:16fit the training data very well, you
- 27:08:19risk not being able to generalize. In
- 27:08:22other words, you risk overfitting, and
- 27:08:24it's hard to um, it's hard to combat
- 27:08:28that in a way. Um, and um, on the
- 27:08:33reverse side, if you have something
- 27:08:34really simple, um, you risk not learning
- 27:08:38enough. Even if you're trying to combat
- 27:08:40that overfitting, you risk not learning
- 27:08:43enough and your model just doesn't
- 27:08:45perform as well as it could. So, there's
- 27:08:47a bit of a trade-off there of trying to
- 27:08:49find the right balance between something
- 27:08:51complex enough to learn, but something
- 27:08:54not overly complex that it's going to
- 27:08:58not generalize to new data. That's the
- 27:09:01challenge. Um, like I said, we are going
- 27:09:05to have techniques to overcome this. So
- 27:09:08luckily there are things to basically
- 27:09:10overcome this trade-off and um and help
- 27:09:15us along the way so that we don't
- 27:09:17overfit. They basically prevent
- 27:09:18overfitting
- 27:09:20um and allow us to use complex enough
- 27:09:23models um that that won't be overfit.
- 27:09:27This is in the um this was in our uh
- 27:09:31lesson 3.2 notebook. So you want to pull
- 27:09:34that one back up. We were working on
- 27:09:35Monday.
- 27:09:37Um, and just to recap this a little bit,
- 27:09:40remember we were building a linear
- 27:09:42regression, I wanted to recap some of
- 27:09:44the steps we took there, um, that we
- 27:09:48will be doing over and over again. And
- 27:09:50really the same kind of steps, uh, that
- 27:09:53we do here, we'll do in a lot of our
- 27:09:55model building. Pretty much all of our
- 27:09:57model building um, that we do, whether
- 27:09:59it's regression or classification,
- 27:10:01doesn't really matter. um we'll still be
- 27:10:04doing a lot of these steps which are um
- 27:10:07remember first we split apart our data
- 27:10:09into kind of a features and a label
- 27:10:13uh x and y and the reason that's
- 27:10:16important is because um the model
- 27:10:19training uses the features and the label
- 27:10:23um to help train the model, right? They
- 27:10:26use those separately. Um so we want to
- 27:10:29split those apart whenever we can. And
- 27:10:31so we have usually uh it's a good
- 27:10:33practice to call your features capital X
- 27:10:35and your labels lowercase Y. And what we
- 27:10:39do with that is remember we immediately
- 27:10:42split that into what we call the
- 27:10:44training in a test set. And the picture
- 27:10:47we had for that was something like this
- 27:10:51where we had about 70% of the data
- 27:10:55we used to train the model against and
- 27:10:58then the other 30% of the data we use to
- 27:11:01test the model against. Meaning that we
- 27:11:04build a model over here and we apply it
- 27:11:07to this set over here um to make
- 27:11:10predictions. And then the that's where
- 27:11:12the supervised learning really comes
- 27:11:14into play, right? is on this test set.
- 27:11:17We already have the answers. We already
- 27:11:19have the label. And so we can apply our
- 27:11:21model to this to the features over here.
- 27:11:24Predict uh what the the label should be
- 27:11:27and compare that. We can get a a metric,
- 27:11:30right, that compares how close we are in
- 27:11:33our prediction to the actual values. Um
- 27:11:36and that was some of our performance
- 27:11:38metrics. I'll recap some of those that
- 27:11:40kind of measure that distance away from
- 27:11:42our predictions to what the actual label
- 27:11:45is. Um, but remember we had this train
- 27:11:49test split function which helps us split
- 27:11:52apart our features and our labels into
- 27:11:55these uh four sets of data. So we have
- 27:11:58our training features, our testing
- 27:12:00features and then our training labels
- 27:12:02and our testing labels. So we have all
- 27:12:05of those and um really these two guys
- 27:12:08are going to be used to train the model.
- 27:12:10That's why they're called underscore
- 27:12:12train. They're going to be used to train
- 27:12:13that model. And then the then we're
- 27:12:16going to predict on these set of
- 27:12:18features and then com use those
- 27:12:20predictions to compare to this set of
- 27:12:23labels, right? That's on the test test
- 27:12:25set. Um and you notice here our test
- 27:12:28size is set to 30%. Um, that's a pretty
- 27:12:32standard number. Anywhere between like
- 27:12:3320 to 30% is pretty standard. Um, we'll
- 27:12:37typically use.3, but it could be 02.
- 27:12:40Anywhere in between is fine.
- 27:12:44Okay, so we had that. Hopefully that uh
- 27:12:47we remember that from Monday.
- 27:12:49So we had a train and a test set. And
- 27:12:51then building the model was actually
- 27:12:53really really easy. Once you have those
- 27:12:55train and test sets, um, we just import
- 27:12:57our model object. So from uh scikitlearn
- 27:13:00sklearn
- 27:13:02um linear model uh module from that
- 27:13:05package we import the linear regression
- 27:13:08model and then we do um linear
- 27:13:11regression.fit
- 27:13:12and we pass in our features and our
- 27:13:14labels and this is again this is where
- 27:13:17that supervised learning is really
- 27:13:18coming into play because we're passing
- 27:13:21in these labels.
- 27:13:23That's really what makes this work,
- 27:13:24right? We need those labels to help
- 27:13:26guide the model to make those updates.
- 27:13:29If you guys remember, the model is
- 27:13:32something that looks like this.
- 27:13:37So, this was a bunch of different
- 27:13:40coefficients
- 27:13:41um times the features,
- 27:13:44however many we have. Um, and so these
- 27:13:49labels are really taking the place of
- 27:13:51this and they're helping us um make the
- 27:13:56correct updates to these to these
- 27:13:58coefficients or sometimes we call them
- 27:14:00weights. Um, these B 0, B1, B2. Um, we
- 27:14:05find out what the optimal one is to get
- 27:14:07the best fit, right? To get the line of
- 27:14:09best fit. Um that's what the model
- 27:14:13training when we call this ffit ffit
- 27:14:15that's really what it's doing in the
- 27:14:16background is finding all those
- 27:14:18coefficients right to end up with the
- 27:14:20line of best fit that has the lowest
- 27:14:21amount of error.
- 27:14:26Okay so hopefully that makes sense.
- 27:14:28That's just a dofit fit um to train our
- 27:14:31models. And that's really going to be um
- 27:14:33the case for
- 27:14:36uh pretty much every single model that
- 27:14:39we uh train with scikitlearn. It's
- 27:14:42pretty much going to be a fit. We pass
- 27:14:43in our training uh features and our
- 27:14:46training labels.
- 27:14:49Okay, so we had that and this was the
- 27:14:52visualization of that where we had our
- 27:14:54test points kind of scattered and we see
- 27:14:57our line of best fit is the one that
- 27:14:59goes through there with that minimal
- 27:15:01error. That's that's the whole goal.
- 27:15:05Pretty decent predictor.
- 27:15:10Okay. And then we talked about
- 27:15:13overfitting underfitting. So just to
- 27:15:15recap this overfitting is the concept of
- 27:15:17our model basically memorizing our
- 27:15:19training data. It performs really well
- 27:15:21on that training set but it is not able
- 27:15:24to generalize outside of that. So it
- 27:15:27performs poorly on the test set or data
- 27:15:30that it's never seen before. Um and
- 27:15:33that's overfitting. So the reason that
- 27:15:36it overfits is generally the model is
- 27:15:38too complex and it needs to be um it
- 27:15:42needs to be simplified a bit. And one of
- 27:15:44the things we're going to do today is
- 27:15:46see a couple of ways we can alter the
- 27:15:48linear regression model um if we are
- 27:15:51overfitting to prevent overfitting. Um
- 27:15:55so there's going to be ways to handle
- 27:15:57this. Um and so we're going to explore
- 27:16:00some of those today.
- 27:16:02uh underfitting is kind of the reverse
- 27:16:04of that. Remember, it's where the model
- 27:16:06is not learning enough. So, the
- 27:16:07performance is poor even on the training
- 27:16:09data. It's not good on the test data
- 27:16:12either. Um that is a sign that the model
- 27:16:16is probably too simple and maybe we
- 27:16:18should use something more complex like
- 27:16:20go from a linear regression maybe to use
- 27:16:22a polomial regression. Um or maybe use
- 27:16:25an entirely different model altogether.
- 27:16:28um if we're underfitting, our
- 27:16:29performance is poor, it's a good signal
- 27:16:31we should try something else. Um
- 27:16:36okay,
- 27:16:37so we talked about those
- 27:16:41and one of the things we also talked
- 27:16:43about was evaluations. If you guys
- 27:16:45remember, we had different metrics that
- 27:16:47we could compute to get a gauge of how
- 27:16:50good our model is actually performing.
- 27:16:52Um one of those was MSE, which is this
- 27:16:55mean squared error function. Um so we
- 27:16:57did this example during class last time
- 27:16:59on Monday um where we uh were able to
- 27:17:04generate the mean squared error. That's
- 27:17:07one of our metrics. And we can see what
- 27:17:09the mean squared error is on the
- 27:17:12training set and see what it is on the
- 27:17:13test set by um just passing in our um
- 27:17:17training predictions and our training
- 27:17:18labels, our test predictions and our
- 27:17:21test labels. pass those into this mean
- 27:17:23squared error function and it computes
- 27:17:25the MSE and that's that's a helpful
- 27:17:27function from the scikitlearn metrics
- 27:17:31um package um or module I should say and
- 27:17:35we'll be using that quite a bit to do
- 27:17:38you know evaluation of of especially of
- 27:17:40regression right mean squared error is
- 27:17:42pretty is probably the most common uh
- 27:17:46performance metric we can have and if
- 27:17:48you guys remember what it's really doing
- 27:17:50is measuring these distances So mean
- 27:17:52squared error is kind of like the
- 27:17:54average distance away from our our
- 27:17:56points to the actual um to the
- 27:17:59predictions which the predictions are
- 27:18:02all on this line. Um so it's like
- 27:18:05measuring on average how how much error
- 27:18:07do we have on average right? Um, and the
- 27:18:10idea is the closer to zero the better.
- 27:18:13Generally means that the distance away
- 27:18:15from our prediction to our points is
- 27:18:18pretty low. The closer to zero it is.
- 27:18:20Um, which is pretty desirable.
- 27:18:23So a low MSE is kind of what we're
- 27:18:25looking for. Um, closer to zero the
- 27:18:28better. And so um if one model has if
- 27:18:32one model has um a low lower MSE than
- 27:18:36another, it's it's a better performing
- 27:18:38model, right? It has less error.
- 27:18:42Okay. And then we also looked at the R r
- 27:18:45squared or sometimes known as R2 um
- 27:18:48score. Um this is another metric that we
- 27:18:52could use that measures the the
- 27:18:55variability
- 27:18:56um of uh the predictions and if our
- 27:19:01model is capturing that variability um
- 27:19:03well um and so R squar is has a range of
- 27:19:070 to one one is better that means the
- 27:19:10model is capturing the the changes in in
- 27:19:12the um output it um our predictions
- 27:19:16follow along with those same changes um
- 27:19:18so they're pretty close um so closer to
- 27:19:22one would be a better score. So we have
- 27:19:25those kind of metrics. So like on this
- 27:19:27data um this would this would show that
- 27:19:30this model was underfitting remember
- 27:19:32because this
- 27:19:34mean this MSE was bad and this MSE was
- 27:19:38bad.
- 27:19:40Um and what we should think of these in
- 27:19:42the units of what our labels are. um
- 27:19:47especially if we take the square root of
- 27:19:49this the RMSSE that was another metric
- 27:19:51we had um the square root of this is
- 27:19:54actually in the exact units that we um
- 27:19:58have for our labels. So uh in this
- 27:20:00example this was the um this was the the
- 27:20:04units or the sales versus the TV
- 27:20:07products, right? Um and so this would
- 27:20:11indicate that on average if we take the
- 27:20:13square root of this um
- 27:20:16and the square root of this um we have
- 27:20:19uh
- 27:20:20um we're on average about 11 sales units
- 27:20:24off squared. So if we take the square
- 27:20:26root of that um it's somewhere around 3
- 27:20:27to four um somewhere in between three
- 27:20:31and four units off. And this is as well.
- 27:20:35Um, and because both of these are still
- 27:20:37not close to zero. Um, this would be
- 27:20:39under fit. And this shows that as well.
- 27:20:42This isn't that close to one. It's
- 27:20:44decent, but it's not um not that close
- 27:20:47to one. So, we would say and performance
- 27:20:49is poor on both training and test sets.
- 27:20:52That's the key indicator of
- 27:20:53underfitting. It's poor on both.
- 27:21:01Yeah. Exactly. High MSE correlates to
- 27:21:03underfitting. Yes. Yes. And it what's
- 27:21:05key is it's high MSE on both on both the
- 27:21:10training and the test sets.
- 27:21:13If you have a high MSE on your test set
- 27:21:15but a low MSE on your training set,
- 27:21:17that's overfitting, right? Where it's
- 27:21:20not generalizing from the training set
- 27:21:22to the test data that it hasn't seen
- 27:21:24before. That's overfitting. So the key
- 27:21:27is high MSE on both sets.
- 27:21:33All right. So we talked about that. Um
- 27:21:35we did polomial regression last time. So
- 27:21:38that was um doing
- 27:21:42that was uh making a curved graph um by
- 27:21:46transforming the features into polomial
- 27:21:48features and then doing linear
- 27:21:49regression with that. So you guys
- 27:21:51remember from Monday we did this where
- 27:21:54um we took our features and uh transform
- 27:21:58them according to this polomial features
- 27:22:00from scikitlearn. So we can go all the
- 27:22:02way up to degree whatever degree we
- 27:22:04want. So we put in four here, but
- 27:22:06there's nothing special about four
- 27:22:07really. This is just testing it out. Um
- 27:22:10and we generate the the polomial
- 27:22:12features and we can fit a linear
- 27:22:14regression on those polomial features
- 27:22:17and we get a slightly better model,
- 27:22:20right? Um it fits the data a little bit
- 27:22:23better than just a straight line. this
- 27:22:25curved line with the polomial
- 27:22:28features um performs a little bit better
- 27:22:30and we could see that with the MSE right
- 27:22:31we could evaluate the MSE of this um and
- 27:22:35it would be lower
- 27:22:37it would be lower than the curve line
- 27:22:39and that's something we could do um we
- 27:22:42would just have to pass in these test
- 27:22:44predictions training predictions and
- 27:22:45then the the test labels and training
- 27:22:48labels and passes into the mean squared
- 27:22:50error function and we could compute that
- 27:22:51right wouldn't be hard to
- 27:22:56All right. And then finally, where we
- 27:22:58left off, um, you know, is on our
- 27:23:01performance metrics. So, we talked about
- 27:23:03mean squared error. That's that average
- 27:23:05distance away from the labels to our
- 27:23:08predictions. Um, and we take the square
- 27:23:11root of that. It's it's basically
- 27:23:13measuring the same thing, but it's the
- 27:23:15square root of it is um more
- 27:23:17interpretable because it's in the same
- 27:23:19units as our label.
- 27:23:21um mean absolute error is is the average
- 27:23:25distance of the absolute value. So it's
- 27:23:27not the squared distance formula like a
- 27:23:29uklidian distance but it is a absolute
- 27:23:32value. So it's a little bit um less
- 27:23:34sensitive to outliers. They don't get
- 27:23:36magnified as much. Um but it's not
- 27:23:41typically used as much as a mean squared
- 27:23:43error would be with regression. um we
- 27:23:46talked about the last time because um
- 27:23:48the distance formula or that distance is
- 27:23:51actually what's used to train the model.
- 27:23:53So it's a more natural um fit for a
- 27:23:57performance metric for it.
- 27:24:02All right. And then we had R square. We
- 27:24:03just talked about that closer to zero
- 27:24:05would be um worse. Closer to one would
- 27:24:08be better. That means that the model
- 27:24:10explains um all the variability in the
- 27:24:13in the predictions. Uh it captures those
- 27:24:16predictions um closely to the labels
- 27:24:21um very well. So uh one would be better.
- 27:24:26Closer to one would be better.
- 27:24:29All right. So that's where we left off.
- 27:24:31Um we're gonna pick up from there with
- 27:24:33cross validation. Um, we've actually
- 27:24:36already seen one method of cross
- 27:24:38validation. So, we're going to study um
- 27:24:40we're going to kind of recap that and
- 27:24:41and then um talk about cross validation
- 27:24:44in general um and look at some more
- 27:24:48sophisticated techniques of it um coming
- 27:24:51up next. But before I do that, any
- 27:24:53questions about anything we've covered
- 27:24:56um to this point in in the recap or
- 27:24:59anything from Monday? Any questions on
- 27:25:02that?
- 27:25:04All right. So let's talk about uh cross
- 27:25:07validation. Um now this term cross
- 27:25:12validation refers to a technique that
- 27:25:16evaluates performance. And what it does
- 27:25:19is it divides our data into essentially
- 27:25:23um training and test sets which we've
- 27:25:25kind of already seen. And then we are
- 27:25:27able to train a model on on the training
- 27:25:30set, evaluate it on the test set, and
- 27:25:32that's where that's where we get the
- 27:25:34name cross validation because we're
- 27:25:36crossing over our model from one batch
- 27:25:38of data used to train it over to another
- 27:25:41set of data used to validate those
- 27:25:43predictions. Um, and there's actually
- 27:25:47different ways to do cross validation.
- 27:25:48So cross validation is a bit of an
- 27:25:50umbrella term for multiple ways to do
- 27:25:52that. We've already seen one way of
- 27:25:54doing that um which I'm going to scroll
- 27:25:56down to is um known as a hold out cross
- 27:26:01validation. So that's um what we've been
- 27:26:03doing so far. So this is just um
- 27:26:06generating a train and a test set
- 27:26:10train um split.
- 27:26:13Um that's the that's what's known as the
- 27:26:16hold out cross validation method. Um and
- 27:26:19and this is exactly what we've been
- 27:26:21doing so far, which is you split your
- 27:26:24data into some type of split, usually
- 27:26:267030,
- 27:26:28um of a train and test
- 27:26:31and then you um train your model on this
- 27:26:34section of data and then apply it to
- 27:26:37this to evaluate performance. Right? So
- 27:26:39that's that's what's known as the hold
- 27:26:41out method. Um it is uh you know
- 27:26:46relatively simple. It's pretty fast to
- 27:26:49do. Um, but there are more robust ways
- 27:26:53to try to divide up our data a little
- 27:26:56bit uh more evenly. Instead of just
- 27:26:59having one split, we can actually do
- 27:27:01many splits, which is the idea of um the
- 27:27:04next kind of cross validation I'll
- 27:27:06cover. But hold out method is one that
- 27:27:09we've already studied. It's the most
- 27:27:11basic type of cross validation you can
- 27:27:13have. Um so hold out this is the most
- 27:27:16basic
- 27:27:18and we we've already been we've already
- 27:27:21been uh working with this type. Okay.
- 27:27:26So we've we've already seen hold out
- 27:27:28method. Let me uh explain to you a more
- 27:27:31sophisticated method a little bit more
- 27:27:33advanced of a cross validation um which
- 27:27:36is known as Kfold cross validation. So
- 27:27:39this is um going to be a little bit more
- 27:27:42advanced of a technique but this is the
- 27:27:45idea of kfold is that you take your data
- 27:27:47set
- 27:27:49and you split it into k number of what
- 27:27:54are called splits or folds. So you take
- 27:27:57your data and you let's say it was let's
- 27:27:59say k equals 5. So we have five splits
- 27:28:02here.
- 27:28:05Okay. So let's say k equals 5. we have
- 27:28:08five splits. So what we're going to do
- 27:28:13is we're going to we're going to train
- 27:28:15our model on K minus one of those folds.
- 27:28:19So if K was five, we had five splits.
- 27:28:22We're going to take our model and train
- 27:28:24it on four out of five of those uh
- 27:28:28splits. So let's say it's these four.
- 27:28:32We train it on these four.
- 27:28:36Okay. And then what we do is the one
- 27:28:39split that's left over we will we will
- 27:28:43test our model against that split. So
- 27:28:45we'll test here.
- 27:28:50Okay. Now, this sounds very similar to
- 27:28:52the hold out method where we're doing a
- 27:28:54train test split, but it's a little bit
- 27:28:56this kful cross validation is a little
- 27:28:58bit more sophisticated because we repeat
- 27:29:00this process that I just mentioned over
- 27:29:03and over for all combinations of the
- 27:29:06splits. So then what we'll do, this is
- 27:29:08just one trial that we'll do it again,
- 27:29:13but this time we will pick um four
- 27:29:16different splits.
- 27:29:18So, this time we might pick,
- 27:29:21let me do blue. This time we might pick
- 27:29:24this one, this one,
- 27:29:27um,
- 27:29:29this one,
- 27:29:32and this one.
- 27:29:35And then those four we will train our
- 27:29:37data on. And then we will test against
- 27:29:39this one. Okay? And we'll do we'll
- 27:29:42repeat this
- 27:29:45repeat for all combos of the folds.
- 27:29:55Okay. So we'll repeat that. So
- 27:29:58essentially what we're doing is rotating
- 27:29:59through. Every time we rotate through
- 27:30:02one of the folds is going to be left out
- 27:30:03as a test set. Now this is a little bit
- 27:30:07more robust than just a train test
- 27:30:09split, right? because we are exposing
- 27:30:12our model to more of the data in in
- 27:30:15doing this, right? Because we're going
- 27:30:17to split it evenly into five or 10
- 27:30:20splits. Those are pretty common um
- 27:30:22number of folds to use. 10 or five. Um
- 27:30:25those are the ones I've most commonly
- 27:30:27seen. Um but we're going to by rotating
- 27:30:32through which folds are being used for
- 27:30:33training, which ones being left out. um
- 27:30:36we are exposing our our model to more of
- 27:30:39the data this way than just doing a
- 27:30:41single train test split. Right? So now
- 27:30:44what do we do with with the results is
- 27:30:47every time we do this we we generate um
- 27:30:50an MSE let's say or some type of
- 27:30:52performance metric. So let's say we
- 27:30:54generate an MSE from this guy,
- 27:30:57we generate an MSE from this version and
- 27:31:00we generate an MSE for all combos.
- 27:31:04each combo we generate MSE and then what
- 27:31:07we do is we average
- 27:31:10the metrics
- 27:31:13or the in this case uh if we use MSE we
- 27:31:16would average those together. So every
- 27:31:19time we do a fold combination and we
- 27:31:21keep four of them for training, one for
- 27:31:23test and we rotate through all those
- 27:31:25combinations, we are going to generate
- 27:31:27an MSE for every combination
- 27:31:30then we're just going to average those
- 27:31:32MSE's to get a final. So the final MSE
- 27:31:36of cross val of this kffold.
- 27:31:40So the final metric
- 27:31:43is just the average of the uh
- 27:31:46performance on all of the fold
- 27:31:48combinations. Okay. So our final MSE, we
- 27:31:52just average all those MSE's from all of
- 27:31:54our combinations.
- 27:31:56Okay.
- 27:31:58Now, what's the advantage to doing this?
- 27:32:01It's way more robust of a estimate of
- 27:32:04the of the performance of the model
- 27:32:06because we're exposing it to all
- 27:32:09basically all of our data, right? We're
- 27:32:11getting a sense of how it performs
- 27:32:12across all those different folds. Um
- 27:32:16rather than just doing a single train
- 27:32:18test split, which is a bit it's basic,
- 27:32:20it works, but it's a bit basic. Um so
- 27:32:23this is more robust estimate of the
- 27:32:26performance.
- 27:32:28Now, what's the drawback to doing this
- 27:32:30is that it's more intensive. So, if you
- 27:32:32have a lot of data, this is going to be
- 27:32:34pretty expensive to do because you're
- 27:32:36going to have to especially you have a
- 27:32:37high number of folds, right? You're
- 27:32:39going to have to divide your data into k
- 27:32:41number of folds and you're going to have
- 27:32:43to do this over and over again. Um, and
- 27:32:45if it's a large data set, it might take
- 27:32:47your model a long time to train. It's
- 27:32:49going to be a little bit more uh
- 27:32:52computationally intense than if we just
- 27:32:55did a train test split.
- 27:32:57Okay, we just did a single like 7030
- 27:32:59split. We only do that once. We only
- 27:33:02train the model once, right? We train it
- 27:33:04on the 70, apply it to the 30% test data
- 27:33:08and evaluate performance that way. Um,
- 27:33:11so we're only really using the model and
- 27:33:14training the model once, but in this
- 27:33:16kfold, we're going to do it um, you
- 27:33:19know, k number of times essentially
- 27:33:23or I should say one for every
- 27:33:24combination that we have to work through
- 27:33:27of of all the folds.
- 27:33:32Okay.
- 27:33:35All right. Does that make sense? Any any
- 27:33:38questions on Kfold cross validation? So
- 27:33:41K K K K K K K K K K K K K K K K K K K K
- 27:33:42K K K K K K K K K K K K K K K K K K K K
- 27:33:42K is an important uh number here. It
- 27:33:45it's how many folds, how many splits do
- 27:33:48you have? A typical value for K is going
- 27:33:50to be somewhere like five or 10.
- 27:33:54So 10 folds or five folds. Those are
- 27:33:57pretty pretty standard
- 27:34:01from what from what I've seen.
- 27:34:05But does the does the concept make sense
- 27:34:07or is there any questions on it on in
- 27:34:09terms of um you're always going to leave
- 27:34:11one fold out. You're going to split it
- 27:34:13up into K number of folds. Always leave
- 27:34:15one out. Train on the rest of it.
- 27:34:18Evaluate on that one that gets left out
- 27:34:19and then rotate those through. And
- 27:34:22you're going to do that for every
- 27:34:23combination and average all those
- 27:34:25metrics.
- 27:34:32And by the way, there's going to be an
- 27:34:33easy function in scikitlearn that will
- 27:34:36do this for us. So managing all these
- 27:34:38combinations will be really easy. It's
- 27:34:41actually just built into scikitlearn. So
- 27:34:43we don't have to um we don't have to do
- 27:34:46this all by hand. Okay, this will be in
- 27:34:48scikitlearn. It'll handle doing all
- 27:34:50these combinations of folds for us and
- 27:34:53computing the average metric will be
- 27:34:55really easy. So um
- 27:34:59we don't have to worry about that. We're
- 27:35:00going to see an example of this coming
- 27:35:02up shortly.
- 27:35:04All right, of kfold cross validation.
- 27:35:08But this is a this is a really widely
- 27:35:10used technique. And again like the
- 27:35:12purpose you may be wondering like what's
- 27:35:13the purpose ultimately of doing this?
- 27:35:15It's to get a sense of if our model is
- 27:35:18going to perform well on new data.
- 27:35:20That's really what we want to know. like
- 27:35:22is the model going to perform well when
- 27:35:23I start to use it on new data that it's
- 27:35:26never seen before and this kffold is a
- 27:35:30decent indicator of that because we are
- 27:35:34varying which data it sees across many
- 27:35:37different folds right so it's a it's
- 27:35:40kind of a good um proxy to exposing it
- 27:35:44to different kinds of data each time and
- 27:35:46seeing how it performs
- 27:35:49right all right because we're working
- 27:35:50our way through each one of the folds
- 27:35:51there's always going to be one fold left
- 27:35:53out. We're going to change which fold
- 27:35:55gets left out each time. And um that's
- 27:35:58sort of mimicking the idea of we're
- 27:36:00going to apply our model to new data and
- 27:36:02see how it performs. And it's it's new
- 27:36:05data every fold.
- 27:36:25um how we know which model is best suits
- 27:36:29for which scenario because we have Yeah,
- 27:36:32that's a good question. Um, so my we're
- 27:36:36going to learn this as we go along
- 27:36:38because we haven't covered all the
- 27:36:39models yet, but generally the best
- 27:36:43advice I can give on that is
- 27:36:46you you generally want to start as
- 27:36:49simple as you can get and then if it's
- 27:36:51not performing well then work your way
- 27:36:53up to something more complex.
- 27:36:56So we are going to have models that are
- 27:36:58simpler. We're going to have models that
- 27:37:00are more complex. The rule of thumb is
- 27:37:02to start with the most simple model that
- 27:37:04works.
- 27:37:07So you're usually going to have the same
- 27:37:10ones that you're going to try in the
- 27:37:12beginning. And linear regression is a
- 27:37:14very simple model. It's usually the
- 27:37:16first one you want to try for regression
- 27:37:18because it's the simplest.
- 27:37:20Um, and for classification, we're going
- 27:37:22to have a similar like logistic
- 27:37:24regression is the simplest kind of
- 27:37:26classification model we could have. So
- 27:37:29usually want to start with that and then
- 27:37:31if it underfits like if we see it's
- 27:37:34producing a lot of error then we work
- 27:37:37our way up to a more sophisticated
- 27:37:39model.
- 27:37:41So um that's the way we that's the way
- 27:37:45it should usually go is simple to
- 27:37:47complex it based on their performance.
- 27:37:50So we evaluate it and then we can repeat
- 27:37:52the process. If it's not performing well
- 27:37:53we can try something different that's
- 27:37:55more complex if it's underfitting.
- 27:38:05Uh this is a good question. Does a model
- 27:38:06reset after training each k minus one
- 27:38:09fold? Um yeah, it's essentially like a
- 27:38:12blank model every time uh every fold. So
- 27:38:15um we imagine like you have a brand you
- 27:38:19have a fresh model every um k minus one
- 27:38:22combination. Yes.
- 27:38:30And the reason the reason it has to be
- 27:38:32that way is because you don't want the
- 27:38:35other folds influencing the model that
- 27:38:39like on on the next combination. You
- 27:38:42don't want the previous combination to
- 27:38:43influence the results on the next one,
- 27:38:45right? Um you want it to be a fresh
- 27:38:48evaluation on every combination of
- 27:38:51folds.
- 27:39:10Okay.
- 27:39:12All right. So, let me describe to you a
- 27:39:15variation on what we just um talked
- 27:39:18about with the K-fold. So, there's
- 27:39:20another cross validation known as
- 27:39:22stratified K-fold. And um this is the
- 27:39:26same exact procedure as kfold except
- 27:39:29that when we this is used for
- 27:39:31classification.
- 27:39:32Um so when we do classification
- 27:39:35uh we want to make sure that the
- 27:39:37different categories are going to be um
- 27:39:40split amongst those folds in a
- 27:39:43proportional way. So we don't what we
- 27:39:45don't want to happen is um when we split
- 27:39:48apart the data. So, let's say we have
- 27:39:50let's say we're predicting um spam not
- 27:39:53spam. What we don't want to have happen
- 27:39:56when we do our splits is we don't want
- 27:39:58to have all of the spams end up in one
- 27:40:01fold and then every other fold has no
- 27:40:04spam, no spam, no spam, no spam, right?
- 27:40:08That's not very good. Um because if we
- 27:40:11if we train against all these guys, we
- 27:40:13have no shot at predicting spam when
- 27:40:16they've never seen spam before. So
- 27:40:18stratify kayfold is is used in
- 27:40:20classification
- 27:40:23and it's to um it's to make our splits
- 27:40:27ensure that they have basically a
- 27:40:29balanced number of categories for each
- 27:40:32split. Um so that we don't end up with
- 27:40:34certain splits with way more spams than
- 27:40:37not spams. Um so we we do what's called
- 27:40:40stratifying where we make sure the
- 27:40:42proportions are balanced across each uh
- 27:40:45split. So this is only really useful in
- 27:40:47classification, not really necessary in
- 27:40:50regression because we're predicting a
- 27:40:51value. But if we were predicting a
- 27:40:53category,
- 27:40:55like in classification like fraud, not
- 27:40:58fraud, we don't want to do the split and
- 27:41:00have every single fraud example um by
- 27:41:03bad luck in our shuffling and split end
- 27:41:05up in one split and every other um every
- 27:41:09other split has no examples of fraud.
- 27:41:11Right? So we want to stratify this to
- 27:41:13spread out those um frauds against all
- 27:41:16the other splits. Um so uh again um
- 27:41:21scikitlearn will take care of that for
- 27:41:23you. Um but if you're doing
- 27:41:24classification and you have an
- 27:41:26imbalanced data set um you you really
- 27:41:29want to make sure you stratify k-fold.
- 27:41:32um imbalanced meaning that you have a a
- 27:41:36um different number. Like if you're
- 27:41:38doing fraud, not fraud, you have way
- 27:41:39more not frauds than frauds. Um where
- 27:41:42where that category is imbalanced,
- 27:41:45you want to make sure it's balanced
- 27:41:46across all your splits.
- 27:41:49Um so this is this is useful in
- 27:41:52classification only, not really
- 27:41:53regression, which is what we're talking
- 27:41:54about right now. Um but it's just a
- 27:41:57variation on this that ensures when we
- 27:41:59do those folds um the data is
- 27:42:02distributed evenly amongst those folds
- 27:42:04as much as we can. The labels are I
- 27:42:06should say.
- 27:42:08Okay. So that's stratified kfold. It's
- 27:42:12the same same procedure once we have our
- 27:42:14splits. It's the same where we do k
- 27:42:16minus one of them. We train test on that
- 27:42:18last fold um and then rotate through all
- 27:42:22the folds and and average all the
- 27:42:24metrics. the same exact procedure. It's
- 27:42:26just the splitting itself um is going to
- 27:42:29be balanced in a stratified kfold.
- 27:42:35Okay. So, hold out we've already talked
- 27:42:36about um is just doing a single train
- 27:42:39test split. We've talked about that. One
- 27:42:42more variation that is a bit of an
- 27:42:44extreme version of K-fold. So it's
- 27:42:46actually the same process as Kfold, but
- 27:42:48it's an extreme version is if you set K
- 27:42:52equal to the number of data points. So
- 27:42:54you basically are um this is a really
- 27:42:57really extreme kfold where you um
- 27:43:00basically are training on all the data.
- 27:43:03Um so you're training on all the data
- 27:43:06except one point and then you test
- 27:43:10against that one point. Um now why would
- 27:43:13you ever do this? Um it's mainly so for
- 27:43:16this reason here. It's to um maximize
- 27:43:20the amount of training data that your
- 27:43:22model gets exposed to because instead of
- 27:43:24just doing instead of just doing five
- 27:43:26splits
- 27:43:28um which would be like
- 27:43:31you know these four folds are going to
- 27:43:33be used and then we um test against one
- 27:43:36fold. um we're essentially going to use
- 27:43:3999% of the data, right? One point is
- 27:43:42going to be left out. 99% of the data
- 27:43:45gets used to train. Um and then we're
- 27:43:47always going to leave out one point. And
- 27:43:49and the issue is we're actually going to
- 27:43:51do that over and over and over again and
- 27:43:53rotate that one point to cover the whole
- 27:43:55data set. So, we're going to train on
- 27:43:5899%, leave one that one point out,
- 27:44:01and then rotate through every
- 27:44:03combination of points until we've left
- 27:44:05out every single point, and then average
- 27:44:08all those together. Um, so this is a
- 27:44:10this is an extreme kfold. Again, the
- 27:44:13number of folds is actually equal to the
- 27:44:15number of data points in this case. So,
- 27:44:16we have every point is its own fold and
- 27:44:19we train on everything but one. Test on
- 27:44:23that one. This gets you the maximum size
- 27:44:26of your training data because you're
- 27:44:28basically gonna have every point but one
- 27:44:30used in the training.
- 27:44:32This gets you the maximum size. However,
- 27:44:34it gets you the maximum uh expense
- 27:44:38especially for large data sets. This is
- 27:44:40going to be usually you're not going to
- 27:44:42use this um especially for large data
- 27:44:44sets because it's just too extreme. It's
- 27:44:48going to take you a really long time to
- 27:44:49work through every single point being
- 27:44:52left out. um it's just going to take a
- 27:44:55while to do.
- 27:44:57So, for that reason, the leave one out
- 27:45:00um that that's why it's called leave one
- 27:45:02out because it's you're leaving one out
- 27:45:04every single time. Um is rarely used. I
- 27:45:08I have don't really see it used that
- 27:45:09often, but it is an extreme version of
- 27:45:12kful cross validation.
- 27:45:16Okay. But rarely ever actually used. I
- 27:45:19think the the ones that get used the
- 27:45:20most are definitely the hold out method
- 27:45:22with just a regular train test split. Um
- 27:45:25and then uh the other one that gets used
- 27:45:27quite a bit is is kfold
- 27:45:30or stratified kfold if you're if you're
- 27:45:32doing classification,
- 27:45:34but certainly kfold in the in a
- 27:45:36regression case.
- 27:45:38Okay.
- 27:45:43All right. Um we're going to do an
- 27:45:45example with these guys. So we'll do
- 27:45:47that next.
- 27:45:48um with with the different cross
- 27:45:50validation techniques. Um but any
- 27:45:54questions on what they are doing
- 27:45:57conceptually before we actually do the
- 27:45:59code example.
- 27:46:16Okay.
- 27:46:18Very good.
- 27:46:25All right. So, let's see some examples.
- 27:46:27Um, let's go into our code and build a
- 27:46:32model and do the different cross
- 27:46:34validation techniques on it. Um, you're
- 27:46:37going to see it's actually going to be
- 27:46:38really easy to do and we it sounds
- 27:46:40complex like doing the kfold and leaving
- 27:46:43one out and testing. It sounds kind of
- 27:46:45complex, but I promise you scikitlearn
- 27:46:47makes it really easy to do. Um,
- 27:46:51and so, uh, we won't need to do too much
- 27:46:54besides just use the right, uh, tools
- 27:46:57from scikitlearn. Uh, so we're going to
- 27:46:59we're going to see that. Um, so here we
- 27:47:01have some imports. The, um, primary, uh,
- 27:47:05thing that's a little bit new for us is
- 27:47:07going to be these, um, different kinds
- 27:47:09of cross validation techniques. So we
- 27:47:11have our kfold, we have our stratified
- 27:47:13kfold, leave one out. Um, which are
- 27:47:15those different cross validation
- 27:47:17techniques. Um, these are going to be
- 27:47:19used in combination with this cross val
- 27:47:24score which is going to keep track of
- 27:47:27the different um metrics and then
- 27:47:29average them
- 27:47:31uh while we do one of these um cross
- 27:47:35validation techniques. So this guy gets
- 27:47:38used in combination with one of these to
- 27:47:42um as as we're going to see in the code
- 27:47:44uh to average those metrics um doing the
- 27:47:48different folds, right? Perform doing
- 27:47:49performance against the different folds.
- 27:47:52Okay. And then of course we need a model
- 27:47:55using a linear regression. That's that's
- 27:47:57the one we've studied so far. Um and
- 27:48:00then we have just a regular metrics. If
- 27:48:02we want to compute those um using maybe
- 27:48:05just hold out, right? And hold out um
- 27:48:08which which is just a regular train test
- 27:48:10split um we could use these guys to
- 27:48:12evaluate performance.
- 27:48:14But in a more sophisticated kfold style
- 27:48:17of cross validation, we're going to use
- 27:48:19this to evaluate the the performance.
- 27:48:25Okay, let's see.
- 27:48:28So, we're going to be working with this
- 27:48:30housing with ocean proximity data. Um,
- 27:48:33you guys should have this one. Uh,
- 27:48:37so you guys should have this one. So, if
- 27:48:40you want to follow along and run it
- 27:48:41yourself, um, you can load that one in.
- 27:48:46Um, I want to make sure that I have it.
- 27:48:51Let me pull that one in. So, it should
- 27:48:52be this guy.
- 27:49:03I'm going to load that in so I can make
- 27:49:05sure I run it with you guys.
- 27:49:08Um,
- 27:49:17let me run this.
- 27:49:21Do you guys have that data?
- 27:49:24the housing with ocean proximity.
- 27:49:28It's another it's another housing data
- 27:49:30set. Um
- 27:49:33but it it's a little bit different than
- 27:49:35the ones we've seen before. It has a a
- 27:49:38special feature for how close it is to
- 27:49:40the ocean at different locations.
- 27:49:50So it looks kind of like this. If we
- 27:49:52load it in and do our head, which is
- 27:49:54usually what we do, right? We can see um
- 27:49:57we can see that it's got these features.
- 27:50:00So it's got uh uh bedrooms, total rooms,
- 27:50:05um it's got uh median age. Now this is
- 27:50:08this is looks a little strange for total
- 27:50:10rooms and um uh bedrooms and population
- 27:50:15etc. But it's um
- 27:50:19it's it's got those uh it's got those
- 27:50:22because it's representing an entire
- 27:50:24neighborhood. So it's an entire
- 27:50:27neighborhood. And we're looking at this
- 27:50:29um this is actually going to be our
- 27:50:30label is this median house value for the
- 27:50:32entire neighborhood. So what's that
- 27:50:34median value uh in the neighborhood? And
- 27:50:38this is the total number of bedrooms,
- 27:50:40total number of rooms, um population,
- 27:50:43households. So, how many houses are
- 27:50:46there? Um, median income. And of course,
- 27:50:49these are scaled. So, these are um
- 27:50:51likely times, you know, uh thousands. Um
- 27:50:58but um that's our data. We could
- 27:51:02describe it.
- 27:51:07So we can see the average age, average
- 27:51:10median age. Um which sounds a little um
- 27:51:13weird, but that's it's because again
- 27:51:15this is the median of data within a
- 27:51:18neighborhood. Um so the average of those
- 27:51:22is about 28 or 29. Um we have
- 27:51:27u
- 27:51:31total bedrooms. The we can look at the
- 27:51:33men. There's some data that only has
- 27:51:36one. So, it's likely only one house in
- 27:51:38there. Um, which is what this
- 27:51:41represents. There's only one house. So,
- 27:51:43there there is some neighborhood that
- 27:51:45only has one house. Um, and we see the
- 27:51:48median um we see the minimum uh median
- 27:51:53house values there. And then the maximum
- 27:51:55down here um is a pretty big number.
- 27:52:016,000 households is the largest that we
- 27:52:03have in any any one of these
- 27:52:05neighborhoods.
- 27:52:07Okay. So, just a little bit of
- 27:52:09description of the data.
- 27:52:25Okay. So, then we can run.info. So, this
- 27:52:28is um let me ask you guys, were you able
- 27:52:30to load this? Were you able to run this?
- 27:52:36If you're following along, were you able
- 27:52:38to
- 27:52:39load it and take a look at
- 27:52:49Okay, great. Great.
- 27:52:52Okay, so we're able to load that and
- 27:52:55then look at head. Perfect. Um
- 27:53:00Okay.
- 27:53:02Um and then we run describe which gives
- 27:53:05us that uh usual kind of statistical
- 27:53:07description. Uh so we can see some
- 27:53:10interesting stats about those.
- 27:53:18What do you guys notice about the info?
- 27:53:20Anything interesting that we see from
- 27:53:22there?
- 27:53:36Is there any missing data
- 27:53:42any features that have missing data? Can
- 27:53:44we see
- 27:53:47object? Yeah, object type usually is
- 27:53:49string. If it's an object type, that
- 27:53:51usually means string. Python when we
- 27:53:54read it into pandas it usually is just a
- 27:53:56string.
- 27:53:59So that that makes sense like we have
- 27:54:00mostly numerical features but then we
- 27:54:03have a this ocean proximity which is a
- 27:54:05string.
- 27:54:14Yeah. Total bedrooms has nles. That's
- 27:54:16right. Because you can see here this
- 27:54:18does not equal the number of uh rows
- 27:54:21that we have. So this is the number of
- 27:54:23rows which about 20,000 rows. That's a
- 27:54:24good size data set, right? 20,000 rows.
- 27:54:27That's decent. Um we're definitely
- 27:54:30missing some data here for sure. Um we
- 27:54:34could count how much we're missing
- 27:54:35exactly by running this is NATO sum. Um
- 27:54:40and so we see that total bedrooms is
- 27:54:42missing about 200 uh 200 rows are
- 27:54:46missing total bedroom uh value.
- 27:54:52Okay. And then one thing I wanted to
- 27:54:54look at is yes, this is a string. So
- 27:54:56what remember what we can do with those?
- 27:54:58That's a categorical.
- 27:55:01So ocean proximity
- 27:55:06is a categorical
- 27:55:09string
- 27:55:11feature.
- 27:55:13So we can take a look at its value
- 27:55:15counts, which is usually a good idea to
- 27:55:17take a look and see what possible values
- 27:55:20that feature could be. So if we look at
- 27:55:23our
- 27:55:25um what are we calling this? Housing
- 27:55:27data.
- 27:55:31housing data
- 27:55:34ocean
- 27:55:36proximity
- 27:55:42value counts.
- 27:55:49So, here's the different types that that
- 27:55:51one can be. So, there's some
- 27:55:53neighborhoods that are less than 1 hour
- 27:55:54from the ocean. There's some that are
- 27:55:56inland. There's some that are near the
- 27:55:58ocean. There's some that are near a bay.
- 27:56:01There's even five of them that are on an
- 27:56:03island. So, these are the different
- 27:56:05values of the ocean proximity. So,
- 27:56:08remember, you can always do that. If you
- 27:56:09see a string feature, you can always
- 27:56:12take a look at what its um categories
- 27:56:14are. And it looks like most things are
- 27:56:17less than 1 hour from the ocean, but
- 27:56:18it's kind of evenly distributed here. Um
- 27:56:22otherwise
- 27:56:25very few islands.
- 27:56:33But as you can imagine like this feature
- 27:56:34is probably going to be important for
- 27:56:36determining um what the value is, right?
- 27:56:40Probably going to be important.
- 27:56:47Okay. So, um, we need to deal with these
- 27:56:51NLES. If we're going to build a model,
- 27:56:53right? So, um, this is all of our
- 27:56:55typical data prep. If we want to build a
- 27:56:57model, we're going to have to deal with
- 27:56:58these NLES. What do you guys think we
- 27:57:00should do with the NLES? What would you
- 27:57:03what do you think for total bedrooms?
- 27:57:05What do you think is a good strategy to
- 27:57:06do? Keep in mind, we have 20,000 points,
- 27:57:1220,000 rows I should say, and about 200
- 27:57:15of them are null.
- 27:57:22Right. So about 200 are null. Um so what
- 27:57:27do you what do you guys think would be
- 27:57:28like a good strategy to deal with those
- 27:57:29NLES in that case?
- 27:57:51average. We can't ignore it because we
- 27:57:55can't ignore that column.
- 27:57:58We can't ignore the whole column. So,
- 27:58:00something needs to go there.
- 27:58:09Probably don't want to make it zero.
- 27:58:11I think average is a decent average is a
- 27:58:14decent idea. Probably don't want to make
- 27:58:16it zero because um that would indicate
- 27:58:19that there's no bedrooms and yet we
- 27:58:21still have a bunch of total rooms. So it
- 27:58:24probably doesn't make sense to do zero.
- 27:58:30Average, I think average could be a
- 27:58:32decent one.
- 27:58:34Now in this example, what we're actually
- 27:58:37going to do is we're
- 27:58:41rows.
- 27:58:43We're actually going to drop the rows al
- 27:58:45together. Now, why are we doing that?
- 27:58:46It's because we have so much data and
- 27:58:50only 200 of them are null.
- 27:58:54Okay, only 200 of them are null. So,
- 27:58:56we're actually just going to drop the
- 27:58:57rows. Now, that's a choice.
- 27:59:02Um, that's a choice, right? Is that we
- 27:59:05could fill in with the average like you
- 27:59:07guys are suggesting. What we're actually
- 27:59:09going to do is just drop the rows. It it
- 27:59:11makes up less. It makes up about 1% of
- 27:59:15the whole data. So it's not that much of
- 27:59:18it is missing. We can drop those rows.
- 27:59:21So that's actually what we're going to
- 27:59:22do here is we remove all the roles with
- 27:59:26the NLES by doing drop NA. So this just
- 27:59:28drops them. So those rows are cut out.
- 27:59:31Um, it's arguable that we could replace
- 27:59:37it's arguable that we could just replace
- 27:59:38it with something and I think you guys
- 27:59:40have good thoughts which is the average
- 27:59:42a default
- 27:59:45um assume total bedrooms. We could we
- 27:59:47could try that. Yeah.
- 27:59:51Assign a value based on comparable home
- 27:59:53value. Yes, you could do that too.
- 27:59:54That's a good strategy is to look at the
- 27:59:57other rows that are similar to it and
- 27:59:59fill in a value. That's absolutely fair.
- 28:00:02Um, in this example, we're actually just
- 28:00:04going to drop those rows,
- 28:00:09but I think that's totally um totally
- 28:00:11valid.
- 28:00:13This is a choice.
- 28:00:16We could fill NA with different values
- 28:00:22such as the average
- 28:00:26total bedrooms
- 28:00:29um derive a value etc. So we could
- 28:00:33derive something which I think Brent you
- 28:00:35have a good suggestion that's a good
- 28:00:37suggestion. Um we could derive something
- 28:00:40like that uh and fill in the blank and
- 28:00:42that's I think that's totally valid. Um,
- 28:00:44we could take the average of the um
- 28:00:48bedrooms. Uh, I meant total rooms here.
- 28:00:51Sorry, total rooms. Um, we could fill in
- 28:00:54we could fill it in with the total rooms
- 28:00:56for that category um or for that row.
- 28:01:00Um, many options. In this case, we're
- 28:01:03actually just going to drop those rows
- 28:01:05because they make up such a small
- 28:01:07percentage relative to the 20,000 rows
- 28:01:10that we have. It's about 1%. Right? 200
- 28:01:14rows is about 1% of 20,000.
- 28:01:17So, we're just going to drop them. But
- 28:01:19that's a choice. We don't have to drop
- 28:01:21them. We could fill in with something.
- 28:01:24Um, and if we did that, we would use
- 28:01:26fill NA rather than drop NA, right?
- 28:01:42Uh after dropping the rows, how many? So
- 28:01:45it's just so after we drop the rows, um
- 28:01:48after we drop the rows, it's just going
- 28:01:50to be we still have all our other rows
- 28:01:53are intact, right? So if we look at this
- 28:01:54now,
- 28:02:03we now have um slightly uh slightly less
- 28:02:07entries.
- 28:02:09So now we have this this many um rather
- 28:02:12than rather than this many,
- 28:02:16right? We dropped those 200
- 28:02:25But they're all filled in. Yeah, they're
- 28:02:27So all the other columns are still
- 28:02:28filled in. We're just we're we're
- 28:02:30cutting out the whole row. So if you
- 28:02:32think about our data set, um we have all
- 28:02:35these rows and all these columns. What
- 28:02:38we're doing is like if there's a null
- 28:02:40here, we're just we're just getting rid
- 28:02:42of that whole row, right? And so we
- 28:02:45still have all the other rows intact.
- 28:02:54Uh, we can drop them because we have a
- 28:02:56good sample size. Yes,
- 28:03:02that's exactly right, Ronald. Yep, we
- 28:03:04can drop them because we have we have
- 28:03:0520,000 rows and only 200 are missing
- 28:03:08values. So, that's totally fine.
- 28:03:15Uh, drop a removes all rows that has any
- 28:03:17null. Yes, that's true. It it will go
- 28:03:20ahead and just drop any row where
- 28:03:22there's any null, no matter what column
- 28:03:24it's in. Yes,
- 28:03:35index. Yeah, the index is not getting
- 28:03:37reset. Um, that's true. So, um, what we
- 28:03:43what you can always do is you can reset
- 28:03:45the index. So, um, if you want to, it's
- 28:03:48optional. We we're not really going to
- 28:03:50use the index for anything that
- 28:03:52important, right? But what we could do
- 28:03:54is, uh, reset index.
- 28:03:59Uh,
- 28:04:01we could do that, right? Which will
- 28:04:02reset it.
- 28:04:17So now now it gets reset.
- 28:04:36But um let me actually I don't I don't
- 28:04:39really want to do that. I'm going to
- 28:04:41reset this.
- 28:04:48Um,
- 28:04:57yeah, we could do that.
- 28:05:18Okay.
- 28:05:20So now importantly there should be uh no
- 28:05:23missing data of this of this new one
- 28:05:25where we've dropped NAS. Right. So now
- 28:05:27this is good. If you now the reason we
- 28:05:29had to do this is because if we try to
- 28:05:31build a linear regression and we have
- 28:05:33NLES in there. Um the the issue is like
- 28:05:37how do you build a model where you have
- 28:05:40something like this
- 28:05:48and these are null? Like what do how do
- 28:05:51you multiply a number by a null?
- 28:05:54Um we can't really do that, right?
- 28:05:58we can't really do that. So, um,
- 28:06:07so therefore, uh, we need to get rid of
- 28:06:10NLES like the the null is not really
- 28:06:12going to work in there. So, uh, we need
- 28:06:15to get rid of them for linear regression
- 28:06:17to to really have a chance to work,
- 28:06:18right? To train it and be able to use
- 28:06:20it.
- 28:06:22You got to get rid of those nles.
- 28:06:30All right,
- 28:06:33any questions so far? So, we haven't
- 28:06:34done any modeling yet. We're doing some
- 28:06:35We're doing some data preparation before
- 28:06:37we get to the modeling. And we haven't
- 28:06:39done any cross validation yet. We
- 28:06:41haven't set that up. We're just doing
- 28:06:42our data preparation before we get to
- 28:06:44the modeling. Right? So, we've dropped
- 28:06:47some NAS. We've checked it. Um, we're
- 28:06:50going to do one more prep step, which is
- 28:06:52to um change that ocean proximity
- 28:06:57feature into something numerical because
- 28:06:58again, how do you build a model where
- 28:07:00you're inserting a string into those
- 28:07:03like beta 1, beta 2, beta 3 times of
- 28:07:05features? You can't really do that when
- 28:07:07it's a string. Um, so what we're going
- 28:07:10to do, I'm going to get rid of this
- 28:07:12because I don't think we really need
- 28:07:13that. um is we are going to uh run this
- 28:07:18get dummies function which is our um our
- 28:07:22get dummies function is our usual one to
- 28:07:26uh our git dummies one is our usual one
- 28:07:29to um
- 28:07:32uh get our one hot encoding.
- 28:07:34So this is our uh one hot encoding here.
- 28:07:42We now are going to have data that's
- 28:07:45like this, right? So we have ocean. So
- 28:07:48So by the way, this prefix
- 28:07:50um this prefix is OP, which which is
- 28:07:54short for ocean proximity, right? So we
- 28:07:56have ocean proximity uh less than 1 hour
- 28:07:59from the ocean, ocean proximity inland,
- 28:08:02ocean proximity island, near bay, near
- 28:08:04ocean. So these first five rows are near
- 28:08:07the bay. Um so they have a one there and
- 28:08:10a zero in the other spots. So this is
- 28:08:12good. This one hot encodes that feature
- 28:08:15into these numerical uh values,
- 28:08:19right?
- 28:08:21Were you guys able to run that one? they
- 28:08:24get dummies.
- 28:08:31So the reason that Yeah, that's a great
- 28:08:33question. How did it go ocean proximity?
- 28:08:35It's because um that is the only uh
- 28:08:38string feature we have. That's the only
- 28:08:41one we have. So it it's going to look
- 28:08:43for any non-numericals and one hot
- 28:08:45encode those however many however many
- 28:08:47there are. So whatever objects we have
- 28:08:51which are strings, it's going to
- 28:08:52automatically oneh hot encode those.
- 28:09:03Yeah, we could have Right. We could have
- 28:09:05went here and did Right. We could have
- 28:09:07done ocean
- 28:09:11proximity,
- 28:09:13but we only have one of those features.
- 28:09:17So it's just going to do that to the
- 28:09:18whole data frame
- 28:09:20uh on that one feature. So what we're
- 28:09:23going to do is um go ahead and split it
- 28:09:26into an x and a y um which the x is
- 28:09:30always what includes our features. The y
- 28:09:34is what we are trying to predict which
- 28:09:35is the label. Now, um, in order to
- 28:09:40separate those out, what we're going to
- 28:09:41do is assign X to be the variable that
- 28:09:44is, um, our data frame minus this median
- 28:09:48house value column. So what this is
- 28:09:50doing is um uh it's not permanently
- 28:09:55dropping because we're not uh dropping
- 28:09:57it in place but it is returning us a
- 28:10:01copy of the data frame with the median
- 28:10:04house value column left out right it's
- 28:10:06dropped. So this is this is uh something
- 28:10:09we want to do because that will the rest
- 28:10:11of it will contain our features right.
- 28:10:14So, um this will temporarily or I should
- 28:10:18say return a copy of the DF with um
- 28:10:24median house value
- 28:10:28dropped,
- 28:10:30right? Median house value dropped. Um so
- 28:10:33we go ahead and drop that one. Uh now
- 28:10:37remember it's not permanent. It's just
- 28:10:38giving us uh the remainder of it which
- 28:10:41is this housing data. dropping this and
- 28:10:44it's assigning that to X and then we're
- 28:10:46taking the actual median house value
- 28:10:49column from the original data and
- 28:10:52assigning that to Y. So this is going to
- 28:10:53be our labels,
- 28:10:57right? So this is what we are trying to
- 28:11:02predict.
- 28:11:05Okay, so that is our Y and that's always
- 28:11:07how it is. X is our features, Y is our
- 28:11:10labels. Um hopefully that makes sense.
- 28:11:12What this is doing is this is going to
- 28:11:15get rid of that label column and
- 28:11:17everything else will be our features and
- 28:11:19then this will get rid of this will just
- 28:11:22assign the label column to Y.
- 28:11:26All right. And then what we can do is
- 28:11:28pass X and Y into our train test split
- 28:11:32function and this will generate the hold
- 28:11:34out set. So if we want to do the hold
- 28:11:36out cross validation this is how we
- 28:11:39would do it is we would split the data
- 28:11:41into X train X test Y train Y test um
- 28:11:45using train test split. So this is what
- 28:11:48we did last time. This would be this
- 28:11:51would be for hold out cross validation
- 28:11:58right where we are uh uh just have that
- 28:12:03one one set for testing one set for uh
- 28:12:06one set for training one test one set
- 28:12:08for testing I should say right so this
- 28:12:11is pretty standard train test split um
- 28:12:14we pass in that x we pass in the y we
- 28:12:16use a 30% test size which pretty
- 28:12:18standard
- 28:12:19and random state so that we get the
- 28:12:21consistent shuffling if we were to run
- 28:12:23this multiple times. Um we we get that
- 28:12:26uh consistent randomization.
- 28:12:30Okay,
- 28:12:31so we have that and so now our X train
- 28:12:36is a percentage um of the data frame of
- 28:12:39the 20,000 uh rows and the X test is uh
- 28:12:4430% of that. So it's only about 6,000
- 28:12:46rows, which is what um the shape of that
- 28:12:48is.
- 28:12:52Yeah. X. So X is our features. So we're
- 28:12:55we're putting all of our data in that is
- 28:12:57our features into X. And so the the um
- 28:13:01most efficient way of doing that is um
- 28:13:05the most efficient way of doing that is
- 28:13:06to
- 28:13:08uh just take our data and drop the
- 28:13:12median house value column because that's
- 28:13:13our label column. So we just remove
- 28:13:16that. The rest of the data is our
- 28:13:18features. So that's what that's what
- 28:13:20this X is, right? It's all of our
- 28:13:22feature data. All of our columns that is
- 28:13:24not the label column essentially is what
- 28:13:27that's doing. And then Y is our label
- 28:13:30column from our original data,
- 28:13:34right? Y is our label column. And so
- 28:13:38this this will um contain all of our
- 28:13:41labels which is the median house value.
- 28:13:44X X contains every column but the one
- 28:13:47we're going to so we we ultimately
- 28:13:50decide that but X contains um X is
- 28:13:54everything that is not our dependent
- 28:13:57variable which is what we're predicting.
- 28:14:00So we're removing what we are trying to
- 28:14:02predict from X. X should be everything
- 28:14:04else. That's always how it's going to
- 28:14:06be. X is X is always going to be all of
- 28:14:09those independent variables that we're
- 28:14:12using to predict the median house value.
- 28:14:15So we are going to predict the median
- 28:14:18house value. We need to remove it from
- 28:14:20X.
- 28:14:22So we're we're taking everything but
- 28:14:24that column.
- 28:14:29So it's the whole data frame. It's the
- 28:14:32whole data frame minus this one column
- 28:14:35with just the dependent variable. Right.
- 28:14:39Exactly right. Removing the dependent
- 28:14:41variable and keeping all the
- 28:14:42independence. That's exactly right.
- 28:14:44Exactly right. So think about it in
- 28:14:47terms of the model. Let's go back to the
- 28:14:49features. Right. Think about it in terms
- 28:14:50of the model. We are trying to predict
- 28:14:53this this value. We're building a model
- 28:14:57to try to predict this. So we are going
- 28:15:00to make sure x is everything but this
- 28:15:04right. So this is actually just y.
- 28:15:07That's our label. That's our dependent
- 28:15:09variable. Right? That's y. Everything
- 28:15:12else is belongs to x. Everything else
- 28:15:15belongs to x including all of these.
- 28:15:22Right? We choose this one to be y
- 28:15:24because we're building a model to
- 28:15:26predict that. That's our label.
- 28:15:29All right. So, we have our we use X and
- 28:15:32Y to do our train test split. So, we
- 28:15:34have our our training features and our
- 28:15:37test features and then our training
- 28:15:38label and test labels here. Um, pretty
- 28:15:42standard there.
- 28:15:44Um, okay. So, this is what's new is if
- 28:15:47we want to do k-fold uh validation, what
- 28:15:50we're going to do is create a kfold
- 28:15:52object. So, we have this kfold from
- 28:15:54scikitlearn that we already imported. we
- 28:15:57are going to create a kfold um where we
- 28:16:02are going to specify how many folds we
- 28:16:04want. So that is the in uh inslits
- 28:16:07parameter as this says um this is going
- 28:16:10to be uh uh in this case we're going to
- 28:16:14do 10 folds. That's pretty standard. So
- 28:16:16I think the typical number of folds that
- 28:16:18I've seen and I've worked with in my in
- 28:16:20my career is usually five or 10.
- 28:16:24Five or 10 folds is the standard.
- 28:16:29Okay. So, we're doing 10 folds in this
- 28:16:31case and we're setting a random state
- 28:16:33because we're going to do shuffling. So,
- 28:16:35in order to produce those folds, we're
- 28:16:37going to shuffle the data first and then
- 28:16:39split it into five folds, right? So,
- 28:16:42this this kffold object is going to
- 28:16:44manage creating these splits for us,
- 28:16:48right? These even splits. I know I I
- 28:16:50didn't draw it even, but um it's going
- 28:16:53to manage these five folds for us and
- 28:16:55it's going to shuffle the data and
- 28:16:57assign them to these different folds and
- 28:16:59we're and then what we're going to do is
- 28:17:01use those to do our training.
- 28:17:04We're going to execute the cross
- 28:17:05validation using this kfold object.
- 28:17:09Okay, so we create the kfold
- 28:17:12um we initialize our model as well. So,
- 28:17:15of course, in order to train something
- 28:17:18uh in the K-folds, we're going to need a
- 28:17:20model. In this case, we're using linear
- 28:17:22regression, right? Which is which is the
- 28:17:24model we've been studying so far. So,
- 28:17:26you have a linear regression. Um now,
- 28:17:29look how easy it's going to be in order
- 28:17:31to execute cross validation. All we need
- 28:17:34to do is um all we need to do is create
- 28:17:38a cross file score function
- 28:17:42um or I should say use the cross file
- 28:17:44score function from scikitlearn. So we
- 28:17:46use that with the model we want to
- 28:17:48train. So our model goes first. So
- 28:17:51that's the linear regression object.
- 28:17:54Then our data. So our extra our features
- 28:17:57and our label for our training.
- 28:18:00And then um let me skip over this for a
- 28:18:03second. I'll explain what this is in a
- 28:18:05second. Um but then we are using uh the
- 28:18:10cross validation technique is our
- 28:18:12K-fold. So this is where our K-fold
- 28:18:14object goes in the CV parameter which is
- 28:18:17cross validation. So what cross
- 28:18:20validation strategy are you using? We're
- 28:18:21using Kfold and the K-fold we're using
- 28:18:24is this one we defined up here KF. So
- 28:18:27we're putting that right here for this.
- 28:18:29And then um in jobs um allows us to
- 28:18:34parallelize this. So if we set it to
- 28:18:36negative one that's the that that's the
- 28:18:38default um it will do it will actually
- 28:18:41train across the different combinations
- 28:18:43in parallel um which speeds it up. So
- 28:18:46you want to you want to keep this to
- 28:18:48negative one if you can. So um now let
- 28:18:52me describe the scoring. So what this
- 28:18:55means is we put in our metric here. Um
- 28:19:00and so you can put mean absolute error,
- 28:19:03you can put in mean squared error. Um
- 28:19:06those are the two that we can use. And
- 28:19:09um the reason we it has a negative in
- 28:19:11front of it is because we want to find
- 28:19:15the one that has the lowest score.
- 28:19:19That's going to be our best model is the
- 28:19:21one that has the lowest score. So, we
- 28:19:24take the absolute value.
- 28:19:26I'm sorry. We take the abs the the the
- 28:19:29metric and we take the negative of it.
- 28:19:32Um because the highest scoring one is
- 28:19:36going to be the closest to zero. Um so
- 28:19:39it's just a we use the we use the
- 28:19:41negative of the of the metric. Um
- 28:19:44because on the number line like the the
- 28:19:47highest um scoring one should be the
- 28:19:50least um or I should say the maximum
- 28:19:53negative that we can get. That's going
- 28:19:55to be closest to zero. So if here's
- 28:19:57zero, this will be like -1 is better
- 28:20:00than -10. Right? So something that
- 28:20:04scores um the maximum negative uh
- 28:20:08absolute error would be closest to zero.
- 28:20:12And something that has more is going to
- 28:20:14be on this side.
- 28:20:17So this is only the reason we need this
- 28:20:19is only just to keep track of the scores
- 28:20:22of each individual um fold. Okay.
- 28:20:27So the one so the reason we can do that
- 28:20:29is at the end we can kind of see which
- 28:20:31which combination performed the best. um
- 28:20:35it's going to be the one that has the
- 28:20:37highest uh highest value of the negative
- 28:20:41which is closest to zero.
- 28:20:48That's just a convention.
- 28:20:51Yeah, it's just because um it's because
- 28:20:54the cross validation is looking to
- 28:20:56maximize the metric. So whatever has the
- 28:20:58best score
- 28:21:00um whatever has the best score is
- 28:21:03considered the best uh performance. Um
- 28:21:05but we are using uh something where
- 28:21:09lower is better. So we we take the
- 28:21:11negative and like the the highest
- 28:21:13negative would be closest to zero,
- 28:21:17right? The highest negative is going to
- 28:21:18be closest to zero.
- 28:21:22So that so it's it's just because like
- 28:21:25we want the lower score to be the best.
- 28:21:29The lowest score should be the best.
- 28:21:32So we take the negative of it. Um and so
- 28:21:36something that is more negative is going
- 28:21:38to be worse. Yeah, that's the reason.
- 28:21:43So something that's down this way is
- 28:21:45going to be worse.
- 28:21:47Okay. So it runs this
- 28:21:51and what you can see is if we actually
- 28:21:53print this out, if we print out our
- 28:21:55k-fold scores, what we should get is 10
- 28:21:57different scores.
- 28:22:01And you can see um we have 10 different
- 28:22:04uh scores here, which are all negative
- 28:22:07because we're taking the negative of the
- 28:22:08absolute of the mean absolute error. Um
- 28:22:12so what we would be looking for here is
- 28:22:17um we want to take the average of these
- 28:22:20scores but take the absolute value of
- 28:22:22them to get the best performance. So
- 28:22:24this is capturing like this is the score
- 28:22:26on the first fold combination. This is
- 28:22:29the score on the second fold
- 28:22:30combination. This is the score on the
- 28:22:32third fold combination and on and on and
- 28:22:35on. And these are the absolute errors.
- 28:22:38Okay, these are the absolute errors. Um,
- 28:22:42so if we take a look at computing the uh
- 28:22:45average, which by the way, we don't need
- 28:22:47this import because we're using the
- 28:22:49numpy average. So that's fine. Um, we
- 28:22:52can take the absolute value of those um
- 28:22:55and take a look at the average MSE
- 28:23:00or sorry MAE. Now I want you to think
- 28:23:04about this this uh average performance.
- 28:23:07So this is our performance right here on
- 28:23:09the cross validation.
- 28:23:11This is our average
- 28:23:13M AE across all of our fold
- 28:23:16combinations. So that's a that's an
- 28:23:18indicator of our performance, right? Um
- 28:23:21for the cross validation.
- 28:23:24Now what are the units of our original
- 28:23:29uh the original median value? They're
- 28:23:33already in the thousands, right? So if
- 28:23:36we go to that feature, they're already
- 28:23:40in these hundreds of thousands. So this
- 28:23:43is not a very good error. It's it's kind
- 28:23:48of high, right? Because it's in this is
- 28:23:5049,000.
- 28:23:52Um that's that's how far away we are in
- 28:23:55absolute value on average is 49,000 um
- 28:23:59dollars on the median value. That's not
- 28:24:02very good. So this score
- 28:24:06this score is
- 28:24:09um not very good. So this model is not
- 28:24:12performing that well and we can see that
- 28:24:14by comparing this error to our actual uh
- 28:24:17data. So this is right around 50,000
- 28:24:22and our median uh house values are in
- 28:24:25the hundreds of thousands. So on average
- 28:24:28we're 50,000 off when we make a
- 28:24:31prediction. That's a significant amount,
- 28:24:34right? That's a significant amount on
- 28:24:36average um when our when our data is in
- 28:24:39about the hundreds of thousands here.
- 28:24:43So we are um we have a significant
- 28:24:46amount of error 50,000 relative to the h
- 28:24:49to our units that our our data is in.
- 28:24:52Right? Um so this score is not very
- 28:24:54good. Um
- 28:24:57and so we see that from the cross
- 28:24:59validation. So look how easy the cross
- 28:25:00valid is. Again we just do cross file
- 28:25:03score. We put in our model. We put in
- 28:25:04our data. We put in our cross validation
- 28:25:08uh strategy here which is kfold. And we
- 28:25:10can generate these metrics across all
- 28:25:13the fold combinations. So it's this
- 28:25:15function is taking care of rotating
- 28:25:17those and doing every combo with just
- 28:25:19the 10 different combinations here of
- 28:25:22the of the folds.
- 28:25:2510 different instances where you have
- 28:25:26you know 10 different folds are the ones
- 28:25:28that are left out for evaluation.
- 28:25:30Um so it's managing that for us using
- 28:25:33this data right using this training data
- 28:25:35here. Um and we uh we generate these um
- 28:25:42generate these scores.
- 28:25:46Okay. So that's kf fold. It's not hard
- 28:25:48to do. All you have to do is um just use
- 28:25:52a cross file score. And we could change
- 28:25:54this to mean squared error. That's you
- 28:25:57know we could do that too. That'd be
- 28:25:58pretty easy. Um, so that'd be no issue.
- 28:26:03We just happen to be using the absolute
- 28:26:05error here. Of course, we could use
- 28:26:07squared error.
- 28:26:09Were you guys able to get this to run?
- 28:26:12K-fold scores.
- 28:26:14It produces an array of 10 10 different
- 28:26:17scores, which should make sense because
- 28:26:19those are these are the um we're
- 28:26:21splitting our data into 10 different
- 28:26:23folds,
- 28:26:25right?
- 28:26:2710 different folds. than leaving one out
- 28:26:29to do our evaluation on. So the one that
- 28:26:31gets left out every time is what's
- 28:26:33producing these scores. So it's 10
- 28:26:35different ones get left out when we
- 28:26:37rotate through all the combinations.
- 28:26:42And so we average these scores
- 28:26:51and we get this amount. We get about
- 28:26:5350,000 in error on average.
- 28:27:02Um, what do you think would be what do
- 28:27:04you think would be acceptable? So, if
- 28:27:05our if we're predicting the price, like
- 28:27:07if we're a real estate agent and we're
- 28:27:09predicting these prices and they
- 28:27:12typically are
- 28:27:14Yeah, close to zero would be great.
- 28:27:16That'd be fantastic. Closer to zero
- 28:27:18would be better. The average is um
- 28:27:20206,000.
- 28:27:23So 50,000 is a decent percentage of
- 28:27:26that. Um so you know you can compute it
- 28:27:29as a percentage right. So 50,000 is a
- 28:27:32decent percentage of that. Um probably
- 28:27:35you want this to be less than 20,000
- 28:27:38would be about 10% error. 20,000
- 28:27:43right? So maybe like 30,000 somewhere in
- 28:27:48there.
- 28:27:52Yeah. 10% would be 5% error. 10,000
- 28:27:55would be 5% error. That's true. That's
- 28:27:56true. So that would be that would be
- 28:27:59much better. So being closer to zero,
- 28:28:00like the smaller the better, of course.
- 28:28:03Of course. Um but yeah, I would say an
- 28:28:06acceptable percentage of error is
- 28:28:08probably 20%.
- 28:28:11Probably 20%, which would be um like
- 28:28:1440,000 or less would probably be
- 28:28:16acceptable.
- 28:28:19Usually when we usually when you build
- 28:28:21models um 80% accuracy is usually uh
- 28:28:26considered decent.
- 28:28:29Usually considered decent
- 28:28:3180%. So I'd say 40,000 or less would be
- 28:28:35kind of ideal.
- 28:28:40Does that make sense
- 28:28:42to answer the question?
- 28:28:45That's a good question. What value is
- 28:28:46acceptable? I think probably less than
- 28:28:4940,000 would be ideal. That's right
- 28:28:52around 20% error.
- 28:28:56All right, so that's K-fold. Um let's do
- 28:28:59just a regular hold out now. So this is
- 28:29:01just using our training and test data.
- 28:29:03Um doing model.fit and calculating an
- 28:29:06MSE on the test data. So this is this is
- 28:29:09just the um hold out strategy here where
- 28:29:12we just have um this is less robust but
- 28:29:15it's a lot quicker to do and easier to
- 28:29:17set up. Right? So um this is using the
- 28:29:22hold out strategy. So just a regular
- 28:29:26um train test split.
- 28:29:30Are we going to rebuild the model? No,
- 28:29:32not necessarily. There's some things we
- 28:29:34could do most likely. And like one thing
- 28:29:38we did not do was scale our features.
- 28:29:41Remember I said that's a pretty
- 28:29:42important thing to do is to scale our
- 28:29:45features. We did not do that. So that
- 28:29:47would be an enhancement to this that
- 28:29:49we're going to So I I actually do think
- 28:29:50we'll do that later. Yes. So I think we
- 28:29:53will actually do that now that I'm
- 28:29:55thinking about it. Yes. One of the
- 28:29:57things we can do is scale these features
- 28:29:59using like a minmax scaler or a standard
- 28:30:01scaler. that's actually going to help us
- 28:30:03um that's going to help us do better
- 28:30:06predictions.
- 28:30:10So that that's one thing we could do. Um
- 28:30:13but yeah, we will we'll try to see if we
- 28:30:16can get better.
- 28:30:18It should help it. Yeah, usually you
- 28:30:20want to scale you want to scale the
- 28:30:22data. That's something we didn't do in
- 28:30:24our preparation step. We did a lot of
- 28:30:26the things we should do. We removed nles
- 28:30:27and we did one hot encoding to the
- 28:30:30proximity feature like this one. Um
- 28:30:33those are good to do but we didn't scale
- 28:30:36any of these other we didn't scale any
- 28:30:38of the features right we didn't scale
- 28:30:40any of them. Um it you it will have an
- 28:30:44effect. It usually when we scale it
- 28:30:46it'll be a better model.
- 28:30:49It'll it'll learn a little bit better if
- 28:30:51we can scale the data. Um so that way
- 28:30:55like these
- 28:30:58um like ages aren't you know drastically
- 28:31:01different than like in scale than total
- 28:31:03bedrooms or income
- 28:31:06uh those kind of things. So we usually
- 28:31:08want these to be in a similar scale
- 28:31:10range.
- 28:31:13So we'll we will I think we'll scale
- 28:31:15them coming up in a bit and it should
- 28:31:18help the model.
- 28:31:21We've talked about that before, right?
- 28:31:22Scaling usually is a good idea to do
- 28:31:25when you're prepping your data for
- 28:31:26modeling.
- 28:31:29No, you want to you want to scale your
- 28:31:31test data as well. You're going to do
- 28:31:33both. You're going to scale your
- 28:31:34training data. You're going to scale it.
- 28:31:36So that's actually a good point you
- 28:31:38bring up is any transformations you do
- 28:31:40on your training to build your model,
- 28:31:43you should also do on your test set so
- 28:31:45you get an applesto apples comparison.
- 28:31:48You should always do the same
- 28:31:49transformations.
- 28:31:51Yes. Would scaling data impact K? Yeah,
- 28:31:54it could. It could make it better. It
- 28:31:56could uh Yeah, it should impact it. We
- 28:31:58should get a better model. So, when we
- 28:32:00do the different folds, we'll get
- 28:32:01different we'll get better scores. Yeah,
- 28:32:04it it will impact
- 28:32:08uh yeah, if they're so that's a good
- 28:32:11point. If they're going to use our
- 28:32:12model, then yes, they have to scale the
- 28:32:14data as well. If they're going to if we
- 28:32:16build the model on the assumption that
- 28:32:17the input is scaled, then yes, they have
- 28:32:21to also scale their data when they're
- 28:32:22using it with our model. That's true.
- 28:32:29I mean, not really. I'll show you why.
- 28:32:32There's something that's actually going
- 28:32:33to make it easier um that that will
- 28:32:36automate doing the scaling for them. So,
- 28:32:38they don't they don't have to do the
- 28:32:40scaling manually. it'll just it'll
- 28:32:42happen automatically when they use the
- 28:32:44model. I'm going to show you something
- 28:32:45that's going to automate that which is
- 28:32:47going to be called a pipeline.
- 28:32:49So that part will be automated and they
- 28:32:52won't have to do that. So it won't be
- 28:32:53heavy on the user. No, in theory it is,
- 28:32:57but
- 28:32:58has a really helpful tool to make it
- 28:33:00easy to do that. So I'm going to I'm
- 28:33:02going to show us that um later on in the
- 28:33:04notebook.
- 28:33:07No, the data data is not for a single
- 28:33:09house. It's for like a neighborhood. So
- 28:33:12there's a certain number of households
- 28:33:14in the neighborhood. And this is the
- 28:33:16we're predicting the median house value
- 28:33:18of that neighborhood.
- 28:33:21Yeah. So there's a there's certain
- 28:33:22number of households. There's there's
- 28:33:23like an a median income, a population,
- 28:33:26certain number of people that live
- 28:33:28there. Um proximity generally of where
- 28:33:32that location is. It also has a latitude
- 28:33:34and longitude.
- 28:33:38So,
- 28:33:41and a median age in that neighborhood.
- 28:33:43So, yeah, it's not just a single house.
- 28:33:54Okay, let's go back to this was the hold
- 28:33:58out strategy. So, this is a lot simpler.
- 28:34:00This is just model.fit, right? This is
- 28:34:02just model.fit on the training uh data.
- 28:34:05And then we um can predict on the test
- 28:34:07features and generate test predictions.
- 28:34:10And then we can compute our error on
- 28:34:12those um we can compute our error
- 28:34:15amongst the test predictions and our
- 28:34:16test uh label. So that's our useful mean
- 28:34:21squared error function, right? To to
- 28:34:23compute the MSE. Um let's see what the
- 28:34:27MSE is. So MSE is right here.
- 28:34:33Um now what we could do is we can take
- 28:34:36the MSE
- 28:34:38and we can take the square root of it.
- 28:34:39So let's actually do that. Let's um do
- 28:34:42MP. Square root of the
- 28:34:45um test
- 28:34:48MSE
- 28:34:51and we get um 67 we get 67,000.
- 28:34:56So that's pretty high on this. So when
- 28:34:59we just now look at the difference of
- 28:35:01that, right? When we just do a train
- 28:35:03test split,
- 28:35:05um
- 28:35:08when we just do a train test split, we
- 28:35:10get a worse score because it's not as
- 28:35:12it's not as robust, right? We're not
- 28:35:14showing that to many of the other uh
- 28:35:17folds. So, we get a lot more error this
- 28:35:20way on the test data.
- 28:35:23So, this is um actually worse
- 28:35:25performance just doing the train test
- 28:35:27split.
- 28:35:30This is a really higher.
- 28:35:42Yeah, we can. We can. I'm going to I'm
- 28:35:45going to show us how to how the scaling
- 28:35:46will be done automatically. Yes, we can.
- 28:35:50Um there's there's a really easy tool to
- 28:35:53do that will scale it automatically.
- 28:36:00It's going to be later in this notebook.
- 28:36:01I'll show us it.
- 28:36:10All right. So, just to recap this, this
- 28:36:12is fitting the model.
- 28:36:14This is fitting the model. This is
- 28:36:16making the predictions, right?
- 28:36:18Model.predict.
- 28:36:20So, this is making the predictions. And
- 28:36:21then this is calculating the error, the
- 28:36:24mean squared error, which is looking at
- 28:36:25our test labels versus our test
- 28:36:27predictions, right? And this is
- 28:36:29computing the distance, the average
- 28:36:32distance away from these values to these
- 28:36:34values,
- 28:36:36right?
- 28:36:38And then we can also compute the R squar
- 28:36:40R R squar and we see that it's not a
- 28:36:43very good R squar 65 uh is not a very
- 28:36:46great model
- 28:36:48um because it closer to one would be
- 28:36:50better. So this is still this is not
- 28:36:53very good.
- 28:36:54We know that we knew that from the cross
- 28:36:56file score but this is just doing um
- 28:36:58this is just doing a hold out uh where
- 28:37:01we do a train and test split. Right? So
- 28:37:03it's a little bit simpler but it's not
- 28:37:05quite as robust. Um,
- 28:37:08it's not quite as robust as the cross
- 28:37:10valve, but it works. Um, it's, you know,
- 28:37:14we can do hold out. Um,
- 28:37:17we can do hold out, uh, to to quickly
- 28:37:19evaluate a model and see if we need to
- 28:37:22make any adjustments.
- 28:37:24It's a little bit quicker to run.
- 28:37:30Okay. And any questions on it? Does it
- 28:37:32make sense what we're doing here?
- 28:37:33Model.fit fit to train it predict to get
- 28:37:37our predictions. Um this is pretty
- 28:37:40standard, right? To train is the
- 28:37:41model.fit and then to use the model to
- 28:37:44predict we predict on the test features.
- 28:37:47Um so this is passing on on all of our
- 28:37:49features into this model to generate
- 28:37:53predictions for every row. That's
- 28:37:55something I also want to point out that
- 28:37:56may be a little bit confusing is this is
- 28:37:58a data frame. So we're passing in a
- 28:38:02bunch of rows of features with columns,
- 28:38:04right? So um we're passing in a bunch of
- 28:38:07data that looks like this. And what
- 28:38:10we're doing is essentially making a
- 28:38:12prediction for every row. So this will
- 28:38:14generate a prediction. This row will
- 28:38:17generate a prediction. This row will
- 28:38:19generate a prediction and on and on and
- 28:38:21on. So this this predict will predict
- 28:38:24for every row. And so we end up with
- 28:38:27this collection of predictions here for
- 28:38:29each row. and we're comparing those to
- 28:38:32the labels that we have for those rows
- 28:38:35from our from our supervised learning,
- 28:38:37right? From our data set. So that's
- 28:38:40truly supervised learning, right? We
- 28:38:42have the examples and we're comparing
- 28:38:44those to what our model is predicting to
- 28:38:47to get our performance.
- 28:38:56All right.
- 28:38:58So let's uh let's try the other just so
- 28:39:01you can see it. The leave one out. Now
- 28:39:03the leave one out cross validation is
- 28:39:05going to actually work the same way
- 28:39:07where we put in the leave one out um
- 28:39:10strategy inside of the cross file score.
- 28:39:13Now here we don't need to specify how
- 28:39:14many folds there are because we know how
- 28:39:17many they're going to be. It's going to
- 28:39:18be the number of data points, right? So
- 28:39:21which is actually going to be quite
- 28:39:22large because there's 20,000 rows. So
- 28:39:25this is going to be extremely
- 28:39:28uh extremely um intensive because we are
- 28:39:33doing um you know 20,000 examples and
- 28:39:37leaving one example out to be our
- 28:39:40validation and then um doing that across
- 28:39:43every 20,000 uh examples.
- 28:39:47So we could do it though just to see how
- 28:39:48it works. Um we have this again leave
- 28:39:51one out. We generate our crossfile score
- 28:39:53from our model our data and then same
- 28:39:56scoring that we had before and but this
- 28:39:58time we change our cross file to be
- 28:40:00instead of our kfold object we have our
- 28:40:03leave one out object which is this
- 28:40:06um and then we could run this. We can
- 28:40:09compute our average uh across the all
- 28:40:12the folds. Now this is going to be a lot
- 28:40:14bigger of an array. It's going to be a
- 28:40:1620,000 size array and we're going to
- 28:40:19compute the average across it.
- 28:40:23So, let's do that. It's going to take a
- 28:40:25moment because there's lots. So, if you
- 28:40:27notice it when you run, it's going to
- 28:40:28take a little bit of time to run because
- 28:40:30it's running across all 20,000 examples
- 28:40:34and leaving one out. So, you have 20,000
- 28:40:37and then one left out to uh test
- 28:40:41against. So, it's quite intensive. You
- 28:40:43can see it's taking a lot more time.
- 28:40:52It's still running. It's taking a while.
- 28:41:03Okay, just let that run. Still running.
- 28:41:08So, if you guys try running this, it's
- 28:41:10going to take a little bit of time.
- 28:41:11Hopefully, that makes sense why it's
- 28:41:13taking so long, right? It's because it's
- 28:41:16instead of doing 10 folds, it's it's
- 28:41:19putting every data point but one is the
- 28:41:21training set and then iterating through
- 28:41:23all 20,000 points.
- 28:41:27This takes a while to do.
- 28:41:44Let's see what our
- 28:41:49RAM our memory is a little increased.
- 28:42:00Okay,
- 28:42:02still running. That's okay. I'll let it
- 28:42:04run.
- 28:42:07Come back when it's finished.
- 28:42:15Yeah, exactly. This is a this is for
- 28:42:17this is giving us a performance
- 28:42:18evaluation. This is like the average
- 28:42:21error across all of our uh different
- 28:42:24folds. Um now this is the extreme case
- 28:42:26where we have the number of folds equals
- 28:42:28the number of points.
- 28:42:31Right? So it's an extreme case but yes
- 28:42:33it's just like kfold. It's giving us
- 28:42:35that performance estimate.
- 28:42:44Okay. It's about the same. Right. This
- 28:42:47is still around 50,000.
- 28:42:50Not much difference, right? Still right
- 28:42:52around there. But look how much longer
- 28:42:55it took. That took 2 minutes to run. The
- 28:42:57other one was pretty instant, right? So
- 28:42:59this this took about 2 minutes to run.
- 28:43:02So um definitely uh
- 28:43:09yeah, definitely don't want to run this
- 28:43:11uh too often. I think that it's
- 28:43:14generally preferred to do k-fold. If
- 28:43:16you're going to do cross validation,
- 28:43:17generally want to do k-fold or just the
- 28:43:19regular hold out train test split. Uh
- 28:43:22generally better than doing leave one
- 28:43:24out. It's just going to take too long
- 28:43:26and um it results in about the same kind
- 28:43:29of score as the kfold.
- 28:43:41Okay,
- 28:43:44any questions about um the cross
- 28:43:47validation that we just did.
- 28:44:08Okay,
- 28:44:10good. And as it says here that the
- 28:44:12stratified kfold is usually used for
- 28:44:14classification. Again, we're not doing
- 28:44:15classification yet. That's in going to
- 28:44:17be in lesson four. So, we don't need to
- 28:44:19worry too much about that. Just for
- 28:44:20regression, um regular k-fold is
- 28:44:23preferred, right? Because we don't need
- 28:44:25to um worry about distributing
- 28:44:28categories amongst our folds uh in any
- 28:44:31regression problems.
- 28:44:34And as we see the error is kind of high.
- 28:44:36Um there's going to be some things we
- 28:44:37can do to improve that which will be uh
- 28:44:40later on we'll learn about some more
- 28:44:42advanced models. This signals that the
- 28:44:45performance is bad. We probably need a
- 28:44:47more complex model. Um one thing we
- 28:44:50could try before we try a complex model
- 28:44:52is to do scaling. We will try to do
- 28:44:55scaling. I'm going to show us how we can
- 28:44:57do that coming up um in a in a nice
- 28:45:00streamlined fashion. Um, but uh outside
- 28:45:04of that, if we still had bad
- 28:45:06performance, we would likely need to use
- 28:45:07a more advanced model. And we'll learn
- 28:45:10about more advanced models uh in the
- 28:45:13next lesson. And what's great is some of
- 28:45:15those advanced models can actually be
- 28:45:17used for regression. So they have
- 28:45:19variations that can be used for both
- 28:45:21classification and regression, which is
- 28:45:23pretty cool. So I'll point those out
- 28:45:25when we get to them. Um, okay.
- 28:45:30So what I want to talk about now is a
- 28:45:33way we can combat overfitting. So if we
- 28:45:36have overfitting which remember that is
- 28:45:38the case where the uh the we see good
- 28:45:43performance on the training data but
- 28:45:45then um it doesn't generalize over to
- 28:45:47the test data. We get poor performance
- 28:45:49on the test data. Um there's there's a
- 28:45:52drop off there. Um that would signal
- 28:45:55overfitting.
- 28:45:58overfitting
- 28:46:00and one way of um combating overfitting
- 28:46:03is to do something called regularization
- 28:46:06which we're going to talk about next. So
- 28:46:09the key idea in regularization
- 28:46:13is to
- 28:46:15change our uh the change the way we
- 28:46:19train. Essentially, what we're going to
- 28:46:21do is modify our training
- 28:46:27uh error function or sometimes called
- 28:46:30the objective function or loss function.
- 28:46:33We're going to change that to add a
- 28:46:36penalty to penalize excessive complex
- 28:46:40complexity. Essentially the the way that
- 28:46:43we're going to penalize is by making
- 28:46:45sure the size of the coefficients
- 28:46:48doesn't grow too much which should
- 28:46:51mitigate overfitting because remember in
- 28:46:54linear regression what we are learning
- 28:46:56are the coefficients right we're
- 28:46:58learning the beta 0 the beta 1 the beta
- 28:47:012 and on and on however many betas there
- 28:47:04are beta n we're learning all of those
- 28:47:06guys um through the regression error
- 28:47:10function we're trying to minimize that
- 28:47:11error function. That's how it trains. We
- 28:47:13talked about that on Monday.
- 28:47:16Um so what we're going to do is um
- 28:47:21basically penalize the these guys
- 28:47:24growing too big and making sure we kind
- 28:47:27of keep them small so that no one
- 28:47:31coefficient has a dominant uh effect on
- 28:47:34the model. And this should help with
- 28:47:36overfitting and complexity. It should
- 28:47:38make the model simpler because all the
- 28:47:40coefficients are going to be encouraged
- 28:47:42to be smaller. They're not going to grow
- 28:47:44too big. Um, and this this has the
- 28:47:47effect of making the model so basically
- 28:47:50make the model simpler.
- 28:47:54Make the model simpler is what these
- 28:47:57regularization techniques are
- 28:47:58essentially trying to achieve is is
- 28:48:00remove complexity, make them a little
- 28:48:02bit simpler, make these coefficients
- 28:48:04smaller so that you can generalize a bit
- 28:48:07better and and prevent overfitting. So
- 28:48:10we want to prevent
- 28:48:13uh overfitting,
- 28:48:15right, is what we want to do. Um so
- 28:48:19there's going to be a penalty and I'll
- 28:48:20show you where that penalty gets added
- 28:48:22and kind of what it looks like.
- 28:48:24Um but uh to control the level of that
- 28:48:28penalty we are actually going to
- 28:48:29introduce another parameter to our model
- 28:48:33um called alpha.
- 28:48:35Alpha is going to scale the penalty. So
- 28:48:38if alpha is really high that imposes a
- 28:48:42stronger penalty on the coefficients um
- 28:48:45which will make the model a lot simpler.
- 28:48:48So the higher the alpha the simpler the
- 28:48:51model we will get and we the the risk
- 28:48:54with that is we actually underfit. So if
- 28:48:57alpha is too big we may underfit the
- 28:49:00training data
- 28:49:02um a bit too much because it will make
- 28:49:04the model way too simple. Um and again
- 28:49:08I'll show you what this means
- 28:49:08mathematically in a moment. Um but on
- 28:49:12the other hand if we have a lower alpha
- 28:49:14this will have a lower penalty. it's a
- 28:49:17weaker penalty term and that'll lead to
- 28:49:20a model that is um a bit more complex.
- 28:49:24Um which could um risk some level of
- 28:49:27overfitting. Um so there's so there's
- 28:49:31still the risk of overfitting if you
- 28:49:33have a low alpha. And of course if alpha
- 28:49:35goes all the way to zero there's no
- 28:49:37penalty at all. So you're back to your
- 28:49:39original linear regression um which
- 28:49:42could risk a lot of overfitting.
- 28:49:45Right? So you you generally want to pick
- 28:49:47an alpha um effectively and actually
- 28:49:50we're going to see h what's the best way
- 28:49:52to pick alpha. Um we're actually going
- 28:49:54to learn how to do that. I'm going to
- 28:49:56show us how doing some tuning techniques
- 28:49:58to pick what alpha should be. Um but um
- 28:50:03a a pretty industry standard alpha that
- 28:50:05most people default to is alpha equals
- 28:50:08to one. So just just one which signals
- 28:50:12that there should be some penalty. we
- 28:50:14just have alpha equal to one is a
- 28:50:16standard penalty. We don't want it to be
- 28:50:18too high. We don't want it to be too
- 28:50:19low. Like we don't want it to be a
- 28:50:20fraction. Um but a penalty of one is
- 28:50:23usually uh good enough.
- 28:50:27Okay, I'm going to show you where that
- 28:50:29comes into play in a moment.
- 28:50:32Um but the whole purpose of doing this
- 28:50:34is to mitigate overfitting, right? Um
- 28:50:37that's what and and doing this penalty
- 28:50:40is is called regularization. So adding
- 28:50:43so going beyond just regular linear
- 28:50:45regression adding this extra penalty to
- 28:50:48to the training process um to penalize
- 28:50:51large weights large coefficients
- 28:50:54um is known as regularization.
- 28:50:58Okay. Um and there's two common
- 28:51:01penalties that are added. Um so there's
- 28:51:04actually two different variations on the
- 28:51:05penalty. Um we're going to study both of
- 28:51:07them and um they're they're known as
- 28:51:10lasso. So if you take linear regression
- 28:51:12and add a particular type of penalty,
- 28:51:14it's known as lasso. If you add another
- 28:51:17type of penalty, it's known as ridge
- 28:51:19regression. We're going to study both of
- 28:51:21those and what their differences are.
- 28:51:23But these are the primary two
- 28:51:26uh regularization tech uh models that
- 28:51:29are used um to take a regular both of
- 28:51:32these take regular linear regression and
- 28:51:34just modify the training process a
- 28:51:37little bit in different ways. Two
- 28:51:39different ways. um using that alpha
- 28:51:43um to penalize the terms in slightly
- 28:51:46different mathematical ways. So we're
- 28:51:48going to learn about these two guys.
- 28:51:50Lasso regression there. Both of these
- 28:51:52are just offshoots of linear regression.
- 28:51:54So underlying model is still linear
- 28:51:56regression. It just adds different types
- 28:51:59of penalties to the training process.
- 28:52:02So both of these are still in the family
- 28:52:05of linear regression. In fact, in um in
- 28:52:09scikitlearn, they both come from they
- 28:52:11both are still from the linear model
- 28:52:13family in inside of the linear model
- 28:52:16module, which is where linear regression
- 28:52:18comes from. So there's still linear
- 28:52:19regression. They just have different
- 28:52:22styles of penalties added to them. Um
- 28:52:25which we're going to see.
- 28:52:28Okay, so just to recap that
- 28:52:32regularization is the process of adding
- 28:52:34a penalty to the training to discourage
- 28:52:38complexity. In this case, we're going to
- 28:52:40discourage large coefficients.
- 28:52:44And um this should help prevent
- 28:52:47overfitting.
- 28:52:49And so uh these are going to lead us to
- 28:52:52two different offshoots of linear
- 28:52:53regression that have two different
- 28:52:55penalties.
- 28:52:56lasso and ridge regression, which we're
- 28:52:59going to uh study next,
- 28:53:01but they they function the same way as
- 28:53:03linear regression. They will just have
- 28:53:06different penalty terms added onto their
- 28:53:08training process um to discourage
- 28:53:12uh discourage um again those large
- 28:53:15weights.
- 28:53:20Okay, any questions about regularization
- 28:53:22before we first look at our we're going
- 28:53:24to look at our first uh variation on on
- 28:53:27our first regularization technique which
- 28:53:28is going to be called lasso regression.
- 28:53:48Okay, let's look at lasso regression. So
- 28:53:50what is lasso regression? It's actually
- 28:53:53lasso is short for least absolute
- 28:53:56shrinkage and selection operator
- 28:53:58regression. Um and this will function by
- 28:54:03adding a particular penalty to the
- 28:54:07linear regression model. So again, it's
- 28:54:09based on linear regression. That's the
- 28:54:11underlying model. It's just that during
- 28:54:13the training process, we are going to um
- 28:54:16add a penalty which has the effect of
- 28:54:21shrinkage of the weights. That's why
- 28:54:23it's called shrinkage. It encourages
- 28:54:25smaller weights through that penalty.
- 28:54:28And it also will shrink some of them so
- 28:54:30much that they'll become zero. And so it
- 28:54:33has has an effect of kind of selection
- 28:54:36which means that some of them get wiped
- 28:54:38out to zero.
- 28:54:40And this means that whatever is left
- 28:54:43over is kind of what's selected as our
- 28:54:45features because the other ones will
- 28:54:48have zero weight applied to them. So
- 28:54:50this penalty will really favor small
- 28:54:54weights um and penalize really large
- 28:54:58weights. In fact, it will favor small
- 28:55:00weight so much that some of them will
- 28:55:02actually um be shrunk to zero um during
- 28:55:06the training process. And the ones that
- 28:55:08are left over are the ones that um are
- 28:55:12the ones that are what we call selected
- 28:55:15because they are the ones that remain in
- 28:55:17in the training um after the other ones
- 28:55:20get uh coefficients of zero. Um now when
- 28:55:24you make some of the coefficient zero
- 28:55:27you are inherently making the model
- 28:55:29simpler right there's less features
- 28:55:31involved in the prediction that or less
- 28:55:33features that have an effect on the
- 28:55:35prediction. So this definitely makes the
- 28:55:37model simpler. This lasso this shrinkage
- 28:55:41and selection uh process makes makes the
- 28:55:45model simpler for sure. Um
- 28:55:48and this is supposed to reduce
- 28:55:50overfitting. Right? If you make the
- 28:55:52model simpler, it's not as complex. It
- 28:55:54has less of a chance of memorizing
- 28:55:56training data and not generalizing over
- 28:55:59to test data. So our whole goal with uh
- 28:56:03regularization is to make our model
- 28:56:05better at generalization, right? Over to
- 28:56:08test data from the original training
- 28:56:10data.
- 28:56:12Um so how does this happen? We have to
- 28:56:15go back to the
- 28:56:18uh training process. If you guys
- 28:56:20remember, I I wrote out this equation a
- 28:56:22little bit earlier, which is the
- 28:56:24distance. This is the sum of squared
- 28:56:27distance between our labels and our
- 28:56:28prediction.
- 28:56:30This is basically the mean squared error
- 28:56:32uh calculation that we're trying to
- 28:56:34reduce when we build our model using the
- 28:56:36training data. Um so this is just in
- 28:56:38standard linear regression. This is the
- 28:56:41um uh sum of squares uh distance, right?
- 28:56:45So this is this is what the model is
- 28:56:47trying to minimize when it learns these
- 28:56:50coefficients.
- 28:56:51So when it learns these coefficients,
- 28:56:53it's trying to minimize this guy
- 28:56:58minimize. It's trying to find the betas
- 28:57:01that minimize this quantity
- 28:57:03mathematically. That's what it's doing.
- 28:57:05Um and there's there's a algorithm that
- 28:57:08will discover what the best betas are
- 28:57:11that actually minimize uses that gives
- 28:57:12us a line of best fit, right? That's
- 28:57:14what we've been talking about for
- 28:57:15regression.
- 28:57:17Now, in regularization,
- 28:57:21here's, by the way, here is that same
- 28:57:22thing, but we've just inserted our model
- 28:57:24for the predictions. This is our model.
- 28:57:28Just a fancy way of writing down our
- 28:57:29model, right? It's the beta 0 plus all
- 28:57:32of these betas. So, beta 1 x1 plus beta
- 28:57:362 x2
- 28:57:38plus on and on and on, right? That's
- 28:57:41that's what this uh means. If you're
- 28:57:43unfamiliar with the sigma notation, it
- 28:57:45just means sum. So it's the sum of all
- 28:57:47these guys or this term. Um, so this is
- 28:57:52this here is just a regular linear
- 28:57:54regression
- 28:57:58uh training regular linear regression
- 28:58:01training. So we the training process
- 28:58:05solves for these parameters, right? It
- 28:58:07solves for these weights. We discover
- 28:58:09what those are by minimizing this
- 28:58:11quantity. That's the whole training
- 28:58:13process. Um, but when we do lasso,
- 28:58:18we add a penalty which is this.
- 28:58:23Here is our penalty.
- 28:58:27So basically um we take our linear
- 28:58:31regression training which is this and we
- 28:58:34add on a penalty which is this. And you
- 28:58:37can see exactly what this penalty when
- 28:58:40when you minimize this penalty. It's
- 28:58:42when these weights are small. So this
- 28:58:45encourages
- 28:58:46So minimizing this quantity encourages
- 28:58:50small weights
- 28:58:53encourages small betas
- 28:58:57beta I
- 28:58:59right you or in this case beta j sorry
- 28:59:04this encourages small beta js uh because
- 28:59:07we want this thing to be minimized
- 28:59:11minimized
- 28:59:13so Um, what's going to make this minimal
- 28:59:16is of course the line of best fit and
- 28:59:18small weights, right? Are going to make
- 28:59:20are going to bring this error down the
- 28:59:24most.
- 28:59:26So, um, and here's our alpha, right?
- 28:59:28Here's our alpha. So, you can encourage
- 28:59:30a higher penalty with a larger alpha or
- 28:59:33a lower penalty. If alpha equals zero,
- 28:59:36what happens to that term? It just goes
- 28:59:39away. So if alpha equals zero, there's
- 28:59:41no penalty and we're back to uh we're
- 28:59:45back to regular
- 28:59:47linear regression.
- 28:59:50We just have regular linear regression
- 28:59:51because we have no penalty at that point
- 28:59:52when alpha equals zero. So the smaller
- 28:59:56alpha is, the less penalty we're
- 28:59:59enforcing and in the regularization.
- 29:00:03Okay.
- 29:00:04Now what happens is in reality when you
- 29:00:07train with lasso. So this is lasso is
- 29:00:10this particular penalty. This is called
- 29:00:12the lasso penalty
- 29:00:14or sometimes um people call this the L1
- 29:00:17penalty.
- 29:00:19Um L1 just comes from the fact that this
- 29:00:23is the first power or absolute value. Um
- 29:00:26so it's not a squared penalty, it's a
- 29:00:28single uh single power penalty
- 29:00:31um there. But when you add this lasso
- 29:00:35penalty, what can happen is it it does
- 29:00:38because the because you're minimizing
- 29:00:40this, it does encourage some of these
- 29:00:42weights to become zero.
- 29:00:46So some if you're really trying to get
- 29:00:48the lowest quantity of this,
- 29:00:51the lower the better.
- 29:00:54What makes this thing lower is of course
- 29:00:57if some of these go away if some of
- 29:00:58these go to zero then that of course
- 29:01:01will lower this as much as we as much as
- 29:01:03possible right so what happens during
- 29:01:05the training is some of these
- 29:01:08coefficients actually they're encouraged
- 29:01:10to be small because of this penalty but
- 29:01:12some of them will actually become will
- 29:01:15actually become zero um in order to get
- 29:01:18the best model the best fit some of
- 29:01:20these will actually get so small that
- 29:01:22they'll basically become zero
- 29:01:24And that means that that that feature
- 29:01:27basically has no effect anymore. It's
- 29:01:30it's been the model has been simplified,
- 29:01:33right? That feature no longer really has
- 29:01:34an effect.
- 29:01:40So just to call out the alpha again, um
- 29:01:42if alpha zero some code, uh basically
- 29:01:46you have your linear regression, you're
- 29:01:48back to linear regression because alpha
- 29:01:500 is just wiping this out and you're
- 29:01:51back to linear regression.
- 29:01:54um if alpha is infinity. Now if alpha is
- 29:01:56infinity that's an extreme. So if alpha
- 29:01:58is infinity the only way to make this
- 29:02:00minimize is if all your coefficients are
- 29:02:02zero. If every beta is zero then this
- 29:02:05will lower the the error as as much as
- 29:02:07possible. So you basically have no
- 29:02:09model. So if all coefficients are zero
- 29:02:12you have no model and that's useless. So
- 29:02:14you don't want your penalty you don't
- 29:02:17want your alpha to be huge is what this
- 29:02:19is saying. You also don't want your
- 29:02:20alpha to be small. you're basically back
- 29:02:22to linear regression. So you want
- 29:02:23something in between. Um and the typical
- 29:02:27typical value is alpha equals 1.
- 29:02:31Typical is alpha equals 1
- 29:02:35to have some level of penalty there. So
- 29:02:38just a regular kind of regular penalty
- 29:02:40term.
- 29:02:48But we are actually going to have a way
- 29:02:50to test and evaluate which alphas are
- 29:02:52the best.
- 29:03:07Um,
- 29:03:08basically you can yeah you can have a
- 29:03:12you can have a penalty that's close to
- 29:03:13zero. You can get rid of this if just a
- 29:03:16regular linear regression performs
- 29:03:18pretty well. You can basically have no
- 29:03:20penalty in that case.
- 29:03:23Yeah. So nearer zero or like it could be
- 29:03:26that adding a little bit of penalty
- 29:03:28actually helps the overfitting and it
- 29:03:30could be really small. One thing that
- 29:03:32we're basically going to do is have a
- 29:03:34strategy to try out different alphas.
- 29:03:40try different alphas
- 29:03:43and evaluate performance
- 29:03:47and then we can decide which so that's
- 29:03:49what we're going to do is have a
- 29:03:51strategy to just plug in different
- 29:03:52alphas generate the like train the model
- 29:03:56and then see what its performance is and
- 29:03:58see if those alphas are good what what
- 29:04:00which alpha is the best we can evaluate
- 29:04:03that
- 29:04:04because we can train the model and see
- 29:04:06what it performance is
- 29:04:11right.
- 29:04:14Yeah. Yeah. So, we'll do that. We'll
- 29:04:15practice that.
- 29:04:24Okay. Great. Any other questions about
- 29:04:26this lasso regression? So, remember this
- 29:04:28is linear regression here. This is the
- 29:04:30this is how you're training to find the
- 29:04:33betas in linear regression. So this is
- 29:04:35just linear regression uh um training
- 29:04:40function there.
- 29:04:42We're adding a penalty which is this is
- 29:04:44the lasso penalty
- 29:04:47lasso penalty there right we're adding
- 29:04:49that this is known as regularization
- 29:04:52and the goal of regularization is to
- 29:04:56prevent overfitting. So you add a
- 29:04:57penalty here this makes the model
- 29:05:00simpler which prevents overfitting.
- 29:05:03helps you generalize better when it's
- 29:05:06simpler.
- 29:05:16Any questions conceptually on this?
- 29:05:18We're going to do a code example with it
- 29:05:19coming up, but any questions on this?
- 29:05:37Uh yeah, you you so that's the thing,
- 29:05:39Ronald, is you may be willing to
- 29:05:41sacrifice some accuracy in order to
- 29:05:44generalize to unseen data because
- 29:05:46remember that's what we're really trying
- 29:05:48to get after is we may be willing to
- 29:05:51sacrifice some accuracy on this training
- 29:05:52data in order to have it perform better
- 29:05:54on the test data, right? we may be
- 29:05:57willing to do that. That's a willing
- 29:05:59that's an okay sacrifice
- 29:06:02as long like if if it generalizes
- 29:06:05better. That's what we want. That's what
- 29:06:07we're trying to do here is add a
- 29:06:09penalty, make the model simpler, and
- 29:06:13help it generalize better to new and
- 29:06:16unseen data. Right?
- 29:06:20That's that picture I've been using with
- 29:06:21the with the um train and test split.
- 29:06:25Where is the square?
- 29:06:27So in the model there's no square. So
- 29:06:30remember the model is the model is this
- 29:06:36um equation uh that has no squares in
- 29:06:38it, right? It's beta 0 plus beta 1 x1
- 29:06:43plus beta 2 x2 plus beta n xn.
- 29:06:50That's the that's the linear regression
- 29:06:52model. This is the now this this is the
- 29:06:55model but this is the equation that
- 29:06:58helps us train and find the betas. This
- 29:07:01is how this is what we find the betas
- 29:07:03with. So we'll continue. Um we were
- 29:07:07talking about the lasso regression which
- 29:07:10uh adds it takes linear regression right
- 29:07:13which is this optimization and adds in a
- 29:07:16penalty um scaled by the alpha. Um, and
- 29:07:20what that does in order to minimize this
- 29:07:23whole thing, it encourages these to be
- 29:07:26small uh as possible. Um, which makes
- 29:07:30the model simpler, right? The weights
- 29:07:32don't get overly big and complex. Um,
- 29:07:35they they tend to stay small. In fact,
- 29:07:37some of them can even go all the way to
- 29:07:39zero. Um, which makes the model even
- 29:07:41more simpler,
- 29:07:43right? Um, so let's practice uh using it
- 29:07:47in code. It's actually really easy to
- 29:07:48use. It's going to be essentially the
- 29:07:51same uh style and and code as linear
- 29:07:55regression except we are um just going
- 29:07:58to have to uh put in our alpha parameter
- 29:08:01um when we use the lasso. So here we are
- 29:08:06um from the linear model family right
- 29:08:10which makes sense. It's a linear
- 29:08:12regression offshoot that has this
- 29:08:13penalty in it during the training. um we
- 29:08:15are grabbing our lasso regression. Um it
- 29:08:19also has a version of the lasso that
- 29:08:22we're going to take a look at that is
- 29:08:23used for cross validation which is
- 29:08:26really um convenient as well. So it has
- 29:08:29a cross validation lasso which is a
- 29:08:31really convenient um combination of
- 29:08:34basically cross val score and lasso um
- 29:08:37all in one. So it actually is really
- 29:08:39nice to use that way. Um so we'll take a
- 29:08:42look at that example. Um, but we are
- 29:08:45importing it. The main thing is going to
- 29:08:47be the lasso model here. Um, we're going
- 29:08:50to be using a different data set for
- 29:08:51this one. So, not the ocean uh data, but
- 29:08:54this hitters data, which is a baseball
- 29:08:56data set. Um, so it has 322 rows um with
- 29:09:0120 different columns and it looks like
- 29:09:03this. So, you want to download that one.
- 29:09:06Um, hopefully you guys have access to
- 29:09:08that one.
- 29:09:11Um,
- 29:09:13so I will upload it into
- 29:09:16this.
- 29:09:19So give me a moment.
- 29:09:25There's that. And then we can run this.
- 29:09:29Okay. So we are displaying the data and
- 29:09:32so it has um the the hitters names and
- 29:09:36then it has a bunch of different
- 29:09:37statistics. These are all baseball
- 29:09:38statistics.
- 29:09:40Um, if you're unfamiliar with with them,
- 29:09:42that's okay. It's not a big deal. Um,
- 29:09:44but just different baseball stats here.
- 29:09:49Okay. Were you guys able to load that?
- 29:09:51Um, if you're following along, were you
- 29:09:52able to load that? You should have
- 29:09:54access to this data. The hitters CSV.
- 29:09:58This is the one we're going to use for
- 29:09:59the lasso model
- 29:10:02to build a lasso model.
- 29:10:12Yeah.
- 29:10:19Okay. Able to load that one. Perfect.
- 29:10:22Okay. So, able to load that one. Um, and
- 29:10:25we take a look at the the head. Um, so
- 29:10:28we're actually going to uh drop this
- 29:10:31unnamed column because we don't care
- 29:10:33about their name. it's actually just the
- 29:10:35batter's name which is not going to be
- 29:10:37useful in modeling. Um so and remember
- 29:10:40that's generally true like an ID, a user
- 29:10:44ID, like a customer ID, a name, that's
- 29:10:47usually not going to be useful in any
- 29:10:48kind of modeling. So we're actually just
- 29:10:50going to drop that uh column and we're
- 29:10:52going to do it in place.
- 29:10:54And access equals 1 means we're dropping
- 29:10:56that column. Um, so we're going to drop
- 29:10:59that and we should no longer have that
- 29:11:02column and we have all of these guys
- 29:11:03now. So you want to run that. This will
- 29:11:05drop that. Um, this will drop drops the
- 29:11:10column in place.
- 29:11:14Um, and now we can see we have uh all we
- 29:11:18have this data where um we have this
- 29:11:21data where it's now removed. So, this
- 29:11:24that column is now gone and now we have
- 29:11:26these guys. Um, do you notice anything
- 29:11:30about this
- 29:11:32from the info?
- 29:11:38Looks like we have a couple categorical
- 29:11:39features, a few of them, league and
- 29:11:42division
- 29:11:44and new league. What do you notice about
- 29:11:47this
- 29:11:56nullles? Yep. So, there's definitely
- 29:11:57some missing data there um that we're
- 29:12:00going to have to deal with.
- 29:12:05So, it looks like there are uh there are
- 29:12:0959.
- 29:12:11Um there are 59. Now we could we the
- 29:12:15alternative to doing that is we could uh
- 29:12:17we could just use our usual code where
- 29:12:19we do dfis
- 29:12:21uh isnull.
- 29:12:24Um and then we do uh dot sum to total
- 29:12:29those up across our different columns.
- 29:12:31And we can see that uh we have 59 of
- 29:12:34those in the salary column. That's this
- 29:12:37is the standard way of doing that,
- 29:12:38right?
- 29:12:42standard way of doing that. And we have
- 29:12:43so we have 59 of those.
- 29:12:5359 of those. So we have to deal with it.
- 29:12:55Any ideas on how to deal with it?
- 29:13:02Any ideas on how to deal with it? This
- 29:13:04is now this is 59 out of 300.
- 29:13:08So,
- 29:13:09what do you guys think about that? It's
- 29:13:11a little bit different than 200 out of
- 29:13:1220,000. A little bit different. We have
- 29:13:16We have about 60 out of 300.
- 29:13:21There's a decent amount.
- 29:13:25Any ideas on how to handle this one?
- 29:13:30Replace. Yep, we should replace. What do
- 29:13:32you think we should replace with?
- 29:13:35It's a float. It's a floating point uh
- 29:13:38value.
- 29:13:49By the way, something unique about this
- 29:13:51that's a little different than usual,
- 29:13:52too, is that the uh this is actually the
- 29:13:56column we're going to use as our label.
- 29:13:58So, we're actually going to predict the
- 29:14:00salary based on the uh based on the um
- 29:14:04rest of the features. So, we definitely
- 29:14:07need to fill in these nles, right?
- 29:14:09Because they're actually going to be the
- 29:14:11labels.
- 29:14:12We're missing some labels uh in our
- 29:14:15data.
- 29:14:16We definitely need to fill them in.
- 29:14:17Yeah. So, we're going to replace them.
- 29:14:29All right. So, we'll we will replace
- 29:14:31them down below. That's going to be
- 29:14:32coming up. Uh we'll come back and
- 29:14:34replace them. um before we replace them,
- 29:14:37we're actually going to get our uh one
- 29:14:40hot encodings for those three different
- 29:14:42um features we have. Um so we do uh get
- 29:14:49dummies with this. Now um of course we
- 29:14:52don't need to do this if we just so this
- 29:14:55code we don't need to do if we just pass
- 29:14:57in the dype here
- 29:15:01um which is uh then we don't need to do
- 29:15:04this. So we can comment this out.
- 29:15:08Um so now what I want you guys to notice
- 29:15:11is this is the alternative to what we
- 29:15:13did before where we are purposely just
- 29:15:15doing these columns not the whole data
- 29:15:17frame but just doing these columns and
- 29:15:20then we can um concatenate those these
- 29:15:26one hot encodings. We're going to
- 29:15:27concatenate back to the data frame.
- 29:15:30Right? So if we do our dummies and then
- 29:15:33do dummies.info info. Um, we can see
- 29:15:36that we end up with six new columns. And
- 29:15:39in fact, we can do dummies.head
- 29:15:43and take a look at what those are.
- 29:15:45Right? So, these are league A, league
- 29:15:48uh, N, division E, W, division W, new
- 29:15:51league A, new league N.
- 29:15:54Okay.
- 29:15:56So, um these are uh these are our one
- 29:16:00hot encodings for these three different
- 29:16:02features which are strings, right? So,
- 29:16:04those those features were strings. If
- 29:16:06you go back up, those were our only
- 29:16:07string features we had. So, we've one
- 29:16:10hot encoded those so we can use them in
- 29:16:11our model. What we need to do is just
- 29:16:14concatenate this back to our data frame.
- 29:16:17Right? So, we just need to concatenate
- 29:16:19it back into our data.
- 29:16:28Okay. So, what we're going to do then is
- 29:16:31we're going to grab um we're going to
- 29:16:35grab Y as our salary. And of course,
- 29:16:37we're going to fill nles on that Y
- 29:16:39coming up shortly. But we're going to
- 29:16:42grab Y as our salary and X new. Now
- 29:16:45before building a full X, we're going to
- 29:16:48take a look at X numerical as our data
- 29:16:51frame minus these columns. The reason
- 29:16:54we're doing minus those is because we
- 29:16:57are going to concatenate our dummy
- 29:17:00variables back into this that are going
- 29:17:02to replace these guys. So we're going to
- 29:17:05replace these anyways with our one hot
- 29:17:07encodings. We don't want the strings. So
- 29:17:09we're going to get rid of those. And
- 29:17:12we're also going to get rid of the
- 29:17:13salary because that's going to be part
- 29:17:15of our that's just a label. So we don't
- 29:17:17want that in the X, the eventual X.
- 29:17:24Are you guys able to run this one?
- 29:17:28Hope I'm not going too fast. You guys
- 29:17:30able to run this? And does it make
- 29:17:32sense? What we're doing is we're putting
- 29:17:34our labels in Y, which is what we
- 29:17:36usually do. So, we're going to predict
- 29:17:38the salary
- 29:17:40and we're getting ready to build the X.
- 29:17:42But before we first want to get rid of
- 29:17:44those one hot the strings. This is
- 29:17:46getting rid of the strings
- 29:17:49and this is getting rid of the label.
- 29:17:51And that's going to be part of our
- 29:17:52features. What we need to do is build
- 29:17:54our final X by concatenating our dummies
- 29:17:56with this. Do you guys see that? We're
- 29:17:59going to concatenate our dummies with
- 29:18:00this to build our final X.
- 29:18:04But but prior to doing that, we need to
- 29:18:06get rid of these string columns here. So
- 29:18:09we're dropping those
- 29:18:11dropping those from the uh data frame uh
- 29:18:15and getting a numerical uh x numerical
- 29:18:18here.
- 29:18:20You can see the columns of that are just
- 29:18:22these guys here. So the the results we
- 29:18:25need to concatenate our we need to
- 29:18:27concatenate this guy um into this and
- 29:18:30then that'll be our full x all of our
- 29:18:32features.
- 29:18:37Okay. So you can see x is going to be
- 29:18:40pd.con
- 29:18:42of this with our dummies.
- 29:18:46This with our dummies. And um
- 29:18:51uh instead of doing this, I'm actually
- 29:18:53going to do the full dummies. We don't
- 29:18:55need to
- 29:18:56um pick just a few columns. We're
- 29:18:59actually going to do our full dummies
- 29:19:01here and um do x equals 1. Now, the
- 29:19:05reason that's the case is because um
- 29:19:08this will get rid of one column per
- 29:19:12feature and basically assume that if you
- 29:19:15have a if you have a zero, the other one
- 29:19:17should be a one. If you have a one, the
- 29:19:19other one should be a zero. Um so it
- 29:19:22basically makes that assumption because
- 29:19:24we only have two of them. Um so whenever
- 29:19:28there's a one, the other should be zero.
- 29:19:30Um, so you can get away with just having
- 29:19:33these three, but um I think it makes
- 29:19:35more sense to just have to have the full
- 29:19:38dummies,
- 29:19:41but by process of elimination, you can
- 29:19:43get away with just using two of them
- 29:19:44because anytime you have a zero, the
- 29:19:47other one should be the other feature
- 29:19:48would have been would have been a one,
- 29:19:51right? And vice versa, when there's a
- 29:19:53one, the other feature would have been a
- 29:19:54zero.
- 29:20:00So we do that one.
- 29:20:11And you can see all of our uh all of our
- 29:20:14one hot encoding features end up back in
- 29:20:16there.
- 29:20:20So this is the code that I want you guys
- 29:20:22to run. I think it makes more sense. It
- 29:20:25follows along what we've been doing.
- 29:20:27um which will concatenate our dummies
- 29:20:30back to our features here to build out
- 29:20:32our full X. So now X is all of our
- 29:20:34features. Um remember X
- 29:20:38X contains all of our features
- 29:20:44now.
- 29:20:46So X contains all of our features and so
- 29:20:49we have all of this now.
- 29:20:56Okay. Were you guys able to run this
- 29:20:57one?
- 29:21:00Damn. We have y, we have x. We still
- 29:21:03need to deal with the nles in y. So that
- 29:21:06something we still need to deal with.
- 29:21:14But hopefully you have this. Now
- 29:21:18all these are numerical.
- 29:21:21So that should be good with the model.
- 29:21:25That's one thing about X is you should
- 29:21:27you our X should have all numerical
- 29:21:31features, right? Because it's going to
- 29:21:33go into a model to to learn those betas.
- 29:21:36So it needs to have all numerical
- 29:21:38features,
- 29:21:40right? These are going to be all
- 29:21:41numerical, which makes sense. We change
- 29:21:43we did one hot encoding to change all
- 29:21:45those guys to numerical.
- 29:21:55Sorry, I'm scrolling down.
- 29:22:00Okay, we do fill in the nil later. Okay.
- 29:22:13Okay.
- 29:22:15Any questions so far? So, we're just
- 29:22:16getting our data ready. We haven't
- 29:22:17applied the lasso yet, but we're just
- 29:22:19doing some prep. Now, hopefully you guys
- 29:22:22recognize th these are some standard
- 29:22:25steps that we're taking when we do our
- 29:22:28modeling. We have to do these data prep
- 29:22:31steps. They're necessary. And so, if it
- 29:22:34seems like it's a lot of work, that's
- 29:22:35because it is. It is work that you do to
- 29:22:39prepare your data to get ready for
- 29:22:41modeling. You have to do that. Okay.
- 29:22:47So, we're doing that here. Um, now we're
- 29:22:50going to do our train test split because
- 29:22:52we're just going to do uh we're going to
- 29:22:54do hold out here. So, we're doing a
- 29:22:56train test split with about with a test
- 29:22:58size of about 0.25. So, again, anywhere
- 29:23:00between 0.2 to.3 would be okay.
- 29:23:04Um, so uh
- 29:23:08it's our choice. We could do 0 2. We
- 29:23:10could do 3. We could do anywhere in
- 29:23:11between there. We're doing 0.25. That's
- 29:23:14fine. Um, that's okay. So, we we build
- 29:23:17our train test split right there. Um, so
- 29:23:21pretty pretty simple and we've seen that
- 29:23:23a bunch of times with our X and our Y
- 29:23:27data frames. There we have our train
- 29:23:30test split.
- 29:23:33Okay.
- 29:23:35Um, now what we're going to do is do our
- 29:23:38our scaling. So, we're we didn't do this
- 29:23:41last time, but we're going to do this
- 29:23:43now as uh because we should get in the
- 29:23:45habit of doing that. Um is um we're
- 29:23:50going to um go ahead and scale our
- 29:23:54features and we're going to use the
- 29:23:56standard scaler here uh to do that
- 29:23:59scaling. Okay. Now, we could use minmax
- 29:24:02scaler. That's fine, too. We're just
- 29:24:04going to use the standard scaler here.
- 29:24:06Um and remember we are going to uh um
- 29:24:11use the standard scaler from sklearn and
- 29:24:14we're going to transform our features uh
- 29:24:18uh according to our um according to our
- 29:24:23training data. So we have our
- 29:24:27pre-processing standard scaler here. So
- 29:24:29we import that guy and then we um build
- 29:24:33our standard scaler and fit it on the
- 29:24:36training data only on the numerical
- 29:24:39features. Um so that's which is going to
- 29:24:44be uh all of these guys. So we're doing
- 29:24:48the scaling on all of these guys. Now
- 29:24:50something to note is that we are not
- 29:24:53scaling all of these one hot encodings
- 29:24:56mainly because it doesn't make sense to
- 29:24:58scale those really. They're zero or one.
- 29:25:01They don't need to be scaled, right?
- 29:25:03They're already zero and one. So they're
- 29:25:06they don't need to be even if we were
- 29:25:07doing minmax scaling, it's going to put
- 29:25:09them between zero and one. It wouldn't
- 29:25:11affect it really, right? So these one
- 29:25:14hot encoding features, we're not going
- 29:25:15to scale because they're they're always
- 29:25:17going to be zero or one.
- 29:25:20There's no need to scale them really.
- 29:25:22Um, but we're going to scale all the
- 29:25:24other features here that are floats.
- 29:25:27So that's these guys here. These
- 29:25:30numerical features we're going to scale.
- 29:25:33Okay.
- 29:25:34Don't need to we don't really need to
- 29:25:36scale the one hot encoding. Uh, it's
- 29:25:38pretty much already scaled.
- 29:25:46Oh, you should change that. Um, go back
- 29:25:48and rerun go back and rerun this, but
- 29:25:51make sure you have your data type as int
- 29:25:53here.
- 29:25:55Make sure you add that in there to
- 29:25:57change that over to integer and rerun
- 29:25:59that and then rerun the rerun the
- 29:26:03concatenation.
- 29:26:04So, make sure you run this
- 29:26:07and then u make sure you rerun this and
- 29:26:10rerun the concatenation part which is uh
- 29:26:14this
- 29:26:27Okay. So, we go ahead and fit the um
- 29:26:32scaler to this data and then we're going
- 29:26:34to transform our training features,
- 29:26:36those numerical features. um we're and
- 29:26:40then we're going to uh transform these
- 29:26:42features uh uh the test features in the
- 29:26:46same way. So we're going to perform the
- 29:26:48same transformation from the scaler on
- 29:26:50the test data. So that's something
- 29:26:52really important I want to note here is
- 29:26:53that we always scale both the training
- 29:26:59and test data. We always scale both. Of
- 29:27:02course, we're going to train the model
- 29:27:03on the training data. Um, but we are
- 29:27:07going to also test it on the testing
- 29:27:10data and it also needs to be scaled
- 29:27:12because our model that we build is going
- 29:27:14to assume scaled features. The
- 29:27:17coefficients that it learns are going to
- 29:27:19be assuming scaled features.
- 29:27:22So, we need to also scale our test data
- 29:27:26in the same way. So, we're doing that as
- 29:27:29well.
- 29:27:33So, we scale that and now we have our uh
- 29:27:37training and testing features have been
- 29:27:39scaled.
- 29:27:41No, we haven't replaced. We're going to
- 29:27:42do that. We have not yet. We're going to
- 29:27:45do that coming up in a minute. Yeah, we
- 29:27:48haven't done that. Um, it is it is the
- 29:27:51label. We definitely need to replace
- 29:27:53NLES. We just haven't done it yet
- 29:27:55because it's not in the features and
- 29:27:56we're doing all of our uh uh
- 29:27:58pre-processing to our pre-processing to
- 29:28:00our features.
- 29:28:10Yeah. So, we're definitely we need to
- 29:28:12we're going to in a minute.
- 29:28:16Okay. So, if you look at the data now,
- 29:28:18it's all been scaled. So, these are all
- 29:28:20um zcores. These are all on a much
- 29:28:23better scale now. Um, and these are we
- 29:28:28still have our one hot encoding features
- 29:28:29which are zero or one. So this scaling
- 29:28:34should lead to a better model than if we
- 29:28:36didn't scale. So scaling is really
- 29:28:39important. We can see that here.
- 29:28:46Okay.
- 29:28:48Now, um, let me ask you guys, were you
- 29:28:50able to run the scaling? Are you caught
- 29:28:53up to here? If you're following along,
- 29:28:55were you able to run the scaling?
- 29:29:08Okay, great. Great.
- 29:29:16Awesome.
- 29:29:19Okay. So, uh what we're going to do now
- 29:29:22is replace NLES in the uh replace NLES
- 29:29:28by calculating the median of the data.
- 29:29:32So, what I want you to notice is that we
- 29:29:35are taking the NLES now this is um this
- 29:29:39is on purpose is we are purposely taking
- 29:29:42the NLES um out of the median
- 29:29:45calculation. So we're skipping the NLES
- 29:29:47when we compute the median because we
- 29:29:49don't want those NLES to affect the
- 29:29:51median calculation.
- 29:29:53Um so we compute a median salary here
- 29:29:56and then we fill our NLES with the
- 29:29:59median salary um from the training data.
- 29:30:03So this is our choice. This is a choice
- 29:30:06um to use the median and it's also a
- 29:30:10choice to use the training set median
- 29:30:14for both train and test. What we could
- 29:30:18have done this is an alternative that we
- 29:30:20could have done is use the entire column
- 29:30:24and then um use the median of all of the
- 29:30:28data to replace. That's really up to us.
- 29:30:31Um this is one way of doing it. We could
- 29:30:34have done before we did the split. We
- 29:30:37could have um filled in with the median
- 29:30:41earlier. We chose to do it here mainly
- 29:30:44because it doesn't affect the features.
- 29:30:46So we could have done this earlier and
- 29:30:48did it before we did the split and
- 29:30:50filled the NAS. Um really doesn't it's
- 29:30:53doesn't matter that much which way we do
- 29:30:54it. Um but we do need to fill in NLES.
- 29:30:58We cannot have those be null when we
- 29:31:00when we put it into our model. So some
- 29:31:02way we need to fill in nulls. Um and so
- 29:31:06in this strategy we're filling in our y
- 29:31:08train um with the median salary from our
- 29:31:12training data. And same with this we're
- 29:31:14filling in with the median salary of the
- 29:31:16training data as well. But that's a
- 29:31:18choice. We could fill in with the mean
- 29:31:21with the average. Um we could fill in
- 29:31:25with the we could do it with all the
- 29:31:27data together before we split it. we
- 29:31:29could have filled in with all of the the
- 29:31:31median across the whole data set. Um
- 29:31:34either one works. You can do it either
- 29:31:36way, but we we did it um later here to
- 29:31:40show that it doesn't really affect the
- 29:31:42features. So, we can choose when we do
- 29:31:44it, right? It doesn't affect the
- 29:31:46features at all. So, we can do all of
- 29:31:48our pre-processing on the features and
- 29:31:50then do our label uh filling in NLES um
- 29:31:54if if we have them.
- 29:32:00uh x numerical. Um make sure you're
- 29:32:03running uh this
- 29:32:07uh x numerical was defined here
- 29:32:12when we split it apart um from
- 29:32:17uh when we dropped these columns here.
- 29:32:19So make sure you're running this. This
- 29:32:21is x numerical.
- 29:32:23It's defined there.
- 29:32:25So go back up this uh this cell
- 29:32:29where we split apart the y and we and we
- 29:32:31have the x here x numerical.
- 29:32:35Make sure you run this.
- 29:32:44Make sure you run this. And then you can
- 29:32:46run these. Then you run this to build x.
- 29:32:59All right.
- 29:33:01Are we up to here with this filling in
- 29:33:04the labels?
- 29:33:06Uh because then we can build our model
- 29:33:09once we're up to here. We've scaled
- 29:33:11everything. We filled in our NLES.
- 29:33:14We've gotten one hot encoding.
- 29:33:21Yeah, it is. That's why you know that's
- 29:33:23why we spend a lot of uh time on model
- 29:33:25on data preparation with pandas, right?
- 29:33:27That's why we did all that pandas work
- 29:33:29for sure. Yes, there is a lot of work
- 29:33:31before we can build a model.
- 29:33:34Yes, the mo do you guys notice that like
- 29:33:37the modeling is relatively easy. It's
- 29:33:38just a fit and predict. The modeling is
- 29:33:41actually really easy. It's all the other
- 29:33:43work that's that's more involved, right?
- 29:33:47more code.
- 29:33:50The modeling itself is really easy.
- 29:33:53It's just it's just one line of like
- 29:33:55ffit.
- 29:33:58Yeah, pretty easy to do.
- 29:34:03And then you do evaluation which is a
- 29:34:06couple lines.
- 29:34:27Yep. There's these are all the these are
- 29:34:30the common steps. All these steps we're
- 29:34:32doing are very very prototypical in
- 29:34:34model building is you let's just go back
- 29:34:37through this to see what we did, right?
- 29:34:38We imported our data. Um we analy we
- 29:34:43dropped this name column because it's
- 29:34:45not useful to us. So we dropped that. Um
- 29:34:48we filled in the NLES eventually. Um but
- 29:34:52you know if there were any nulls in our
- 29:34:54features we would have to deal with
- 29:34:55those as well by replacing them or
- 29:34:57dropping the rows like we did earlier.
- 29:35:00Um
- 29:35:01and then we do one hot encoding because
- 29:35:04of course we can't have any string
- 29:35:05columns going into our models. We got a
- 29:35:07one hot encode.
- 29:35:09Um we uh then build our X and Y by
- 29:35:13concatenating the one hot encoded back
- 29:35:16to the numerical features.
- 29:35:19Then we train test split. Right? That's
- 29:35:22pretty common. Or we could do cross
- 29:35:23validation either way. Um a kfold cross
- 29:35:27validation. Then we scale. So we didn't
- 29:35:30do this last time, but this is something
- 29:35:31we should get in the habit of is scaling
- 29:35:33um our features. So we do that. And now
- 29:35:37we're ready to model. So now we're ready
- 29:35:39to model. Um so that's this part.
- 29:35:44Okay. So let's do the model. Um the
- 29:35:48model is actually uh pretty easy to do.
- 29:35:50So we're going to use a lasso. So we
- 29:35:52have a lasso model here. Notice what
- 29:35:54we're setting our alpha to. So the big
- 29:35:56parameter we really need, ignore this
- 29:35:59iterations. We actually don't really
- 29:36:00need the we don't really need that
- 29:36:02parameter. Um so just ignore it for the
- 29:36:04moment. But the big one that we're
- 29:36:06setting here is the alpha. So when we
- 29:36:09did linear regression, we didn't need
- 29:36:11any parameters to go inside the linear
- 29:36:13regression object. We didn't need any
- 29:36:15parameters, right? Because there are
- 29:36:17really no parameters of it. But for
- 29:36:20lasso, the important one is the alpha.
- 29:36:22And so we need to know what to set alpha
- 29:36:24to. Um let's start with alpha equals 1.
- 29:36:29That's a good starting place. So a
- 29:36:31typical um starting point
- 29:36:35for alpha
- 29:36:38um
- 29:36:39is uh is one. So that's a typical
- 29:36:43starting point. And so we can set alpha
- 29:36:45equals to one. This max iterations is
- 29:36:48the the parameter that governs the
- 29:36:51training process because it is
- 29:36:53iterative. So if for some reason we we
- 29:36:56can't converge to the right betas and
- 29:36:59we've run it for 10,000 steps once we
- 29:37:01pass 10,000 steps, uh it will stop and
- 29:37:04just give us the betas at that point.
- 29:37:07But it will likely never hit this
- 29:37:09number. It'll converge before then. So
- 29:37:12um we don't really need to um specify
- 29:37:15it. So, I'm actually just going to get
- 29:37:16rid of it. Um, it's not really a big
- 29:37:18deal. It should converge before then.
- 29:37:21Um, but if if we want to set like a
- 29:37:24maximum step size in the optimization,
- 29:37:26we definitely could there. Uh, but not
- 29:37:30concerned about that too much. But
- 29:37:32here's our lasso. And then we're just
- 29:37:33going to do a fit on our data. So, look
- 29:37:36how easy that is. Just like a linear
- 29:37:38regression, lasso.fit,
- 29:37:42right? So, we do fit. Um,
- 29:37:50oh, I didn't run this. I'm sorry. I got
- 29:37:52to run this. Okay. Actually, that's a
- 29:37:56good example of what happens when you
- 29:37:57don't when you have nles, right? So, it
- 29:37:59says our our null contains nan. That's
- 29:38:02because I didn't run this. But now that
- 29:38:04should be filled in. Now, we should be
- 29:38:06able to run this. Okay, perfect. So it
- 29:38:09runs.
- 29:38:16Okay. So you can see what the intercept
- 29:38:18is. Um this is one of our coefficients,
- 29:38:20right? The intercept is 457. And look
- 29:38:23now what's really interesting about the
- 29:38:24coefficients is look at what some of the
- 29:38:27coefficients are. Some of them are
- 29:38:32actually zero, which is really So some
- 29:38:34of them ended up being at zero, which is
- 29:38:36very very interesting. that means that
- 29:38:39those features get cancelceled out and
- 29:38:42they're basically not part of the model
- 29:38:46which is really interesting. Um, so we
- 29:38:48have all these coefficients and some of
- 29:38:50them are zero.
- 29:38:56Yeah, negative0 is just because of the
- 29:38:59convergence like they started out
- 29:39:01negative and worked their way up to
- 29:39:03zero. it. Negative Z really just means
- 29:39:07zero, but they just were coming from
- 29:39:09they were like small negatives and ended
- 29:39:12up at zero
- 29:39:14during the training process. They were
- 29:39:16negative at one point and it ended up
- 29:39:18zero. Um
- 29:39:20so yeah, negative 0 just obviously means
- 29:39:23zero. Um it's still still zero there.
- 29:39:34So what's interesting is some of these
- 29:39:36features ended up uh being zero which
- 29:39:39you don't usually see in a linear
- 29:39:41regression. So if we were to train this
- 29:39:43using a linear regression we typically
- 29:39:45wouldn't see that but some of these
- 29:39:47turned out to be zero because again
- 29:39:49we're encouraging those betas to be
- 29:39:53small. we're encouraging them to be uh
- 29:39:55small and so um you know what happens is
- 29:40:00some of them can be shrunk all the way
- 29:40:02down to zero meaning those features
- 29:40:04don't contribute that's a really simple
- 29:40:06model at that point right so we've taken
- 29:40:09something complex that includes all of
- 29:40:12these features and actually reduced it
- 29:40:14into something simple that only includes
- 29:40:16these features
- 29:40:18right
- 29:40:21so that's what it does um now we need to
- 29:40:24evaluate this to see how good of a model
- 29:40:26it is. But that's what this is saying
- 29:40:29here in this text is that um a positive
- 29:40:32uh coefficient indicates that as the
- 29:40:34independent variable increases the
- 29:40:36dependent variable also increases.
- 29:40:38Negative coefficient means as the
- 29:40:40independent variable increases dependent
- 29:40:43decreases because it's reducing the
- 29:40:44value. Um and lasso is known for feature
- 29:40:49selection by shrinking some of them to
- 29:40:51zero effectively removing those
- 29:40:53variables from the model from the
- 29:40:55equation right
- 29:40:57um
- 29:40:59so that's what happens
- 29:41:03some of them end up being zero
- 29:41:07were you guys able to run this this
- 29:41:10lasso uh fit which is the training of
- 29:41:13the lasso Control.
- 29:41:32No, it doesn't ensure there's no
- 29:41:34overfit, but it helps with overfitting.
- 29:41:36It's supposed to help by making the
- 29:41:38model simpler. And this is definitely a
- 29:41:40simpler model because it's removing some
- 29:41:42of the features from the model
- 29:41:43essentially, right? Because some of the
- 29:41:46features aren't going to contribute.
- 29:41:47It's a simpler model.
- 29:41:50It doesn't it doesn't mean there's not
- 29:41:51going to be any overfitting, but it
- 29:41:53helps prevent it. That's what it's
- 29:41:55designed to do to help prevent it.
- 29:41:59Yeah. So higher coefficient. Yes. The
- 29:42:02higher coefficient means it's a more
- 29:42:04important feature towards the
- 29:42:06prediction. Yes. That's what it means
- 29:42:09for sure. The higher the magnitude, the
- 29:42:12more of a contributor towards that
- 29:42:14prediction. Uh it is. Yes.
- 29:42:23And it's not just it's it could be
- 29:42:25higher positive or negative there. Like
- 29:42:27a higher negative is also a pretty big
- 29:42:30factor,
- 29:42:32right? So So you want to think about it
- 29:42:33in terms of absolute value.
- 29:42:41does not guarantee but helps. Yes, it
- 29:42:43doesn't guarantee it but it's designed
- 29:42:45to help overfitting, help prevent it.
- 29:42:47Yes, absolutely.
- 29:43:13Okay.
- 29:43:15So let's do some evaluation. Um so let's
- 29:43:19do in this case we are going to do our
- 29:43:22predict
- 29:43:42Oh, yeah. I'm not sure why that's the
- 29:43:45case.
- 29:43:47Interesting.
- 29:44:07We could try increasing the um max
- 29:44:11iterations.
- 29:44:23Okay, that's why. Yeah. So then you get
- 29:44:25that result with the with the higher max
- 29:44:27iterations.
- 29:44:29It doesn't get cut off there.
- 29:44:33I think that's why you probably left
- 29:44:35this in there,
- 29:44:37which is fine. You get about the same
- 29:44:38numbers.
- 29:44:46Yeah.
- 29:44:56All right. Let's evaluate this. So,
- 29:44:57we're going to to to do evaluation. I
- 29:44:59want you guys to see again. We should
- 29:45:01get in the habit of doing evaluation,
- 29:45:04which is taking our model and predicting
- 29:45:07on the training and predicting on the
- 29:45:09test sets, right? So we predict on the
- 29:45:11train set and calculate our MSE
- 29:45:15and we um calculate our R2 score um or R
- 29:45:22squar score I should say. Uh but again
- 29:45:25the MSE is the one we're really going to
- 29:45:26use mostly. Um but we calculate so we do
- 29:45:30our predictions and then we compare that
- 29:45:32into our mean squared error with our
- 29:45:34labels
- 29:45:36and we uh go ahead and do the same thing
- 29:45:40with the test. Right? So we do uh
- 29:45:42lasso.predict
- 29:45:44on our test features and we go ahead and
- 29:45:47compare that with the test labels. And
- 29:45:51so what we're doing there is generating
- 29:45:52our MSE.
- 29:45:55So, we we take a look at our MSE and we
- 29:45:58get uh 84,000
- 29:46:01MSE. Um, and so, of course, we could
- 29:46:05take the um what we could do with that
- 29:46:08is take a look at the um MSE on the uh
- 29:46:14we could do um MP. Square root
- 29:46:19and do the square root of the MSE test.
- 29:46:25and we get um 340. So this would be in
- 29:46:29the units of our label. So, we go back
- 29:46:32and look at our label um for some of
- 29:46:34those um
- 29:46:48so uh we are in 300s and our data is
- 29:46:52like right around the 500. So, of
- 29:46:54course, if we describe this um we could
- 29:46:56see what the statistics are of it. So,
- 29:46:59we could do df.describe describe and
- 29:47:01generate that. But that doesn't look
- 29:47:02like a very good error, right? If these
- 29:47:04are in the 400s, um that's that's not a
- 29:47:07very good error.
- 29:47:09So again, it's not a very great model.
- 29:47:12But one thing I want you to see is that
- 29:47:13it's it's not overfitting.
- 29:47:16Um if anything, it's actually
- 29:47:18underfitting, which is what this kind of
- 29:47:21um MSE suggests, right? because our
- 29:47:23error here is 84 uh excuse me 84,000.
- 29:47:29Um
- 29:47:34our our area here is 84,000
- 29:47:38excuse me and on the test set it's
- 29:47:41116,000.
- 29:47:43Um so these two errors are both bad. So
- 29:47:48it's not overfitting. This is actually
- 29:47:51underfitting. So it's not overfitting,
- 29:47:53it's actually underfitting. Um, and so
- 29:47:56that's the risk with something like
- 29:47:57lasso is that it's making the model a
- 29:48:00bit too simple and we actually risk
- 29:48:03underfitting, which is what happens. We
- 29:48:06have too much error across both the
- 29:48:09training and the test set. Overfitting
- 29:48:12is when we do we have really good
- 29:48:14performance on the training set, but bad
- 29:48:16performance on the test set. We're not
- 29:48:18overfitting.
- 29:48:20um we are uh underfitting because our
- 29:48:24performance is not good either way. Even
- 29:48:26this R squar is pretty low. It's not
- 29:48:28even at 50%.
- 29:48:36Okay, so that's so we we do the
- 29:48:39evaluation and again the evaluation just
- 29:48:40comes down to making predictions and
- 29:48:43computing our error amongst those
- 29:48:45predictions to our labels. That's always
- 29:48:47what the uh evaluation is going to be
- 29:48:54for MSE.
- 29:48:57What's the ideal MSE? What do you think
- 29:49:00it should be? What is So, think about it
- 29:49:03like this. The MSE represents the
- 29:49:05average distance between our predictions
- 29:49:10and the labels.
- 29:49:13So, if we're getting it right all the
- 29:49:16time, what's that distance going to be
- 29:49:18if we're always right? What's our
- 29:49:20distance from what's our distance from
- 29:49:23our predictions to our labels going to
- 29:49:26be if we're always getting it right?
- 29:49:29Zero. Yeah, there's not going to be any
- 29:49:31distance. It's going to be right. It's
- 29:49:33going to be perfectly aligned, right?
- 29:49:35There's going to be no distance there.
- 29:49:37So, yeah, an ideal MSE is zero. That's
- 29:49:41an ideal MSE.
- 29:49:43So, anything close to like the smaller
- 29:49:46the better for MSE. The smaller the
- 29:49:49better. Um, for this R squared, uh, it's
- 29:49:53it's a scale between 0 to one where one
- 29:49:56is the best. So, one would be perfectly
- 29:49:58aligned predictions. Um, so, and again,
- 29:50:02this this is we actually multiply by 100
- 29:50:05to get uh because it's it's a number
- 29:50:07between 0 and one. So we get about 47%
- 29:50:10which is not good.
- 29:50:21Okay.
- 29:50:36All right. Any questions on this
- 29:50:39evaluation?
- 29:50:54All right. I want to show you something
- 29:50:56which is
- 29:50:58Yeah, this that's true. The scale of it
- 29:51:01matters on the data because we should be
- 29:51:03you should always interpret your MSE in
- 29:51:05the scale of
- 29:51:07um your your labels because your labels
- 29:51:12like in this case our labels um you know
- 29:51:14we could take uh for example we could
- 29:51:17easily let's actually do that let's take
- 29:51:20the average
- 29:51:22let's take the average of our labels on
- 29:51:25the training data
- 29:51:29and and we could see what those are. Um,
- 29:51:32so the average is 500,
- 29:51:35right? The average is 500. And look at
- 29:51:38what our uh square root of our MSE is,
- 29:51:41which is in the same units as our
- 29:51:43original. Um, so we have uh quite a bit
- 29:51:47of error. 340 when our units are right
- 29:51:50around 500.
- 29:51:53So that's quite a bit of error.
- 29:52:03Yeah, MSE of zero means our our uh our
- 29:52:06predictions are nearly identical to the
- 29:52:10test labels. Yes, that's what MSE of
- 29:52:13zero means. There's zero distance.
- 29:52:18So closer to zero, the better.
- 29:52:27But we talked about it as you you really
- 29:52:29so the rule of thumb should be what is
- 29:52:34your RMSSE as a percentage of your
- 29:52:37typical value. So your typical value is
- 29:52:40in the 500s. Our our RMSSE is 340.
- 29:52:45That's just really high. That's over
- 29:52:47like 60% of that value.
- 29:52:51So that's just a lot. That's too much
- 29:52:54error. What we would love this RMSSE to
- 29:52:56be is under 20% of the typical value. So
- 29:53:00that means on average we are 20% or less
- 29:53:05off in our prediction. That would be
- 29:53:08good. That would be pretty good. That
- 29:53:10means we're like 80% accurate,
- 29:53:14right? That'd be pretty ideal. So you
- 29:53:16got to think about it in terms of this
- 29:53:17RMSSE, which is in the same units as
- 29:53:20your labels.
- 29:53:23This is the
- 29:53:27RMSSE
- 29:53:30which is in the same units as the
- 29:53:34labels.
- 29:53:37So and then to interpret this we have
- 29:53:40340
- 29:53:42is compared to
- 29:53:45typical
- 29:53:47um salary unit of 500
- 29:53:52right so this is uh quite a bit when the
- 29:53:55typical value is 500 and we are off on
- 29:53:58average by 340 units
- 29:54:01that's so much relative to the typical
- 29:54:04value
- 29:54:06that's just too. That's a lot of error.
- 29:54:08That's not a very good model, right?
- 29:54:11It's underfitting. It's definitely
- 29:54:13underfitting.
- 29:54:25Yeah. So, that's a great question. What
- 29:54:26should we do from here? So, um because
- 29:54:29we're underfitting
- 29:54:37um we should use a more complex model.
- 29:54:42So uh we're going to learn about those
- 29:54:45in lesson four, but we should use
- 29:54:47something different. This linear
- 29:54:48regression is still too basic. Even with
- 29:54:50lasso, it's still too basic.
- 29:54:56Yeah, we're underfitting because we But
- 29:54:59it could also be we're underfitting with
- 29:55:00a regular linear regression. We should
- 29:55:02test that out. Um, and maybe it would be
- 29:55:04an exercise for you guys um to test that
- 29:55:08out yourself. It shouldn't be hard to
- 29:55:10do. Um, you already have all the data
- 29:55:12scaled. You So, do you see how you would
- 29:55:15do that? You would just come in here and
- 29:55:17build a linear regression rather than a
- 29:55:18lasso and dofit and then you would
- 29:55:21evaluate it the same way with a
- 29:55:22dotpredict. It's really easy to do that.
- 29:55:26And then we can compare that um to to
- 29:55:29this. It shouldn't be that hard to do
- 29:55:32that, right?
- 29:55:34And something you guys could do for
- 29:55:35sure. Um,
- 29:55:38is build the linear regression and
- 29:55:41actually compare it and see what kind of
- 29:55:44difference it makes. I mean, we honestly
- 29:55:46we could do it ourselves. We could do it
- 29:55:48right now. Maybe it's worth trying that.
- 29:55:52So, let's build a linear regression
- 29:55:56for comparison.
- 29:56:00So we have our linear regression
- 29:56:04uh is linear regression and then we do
- 29:56:08ffit linear regression.fit fit
- 29:56:12right so so this will train it um and
- 29:56:16then we can evaluate it so lin MSE is
- 29:56:21mean squared error
- 29:56:24and then we can do our um let's do our
- 29:56:27training let's do the training and then
- 29:56:32um let's predict
- 29:56:34actually let me do that here
- 29:56:37uh y prediction
- 29:56:40train
- 29:56:43linear
- 29:56:45equals um linear regression.predict
- 29:56:50and then we're going to predict on our
- 29:56:52training features.
- 29:56:57Okay, do you guys see what I'm doing?
- 29:56:58I'm building a linear regression for
- 29:57:00comparison.
- 29:57:02I'm doing dofit here to train it and
- 29:57:05then I'm making some predictions on the
- 29:57:06training set and we're going to evaluate
- 29:57:09those. I'm going to replace that here
- 29:57:10with y prred
- 29:57:14uh train
- 29:57:17linear. So these predictions
- 29:57:42Okay. So, if you guys want this code, I
- 29:57:44can paste it in.
- 29:57:55So, let's see what the RMSSE for just a
- 29:57:58linear model is.
- 29:58:01It's a little bit better. It's better
- 29:58:03for sure.
- 29:58:06So 289 is better than this 340. It's
- 29:58:10better. It's getting closer to zero.
- 29:58:13It's still underfitting though,
- 29:58:17right? And that's just on the training
- 29:58:18set. Let's look at the Let's do the same
- 29:58:21thing, but on
- 29:58:25Let's change this. Let's swap this out
- 29:58:27for um test.
- 29:58:31And then let's do test.
- 29:58:35And then let's do test
- 29:58:41test.
- 29:58:44and then
- 29:58:47test.
- 29:58:58Okay, so this is producing test
- 29:58:59predictions on the test set.
- 29:59:02We are generating an MSE test
- 29:59:07and then we're doing MSE test
- 29:59:10which is using the test labels and our
- 29:59:12test predictions
- 29:59:14and then we take the square root of that
- 29:59:15for RMSSE and then we're going to
- 29:59:17generate that. So it's still under fit.
- 29:59:20I mean this is still high. This is still
- 29:59:23high um on the test set and versus on
- 29:59:25the training set. So it's still pretty
- 29:59:27high. Um, even the basic linear
- 29:59:29regression is under is still
- 29:59:31underfitting. Still underfitting, right?
- 29:59:34Even without the lasso, which is lasso
- 29:59:37is supposed to help with overfitting.
- 29:59:39It's definitely not overfitting. Um,
- 29:59:42it's definitely underfitting,
- 29:59:44but this is a signal that it's kind of
- 29:59:46overfitting because this is performing
- 29:59:47better on the training data and then it
- 29:59:49gets worse on the test data.
- 29:59:53Definitely gets worse, right?
- 30:00:03Did you guys follow?
- 30:00:07I'm just running this above I'm running
- 30:00:09this above this. It doesn't matter where
- 30:00:12you put it. We could uh we could move it
- 30:00:13down.
- 30:00:20We could move it down to I just ran I
- 30:00:23just picked a new cell right here.
- 30:00:25and ran it. But we could move it.
- 30:00:27Actually, let's do that. Let's move it
- 30:00:29down
- 30:00:33to
- 30:00:36after the lasso evaluation.
- 30:00:41Okay. So, I just moved it there.
- 30:00:44And then let's move
- 30:00:47this down.
- 30:00:49So, I just put it here after the um
- 30:00:52after this. So this is the um this is
- 30:00:56basically the objective function right
- 30:00:59of the training process. So during the
- 30:01:02algorithm that runs when we call ffit in
- 30:01:05scikitlearn it's going to find these
- 30:01:08betas right it's actually going to learn
- 30:01:10what these best betas are for our model.
- 30:01:14Um this is our model here right it's the
- 30:01:16combination of betas times our features
- 30:01:19um plus an intercept beta. Uh so that's
- 30:01:23our model but um we penalize those large
- 30:01:26uh weights in absolute value by um
- 30:01:30adding a penalty term like this um where
- 30:01:33alpha is some level of penalty that we
- 30:01:37want to provide. Usually alpha equals 1
- 30:01:39is okay. But um actually what we're
- 30:01:41going to learn uh to finish out this
- 30:01:43section is there's going to be a
- 30:01:44systematic way we can test out different
- 30:01:46alphas um that represent the level of
- 30:01:49penalty we want to uh apply to lasso or
- 30:01:52even ridge
- 30:01:54uh regression. So that was the lasso and
- 30:01:57um if you guys remember using it was
- 30:01:59super easy. Uh we worked through this
- 30:02:02problem with this um baseball data um
- 30:02:06and we had uh
- 30:02:09let's see scrolling down we um split out
- 30:02:12our numerical data and we did uh we one
- 30:02:16hot encoded our our categorical data
- 30:02:19combined it back together. Hopefully
- 30:02:21that um rings a bell there. Um and we
- 30:02:24actually scaled our data which is pretty
- 30:02:26standard to do is we do some type of
- 30:02:28scaling to our features especially our
- 30:02:30numerical features right want to scale
- 30:02:32those in some way whether it's minmax
- 30:02:34scale or standard scaler um want to do
- 30:02:37that and so we did that for this example
- 30:02:39and then we um ran the lasso regression
- 30:02:44which is pretty easy to use. You just
- 30:02:45use the lasso object and you pick an
- 30:02:47alpha here. Um, again, we are going to
- 30:02:51have a way to test out different alphas
- 30:02:54that could be candidates and we can see
- 30:02:56which one's the best. Um, so I'm going
- 30:02:59to show us that today coming up shortly.
- 30:03:03But that was that was the lasso. If you
- 30:03:05guys remember, we did that. Um, this it
- 30:03:08we compared that to a basic linear
- 30:03:10regression which is just this pretty
- 30:03:12straightforward just a fit and then
- 30:03:14predict and then we can generate mean
- 30:03:16squared error. Um, still not a very good
- 30:03:19mean squared error on this data, it's
- 30:03:21still fairly large. Um, so it's still
- 30:03:25not, no matter which model we use, it's
- 30:03:27still not very good, but at least we can
- 30:03:29practice doing that comparison. That's
- 30:03:30what we did last time. We did this on
- 30:03:33Wednesday.
- 30:03:34Um
- 30:03:36and then
- 30:03:38we saw that the effect of different
- 30:03:40alphas we had a lasso um
- 30:03:43we had a lasso uh cross validation
- 30:03:46example here. So beyond just using a
- 30:03:48regular lasso model that um scikitlearn
- 30:03:50has a lasso cv which allows you to try
- 30:03:53out different alphas uh with cross
- 30:03:56validation and um figure out what the
- 30:03:58best alpha is. Um, now we're actually
- 30:04:01going to have a different strategy
- 30:04:03that'll instead of just picking random
- 30:04:05ones, we can actually um supply multiple
- 30:04:08parameters that we may want to test um
- 30:04:11as many as the models may support. And
- 30:04:13in some more complex models will have
- 30:04:16more than one parameter like lasso only
- 30:04:18has the alpha. Um, technically it also
- 30:04:21has its max iterations, but really the
- 30:04:23only one that matters is this alpha.
- 30:04:26Other models have many more
- 30:04:27hyperparameters that we can um uh change
- 30:04:32and so we want a way to systematically
- 30:04:34test out those different combinations
- 30:04:37and to see which one leads to the best
- 30:04:39uh version of that model. Let's say the
- 30:04:41best results. So um we're going to
- 30:04:44explore that coming up. So we had lasso.
- 30:04:48Um now this is where we ended last time.
- 30:04:50We had ridge regression. If you guys
- 30:04:52remember, this one is just a slightly
- 30:04:55different penalty. Um,
- 30:04:58it takes the it I drew it out for us. It
- 30:05:01takes the same penalty we had before.
- 30:05:03So, it has that um residual sum of
- 30:05:06squares error, which is the main one we
- 30:05:09used for linear regression, but it has a
- 30:05:11penalty with an alpha and then it has
- 30:05:13the sum of the beta squares
- 30:05:17beta squares. So it penalizes it has a
- 30:05:22penalty but it penalizes slightly
- 30:05:24differently where it uses the square not
- 30:05:26the absolute value. That's the ridge
- 30:05:28regression. And this has the similar
- 30:05:30effect of you don't in order to minimize
- 30:05:33this right because our goal in training
- 30:05:34a model was to minimize this thing
- 30:05:38minimize this um quantity and find the
- 30:05:41best betas that minimize this. Um so
- 30:05:44generally yes you want to encourage
- 30:05:46lower values but the um once you get
- 30:05:50values that are a fraction if you square
- 30:05:52them they actually get smaller. Um so uh
- 30:05:56it's it's not um it's not necessary to
- 30:06:01shrink them all the way to zero. They
- 30:06:03will get smaller as soon as they're kind
- 30:06:05of below one. Um so they don't encourage
- 30:06:08it to completely go away uh like the
- 30:06:12absolute value does. It's just slightly
- 30:06:13different minimization. Um so what we
- 30:06:15see with the ridge is we don't see the
- 30:06:18features kind of get wiped out
- 30:06:19completely like we do with a lasso. In
- 30:06:21lasso they get encouraged to be um to
- 30:06:24become zero because that's kind of the
- 30:06:26only way to minimize an absolute value.
- 30:06:28But with squares they can keep getting
- 30:06:30smaller and smaller and smaller um
- 30:06:32fractions and they don't have to become
- 30:06:35zero. It's not as harsh of a of a
- 30:06:38penalty.
- 30:06:39Um,
- 30:06:41so, uh, the ridge was easy to use as
- 30:06:45well. Um, and it also has an alpha that
- 30:06:50we can set. So, it's literally the same
- 30:06:52exact code, just a different model.
- 30:06:55There's slightly different penalty and
- 30:06:57it results in different coefficients.
- 30:06:59You notice that none of them are exactly
- 30:07:00zero. Like with the lasso, you can get
- 30:07:02ones that are exactly zero. We don't see
- 30:07:05that with the ridge. You remember that?
- 30:07:08Um so we we s pointed out that last
- 30:07:10time. Notice the coefficients aren't
- 30:07:12zero. Um and then we can evaluate it. So
- 30:07:15we did our MSE calculation which is a
- 30:07:17pretty standard thing where we use our
- 30:07:19model to predict on a training set,
- 30:07:21predict on a test set, evaluate those um
- 30:07:25by computing the metric like the mean
- 30:07:27squed error and we can see if we're
- 30:07:29overfitting, underfitting. This is
- 30:07:31definitely the same kind of story we've
- 30:07:33seen with all these models is
- 30:07:34underfitting because the error is so big
- 30:07:36across both sets
- 30:07:38across training and tests. So it's it's
- 30:07:40definitely underfitting.
- 30:07:42Um
- 30:07:44and same thing as lasso, it has a cross
- 30:07:46validation uh variation on it that
- 30:07:49allows you to try out different alphas
- 30:07:52and um do different folds. So 10fold,
- 30:07:56fivefolds, whatever and compute the um
- 30:08:00try to find the best alpha that way.
- 30:08:03Okay.
- 30:08:07All right.
- 30:08:09Any questions on this so far from last
- 30:08:12time from reviewing that a little bit?
- 30:08:15Hopefully that uh hopefully that is
- 30:08:18jogging your memory a little bit on
- 30:08:20ridge and lasso. Um, you know, where
- 30:08:23we're going to pick it up today is to
- 30:08:25finish out this lesson with one more
- 30:08:27model,
- 30:08:29which is going to be a combination of
- 30:08:32ridge and lasso. So, you can actually
- 30:08:34combine them together
- 30:08:37um in a linear fashion, those penalties.
- 30:08:40So, you can actually have both
- 30:08:41penalties, the absolute value and the
- 30:08:43square. And when you have both penalties
- 30:08:47um that's a special model called the
- 30:08:49elastic net uh regression or elastic net
- 30:08:53model. Um so this is a combination of
- 30:08:57lasso and ridge together. So you have
- 30:08:59lasso, you have ridge and then you have
- 30:09:01elastic net which combines both of those
- 30:09:03penalties. Um let me show you the
- 30:09:06equation.
- 30:09:08So here is the uh so here is the the
- 30:09:13model. This is the same that we've
- 30:09:15always had. This is our usual u model
- 30:09:20fitting for linear. This is a basic
- 30:09:22linear regression um loss function or
- 30:09:25objective function that we're trying to
- 30:09:26minimize to find the betas. Notice how
- 30:09:29we have both of our penalties though
- 30:09:30this time. So instead of just having one
- 30:09:32of the penalties, we actually have both.
- 30:09:34So we have the lasso penalty
- 30:09:38and then we have the ridge penalty here.
- 30:09:40So we actually use both of them and um
- 30:09:44try to find a balance of minimizing
- 30:09:47those two uh those two penalties.
- 30:09:51Okay. And notice how they instead of
- 30:09:53just a single alpha, we kind of have a
- 30:09:55balance on both of them.
- 30:09:58So we can actually weight the lasso one
- 30:10:01more. We can weight the ridge one more.
- 30:10:03We can weight them the same. Uh we can
- 30:10:07um change that around as much as we
- 30:10:08want. So they have two different weights
- 30:10:10there um that they could be.
- 30:10:14Um now what happens in reality is uh
- 30:10:19we're going to see this in the model is
- 30:10:21that um usually what happens is these
- 30:10:24get combined into a fraction. So there's
- 30:10:28usually a ratio of lambda 1 to lambda 2
- 30:10:32and this is known as the um this is
- 30:10:35sometimes known as the L1 ratio
- 30:10:38and this is a this is a a parameter
- 30:10:40inside the model that we'll be able to
- 30:10:42set um along with alpha. So we'll be
- 30:10:45able to set an alpha and then this
- 30:10:47ratio. Um the idea is is that um the
- 30:10:52ratio will uh allow us to control which
- 30:10:56one is more dominant. So if this number
- 30:10:59is bigger the um this lasso penalty will
- 30:11:03will be weighted more. If this ratio is
- 30:11:06smaller if it's less than one for
- 30:11:09example that means that the um ridge
- 30:11:11regression is more uh dominant. Um but
- 30:11:16the so we'll have this we'll have really
- 30:11:18this and this at our disposal and alpha
- 30:11:23is um
- 30:11:26alpha is kind of like a a you can think
- 30:11:28of it as a scale that is um so lambda 1
- 30:11:33kind of like lambda 1 plus lambda 2 um
- 30:11:37combined to equal alpha.
- 30:11:40So it's like our total level of penalty
- 30:11:43um our total level of penalty and we can
- 30:11:46set that equal to one. We can set it
- 30:11:48equal to whatever we want. Um and so
- 30:11:51these will be in this ratio and there'll
- 30:11:53be a total level of penalty that we can
- 30:11:55apply. So the model will actually use
- 30:11:58these two parameters when we when we do
- 30:12:00it. But that's how they're that's how
- 30:12:02they're all related.
- 30:12:05Okay. So ridge uses both penalties.
- 30:12:08That's the only difference between lasso
- 30:12:10or sorry elastic net uses both
- 30:12:12penalties. Um so one thing I want you to
- 30:12:15notice is that uh if we um if we want we
- 30:12:21could set this L1 ratio all the way to
- 30:12:23zero.
- 30:12:25Um which uh if we do that um the only
- 30:12:30way this L1 ratio could be zero would be
- 30:12:32if lambda 1 is zero. So it would just
- 30:12:34revert back to ridge regression. So it
- 30:12:36complet if if this is zero this will
- 30:12:39wipe out this term and we'll be back to
- 30:12:40ridge if the L1 ratio is zero.
- 30:12:46Okay.
- 30:12:50All right. So we have a elastic net
- 30:12:53model. Um now it's used the exact same
- 30:12:57way as we did the other models in the
- 30:12:59code. So we have elastic net um uh from
- 30:13:03the scikitlearn linear model family just
- 30:13:06exactly where we had linear regression
- 30:13:09lasso ridge all of those came from this
- 30:13:12linear model um elastic net also comes
- 30:13:15from there and then the cross validation
- 30:13:16version also comes from there um
- 30:13:21so let's see so when we build our model
- 30:13:23it's going to be um very very simple
- 30:13:26easy stuff because it's the same code
- 30:13:28that we always have um we just use the
- 30:13:31elastic net. We set an alpha alpha
- 30:13:34equals 1 is pretty standard um just like
- 30:13:37it is in in the last one ridge that's
- 30:13:39industry standard is one and then an L
- 30:13:42L1 ratio of.5
- 30:13:44that's pretty standard as well. What the
- 30:13:45L1 ratio of.5 is is kind of a um
- 30:13:51uh kind of a that means that the lambda
- 30:13:541 to lambda 2 ratio is 1/2. Um, so
- 30:13:58that's that's a pretty standard uh ratio
- 30:14:00as well. But again, we could set this
- 30:14:03equal to one and they'd be kind of
- 30:14:05equally weighted. Um, L1 ratio of a half
- 30:14:08means that the uh ridge regard the the
- 30:14:12ridge penalty is a little bit more
- 30:14:14weighted uh in that in that situation.
- 30:14:19Okay.
- 30:14:21So uh once we have this model um we can
- 30:14:24do ffit and we can run that on our
- 30:14:26training data and we can um get we can
- 30:14:30figure out what our parameters are like
- 30:14:32our coefficients and our intercepts. Our
- 30:14:33model will have that but more
- 30:14:35importantly we can use our model to
- 30:14:36predict right so we can predict on the
- 30:14:38test set. Um let me go back and load our
- 30:14:42data and actually run this.
- 30:14:47So, we're going to be using the same
- 30:14:48data that we did for uh lasso,
- 30:14:54which is the I'm scrolling back up so I
- 30:14:57can load it. It's the baseball data
- 30:14:58here.
- 30:15:02Um,
- 30:15:05just run it from there.
- 30:15:07It's this hitters.csv. So, hopefully you
- 30:15:10have that one.
- 30:15:22Let me load this.
- 30:15:30Okay, so we loaded that and then that
- 30:15:32should load.
- 30:15:34Drop that unnamed column.
- 30:15:41We will get our dummies
- 30:15:50and then concatenate those split
- 30:15:56scale. I'm just rerunning things. I'm
- 30:15:58rerunning things so we can see our model
- 30:16:00one more time.
- 30:16:02So rerun that. Take a look at that. That
- 30:16:04looks good. and then
- 30:16:07fill in the nles on the on those.
- 30:16:10Okay. So, we should be able to run our
- 30:16:14uh elastic net now.
- 30:16:23Okay. So, let's import that and then
- 30:16:26let's build our model. So, there we go.
- 30:16:28We build our model and the intercept is
- 30:16:31that. Now, of course, we can look at our
- 30:16:33coefficients. Let's look at that.
- 30:16:39Look at our coefficients. So remember
- 30:16:41the coefficients are the uh betas. These
- 30:16:43are our betas that are in our model. Um
- 30:16:46so we can take a look at those. Now um
- 30:16:49they're it's somewhere in between. It's
- 30:16:51not a full lasso where we're going to
- 30:16:52see some of these be zero. It's not a
- 30:16:54full ridge. Um so the coefficients we
- 30:16:57get are different. They're somewhere in
- 30:16:59between there those two models that
- 30:17:01we've already built. So not quite the
- 30:17:03same um somewhere in between there.
- 30:17:10Um and then we can use our model to make
- 30:17:12predictions and and compute the MSE
- 30:17:15uh or the RMSSE I should say as well. So
- 30:17:18we can take the mean squared error, pass
- 30:17:20that into the square root and comput the
- 30:17:21RMSSE. So still pretty bad. Um this is
- 30:17:24right around that 300 range of what
- 30:17:26we've gotten for our other RMSSE. So,
- 30:17:28it's not like elastic net is any better
- 30:17:31than those other like linear or lasso or
- 30:17:34ridge. And that's not surprising because
- 30:17:36it's just adding those extra penalties.
- 30:17:38We don't expect it to magically get
- 30:17:40better. It's actually a more complex
- 30:17:42um when we add when we add those in,
- 30:17:46we're actually reducing it and making it
- 30:17:47simpler. And we need something more
- 30:17:49complex, I should say. So, we're making
- 30:17:51it simpler um by by making penalizing
- 30:17:56our weights a little bit more. And so
- 30:17:58it's still not a good fit. That's not
- 30:18:01really surprising, right? It's still not
- 30:18:03really a great fit.
- 30:18:05And we can we can even double check
- 30:18:07that. We know our RMSSE is pretty bad.
- 30:18:10Um but we can double check it with this
- 30:18:12R2 score. And it's, you know, still not
- 30:18:15good. Remember, a one would be really
- 30:18:16good. Um that'd be like a perfect linear
- 30:18:18model. This is um still pretty bad.
- 30:18:26Okay, so as we said, the alpha controls
- 30:18:28the overall strength. Um, so the higher
- 30:18:31the alpha, the more overall penalty
- 30:18:34we're supplying, which makes the model
- 30:18:37simpler. Um,
- 30:18:40uh, but the L1 um ratio determines the
- 30:18:43mix or that ratio of the lambdas, the
- 30:18:47lasso to the ridge. Um, if you have it
- 30:18:51be um exactly zero, you you revert all
- 30:18:55the way back to um if you if you put it
- 30:18:59at zero, you revert all the way back to
- 30:19:00ridge one would be all the way to pure
- 30:19:02lassos. Somewhere in between like one
- 30:19:04half is is good.
- 30:19:15Okay,
- 30:19:17so this is another example of trying out
- 30:19:21different values of alpha in the CV to
- 30:19:23see which one works. Now again, I'm
- 30:19:25going to show us in a minute a
- 30:19:26systematic way to do this, but this is
- 30:19:29just trying out um different alphas that
- 30:19:31we set up in this uh in this um
- 30:19:36uh range. So we have different uh values
- 30:19:39between minus2 and two um
- 30:19:42logarithmically.
- 30:19:44Um so these are uh logarithm values that
- 30:19:47are between this between minus2 and two
- 30:19:49and we choose 100 different alphas and
- 30:19:52then we choose 100 different um L1
- 30:19:54ratios between 0.01 and one and we run
- 30:19:58that we run this um cross validation
- 30:20:00with 10 folds. So this is quite a bit.
- 30:20:02So we're doing 10 folds and we're trying
- 30:20:05out 100 different um options. Uh every
- 30:20:09time we do an option we're trying out 10
- 30:20:11folds to evaluate it. So, it's going to
- 30:20:13take a minute to run.
- 30:20:26It's still running here. But again, what
- 30:20:29this is doing is trying out different
- 30:20:30alphas and it's it's going to do a cross
- 30:20:34validation. And you guys remember the
- 30:20:36t-fold cross validation is where we take
- 30:20:38our data and we divide it into 10 folds
- 30:20:43and then we um train on nine of those
- 30:20:46and then test on the remaining fold and
- 30:20:48then we rotate all the folds 10 times.
- 30:20:51and that we average those mean squared
- 30:20:54error metrics together um against those
- 30:20:5810 different uh fold options to generate
- 30:21:02a basically like an average performance
- 30:21:05for that value of alpha. And we're doing
- 30:21:07that 100 times for all these different
- 30:21:09100 alphas that there are and 100
- 30:21:12different L1 ratios that we're trying
- 30:21:13with them.
- 30:21:19So that's quite a bit of processing but
- 30:21:22uh it did finish.
- 30:21:26So we can see what our best alpha is and
- 30:21:28our best one ratio. So we get the best
- 30:21:30alpha is this best one ratio is this. Um
- 30:21:34and therefore we can uh build a model
- 30:21:37with those with just these two guys as
- 30:21:39the alpha and the L1 and um see how that
- 30:21:44performs.
- 30:21:46We build that model and then we predict
- 30:21:48on the test set and we generate the
- 30:21:50RMSSE. It's just a little bit better.
- 30:21:52It's still not It's just a little bit
- 30:21:54better, but it's still not good, right?
- 30:21:56It's still 338. It is just way too big.
- 30:22:00Remember, this is RMSSE, so it's in the
- 30:22:03units of our uh target variable. So,
- 30:22:07it's in the units of, if we go back to
- 30:22:10our data, actually, I can just print it
- 30:22:12out here.
- 30:22:14um this RMSSE.
- 30:22:18If I just do this, we could take a look
- 30:22:20at um DF
- 30:22:22or I could look at Y test
- 30:22:27and you can see some of these values.
- 30:22:28These are these salary values in the
- 30:22:30hundreds, right? Some of them are in the
- 30:22:31thousands. Um but an error of like 338
- 30:22:36is just too big. That's a really big
- 30:22:38error. That means we would be off by an
- 30:22:39average of 300 when our our values if we
- 30:22:42just do the mean
- 30:22:46um
- 30:22:48is only 550 as on average is 550 but we
- 30:22:52have this amount of error on average um
- 30:22:55so that's just a way too big of a
- 30:22:57proportion of error right it's not a
- 30:22:59very good model and again we can verify
- 30:23:02that by looking at this R2 for.
- 30:23:11So if we go down here,
- 30:23:20still not very good.
- 30:23:23Here's some of our coefficients. So
- 30:23:25remember, you can always take your
- 30:23:26coefficients and line them up to your
- 30:23:28your data columns. Uh so that you can
- 30:23:31get a sense of what coefficient belongs
- 30:23:33with what feature. So that's all we're
- 30:23:36doing here is just creating a series
- 30:23:37where those coefficients instead of just
- 30:23:39printing out the coefficients, we're
- 30:23:40actually lining them up to the columns.
- 30:23:42So this tells us um remember the larger
- 30:23:45it is the more influence it kind of has
- 30:23:47on the on the final result. Um either
- 30:23:50way, so like this has a big negative
- 30:23:52influence. Um, this has a large positive
- 30:23:55influence.
- 30:24:02Okay, let me pause there. Any questions
- 30:24:05about the
- 30:24:07elastic net model?
- 30:24:11This is a really this model is a really
- 30:24:13good one to use when you are building a
- 30:24:16linear regression and it's performing
- 30:24:18well, but it's overfitting. This is a
- 30:24:20really good one to use because you can
- 30:24:21balance
- 30:24:23lasso and ridge you can get the best of
- 30:24:25both worlds. So the the main strategy is
- 30:24:28if you are using a linear regression and
- 30:24:31you see overfitting
- 30:24:33um meaning that it's performing decently
- 30:24:36so on the training set
- 30:24:39it's performing okay but then on the
- 30:24:41test set like you know it's it's not
- 30:24:44underfitting. it's performing pretty
- 30:24:45well on the training set, but then on
- 30:24:47the test set it's um performance is much
- 30:24:51worse. That's overfitting. If you're
- 30:24:54overfitting, then this is a great model
- 30:24:55to use because we can try basically by
- 30:24:58by rotating through different alphas and
- 30:25:00different L1 ratios, we can try out
- 30:25:03different strengths of penalty and
- 30:25:06different variations on lasso and ridge
- 30:25:08together. This is a really good model to
- 30:25:11to use for those overfitting cases where
- 30:25:13linear regression is doing decently
- 30:25:17um but it's overfitting.
- 30:25:20Right? So far we haven't ran that case
- 30:25:22because so far no matter what model
- 30:25:25we've used it's always underfit. So
- 30:25:28anytime we have those underfitting cases
- 30:25:31it signals that we should likely just
- 30:25:33use a more complex model and we haven't
- 30:25:36learned about those yet.
- 30:25:38um we will coming up in lesson four, but
- 30:25:42um that's for this data. That's ultim
- 30:25:45ultimately what we'd want to do is
- 30:25:46probably use a more advanced model
- 30:25:48because it's underfitting um just using
- 30:25:50a linear regression and and then using
- 30:25:52the the overfitting variations of linear
- 30:25:54regression like lasso ridge and elastic
- 30:25:56net.
- 30:26:03Okay.
- 30:26:06Any questions on this on elastic then
- 30:26:15the TV? Yeah, we Yeah, I think I have
- 30:26:17it. I can share it with you.
- 30:26:24I said that and now I can't find it. I
- 30:26:26thought I had it.
- 30:26:35I don't have it. I thought I had it, but
- 30:26:37I don't.
- 30:26:39If anyone does have that one.
- 30:26:45Yeah, I'll look one more time. I thought
- 30:26:47I had that one.
- 30:26:51Um,
- 30:26:57yeah, it's not in there. I had it. Let
- 30:26:59me see.
- 30:27:12Yeah, I don't have it either. I thought
- 30:27:13I had it in here.
- 30:27:22Yeah, I don't have that one.
- 30:27:27I don't have that one. and I'll have to
- 30:27:28find it. Uh I have this marketing data.
- 30:27:30I don't think this is the same one.
- 30:27:34I have this marketing data. I don't
- 30:27:35think that's the right one, but you can
- 30:27:36take a look at it.
- 30:27:41No, we're using So, for this example,
- 30:27:43we're using the same hitters data set
- 30:27:45that we used earlier for lasso.
- 30:27:48No, that's an earlier one.
- 30:27:54That's from the uh very beginning of the
- 30:27:58notebook. So that's the that's from this
- 30:28:01one.
- 30:28:05Oh, this Oh, this is where it is. Sorry.
- 30:28:08This is where it is. You can find it
- 30:28:09here.
- 30:28:14That's right. It was from a URL.
- 30:28:20It was used in the very beginning of the
- 30:28:21notebook.
- 30:28:23And we did we did this.
- 30:28:27Okay.
- 30:28:30That's right. That's why I didn't have
- 30:28:31it downloaded.
- 30:28:35Okay.
- 30:28:39All right. Any other questions on the
- 30:28:41elastic net before I move I'm going to
- 30:28:43move on to uh finding those a systematic
- 30:28:47way to find the best hyperparameters.
- 30:28:50Um, I'm going to show you a couple
- 30:28:51strategies to doing that. Um, so far
- 30:28:54we've just ran CV with some random
- 30:28:56choices. Um, I'm going to show you a
- 30:28:59better, more systematic approach. That's
- 30:29:00kind of the industry standard for doing
- 30:29:02tuning. Um, so I'm going to I'm going to
- 30:29:05show you that next. But any questions on
- 30:29:07the elastic net?
- 30:29:13Okay. And again like you know
- 30:29:16scikitlearn makes it really easy for you
- 30:29:17guys because it just behaves the same
- 30:29:21way as any other model. You use the
- 30:29:24object and then you do ffit and predict
- 30:29:26right? So the ffit is going to train it
- 30:29:29um and the predict is going to allow you
- 30:29:32to use that model to predict. It's it's
- 30:29:34super easy that way. Every scikitlearn
- 30:29:36model is like that dofitit and predict.
- 30:29:39So it provides a really simple way to
- 30:29:41use basically every model.
- 30:29:47Okay,
- 30:29:49let's talk about let's finish up this
- 30:29:51lesson with a couple things. Um, one of
- 30:29:54those things is going to be
- 30:29:55hyperparameter tuning. So what is this?
- 30:29:59The hyperparameter tuning is a
- 30:30:01systematic way to find the best
- 30:30:05parameters in a machine learning model.
- 30:30:08So a lot of machine learning models have
- 30:30:10what are called hyperparameters.
- 30:30:13These are not the betas that we learn
- 30:30:15during the training that's learned from
- 30:30:17the data. These are settings that we set
- 30:30:20ahead of time like the alpha. That's a
- 30:30:23perfect example like alpha L1 ratio in
- 30:30:26in the elastic net. We set those up
- 30:30:28ahead of time and depending on what we
- 30:30:30pick for those we get different
- 30:30:31performance, right? And so what we
- 30:30:34really need is a systematic way to find
- 30:30:37the best settings for those
- 30:30:40hyperparameters as we are training our
- 30:30:42models. Um the the the main like idea
- 30:30:48behind this process though is going to
- 30:30:50be to systematically try out different
- 30:30:54combinations as many as we want to try.
- 30:30:57And so we're we're basically going to
- 30:30:59have a strategy for tuning that is going
- 30:31:02to exhaust all the combinations of those
- 30:31:06hyperparameters that we want to try
- 30:31:08until we find the one that performs the
- 30:31:11best. Um and and that strategy is known
- 30:31:14as grid search. Um and essentially what
- 30:31:19it does is it sets up a grid um where
- 30:31:22which is basically like a matrix to say
- 30:31:25okay which parameters do you want to
- 30:31:27try? I want to try um alpha and I want
- 30:31:30to try L1 ratio
- 30:31:33um L1 ratio like let's say I want to try
- 30:31:37these two. So we set these up in a grid
- 30:31:39where we say, "Okay, I want to try this
- 30:31:41value. I want to try this value. I want
- 30:31:43to try this value. This one, this one,
- 30:31:44this one, and on and as many as we want
- 30:31:47to try." So we could set up set those up
- 30:31:49systematically like a linear um a
- 30:31:52linearly spaced like I want to try every
- 30:31:55alpha between between 0 and 10 spaced by
- 30:31:58one. Um whatever, you know, we can set
- 30:32:01up different ranges of those, but that's
- 30:32:03going to be in this grid. And then the
- 30:32:05L1 ratio, same thing. We can try out
- 30:32:07different values of these that we want
- 30:32:08to try. Let's say there's many of those.
- 30:32:12Um maybe every um tenth between 0 to one
- 30:32:16I want to try out. Um so you set up your
- 30:32:20parameters and you can set up as many as
- 30:32:21you want in the grid. And then
- 30:32:23essentially what you're going to do to
- 30:32:25do grid search is you're going to work
- 30:32:27your way through every combination of
- 30:32:29those. You're going to try out this
- 30:32:31combo. You're going to try out this
- 30:32:33combo. You're going to try out this
- 30:32:34combo.
- 30:32:36this combo. So the first value of alpha
- 30:32:40with every possible L1 ratio, then go to
- 30:32:43the next, try out the next value of
- 30:32:45alpha with every L1 ratio, and on and on
- 30:32:48and on. So we're going to try
- 30:32:51all combos
- 30:32:54in the grid.
- 30:32:56We're going to try all combos and we're
- 30:32:58going to find the lowest MSE
- 30:33:02combination. find lowest
- 30:33:05MSE
- 30:33:07combo.
- 30:33:08So whatever leads to the best model is
- 30:33:11going to be the um parameters that are
- 30:33:14that are deemed to be the best. And the
- 30:33:16idea is once we have found those we know
- 30:33:20that we can use we can go ahead and
- 30:33:22train a model with those best alpha and
- 30:33:24len ratio and on and on and on.
- 30:33:34Yeah, when you get an So this goes back
- 30:33:36to the error. Remember that for a
- 30:33:39regression,
- 30:33:41the error is this measurement of how far
- 30:33:44off we are, right? So if we have a bunch
- 30:33:46of points and we draw we fit a line
- 30:33:49through there, the the MSE is measuring
- 30:33:52this distance, right? So what do you
- 30:33:54think is a good distance? Like if our
- 30:33:56model is perfect,
- 30:33:59what's the best distance from our
- 30:34:02predictions to the actual points? Zero.
- 30:34:05Yes. So the lower the better. The lower
- 30:34:09the better. Um so for an R RMSSE, the
- 30:34:13lower the closer to zero the better.
- 30:34:15However, the RMSSE can be it's its units
- 30:34:19are interpreted in the units of our
- 30:34:22target.
- 30:34:23So what is deemed to be good is relative
- 30:34:26to our target. Like let's say our target
- 30:34:28is in the thousands. Like it averages in
- 30:34:31the thousands. If we produce an MSE of
- 30:34:3450 or sorry an RMSSE of 50, that's
- 30:34:39pretty good, right? Because our units
- 30:34:41are in the thousands
- 30:34:43and we're only on average we are off by
- 30:34:4750 units,
- 30:34:49right? Our distance away is about 50
- 30:34:51units. That's pretty good. So the RMSSE
- 30:34:55is relative to your target variable.
- 30:34:58Does that make sense? Yeah. It depends
- 30:35:00on the target. It depends on what you're
- 30:35:02trying to predict.
- 30:35:05So that's why we got RMSSE that were in
- 30:35:07the 300s for those hitters, but the
- 30:35:09average was the average of the target
- 30:35:11was in the 500s. So that's a really bad
- 30:35:15proportion of error relative to the
- 30:35:18average target value, right? If our
- 30:35:21RMSSE was 300,
- 30:35:24but the target was sitting in the 500s,
- 30:35:27that's just too much error. Way too much
- 30:35:30error, right? That's just too big of a
- 30:35:32value. Um, our predictions are just off
- 30:35:36way too much in terms of that distance.
- 30:35:39So, this would be this was a bad model.
- 30:35:42It was underfit.
- 30:35:44We know that from the the RMSSE. So
- 30:35:47yeah, the RMSSE closer to zero, no
- 30:35:49matter what is good,
- 30:35:52zero is being perfect. Um, but it to
- 30:35:56know what's good, you need to know what
- 30:35:58your target is on average and then think
- 30:36:01of this as kind of a ratio to that
- 30:36:03average target. I think that's the best
- 30:36:06way to think about it.
- 30:36:17Okay, so going back to this grid idea is
- 30:36:22so the grid is just basically laying out
- 30:36:25all possible parameter combinations and
- 30:36:28trying them all out by fitting and
- 30:36:30predicting until and generating an a
- 30:36:33metric like an MSE
- 30:36:36until we find the one with the lowest
- 30:36:38MSE. So find the lowest MSE combination
- 30:36:42and that will be the best
- 30:36:44that will be the best combo. And then if
- 30:36:47we once we know that best combo we can
- 30:36:49use that we can use that alpha we can
- 30:36:51use that L1 ratio and use that model
- 30:36:54going forward. We can we can use those
- 30:36:56parameters in our model. So this
- 30:36:59strategy it has a name. It's known as
- 30:37:02grid search. So it is a hyperparameter
- 30:37:05tuning process that tries out all
- 30:37:08combinations.
- 30:37:11So what's the what's the uh benefit to
- 30:37:15this is that we get to test out a lot of
- 30:37:17different combo combos of those
- 30:37:18parameters like the alpha and L1. So we
- 30:37:21can be confident what the best model is,
- 30:37:23right? So we can pick the alpha and L1
- 30:37:26perfectly because we're trying out a
- 30:37:27bunch of different combinations on the
- 30:37:29data to see which one's the best. What's
- 30:37:32the downside?
- 30:37:34It's expensive, right? It's an
- 30:37:37exhaustive search. So, if you have many
- 30:37:40different parameters and you're trying
- 30:37:43out many different combinations, it can
- 30:37:46get exponentially
- 30:37:48expensive
- 30:37:49to perform this search. Okay, so grid
- 30:37:53search is great except for the fact that
- 30:37:56it can be expensive if you have many
- 30:37:57parameters with with very wide ranges
- 30:38:00that you're searching over because that
- 30:38:02that's a lot of combinations you have to
- 30:38:03test, right? And especially if you have
- 30:38:06a lot of data, that's going to be
- 30:38:08expensive
- 30:38:10um to do.
- 30:38:13So uh we're going to practice doing grid
- 30:38:16search, but that is that's the pro and
- 30:38:17con. The pro is that we get to try out
- 30:38:19all these combinations and see which
- 30:38:20one's the best. The downside is it can
- 30:38:23be expensive to do that if you have a
- 30:38:24lot of parameters um that you want to
- 30:38:27tune for your model um and you have very
- 30:38:32uh many different choices that you're
- 30:38:34trying to evaluate for those and it just
- 30:38:36creates a really big um collection of
- 30:38:39combinations that you have to try out,
- 30:38:42right? Um that's the only downside to
- 30:38:45grid search.
- 30:38:48Now on the opposite end of the spectrum
- 30:38:50of that is a randomized search or random
- 30:38:54search and this will basically just um
- 30:38:58do a sampling of those parameters from
- 30:39:03um kind of fixed uh specified
- 30:39:05distribution. So essentially what you do
- 30:39:08is similarly you define your range. So
- 30:39:11you say I want to look at alphas um
- 30:39:14between zero or sorry between let's say
- 30:39:18yeah 0 to 10. I want to look at a bunch
- 30:39:20of different alphas um and I want to
- 30:39:22look at a bunch of different L1 ratios
- 30:39:25that are between 0 to 1
- 30:39:290 to one and um what we do is we say
- 30:39:33okay I'm going to restrict only testing
- 30:39:3720 30 40 times. I'm not going to do all
- 30:39:40possible combinations. I'm just going to
- 30:39:43randomly sample something in this range
- 30:39:46and randomly sample something in this
- 30:39:48range. And so, and I'm going to perform
- 30:39:50that experiment a fixed number of times.
- 30:39:53So, let's say I set the uh sampling
- 30:39:56where I'm only going to do um 20
- 30:40:00evaluations.
- 30:40:02And so, 20 times we're going to pick a
- 30:40:04combo randomly. So, I'm going to pick an
- 30:40:08alpha and I'm going to pick an L1 ratio.
- 30:40:14L1 ratio.
- 30:40:16And um we are we are just going to uh
- 30:40:20sample those randomly from this range.
- 30:40:24Um and we're going to use those and test
- 30:40:28those out. And then it's but otherwise
- 30:40:29it's the same as grid search. Whatever
- 30:40:31is the lowest MSE. Um, so whatever is
- 30:40:34the lowest MSE is the best.
- 30:40:37So we evaluate those, we sample, we
- 30:40:40train the model, evaluate it. Whatever
- 30:40:42is the lowest MSE
- 30:40:45is the best is the best combo. Now
- 30:40:49what's the benefit to this is it's a
- 30:40:52much more controlled experiment in the
- 30:40:56sense that we um aren't going to iterate
- 30:40:59through every possible combination in
- 30:41:00the grid. where we basically set up a
- 30:41:03fixed number of times we're going to try
- 30:41:05out stuff.
- 30:41:06The risk to doing this is that you're
- 30:41:08not
- 30:41:10you're not exploring all combinations,
- 30:41:12right? Because you're randomly sampling,
- 30:41:14you may get unlucky and you may not
- 30:41:17stumble into the best. You you can make
- 30:41:21um samples and figure out what's the
- 30:41:22best amongst your samples, but you may
- 30:41:25not be covering all the combinations.
- 30:41:27Does that make sense? The grid search is
- 30:41:29going to try every combo. The random
- 30:41:32search is going to randomly sample those
- 30:41:35combos.
- 30:41:36So, it's not going to try every single
- 30:41:38one. It's going to try a limited number,
- 30:41:40however many you set up. Now, if you set
- 30:41:43that number really, really, really high,
- 30:41:46now you're starting to approach a grid
- 30:41:47search because now you're sampling so
- 30:41:50many of those combos that you basically
- 30:41:52are trying them all at that point,
- 30:41:56right?
- 30:41:58Um
- 30:42:00so so that's the way the random search.
- 30:42:02So by the way both of these use cross
- 30:42:04validation in the sense that when you
- 30:42:07evaluate accommodation you're actually
- 30:42:09doing it with cross validation. So when
- 30:42:11you do an evaluation, you're going to do
- 30:42:14probably 10 or five folds where you
- 30:42:16split your data and then you test it on
- 30:42:19the rest of the folds and evaluate or
- 30:42:21train it on the rest of the folds,
- 30:42:22evaluate it on one of them and generate
- 30:42:24an average MSE to get your evaluation.
- 30:42:29So every evaluation is using cross
- 30:42:31validation. That's why that's and
- 30:42:34hopefully you can see why this would be
- 30:42:35so expensive for a really big grid,
- 30:42:38right? because you're trying out many
- 30:42:40different combinations
- 30:42:42and every combination is going to do a
- 30:42:44cross validation procedure. So, it's
- 30:42:47going to train 10 times and test against
- 30:42:5010 different folds and average those
- 30:42:52together. It's going to be a pretty
- 30:42:53expensive operation
- 30:42:55for a really big grid, right, of of
- 30:42:58parameters.
- 30:43:00Um, but these are the two kind of
- 30:43:01systematic approaches we have at trying
- 30:43:04out different hyperparameters. Remember
- 30:43:07those those things are called
- 30:43:08hyperparameters. These are those choices
- 30:43:11that we have before we train our model.
- 30:43:14Um those choices we have that affect the
- 30:43:17performance of the model like the
- 30:43:18alphas, the L1 ratios, those kind of
- 30:43:20things. Um we have control over what
- 30:43:23they're going to be. This is a
- 30:43:24systematic approach to find out what the
- 30:43:26best
- 30:43:29uh value of those parameters is going to
- 30:43:31be on our data,
- 30:43:34right?
- 30:43:37Okay. So, before we practice this, we're
- 30:43:40going to practice a grid search first.
- 30:43:42Um,
- 30:43:44any questions?
- 30:43:54Uh, I don't know if it has a built-in
- 30:43:56That's a good question. by time limit. I
- 30:43:57don't know if it has a built-in way of
- 30:43:59doing it, but you could certainly set up
- 30:44:00like a a a loop um to like to wrap
- 30:44:05around, you know what I mean? Like you
- 30:44:07could set up a loop where you check the
- 30:44:08time. If it's if if the time elapsed as
- 30:44:11you're doing the search, if the time
- 30:44:12elapsed is greater than the the time
- 30:44:15limit, then you can kind of break early.
- 30:44:18Um so it's not hard to implement that,
- 30:44:20but I don't know if it has that built
- 30:44:22in. I don't think it does
- 30:44:25because I don't think it really cares
- 30:44:26how long every evaluation takes. It's
- 30:44:29just going to exhaust all those
- 30:44:30especially in a grid search.
- 30:44:33But um yeah, I there's probably a way to
- 30:44:36manually kind of set up a time time
- 30:44:38loop.
- 30:44:44So hyperparameters are um settings that
- 30:44:49we have on the model itself and a really
- 30:44:52good example of this is like the alpha
- 30:44:53and L1 ratio in the in the elastic net.
- 30:44:56So they're not things that we um learn
- 30:45:00from the data directly like the betas in
- 30:45:02the model like those get trained
- 30:45:05directly by doing the um lease squares
- 30:45:08process right um by doing that gradient
- 30:45:11descent and all that optimization.
- 30:45:14Um so these are not learned from that.
- 30:45:16They're actually set ahead of time. And
- 30:45:19so what we're saying is
- 30:45:21the best way to understand the effects
- 30:45:23of those is to try out different
- 30:45:25combinations of those until we land on
- 30:45:27the best one. Right? So hyperparameters
- 30:45:30are those options we have in the model
- 30:45:32like the alpha like the alpha and l1
- 30:45:35ratio in the uh elastic net. Many models
- 30:45:40have hyperparameters. Um we're actually
- 30:45:42going to see that in in future models
- 30:45:44that we study. they have options that
- 30:45:46you can set that affect their
- 30:45:48performance.
- 30:45:49And so this this is just a strategy to
- 30:45:52evaluate those different options to see
- 30:45:53which one's the best.
- 30:46:05Yeah. So again, hyperparameters, those
- 30:46:08are settings on the model itself um that
- 30:46:12affect the performance of it.
- 30:46:16And basically we have the two two
- 30:46:18strategies here. We can set up an
- 30:46:20exhaustive grid and search through all
- 30:46:22of those until we find the lowest MSE uh
- 30:46:24option or we can randomly sample
- 30:46:28potential options, try them out and see
- 30:46:30which one's the lowest as well. And do
- 30:46:32that a fixed number of times. Um
- 30:46:37sort of like a fixed number of trials
- 30:46:38almost. um which has a risk of not
- 30:46:42trying out every option but but
- 30:46:44hopefully you try out enough that you've
- 30:46:46explored the space a bit and you get
- 30:46:49some quality choices there but no
- 30:46:52guarantees right no guarantees you try
- 30:46:54everything which is what a grid search
- 30:46:55will do it will try everything
- 30:47:01okay now luckily per usual scikitlearn
- 30:47:06has something to manage this process for
- 30:47:08us in terms of grid search. Um so in
- 30:47:13that way we will not need to manage this
- 30:47:16process ourselves. We can just rely on
- 30:47:18scikitlearn. And so if you're doing
- 30:47:20hyperparameter tuning um this is going
- 30:47:23to come from the model selection module
- 30:47:25inside of sklearn. So we're going to
- 30:47:28import from from skarn the model
- 30:47:30selection module. We have our grid
- 30:47:32search cross validation.
- 30:47:35Okay. That's what the CV stands for.
- 30:47:37grid search cross validation. So, this
- 30:47:40is going to do that grid search
- 30:47:41strategy. Um, we're going to set it up
- 30:47:43with our dictionary essentially of
- 30:47:46choices. So, we're going to say, hey,
- 30:47:48here's the alphas I want to try. Here's
- 30:47:49the L1 ratios I want to try. Um, and
- 30:47:52here's my other settings like uh how
- 30:47:55many folds I want to use, what my random
- 30:47:57state is for the shuffling. So, we'll
- 30:47:59set all that up. Um,
- 30:48:02and then we'll just run the grid search.
- 30:48:04And then what should come out of that is
- 30:48:06the best options for our parameters from
- 30:48:09the grid. And then we can use those
- 30:48:11going forward in the we can build a
- 30:48:13model with those best options, right? So
- 30:48:16we're really doing some evaluation here
- 30:48:19of what is going to be those best
- 30:48:20alphas, those best1 ratios on our data
- 30:48:23set, right? And the only way to really
- 30:48:26know that is to evaluate them because
- 30:48:28they're not things that are learned
- 30:48:30during the training. Hopefully that
- 30:48:32makes sense, right? They're not things
- 30:48:33that we learn directly from training.
- 30:48:36They're things that we have to set and
- 30:48:38then kind of evaluate and see how they
- 30:48:39affect things.
- 30:48:43Okay. So, we have grid search CV. That's
- 30:48:45going to be our primary um tool to do
- 30:48:49the evaluations of the different
- 30:48:50hyperparameter options.
- 30:48:53Grid search CV. Um we're going to set up
- 30:48:56our cross validation uh object here. Now
- 30:48:59I want you to pay attention to this is
- 30:49:01that um it's a slightly different
- 30:49:04version than the kfold we had earlier.
- 30:49:06So we've used kfold before with a
- 30:49:08certain number of folds. This would be
- 30:49:0910 folds and we can set a random state
- 30:49:12for the shuffling um that happens in the
- 30:49:14folds.
- 30:49:16But this is actually a slight different
- 30:49:17variation on it where it is a repeated
- 30:49:19kfold where we do three repeated trials.
- 30:49:23Now why would we do that? It's to be
- 30:49:26extra extra extra careful with the
- 30:49:29shuffling.
- 30:49:31So this what this means is we do three
- 30:49:33different shuffles. So we do k-fold, we
- 30:49:36actually repeat it three times with
- 30:49:38three different shufflings. That's all
- 30:49:39that means. So the repeated k-fold is
- 30:49:43actually a bit beyond the just basic
- 30:49:46kfold. What basic kfold will do will
- 30:49:49we'll shuffle and then do our splits
- 30:49:52into 10 splits and then train on nine of
- 30:49:54those. test on the other one and rotate
- 30:49:57through all the splits.
- 30:49:59We're actually going to do that process
- 30:50:02three different times with three
- 30:50:04different shuffles. So this and we're
- 30:50:06going to average 30 results instead of
- 30:50:09just 10. So repeated kfold is just going
- 30:50:13ab above and beyond to do extra to
- 30:50:16repeat the kfold three different times.
- 30:50:18In this case only three. We could do
- 30:50:19more.
- 30:50:21But um now is that necessary to do? You
- 30:50:24could argue not necessarily. Um but it
- 30:50:27just provides extra robustness
- 30:50:30uh beyond just our single shuffle and
- 30:50:33then split and then rotation of those
- 30:50:36folds. Right? We're doing it actually
- 30:50:37three different shuffles. Um so we're
- 30:50:40repeating our kfold three times uh for
- 30:50:43every now is the thing is we're doing
- 30:50:45that for every evaluation. So it is
- 30:50:48going to be more expensive than just a
- 30:50:49basic kfold.
- 30:51:06So we have three different K-fold trials
- 30:51:08that we're doing essentially.
- 30:51:12Okay, hopefully that makes sense. This
- 30:51:13is the repeated K-fold. We haven't
- 30:51:15really seen that before. We've only
- 30:51:16worked with the Kfold, which would get
- 30:51:19rid of this repeats option and only have
- 30:51:21uh 10 splits in a random state for the
- 30:51:24for the single shuffle that we do. So we
- 30:51:26can um recreate that same shuffle every
- 30:51:29time. Um but now we're actually going to
- 30:51:32do three random shuffles, uh three
- 30:51:34different trials. So one shuffle creates
- 30:51:37the and then create the 10 splits,
- 30:51:38evaluate, then go back and do another
- 30:51:40shuffle, another new 10 splits. So, one
- 30:51:44thing that should be um clear is that we
- 30:51:48get different splits every time because
- 30:51:50we're going to shuffle once, right?
- 30:51:52We're going to shuffle once and generate
- 30:51:54our splits
- 30:51:56and then we're going to shuffle again,
- 30:51:58generate these splits which are going to
- 30:52:00be different and then shuffle one more
- 30:52:01time for for three different times,
- 30:52:04right? And then get get these splits and
- 30:52:06then we're going to get 10 metrics here,
- 30:52:0810 metrics here, 10 metrics here, and
- 30:52:11then average all of those together.
- 30:52:14So, it's a bit more just going up extra
- 30:52:17above and beyond for a kfold. Okay.
- 30:52:22All right. So, here comes the fun of
- 30:52:24when you do grid search. Now, the grid
- 30:52:28is actually just a dictionary. It's a
- 30:52:30Python dictionary where you declare what
- 30:52:34your parameters are going to be inside
- 30:52:36the dictionary and you set up a range of
- 30:52:39values that you're go or a list. It can
- 30:52:42be a list. It can be a range
- 30:52:44but some declaration of what you are
- 30:52:47going to test and evaluate inside of
- 30:52:50your grid search. So the grid is
- 30:52:52initialized as an empty dictionary.
- 30:52:55And then what we do is we say okay in my
- 30:52:58grid I want to check different alphas.
- 30:53:00So we're going to add a collection of
- 30:53:02alphas in here that we're going to test.
- 30:53:06So let me make a comment there. We add a
- 30:53:10add a range of alphas to test. And this
- 30:53:16range is a this is just like the Python
- 30:53:20range. Um
- 30:53:23this is just like a Python range um uh
- 30:53:26operator here where this is going to be
- 30:53:29uh every so it's going to be um every
- 30:53:34uh value between
- 30:53:37zero and one um uh steps with a step
- 30:53:43size
- 30:53:45of 0.1. So it's going to try a bunch of
- 30:53:49different alphas um between uh zero and
- 30:53:530.1
- 30:53:54sorry 0 and one stepping by 0.1. So it's
- 30:53:57going to try zero.1
- 30:53:592.3 point 4.5 6 right all the way up to
- 30:54:02one.
- 30:54:04So that's what this will do and it's a
- 30:54:06numpy range. So it's just all those
- 30:54:08decimals between 0 to one.
- 30:54:11It you can use either that's valid.
- 30:54:14Yeah, you can do you can do that to
- 30:54:16create a dictionary or you can use the
- 30:54:18keyword um dict. You can use either one.
- 30:54:21Either one works.
- 30:54:25Whatever whatever you want to use.
- 30:54:26They're the same.
- 30:54:29Yeah. The the reason people prefer
- 30:54:33dictionary is because um sets are
- 30:54:37created with the same braces.
- 30:54:39So it it makes it clear what you're
- 30:54:41creating as a dictionary. If you use if
- 30:54:43you use this, that's the only advantage
- 30:54:46is it's just plainly obvious what you're
- 30:54:48making. Uh because technically you can
- 30:54:50make a set with the curly braces as
- 30:54:54well.
- 30:54:59Yeah,
- 30:55:02no worries. Um okay, so we have our
- 30:55:05alphas here. So what I want you to
- 30:55:07notice is that we are going to try out
- 30:55:09different alphas and we are that's the
- 30:55:12only parameter we are going to try in
- 30:55:13our ridge regression. So we're going to
- 30:55:16we're going to try ridge but just try
- 30:55:18different alphas in the in this range um
- 30:55:21in our grid search. So the grid search
- 30:55:24CV takes in a model. It takes in our
- 30:55:27grid dictionary which is really
- 30:55:29critical. We need that dictionary to
- 30:55:30declare what we're going to try.
- 30:55:34um we need a scoring to say to find the
- 30:55:36best. Now remember it uses the negative
- 30:55:39to find the lowest which is going to be
- 30:55:42the the least negative option.
- 30:55:46Um otherwise it wouldn't um just based
- 30:55:49on the optimization it would look for
- 30:55:51the highest value. Um so the highest
- 30:55:54would be closest to zero in this
- 30:55:56situation. Um so we use negative and
- 30:55:59again we could use squared error. It's
- 30:56:02using absolute. We could use um squared
- 30:56:05uh either either one works.
- 30:56:09Um more typical would probably be
- 30:56:11squared error, but um absolute is fine.
- 30:56:15Here's where we have our repeated kfold.
- 30:56:17So we pass in our um how we're doing CV.
- 30:56:20That can be it can be a kfold object. It
- 30:56:22can actually just be an integer, which
- 30:56:24is say I just want to do 10 splits or
- 30:56:26five splits um to to do every
- 30:56:29evaluation. But these are the bare
- 30:56:32minimum that you need just really the
- 30:56:33model and the grid and your CV. Um what
- 30:56:37metric you're using to evaluate what's
- 30:56:39going to be the best. And then this end
- 30:56:41jobs is to parallelize. If you have it
- 30:56:43set to minus one, it's going to it's
- 30:56:44going to try out all the grid options in
- 30:56:46parallel. Um which is nice. It's going
- 30:56:49to help speed up the overall search.
- 30:56:52Okay. So let me mark that down as n
- 30:56:57jobs equals minus one.
- 30:57:01tries out the combos in parallel.
- 30:57:06So in this situation, we actually don't
- 30:57:08have more than one parameter. We only
- 30:57:10have the alpha. So we're really just
- 30:57:12going to be systematically working our
- 30:57:14way through every alpha and evaluating
- 30:57:16which one's the best right with this.
- 30:57:19And notice that in order to use this
- 30:57:22grid search, all we have to do is call
- 30:57:23search.fit. So it works kind of like
- 30:57:27every other model does, right? It's the
- 30:57:29grid search.fit.
- 30:57:32And we pass in our data.
- 30:57:34And we um once we're once this prints
- 30:57:38out the results, you get a results
- 30:57:40object
- 30:57:42um which has a best score and then a
- 30:57:45dictionary with your best parameters.
- 30:57:47So, whatever your best grid member was
- 30:57:50or grid members, um it prints that out
- 30:57:52and you can So, for from that, we can um
- 30:57:56grab our best alpha, which it which
- 30:57:58let's confirm what that ends up being.
- 30:58:06Oops. We need to import repeated kfold.
- 30:58:17So we'll import that.
- 30:58:24Oh, I didn't. Let's do from
- 30:58:27sklearn.linear
- 30:58:32model import ridge.
- 30:58:36Okay.
- 30:58:44Okay. So, it completed the search and
- 30:58:46what we found is this is the best score
- 30:58:48is 238 for the mean absolute error and
- 30:58:52the best alpha that we got was 0.9. So,
- 30:58:56the best alpha that worked here, the one
- 30:58:59that gave us the best score was actually
- 30:59:010.9 as the alpha. So what it did is it
- 30:59:04tried out everything between this range
- 30:59:07and 0.9 was the best. So it did cross
- 30:59:10validation tried out every single combo
- 30:59:13in our grid.
- 30:59:15So if we want we could actually print
- 30:59:17out
- 30:59:20print our grid so we can see
- 30:59:23what our combinations were.
- 30:59:29So, it tried out all of these guys and
- 30:59:32the best one that we had was 0.9.
- 30:59:40Okay, so pretty cool how that works. And
- 30:59:44if we had other parameters, like if we
- 30:59:46were doing a elastic net, we could add
- 30:59:48those into our dictionary and it would
- 30:59:50do all combinations of those. So if we
- 30:59:53did um so for instance to add to our
- 30:59:56grid we could do grid
- 30:59:58um L1 ratio
- 31:00:01this would be for like an elastic net
- 31:00:02right now the ridge regression by itself
- 31:00:04doesn't have an L1 ratio parameter but
- 31:00:06just as an example um we could try out
- 31:00:09different ranges
- 31:00:11um similar range different one um maybe
- 31:00:14an exact list whatever we want to do. So
- 31:00:17this is going to try out different ones
- 31:00:18between 0 to one as well. And so it's
- 31:00:22going to try out every combination of
- 31:00:24these from this grid.
- 31:00:27Okay, if we did that. But again, this
- 31:00:29the ridge regression doesn't have an L1
- 31:00:32ratio. The elastic net does. So that the
- 31:00:35ridge regression only has an alpha to to
- 31:00:38as a hyperparameter. So we're only
- 31:00:40testing out that one.
- 31:00:45Okay, so that's grid search CV.
- 31:00:49Pretty useful. This is pretty useful in
- 31:00:51doing parameter tuning again when you
- 31:00:53want to try out ranges of different
- 31:00:55values and you can evaluate those to see
- 31:00:58which one is your best and then we can
- 31:01:01use that best going forward. So we can
- 31:01:04for instance this is what this code does
- 31:01:06below it is it fetches the best. Um you
- 31:01:09can do it this way or you can do it um
- 31:01:12the alternative is to do results.b best
- 31:01:15params
- 31:01:18and then you can just grab it like this
- 31:01:21alpha.
- 31:01:23Either way you can do get or like this
- 31:01:26um and it this is just a dictionary,
- 31:01:29right? And you can grab your alpha. So
- 31:01:31that's the 0.9 um and we can pass that
- 31:01:34alpha into the ridge regression and go
- 31:01:36back and refit it to our data um and
- 31:01:39then use that model going forward. So
- 31:01:41the grid search really just evaluates
- 31:01:44those different options, tells you
- 31:01:46what's the best according to this score,
- 31:01:50right?
- 31:01:52And you should, by the way, you should
- 31:01:54interpret this score in the positive
- 31:01:56sense. It's only negative because we're
- 31:01:59purposely making it negative to find out
- 31:02:02what the lowest option is, right?
- 31:02:04Because the lower is the better. So we
- 31:02:06we purposely make it negative to make it
- 31:02:08whatever is the least negative is the
- 31:02:10winner. Um more negative is worse.
- 31:02:15So it's the really positive version of
- 31:02:18it is the is the true result for the
- 31:02:21error. Um and they are a tool from
- 31:02:24scikitlearn to put together your model
- 31:02:26with your pre-processing steps. So they
- 31:02:29kind of get automated together. Um and
- 31:02:32they combine everything into kind of a
- 31:02:34streamline process. You're going to see
- 31:02:36what that looks like, but it's a really
- 31:02:38nice um feature of scikitlearn. Um why
- 31:02:42would we care about pipelines? They help
- 31:02:45organize our code um so that we ensure
- 31:02:48that we basically always run the
- 31:02:50pre-processing steps before we train and
- 31:02:52use a model to with the predictions. Um,
- 31:02:56so it bundles those steps together,
- 31:02:58minimizes the risk of forgetting a step
- 31:03:00because one of the things that can
- 31:03:02happen is when you do pre-processing, if
- 31:03:04you're doing it on the training set, you
- 31:03:05have to do it on new test data as well
- 31:03:07when you put it through your model
- 31:03:09because your model is training against
- 31:03:11that pre-processed data.
- 31:03:13So in order to make sure you never
- 31:03:15forget that, you can bundle it all
- 31:03:17together in a pipeline which is going to
- 31:03:19make things really really easy to use
- 31:03:22and and make sure that those steps
- 31:03:25happen in a repeatable way. Um and it
- 31:03:29makes things easier to uh deploy that
- 31:03:32model as well because everything is
- 31:03:34together in one pipeline. So in the in
- 31:03:37the industry, I've seen this a lot. Um
- 31:03:40you know, people will do their initial
- 31:03:43exploration steps and initial model
- 31:03:45building. They may not use pipelines
- 31:03:47right away, but as they found their
- 31:03:50model, um they'll generally move it into
- 31:03:53a pipeline and all their steps into a
- 31:03:54pipeline so that it's uh easier to work
- 31:03:57with um when you're when you're
- 31:03:58deploying it and actually using it uh in
- 31:04:02in the real world. Um, so here's what a
- 31:04:06pipeline generally looks like. It's from
- 31:04:07scikitlearn. It's this pipeline object.
- 31:04:10Um, and the pipeline is made up of steps
- 31:04:13that we're going to see that that are
- 31:04:15various um uh basically um kinds of
- 31:04:20pre-processing we've seen before like a
- 31:04:22scaler or um filling in missing values.
- 31:04:26Those kind of things we can put here in
- 31:04:28the steps which is basically a list. um
- 31:04:31steps is just going to be a list of
- 31:04:33scikitlearn functions that we can apply
- 31:04:34to data. One of those being a model. Um
- 31:04:38and then whenever we use the pipeline,
- 31:04:40it's basically um you know it's going to
- 31:04:43be something like pipeline.fit
- 31:04:45or pipeline.predict.
- 31:04:47So the pipeline kind of behaves like a
- 31:04:50model. It's just going to contain many
- 31:04:53more steps than that like the
- 31:04:54pre-processing steps we've worked with
- 31:04:56before. Um, and it also has some
- 31:04:59capabilities for caching. So you can
- 31:05:02like uh cache some of the data in
- 31:05:04memory. Um, so that if you're reusing
- 31:05:06the predictions, it kind of goes faster.
- 31:05:09Um, so there's some options for that
- 31:05:11too. I'm not too concerned about that at
- 31:05:13this stage, but the main thing is going
- 31:05:15to be filling out our steps and then
- 31:05:17using the pipeline.
- 31:05:20Okay.
- 31:05:22Um, so some important bits of
- 31:05:25information about the pipeline is that
- 31:05:27it is going to be a sequence of data
- 31:05:28transformations that will have at the
- 31:05:31very end of the pipeline the model
- 31:05:34because of course we're going to do
- 31:05:35transformations and then train a model
- 31:05:38or predict with a model. So every
- 31:05:43Oh, can you guys hear me? Okay,
- 31:05:46not able to hear me. Thanks for letting
- 31:05:48me know. Can you guys were you able to
- 31:05:49hear me so far?
- 31:05:55Okay. Make sure. Yeah, it might be on
- 31:05:56your internet or your your uh Yeah, it
- 31:06:00seems like seems like it's good. So,
- 31:06:04no, you can't hear me. Check your
- 31:06:06volume. Check your headphones if you're
- 31:06:08wearing headphones. Oh, no issues. Okay,
- 31:06:11perfect.
- 31:06:13Okay. Yeah, local internet issue. Yeah.
- 31:06:17Okay.
- 31:06:19Always let me know. always let me know
- 31:06:20because it could be the case that it is
- 31:06:22me. So, um always always make sure to
- 31:06:26let me know. Um but sounds like yeah,
- 31:06:29you may want to check on that. Um
- 31:06:32so, okay, what I was saying is every
- 31:06:35pipeline is going to have a uh a
- 31:06:37sequence of steps that go first and then
- 31:06:39the model at the end. Um so, the order
- 31:06:42really matters. Um
- 31:06:45uh so the order matters in the sense
- 31:06:48that we want our transformations to go
- 31:06:49first. Things like scaling, things like
- 31:06:52filling in missing values, we want those
- 31:06:53to be first and then we want our uh
- 31:06:57model to be last because we want those
- 31:06:59transformations to happen prior to
- 31:07:01training or prior to prediction. So
- 31:07:04usually what you'll see in these
- 31:07:06pipelines is a model at the end, right?
- 31:07:09a model that's going to be at the end of
- 31:07:11the pipeline because we want basically
- 31:07:14our processing steps then our training
- 31:07:16or processing steps then our
- 31:07:18predictions. Um so everything in the
- 31:07:22pipeline though is going to be from
- 31:07:23scikitlearn. Uh that's how it gets
- 31:07:26automated in the sense that all of those
- 31:07:28things are going to have fit and
- 31:07:29transform functions built into them so
- 31:07:31the pipeline can use them. Uh, and then
- 31:07:34the last step is going to be a model
- 31:07:36that has a fit and a predict. So it's
- 31:07:39pretty standard that the last part of
- 31:07:40the pipeline is just going to be a
- 31:07:41model. Um,
- 31:07:45uh, so we can um, as we do more
- 31:07:50modeling, we're going to play around
- 31:07:52with the pipelines quite a bit and see
- 31:07:53how we can change up some of the
- 31:07:55parameters. like if we want to change a
- 31:07:57model's parameter um we can actually
- 31:07:59adjust it to do things like uh grid
- 31:08:02search or cross validation. So um we're
- 31:08:05going to see some examples of some
- 31:08:07pipelines but for right now mostly what
- 31:08:09we're going to see is how to build one
- 31:08:11and then how to use one. And then as we
- 31:08:14get into lesson four, we'll get some
- 31:08:16more practice with pipelines because
- 31:08:18we're going to start using them quite a
- 31:08:19bit uh to build our models rather than
- 31:08:22do manual steps uh all the manual
- 31:08:25pre-processing
- 31:08:27um and then kind of building a model
- 31:08:29from there. We'll just include all of it
- 31:08:31together in a pipeline.
- 31:08:35Okay, so the example we're going to do
- 31:08:36is with this housing with ocean
- 31:08:39proximity. So we've actually looked at
- 31:08:40this data set before. Um so we have uh
- 31:08:45this ocean proximity data set that has
- 31:08:47the feature of like how close it is to
- 31:08:49the ocean like the bay or the less than
- 31:08:521 hour. Remember we had that and it had
- 31:08:54the median house value for different
- 31:08:56neighborhoods. Um so we're going to work
- 31:08:58with that one again. Let me make sure I
- 31:09:00have that one uploaded.
- 31:09:04You guys should have this one. It should
- 31:09:05be in your uh data sets.
- 31:09:10Um, I'll I can upload it here in case
- 31:09:11you don't have it though.
- 31:09:22Does this use multi-threading? I think
- 31:09:24it does. Yeah, I think in order to do it
- 31:09:26can do uh um I think it can do
- 31:09:29processing in parallel for some of the
- 31:09:31pipeline steps. Um, now does it use that
- 31:09:35all the time? Not necessarily because
- 31:09:37some of it is sequential in nature where
- 31:09:40you have to do one step and then you do
- 31:09:41the next step and then you do the next
- 31:09:43step. So it's not like you can do them
- 31:09:44in parallel.
- 31:09:46Um in terms of the like you need to know
- 31:09:48the output of one step to compute the
- 31:09:50the output of the next step. Um so it
- 31:09:55can but it it doesn't always lend itself
- 31:09:58well. The thing that will use
- 31:10:00multi-threading is is like the training
- 31:10:03process can be parallelized
- 31:10:05like the fit um can be and for some
- 31:10:08models it can be parallelized not every
- 31:10:11model
- 31:10:15it. So long answer is or the short
- 31:10:18answer is that it depends
- 31:10:20depends on what kind of transforms
- 31:10:21you're doing and what kind of model
- 31:10:22you're using. if you can really take
- 31:10:24advantage of that.
- 31:10:31Okay, so we load our data here and take
- 31:10:33a look at that. Um, do you guys have
- 31:10:36this data set? Are you able to load it
- 31:10:38in? If you're following along, are you
- 31:10:40able to load it?
- 31:10:48Okay.
- 31:10:50And and again, we've worked with this
- 31:10:51data before, so hopefully it's somewhat
- 31:10:53familiar. Remember, every row represents
- 31:10:56a neighborhood, and it has a we're going
- 31:10:58to end up trying to predict this median
- 31:11:00house value as our target um variable,
- 31:11:04our dependent variable. Um and we're
- 31:11:06going to use the rest of these features.
- 31:11:08Remember that um this feature is in
- 31:11:11particular going to need to be one hot
- 31:11:13encoded,
- 31:11:15right? It's going to be one hot encoded
- 31:11:17because it is currently a string. and we
- 31:11:19need to turn that into a numerical
- 31:11:22feature which is the one hot encoded
- 31:11:24feature. So we're gonna have to do that
- 31:11:26but we're going to do that as part of
- 31:11:28our pipeline.
- 31:11:30Okay. So we'll be able to include that
- 31:11:32in our pipeline steps uh to to do one
- 31:11:35hot encoding which is nice.
- 31:11:39All right. So we're going to split apart
- 31:11:41our data um as we normally do. So we're
- 31:11:44going to uh create our feature uh data
- 31:11:48frame which is everything but this
- 31:11:50median house value. So we go ahead and
- 31:11:52drop that column and then our target is
- 31:11:54the median house value. So it is just
- 31:11:56that column here. Pretty standard. Um
- 31:12:00and then we're going to train test split
- 31:12:03and um split it into 30%
- 31:12:07uh test data. And again random state you
- 31:12:10can choose whatever you want to be. that
- 31:12:11just affects the shuffling. Um, so
- 31:12:14whatever doesn't really matter what it
- 31:12:16is. It's just so that when you rerun
- 31:12:17this, you get the same result in the in
- 31:12:19the shuffle.
- 31:12:22Okay, so we have our train and our test.
- 31:12:27So you want to make sure you run those.
- 31:12:30All right, so what we're going to do is
- 31:12:32take a look at our data
- 31:12:35and see if we have any null values. Um,
- 31:12:38if you guys remember this data actually
- 31:12:41did have null values. You can see it
- 31:12:42here in this this guy and exactly how
- 31:12:46many there are is from this the sum. So
- 31:12:48we have um 162 nles in in this data. Uh
- 31:12:53and this is just a training data. So of
- 31:12:55course you know the test data could have
- 31:12:57that in there as well. Um so that's
- 31:13:00something we're going to want to make
- 31:13:01sure we fill in the blanks on any data
- 31:13:03set we use whether we're using the
- 31:13:05training or test set. Um, like if we're
- 31:13:08doing training, we want to make sure
- 31:13:09that gets filled in. If we're doing
- 31:13:10predictions with the test set, want to
- 31:13:12make sure that gets filled in. Um, so we
- 31:13:16we should be doing that. Um, now
- 31:13:21what we're going to do is use this data
- 31:13:25to help uh train our pipeline or or use
- 31:13:29with our pipeline. We need to construct
- 31:13:31our pipeline. So far, we've just split
- 31:13:33apart our data. We haven't done anything
- 31:13:36with our processing steps in our model
- 31:13:38yet. Um so roughly
- 31:13:42it this should be the flow of our
- 31:13:43pipeline. What should happen is we
- 31:13:45should be doing some type of feature
- 31:13:47scaling
- 31:13:49um some type of uh feature um
- 31:13:52manipulation. So that could be
- 31:13:53engineering, that could be um that could
- 31:13:57be uh doing the one hot encoding. Um so
- 31:14:01extracting new features like one hot
- 31:14:03encoding,
- 31:14:06one hot encoding. Um we are going to be
- 31:14:09doing that and and by the way this is
- 31:14:11split up into this is when we use our
- 31:14:13pipeline for training.
- 31:14:16Um it's going to look like this where we
- 31:14:17do our scaling, we do one hot encoding.
- 31:14:20Um we have our model here. Um so that
- 31:14:23could be a linear regression, that could
- 31:14:24be a lasso, that could be a ridge, it
- 31:14:26could be elastic net. Whatever model we
- 31:14:28end up using is going to be last in the
- 31:14:30pipeline. And we're going to run this
- 31:14:33pipeline. Ultimately, we're going to run
- 31:14:36pipeline.fit,
- 31:14:40right? We're going to run a fit
- 31:14:41function. and we get a fitted model as
- 31:14:44the result of this pipeline.
- 31:14:47Then when we use it when we use our
- 31:14:50model for prediction,
- 31:14:54we use our model for prediction in this
- 31:14:56lower part, it's the same pipeline, same
- 31:14:59exact pipeline, but it's this model has
- 31:15:02now been trained.
- 31:15:04So we now have a trained model here. So
- 31:15:07the great thing about the pipeline is
- 31:15:08it's the same this is the same pipeline
- 31:15:11that we're using here. So it's just
- 31:15:13going to it's going to repeat those same
- 31:15:15transformations. It's going to do our
- 31:15:17scaling. It's going to do our one hot
- 31:15:19encoding. It's going to use our model
- 31:15:21and it's going to generate predictions
- 31:15:23and generate uh we can we can do
- 31:15:26predictions. We can do evaluation like
- 31:15:28an across validation. Um we can use it
- 31:15:30however we want to use it. Uh but notice
- 31:15:34that the pipeline makes it consistent
- 31:15:37between training and test. We're using
- 31:15:39the exact same transformations
- 31:15:41and the model is last. It's it's either
- 31:15:44being trained or it's being used for
- 31:15:46prediction but it's last. Our
- 31:15:47transformations are up front which are
- 31:15:50things like our scaling, things like our
- 31:15:52one hot encoding, right? Those happen
- 31:15:54first. No matter what data we put
- 31:15:57through there, we put our training data
- 31:15:59through there, we put our test data
- 31:16:00through there, they're going to go
- 31:16:02through the same steps,
- 31:16:05right?
- 31:16:08So that's that's the design of the
- 31:16:09pipeline. That's what it's supposed to
- 31:16:11do. So our job is to create those steps.
- 31:16:16So we need to create those relevant
- 31:16:18steps and then put them together into
- 31:16:20this pipeline. Okay, so that's going to
- 31:16:23be the code we're going to see coming up
- 31:16:24is we're going to build out these steps
- 31:16:27and then put them together into the
- 31:16:29pipeline.
- 31:16:38Um, any questions on this diagram? Does
- 31:16:41it make sense what we're trying to do
- 31:16:42with this pipeline? We want to have
- 31:16:44repeatable steps during the training,
- 31:16:46during a prediction process.
- 31:16:49Okay.
- 31:16:55All right.
- 31:17:01All right. So, um, a couple of things
- 31:17:04we're going to need is, uh, to first of
- 31:17:07all, let's jot down what steps we're
- 31:17:09actually going to do. We're going to
- 31:17:11need to deal with missing values. So,
- 31:17:12we're going to fill in we're going to
- 31:17:13need a pre-process pre-processing step
- 31:17:16that fills in any nles. We always need
- 31:17:19that, right? So, if there's nles, we're
- 31:17:22going to fill them in somehow.
- 31:17:24We're going to define how we do that in
- 31:17:26our in our step. Um, and we also need to
- 31:17:30one hot encode. And we need to scale,
- 31:17:33right? Those are pretty standard steps
- 31:17:36that we've dealt with whenever we're
- 31:17:37building these models, right? So, pretty
- 31:17:40standard things. fill in nles one hot
- 31:17:42encode any categorical data whatever
- 31:17:44however much we have and then go ahead
- 31:17:47and um standardize which is the scaling.
- 31:17:51So this this just is the same word for
- 31:17:53scaling our numeric features. So we're
- 31:17:56going to we're going to define those. Um
- 31:17:59so that's why we're going to go ahead
- 31:18:01and import from pre-processing. We're
- 31:18:03going to import our scaler. Um again we
- 31:18:06could use minmax scaler here. We're
- 31:18:07going to use standard scaler. Um but we
- 31:18:11could use minmax. Um we have our oneh
- 31:18:13hot encoder here. Now usually when we do
- 31:18:17oneh hot encoding we use pd.get dummies.
- 31:18:22This does the same thing as that. But
- 31:18:26because we're going to be building a
- 31:18:27pipeline we actually want the
- 31:18:29scikitlearn version of git dummies. So
- 31:18:32this is the scikitlearn version of git
- 31:18:34dummies here. And it and we have to use
- 31:18:37that version in the pipeline because
- 31:18:39everything in the pipeline needs to be
- 31:18:41an sklearn object. It needs to be an
- 31:18:43sklearn tool or object.
- 31:18:46So um instead of using pandas get
- 31:18:49dummies, we're using one hot encoder
- 31:18:51which is does the same thing. Okay. In
- 31:18:55fact, it just this basically just uses
- 31:18:58pd.get dummies um under the hood.
- 31:19:04Okay. So, it just uses that. Uh,
- 31:19:06anyways, it's just code that builds on
- 31:19:08builds on that.
- 31:19:10Now, what's really nice here is we're
- 31:19:12also going to use from sklearn.impute,
- 31:19:15we're going to use a simple imper. Now,
- 31:19:17what this is is an automated way to fill
- 31:19:20in missing values. So, this is a fancy
- 31:19:22way of basically doing the the fill na
- 31:19:27on a data frame. So, simple imputer um
- 31:19:30we are going to basically fill in the
- 31:19:33blanks. What we're going to do when we
- 31:19:35create this object is give it a strategy
- 31:19:37of how to fill in blanks. Should you use
- 31:19:39the average? Should you use the median?
- 31:19:41Should you use the max? Should use the
- 31:19:43men? Should you use a default value?
- 31:19:45We're going to tell it what to do in
- 31:19:47this object.
- 31:19:50Okay. So, we're going to we're so we're
- 31:19:52going to use this as our automated tool
- 31:19:54for filling in missing values. So,
- 31:19:56that's really nice. it has. So this is
- 31:19:58going to be a critical part of our
- 31:20:00pipeline an imputer that's going to fill
- 31:20:03in missing values.
- 31:20:06So we have that
- 31:20:11yeah coding to reduce coding. Exactly.
- 31:20:14Uh we have our pipeline now. So we have
- 31:20:17our pipeline. So our pipeline is going
- 31:20:18to hold everything. So we need the
- 31:20:20pipeline object um to hold everything
- 31:20:23and that comes from sklearn.pipeline.
- 31:20:25Um, so everything's going to actually go
- 31:20:27into a pipeline object. We're going to
- 31:20:29see how that looks. Um, and finally,
- 31:20:33we're going to from skarn.compose, we're
- 31:20:36going to use a column transformer. The
- 31:20:38reason we're going to do this is because
- 31:20:40we are going to specify for some columns
- 31:20:44like the numerical features, we should
- 31:20:46be scaling.
- 31:20:48For some columns like the categorical
- 31:20:50features, we should be one hot encoding.
- 31:20:53So the column transformer will allow us
- 31:20:55to map different transformations to
- 31:20:58different sections of columns which is
- 31:21:00really useful. So this is actually going
- 31:21:02to be a critical part of our pipeline to
- 31:21:05apply to make sure we only apply this to
- 31:21:07numerical features and only apply this
- 31:21:10to categorical features. Right? So this
- 31:21:14column transformer will help us um to to
- 31:21:18apply pre-processing to particular
- 31:21:20columns. Um like that ocean proximity is
- 31:21:24the only one that really needs this but
- 31:21:26every other column is going to need this
- 31:21:28all the numerical features.
- 31:21:31So we're going to use this column
- 31:21:32transformer and again we're going to see
- 31:21:34how this looks but just trying to give
- 31:21:36you an idea of why we're importing all
- 31:21:37these things.
- 31:21:43Okay, so let's import those.
- 31:21:46Uh, this mentions about the column
- 31:21:48transformer. We just talked about it. It
- 31:21:50allows us to have a particular column or
- 31:21:53group of columns get the right
- 31:21:54transformation. So again, uh, looking
- 31:21:57ahead to our pipeline, the numerical
- 31:22:00features are the ones that are going to
- 31:22:02need scaling, but the categorical
- 31:22:05features are the ones that are going to
- 31:22:06need one hot encoding. However many
- 31:22:07categoricals there are, in this case,
- 31:22:09there's really only one, which is that
- 31:22:10ocean proximity. to go back to our data.
- 31:22:14Um, you can even see that in the info,
- 31:22:16there's just that one. Um, and we see
- 31:22:18that here, right? Just this one string
- 31:22:20column that should be one hot encoded.
- 31:22:22All these other guys should be scaled,
- 31:22:25right? They should all be uh uh standard
- 31:22:27scaled.
- 31:22:29So, this will allow us to specify those
- 31:22:32distinctions.
- 31:22:36All right. So, let's get started
- 31:22:40building our pipeline. So, this is going
- 31:22:42to be really cool. We're going to build
- 31:22:43out the pipeline. Um, let's extract our
- 31:22:48numerical data and our categorical data.
- 31:22:50Now, this is a really neat way of doing
- 31:22:52that that I'm not sure we've seen
- 31:22:53before. Um, so what this does is we'll
- 31:22:58take our data frame, particularly our
- 31:23:00training data frame, and select our
- 31:23:04data.
- 31:23:06That's what this select dtypes does is
- 31:23:08select data from it. Um, which includes
- 31:23:12only the object type columns. So only
- 31:23:15the object types. Now what's that? The
- 31:23:18object type is the string, right? So
- 31:23:21this should select only this column
- 31:23:24because it's in the include.
- 31:23:27We go here include only object types in
- 31:23:30the result. And so this should only have
- 31:23:33our one categorical column which is the
- 31:23:36ocean proximity. So housing cat is going
- 31:23:39to have a reference to our uh it's going
- 31:23:43to be a list that has a a basically just
- 31:23:46our ocean proximity feature because this
- 31:23:50select dtypes will make sure we only
- 31:23:52pick object types and um
- 31:23:56grab those columns. So this is a way to
- 31:24:00neatly grab um our categorical features
- 31:24:04here by including the object types. Now
- 31:24:08on the flip side we can exclude object
- 31:24:10types and get everything else. So this
- 31:24:12is going to be all other columns which
- 31:24:15is excluding the object. So this is
- 31:24:18excluding this meaning we should get all
- 31:24:21of our numerical features that way. So
- 31:24:25this will be all of our numericals
- 31:24:27by excluding the object type and this
- 31:24:31will be our housing num which is short
- 31:24:34for numerical. So this excludes
- 31:24:38the uh object type
- 31:24:42meaning all numerical
- 31:24:46features
- 31:24:49right all numerical features there.
- 31:24:52Okay.
- 31:24:55So, if we were to uh let's double check
- 31:24:58this. Let's sanity check this. If we
- 31:24:59were to print out the housing
- 31:25:03cat, um this should be just the ocean
- 31:25:07proximity feature, which it is. So, just
- 31:25:10that one. If we were to print out the
- 31:25:12housing num, this should be all the
- 31:25:14numerical features, which are all these
- 31:25:17guys. So it's just a reference to those
- 31:25:19columns so that we can uh use those
- 31:25:23later when we're mapping uh this
- 31:25:25transform needs to go to this column
- 31:25:27like the one hot encoding needs to go to
- 31:25:30this column and the scaling needs to go
- 31:25:32to these columns right so we have those
- 31:25:36uh names of those columns already at our
- 31:25:39disposal. So, we're just doing that.
- 31:25:44And this is just a
- 31:25:47simple check.
- 31:25:51Uh, are you guys able to run this?
- 31:25:57If you're following along, let me pause
- 31:25:58there. Make sure I'm not going too fast.
- 31:26:06Uh it so the the issue with a specific
- 31:26:09data type like that is none of these are
- 31:26:11ants. They're actually all floats. So we
- 31:26:14did float. I think that should work. But
- 31:26:16yes, that's the idea.
- 31:26:21Great. I'm glad to hear that right there
- 31:26:23with me. Great. Glad to hear that.
- 31:26:34Okay. So, we have our columns picked out
- 31:26:36here, which we're going to use later.
- 31:26:39Okay.
- 31:26:42All right. So, let's go ahead and build
- 31:26:46out our steps for each of these types.
- 31:26:50So, um for our numerical features, let's
- 31:26:55build out our pipeline steps. So what
- 31:26:57we're going to do is build out a
- 31:26:58numerical pipeline. And it's going to be
- 31:27:01a pipeline with a list
- 31:27:05of tupils. And the reason these are
- 31:27:08tupils is because every tupil has a
- 31:27:11name. So here this is a name that we can
- 31:27:13it can be whatever we want it to be. So
- 31:27:16we're calling it imputer. We could call
- 31:27:18it anything we want. We could call it
- 31:27:19fill in the blanks. We could call it
- 31:27:21null filling. Call it whatever you want.
- 31:27:25We're calling it imputer because it's
- 31:27:26that's a pretty um easy name for it. An
- 31:27:30accurate name to what it's doing. Um but
- 31:27:33the important thing is after the name
- 31:27:36you give it, you put in the scikitlearn
- 31:27:39object that you are going to use to
- 31:27:41operate on your data. So in this case,
- 31:27:44we're using a simple impery
- 31:27:48of median. Now that's a choice. We could
- 31:27:51use a strategy of mean, max.
- 31:27:55Um, we could provide it a constant
- 31:27:58default value. But what this means is we
- 31:28:02are going to fill any blanks we find in
- 31:28:04those columns with the median value of
- 31:28:07that column. That's the strategy for the
- 31:28:09imper. So that's pretty cool. This is
- 31:28:11kind of an automated way to fill in the
- 31:28:13blanks using for any column using its
- 31:28:17median,
- 31:28:19right? And so we could change that. We
- 31:28:21could put mean here or max or min or
- 31:28:23whatever. Um
- 31:28:26but we are filling in the blank on any
- 31:28:28column with its median. And the reason
- 31:28:31this works is because we are going to
- 31:28:33apply this pipeline only to these
- 31:28:35numerical features. So that is fine.
- 31:28:39We're we're not going to apply it to the
- 31:28:41categorical features. We're going to
- 31:28:42apply it to only those numerical. So it
- 31:28:45should have a median value, right? So
- 31:28:48that that's totally fine. So we're going
- 31:28:51to now look at how we're constructing
- 31:28:53the steps. We have a list of tupils.
- 31:28:56Here's one tupil
- 31:28:59which is the imper with a simple imper
- 31:29:02of strategy median. And then we can have
- 31:29:05as many tupils as we want which
- 31:29:07represent processing steps. So every let
- 31:29:10me write that down. Every tupil
- 31:29:14represents
- 31:29:16a pre-processing
- 31:29:18step on our data.
- 31:29:22Okay, so we have an imputer step named
- 31:29:26imputer and the reason it has a name is
- 31:29:29just so you can reference it in the
- 31:29:31pipeline if you need to. So you so it
- 31:29:33has like a a reference name um that you
- 31:29:37give it. Um but this is the more
- 31:29:39important part is the actual scikitlearn
- 31:29:41object that's doing the processing. So
- 31:29:44in this case a simple computer but
- 31:29:46notice that we have a secondary step
- 31:29:48which is our scaling. Now this makes
- 31:29:50sense. This is something we should be
- 31:29:51doing to our features is we should be
- 31:29:54scaling them. So here we we say okay
- 31:29:57let's fill in any blanks first.
- 31:30:00By the way order
- 31:30:03matters.
- 31:30:06So, and what I mean by that is the
- 31:30:10simple imper
- 31:30:13is before the scaler. Now, that's
- 31:30:17important because what that means is we
- 31:30:20should be filling in any blanks before
- 31:30:22we attempt scaling.
- 31:30:25So, that order actually matters. We're
- 31:30:27going to fill in blanks first in this
- 31:30:30list. That's first. We're going to fill
- 31:30:33in blanks. Then we are going to scale
- 31:30:38right then we scale which makes sense
- 31:30:41right so we we fill in blanks first then
- 31:30:44we apply the scaler to scale our
- 31:30:45features so those are our two steps
- 31:30:50so so pretty simple um we are building
- 31:30:54out our two steps now this is just one
- 31:30:57piece of the puzzle we are going to put
- 31:30:59this pipeline together with our one hot
- 31:31:01encoding that's going to be coming up
- 31:31:04next and build out our final pipeline.
- 31:31:07But this is um a a pipeline that has two
- 31:31:10steps that will actually be used with a
- 31:31:12larger pipeline coming up where we we do
- 31:31:15one hot encoding to our categoricals and
- 31:31:17then we put a model in there at the end
- 31:31:20to train and and use for prediction. So
- 31:31:24um pipelines can actually be composed is
- 31:31:28is uh something to realize there is that
- 31:31:30we can have a pipeline that contains a
- 31:31:33few steps. We can have another pipeline
- 31:31:34over here that contains a few steps and
- 31:31:36we can actually um kind of put them
- 31:31:38together into a final pipeline that has
- 31:31:40both pipelines uh kind of merged
- 31:31:43together. Okay. So we're going to see
- 31:31:45that coming up when we construct our
- 31:31:47final one. Our final one, as you can
- 31:31:49imagine, needs to handle this mapping of
- 31:31:52basically saying, let's do one hot
- 31:31:54encoding to these guys and then do this
- 31:31:57pipeline here to these numerical
- 31:32:00features. That's what our final pipeline
- 31:32:03needs to handle, and it will. We're
- 31:32:05going to build that out.
- 31:32:08But let me pause here. Um, were you guys
- 31:32:12able to run this? Are you with me on
- 31:32:15this this pipeline here?
- 31:32:18Does that make sense? Those two steps
- 31:32:20one is filling in blanks with a median
- 31:32:24whatever column. So where so this is
- 31:32:27this is what's so amazing about this is
- 31:32:30this is going to automatically search
- 31:32:32for nles and if you come across a column
- 31:32:36with a null, it's going to use the
- 31:32:39median of that column
- 31:32:42to fill in the blank, right? To fill in
- 31:32:44those nles.
- 31:32:58Okay,
- 31:33:01great. Glad to hear. Glad to hear.
- 31:33:05Okay.
- 31:33:07All right. So we are going to now um put
- 31:33:13this together with a column transformer
- 31:33:17to basically say what steps are going to
- 31:33:20be mapped to what columns.
- 31:33:24Um so now you can see what we're doing
- 31:33:27here is using the column transformer
- 31:33:29which is going to be a list of tupils
- 31:33:31again. So this is another um list of
- 31:33:35tupils.
- 31:33:37But the important thing is um
- 31:33:41each tupil
- 31:33:44has a name
- 31:33:47followed by so it has a name uh which
- 31:33:50again is is generic. You can say
- 31:33:52whatever you want it to be. So here
- 31:33:54we're kind of shortening this to
- 31:33:55numerical. This is short for
- 31:33:56categorical. But the important thing is
- 31:33:59it's followed by a pipeline
- 31:34:03slashstep
- 31:34:06followed by a pipeline slashstep
- 31:34:09um followed by a uh followed by a list
- 31:34:15of columns that it applies to. So you
- 31:34:19can see that pattern here. What we're
- 31:34:22saying is we're going to apply that
- 31:34:24numerical pipeline we just defined. So
- 31:34:26this is saved in a numerical pipeline
- 31:34:29object here. We're going to apply that
- 31:34:32to those numerical features. So this is
- 31:34:34that list
- 31:34:36of numerical features here. So that's
- 31:34:39how we do the mapping. We have a tupil
- 31:34:41here that says okay apply these steps to
- 31:34:44these columns.
- 31:34:46Those go together in that tupole, right?
- 31:34:49Apply these steps to this uh these
- 31:34:52columns. And then apply this step. Now
- 31:34:55what is the step? This is a one hot
- 31:34:58encoder
- 31:34:59which is going to uh uh encode um those
- 31:35:04features and it's going to uh ignore um
- 31:35:09basically nles for now. That's a choice
- 31:35:11but it's going to ignore um uh basically
- 31:35:16ignore nles and and uh skip over them
- 31:35:19for now. We now we know there's no NLES
- 31:35:23because we already did an is NA from
- 31:35:25before and we know there's not any NLES
- 31:35:28in that ocean proximity. So this isn't
- 31:35:30going to be an issue. But that's what
- 31:35:32that would do.
- 31:35:34But we have a one hot encoder here which
- 31:35:37we're going to apply to our categorical
- 31:35:40features. Now of course that's just the
- 31:35:43ocean proximity feature but that but
- 31:35:46again you see the pattern in the tupil
- 31:35:47is apply this transform which is a one
- 31:35:50hot encoding to this column apply these
- 31:35:53numerical transforms which is a whole
- 31:35:55pipeline. So it's two steps in a
- 31:35:59pipeline of um
- 31:36:02uh an imputer and a scaler are going to
- 31:36:05be applied to this
- 31:36:08really nice. So those are going to be
- 31:36:10all together in this column transformer.
- 31:36:12And that is our way to signal that for
- 31:36:14these numerical features, use these
- 31:36:16steps. For our categorical features, use
- 31:36:18this step. And and you know, if we had
- 31:36:21more than one step, we were applying to
- 31:36:23categorical. We could build a pipeline
- 31:36:25for the categorical and it would and do
- 31:36:27the same thing. We have more than one
- 31:36:29step here. And so it's good practice
- 31:36:32when you have more than one step to just
- 31:36:33put that in a pipeline because we have
- 31:36:35more than one step. We'll just put that
- 31:36:37in this list inside of the pipeline and
- 31:36:39we can map that pipeline to those
- 31:36:42features. Here we only have one step. So
- 31:36:45it's okay to just put that there um and
- 31:36:48apply that to the categorical features.
- 31:36:51But if we had more than one step um it
- 31:36:54would be good practice to put that in a
- 31:36:56pipeline
- 31:36:57which is what we do here. Right? This
- 31:36:58pipeline is being mapped to these
- 31:37:00features. This step is being applied to
- 31:37:03this feature.
- 31:37:08Okay,
- 31:37:10how about that? Are you guys able to run
- 31:37:13that one? Does that make sense what we
- 31:37:15have set up so far? So, we're almost
- 31:37:17there. We almost have our final
- 31:37:19pipeline. We have our pre-processing
- 31:37:20basically done to say our numerical
- 31:37:23features should be processed with that
- 31:37:24other pipeline and our categorical
- 31:37:27features should be one hot encoded.
- 31:37:29We're getting close. The only thing
- 31:37:31we're really missing here is a model.
- 31:37:34The only thing we're really missing is
- 31:37:36to have our final model training
- 31:37:39pipeline is to actually include a model
- 31:37:42which should come at the end.
- 31:37:44Right? So it should we should be doing
- 31:37:46these steps first
- 31:37:49then doing modeling which we know right
- 31:37:52we we've done that uh many times. We've
- 31:37:54done our pre-processing and then we do
- 31:37:55our modeling.
- 31:38:00Any questions on that?
- 31:38:17Okay.
- 31:38:18Fantastic.
- 31:38:22All right.
- 31:38:25So, if we wanted to uh see if we wanted
- 31:38:28to test this so far, um we could. So, we
- 31:38:31could run the pre-processing and
- 31:38:33actually run a fit transform on our data
- 31:38:36and this will um basically apply that
- 31:38:39pipeline to the data. Now, this would be
- 31:38:41a sanity check. This is a good This is a
- 31:38:44good kind of um This is a good sanity
- 31:38:47check that our pre-processing
- 31:38:52works. So, it's doing what we expected
- 31:38:55to do. It's not our final pipeline
- 31:38:57because we don't have our model in there
- 31:38:59yet, but this is just to ensure that all
- 31:39:01of the features are kind of behaving as
- 31:39:03we expect. So, we can uh we can do that
- 31:39:07and we can take a look at the um
- 31:39:09results. This looks pretty good. this
- 31:39:11all of our numerical features ended up
- 31:39:13scaled
- 31:39:15which is pretty good and we have one hot
- 31:39:17encoded features for that ocean
- 31:39:19proximity over here.
- 31:39:22Okay, so this looks pretty this looks
- 31:39:24reasonable of those steps being applied
- 31:39:27to the right columns. But this is a good
- 31:39:29kind of sanity check to just run our fit
- 31:39:31transform on our data to ensure those
- 31:39:35steps are actually happening and they
- 31:39:37are. You can see here the result of the
- 31:39:40scaling and the uh the one hot encoding.
- 31:39:44So that that all looks pretty
- 31:39:46reasonable,
- 31:39:49right? And uh what we should also do is
- 31:39:55make sure there are no nulls in this
- 31:39:56which there shouldn't be because we did
- 31:39:58the imputer. So we should be doing uh is
- 31:40:01na dot
- 31:40:05sum
- 31:40:09And there is no NLES anymore. So that
- 31:40:11looks pretty good, right? Those got
- 31:40:13filled in uh by doing our steps. Our
- 31:40:17pipeline steps executed really nicely on
- 31:40:20our training data. Um and and we were
- 31:40:24off and running. And there's nothing
- 31:40:25unique about the training data. We could
- 31:40:27do this to our test data as well
- 31:40:30and verify that those steps are running
- 31:40:32and they would, right? There's nothing
- 31:40:34really that special about running it on
- 31:40:35the training data. Um, it should also
- 31:40:39work on the test features as well, and
- 31:40:40it does. You can check that for
- 31:40:43yourself.
- 31:40:45Okay.
- 31:40:50All right. So, that's pretty cool. We
- 31:40:52can uh verify all that's working.
- 31:41:01Any questions on that?
- 31:41:04We're almost there with our full
- 31:41:06pipeline. This this is this is not the
- 31:41:08full pipeline, but this is something
- 31:41:10that will run during our full pipeline.
- 31:41:13Of course, our features are going to be
- 31:41:14transformed according to those steps and
- 31:41:16then it will be uh put into our model to
- 31:41:19either predict or train with. Um
- 31:41:22so let's do that. Let's actually build
- 31:41:24out our final uh model here. So it's
- 31:41:29actually going to be really easy to do.
- 31:41:30All we need to do is um put in our
- 31:41:34model. So here we're going to import the
- 31:41:36ridge model here. Now we could use any
- 31:41:39we could use linear regression, we could
- 31:41:41use lasso, we could use elastic net. Um
- 31:41:43we're just going to use ridge um uh um
- 31:41:47just to test it out. And um we are going
- 31:41:51to uh now put in a final pipeline. So
- 31:41:55we're going to use our pipeline. And so
- 31:41:58we're going to create a new one here.
- 31:42:00and map our pre-processing to our
- 31:42:03pre-processing that we've already built.
- 31:42:05So, this is a column transformer that
- 31:42:08already has all of our steps. And then
- 31:42:10notice what comes after it is just the
- 31:42:12model. Now, that's pretty pretty basic,
- 31:42:14but it makes sense that it should come
- 31:42:16after that model. Um, and of course,
- 31:42:19this is a generic name. We could we can
- 31:42:21name it whatever we want to. um model
- 31:42:24ridge is pretty reasonable um to because
- 31:42:28it is a ridge uh regression but uh of
- 31:42:31course we could we could change that.
- 31:42:35Okay, so that builds out our uh final um
- 31:42:38pipeline. So now we have a pipeline and
- 31:42:42what's great about that is this signals
- 31:42:45that all of these steps should be
- 31:42:47completed prior to doing anything with
- 31:42:49this model. So all of those processing
- 31:42:52steps are going to run and then we're
- 31:42:54going to do ffit orpredict and that so
- 31:42:57that's really great. It ensures that all
- 31:42:59those steps are running together every
- 31:43:01single time we call predict with this
- 31:43:04with this model. So we're just going to
- 31:43:06use the pipeline in place of the model
- 31:43:10to ensure that all of those steps are
- 31:43:12running together. And this is our this
- 31:43:15is kind of our final pipeline that we
- 31:43:17would use uh with like something like
- 31:43:19ffit or predict.
- 31:43:22So let me make that uh a note of that.
- 31:43:24Now we can use this final pipeline just
- 31:43:30like a regular model i.e. pipeline.fit
- 31:43:36or pipeline.predict.
- 31:43:40We could use it in ei in either fashion
- 31:43:43uh to to train the pipeline would be
- 31:43:47this guy and then use the pipeline to
- 31:43:48predict would be this. And what we
- 31:43:50should realize is under the hood these
- 31:43:52steps are running first and then we
- 31:43:54train it or these steps run first then
- 31:43:57we use it for prediction.
- 31:44:04Okay.
- 31:44:06Questions on that? Does that make sense?
- 31:44:09On this final pipeline here, it's just
- 31:44:12now it it's really cool because we have
- 31:44:14a pipeline
- 31:44:16made up of a of a pipeline really,
- 31:44:19right? A pipeline made up of a pipeline.
- 31:44:21But that's scikitlearn allows you to do
- 31:44:22that to compose pipelines in this way.
- 31:44:26That's that's pretty uh pretty uh normal
- 31:44:29there.
- 31:44:39Okay,
- 31:44:41what I want to show you is we can
- 31:44:44actually use this pipeline in a grid
- 31:44:46search. So that's pretty amazing. We can
- 31:44:48use this pipeline in any way we can use
- 31:44:51a mo like a regular model. It's just
- 31:44:53that now our pre-processing steps have
- 31:44:56kind of been packaged together with our
- 31:44:58model to ensure that they always run
- 31:45:01anytime we do any processing with this
- 31:45:03model. Um so for instance we can do a
- 31:45:07grid search just like we did with a
- 31:45:09regular with with just a model right
- 31:45:11with just this. Um we can do the same
- 31:45:14thing with the whole pipeline. Um, so
- 31:45:17the only catch is that you want to make
- 31:45:20sure in your grid you name things in the
- 31:45:24appropriate way inside of your your uh
- 31:45:26keys in your dictionary. So uh for
- 31:45:30instance um inside of the grid uh we're
- 31:45:34going to set up the alpha that would be
- 31:45:36used with this ridge regression by
- 31:45:39referencing its name. So this is model
- 31:45:41ridge is this is the name of the model
- 31:45:44inside of the pipeline. So you want to
- 31:45:46make sure that goes first.
- 31:45:48And then what scikitlearn does is it
- 31:45:51recognizes parameters that belong with
- 31:45:53this model by using a double underscore.
- 31:45:57So the so you have underscore underscore
- 31:46:00alpha um here. So the double
- 31:46:05uh underscore
- 31:46:08signals a parameter
- 31:46:11belonging to model ridge. in the in the
- 31:46:17pipeline.
- 31:46:19Okay, so we have a model ridge is just a
- 31:46:22reference to the model in our pipeline.
- 31:46:24That's the one we're going to test out
- 31:46:26these parameters with. And
- 31:46:28underscore_pha is just a way to say this
- 31:46:31alpha belongs to this model. Okay, it
- 31:46:36belongs so it's going to be used with
- 31:46:38that model in our pipeline. Um otherwise
- 31:46:42it's going to work exactly the same way.
- 31:46:44It's just we need to line up this naming
- 31:46:46convention of of scikitlearn.
- 31:46:48You just have to reference this to
- 31:46:51whatever name you provided here and then
- 31:46:54underscore parameter. So L1 ratio alpha
- 31:46:58whatever right would go there.
- 31:47:01Okay. So there is a range from 0.1 to
- 31:47:05two uh step size of 0.1
- 31:47:09um and then we do our grid search CV. So
- 31:47:12this is exactly the same setup as we had
- 31:47:14before. It's just that our model is now
- 31:47:18the pipeline. So our pipeline is going
- 31:47:20in there. Um we have our grid going in
- 31:47:23there. We have our scoring is the same,
- 31:47:26you know, negative absolute error. Um
- 31:47:28we're using five-fold cross validation
- 31:47:31and we're parallelizing that search. Um,
- 31:47:34so we're going to search through these
- 31:47:35alphas and uh basically fit this to our
- 31:47:40um data and find the best um find the
- 31:47:45best alpha.
- 31:47:47So it's going to try out all those
- 31:47:49combinations and try to come up with the
- 31:47:51best alpha.
- 31:47:54So looks like the best alpha was 0.1 for
- 31:47:57the ridge.
- 31:48:00Okay, is the best. So then um if we
- 31:48:04wanted to we could uh then predict using
- 31:48:08the model um which would be doing
- 31:48:11something like this. Um and we could
- 31:48:14also go back and do something like so we
- 31:48:18could
- 31:48:22now use um this param. So we could do
- 31:48:27model
- 31:48:30um equals ridge
- 31:48:35and then we could put in our alpha
- 31:48:38um alpha is our results our best
- 31:48:41parameters and then we get that model
- 31:48:43ridge alpha and then we just rebuild our
- 31:48:46our pipeline
- 31:48:50equals um pipeline and then we uh put in
- 31:48:54this new model here. So we could do
- 31:48:57this. This would be going back and just
- 31:49:00um putting in our best alpha here for
- 31:49:04this model and then uh ensuring that's
- 31:49:07part of our our pipeline. So we're just
- 31:49:09overwriting that pipeline with the best
- 31:49:10model there
- 31:49:16to get the best model in our pipeline.
- 31:49:23Okay,
- 31:49:24so that's all this is doing is just
- 31:49:26initializing a new um let me actually I
- 31:49:30can put this code in here.
- 31:49:36This is actually just getting this is
- 31:49:38just getting a model with the best alpha
- 31:49:40and then reinserting that into our our
- 31:49:43uh we're just overwriting our final
- 31:49:45pipeline there with the best model that
- 31:49:47we have.
- 31:49:50So pretty cool that pipeline can be used
- 31:49:53basically exactly like a model, right?
- 31:49:55It's it's going right here in the grid
- 31:49:57search and being used uh entirely like a
- 31:50:00basic model. So we do ffit
- 31:50:04um and that allows us to use it. We
- 31:50:06could dopredict. We could even do
- 31:50:08pipeline.predict once we we could go
- 31:50:10back and do final pipeline.fit
- 31:50:13um with this and then final
- 31:50:14pipeline.predict with this and evaluate
- 31:50:21Okay,
- 31:50:23so pretty cool that pipeline can be used
- 31:50:25uh basically exactly like how a model
- 31:50:27would be any way we' use a model.fit
- 31:50:30model.predict, we can use a pipeline.
- 31:50:34So grid search is for instance something
- 31:50:37that can use a model in there. Um but
- 31:50:39instead of just a model, we're ensuring
- 31:50:41we have our pre-processing steps kind of
- 31:50:43bundled with that model in this
- 31:50:45pipeline.
- 31:50:47Any
- 31:50:50questions on
- 31:50:52uh this example so far?
- 31:50:58Were you guys able to run it up to here?
- 31:51:00Were you able to run the grid search?
- 31:51:16Okay, great.
- 31:51:25Okay.
- 31:51:27Okay. So, this is this is uh just
- 31:51:30showing you what's actually happening
- 31:51:31underneath the hood is uh you know,
- 31:51:34we're doing some scaling. We're doing
- 31:51:36some one hot encoding um
- 31:51:40and we're doing some uh we're doing a
- 31:51:43model here. And that's all part of our
- 31:51:46pipeline. Um, and then we can use the
- 31:51:50pipeline however we want. So for
- 31:51:52example, I know it's not here, but for
- 31:51:54an example, we could use um once we do
- 31:51:57once we have this final pipeline um we
- 31:52:00can can use the final um
- 31:52:05pipeline to predict. So we can do um
- 31:52:09predictions
- 31:52:11equals final
- 31:52:14pipeline.predict
- 31:52:16and then we can pass in our test data.
- 31:52:18Now what happens on this is once we have
- 31:52:22ran our our pipeline.fit we have a
- 31:52:25trained pipeline and then when we run
- 31:52:27this final pipeline.predict uh this data
- 31:52:30is going to be transformed.
- 31:52:32It's going to go through those
- 31:52:33transformation steps and then we would
- 31:52:35apply our model to it at the end uh to
- 31:52:39to make those predictions and then we
- 31:52:40can evaluate those predictions which is
- 31:52:42what we're doing kind of here.
- 31:52:46Right?
- 31:52:52Okay.
- 31:52:56All right. So in conclusion uh we have
- 31:52:59gone through a lot of stuff here. Um,
- 31:53:02we've gone through regression, we've
- 31:53:04done the regularization on regression.
- 31:53:08So hopefully we have a good foundation
- 31:53:09on regression. Um, what we're going to
- 31:53:11do in a little bit is actually do some
- 31:53:13additional practice with regression on a
- 31:53:15new problem. We're going to do a
- 31:53:16capstone problem and do some additional
- 31:53:20regression work with that. Um, so we'll
- 31:53:23do that next. Um but the other thing we
- 31:53:27learned is how to evaluate the
- 31:53:28regression using things like mean
- 31:53:30squared error, RMSSE which is square
- 31:53:32root of that. Um which is which is
- 31:53:35really cool. So we have a sense of that
- 31:53:38error which is our distance from our
- 31:53:40prediction to the actual value. That's
- 31:53:42always what these uh that's always what
- 31:53:45these things are doing like this, right?
- 31:53:48This mean absolute error metric from
- 31:53:50scikitlearn is computing the average
- 31:53:53distance from these predictions to these
- 31:53:55test labels that we have, right? Those
- 31:53:58actual values. Um, and that gives us a
- 31:54:00sense of on average how far away are our
- 31:54:03predictions
- 31:54:04um to see how good of a model that we
- 31:54:07have, right? And we should be evaluating
- 31:54:10that error generally
- 31:54:12um against the scale of our targets to
- 31:54:16see, you know,
- 31:54:19uh how far off we typically are.
- 31:54:23Okay. Any questions at all on this
- 31:54:25lesson on regression? Uh anything we
- 31:54:27covered up to this point? We're going to
- 31:54:30do some more practice with the next
- 31:54:32we'll do the capstone. So we get so we
- 31:54:35just do some more regression problems.
- 31:54:47Yeah, it's a that's another bad score.
- 31:54:49It's a little bit hard to interpret this
- 31:54:50though because it's m ae. Um, so one
- 31:54:54thing we could do is is compute mean
- 31:54:57squared error and then take the square
- 31:55:00root of it to get the RMSSE which is a
- 31:55:02much better uh evaluation metric in
- 31:55:05terms of our target. Um so we could
- 31:55:08actually run that. Uh if we go back here
- 31:55:11and um we could generate for instance we
- 31:55:14could generate the MSE which is the mean
- 31:55:17squared
- 31:55:20error
- 31:55:21and it's it's the same exact function uh
- 31:55:25of using our predictions.
- 31:55:28Um
- 31:55:30and then we could just print that out.
- 31:55:32Mean squared error.
- 31:55:36So we have mean squared error and then
- 31:55:38what we can do is let's take the um MP.
- 31:55:43square root of that.
- 31:55:46So that way we can generate the RMSSE.
- 31:55:49So yeah that I mean that's pretty bad.
- 31:55:50That's uh pretty bad. Uh now let's let's
- 31:55:55go back and look at our
- 31:55:58uh data though. So let's take a look at
- 31:56:00the average for our y. Um remember one
- 31:56:03thing we should be doing is taking a
- 31:56:05look at um what our uh let's take a look
- 31:56:08at y test mean
- 31:56:11to get an average value. So the average
- 31:56:14value is in the 200,000s. So
- 31:56:18this isn't this isn't awful. This is
- 31:56:2170,000. It's still a decent amount of
- 31:56:23error. It's not as bad as the models we
- 31:56:25have before though, right? This is an
- 31:56:28average median price of the house is in
- 31:56:32the 206,000 range and our error is off
- 31:56:36by like 70,000,
- 31:56:39right?
- 31:56:41So, it's not good. Um, but it's not
- 31:56:47hor like as bad as the it's not as
- 31:56:50horrible as we've seen so far. Right.
- 31:56:52This is a little bit better of a model.
- 31:56:54A little bit better. closer to zero
- 31:56:56would be better, right? Um but the
- 31:56:59smaller the better. But uh remember this
- 31:57:02is the um these even the mean absolute
- 31:57:06error is is technically in similar units
- 31:57:09as the as the uh um
- 31:57:14as the target. So 50,000 60,000 here
- 31:57:1770,000 it's still a decent amount of
- 31:57:19error in terms of 200,000.
- 31:57:24Uh so far we only come up with models
- 31:57:25and test their accuracy with available
- 31:57:27data. We haven't used a model to make
- 31:57:28completely new predictions on No, we
- 31:57:31haven't done that. Uh except we know how
- 31:57:33to do that. Um it would so to make
- 31:57:36predictions on new data would be exactly
- 31:57:38how we're making them on our available
- 31:57:40data because we actually do that all the
- 31:57:43time. If we go back down to our model
- 31:57:45building,
- 31:57:47um it's it looks just like this, right?
- 31:57:49where we take so for instance we do
- 31:57:53predictions all the time on test data
- 31:57:56that was never involved in the training.
- 31:57:58So it's it's as if this data mimics new
- 31:58:02data that we've never seen before. So if
- 31:58:06we had new raw data it would just it
- 31:58:08would be the same exact process. the new
- 31:58:11now with our pipeline it makes it a
- 31:58:14little bit easier because with the
- 31:58:15pipeline
- 31:58:17um the raw data will go through those
- 31:58:19transformations which it should right
- 31:58:21the raw data should because if it's
- 31:58:22missing data it needs to be filled in if
- 31:58:24it has categoricals it needs to be one
- 31:58:26hot encoded so that's the purpose of the
- 31:58:29pipeline actually is to make sure that
- 31:58:33if we're dealing with raw data um those
- 31:58:37steps can happen on the data before it
- 31:58:40goes into the model. Right?
- 31:58:43So we so
- 31:58:45that's kind of the purpose of the
- 31:58:47pipeline
- 31:58:48is to ensure that we run those steps
- 31:58:51ahead of using it using a model with it.
- 31:58:56But but ultimately that's how it uh any
- 31:58:58scikitlearn model is going to be doing
- 31:59:00the predict even if it's a pipeline
- 31:59:02right it's going to be uh we just go
- 31:59:05back down here it's going to be um
- 31:59:08predict it's always going to be that on
- 31:59:10new data
- 31:59:11>> hello everyone in this session we will
- 31:59:14cover all the important supervised and
- 31:59:16unsupervised learning algorithms that
- 31:59:17are widely used with hands-on
- 31:59:18demonstrations in Python our instructors
- 31:59:21with rich experience in machine learning
- 31:59:23will take us through this course but
- 31:59:25before Before we begin, make sure to
- 31:59:26subscribe to the SimplyLearn channel and
- 31:59:28hit the bell icon to never miss an
- 31:59:29update.
- 31:59:31So, we will start by understanding the
- 31:59:33basics of machine learning from a short
- 31:59:35animated video followed by the
- 31:59:37difference between supervised and
- 31:59:38unsupervised learning. We will then jump
- 31:59:40into learning the various algorithms
- 31:59:43from scratch. So, we will understand
- 31:59:45linear regression, logistic regression,
- 31:59:47decision tree, and random forest. We'll
- 31:59:50then look at support vector machines and
- 31:59:52K nearest neighbor algorithm with a
- 31:59:53hands-on demonstration in Python.
- 31:59:56Finally, we'll get an idea about
- 31:59:58unsupervised learning algorithms such as
- 32:00:00K means clustering and principal
- 32:00:01component analysis. We will conclude
- 32:00:04this session with regularization in
- 32:00:06machine learning. So let's get started.
- 32:00:09>> We know humans learn from their past
- 32:00:11experiences and machines follow
- 32:00:13instructions given by humans.
- 32:00:16But what if humans can train the
- 32:00:18machines to learn from their past data
- 32:00:20and do what humans can do and much
- 32:00:22faster? Well, that's called machine
- 32:00:24learning. But it's a lot more than just
- 32:00:26learning. It's also about understanding
- 32:00:28and reasoning. So today we will learn
- 32:00:30about the basics of machine learning. So
- 32:00:33that's Paul. He loves listening to new
- 32:00:36songs.
- 32:00:38He either likes them or dislikes them.
- 32:00:40Paul decides this on the basis of the
- 32:00:42song's tempo, genre, intensity, and the
- 32:00:46gender of voice. For simplicity, let's
- 32:00:49just use tempo and intensity for now.
- 32:00:52So, here tempo is on the x-axis, ranging
- 32:00:55from relaxed to fast, whereas intensity
- 32:00:58is on the y-axis, ranging from light to
- 32:01:01soaring. We see that Paul likes the song
- 32:01:04with fast tempo and soaring intensity
- 32:01:07while he dislikes the song with relaxed
- 32:01:10tempo and light intensity. So now we
- 32:01:13know Paul's choices. Let's say Paul
- 32:01:15listens to a new song. Let's name it as
- 32:01:17song A. Song A has fast tempo and a
- 32:01:20soaring intensity. So it lies somewhere
- 32:01:23here. Looking at the data, can you guess
- 32:01:25whether Paul will like the song or not?
- 32:01:27Correct. So Paul likes this song. By
- 32:01:30looking at Paul's past choices, we were
- 32:01:32able to classify the unknown song very
- 32:01:35easily, right? Let's say now Paul
- 32:01:37listens to a new song. Let's label it as
- 32:01:40song B. So song B lies somewhere here
- 32:01:44with medium tempo and medium intensity.
- 32:01:47Neither relaxed nor fast, neither light
- 32:01:50nor soaring. Now, can you guess whether
- 32:01:52Paul likes it or not? Not able to guess
- 32:01:54whether Paul will like it or dislike it.
- 32:01:56Are the choices unclear? Correct. We
- 32:01:59could easily classify song A. But when
- 32:02:02the choice became complicated as in the
- 32:02:04case of song B. Yes. And that's where
- 32:02:06machine learning comes in. Let's see
- 32:02:08how. In the same example for song B, if
- 32:02:11we draw a circle around the song B, we
- 32:02:13see that there are four votes for like
- 32:02:15whereas one vote for dislike. If we go
- 32:02:18for the majority votes, we can say that
- 32:02:20Paul will definitely like the song.
- 32:02:22That's all. This was a basic machine
- 32:02:24learning algorithm also. It's called K
- 32:02:26nearest neighbors. So this is just a
- 32:02:28small example in one of the many machine
- 32:02:30learning algorithms quite easy right
- 32:02:33believe me it is but what happens when
- 32:02:36the choices become complicated as in the
- 32:02:39case of song B that's when machine
- 32:02:40learning comes in it learns the data
- 32:02:43builds the prediction model and when the
- 32:02:45new data point comes in it can easily
- 32:02:47predict for it more the data better the
- 32:02:49model higher will be the accuracy there
- 32:02:52are many ways in which the machine
- 32:02:54learns it could be either supervised
- 32:02:56learning unsupervised learning or
- 32:02:58reinforcement learning. Let's first
- 32:03:00quickly understand supervised learning.
- 32:03:02Suppose your friend gives you 1 million
- 32:03:05coins of three different currencies. Say
- 32:03:071 rupee, 1 and 1 dirham. Each coin has
- 32:03:10different weights. For example, a coin
- 32:03:12of 1 rupee weighs 3 g. 1 euro weighs 7 g
- 32:03:16and 1 dirham weighs 4 g. Your model will
- 32:03:19predict the currency of the coin. Here
- 32:03:21your weight becomes the feature of coins
- 32:03:24while currency becomes their label. When
- 32:03:26you feed this data to the machine
- 32:03:28learning model, it learns which feature
- 32:03:30is associated with which label. For
- 32:03:33example, it will learn that if a coin is
- 32:03:35of 3 g, it will be a 1 rupee coin. Let's
- 32:03:38give a new coin to the machine. On the
- 32:03:40basis of the weight of the new coin,
- 32:03:42your model will predict the currency.
- 32:03:44Hence, supervised learning uses labeled
- 32:03:46data to train the model. Here, the
- 32:03:48machine knew the features of the object
- 32:03:50and also the labels associated with
- 32:03:53those features. On this note, let's move
- 32:03:55to unsupervised learning and see the
- 32:03:57difference. Suppose you have cricket
- 32:03:58data set of various players with their
- 32:04:00respective scores and wickets taken.
- 32:04:02When we feed this data set to the
- 32:04:05machine, the machine identifies the
- 32:04:07pattern of player performance. So, it
- 32:04:09plots this data with the respective
- 32:04:10wickets on the x-axis while runs on the
- 32:04:13y-axis. While looking at the data,
- 32:04:14you'll clearly see that there are two
- 32:04:16clusters. The one cluster are the
- 32:04:18players who scored high runs and took
- 32:04:21less wickets while the other cluster is
- 32:04:23of the players who scored less runs but
- 32:04:26took many wickets. So here we interpret
- 32:04:28these two clusters as batsmen and
- 32:04:30bowlers. The important point to note
- 32:04:32here is that there were no labels of
- 32:04:34batsmen and bowlers. Hence the learning
- 32:04:37with unlabeled data is unsupervised
- 32:04:39learning. So we saw supervised learning
- 32:04:41where the data was labeled and the
- 32:04:43unsupervised learning where the data was
- 32:04:45unlabeled. And then there is
- 32:04:47reinforcement learning which is a
- 32:04:48reward-based learning or we can say that
- 32:04:50it works on the principle of feedback.
- 32:04:52Here let's say you provide the system
- 32:04:54with an image of a dog and ask it to
- 32:04:56identify it. The system identifies it as
- 32:04:59a cat. So you give a negative feedback
- 32:05:01to the machine saying that it's a dog's
- 32:05:03image. The machine will learn from the
- 32:05:04feedback and finally if it comes across
- 32:05:06any other image of a dog, it'll be able
- 32:05:09to classify it correctly. That is
- 32:05:11reinforcement learning. To generalize
- 32:05:13machine learning model, let's see a
- 32:05:14flowchart. Input is given to a machine
- 32:05:16learning model which then gives the
- 32:05:18output according to the algorithm
- 32:05:20applied. If it's right, we take the
- 32:05:22output as a final result. Else we
- 32:05:24provide feedback to the training model
- 32:05:26and ask it to predict until it learns. I
- 32:05:29hope you've understood supervised and
- 32:05:31unsupervised learning. So let's have a
- 32:05:33quick quiz. You have to determine
- 32:05:35whether the given scenarios uses
- 32:05:37supervised or unsupervised learning.
- 32:05:38Simple, right? Scenario one. Facebook
- 32:05:41recognizes your friend in a picture from
- 32:05:43an album of tagged photographs.
- 32:05:46Scenario two, Netflix recommends new
- 32:05:49movies based on someone's past movie
- 32:05:51choices.
- 32:05:53Scenario three, analyzing bank data for
- 32:05:55suspicious transactions and flagging the
- 32:05:58fraud transactions. Think wisely and
- 32:06:00comment below your answers. Moving on,
- 32:06:02don't you sometimes wonder how is
- 32:06:04machine learning possible in today's
- 32:06:06era? Well, that's because today we have
- 32:06:08humongous data available. Everybody's
- 32:06:11online either making a transaction or
- 32:06:13just surfing the internet and that's
- 32:06:15generating a huge amount of data every
- 32:06:17minute. And that data my friend is the
- 32:06:20key to analysis. Also, the memory
- 32:06:22handling capabilities of computers have
- 32:06:24largely increased which helps them to
- 32:06:26process such huge amount of data at hand
- 32:06:29without any delay. And yes, computers
- 32:06:31now have great computational powers. So
- 32:06:34there are a lot of applications of
- 32:06:36machine learning out there. To name a
- 32:06:37few, machine learning is used in
- 32:06:39healthcare where diagnostics are
- 32:06:41predicted for doctor's review. The
- 32:06:43sentiment analysis that the tech giants
- 32:06:45are doing on social media is another
- 32:06:47interesting application of machine
- 32:06:49learning. Fraud detection in the finance
- 32:06:51sector and also to predict customer
- 32:06:53churn in the e-commerce sector. While
- 32:06:54booking a cab, you must have encountered
- 32:06:57search pricing often where it says the
- 32:06:59fair of your trip has been updated.
- 32:07:01Continue booking. Yes, please. I'm
- 32:07:03getting late for office. Well, that's an
- 32:07:06interesting machine learning model which
- 32:07:08is used by global taxi giant Uber and
- 32:07:10others where they have differential
- 32:07:12pricing in real time based on demand,
- 32:07:14the number of cars available, bad
- 32:07:16weather, rush hour, etc. So they use the
- 32:07:19search pricing model to ensure that
- 32:07:21those who need a cab can get one. Also,
- 32:07:24it uses predictive modeling to predict
- 32:07:27where the demand will be high with a
- 32:07:29goal that drivers can take care of the
- 32:07:31demand and search pricing can be
- 32:07:33minimized. Great. Hey Siri, can you
- 32:07:35remind me to book a cab at 6 p.m. today?
- 32:07:37>> Okay, I'll remind you.
- 32:07:39>> Thanks.
- 32:07:40>> No problem.
- 32:07:41>> Comment below some interesting everyday
- 32:07:43examples around you where machines are
- 32:07:45learning and doing amazing jobs. Hi
- 32:07:48guys, this is Acha from SimplyLearn and
- 32:07:51we're going to talk about the two types
- 32:07:52of machine learning supervised and
- 32:07:55unsupervised learning their types and
- 32:07:57applications. But before we talk about
- 32:07:59them, let's quickly understand what is
- 32:08:01machine learning. These days
- 32:08:02applications use artificial intelligence
- 32:08:05in machine learning to optimize speech
- 32:08:07recognition. I usually ask Siri things I
- 32:08:10want to know like hey Siri how far is
- 32:08:12the nearest subway? So whenever we ask
- 32:08:14something to Siri, a powerful speech
- 32:08:17recognition kicks off and converts the
- 32:08:19audio into its corresponding textual
- 32:08:21form which is then sent to the Apple
- 32:08:23servers for further processing. Then
- 32:08:26neural language processing algorithms
- 32:08:28are run to understand the user's intent
- 32:08:31and then finally Siri tells you the
- 32:08:33answer. Well, this is what machine
- 32:08:35learning is all about. making the
- 32:08:37machines learn and act like humans by
- 32:08:39feeding them with data and information
- 32:08:42without being explicitly programmed. As
- 32:08:44we saw in the previous example, when the
- 32:08:46data comes in, machines immediately
- 32:08:48starts analyzing the data and eventually
- 32:08:51gets trained on it and learns it. Now
- 32:08:53when a new data point comes in, machine
- 32:08:56accurately makes prediction and
- 32:08:58decisions based on the past data. Now
- 32:09:00that you know what is machine learning,
- 32:09:02let's talk about supervised and
- 32:09:03unsupervised learning. Supervised
- 32:09:05learning as the name suggests works
- 32:09:07under supervision that is it's a
- 32:09:09learning in which machine is trained
- 32:09:11with data which is welllabeled and then
- 32:09:13predicts with the help of the label data
- 32:09:15set. But what is a label data set? Data
- 32:09:18for which you already know the target
- 32:09:19answer is called a label data. Like I
- 32:09:22show you an image and tell you that it's
- 32:09:25a dog then it's a label data. While if I
- 32:09:28show you an image without telling you
- 32:09:30what exactly it is, then it's an
- 32:09:32unlabelled data. Now let's say we have
- 32:09:35images which are labeled as spoon or
- 32:09:37knife. We then feed it to the machine
- 32:09:39which analyzes and learns the
- 32:09:41association of these images with its
- 32:09:43labels based on its features such as
- 32:09:46shape, size, sharpness etc. Now when a
- 32:09:50new image is fed to the machine without
- 32:09:52any label with the help of the past data
- 32:09:54the machine is able to predict
- 32:09:56accurately and tell that it's a spoon.
- 32:09:58Hence in supervised machine learning the
- 32:10:00algorithm teaches the model to learn
- 32:10:02from the labeled example that we
- 32:10:04provide. So supervised learning can be
- 32:10:07further divided into classification and
- 32:10:09regression. It is a classification
- 32:10:11problem when the output variable is
- 32:10:13categoric such as red or blue, disease
- 32:10:16or no disease, male or female. Whereas
- 32:10:19it's a regression problem when the
- 32:10:20output variable is a real or continuous
- 32:10:23value. For example, salary based on work
- 32:10:26experience, weight based on height. So
- 32:10:29it creates a predictive model showing
- 32:10:31trends in data. So now if I say will I
- 32:10:34get a salary raise or not, that's
- 32:10:36classification. But if I say how much
- 32:10:38salary raise will I get, that is
- 32:10:41regression. Now let's understand
- 32:10:42classification with the help of an
- 32:10:44example. In order to predict whether an
- 32:10:46email is a spam or not, first we need to
- 32:10:49teach our machine what a spam mail looks
- 32:10:51like. This is done based on a lot of
- 32:10:54spam filters like firstly reviewing the
- 32:10:57content of the email then review the
- 32:10:59email header and search if it contains
- 32:11:01any falsified information. This is done
- 32:11:03based on some keywords like free lottery
- 32:11:08prize claim etc. Then general blacklist
- 32:11:12filters to stop emails that come from
- 32:11:14already blacklisted known spammers and
- 32:11:18etc. So all these filters scores the
- 32:11:20email which is known as spam score. The
- 32:11:23lower the total spam score of the email,
- 32:11:26it is more likely that the email will
- 32:11:28land in subscribers inboxes. So based on
- 32:11:31the content labels and spam score of the
- 32:11:34new incoming mail, the algorithm decides
- 32:11:37whether it should land in inbox or the
- 32:11:40spam folder. Now let's quickly
- 32:11:42understand regression. So let's say we
- 32:11:44have two variables that is temperature
- 32:11:46and humidity where temperature is the
- 32:11:49independent variable and humidity is the
- 32:11:51dependent variable such that as the
- 32:11:54temperature increases humidity decreases
- 32:11:56hence they're correlated. When we feed
- 32:11:58this data to a regression model it will
- 32:12:01understand the relationship between
- 32:12:02these two variables and how one variable
- 32:12:05depends on the other. After the machine
- 32:12:07is trained it can easily predict the
- 32:12:10humidity based on the given temperature.
- 32:12:12Well, that was about regression. Now,
- 32:12:15let's see some real life applications
- 32:12:17where supervised learning is used. So,
- 32:12:19supervised learning is used in risk
- 32:12:21assessment to assess risk in financial
- 32:12:24services or an insurance domain to
- 32:12:26minimize the risk portfolio of the
- 32:12:28companies. Image classification.
- 32:12:30Facebook recognizes your friend in a
- 32:12:32picture from an album of tact photos.
- 32:12:34So, image classification is one of the
- 32:12:36key use cases of demonstrating
- 32:12:38supervised machine learning algorithms.
- 32:12:40Obviously a lot more goes into all these
- 32:12:43like convolutional neural networks etc.
- 32:12:46Fraud detection whether the transactions
- 32:12:48made by the user are authentic or not
- 32:12:50and visual recognition the ability of a
- 32:12:53machine learning model to identify
- 32:12:55objects places people and actions in
- 32:12:58images. Now let's quickly talk about
- 32:13:00unsupervised learning. In unsupervised
- 32:13:02learning, there is no supervision that
- 32:13:04is no training will be given to the
- 32:13:06machine allowing it to act on the data
- 32:13:09which is not labeled. Hence, machine
- 32:13:12tries to identify patterns and gives the
- 32:13:14response. Let's take a similar example
- 32:13:17as before. But this time we do not tell
- 32:13:19the machine whether it's a spoon or a
- 32:13:21knife. The machine identifies patterns
- 32:13:24from the given set and groups them based
- 32:13:27on their patterns, similarities, etc.
- 32:13:30Again unsupervised learning can be
- 32:13:32further grouped into clustering and
- 32:13:34association. Clustering is basically
- 32:13:36where the machine forms groups based on
- 32:13:39the behavior of the data. Secondly,
- 32:13:41association. It is a rule-based machine
- 32:13:44learning to discover interesting
- 32:13:45relation between variables in large data
- 32:13:48sets. For example, which customer made
- 32:13:51similar product purchases is clustering.
- 32:13:54Whereas association is which products
- 32:13:56were purchased together. Now let's
- 32:13:59understand clustering with the help of
- 32:14:00an example. To reduce their churn rate,
- 32:14:03a telecom company studies the behavior
- 32:14:05of the customers based on average call
- 32:14:08duration and internet usage and observes
- 32:14:11that while some customers call duration
- 32:14:13is quite high, others have heavy
- 32:14:15internet usage. The customers are
- 32:14:17grouped based on their observed behavior
- 32:14:19and a strategy is adopted to minimize
- 32:14:22churn rate and maximize profit via
- 32:14:24suitable promotions and campaigns. As
- 32:14:27you can see in the chart on the right
- 32:14:28hand side, customers in group A uses
- 32:14:31more data and also have high
- 32:14:33qualuration. Group B customers are heavy
- 32:14:35internet users while group C customers
- 32:14:38have high qualation. So group B will be
- 32:14:41given more data benefit plans while
- 32:14:43group C will be given cheaper call rates
- 32:14:45to buy their loyalty. So this was the
- 32:14:48example of clustering. Now let's
- 32:14:50understand association with another
- 32:14:52example. Let's say customer one goes to
- 32:14:54a supermarket and buys these products.
- 32:14:56say bread, milk, fruits, wheat. Then
- 32:15:00customer two goes and buys bread, milk,
- 32:15:03rice and butter. Now when customer three
- 32:15:06goes and buys bread, it is highly likely
- 32:15:09that he will also buy milk. Hence
- 32:15:12relationship is established based on
- 32:15:14customer behavior and recommendations
- 32:15:16are made. Now let's look at some real
- 32:15:19life applications of unsupervised
- 32:15:20learning. Market basket analysis is a
- 32:15:23machine learning model based on the
- 32:15:25algorithm that if you buy a certain
- 32:15:26group of items, you are less or more
- 32:15:28likely to buy another group of items.
- 32:15:30Semantic clustering. Semantically
- 32:15:32similar words share similar context.
- 32:15:35People post their queries on websites in
- 32:15:37their own ways. Semantic clustering
- 32:15:39groups all responses in a cluster with
- 32:15:42same meaning to ensure that the customer
- 32:15:44finds the information they want quickly
- 32:15:47and easily. It plays an important role
- 32:15:49in information retrieval. Good browsing
- 32:15:51experience and comprehension. Delivery
- 32:15:54store optimization. Machine learning
- 32:15:56models are used to predict the demand
- 32:15:58and keep up with the supply also to open
- 32:16:00stores where demand is more and
- 32:16:03optimizing routes for more efficient
- 32:16:04deliveries according to past data and
- 32:16:07behavior. We can also use unsupervised
- 32:16:09machine learning models to identify
- 32:16:11accidentprone areas based on the
- 32:16:13intensity of those accidents and the
- 32:16:15area in order to introduce safety
- 32:16:17measures. By now I hope you've
- 32:16:19understood supervised and unsupervised
- 32:16:21learning. For a quick recap, let's see a
- 32:16:24few differences between the two. The
- 32:16:25most fundamental difference is that
- 32:16:27supervised learning uses known and
- 32:16:29labelled data and unsupervised learning
- 32:16:31uses unlabelled data as their input.
- 32:16:33Secondly, supervised learning follows a
- 32:16:36feedback mechanism while unsupervised
- 32:16:38learning does not. Also, the most
- 32:16:40commonly used algorithms in supervised
- 32:16:42learning are decision tree, logistic
- 32:16:44regression, support vector machine, etc.
- 32:16:47And in unsupervised learning there K
- 32:16:50means clustering, hierarchical
- 32:16:51clustering, a priori algorithm and many
- 32:16:54more.
- 32:16:55>> Welcome to linear regression. My name is
- 32:16:57Richard Kersner. I'm with SimplyLearn.
- 32:16:59Let's look at an example of a common use
- 32:17:01for linear regression, profit estimation
- 32:17:04of a company. If I was going to invest
- 32:17:06in a company, I would like to know how
- 32:17:08much money I could expect to make. So
- 32:17:10we'll take a look at a venture
- 32:17:11capitalist firm and try to understand
- 32:17:14which companies they should invest in.
- 32:17:16So we'll take the idea that we need to
- 32:17:18decide the companies to invest in. We
- 32:17:20need to predict the profit the company
- 32:17:22makes and we're going to do it based on
- 32:17:24the company's expenses and even just a
- 32:17:27specific expense. In this case we have
- 32:17:29our company, we have the different
- 32:17:31expenses. So we have our R&D which is
- 32:17:33your research and development. We have
- 32:17:35our marketing. Uh we might have the
- 32:17:37location. We might have what kind of
- 32:17:39administration it's going through. Based
- 32:17:41on all this different information, we
- 32:17:43would like to calculate the profit. Now,
- 32:17:45in actuality, there's usually about 23
- 32:17:48to 27 different markers that they look
- 32:17:50at if they're a heavy duty investor.
- 32:17:52We're only going to take a look at one
- 32:17:54basic one. We're going to come in and
- 32:17:56for simplicity, let's consider a single
- 32:17:58variable, R&D, and find out which
- 32:18:00companies to invest in based on that.
- 32:18:02So, we take our R&D and we're plotting
- 32:18:04the profit based on the R&D expenditure,
- 32:18:06how much money they put into the
- 32:18:07research and development. And then we
- 32:18:09look at the profit that goes with that.
- 32:18:11We can predict a line to estimate the
- 32:18:13profit. So we can draw a line right
- 32:18:15through the data. And when you look at
- 32:18:16that, you can see how much they invest
- 32:18:18in the R&D is a good marker as to how
- 32:18:20much profit they're going to have. We
- 32:18:22can also note that companies spending
- 32:18:23more on R&D make good profit. So let's
- 32:18:26invest in the ones that spend a higher
- 32:18:27rate in their R&D. What's in it for you?
- 32:18:30First, we'll have an introduction to
- 32:18:32machine learning followed by machine
- 32:18:34learning algorithms. These will be
- 32:18:36specific to linear regression and where
- 32:18:38it fits into the larger model. Then
- 32:18:40we'll take a look at applications of
- 32:18:42linear regression, understanding linear
- 32:18:44regression, and multiple linear
- 32:18:46regression. Finally, we'll roll up our
- 32:18:48sleeves and do a little programming in
- 32:18:50use case profit estimation of companies.
- 32:18:53Let's go ahead and jump in. Let's start
- 32:18:55with our introduction to machine
- 32:18:56learning along with some machine
- 32:18:58learning algorithms and where that fits
- 32:19:00in with linear regression. Let's look at
- 32:19:02another example of machine learning.
- 32:19:04Based on the amount of rainfall, how
- 32:19:05much would be the crop yield? So we here
- 32:19:08we have our crops, we have our rainfall,
- 32:19:10and we want to know how much we're going
- 32:19:12to get from our crops this year. So
- 32:19:14we're going to introduce two variables,
- 32:19:16independent and dependent. The
- 32:19:18independent variable is a variable whose
- 32:19:20value does not change by the effect of
- 32:19:22other variables and is used to
- 32:19:24manipulate the dependent variable. It is
- 32:19:26often denoted as X. In our example,
- 32:19:28rainfall is the independent variable.
- 32:19:30This is a wonderful example because you
- 32:19:32can easily see that we can't control the
- 32:19:34rain, but the rain does control the
- 32:19:36crop. So we talk about the independent
- 32:19:38variable controlling the dependent
- 32:19:40variable. Let's define dependent
- 32:19:42variable as a variable whose value
- 32:19:43change when there is any manipulation in
- 32:19:46the values of the independent variables.
- 32:19:48It is often denoted as y. And you can
- 32:19:50see here our crop yield is dependent
- 32:19:52variable and it is dependent on the
- 32:19:54amount of rainfall received. Now that
- 32:19:56we've taken a look at a real life
- 32:19:57example, let's go a little bit into the
- 32:19:59theory and some definitions on machine
- 32:20:02learning and see how that fits together
- 32:20:04with linear regression. numerical and
- 32:20:06categorical values. Let's take our data
- 32:20:09coming in and this is kind of random
- 32:20:11data from any kind of project. We want
- 32:20:14to divide it up into numerical and
- 32:20:16categorical. So numerical is numbers,
- 32:20:19age, salary, height, where categorical
- 32:20:23would be a description, the color, a
- 32:20:26dog's breed, gender. Categorical is
- 32:20:29limited to very specific items where
- 32:20:30numerical is a range of information. Now
- 32:20:34that you've seen the difference between
- 32:20:36numerical and categorical data, let's
- 32:20:38take a look at some different machine
- 32:20:40learning definitions. When we look at
- 32:20:42our different machine learning
- 32:20:43algorithms, we can divide them into
- 32:20:46three areas. Supervised, unsupervised,
- 32:20:50reinforcement. We're only going to look
- 32:20:51at supervised today. Unsupervised means
- 32:20:54we don't have the answers and we're just
- 32:20:56grouping things. Reinforcement is where
- 32:20:58we give positive and negative feedback
- 32:21:01to our algorithm to program it. and it
- 32:21:03doesn't have the information till after
- 32:21:04the fact. But today we're just looking
- 32:21:06at supervised because that's where
- 32:21:07linear regression fits in. In supervised
- 32:21:10data, we have our data already there and
- 32:21:12our answers for a group. And then we use
- 32:21:14that to program our model and come up
- 32:21:17with an answer. The two most common uses
- 32:21:19for that is through the regression and
- 32:21:21classification. Now, we're doing linear
- 32:21:23regression. So, we're just going to
- 32:21:24focus on the regression side. And in the
- 32:21:27regression we have simple linear
- 32:21:29regression, we have multiple linear
- 32:21:31regression and we have polomial linear
- 32:21:34regression. Now on these three simple
- 32:21:36linear regression is the examples we've
- 32:21:38looked at so far where we have a lot of
- 32:21:40data and we draw a straight line through
- 32:21:42it. Multiple linear regression means we
- 32:21:44have multiple variables. Remember where
- 32:21:47we had the rainfall and the crops. We
- 32:21:49might add additional variables in there
- 32:21:51like how much food do we give our crops?
- 32:21:53When do we harvest them? Those would be
- 32:21:55additional information add into our
- 32:21:57model and that's why it' be multiple
- 32:21:59linear regression. And finally we have
- 32:22:00polomial linear regression that is
- 32:22:02instead of drawing a line we can draw a
- 32:22:04curved line through it. Now that you see
- 32:22:06where regression model fits into the
- 32:22:08machine learning algorithms and we're
- 32:22:11specifically looking at linear
- 32:22:12regression. Let's go ahead and take a
- 32:22:14look at applications for linear
- 32:22:16regression. Let's look at a few
- 32:22:18applications of linear regression.
- 32:22:21Economic growth used to determine the
- 32:22:23economic growth of a country or a state
- 32:22:26in the coming quarter can also be used
- 32:22:28to predict the GDP of a country. Product
- 32:22:31price can be used to predict what would
- 32:22:33be the price of a product in the future.
- 32:22:34We can guess whether it's going to go up
- 32:22:36or down or should I buy today. Housing
- 32:22:38sales to estimate the number of houses a
- 32:22:40builder would sell and what price in the
- 32:22:42coming months. Score predictions.
- 32:22:44Cricket fever to predict the number of
- 32:22:46runs a player would score in the coming
- 32:22:48matches based on the previous
- 32:22:50performance. I'm sure you can figure out
- 32:22:52other applications you could use linear
- 32:22:54regression for. So let's jump in and
- 32:22:57let's understand linear regression and
- 32:22:59dig into the theory. Understanding
- 32:23:01linear regression. Linear regression is
- 32:23:04the statistical model used to predict
- 32:23:07the relationship between independent and
- 32:23:09dependent variables by examining two
- 32:23:11factors. The first important one is
- 32:23:14which variables in particular are
- 32:23:16significant predictors of the outcome
- 32:23:18variable. And the second one that we
- 32:23:20need to look at closely is how
- 32:23:22significant is the regression line to
- 32:23:24make predictions with the highest
- 32:23:25possible accuracy. If it's inaccurate,
- 32:23:28we can't use it. So, it's very important
- 32:23:30we find out the most accurate line we
- 32:23:32can get. Since linear regression is
- 32:23:35based on drawing a line through data,
- 32:23:37we're going to jump back and take a look
- 32:23:39at some uklitian geometry. The simplest
- 32:23:41form of a simple linear regression
- 32:23:43equation with one dependent and one
- 32:23:45independent variable is represented by y
- 32:23:49= m * x + c. And if you look at our
- 32:23:52model here, we plotted two points on
- 32:23:54here. Uh x1 and y1, x2 and y2. y being
- 32:24:00the dependent variable, remember that
- 32:24:02from before. And x being the independent
- 32:24:05variable. So y depends on whatever x is.
- 32:24:08M in this case is the slope of the line
- 32:24:11where M equals the difference in the Y2
- 32:24:15- Y1 and X2 - X1. And finally we have C
- 32:24:19which is the coefficient of the line or
- 32:24:22where happens to cross the zero axis.
- 32:24:25Let's go back and look at an example we
- 32:24:27used earlier of linear regression. We're
- 32:24:29going to go back to plotting the amount
- 32:24:31of crop yield based on the amount of
- 32:24:33rainfall. And here we have our rainfall.
- 32:24:36Remember, we cannot change rainfall. And
- 32:24:39we have our crop yield, which is
- 32:24:41dependent on the rainfall. So, we have
- 32:24:43our independent and our dependent
- 32:24:44variables. We're going to take this and
- 32:24:46draw a line through it as best we can
- 32:24:49through the middle of the data. And then
- 32:24:50we look at that. We put the red point on
- 32:24:52the y ais is the amount of crop yield
- 32:24:55you can expect for the amount of
- 32:24:56rainfall represented by the green dot.
- 32:24:59So, if we have an idea what the rainfall
- 32:25:00is for this year and what's going on,
- 32:25:03then we can guess how good our crops are
- 32:25:05going to be. and we've created a nice
- 32:25:06line right through the middle to give us
- 32:25:08a nice mathematical formula. Let's take
- 32:25:10a look and see what the math looks like
- 32:25:12behind this. Let's look at the intuition
- 32:25:15behind the regression line. Now, before
- 32:25:17we dive into the math and the formulas
- 32:25:20that go behind this and what's going on
- 32:25:22behind the scenes, I want you to note
- 32:25:25that when we get into the case study and
- 32:25:28we actually apply some Python script
- 32:25:30that this math that you're going to see
- 32:25:31here is already done automatically for
- 32:25:33you. You don't have to have it
- 32:25:34memorized. It is, however, good to have
- 32:25:37an idea what's going on so if people
- 32:25:40reference the different terms, you'll
- 32:25:41know what they're talking about. Let's
- 32:25:43consider a sample data set with five
- 32:25:46rows and find out how to draw the
- 32:25:48regression line. We're only going to do
- 32:25:49five rows because if we did like the
- 32:25:51rainfall with hundreds of points of
- 32:25:53data, that would be very hard to see
- 32:25:55what's going on with the mathematics.
- 32:25:57So, we'll go ahead and create our own
- 32:25:58two sets of data. And we have our
- 32:26:01independent variable x and our dependent
- 32:26:03variable y. And when x was 1, we got y =
- 32:26:072. When x was uh 2, y was 4. And so on
- 32:26:12and so on. If we go ahead and plot this
- 32:26:14data on a graph, we can see how it forms
- 32:26:17a nice line through the middle. You can
- 32:26:19see where it's kind of grouped going
- 32:26:21upwards to the right. The next thing we
- 32:26:23want to know is what the means is of
- 32:26:25each of the data coming in, the x and
- 32:26:28the y. The means doesn't mean anything
- 32:26:30other than the average. So, we add up
- 32:26:32all the numbers and divide by the total.
- 32:26:34So, 1 + 2 + 3 + 4 + 5 over 5 equals 3.
- 32:26:39And the same for y, we get four. If we
- 32:26:41go ahead and plot the means on the
- 32:26:43graph, we'll see we get 3a 4, which
- 32:26:46draws a nice line down the middle, a
- 32:26:48good estimate. Here, we're going to dig
- 32:26:50deeper into the math behind the
- 32:26:52regression line. Now, remember before I
- 32:26:54said you don't have to have all these
- 32:26:56formulas memorized or fully understand
- 32:26:59them, even though we're going to go into
- 32:27:00a little more detail of how it works.
- 32:27:02And if you're not a math wiz and you
- 32:27:04don't know if you've never seen the
- 32:27:05sigma character before, which looks a
- 32:27:08little bit like an e that's opened up,
- 32:27:10that just means summation. That's all
- 32:27:12that is. So, when you see the sigma
- 32:27:14character, it just means we're adding
- 32:27:15everything in that row. And for
- 32:27:17computers, this is great because as a
- 32:27:19programmer, you can easily iterate
- 32:27:21through each of the XY points and create
- 32:27:24all the information you need. So in the
- 32:27:26top half, you can see where we've broken
- 32:27:27that down into pieces. And as it goes
- 32:27:29through the first two points, it
- 32:27:31computes the squared value of X, the
- 32:27:33squared value of Y, and X * Y. And then
- 32:27:36it takes all of X and adds them up. All
- 32:27:38of Y adds them up. All of X squar adds
- 32:27:40them up. And so on and so on. And you
- 32:27:42can see we have the sum of equal to 15.
- 32:27:45The sum is equal to 20. all the way up
- 32:27:47to x * y where the sum equals 66. This
- 32:27:50all comes from our formula for
- 32:27:52calculating a straight line where y
- 32:27:54equals the slope* x plus the coefficient
- 32:27:57c. So we go down below and we're going
- 32:28:00to compute more like the averages of
- 32:28:02these and we're going to explain exactly
- 32:28:04what that is in just a minute and where
- 32:28:06that information comes from. It's called
- 32:28:07the square means error, but we'll go
- 32:28:09into that in detail in a few minutes.
- 32:28:11All you need to do is look at the
- 32:28:12formula and see how we've gone about
- 32:28:14computing it line by line instead of
- 32:28:17trying to have a huge set of numbers
- 32:28:20pushed into it. And down here you'll see
- 32:28:22where the slope m equals and then the
- 32:28:25top part if you read through the
- 32:28:27brackets you have the number of data
- 32:28:29points times the sum of x * y which we
- 32:28:34computed one line at a time there. And
- 32:28:36that's just the 66. and take all that
- 32:28:38and you subtract it from the sum of x
- 32:28:41times the sum of y and those have both
- 32:28:43been computed. So you have 15 * 20. And
- 32:28:45on the bottom we have the number of
- 32:28:48lines times the sum of x^2 easily
- 32:28:51computed as 86 for the sum minus I'll
- 32:28:54take all that and subtract the sum of
- 32:28:57x^2. And we end up as we come across
- 32:29:00with our formula. You can plug in all
- 32:29:02those numbers which is very easy to do
- 32:29:03on the computer. You don't have to do
- 32:29:05the math on a piece of paper or
- 32:29:07calculator. And you'll get a slope of 6
- 32:29:09and you'll get your C coefficient. If
- 32:29:11you continue to follow through that
- 32:29:12formula, you'll see it comes out as
- 32:29:14equal to 2.2. Continuing deeper into
- 32:29:17what's going behind the scenes, let's
- 32:29:19find out the predicted values of y for
- 32:29:21corresponding values of x using the
- 32:29:23linear equation where m=6 and c = 2.2.
- 32:29:28We're going to take these values and
- 32:29:30we're going to go ahead and plot them.
- 32:29:32We're going to predict them. So y =6 * x
- 32:29:35= 1 + 2.2 = 2.8 so on and so on. And
- 32:29:39here the blue points represent the
- 32:29:41actual y values and the brown points
- 32:29:44represent the predicted yv values based
- 32:29:46on the model we created. The distance
- 32:29:48between the actual and predicted values
- 32:29:50is known as residuals or errors. The
- 32:29:54best fit line should have the least sum
- 32:29:56of squares of these errors also known as
- 32:29:59equare. If we put these into a nice
- 32:30:01chart where you can see X and you can
- 32:30:04see Y what the actual values were and
- 32:30:06you can see Y predicted you can easily
- 32:30:08see where we take Y minus Y predicted
- 32:30:11and we get an answer. What is the
- 32:30:12difference between those two and if we
- 32:30:14square that Y - Y prediction squared we
- 32:30:18can then sum those squared values.
- 32:30:20That's where we get the 64 plus the.36 +
- 32:30:231 all the way down until we have a
- 32:30:25summation equals 2.4. So the sum of
- 32:30:28squared errors for this regression line
- 32:30:30is 2.4. We check this error for each
- 32:30:32line and conclude the best fit line
- 32:30:34having the least e value. In a nice
- 32:30:37graphical representation, we can see
- 32:30:39here where we keep moving this line
- 32:30:41through the data points to make sure the
- 32:30:43best fit line has the least squared
- 32:30:45distance between the data points and the
- 32:30:47regression line. Now we only looked at
- 32:30:50the most commonly used formula for
- 32:30:52minimizing the distance. There are lots
- 32:30:55of ways to minimize the distance between
- 32:30:57the line and the data points like sum of
- 32:30:59squared errors, sum of absolute errors,
- 32:31:02root mean square error, etc. What you
- 32:31:04want to take away from this is whatever
- 32:31:06formula is being used, you can easily
- 32:31:09using a computer programming and
- 32:31:11iterating through the data calculate the
- 32:31:13different parts of it. That way, these
- 32:31:16complicated formulas you see with the
- 32:31:17different summations and absolute values
- 32:31:20are easily computed one piece at a time.
- 32:31:22Up until this point, we've only been
- 32:31:24looking at two values, X and Y. Well, in
- 32:31:28the real world, it's very rare that you
- 32:31:29only have two values when you're
- 32:31:31figuring out a solution. So, let's move
- 32:31:33on to the next topic, multiple linear
- 32:31:36regression. Let's take a brief look at
- 32:31:38what happens when you have multiple
- 32:31:39inputs. So, in multiple linear
- 32:31:41regression, we have uh well, we'll start
- 32:31:43with the simple linear regression where
- 32:31:45we had y = m + x + c and we're trying to
- 32:31:49find the value of y. Now with multiple
- 32:31:51linear regression we have multiple
- 32:31:53variables coming in. So instead of
- 32:31:55having just x we have x1 x2 x3 and
- 32:32:00instead of having just one slope each
- 32:32:02variable has its own slope attached to
- 32:32:04it. As you can see here we have m1 m2 m3
- 32:32:08and we still just have the single
- 32:32:09coefficient. So when you're dealing with
- 32:32:11multiple linear regression you basically
- 32:32:13take your single linear regression and
- 32:32:15you spread it out. So you have y = m1 *
- 32:32:18x1 + m2 * x2 so on all the way to m x to
- 32:32:24the nth and then you add your
- 32:32:25coefficient on there. Implementation of
- 32:32:28linear regression. Now we get into my
- 32:32:30favorite part. Let's understand how
- 32:32:32multiple linear regression works by
- 32:32:34implementing it in Python. If you
- 32:32:36remember before we were looking at a
- 32:32:38company and just based on its R&D trying
- 32:32:41to figure out its profit. We're going to
- 32:32:43start looking at the expenditure of the
- 32:32:44company. We're going to go back to that.
- 32:32:46We're going to predict his profit, but
- 32:32:48instead of predicting it just on the
- 32:32:50R&D, we're going to look at other
- 32:32:51factors like administration costs,
- 32:32:54marketing costs, and so on. And from
- 32:32:56there, we're going to see if we can
- 32:32:58figure out what the profit of that
- 32:32:59company is going to be. To start our
- 32:33:01coding, we're going to begin by
- 32:33:03importing some basic libraries. And
- 32:33:05we're going to be looking through the
- 32:33:06data before we do any kind of linear
- 32:33:08regression. We're going to take a look
- 32:33:09at the data to see what we're playing
- 32:33:10with. Then we'll go ahead and format the
- 32:33:13data to the format we need to be able to
- 32:33:15run it in the linear regression model.
- 32:33:17And then from there we'll go ahead and
- 32:33:18solve it and just see how valid our
- 32:33:21solution is. So let's start with
- 32:33:22importing the basic libraries. Now I'm
- 32:33:24going to be doing this in Anaconda
- 32:33:27Jupyter notebook, a very popular IDE. I
- 32:33:30enjoy it because it's such a visual to
- 32:33:32look at and so easy to use. Um just any
- 32:33:34ID for Python will work just fine for
- 32:33:36this. So break out your favorite Python
- 32:33:38IDE. So, here we are in our Jupyter
- 32:33:41notebook. Let me go ahead and paste our
- 32:33:42first piece of code in there. And let's
- 32:33:44walk through what libraries we're
- 32:33:46importing. First, we're going to import
- 32:33:48numpy as np. And then I want you to skip
- 32:33:51one line and look at import pandas as
- 32:33:53pd. These are very common tools that you
- 32:33:55need with most of your linear
- 32:33:56regression. The numpy, which stands for
- 32:33:59number python, is usually denoted as np,
- 32:34:01and you have to almost have that for
- 32:34:03your sklearn toolbox. So, you always
- 32:34:05import that right off the beginning.
- 32:34:06pandas. Although you don't have to have
- 32:34:08it for your sklearn libraries, it does
- 32:34:10such a wonderful job of importing data,
- 32:34:12setting it up into a data frame so we
- 32:34:14can manipulate it rather easily and it
- 32:34:16has a lot of tools also in addition to
- 32:34:18that. So we usually like to use the
- 32:34:20pandas when we can and I'll show you
- 32:34:21what that looks like. The other three
- 32:34:23lines are for us to get a visual of this
- 32:34:26data and take a look at it. So we're
- 32:34:28going to import mapplot library.pipplot
- 32:34:30as plt and then seabour as sns. Seabor
- 32:34:35works with the mattplot library. So you
- 32:34:37have to always import mapplot library
- 32:34:39and then seabour sits on top of it. And
- 32:34:41we'll take a look at what that looks
- 32:34:42like. You could use any of your own
- 32:34:44plotting libraries you want. There's all
- 32:34:46kinds of ways to look at the data. These
- 32:34:48are just very common ones. And the
- 32:34:49seabor is so easy to use. It just looks
- 32:34:52beautiful. It's a nice representation
- 32:34:53that you can actually take and show
- 32:34:55somebody. And the final line is the
- 32:34:57amberigned mattplot library inline. That
- 32:35:01is only because I'm doing an inline IDE.
- 32:35:03My interface in the Anaconda Jupiter
- 32:35:05notebook requires I put that in there or
- 32:35:08you're not going to see the graph when
- 32:35:10it comes up. Let's go ahead and run
- 32:35:11this. It's not going to be that
- 32:35:12interesting because we're just setting
- 32:35:13up variables. In fact, it's not going to
- 32:35:15do anything that we can see, but it is
- 32:35:17importing these different libraries and
- 32:35:19setup. The next step is load the data
- 32:35:22set and extract independent and
- 32:35:24dependent variables. Now, here in the
- 32:35:27slide, you'll see companies equals PD
- 32:35:29read CSV. And it has a long line there
- 32:35:32with the file at the end. 10,00
- 32:35:34companies.csv. You're going to have to
- 32:35:36change this to fit whatever setup you
- 32:35:38have. And the file itself, you can
- 32:35:41request. Just go down to the commentary
- 32:35:43below this video and put a note in there
- 32:35:45and SimplyLearn will try to get in
- 32:35:47contact with you and supply you with
- 32:35:48that file so you can try this coding
- 32:35:50yourself. So, we're going to add this
- 32:35:52code in here. And we're going to see
- 32:35:53that I have companies equals
- 32:35:55PD.reader_csv.
- 32:35:57And I've changed this path to match my
- 32:35:59computer. C/simplylearn/1000
- 32:36:03companies.csv. And then below there,
- 32:36:05we're going to set the x equals to
- 32:36:07companies under the i location. And
- 32:36:09because this is companies is a pd data
- 32:36:12set, I can use this nice notation that
- 32:36:14says take every row, that's what the
- 32:36:16colon the first colon is, comma, except
- 32:36:19for the last column. That's what the
- 32:36:20second part is where we have a colon
- 32:36:22minus one and we want the values set
- 32:36:24into there. So x is no longer a data set
- 32:36:27a pandas data set but we can easily
- 32:36:29extract the data from our pandas data
- 32:36:31set with this notation and then y we're
- 32:36:33going to set equal to the last row. Well
- 32:36:36the question is going to be what are we
- 32:36:37actually looking at? So let's go ahead
- 32:36:39and take a look at that and we're going
- 32:36:41to look at the companies.
- 32:36:43Which lists the first five rows of data
- 32:36:45and I'll open up the file in just a
- 32:36:47second so you can see where that's
- 32:36:48coming from. But let's look at the data
- 32:36:50in here as far as the way the pandas
- 32:36:52sees it. When I hit run, you'll see it
- 32:36:54breaks it out into a nice setup. This is
- 32:36:56what pandas, one of the things pandas is
- 32:36:58really good about is it looks just like
- 32:36:59an Excel spreadsheet. You have your rows
- 32:37:01and remember when we're programming, we
- 32:37:03always start with zero. We don't start
- 32:37:05with one. So it shows the first five
- 32:37:08rows 0 1 2 3 4 and then it shows your
- 32:37:11different columns. R&D spend,
- 32:37:14administration, marketing spend, state,
- 32:37:17profit. It even notes that the top are
- 32:37:19column names. It was never told that,
- 32:37:21but Pandas is able to recognize a lot of
- 32:37:23things that they're not the same as the
- 32:37:25data rows. Why don't we go ahead and
- 32:37:26open this file up in a CSV so you can
- 32:37:29actually see the raw data. So here I've
- 32:37:31opened it up as a text editor. And you
- 32:37:33can see at the top we have R&D spend,
- 32:37:36administration, marketing spend, state,
- 32:37:39profit, carriage return. I don't know
- 32:37:41about you, but I'd go crazy trying to
- 32:37:42read files like this. That's why we use
- 32:37:45the pandas. You could also open this up
- 32:37:47in an Excel and it would separate it
- 32:37:48since it is a comma separated variable
- 32:37:50file. But we don't want to look at this
- 32:37:51one. We want to look at something we can
- 32:37:53read rather easily. So let's flip back
- 32:37:55and take a look at that top part, the
- 32:37:57first five row. Now, as nice as this
- 32:37:59format is where I can see the data, to
- 32:38:01me it doesn't mean a whole lot. Maybe
- 32:38:03you're an expert in business and
- 32:38:05investments and you understand what
- 32:38:08$165,34920
- 32:38:11compared to the administration cost of
- 32:38:14$136,897.80
- 32:38:17so on so on helps to create the profit
- 32:38:19of $192,26183.
- 32:38:23That makes no sense to me whatsoever. No
- 32:38:26pun intended. So let's flip back here
- 32:38:27and take a look at our next set of code
- 32:38:29where we're going to graph it so we can
- 32:38:30get a better understanding of our data
- 32:38:32and what it mean. So at this point we're
- 32:38:35going to use a single line of code to
- 32:38:37get a lot of information so we can see
- 32:38:39where we're going with this. Let's go
- 32:38:41ahead and paste that into our uh
- 32:38:43notebook and see what we got going. And
- 32:38:45so we have the visualization and again
- 32:38:47we're using SNS which is pandas. As you
- 32:38:50can see, we imported the mapplot
- 32:38:51library.pipplot as plt, which then the
- 32:38:55seabor uses and we imported the seabour
- 32:38:57as sns. And then that final line of code
- 32:39:00helps us show this in our um inline
- 32:39:03coding. Without this, it wouldn't
- 32:39:05display and you could display it to a
- 32:39:06file and other means. And that's the map
- 32:39:08plot library in line with the amber sign
- 32:39:10at the beginning. So here we come down
- 32:39:12to the single line of code. Seabor is
- 32:39:14great because it actually recognizes the
- 32:39:16panda data frame. So I can just take the
- 32:39:18companies.core
- 32:39:20for coordinates and I can put that right
- 32:39:22into the seaborn. And when we run this,
- 32:39:25we get this beautiful plot. And let's
- 32:39:27just take a look at what this plot
- 32:39:28means. If you look at this plot on mine,
- 32:39:31the colors are probably a little bit
- 32:39:32more purplish and blue than the original
- 32:39:34one. Uh we have the columns and the
- 32:39:36rows. We have R and D spending. We have
- 32:39:38administration. We have marketing
- 32:39:40spending and profit. And if you cross
- 32:39:42index any two of these, since we're
- 32:39:44interested in profit, if you cross-index
- 32:39:46profit with profit, it's going to show
- 32:39:48up, if you look at the scale on the
- 32:39:49right, way up in the dark. Why? Because
- 32:39:52those are the same data. They have an
- 32:39:54exact correspondence. So R&D spending is
- 32:39:57going to be the same as R&D spending.
- 32:40:00And the same thing with administration
- 32:40:01costs. But right down the middle, you
- 32:40:03get this dark row or dark um diagonal
- 32:40:06row that shows that this is the highest
- 32:40:09corresponding data. That's exactly the
- 32:40:11same. And as it becomes lighter, there's
- 32:40:13less connections between the data. So we
- 32:40:16can see with profit, obviously profit is
- 32:40:18the same as profit. And next, it has a
- 32:40:20very high correlation with R&D spending,
- 32:40:23which we looked at earlier. And it has a
- 32:40:25slightly less connection to marketing
- 32:40:27spending and even less to how much money
- 32:40:29we put into the administration. So now
- 32:40:31that we have a nice look at the data,
- 32:40:33let's go ahead and dig in and create
- 32:40:35some actual useful linear regression
- 32:40:37models so that we can predict values and
- 32:40:39have a better profit. Now that we've
- 32:40:42taken a look at the visualization of
- 32:40:43this data, we're going to move on to the
- 32:40:45next step. Instead of just having a
- 32:40:47pretty picture, we need to generate some
- 32:40:49hard data, some hard values. So let's
- 32:40:52see what that looks like. We're going to
- 32:40:54set up our linear regression model in
- 32:40:56two steps. The first one is we need to
- 32:40:58prepare some of our data so it fits
- 32:41:01correctly. And let's go ahead and paste
- 32:41:02this code into our Jupyter notebook. And
- 32:41:04what we're bringing in is we're going to
- 32:41:05bring in the sklearn pre-processing
- 32:41:08where we're going to import the label
- 32:41:10encoder and the one hot encoder. To use
- 32:41:13the label encoder, we're going to create
- 32:41:14a variable called label encoder and set
- 32:41:16it equal to capital L label capital E
- 32:41:19encoder. This creates a class that we
- 32:41:21can reuse for transferring the labels
- 32:41:24back and forth. Now about now you should
- 32:41:25ask what labels are we talking about.
- 32:41:27Let's go take a look at the data we
- 32:41:29processed before and see what I'm
- 32:41:31talking about here. If you remember when
- 32:41:32we did the companies.head and we printed
- 32:41:34the top five rows of data. We have our
- 32:41:37columns going across. We have column
- 32:41:39zero which is R&D spending, column one
- 32:41:42which is administration, column two
- 32:41:44which is marketing spending and column
- 32:41:46three is state. And you'll see under
- 32:41:49state we have New York, California,
- 32:41:51Florida. Now to do a linear regression
- 32:41:53model, it doesn't know how to process
- 32:41:55New York. It knows how to process a
- 32:41:57number. So the first thing we're going
- 32:41:58to do is we're going to change that New
- 32:42:00York, California, and Florida. And we're
- 32:42:02going to change those to numbers. That's
- 32:42:04what this line of code does here. X
- 32:42:06equals and then it has the colon, 3 in
- 32:42:09brackets. The first part, the colon,
- 32:42:11comma, means that we're going to look at
- 32:42:13all the different rows. So we're going
- 32:42:15to keep them all together. But the only
- 32:42:17row we're going to edit is the third
- 32:42:18row. And in there, we're going to take
- 32:42:20the label coder and we're going to fit
- 32:42:22and transform the x also the third row.
- 32:42:25So, we're going to take that third row,
- 32:42:26we're going to set it equal to a
- 32:42:28transformation. And that transformation
- 32:42:29basically tells it that instead of
- 32:42:31having a uh New York, it has a zero or a
- 32:42:34one or a two. And then finally, we need
- 32:42:36to do a one hot encoder, which equals
- 32:42:39one hot encoder categorical features
- 32:42:42equals three. And then we take the X and
- 32:42:44we go ahead and do that equal to one hot
- 32:42:46encoder fit transform X to array. This
- 32:42:49final transformation preps our data for
- 32:42:52us. So it's completely set the way we
- 32:42:54need it as just a row of numbers. Even
- 32:42:56though it's not in here, let's go ahead
- 32:42:57and print X and just take a look what
- 32:43:00this data is doing. You'll see I have an
- 32:43:01array of arrays and then each array is a
- 32:43:04row of numbers. And if I go ahead and
- 32:43:06just do row zero, you'll see I have a
- 32:43:08nice organized row of numbers that the
- 32:43:10computer now understands. We'll go ahead
- 32:43:12and take this out there because it
- 32:43:14doesn't mean a whole lot to us. It's
- 32:43:15just a row of numbers. Next on setting
- 32:43:18up our data, we have avoiding dummy
- 32:43:21variable trap. This is very important.
- 32:43:24Why? Because the computer's
- 32:43:26automatically transformed our header
- 32:43:28into the setup and it's automatically
- 32:43:30transformed all these different
- 32:43:31variables. So when we did the encoder,
- 32:43:34the encoder created two columns. And
- 32:43:37what we need to do is just have the one
- 32:43:39because it has both the variable and the
- 32:43:41name. That's what this piece of code
- 32:43:43does here. Let's go ahead and paste this
- 32:43:45in here. And we have x= x colon, one
- 32:43:49colon. All this is doing is removing
- 32:43:51that one extra column we put in there
- 32:43:54when we did our one hot encoder and our
- 32:43:55label encoding. Let's go ahead and run
- 32:43:57that. And now we get to create our
- 32:43:59linear regression model. And let's see
- 32:44:01what that looks like here. And we're
- 32:44:03going to do that in two steps. The first
- 32:44:06step is going to be in splitting the
- 32:44:08data. Now, whenever we create a uh
- 32:44:11predictive model of data, we always want
- 32:44:14to split it up. So, we have a training
- 32:44:15set and we have a testing set. That's
- 32:44:17very important. Otherwise, we'd be very
- 32:44:19unethical without testing it to see how
- 32:44:21good our fit is. And then we'll go ahead
- 32:44:23and create our multiple linear
- 32:44:25regression model and train it and set it
- 32:44:27up. Let's go ahead and paste this next
- 32:44:29piece of code in here. And I'll go ahead
- 32:44:30and shrink it down a size or two so it
- 32:44:32all fits on one line. So from the
- 32:44:34sklearn module selection, we're going to
- 32:44:37import train test split. And you'll see
- 32:44:40that we've created four completely
- 32:44:42different variables. We have capital X
- 32:44:44train capital X test smallercase Y train
- 32:44:48smallerase Y test. That is the standard
- 32:44:51way that they usually reference these
- 32:44:53when we're doing different uh models.
- 32:44:56usually see that a capital X and you see
- 32:44:58the train and the test and the lowercase
- 32:45:00Y. What this is is X is our data going
- 32:45:02in. That's our R&D spin, our
- 32:45:04administration, our marketing. And then
- 32:45:07Y, which we're training, is the answer.
- 32:45:09That's the profit because we want to
- 32:45:10know the profit of an unknown entity. So
- 32:45:12that's what we're going to shoot for in
- 32:45:14this tutorial. The next part, train,
- 32:45:16test, split. We take X and we take Y.
- 32:45:20We've already created those. X has the
- 32:45:22columns with the data in it and Y has a
- 32:45:25column with profit in it. And then we're
- 32:45:26going to set the test size equals 0.2.
- 32:45:30That basically means 20%. So 20% of the
- 32:45:33rows are going to be tested. We're going
- 32:45:34to put them off to the side. So since
- 32:45:36we're using a thousand lines of data,
- 32:45:38that means that 200 of those lines we're
- 32:45:41going to hold off to the side to test
- 32:45:42for later. And then the random state
- 32:45:44equals zero. We're going to randomize
- 32:45:46which ones it picks to hold off to the
- 32:45:48side. We'll go ahead and run this. It's
- 32:45:50not overly exciting because it's setting
- 32:45:51up our variables. But the next step is
- 32:45:54the next step we actually create our
- 32:45:55linear regression model. Now that we got
- 32:45:57to the linear regression model, we get
- 32:45:59that next piece of the puzzle. Let's go
- 32:46:01ahead and put that code in there and
- 32:46:02walk through it. So here we go. We're
- 32:46:04going to paste it in there. And let's go
- 32:46:05ahead and uh since this is a shorter
- 32:46:07line of code, let's zoom up there so we
- 32:46:09can get a good look. And we have from
- 32:46:10the sklearn.linear_model,
- 32:46:13we're going to import linear regression.
- 32:46:15Now, I don't know if you recall from
- 32:46:17earlier when we were doing all the math.
- 32:46:19Let's go ahead and flip back there and
- 32:46:20take a look at that. Do you remember
- 32:46:22this where we had this long formula on
- 32:46:24the bottom and we were doing all this
- 32:46:25summization and then we also looked at
- 32:46:28setting it up with the different lines
- 32:46:30and then we also looked all the way down
- 32:46:32to multiple linear regression where
- 32:46:34we're adding all those formulas
- 32:46:36together. All of that is wrapped up in
- 32:46:38this one section. So what's going on
- 32:46:40here is I'm going to create a variable
- 32:46:42called regressor. And the regressor
- 32:46:44equals the linear regression. That's a
- 32:46:46linear regression model that has all
- 32:46:47that math built in. So we don't have to
- 32:46:50have it all memorized or have to compute
- 32:46:51it individually. And then we do the
- 32:46:53regressor.fit.
- 32:46:54In this case, we do xrain and y train
- 32:46:57because we're using the training data. X
- 32:46:59being the data in and y being profit
- 32:47:02what we're looking at. And this does all
- 32:47:03that math for us. So within one click
- 32:47:06and one line, we've created the whole
- 32:47:07linear regression model and we fit the
- 32:47:10data to the linear regression model. And
- 32:47:12you can see that when I run the
- 32:47:13regressor, it gives an output linear
- 32:47:15regression. It says copy X equals true,
- 32:47:18fit intercept equals true, in jobs equal
- 32:47:201, normalize equals false. It's just
- 32:47:22giving you some general information on
- 32:47:24what's going on with that regressor
- 32:47:26model. Now that we've created our linear
- 32:47:28regression model, let's go ahead and use
- 32:47:30it. And if you remember, we kept a bunch
- 32:47:32of data aside. So, we're going to do a Y
- 32:47:35predict variable and we're going to put
- 32:47:37in the X test. And let's see what that
- 32:47:39looks like. Scroll up a little bit.
- 32:47:41Paste that in here. Predicting the test
- 32:47:44set results. So, here we have Y predict
- 32:47:47equals regressor.predict
- 32:47:49X test going in. And this gives us Y
- 32:47:52predict. Now, because I'm in Jupiter in
- 32:47:54line, I can just put the variable up
- 32:47:56there. And when I hit the run button,
- 32:47:58it'll print that array out. I could have
- 32:48:00just as easily done print y predict. So
- 32:48:03if you're in a different IDE that's not
- 32:48:05an inline setup like the Jupyter
- 32:48:07notebook, you can do it this way. Print
- 32:48:09y predict. And you'll see that for the
- 32:48:11200 different test variables we kept off
- 32:48:13to the side, it's going to produce 200
- 32:48:16answers. This is what it says the profit
- 32:48:18are for those 200 predictions. But let's
- 32:48:20don't stop there. Let's keep going and
- 32:48:22take a couple look. We're going to take
- 32:48:24just a short detail here and calculating
- 32:48:26the coefficients and the intercepts.
- 32:48:28This gives us a quick flash at what's
- 32:48:30going on behind the line. We're going to
- 32:48:32take a short detour here and we're going
- 32:48:34to be calculating the coefficient and
- 32:48:36intercepts. So you can see what those
- 32:48:38look like. What's really nice about our
- 32:48:40regressor we created is it already has
- 32:48:42the coefficients for us. And we can
- 32:48:44simply just print regressor.coefficient
- 32:48:47underscore. When I run this, you'll see
- 32:48:49our coefficients here. And if we can do
- 32:48:51the regressor coefficient, we can also
- 32:48:54do the regressor intercept. And let's
- 32:48:57run that and take a look at that. This
- 32:48:58all came from the multiple regression
- 32:49:00model. And we'll flip over so you can
- 32:49:02remember where this is going into where
- 32:49:04it's coming from. You can see the
- 32:49:05formula down here where y = m1 * x1 + m2
- 32:49:10* x2 and so on and so on plus c the
- 32:49:12coefficient. So these variables fit
- 32:49:14right into this formula. Y equ= slope 1
- 32:49:18* column 1 variable plus slope 2 *
- 32:49:21column 2 variable all the way to the m
- 32:49:23into the n and x to the n + c the
- 32:49:26coefficient or in this case you have -
- 32:49:288.89 8 9 to the power of two etc etc
- 32:49:32times the first column and the second
- 32:49:34column and the third column and then our
- 32:49:36intercept is the minus one3009
- 32:49:39point. Boy, it gets kind of complicated
- 32:49:41when you look at it. This is why we
- 32:49:42don't do this by hand anymore. This is
- 32:49:44why we have the computer to make these
- 32:49:46calculations easy to understand and
- 32:49:48calculate. Now, I told you that was a
- 32:49:50short detour and we're coming towards
- 32:49:52the end of our script. As you remember
- 32:49:54from the beginning, I said if we're
- 32:49:55going to divide this information, we
- 32:49:57have to make sure it's a valid model,
- 32:49:59that this model works and understand how
- 32:50:01good it works. So calculating the R
- 32:50:03squar value, that's what we're going to
- 32:50:05use to predict how good our prediction
- 32:50:08is. And let's take a look at what that
- 32:50:09looks like in code. And so we're going
- 32:50:11to use this from sklearn.metrics.
- 32:50:14We're going to import R2 score. That's
- 32:50:16the R squared value. We're looking at
- 32:50:18the error. So in the R2 score, we take
- 32:50:22our Y test versus our Y predict. Y test
- 32:50:26is the actual values we're testing. That
- 32:50:28was the one that was given to us. So we
- 32:50:29know are true. The Y predict of those
- 32:50:32200 values is what we think it was true.
- 32:50:34And when we go ahead and run this, we
- 32:50:36see we get a 9352.
- 32:50:39That's the R2 score. Now, it's not
- 32:50:41exactly a straight percentage. So it's
- 32:50:43not saying it's 93% correct, but you do
- 32:50:46want that in the upper 90s. O and higher
- 32:50:49shows that this is a very valid
- 32:50:51prediction based on the R2 score. And if
- 32:50:53R squar value of N1 or 92 as we got on
- 32:50:57our model remember it does have a random
- 32:50:59generation involved. This proves the
- 32:51:01model is a good model which means
- 32:51:03success. Yay. We successfully trained
- 32:51:06our model with certain predictors and
- 32:51:08estimated the profit of the companies
- 32:51:10using linear regression. What is
- 32:51:12logistic regression? Let's say we have
- 32:51:15to build a predictive model or a machine
- 32:51:17learning model to predict whether the
- 32:51:20passengers of the Titanic ship have
- 32:51:23survived or not the shipwreck. So how do
- 32:51:24we do that? So we use logistic
- 32:51:26regression to build a model for this.
- 32:51:29How do we use logistic regression? So we
- 32:51:32have the information about the
- 32:51:34passengers, their ID, whether they have
- 32:51:36survived or not, their class and name
- 32:51:39and so on and so forth. And we use this
- 32:51:42information where we already know
- 32:51:44whether the person has survived or not.
- 32:51:46That is the labeled information and we
- 32:51:48help the system to train based on this
- 32:51:52information with based on this labeled
- 32:51:54data. This is known as labeled data. And
- 32:51:56during the process of building the
- 32:51:58model, we probably will remove some of
- 32:52:01the non-essential parameters or
- 32:52:03attributes here. We only take those
- 32:52:05attributes which are really required to
- 32:52:07make these predictions. And once we
- 32:52:09train the model, we run new data through
- 32:52:12it whereby the model will predict
- 32:52:14whether the passenger has survived or
- 32:52:16not. All right. What is logistic
- 32:52:18regression? As I mentioned earlier,
- 32:52:20logistic regression is an algorithm for
- 32:52:22performing binary classification. So
- 32:52:25let's take an example and see how this
- 32:52:27works. Let's say your car has not been
- 32:52:30serviced for quite a few years and now
- 32:52:32you want to find out if it is going to
- 32:52:34break down in the near future. So this
- 32:52:37is like a classification problem. Find
- 32:52:39out whether your car will break down or
- 32:52:41not. So how are we going to perform this
- 32:52:44classification? So here's how it looks.
- 32:52:47If we plot the information along the X
- 32:52:50and Y axis, X is the number of years
- 32:52:52since the last service was performed and
- 32:52:55Y is the probability of your car
- 32:52:58breaking down. And let's say this
- 32:53:00information was this data rather was
- 32:53:03collected from several car users. It's
- 32:53:06not just your car but several car users.
- 32:53:09So that is our labeled data. So the data
- 32:53:11has been collected and um for for the
- 32:53:15number of years and when the car broke
- 32:53:18down and what was the probability and
- 32:53:20that has been plotted along x and y
- 32:53:22axis. So this provides an idea or from
- 32:53:26this graph we can find out whether your
- 32:53:29car will break down or not. We'll see
- 32:53:31how. So first of all the probability can
- 32:53:34go from 0 to one. As you all aware
- 32:53:37probability can be between 0 and one.
- 32:53:39And as we can imagine it is intuitive as
- 32:53:43well. As the number of years are on the
- 32:53:46lower side maybe 1 year, 2 years or 3
- 32:53:49years till after the service the chances
- 32:53:52of your car breaking down are very
- 32:53:54limited. Right? So for example, chances
- 32:53:57of your car breaking down or the
- 32:53:58probability of your car breaking down
- 32:54:00within 2 years of your last service are
- 32:54:030.1 probability. Similarly 3 years is
- 32:54:06maybe.3 and so on. But as the number of
- 32:54:09years increases let's say if it was 6 or
- 32:54:127 years there is almost a certainty that
- 32:54:15your car is going to break down. That is
- 32:54:17what this graph shows. So this is an
- 32:54:19example of a application of the
- 32:54:21classification algorithm and we will see
- 32:54:24in little details how exactly logistic
- 32:54:27regression is applied here. One more
- 32:54:29thing needs to be added here is that the
- 32:54:31dependent variables outcome is discrete.
- 32:54:34So if we are talking about whether the
- 32:54:37car is going to break down or not. So
- 32:54:40that is a discrete value. The y that we
- 32:54:42are talking about the dependent variable
- 32:54:44that we are talking about what we are
- 32:54:46looking at is whether the car is going
- 32:54:48to break down or not yes or no that is
- 32:54:51what we are talking about. So here the
- 32:54:53outcome is discrete and not a continuous
- 32:54:56value. So this is how the logistic
- 32:54:58regression curve looks. Let me explain a
- 32:55:00little bit what exactly and how exactly
- 32:55:03we are going to uh determine the class
- 32:55:07the outcome rather. So for a logistic
- 32:55:10regression curve a threshold has to be
- 32:55:12set saying that because this is a
- 32:55:14probability calculation remember this is
- 32:55:16a probability calculation and the
- 32:55:19probability itself will not be zero or
- 32:55:22one but based on the probability we need
- 32:55:25to decide what the outcome should be. So
- 32:55:28there has to be a threshold like for
- 32:55:29example 0.5 can be the threshold let's
- 32:55:32say in this case. So any value of the
- 32:55:35probability below 0.5 is considered to
- 32:55:38be zero and any value above.5 is
- 32:55:41considered to be one. So an output of
- 32:55:44let's say8
- 32:55:46will mean that the car will break down.
- 32:55:48So that is considered as an output of 1
- 32:55:51and let's say an output of 29 is
- 32:55:54considered as zero which means that the
- 32:55:57car will not break down. So that's the
- 32:55:59way logistic regression works. Now let's
- 32:56:02do a quick comparison between logistic
- 32:56:04regression and linear regression because
- 32:56:06they both have the term regression in
- 32:56:09them. So it can cause confusion. So
- 32:56:11let's try to remove that confusion. So
- 32:56:13what is linear regression? Linear
- 32:56:15regression is a process is once again an
- 32:56:18algorithm for supervised learning.
- 32:56:21However, here you're going to find a
- 32:56:24continuous value. You're going to
- 32:56:25determine a continuous value. It could
- 32:56:27be the price of a real estate property.
- 32:56:30It could be your hike, how much hike
- 32:56:32you're going to get or it could be a
- 32:56:33stock price. These are all continuous
- 32:56:36values. These are not discrete compared
- 32:56:38to a yes or a no kind of a response that
- 32:56:40we are looking for in logistic
- 32:56:42regression. So this is one example of a
- 32:56:44linear regression. Let's say the HR team
- 32:56:47of a company tries to find out what
- 32:56:49should be the salary hike of an
- 32:56:51employee. So they collect all the
- 32:56:53details of their existing employees,
- 32:56:55their ratings and their salary hikes,
- 32:56:57what has been given and that is the
- 32:56:59labeled information that is available
- 32:57:01and the system learns from this. It is
- 32:57:04trained and it learns from this labeled
- 32:57:06information so that when a new employees
- 32:57:09information is fed based on the rating
- 32:57:11it will determine what should be the
- 32:57:13high. So this is a linear regression
- 32:57:15problem and a linear regression example.
- 32:57:18Now salary is a continuous value. You
- 32:57:21can get 5,000, 5,500,
- 32:57:255,600. It is not discrete like a cat or
- 32:57:29a dog or an apple or a banana. These are
- 32:57:32discrete or a yes or a no. These are
- 32:57:34discrete values, right? So this where
- 32:57:37you're trying to find continuous values
- 32:57:40is where we use linear regression. So
- 32:57:42let's say just to extend on this
- 32:57:44scenario, we now want to find out
- 32:57:47whether this employee is going to get a
- 32:57:49promotion or not. So we want to find out
- 32:57:53that is a discrete problem, right? A yes
- 32:57:55or no kind of a problem. In this case,
- 32:57:58we actually cannot use linear regression
- 32:58:01even though we may have labeled data. So
- 32:58:04this is the label data. So based on the
- 32:58:06employee rating these are the ratings
- 32:58:08and then some people got the promotion
- 32:58:11and this is the ratings for which people
- 32:58:13did not get promotion that is a no and
- 32:58:16this is the rating for which people got
- 32:58:18promotion we just plotted the data about
- 32:58:21whether a person has got an employee has
- 32:58:23got promotion or not yes no right so
- 32:58:26there is nothing in between and what is
- 32:58:28the employees rating okay and ratings
- 32:58:30can be continuous that is not an issue
- 32:58:32but the output is discrete in In this
- 32:58:35case whether employee got promotion yes
- 32:58:37no okay so if we try to plot that and we
- 32:58:40try to find a straight line this is how
- 32:58:43it would look and as you can see it
- 32:58:45doesn't look very right because looks
- 32:58:47like there will be lot of errors this
- 32:58:49root mean square error if you remember
- 32:58:51for linear regression would be very very
- 32:58:54high and also the the values cannot go
- 32:58:57beyond zero or beyond one. So the graph
- 32:59:00should probably look somewhat like this
- 32:59:02clipped at 0 and one. But still the
- 32:59:06straight line doesn't look right.
- 32:59:09Therefore instead of using a linear
- 32:59:12equation we need to come up with
- 32:59:14something different and therefore the
- 32:59:16logistic regression model looks somewhat
- 32:59:19like this. So we calculate the
- 32:59:21probability and if we plot that
- 32:59:23probability not in the form of a
- 32:59:25straight line but we need to use some
- 32:59:27other equation. And we will see very
- 32:59:28soon what that equation is. Then it is a
- 32:59:31gradual process. Right? So you see here
- 32:59:34people with some of these ratings are
- 32:59:35not getting any promotions and then
- 32:59:38slowly uh at certain rating they get
- 32:59:41promotion. So that is a gradual process
- 32:59:44and uh this is how the math behind
- 32:59:47logistic regression looks. So we are
- 32:59:49trying to find the odds for a particular
- 32:59:53event happening and this is the formula
- 32:59:55for finding the odds. So the probability
- 32:59:57of an event happening divided by the
- 33:00:00probability of the event not happening.
- 33:00:03So P if it is the probability of the
- 33:00:05event happening probability of the
- 33:00:07person getting a promotion and divided
- 33:00:09by the probability of the person not
- 33:00:11getting a promotion that is 1 minus P.
- 33:00:15So this is how you measure the odds. Now
- 33:00:17the values of the odds range from 0 to
- 33:00:21infinity. So when this probability is
- 33:00:24zero then the odds will the value of the
- 33:00:27odds is equal to zero and when the
- 33:00:29probability becomes 1 then the value of
- 33:00:32the odds is 1 by 0 that will be infinity
- 33:00:35but the probability itself remains
- 33:00:37between 0 and 1. Now this is how an
- 33:00:40equation of a straight line looks. So y
- 33:00:42is equal to beta 0 plus beta 1x where
- 33:00:45beta 0 is the y intercept and beta 1 is
- 33:00:48the slope of the line. If we take the
- 33:00:51odds equation and take a log of both
- 33:00:54sides, then this would look somewhat
- 33:00:56like this. And the term logistic is
- 33:00:59actually derived from the fact that we
- 33:01:01are doing this. We take a log of px by 1
- 33:01:04minus px. This is an extension of the
- 33:01:07calculation of odds that we have seen,
- 33:01:09right? And that is equal to beta 0 plus
- 33:01:12beta 1x which is the equation of the
- 33:01:13straight line. And now from here if you
- 33:01:16want to find out the value of px we will
- 33:01:20see we can take the exponential on both
- 33:01:22sides and then if we solve that equation
- 33:01:25we will get the equation of px like this
- 33:01:28px is equal to 1 by 1 + e ^ of minus
- 33:01:33beta 0 + beta 1x and recall this is
- 33:01:36nothing but the equation of the line
- 33:01:37which is equal to y is equal to beta 0 +
- 33:01:40beta 1x. So that this is the equation
- 33:01:44also known as the sigmoid function and
- 33:01:46this is the equation of the logistic
- 33:01:48regression alg. All right and if this is
- 33:01:50plotted this is how the sigmoid curve is
- 33:01:53obtained. So let's compare linear and
- 33:01:57logistic regression how they are
- 33:01:59different from each other. Let's go
- 33:02:00back. So linear regression is solved or
- 33:02:03used to solve regression problems and
- 33:02:07logistic regression is used to solve
- 33:02:10classification problems. So both are
- 33:02:12called regression. But linear regression
- 33:02:14is used for solving regression problems
- 33:02:17where we predict continuous values.
- 33:02:19Whereas logistic regression is used for
- 33:02:22solving classification problems where we
- 33:02:24have had to predict discrete values. The
- 33:02:27response variables in case of linear
- 33:02:29regression are continuous in nature.
- 33:02:31Whereas here they are categorical or
- 33:02:33discrete in nature. And the linear
- 33:02:36regression helps to estimate the
- 33:02:38dependent variable when there is a
- 33:02:40change in the independent variable.
- 33:02:42Whereas here in case of logistic
- 33:02:44regression it helps to calculate the
- 33:02:46probability or the possibility of a
- 33:02:49particular event happening. And linear
- 33:02:51regression as the name suggests is a
- 33:02:53straight line. That's why it's called
- 33:02:54linear regression. Whereas logistic
- 33:02:57regression is a sigmoid function and the
- 33:03:00curve is the shape of the curve is S.
- 33:03:03It's an S-shaped curve. This is another
- 33:03:05example of application of logistic
- 33:03:07regression in weather prediction.
- 33:03:09Whether it's going to rain or not rain.
- 33:03:12Now keep in mind both are used in
- 33:03:14weather prediction. If we want to find
- 33:03:16the discrete values like whether it's
- 33:03:18going to rain or not rain that is a
- 33:03:21classification problem. We use logistic
- 33:03:23regression. But if we want to determine
- 33:03:25what is going to be the temperature
- 33:03:26tomorrow, then we use linear regression.
- 33:03:29So just keep in mind that in weather
- 33:03:31prediction, we actually use both. But
- 33:03:33these are some examples of logistic
- 33:03:34regression. So we want to find out
- 33:03:36whether it's going to be rain or not,
- 33:03:38it's going to be sunny or not, whether
- 33:03:40it's going to snow or not. These are all
- 33:03:42logistic regression examples. A few more
- 33:03:44examples. Classification of objects.
- 33:03:47This is a again another example of
- 33:03:49logistic regression. Now here of course
- 33:03:52one distinction is that these are
- 33:03:54multiclass classification. So logistic
- 33:03:57regression is not used in its original
- 33:04:01form but it is used in a slightly
- 33:04:03different form. So we say whether it is
- 33:04:05a dog or not a dog. I hope you
- 33:04:08understand. So instead of saying is it a
- 33:04:09dog or a cat or elephant we convert this
- 33:04:12into saying so because we need to keep
- 33:04:15it to binary classification. So we say
- 33:04:18is it a dog or not a dog? Is it a cat or
- 33:04:22not a cat? So that's the way logistic
- 33:04:24regression can be used for classifying
- 33:04:26objects. Otherwise there are other
- 33:04:28techniques which can be used for
- 33:04:30performing multiclass classification. In
- 33:04:32healthcare logistic regression is used
- 33:04:34to find the survival rate of a patient.
- 33:04:37So they take multiple parameters like
- 33:04:40trauma score and age and so on and so
- 33:04:42forth and they try to predict the rate
- 33:04:45of survival. All right. Now finally
- 33:04:48let's take an example and see how we can
- 33:04:50apply logistic regression to predict the
- 33:04:53number that is shown in the image. So
- 33:04:56this is actually a live demo. I will
- 33:04:58take you into Jupyter notebook and u
- 33:05:01show the code. But before that let me
- 33:05:03take you through a couple of slides to
- 33:05:05explain what we're trying to do. So
- 33:05:07let's say you have an 8x8 image and the
- 33:05:10the image has a number 1 2 3 4 and you
- 33:05:13need to train your model to predict what
- 33:05:16this number is. So how do we do this? So
- 33:05:19the first thing is obviously in any
- 33:05:21machine learning process you train your
- 33:05:23model. So in this case we are using
- 33:05:25logistic regression. So and then we
- 33:05:28provide a training set to train the
- 33:05:30model and then we test how accurate our
- 33:05:33model is with the test data which means
- 33:05:36that like any machine learning process.
- 33:05:38We split our initial data into two parts
- 33:05:41training set and test set. With the
- 33:05:42training set we train our model and then
- 33:05:44with the test set we we test the model
- 33:05:46till we get good accuracy and then we
- 33:05:49use it for for inference. Right? So that
- 33:05:52is typical methodology of uh uh
- 33:05:56training, testing and then deploying of
- 33:05:58machine learning models. So let's uh
- 33:06:00take a look at the code and uh see what
- 33:06:02we are doing. So I'll not go line by
- 33:06:04line but just take you through some of
- 33:06:06the blocks. So first thing we do is
- 33:06:08import all the libraries and then we
- 33:06:11basically take a look at the images and
- 33:06:14see what is the total number of images.
- 33:06:16We can display using mattplot lip some
- 33:06:19of the images or a sample of these
- 33:06:21images and um then we split the data
- 33:06:24into training and test as I mentioned
- 33:06:25earlier and we can do some exploratory
- 33:06:28analysis and uh then we build our model.
- 33:06:32We train our model with the training set
- 33:06:34and then we test it with our test set
- 33:06:37and find out how accurate our model is
- 33:06:40using the confusion matrix the heat map
- 33:06:43and use heat map for visualizing this
- 33:06:45and uh I will show you in the code what
- 33:06:48exactly is the confusion matrix and how
- 33:06:50it can be used for finding the accuracy
- 33:06:53in our example we got we get an accuracy
- 33:06:56of about 94 which is pretty good or 94%
- 33:06:59which is pretty good all right so what
- 33:07:01is the confusion matrix. This is an
- 33:07:03example of a confusion matrix and uh
- 33:07:06this is used for identifying the
- 33:07:09accuracy of a classification model or
- 33:07:12like a logistic regression model. So the
- 33:07:15most important part in a confusion
- 33:07:17matrix is that first of all this as you
- 33:07:19can see this is a matrix and the size of
- 33:07:22the matrix depends on how many outputs
- 33:07:24uh we are expecting right. So the the
- 33:07:27most important part here is that the
- 33:07:29model will be most accurate when we have
- 33:07:32the maximum numbers in its diagonal like
- 33:07:36in this case that's why it has almost 93
- 33:07:3994% because the diagonals should have
- 33:07:42the maximum numbers and the others other
- 33:07:45than diagonals the cells other than the
- 33:07:47diagonals should have very few numbers.
- 33:07:50So here that's what is happening. So
- 33:07:51there is a two here. There are there's a
- 33:07:53one here. But most of them are along the
- 33:07:57diagonal. This what does this mean? This
- 33:07:59means that the number that has been fed
- 33:08:03is zero and the number that has been
- 33:08:06detected is also zero. So the predicted
- 33:08:09value and the actual value are the same.
- 33:08:11So along the diagonals that is true.
- 33:08:14Which means that let's let's take this
- 33:08:16diagonal right. If if the maximum number
- 33:08:18is here that means that uh like here in
- 33:08:21this case it is 34 which means that 34
- 33:08:24of the images that have been fed or
- 33:08:26rather actually there are two
- 33:08:27mclassifications in there. So 36 images
- 33:08:31have been fed which have number four and
- 33:08:33out of which 34 have been predicted
- 33:08:36correctly as number four and one has
- 33:08:39been predicted as number eight and
- 33:08:41another one has been predicted as number
- 33:08:43nine. So these are two mclassifications.
- 33:08:46Okay. So that is the meaning of saying
- 33:08:48that the maximum number should be in the
- 33:08:50diagonal. So if you have all of them so
- 33:08:52for an ideal model which has let's say
- 33:08:55100% accuracy everything will be only in
- 33:08:58the diagonal. There will be no numbers
- 33:09:00other than zero in all other cells. So
- 33:09:02that is like a 100% accurate model.
- 33:09:05Okay. So that's the gist of how to use
- 33:09:08this matrix. How to use this uh
- 33:09:10confusion matrix. So I know the name uh
- 33:09:13is a little funny sounding confusion
- 33:09:15matrix but actually it is not very
- 33:09:17confusing. It's very straightforward. So
- 33:09:19you are just plotting what has been
- 33:09:21predicted and what is the labeled
- 33:09:24information or what is the actual data
- 33:09:26that's also known as the ground truth
- 33:09:28sometimes. Okay, these are some fancy
- 33:09:29terms that are used. So predicted label
- 33:09:31and the actual label that's all it is.
- 33:09:34Okay. Yeah. So we are showing a little
- 33:09:35bit more information here. So 38 have
- 33:09:38been predicted and here you will see
- 33:09:40that all of them have been predicted
- 33:09:41correctly. There have been 38 zeros and
- 33:09:44the predicted value and the actual value
- 33:09:47is is exactly the same. Whereas in this
- 33:09:49case right it has uh there are I think
- 33:09:5237 + 5 yeah 42 have been fed the images
- 33:09:5742 images are of digit three and uh the
- 33:10:01accuracy is only 37 of them have been
- 33:10:04accurately predicted. Three of them have
- 33:10:07been predicted as number seven and two
- 33:10:09of them have been predicted as number
- 33:10:11eight and so on and so forth. Okay. All
- 33:10:13right. So with that let's go into
- 33:10:15Jupyter notebook and see how the code
- 33:10:17looks. So this is the code in in Jupyter
- 33:10:22notebook for logistic regression. In
- 33:10:25this particular demo, what we are going
- 33:10:27to do is train our model to recognize
- 33:10:31digits which are the images which have
- 33:10:34digits from let's say 0 to 5 or 0 to 9
- 33:10:38and um and then we will see how well it
- 33:10:41is trained and whether it is able to
- 33:10:43predict these numbers correctly or not.
- 33:10:46So let's get started. So the first part
- 33:10:48is as usual we are importing some
- 33:10:51libraries that are required and uh then
- 33:10:55the last line in this block is to load
- 33:10:58the digits. So let's go ahead and run
- 33:11:02this code. Then here we will visualize
- 33:11:06the shape of these uh digits. So we can
- 33:11:08see here if we take a look this is how
- 33:11:11the shape is 1797 by 64. These are like
- 33:11:158 by8 images. So that's that's what is
- 33:11:17reflected in this uh shape. Now from
- 33:11:20here onwards we are basically once again
- 33:11:22importing some of the libraries that are
- 33:11:24required like numpy and map plot and we
- 33:11:27will take a look at uh some of the
- 33:11:29sample images that we have loaded. So th
- 33:11:33this one for example creates a figure uh
- 33:11:36and then we go ahead and take a few
- 33:11:38sample images to see how they look. So
- 33:11:41let me run this code and so that it
- 33:11:43becomes easy to understand. So these are
- 33:11:46about five images sample images that we
- 33:11:48are looking at 0 1 2 3 4. So this is how
- 33:11:52the images this is how the data is.
- 33:11:54Okay. And uh based on this we will
- 33:11:57actually train our logistic regression
- 33:12:00model and then we will test it and see
- 33:12:03how well it is able to recognize. So the
- 33:12:06way it works is the pixel information.
- 33:12:09So as you can see here this is an 8x 8
- 33:12:12pixel kind of a image and uh the each
- 33:12:16pixel whether it is activated or not
- 33:12:19activated that is the information
- 33:12:20available for each pixel. Now based on
- 33:12:23the pattern of this activation and
- 33:12:26non-activation of the various pixels
- 33:12:28this will be identified as a zero for
- 33:12:31example right similarly as you can see
- 33:12:34so overall each of these numbers
- 33:12:37actually has a different pattern of the
- 33:12:40pixel activation and that's pretty much
- 33:12:42that our model needs to learn for which
- 33:12:45number what is the pattern of the
- 33:12:47activation of the pixels right so that
- 33:12:50is what we are going to train our model.
- 33:12:52Okay. So the first thing we need to do
- 33:12:55is to split our data into training and
- 33:12:59test data set. Right? So whenever we
- 33:13:01perform any training, we split the data
- 33:13:04into training and test. So that the
- 33:13:06training data set is used to train the
- 33:13:08system. So we pass this probably
- 33:13:11multiple times. Uh and then we test it
- 33:13:14with the test data set. And the split is
- 33:13:16usually in the form of there and there
- 33:13:18are various ways in which you can split
- 33:13:20this data. It is up to the individual
- 33:13:22preferences. In our case here we are
- 33:13:25splitting in the form of 23 and 77. So
- 33:13:29when we say test size as 2023
- 33:13:33that means 23% of the entire data is
- 33:13:38used for testing and the remaining 77%
- 33:13:41is used for training. So there is a
- 33:13:43readily available function which is uh
- 33:13:46called train test split. So we don't
- 33:13:49have to write any special code for the
- 33:13:51splitting. It will automatically split
- 33:13:53the data based on the proportion that we
- 33:13:56give here which is test size. So we just
- 33:13:58give the test size automatically
- 33:14:00training size will be determined and uh
- 33:14:02we pass the data that we want to split
- 33:14:05and the the results will be stored in x
- 33:14:09train and y train for the training data
- 33:14:13set. And what is x train? This are these
- 33:14:16are the features right which is like the
- 33:14:18independent variable and y train is the
- 33:14:22label right so in this case what happens
- 33:14:25is we have the input value which is or
- 33:14:28the features value which is in x train
- 33:14:30and since this is a labeled data for
- 33:14:33each of them each of the observations we
- 33:14:36already have the label information
- 33:14:38saying whether this digit is a zero or a
- 33:14:41one or a two so that this this is what
- 33:14:43will be used for comparison to find out
- 33:14:46whether the the system is able to
- 33:14:48recognize it correctly or there is an
- 33:14:50error for each observation it will
- 33:14:52compare with this right so this is the
- 33:14:54label so the same way x train y train is
- 33:14:58for the training data set x test y test
- 33:15:02is for the test data set okay so let me
- 33:15:05go ahead and execute this code as well
- 33:15:06and then we can go and check quickly
- 33:15:09what is the how many entries are there
- 33:15:11and in each of this so x train the shape
- 33:15:14is 1383x
- 33:15:1764 and y train has 1383 because there is
- 33:15:22uh nothing like the second part is not
- 33:15:24required here and then x test shape we
- 33:15:28see is 414 so actually there are 414
- 33:15:32observations in test and 1383
- 33:15:35observations in train so that's
- 33:15:37basically what these four lines of code
- 33:15:39are are saying okay then we import the
- 33:15:42uh logistic regression
- 33:15:44library and uh which is a part of
- 33:15:47scikitlearn. So we we don't have to
- 33:15:50implement the logistic regression
- 33:15:51process itself. We just call these uh
- 33:15:53the function and uh let me go ahead and
- 33:15:56execute that so that uh we have the
- 33:15:59logistic regression library imported.
- 33:16:01Now we create an instance of logistic
- 33:16:04regression. Right? So logistic regr is a
- 33:16:07is an instance of logistic regression
- 33:16:10and then we use that for training our
- 33:16:12model. So let me first execute this
- 33:16:15code. So these two lines. So the first
- 33:16:17line basically creates an instance of
- 33:16:19logistic regression model and then the
- 33:16:22second line is where we are passing our
- 33:16:25data the training data set. Right? This
- 33:16:27is our the the predictors and uh this is
- 33:16:31our target. We are passing this data set
- 33:16:33to train our model. All right. So once
- 33:16:36we do this in this case the data is not
- 33:16:39large but by and large uh the training
- 33:16:42is what takes usually a lot of time. So
- 33:16:44we spend in machine learning activities
- 33:16:47in machine learning projects we spend a
- 33:16:50lot of time for the training part of it.
- 33:16:52Okay. So here the data set is relatively
- 33:16:54small so it was pretty quick. So all
- 33:16:56right so now our model has been trained
- 33:16:59using the training data set and uh we
- 33:17:02want to see how accurate this is. So
- 33:17:04what we'll do is we will test it out in
- 33:17:07probably faces. So let me first try out
- 33:17:10how well this is working for one image.
- 33:17:14Okay, I will just try it out with one
- 33:17:16image my the first entry in my test data
- 33:17:19set and see whether it is uh correctly
- 33:17:21predicting or not. So and in order to
- 33:17:24test it so for training purpose we use
- 33:17:26the fit method. There is a method called
- 33:17:29fit which is for training the model and
- 33:17:32once the training is done if you want to
- 33:17:34test for uh a particular value new input
- 33:17:37you use the predict method. Okay. So
- 33:17:39let's run the predict method and we pass
- 33:17:42this particular image and uh we see that
- 33:17:46the shape is or the prediction is four.
- 33:17:50So let's try a few more. Let me see for
- 33:17:53the next 10 uh seems to be fine. So let
- 33:17:56me just go ahead and test the entire
- 33:17:58data set. Okay, that's basically what we
- 33:18:00will do. So now we want to find out how
- 33:18:02accurately this has u performed. So we
- 33:18:06use the score method to find what is the
- 33:18:09percentages of accuracy and we see here
- 33:18:11that it has performed up to 94%
- 33:18:14accurate. Okay. So that's uh on this
- 33:18:17part. Now what we can also do is we can
- 33:18:20um also see this accuracy using what is
- 33:18:23known as confusion matrix. So let us go
- 33:18:26ahead and uh try that as well. Uh so
- 33:18:29that we can also visualize how well uh
- 33:18:32this model has uh done. So let me
- 33:18:35execute this piece of code which will
- 33:18:37basically import some of the libraries
- 33:18:39that are required and um we we basically
- 33:18:42create a confusion matrix an instance of
- 33:18:46confusion matrix by running confusion
- 33:18:48matrix and passing these uh values. So
- 33:18:52we have so this confusion_matrix
- 33:18:55method takes two parameters one is the y
- 33:18:59test and the other is uh the prediction.
- 33:19:02So what is a y test? These are the
- 33:19:04labeled values which we already know for
- 33:19:06the test data set and predictions are
- 33:19:10what the system has predicted for the
- 33:19:13test data set. Okay. So this is known to
- 33:19:16us and this is what the system has uh
- 33:19:19the model has generated. So we kind of
- 33:19:22create the confusion matrix and we will
- 33:19:24print it. And uh this is how the
- 33:19:26confusion matrix looks. As the name
- 33:19:28suggests it is a matrix and um the key
- 33:19:32point out here is that the accuracy of
- 33:19:34the model is determined by how many
- 33:19:38numbers are there in the diagonal. The
- 33:19:40more the numbers in the diagonal, the
- 33:19:43better the accuracy is. Okay. And first
- 33:19:46of all, the total sum of all the numbers
- 33:19:48in this whole matrix is equal to the
- 33:19:50number of observations in the test data
- 33:19:53set. That is the first thing, right? So
- 33:19:55if you add up all these numbers, that
- 33:19:57will be equal to the number of
- 33:19:59observations in the test data set. And
- 33:20:01then out of that, the maximum number of
- 33:20:04them should be in the diagonal. That
- 33:20:06means the accuracy is pretty good. If
- 33:20:08the the numbers in the diagonal are less
- 33:20:10and in all other places there are a lot
- 33:20:12of numbers uh which means the accuracy
- 33:20:15is very low. The diagonal indicates a
- 33:20:17correct prediction that this means that
- 33:20:19the actual value is same as the
- 33:20:22predicted value. Here again actual value
- 33:20:24is same as the predicted value and so
- 33:20:25on. Right? So the moment you see a
- 33:20:27number here that means the actual value
- 33:20:29is something and the predicted value is
- 33:20:32something else. Right? Similarly here
- 33:20:34the actual value is something and the
- 33:20:36predicted value is something else. So
- 33:20:38that is basically how we read the
- 33:20:41confusion matrix. Now how do we find the
- 33:20:45accuracy? You can actually add up the
- 33:20:47total values in the diagonal. So it it's
- 33:20:50like 38 + 44 + 43 and so on and divide
- 33:20:54that by the total number of test
- 33:20:56observations that will give you the
- 33:20:58percentage accuracy using a confusion
- 33:21:01matrix. Now let us visualize this
- 33:21:03confusion matrix in a slightly more
- 33:21:06sophisticated way uh using a heat map.
- 33:21:09So we will create a heat map with some
- 33:21:11we'll add some colors as well. It's uh
- 33:21:14it's like a more visually visually more
- 33:21:17appealing. So that's the whole idea. So
- 33:21:19if we let me run this piece of code and
- 33:21:21this is how the heat map looks. Uh and
- 33:21:25as you can see here the diagonals again
- 33:21:28are all the values are here most of the
- 33:21:30values. So which means reasonably this
- 33:21:32seems to be reasonably accurate and yeah
- 33:21:35basically the accuracy score is 94%.
- 33:21:38This is calculated as I mentioned by
- 33:21:40adding all these numbers divided by the
- 33:21:43total test values or the total number of
- 33:21:45observations in test data set. Okay. So
- 33:21:49this is the confusion matrix for
- 33:21:51logistic regression.
- 33:21:54All right. So now that we have seen the
- 33:21:57confusion matrix, let's take a quick
- 33:22:00sample and see how well uh the system
- 33:22:02has classified and we will take a a few
- 33:22:05examples of the data. So if we see here
- 33:22:09we we picked up randomly a few of them.
- 33:22:11So this is uh number four which is the
- 33:22:14actual value and also the predicted
- 33:22:16value both are four. This is an image of
- 33:22:19zero. So the predicted value is also
- 33:22:22zero. Actual value is of course zero.
- 33:22:24Then this is the image of nine. So this
- 33:22:27has also been predicted correctly 9 and
- 33:22:30actual value is 9. And this is the image
- 33:22:32of one. And again this has been
- 33:22:34predicted correctly as like the actual
- 33:22:37value. Okay. So this was a quick demo of
- 33:22:40logistic regression. How to use logistic
- 33:22:43regression to identify images.
- 33:22:46>> What is a decision tree? Let's go
- 33:22:48through a very simple example before we
- 33:22:50dig in deep. Decision tree is a
- 33:22:52treeshaped diagram used to determine a
- 33:22:54course of action. Each branch of the
- 33:22:56tree represents a possible decision or
- 33:22:58occurrence or reaction. Let's start with
- 33:23:00a simple question. How to identify a
- 33:23:02random vegetable from a shopping bag?
- 33:23:04So, we have this group of vegetables in
- 33:23:06here. And we can start off by asking a
- 33:23:07simple question. Is it red? And if it's
- 33:23:09not, then it's going to be the purple
- 33:23:12fruit to the left, probably an eggplant.
- 33:23:14If it's true, it's going to be one of
- 33:23:15the red fruits. Is the diameter greater
- 33:23:17than two? If false, it's going to be a
- 33:23:20what looks to be a red chili. And if
- 33:23:21it's true, it's going to be a bell
- 33:23:24pepper from the capsicum family. So,
- 33:23:26it's a capsicum.
- 33:23:28Problems that decision tree can solve.
- 33:23:30So, let's look at the two different
- 33:23:31categories the decision tree can be used
- 33:23:33on. It can be used on the
- 33:23:35classification, the true false, yes, no,
- 33:23:37and it can be used on regression where
- 33:23:39we figure out what the next value is in
- 33:23:41a series of numbers or a group of data.
- 33:23:43In classification, the classification
- 33:23:45tree will determine a set of logical if
- 33:23:49then conditions to classify problems.
- 33:23:51For example, discriminating between
- 33:23:53three types of flowers based on certain
- 33:23:55features. In regression, a regression
- 33:23:58tree is used when the target variable is
- 33:24:00numerical or continuous in nature. We
- 33:24:02fit the regression model to the target
- 33:24:04variable using each of the independent
- 33:24:06variables. Each split is made based on
- 33:24:08the sum of squared error. Before we dig
- 33:24:12deeper into the mechanics of the
- 33:24:14decision tree, let's take a look at the
- 33:24:16advantages of using a decision tree and
- 33:24:18we'll also take a glimpse at the
- 33:24:20disadvantages. The first thing you'll
- 33:24:22notice is that it's simple to
- 33:24:24understand, interpret, and visualize. It
- 33:24:27really shines here because you can see
- 33:24:28exactly what's going on in a decision
- 33:24:30tree. Little effort is required for data
- 33:24:32preparation. So, you don't have to do
- 33:24:34special scaling. There's a lot of things
- 33:24:36you don't have to worry about when using
- 33:24:37a decision tree. It can handle both
- 33:24:39numerical and categorical data as we
- 33:24:42discovered earlier and nonlinear
- 33:24:44parameters don't affect its performance.
- 33:24:46So even if the data doesn't fit an easy
- 33:24:49curved graph, you can still use it to
- 33:24:51create an effective decision or
- 33:24:54prediction. If we're going to look at
- 33:24:56the advantages of a decision tree, we
- 33:24:59also need to understand the
- 33:25:00disadvantages of a decision tree. The
- 33:25:02first disadvantage is overfitting.
- 33:25:04Overfitting occurs when the algorithm
- 33:25:06captures noise in the data. That means
- 33:25:08you're solving for one specific instance
- 33:25:10instead of a general solution for all
- 33:25:13the data. High variance. The model can
- 33:25:15get unstable due to small variation in
- 33:25:17data. Low bias tree. A highly
- 33:25:20complicated decision tree tends to have
- 33:25:22a low bias which makes it difficult for
- 33:25:24the model to work with new data.
- 33:25:26Decision tree important terms. Before we
- 33:25:30dive in further, we need to look at some
- 33:25:32basic terms. We need to have some
- 33:25:35definitions to go with our decision tree
- 33:25:37in the different parts we're going to be
- 33:25:38using. We'll start with entropy. Entropy
- 33:25:40is a measure of randomness or
- 33:25:42unpredictability in the data set. For
- 33:25:45example, we have a group of animals in
- 33:25:47this picture. There's four different
- 33:25:48kinds of animals. And this data set is
- 33:25:50considered to have a high entropy. You
- 33:25:52really can't pick out what kind of
- 33:25:54animal it is based on looking at just
- 33:25:55the four animals as a big clump of of uh
- 33:25:58entities. So as we start splitting it
- 33:26:02into subgroups, we come up with our
- 33:26:04second definition which is information
- 33:26:07gain. Information gain it is a measure
- 33:26:09of decrease in entropy after the data
- 33:26:12set is split. So in this case based on
- 33:26:14the color yellow, we've split one group
- 33:26:16of animals on one side as true and those
- 33:26:19who aren't yellow as false. As we
- 33:26:21continue down the yellow side, we split
- 33:26:23based on the height. True or false
- 33:26:24equals 10. And on the other side, height
- 33:26:26is less than 10. True or false? And as
- 33:26:29you see as we split it, the entropy
- 33:26:31continues to be less and less and less.
- 33:26:33And so our information gain is simply
- 33:26:35the entropy E1 from the top and how it's
- 33:26:38changed to E2 in the bottom. And we'll
- 33:26:40look at the uh deeper math, although you
- 33:26:42really don't need to know a huge amount
- 33:26:44of math when you actually do the
- 33:26:45programming in Python because it'll do
- 33:26:47it for you. But we'll look on the actual
- 33:26:48math of how they compute entropy.
- 33:26:50Finally, we want to know the different
- 33:26:51parts of our tree and they call the leaf
- 33:26:54node. Leaf node carries the
- 33:26:55classification or the decision. So it's
- 33:26:57the final end at the bottom. The
- 33:26:59decision node has two or more branches.
- 33:27:02This is where we're breaking the group
- 33:27:03up into different parts. And finally,
- 33:27:06you have the root node. The topmost
- 33:27:08decision node is known as the root node.
- 33:27:11How does a decision tree work? Wonder
- 33:27:13what kind of animals I'll get in the
- 33:27:15jungle today? Maybe you're the hunter
- 33:27:17with the gun. Or if you're more into
- 33:27:18photography, you're a photographer with
- 33:27:20a camera. So let's look at this group of
- 33:27:22animals and let's try to classify
- 33:27:24different types of animals based on
- 33:27:26their features using a decision tree. So
- 33:27:28the problem statement is to classify the
- 33:27:31different types of animals based on
- 33:27:32their features using a decision tree.
- 33:27:34The data set is looking quite messy and
- 33:27:36the entropy is high in this case. So
- 33:27:38let's look at a training set or a
- 33:27:40training data set and we're looking at
- 33:27:42color. We're looking at height and then
- 33:27:44we have our different animals. We have
- 33:27:46our elephants, our giraffes, our
- 33:27:48monkeys, and our tigers. And they're of
- 33:27:50different colors and shapes. Let's see
- 33:27:52what that looks like. And how do we
- 33:27:53split the data? We have to frame the
- 33:27:55conditions that split the data in such a
- 33:27:58way that the information gain is the
- 33:28:00highest. Note, gain is the measure of
- 33:28:02decrease in entropy after splitting. So
- 33:28:05the formula for entropy is the sum
- 33:28:08that's what this symbol looks like. That
- 33:28:09looks like kind of like a uh e funky e
- 33:28:12of k where i equals 1 to k. K would
- 33:28:15represent the number of animal the
- 33:28:17different animals in there where value
- 33:28:19or P value of I would be the percentage
- 33:28:23of that animal times the log base 2 of
- 33:28:26the same the percentage of that animal.
- 33:28:28Let's try to calculate the entropy for
- 33:28:29the current data set and take a look at
- 33:28:32what that looks like. And don't be
- 33:28:33afraid of the math. You don't really
- 33:28:35have to memorize this math. Just be
- 33:28:37aware that it's there and this is what's
- 33:28:38going on in the background. And so we
- 33:28:40have three giraffes, two tigers, one
- 33:28:43monkey, two elephants, a total of eight
- 33:28:45animals gathered. And if we plug that
- 33:28:46into the formula, we get an entropy that
- 33:28:49equals 3 over8. So we have three
- 33:28:51giraffes, a total of 8 times the log.
- 33:28:54Usually they use base 2 on the log. So
- 33:28:56log base 2 of 3 over8 plus in this case,
- 33:29:00let's say it's the elephants, 2 over 8.
- 33:29:02Two elephants over total of 8 time log
- 33:29:04base 2 2 over 8 plus one monkey over
- 33:29:07total of 8. log base 2 1 over 8 and plus
- 33:29:112 over 8 of the tigers log base 2 over 8
- 33:29:15and if we plug that into our computer or
- 33:29:17calculator I obviously can't do logs in
- 33:29:19my head we get an entropy equal to.571
- 33:29:23the program will actually calculate the
- 33:29:25entropy of the data set similarly after
- 33:29:27every split to calculate the gain now
- 33:29:30we're not going to go through each set
- 33:29:31one at a time to see what those numbers
- 33:29:34are just want you to be aware that this
- 33:29:36is a formula or the mathematics behind
- 33:29:38It gain can be calculated by finding the
- 33:29:39difference of the subsequent entropy
- 33:29:41values after a split. Now we will try to
- 33:29:43choose a condition that gives us the
- 33:29:45highest gain. We will do that by
- 33:29:47splitting the data using each condition
- 33:29:49and checking that the gain we get out of
- 33:29:51them. The condition that gives us the
- 33:29:52highest gain will be used to make the
- 33:29:54first split. Can you guess what that
- 33:29:56first split will be just by looking at
- 33:29:58this image? As a human, it's probably
- 33:30:00pretty easy to split it. Let's see if
- 33:30:02you're right. If you guessed the color
- 33:30:04yellow, you're correct. Let's say the
- 33:30:06condition that gives us the maximum gain
- 33:30:08is yellow. So we will split the data
- 33:30:10based on the color yellow. If it's true,
- 33:30:12that group of animals goes to the left.
- 33:30:14If it's false, it goes to the right. The
- 33:30:16entropy after the splitting has
- 33:30:18decreased considerably. However, we
- 33:30:21still need some splitting at both the
- 33:30:23branches to attain an entropy value
- 33:30:25equal to zero. So we decide to split
- 33:30:27both the nodes using height as a
- 33:30:29condition. Since every branch now
- 33:30:30contains single label type, we can say
- 33:30:33that entropy in this case has reached
- 33:30:35the least value. And here you see we
- 33:30:37have the giraffes, the tigers, the
- 33:30:39monkey and the elephants all separated
- 33:30:40into their own groups. This tree can now
- 33:30:42predict all the classes of animals
- 33:30:44present in the data set with 100%
- 33:30:46accuracy. That was easy. Use case loan
- 33:30:50repayment prediction. Let's get into my
- 33:30:52favorite part and open up some Python
- 33:30:54and see what the programming code and
- 33:30:56the scripting looks like. In here, we're
- 33:30:58going to want to do a prediction. And we
- 33:31:00start with this individual here who's
- 33:31:01requesting to find out how good his
- 33:31:03customers are going to be, whether
- 33:31:04they're going to repay their loan or not
- 33:31:06for this bank. And from that, we want to
- 33:31:08generate a problem statement to predict
- 33:31:11if a customer will repay loan amount or
- 33:31:13not. And then we're going to be using
- 33:31:14the decision tree algorithm in Python.
- 33:31:16Let's see what that looks like. And
- 33:31:18let's dive into the code. In our first
- 33:31:20few steps of implementation, we're going
- 33:31:22to start by importing the necessary
- 33:31:24packages that we need from Python. and
- 33:31:26we're going to load up our data and take
- 33:31:28a look at what the data looks like. So,
- 33:31:29the first thing I need is I need
- 33:31:31something to edit my Python and run it
- 33:31:33in. So, let's flip on over. And here I'm
- 33:31:35using the Anaconda Jupiter notebook.
- 33:31:39Now, you can use any Python IDE you like
- 33:31:41to run it in, but I find the Jupyter
- 33:31:43Notebook's really nice for doing things
- 33:31:44on the fly. And let's go ahead and just
- 33:31:46paste that code in the beginning. And
- 33:31:48before we start, let's talk a little bit
- 33:31:50about what we're bringing in. And then
- 33:31:52we're going to do a couple things in
- 33:31:53here. where I have to make a couple
- 33:31:54changes as we go through this first part
- 33:31:56of the import. The first thing we bring
- 33:31:57in is numpy as np. That's very standard
- 33:32:01when we're dealing with mathematics,
- 33:32:03especially with uh very complicated
- 33:32:05machine learning tools. You'll almost
- 33:32:06always see the numpy come in for your
- 33:32:08num your numbers. It's called number
- 33:32:10python. It has your mathematics in
- 33:32:12there. In this case, we actually could
- 33:32:14take it out, but generally you'll need
- 33:32:15it for most of your different things you
- 33:32:17work with. And then we're going to use
- 33:32:18pandas as pd. That's also a standard.
- 33:32:21The pandas is a dataf frame setup and
- 33:32:24you can liken this to uh taking your
- 33:32:26basic data and storing it in a way that
- 33:32:28looks like an Excel spreadsheet. So as
- 33:32:31we come back to this when you see np or
- 33:32:33pd those are very standard uses you'll
- 33:32:35know that that's the pandas and I'll
- 33:32:37show you a little bit more when we
- 33:32:38explore the data in just a minute. Then
- 33:32:40we're going to need to split the data.
- 33:32:41So I'm going to bring in our train test
- 33:32:43and split and this is coming from the
- 33:32:45sklearn package cross validation. In
- 33:32:48just a minute, we're going to change
- 33:32:50that and we'll go over that, too. And
- 33:32:52then there's also the sktree import
- 33:32:54decision tree classifier. That's the
- 33:32:56actual tool we're using. Remember, I
- 33:32:57told you don't be afraid of the
- 33:32:58mathematics. It's going to be done for
- 33:33:00you. Well, the decision tree classifier
- 33:33:02has all that mathematics in there for
- 33:33:03you, so you don't have to figure it back
- 33:33:05out again. And then we have
- 33:33:06sklearn.metrics
- 33:33:08for accuracy score. We need to score our
- 33:33:10our setup. That's the whole reason we're
- 33:33:12splitting it between the training and
- 33:33:13testing data. And finally, we still need
- 33:33:15the sklearn import tree. And that's just
- 33:33:18the basic tree function that's needed
- 33:33:20for the decision tree classifier. And
- 33:33:22finally, we're going to load our data
- 33:33:23down here. And I'm going to run this and
- 33:33:25we're going to get two things on here.
- 33:33:27One, we're going to get an error. And
- 33:33:29two, we're going to get a warning. Let's
- 33:33:30see what that looks like. So the first
- 33:33:32thing we had is we have an error. Why is
- 33:33:35this error here? Well, it's looking at
- 33:33:36this. It says I need to read a file. And
- 33:33:38when this was written, the person who
- 33:33:41wrote it, this is their path where they
- 33:33:43stored the file. So let's go ahead and
- 33:33:45fix that.
- 33:33:48And I'm going to put in here my file
- 33:33:50path. I'm just going to call it full
- 33:33:52file name. And you'll see it's on my C
- 33:33:54drive. And there's this very lengthy
- 33:33:56setup on here where I stored the data
- 33:33:582.csv file.
- 33:34:01Don't worry too much about the full path
- 33:34:02because on your computer it'll be
- 33:34:04different. The data.2 CSV file was
- 33:34:08generated by SimplyLearn. If you want a
- 33:34:11copy of that, you can comment down below
- 33:34:13and request it here in the YouTube.
- 33:34:16And then if I'm going to give it a name,
- 33:34:18full file name, I'm going to go ahead
- 33:34:21and change it here to full
- 33:34:25file name. So let's go ahead and run it
- 33:34:27now and see what happens.
- 33:34:32And we get a warning
- 33:34:37when you're coding. Understanding these
- 33:34:39different warnings and these different
- 33:34:41errors that come up is probably the
- 33:34:43hardest lesson to learn. So let's just
- 33:34:45go ahead and take a look at this and use
- 33:34:47this as a uh opportunity to understand
- 33:34:49what's going on here. If you read the
- 33:34:52warning, it says the cross validation is
- 33:34:55depreciated. So it's a warning on it's
- 33:34:57being removed and it's going to be moved
- 33:34:59in favor of the model selection. So if
- 33:35:02we go up here, we have
- 33:35:03sklearn.crossvalidation.
- 33:35:06And if you research this and go to
- 33:35:07sklearn site, you'll find out that you
- 33:35:10can actually just swap it right in there
- 33:35:11with model selection.
- 33:35:15And so when I come in here and I run it
- 33:35:17again, that removes a warning. What
- 33:35:20they've done is they've had two
- 33:35:22different developers develop it in two
- 33:35:24different branches and then they decided
- 33:35:26to keep one of those and eventually get
- 33:35:28rid of the other one. That's all that is
- 33:35:31and very easy and quick to fix.
- 33:35:34Before we go any further, I went ahead
- 33:35:36and opened up the data from this file.
- 33:35:40Remember the the data file we just
- 33:35:41loaded on here, the data_2.c
- 33:35:44CSV. Let's talk a little bit more about
- 33:35:46that and see what that looks like both
- 33:35:47as a text file because it's a
- 33:35:49commaepparated variable file and in a
- 33:35:52spreadsheet. This is what it looks like
- 33:35:54as a basic text file. You can see at the
- 33:35:56top they've created a header and it's
- 33:35:58got 1 2 3 4 five columns and each column
- 33:36:01has data in it. And let me flip this
- 33:36:03over cuz we're also going to look at
- 33:36:04this uh in an actual spreadsheet so you
- 33:36:06can see what that looks like. And here
- 33:36:08I've opened it up in the open office
- 33:36:10calc, which is pretty much the same as
- 33:36:11um Excel and zoomed in. And you can see
- 33:36:14we've got our columns and our rows of
- 33:36:16data. A little easier to read in here.
- 33:36:18We have a result, yes, yes, no. We have
- 33:36:20initial payment, last payment, credit
- 33:36:22score, house number. If we scroll way
- 33:36:26down,
- 33:36:28we'll see that this occupies a 101 lines
- 33:36:31of code or lines of data with uh the
- 33:36:34first one being a column and then 1,000
- 33:36:36lines of data.
- 33:36:41Now, as a programmer,
- 33:36:43if you're looking at a small amount of
- 33:36:44data, I usually start by pulling it up
- 33:36:46in different sources so I can see what
- 33:36:48I'm working with.
- 33:36:50But in larger data, you won't have that
- 33:36:52option. it would just be um too too
- 33:36:54large. So you need to either bring in a
- 33:36:56small amount that you can look at it
- 33:36:57like we're doing right now or we can
- 33:36:59start looking at it through the Python
- 33:37:01code. So let's go ahead and move on and
- 33:37:03take the next couple steps to explore
- 33:37:05the data using Python. Let's go ahead
- 33:37:07and see what it looks like in Python to
- 33:37:10print the length and the shape of the
- 33:37:12data. So let's start by printing the
- 33:37:14length of the database. We can use a
- 33:37:16simple lin function from Python. And
- 33:37:19when I run this, you'll see that it's a
- 33:37:22thousand long. And that's what we
- 33:37:23expected. There's a thousand lines of
- 33:37:25data in there. If you subtract the
- 33:37:27column head, and this is one of the nice
- 33:37:29things when we did the uh balance data
- 33:37:31from the panda read CSV, you'll see that
- 33:37:35the header is row zero. So, it
- 33:37:37automatically removes a row and then
- 33:37:40shows the data separate. It does a good
- 33:37:42job sorting that data out for us. And
- 33:37:45then we can use a different function.
- 33:37:47And let's take a look at that. And
- 33:37:49again, we're going to utilize the tools
- 33:37:51in Panda.
- 33:37:53And since the balance data was loaded as
- 33:37:55a Panda data frame,
- 33:37:58we can do a shape on it. And let's go
- 33:38:00ahead and run the shape and see what
- 33:38:02that looks like.
- 33:38:04What's nice about the shape is not only
- 33:38:06does it give me the length of the data,
- 33:38:07we have a th00and lines, it also tells
- 33:38:09me there's five columns. So when we were
- 33:38:11looking at the data, we had five columns
- 33:38:13of data. And then let's take one more
- 33:38:15step to explore the data using Python.
- 33:38:18And now that we've taken a look at the
- 33:38:19length and the shape, let's go ahead and
- 33:38:22use the uh pandas module for head.
- 33:38:25Another beautiful thing in the data set
- 33:38:27that we can utilize. So let's put that
- 33:38:29on our sheet here. And we have print
- 33:38:31data set and balance data.head.
- 33:38:35And this is a pandas print statement of
- 33:38:37its own. So it has its own print feature
- 33:38:39in there. And then we went ahead and
- 33:38:41gave a label for our print job here of
- 33:38:43data set. Just a simple print statement.
- 33:38:45And we run that. And let's just take a
- 33:38:48closer look at that. Let me zoom in
- 33:38:49here.
- 33:38:51There we go.
- 33:38:53Pandas does such a wonderful job of
- 33:38:55making this a very clean readable data
- 33:38:59set. So you can look at the data, you
- 33:39:01can look at the column headers, you can
- 33:39:02have it uh when you put it as the head,
- 33:39:05it prints the first five lines of the
- 33:39:07data. And we always start with zero. So
- 33:39:09we have five lines. We have 0 1 2 3 4
- 33:39:12instead of 1 2 3 4 5. That's a standard
- 33:39:15scripting and programming set is you
- 33:39:17want to start with the zero position.
- 33:39:19And that is what the data head does. It
- 33:39:20pulls the first five rows of data. Puts
- 33:39:22it in a nice format that you can look at
- 33:39:24and view. Very powerful tool to view the
- 33:39:27data. So instead of having to flip and
- 33:39:29open up an Excel spreadsheet or open
- 33:39:31Office Cal or trying to look at a word
- 33:39:34doc where it's all scrunched together
- 33:39:36and hard to read, you can now get a nice
- 33:39:38open view of what you're working with.
- 33:39:40We're working with a shape of a thousand
- 33:39:42long, five wide. So we have five columns
- 33:39:45and we do the full data head. You can
- 33:39:47actually see what this data looks like.
- 33:39:48The initial payment, last payment,
- 33:39:50credit scores, house number. So let's
- 33:39:52take this now that we've explored the
- 33:39:54data and let's start digging into the
- 33:39:56decision tree. So in our next step,
- 33:39:59we're going to train and build our data
- 33:40:02tree. And to do that, we need to first
- 33:40:05separate the data out. We're going to
- 33:40:06separate into two groups so that we have
- 33:40:08something to actually train the data
- 33:40:10with. And then we have some data on the
- 33:40:12side to test it to see how good our
- 33:40:14model is. Remember with any of the
- 33:40:15machine learning, you always want to
- 33:40:17have some kind of test set to to weigh
- 33:40:19it against so you know how good your
- 33:40:20model is when you distribute it. Let's
- 33:40:22go ahead and break this code down and
- 33:40:24look at it in pieces. So first we have
- 33:40:27our X and Y.
- 33:40:30Where do X and Y come from? Well, X is
- 33:40:32going to be our data and Y is going to
- 33:40:35be the answer or the target. You can
- 33:40:37look at it source and target. In this
- 33:40:39case, we're using X and Y to denote the
- 33:40:41data in and the data that we're actually
- 33:40:43trying to guess what the answer is going
- 33:40:45to be. And so to separate it, we can
- 33:40:46simply put in X equals the balance of
- 33:40:49the data values. The first brackets
- 33:40:53means that we're going to select all the
- 33:40:55lines in the database. So, it's all the
- 33:40:57data. And the second one says we're only
- 33:41:00going to look at columns 1 through five.
- 33:41:02Remember, always start with zero. Zero
- 33:41:04is a yes or no. And that's whether the
- 33:41:06loan went default or not. So, we want to
- 33:41:08start with one. If we go back up here,
- 33:41:10that's the initial payment and it goes
- 33:41:12all the way through the house number.
- 33:41:15Well, if we want to look at uh 1 through
- 33:41:17five, we can do the same thing for y,
- 33:41:20which is the answers. And we're going to
- 33:41:22set that just equal to the zero row. So,
- 33:41:25it's just the zero row and then it's all
- 33:41:27rows going in there. So, now we've
- 33:41:28divided this into two different data
- 33:41:31sets. One of them with the
- 33:41:34data going in and one with the answers.
- 33:41:40Next, we need to split the data.
- 33:41:43And here you'll see that we have it
- 33:41:45split into four different parts. The
- 33:41:48first one is your X training, your X
- 33:41:51test, your Y train, your Y test.
- 33:41:56Simply put, we have X going in where
- 33:41:58we're going to train it and we have to
- 33:42:00know the answer to train it with. And
- 33:42:02then we have X test where we're going to
- 33:42:04test that data and we have to know in
- 33:42:07the end what the Y was supposed to be.
- 33:42:10And that's where this train test split
- 33:42:12comes in that we loaded earlier in the
- 33:42:14modules. This does it all for us. And
- 33:42:16you can see they set the test size equal
- 33:42:18to.3. So that's roughly 30% will be used
- 33:42:21in the test. And then we use a random
- 33:42:22state. So it's completely random which
- 33:42:24rows it takes out of there. And then
- 33:42:26finally we get to actually build our
- 33:42:28decision tree. And they've called it
- 33:42:29here CLF entropy. That's the actual
- 33:42:33decision tree or decision tree
- 33:42:34classifier. And in here, they've added a
- 33:42:37couple variables which we'll explore in
- 33:42:39just a minute. And then finally, we need
- 33:42:41to fit the data to that. So, we take our
- 33:42:44CLF entropy that we created and we fit
- 33:42:46the X train. And since we know the
- 33:42:48answers for X-ray or the Y train, we go
- 33:42:50ahead and put those in. And let's go
- 33:42:52ahead and run this. And what most of
- 33:42:54these sklearn modules do is when you set
- 33:42:57up the variable, in this case, when we
- 33:42:59set the CLF entropy equal decision tree
- 33:43:01classifier, it automatically prints out
- 33:43:03what's in that decision tree. There's a
- 33:43:05lot of variables you can play with in
- 33:43:06here. And it's quite beyond the scope of
- 33:43:08this tutorial to go through all of these
- 33:43:11and how they work. But we're working on
- 33:43:13entropy. That's one of the options.
- 33:43:14We've added that it's completely a
- 33:43:16random state of 100, so 100%. And we
- 33:43:19have a max depth of three. Now, the max
- 33:43:22depth, if you remember above when we
- 33:43:23were doing the different graphs of
- 33:43:25animals, means it's only going to go
- 33:43:27down three layers before it stops. And
- 33:43:30then we have minimal samples of leaves
- 33:43:31is five. So, it's going to have at least
- 33:43:33five leaves at the end. So, I'll have at
- 33:43:35least three splits or have no more than
- 33:43:37three layers and at least five end
- 33:43:40leaves with the final result at the
- 33:43:42bottom. Now that we've created our
- 33:43:45decision tree classifier, not only
- 33:43:47created it, but trained it, let's go
- 33:43:49ahead and apply it and see what that
- 33:43:51looks like. So, let's go ahead and make
- 33:43:53a prediction and see what that looks
- 33:43:55like. We're going to paste our predict
- 33:43:57code in here. And before we run it,
- 33:43:59let's just take a quick look at what's
- 33:44:01doing here. We have a variable y predict
- 33:44:04that we're going to do. And we're going
- 33:44:06to use our variable CLF entropy that we
- 33:44:10created.
- 33:44:12And then you'll see predict. And it's
- 33:44:14very common in the sklearn modules that
- 33:44:17their different tools have the predict
- 33:44:19when you're actually running a
- 33:44:20prediction. In this case, we're going to
- 33:44:22put our X test data in here. Now, if you
- 33:44:26delivered this for use, an actual
- 33:44:28commercial use, and distributed it, this
- 33:44:31would be the new loans you're putting in
- 33:44:33here to guess whether the person's going
- 33:44:36to be uh pay them back or not. In this
- 33:44:38case though, we need to test out the
- 33:44:40data and just see how good our sample
- 33:44:42is, how good of our tree does at
- 33:44:45predicting the loan payments. And
- 33:44:46finally, since Anaconda Jupyter notebook
- 33:44:49is works as a command line for Python,
- 33:44:51we can simply put the y predict en to
- 33:44:54print it. I could just as easily have
- 33:44:56put the print
- 33:44:58and put brackets around y predict en to
- 33:45:01print it out. We'll go ahead and do
- 33:45:02that. It doesn't matter which way you do
- 33:45:03it. And you'll see right here that it
- 33:45:06runs a prediction. This is roughly 300
- 33:45:09in here. Remember, it's 30% of a
- 33:45:11thousand. So, you should have about 300
- 33:45:13answers in here. And this tells you
- 33:45:16which each one of those lines of ourh
- 33:45:18test went in there. And this is what our
- 33:45:20y predict came out. So, let's move on to
- 33:45:23the next step where we're going to take
- 33:45:25this data and try to figure out just how
- 33:45:27good a model we have. So, here we go.
- 33:45:29Since sklearn does all the heavy lifting
- 33:45:31for you and all the math, we have a
- 33:45:33simple line of code to let us know what
- 33:45:35the accuracy is. And let's go ahead and
- 33:45:37go through that and see what that means
- 33:45:39and what that looks like. Let's go ahead
- 33:45:40and paste this in. And let me zoom in a
- 33:45:42little bit. There we go.
- 33:45:45So you have a nice full picture. And
- 33:45:47we'll see here. We're just going to do a
- 33:45:48print accuracy is.
- 33:45:51And then we do the accuracy score. And
- 33:45:54this was something we imported um
- 33:45:56earlier. If you remember at the very
- 33:45:58beginning, let me just scroll up there
- 33:45:59real quick so you can see where that's
- 33:46:01coming from. That's coming from here
- 33:46:03down here from sklearn.metrics metrics
- 33:46:06import accuracy score. And you could
- 33:46:09probably run a script, make your own
- 33:46:10script to do this very easily. How
- 33:46:12accurate is it? How many out of 300 do
- 33:46:15we get right? And so we put in our y
- 33:46:17test. That's the one we ran the predict
- 33:46:19on. And then we put in our y predict en
- 33:46:22that's the answers we got. And we're
- 33:46:24just going to multiply that by 100
- 33:46:26because this is just going to give us an
- 33:46:27answer as a decimal and we want to see
- 33:46:29it as a percentage. And let's run that
- 33:46:31and see what it looks like. And if you
- 33:46:34see here, we got an accuracy of
- 33:46:3593.666667.
- 33:46:38So when we look at the number of loans
- 33:46:40and we look at how good our model fit,
- 33:46:42we can tell people it has about a 93.6
- 33:46:46fitting to it. So just a quick recap on
- 33:46:49that. We now have accuracy set up on
- 33:46:52here. And so we have created a model
- 33:46:54that uses the decision tree algorithm to
- 33:46:56predict whether a customer will repay
- 33:46:57the loan or not. The accuracy of the
- 33:46:59model is about 94.6%.
- 33:47:02The bank can now use this model to
- 33:47:03decide whether it should approve the
- 33:47:05loan request from a particular customer
- 33:47:07or not. And so this information is
- 33:47:09really powerful. We may not be able to
- 33:47:11as individuals understand all these
- 33:47:12numbers because they have thousands of
- 33:47:14numbers that come in, but you can see
- 33:47:16that this is a smart decision for the
- 33:47:17bank to use a tool like this to help
- 33:47:20them to predict how good their uh
- 33:47:22profit's going to be off of the loan
- 33:47:24balances and how many are going to
- 33:47:25default or not. We're going to be
- 33:47:27looking at random forest, one of the
- 33:47:28many powerful tools in the machine
- 33:47:30learning library. Before we dive into
- 33:47:33the topic, let's start by looking at a
- 33:47:35few of the uses for random forest.
- 33:47:38Currently today, it's used in remote
- 33:47:40sensing. Uh for example, they're used in
- 33:47:43the ETM devices. If you're a space buff,
- 33:47:46that's the enhanced thermatic mapper
- 33:47:48they use on satellites which see uh far
- 33:47:51outside the human spectrum for looking
- 33:47:53at land masses. and they acquire images
- 33:47:55of the earth's surface. The accuracy is
- 33:47:57higher and training time is less than
- 33:48:00many other machine learning tools out
- 33:48:01there. Also, object detection,
- 33:48:03multiclass object detection is done
- 33:48:05using random forest algorithms. A good
- 33:48:07example is a traffic where you're trying
- 33:48:09to sort out the different cars, buses,
- 33:48:11and things. And it provides a better
- 33:48:12detection in complicated environments.
- 33:48:14They're very complicated up there. And
- 33:48:16then we have uh another example connect.
- 33:48:19And let's take a little closer look at
- 33:48:20connect. Connect. They use a random
- 33:48:23forest as part of the game console and
- 33:48:26what it does is it tracks a body
- 33:48:27movements and it recreates it in the
- 33:48:29game and let's see what that looks like.
- 33:48:32Uh we have a user who performs a step.
- 33:48:34In this case it looks like Elvis Presley
- 33:48:36going there that is then recorded so
- 33:48:39that connect registers the movement and
- 33:48:41then it marks the user based on
- 33:48:43accuracy. And it looks like we have uh
- 33:48:46Prince going on this one from Elvis
- 33:48:48Presley to Prince. It's great. Uh so it
- 33:48:51marks user base on the accuracy. If we
- 33:48:53look at that a little closer, we have a
- 33:48:55training set to identify body parts.
- 33:48:57Where are the hands? Where are the feet?
- 33:48:59Uh what's going on with the body? That
- 33:49:02then goes into a random forest
- 33:49:04classifier that learns from it. Once
- 33:49:06we've trained the classifier, it then
- 33:49:09identifies the body parts while the
- 33:49:11person's dancing. It's able to represent
- 33:49:13that in a computer format. And then
- 33:49:16based on that, it scores the game and
- 33:49:18how accurate you are as being Elvis
- 33:49:20Presley or Prince in your dancing. So
- 33:49:22why random forest? It's always important
- 33:49:26to understand why we use this tool over
- 33:49:29the other ones. What are the benefits
- 33:49:31here? And so with the random forest, the
- 33:49:34first one is there's no overfitting. If
- 33:49:36you use of multiple trees, reduce the
- 33:49:39risk of overfitting. Training time is
- 33:49:41less. Overfitting means that we have fit
- 33:49:44the data so close to what we have as our
- 33:49:46sample that we pick up on all the weird
- 33:49:49parts and instead of predicting the
- 33:49:51overall data, you're predicting the
- 33:49:52weird stuff which you don't want. High
- 33:49:55accuracy runs efficiently on large
- 33:49:57database. For large data, it produces
- 33:49:59highly accurate predictions. In today's
- 33:50:02world of uh big data, this is really
- 33:50:04important. And this is probably where it
- 33:50:07really shines. This is where Y random
- 33:50:09forest really comes in. It estimates
- 33:50:11missing data. Data in today's world is
- 33:50:14very messy. So when you have a random
- 33:50:15forest, it can maintain the accuracy
- 33:50:17when a large proportion of the data is
- 33:50:19missing. What that means is if you have
- 33:50:21data that comes in from uh five or six
- 33:50:23different areas and maybe they took one
- 33:50:26set of statistics in one area and they
- 33:50:28took a slightly different set of
- 33:50:30statistics in the other. So they have
- 33:50:31some of the sh same shared data, but one
- 33:50:34is missing like the uh number of
- 33:50:36children in the house if you're doing
- 33:50:37something over demographics. and the
- 33:50:40other one is missing the size of the
- 33:50:42house. It will look at both of those
- 33:50:44separately and build two different trees
- 33:50:46and then it can do a very good job of
- 33:50:48guessing which one fits better even
- 33:50:50though it's missing that data. Let us
- 33:50:52dig deep into the theory of exactly how
- 33:50:55it works. And let's look at what is
- 33:50:58random forest. Random forest or random
- 33:51:01decision forest is a method that
- 33:51:04operates by constructing multiple
- 33:51:05decision trees. The decision of the
- 33:51:08majority of the trees is chosen by the
- 33:51:10random forest as the final decision. And
- 33:51:12let's uh we have some nice graphics
- 33:51:14here. We have a decision tree and they
- 33:51:15actually use a real tree to denote the
- 33:51:17decision tree which I love. And given a
- 33:51:20random some kind of picture of a fruit.
- 33:51:22This decision tree decides that the
- 33:51:24output is it's an apple. And we have a
- 33:51:26decision tree too where we have that
- 33:51:28picture of the fruit goes in and this
- 33:51:30one decides that it's a lemon. And the
- 33:51:31decision three tree gets another image
- 33:51:33and it decides it's an apple. And then
- 33:51:35this all comes together in what they
- 33:51:37call the random forest. And this random
- 33:51:39forest then looks at it and says, "Okay,
- 33:51:41I got two votes for apple, one vote for
- 33:51:43lemon. The majority is apples. So the
- 33:51:46final decision is apples." To understand
- 33:51:49how the random forest works, we first
- 33:51:52need to dig a little deeper and take a
- 33:51:54look at the random forest and the actual
- 33:51:56decision tree and how it builds that
- 33:51:58decision tree. In looking closer at how
- 33:52:00the individual decision trees work,
- 33:52:02we'll go ahead and continue to use the
- 33:52:04fruit example since we're talking about
- 33:52:06trees and forests. A decision tree is a
- 33:52:08treerehaped diagram used to determine a
- 33:52:11course of action. Each branch of the
- 33:52:13tree represents a possible decision,
- 33:52:15occurrence, or reaction. So in here we
- 33:52:17have a bowl of fruit and if you look at
- 33:52:18that it looks like um they switch from
- 33:52:20lemons to oranges. So we have oranges,
- 33:52:23cherries, and apples. And the first
- 33:52:25decision of the decision tree might be
- 33:52:27is a diameter greater than or equal to
- 33:52:29three. And if it says false, it knows
- 33:52:31that they're cherries because everything
- 33:52:32else is bigger than that. So all the
- 33:52:34cherries fall into that decision. So we
- 33:52:36have all that data we're training. We
- 33:52:38can look at that. We know that that's
- 33:52:39what's going to come up. Is the color
- 33:52:41orange? Well, goes, hm, orange or red?
- 33:52:44Well, if it's true, then it comes out as
- 33:52:47the orange. And if it's false, that
- 33:52:49leaves apples. So in this example, it
- 33:52:52sorts out the fruit in the bowl or the
- 33:52:54images of the fruit. A decision tree.
- 33:52:57These are very important terms to know
- 33:52:59because these are very central to
- 33:53:00understanding the decision tree and when
- 33:53:01working with them. The first is entropy.
- 33:53:04Everything on the decision tree and how
- 33:53:06it makes a decision is based on entropy.
- 33:53:09Entropy is a measure of randomness or
- 33:53:11unpredictability in the data set. uh
- 33:53:14then they also have information gain,
- 33:53:17the leaf node, the decision node and the
- 33:53:20root node. We'll cover these other four
- 33:53:22terms as we go down the tree, but let's
- 33:53:25start with entropy. So starting with
- 33:53:28entropy, we have here a high amount of
- 33:53:30randomness. What that means is that
- 33:53:34whatever is coming out of this decision,
- 33:53:35if it was going to guess based on this
- 33:53:38data, it wouldn't be able to tell you
- 33:53:40whether it's a lemon or an apple. it
- 33:53:42would just say it's a fruit. Uh so the
- 33:53:45first thing we want to do is we want to
- 33:53:47split this apart and we take the initial
- 33:53:49data set. We're going to set create a
- 33:53:50data set one and a data set two. We just
- 33:53:53split it in two. And if you look at
- 33:53:54these new data sets after splitting
- 33:53:57them, the entropy of each of those sets
- 33:53:59is much less. So for the first one,
- 33:54:02whatever comes in there, it's going to
- 33:54:04sort that data and it's going to say,
- 33:54:05okay, if this data goes this direction,
- 33:54:07it's probably an apple. And if it goes
- 33:54:09into the other direction, it's probably
- 33:54:11a lemon. So that brings us up to
- 33:54:13information gain. It is the measure of
- 33:54:15decrease in the entropy after the data
- 33:54:17set is split. What that means in here is
- 33:54:20that we've gone from one set which has a
- 33:54:23very high entropy to two lower sets of
- 33:54:26entropy and we've added in the values of
- 33:54:28E1 for the first one and E2 for the
- 33:54:31second two which are much lower. And so
- 33:54:33that information gain is increased
- 33:54:36greatly in this example. And so you can
- 33:54:38find that the information grain simply
- 33:54:40equals uh decision E1 minus E2. As we're
- 33:54:45going down our list of uh definitions,
- 33:54:48we'll look at the leaf node. And the
- 33:54:49leaf node carries the classification or
- 33:54:52the decision. So we look down here to
- 33:54:55the leaf node. We finally get to our set
- 33:54:58one or our set two. When it comes down
- 33:55:00there and it says, "Okay, this object's
- 33:55:02gone into set one." If it's gone into
- 33:55:05set one, it's going to be split by some
- 33:55:08means and we'll either end up with
- 33:55:09apples on the leaf node or a lemon on
- 33:55:12the leaf node. And on the right, it'll
- 33:55:13either be an apple or lemons. Those leaf
- 33:55:16nodes are those final decisions or
- 33:55:18classifications. Uh that's the
- 33:55:20definition of leaf node in here. If
- 33:55:22we're going to have a final leaf where
- 33:55:24we make the decision, we should have a
- 33:55:27name for the nodes above it. And they
- 33:55:29call those decision nodes. A decision
- 33:55:32node. decision node has two or more
- 33:55:34branches and you can see here where we
- 33:55:36have the uh five apples and one lemon
- 33:55:39and in the other case the five lemons
- 33:55:41and one apple. They have to make a
- 33:55:43choice of which tree it goes down based
- 33:55:45on some kind of measurement or
- 33:55:48information given to the tree. And that
- 33:55:51brings us to our last definition. The
- 33:55:53root node, the topmost decision node is
- 33:55:56known as the root node. And this is
- 33:55:58where you have all of your data and you
- 33:56:01have your first decision. it has to make
- 33:56:02or the first split in information. So
- 33:56:05far, we've looked at a very general
- 33:56:07image um with the fruit being split.
- 33:56:10Let's look and see exactly what that
- 33:56:12means to split the data and how do we
- 33:56:14make those decisions on there. Uh let's
- 33:56:17go in there and find out how does a
- 33:56:19decision tree work. So let's try to
- 33:56:22understand this and let's use a simple
- 33:56:24example and we'll stay with the fruit.
- 33:56:27We have a bowl of fruit and so let's
- 33:56:29create a problem statement and the
- 33:56:31problem is we want to classify the
- 33:56:34different types of fruits in the bowl
- 33:56:35based on different features. The data
- 33:56:38set in the bowl is looking quite messy
- 33:56:40and the entropy is high in this case. So
- 33:56:42if this bowl was our decision maker, it
- 33:56:45wouldn't know what choice to make. It
- 33:56:47has so many choices. Which one do you
- 33:56:49pick? Apple, grapes, or lemons. And so
- 33:56:52we look in here. We're going to start
- 33:56:53with a d a training set. So this is our
- 33:56:56data that we're training our data with
- 33:56:58and we have a number of options here. We
- 33:57:00have the color and under the color we
- 33:57:01have red yellow purple uh we have a
- 33:57:04diameter uh 331 331 and we have a label
- 33:57:08apple lemon grapes apple lemon grapes
- 33:57:11and how do we split the data? We have to
- 33:57:13frame the conditions to split the data
- 33:57:15in such a way that the information gain
- 33:57:17is the highest. It's very key to note
- 33:57:20that we're looking for the best gain. We
- 33:57:22don't want to just start sorting out the
- 33:57:23smallest piece in there. We want to
- 33:57:25split it the biggest way we can. And so
- 33:57:27we measure this decrease in entropy.
- 33:57:30That's what they call it, entropy.
- 33:57:31There's our entropy after splitting. And
- 33:57:33now we'll try to choose a condition that
- 33:57:35gives us the highest gain. We will do
- 33:57:37that by splitting the data using each
- 33:57:39condition and checking the gain that we
- 33:57:41get out of them. The conditions that
- 33:57:42give us the highest gain will be used to
- 33:57:44make the first split. So let's take a
- 33:57:46look at these different conditions. We
- 33:57:48have color, we have diameter, and if we
- 33:57:50look underneath that, we have a couple
- 33:57:51different values. is we have diameter
- 33:57:52equals 3, color equals yellow, red,
- 33:57:55diameter equals 1. And when we look at
- 33:57:57that, you'll see over here we have 1 2 3
- 33:58:014 threes. That's a pretty hardy
- 33:58:04selection. So let's say the condition
- 33:58:06gives us a maximum gain of three. So we
- 33:58:09have the most pieces fall into that
- 33:58:12range. So our first split from our
- 33:58:15decision node is we split the data based
- 33:58:18on the diameter. Is it greater than or
- 33:58:20equal to three? If it's not, that's
- 33:58:23false. It goes into the grape bowl. And
- 33:58:25if it's true, it goes into a bowl fold
- 33:58:27of lemon and apples. The entropy after
- 33:58:30splitting has decreased considerably. So
- 33:58:32now we can make two decisions. If you
- 33:58:34look at they're very uh much less chaos
- 33:58:37going on there. This node has already
- 33:58:39attain an entropy value of zero. As you
- 33:58:41can see, there's only one kind of label
- 33:58:43left for this branch. So no further
- 33:58:45splitting is required for this node.
- 33:58:48However, this node on the right is still
- 33:58:51requires a split to decrease the entropy
- 33:58:53further. So, we split the right node
- 33:58:55further based on color. If you look at
- 33:58:57this, if I split it on color, that
- 33:58:59pretty much cuts it right down the
- 33:59:01middle. And it's the only thing we have
- 33:59:02left in our choices of color and
- 33:59:03diameter, too. And if the color is
- 33:59:06yellow, it's going to go to the right
- 33:59:08bowl. And if it's false, it's going to
- 33:59:09go to the left bowl. So, the entropy in
- 33:59:11this case is now zero. So, now we have
- 33:59:14three bowls with zero entropy. There's
- 33:59:16only one type of data in each one of
- 33:59:18those bowls. So, we can predict a lemon
- 33:59:20with 100% accuracy. And we can predict
- 33:59:23the apple also with 100% accuracy along
- 33:59:26with our grapes up there. So, we've
- 33:59:28looked at kind of a basic tree in our
- 33:59:31forest. But what we really want to know
- 33:59:33is how does a random forest work as a
- 33:59:36whole. So to begin our um random forest
- 33:59:40classifier, let's say we already have
- 33:59:42built three trees. And we're going to
- 33:59:44start with the first tree that looks
- 33:59:46like this. Just like we did in the
- 33:59:48example, this tree looks at the
- 33:59:49diameter. If it's greater than or equal
- 33:59:51to three, it's true. Otherwise, it's
- 33:59:53false. So one side goes to the smaller
- 33:59:56diameter, one side goes to larger
- 33:59:58diameter. And if the color is orange,
- 34:00:01it's going to go to the right. True.
- 34:00:02We're using oranges now instead of
- 34:00:04lemons. And if it's red, it's going to
- 34:00:06go to the left. False. We build a second
- 34:00:08tree very similar, but it's split
- 34:00:10differently. Instead of the first one
- 34:00:12being split by a diameter, uh this one
- 34:00:15when they created it, if you look at
- 34:00:16that first bowl, it has a lot of red
- 34:00:18objects. So it says, is the color red?
- 34:00:21Because that's going to bring our
- 34:00:22entropy down the fastest. And so, of
- 34:00:25course, if it's true, it goes to the
- 34:00:26left. If it's false, it goes to the
- 34:00:28right. And then it looks at the shape,
- 34:00:30false or true, and so on and so on. And
- 34:00:33tree three is the diameter equal to one.
- 34:00:36And it came up with this because there's
- 34:00:38a lot of cherries in this bowl. So that
- 34:00:39would be the biggest split on there is
- 34:00:41is the diameter equal to one. That's
- 34:00:43going to drop the entropy the quickest.
- 34:00:45And as you can see, it splits it into
- 34:00:46true. If it goes false, and they've
- 34:00:48added another category, does it grow in
- 34:00:50the summer? And if it's false, it goes
- 34:00:53off to the left. If it's true, it goes
- 34:00:54off to the right. Let's go ahead and
- 34:00:56bring these three trees so you can see
- 34:00:58them all in one image. So this would be
- 34:01:00three completely different trees
- 34:01:02categorizing a fruit. And let's take a
- 34:01:04fruit. Now let's try this. And this
- 34:01:06fruit, if you look at it, we've
- 34:01:08blackened it out. You can't see the
- 34:01:10color on it. So it's missing data.
- 34:01:12Remember one of the things we talked
- 34:01:13about earlier is that a random forest
- 34:01:16works really good if you're missing
- 34:01:18data, if you're missing pieces. So this
- 34:01:20fruit has an image, but maybe it's a
- 34:01:22person had a black and white camera when
- 34:01:24they took the picture. And we're going
- 34:01:25to take a look at this. And it's going
- 34:01:27to have um they put the color in there,
- 34:01:29so ignore the color down there. But the
- 34:01:31diameter equals three. We find out it
- 34:01:33grows in the summer equals yes. And the
- 34:01:35shape is a circle. And if you go to the
- 34:01:37right, you can look at what one of the
- 34:01:39decision trees did. This is the third
- 34:01:41one. Is the diameter greater than equal
- 34:01:43to three? Is a color orange? Well, it
- 34:01:46doesn't really know on this one, but it
- 34:01:48if you look at the value, it' say true,
- 34:01:49and it go to the right. Tree two
- 34:01:51classifies it as cherries. Is a color
- 34:01:54equal red? Is the shape a circle? True.
- 34:01:57It is a circle. So, this would look at
- 34:01:59it and say, "Oh, that's a cherry." And
- 34:02:01then we go to the other classifier and
- 34:02:02it says, "Is the diameter equal one?"
- 34:02:05Well, that's false. Does it grow in the
- 34:02:07summer? True. So, it goes down and looks
- 34:02:09at as oranges. So, how does this random
- 34:02:12forest work? The first one says it's an
- 34:02:14orange. The second one said it was a
- 34:02:16cherry. And the third one says, hm, it's
- 34:02:19an orange. And you can guess that if you
- 34:02:21have two oranges and one says it's a
- 34:02:23cherry, uh, when you add that all
- 34:02:24together, the majority of the vote says
- 34:02:27orange. So, the answer is it's
- 34:02:29classified as an orange, even though we
- 34:02:31didn't know the color and we're missing
- 34:02:32data on it. I don't know about you, but
- 34:02:34I'm getting tired of fruit. So, let's
- 34:02:37switch. And I did promise you we'd start
- 34:02:39looking at a case example and get into
- 34:02:41some Python coding. Today, we're going
- 34:02:43to use the case the iris flower
- 34:02:45analysis.
- 34:02:47This is the exciting part as we roll up
- 34:02:49our sleeves and actually look at some
- 34:02:51Python coding. Before we start the
- 34:02:53Python coding, we need to go ahead and
- 34:02:55create a problem statement. Wonder what
- 34:02:57species of iris do these flowers belong
- 34:02:59to? Let's try to predict the species of
- 34:03:01the flowers using machine learning in
- 34:03:03Python. Let's see how it can be done. So
- 34:03:06here we begin to go ahead and implement
- 34:03:08our Python code. And you'll find that
- 34:03:11the first half of our implementation is
- 34:03:13all about organizing and exploring the
- 34:03:16data coming in. Let's go ahead and take
- 34:03:18this first step, which is loading the
- 34:03:20different modules into Python. And let's
- 34:03:22go ahead and put that in our favorite
- 34:03:24editor, whatever your favorite editor
- 34:03:26is. In this case, I'm going to be using
- 34:03:28the Anaconda Jupiter Notebook, which is
- 34:03:31one of my favorites. Certainly, there's
- 34:03:33Notepad++ and Eclipse and dozens of
- 34:03:36others, or just even using the Python
- 34:03:38terminal window. any of those will work
- 34:03:40just fine to go ahead and explore this
- 34:03:43Python coding. So, here we go. Let's go
- 34:03:45ahead and flip over to our Jupyter
- 34:03:46notebook. And I've already opened up a
- 34:03:49new page for Python 3 code. And I'm just
- 34:03:52going to paste this right in there. And
- 34:03:53let's take a look and see what we're
- 34:03:55bringing into our Python. The first
- 34:03:57thing we're going to do is from the
- 34:03:58sklearn.data sets import load iris. Now,
- 34:04:02this isn't the actual data. So this is
- 34:04:04just the module that allows us to bring
- 34:04:06in the data, the load iris. And the iris
- 34:04:09is so popular. It's been around since
- 34:04:111936 when Ronald Fiser published a paper
- 34:04:14on it. And they're measuring the
- 34:04:16different parts of the flower. And based
- 34:04:18on those measurements, predicting what
- 34:04:19kind of flower it is. And then if we're
- 34:04:21going to do a random forest classifier,
- 34:04:24we need to go ahead and import a random
- 34:04:25forest classifier from the sklearn
- 34:04:27module. So sklearn.semble
- 34:04:30import random forest classifier. And
- 34:04:32then we want to bring in two more
- 34:04:33modules. Um, and these are probably the
- 34:04:36most commonly used modules in Python and
- 34:04:38data science with any of the um, other
- 34:04:42modules that we bring in. And one is
- 34:04:43going to be pandas. We're going to
- 34:04:45import pandas as pd. PD is the common
- 34:04:47term used for pandas. And pandas is
- 34:04:50basically creates a data format for us
- 34:04:53where when you create a pandas data
- 34:04:56frame, it looks like an Excel
- 34:04:58spreadsheet. And you'll see that in a
- 34:05:00minute when we start digging deeper into
- 34:05:01the code. Panda is just wonderful
- 34:05:03because it plays nice with all the other
- 34:05:05modules in there. And then we have
- 34:05:06Numpy, which is our numbers Python. And
- 34:05:09the numbers Python allows us to do
- 34:05:12different mathematical sets on here.
- 34:05:14We'll see right off the bat, we're going
- 34:05:16to take our NP and we're going to go
- 34:05:18ahead and seed the randomness with it
- 34:05:19with zero. So NP.random seed is seeding
- 34:05:22that as zero. This code doesn't actually
- 34:05:24show anything. We're going to go ahead
- 34:05:26and run it because I need to make sure I
- 34:05:28have all those loaded. And then let's
- 34:05:29take a look at the next module on here.
- 34:05:31The next six slides, including this one,
- 34:05:34are all about exploring the data.
- 34:05:36Remember, I told you half of this is
- 34:05:38about looking at the data and getting it
- 34:05:40all set. So, let's go ahead and take
- 34:05:42this code right here, the script, and
- 34:05:44let's get that over into our Jupyter
- 34:05:46notebook. And here we go. We've gone
- 34:05:47ahead and uh run the imports. Now I'm
- 34:05:51going to paste the code down here
- 34:05:54and let's take a look and see what's
- 34:05:55going on. The first thing we're doing is
- 34:05:57we're actually loading the iris data.
- 34:06:00And if you remember up here, we loaded
- 34:06:02the module that tells it how to get the
- 34:06:04iris data. Now we're actually assigning
- 34:06:06that data to the variable iris. And then
- 34:06:08we're going to go ahead and use the df
- 34:06:11to define dataf frame. And that's going
- 34:06:13to equal pd. And if you remember that's
- 34:06:15pandas as pd. So that's our pandas and
- 34:06:19panda dataf frame. And then we're
- 34:06:20looking at iris data and columns equals
- 34:06:24iris feature names. And we're going to
- 34:06:26do the DF head. And let's run this so
- 34:06:29you can understand what's going on here.
- 34:06:32The first thing you want to notice is
- 34:06:34that our DF has created uh what looks
- 34:06:36like an Excel spreadsheet. And in this
- 34:06:39Excel spreadsheet, we have set the
- 34:06:40columns. So up on the top, you can see
- 34:06:42the four different columns. And then we
- 34:06:45have the data iris.data down below. It's
- 34:06:47a little confusing without knowing where
- 34:06:49this data is coming from. So let's look
- 34:06:51at the bigger picture and I'm going to
- 34:06:53go print. I'm just going to change this
- 34:06:55for a moment and we're going to print
- 34:06:57all of Iris and see what that looks
- 34:06:59like. So when I print all of Iris I get
- 34:07:02this long list of information. And you
- 34:07:05can scroll through here and see all the
- 34:07:07different titles on there. What's
- 34:07:10important to notice is that first off
- 34:07:11there's a brackets at the beginning. So
- 34:07:13this is a Python dictionary
- 34:07:16and in a Python dictionary you'll have a
- 34:07:19key or a label and this label pulls up
- 34:07:23whatever information comes after it. So
- 34:07:25feature names which we actually used
- 34:07:27over here under columns is equal to an
- 34:07:30array of sele length sele width pedal
- 34:07:32length pedal width. These are the
- 34:07:34different names they have for the four
- 34:07:36different columns. And if you scroll
- 34:07:38down far enough you'll also see data
- 34:07:40down here. Oh goodness, it came up right
- 34:07:42towards the top. And uh data is equal to
- 34:07:44the different data we're looking at.
- 34:07:47Now, there's a lot of other things in
- 34:07:48here like target. We're going to be
- 34:07:50pulling that up in a minute. And there's
- 34:07:51also the names uh the target names which
- 34:07:54is further down. And we'll show you that
- 34:07:55also in a minute. Let's go ahead and set
- 34:07:57that back to the head. And this is one
- 34:08:01of the neat features of pandas and panda
- 34:08:04dataf frames is when you do df.ad or the
- 34:08:08panda dataf frame. head. It'll print the
- 34:08:11first five lines of the data set in
- 34:08:14there along with the headers if you have
- 34:08:16them. In this case, we have the column
- 34:08:18headers set to iris features. And in
- 34:08:20here, you'll see that we have 0 1 2 3 4.
- 34:08:24In Python, most arrays always start at
- 34:08:26zero. So, when you look at the first
- 34:08:28five, it's going to be 0 1 2 3 4, not 1
- 34:08:312 3 4 5. So, now we've got our iris data
- 34:08:34imported into a data frame. Let's take a
- 34:08:36look at the next piece of code in here.
- 34:08:38And so in this section here of the code,
- 34:08:41we're going to take a look at the
- 34:08:43target. And let's go ahead and get this
- 34:08:45into our notebook, this piece of code,
- 34:08:47so we can discuss it a little bit more
- 34:08:48in detail. So here we are in our Jupyter
- 34:08:51notebook. I'm going to put the code in
- 34:08:52here. And before I run it, I want to
- 34:08:55look at a couple things going on. So we
- 34:08:57have uh DF species. And this is
- 34:09:00interesting because right here you'll
- 34:09:02see where I have DF species in brackets
- 34:09:05which is uh the key code for creating
- 34:09:07another column. And here we have
- 34:09:09iris.target.
- 34:09:11Now these are both in the pandas setup
- 34:09:13on here. So in pandas we can do either
- 34:09:16one. I could have just as easily done
- 34:09:18iris and then in brackets target
- 34:09:21depending on what I'm working on. Both
- 34:09:23are um acceptable. Let's go ahead and
- 34:09:26run this code and see how this changes.
- 34:09:28And what we've done is we've added the
- 34:09:30target from the iris data set as another
- 34:09:33column on the end.
- 34:09:36Now what species is this is what we're
- 34:09:38trying to predict. So we have our data
- 34:09:40which tells us the answer for all these
- 34:09:42different pieces. And then we've added a
- 34:09:44column with the answer. So that way when
- 34:09:46we do our final setup, we'll have the
- 34:09:48ability to program our our neural
- 34:09:50network to look for these this different
- 34:09:52data and know what a satossa is or a
- 34:09:55veraricolor which we'll see in just a
- 34:09:57minute or virginica. Those are the three
- 34:09:59that are in there. And now we're going
- 34:10:00to add one more column. I know we're
- 34:10:03organizing all this data over and over
- 34:10:05again. It's kind of fun. There's a lot
- 34:10:07of ways to organize it. What's nice
- 34:10:09about putting everything onto one data
- 34:10:12frame is I can then do a print out and
- 34:10:15it shows me exactly what I'm looking at.
- 34:10:16And I'll show you where you where that's
- 34:10:18different where you can alter that and
- 34:10:20do it slightly differently. But let's go
- 34:10:22ahead and put this into our script up to
- 34:10:24now. And here we go. We're going to put
- 34:10:26that down here and we're going to run
- 34:10:29that. And let's talk a little bit about
- 34:10:32what we're doing. Now we're exploring
- 34:10:34data. And one of the challenges is
- 34:10:37knowing how good your model is. Did your
- 34:10:40model work? And to do this, we need to
- 34:10:42split the data. And we split it into two
- 34:10:44different parts. They usually call it
- 34:10:46the training and the testing. And so in
- 34:10:48here, we're going to go ahead and put
- 34:10:50that in our database so you can see it
- 34:10:52clearly. And we've set it df. And
- 34:10:55remember, you can put brackets. This is
- 34:10:56creating another column. Is train. So
- 34:10:58we're going to use part of it for
- 34:10:59training. And this equals np. Remember
- 34:11:01that stands for numpy.random.uniform.
- 34:11:04So we're generating a random number
- 34:11:07between zero and one. And we're going to
- 34:11:09do it for each of the rows. That's where
- 34:11:12the length df comes from. So each row
- 34:11:14gets a generated number. And if it's
- 34:11:16less than 75, it's true. And if it's
- 34:11:19greater than 75, it's false. This means
- 34:11:22we're going to take 75% of the data
- 34:11:26roughly because there's a randomness
- 34:11:27involved. And we're going to use that to
- 34:11:29train it. And then the other 25% we're
- 34:11:32going to hold off to the side and use
- 34:11:33that to test it later on. So let's flip
- 34:11:36back on over and see what the next step
- 34:11:37is. So now that we've labeled our
- 34:11:39database for which is training and which
- 34:11:41is testing, let's go ahead and sort that
- 34:11:44into two different variables, train and
- 34:11:46test. And let's take this code and let's
- 34:11:48bring it into our project. And here we
- 34:11:50go. Let's paste it on down here. And
- 34:11:53before I run this, let's just take a
- 34:11:55quick look at what's going on here. is
- 34:11:58we have up above we created remember
- 34:12:00there's our def head which prints the
- 34:12:02first five rows and we've added a column
- 34:12:04is train at the end and so we're going
- 34:12:06to take that we're going to create two
- 34:12:07variables we're going to create two new
- 34:12:09data frames one's called train one's
- 34:12:12called test 75% in train 25% in test and
- 34:12:18then to sort that out we're going to do
- 34:12:21that by doing df our main original data
- 34:12:24frame with the iris data in it and if df
- 34:12:27F is train equals true, that's going to
- 34:12:30go in the train. And if DF is train
- 34:12:32equals false, it goes in the test. And
- 34:12:35so when I run this, we're going to print
- 34:12:37out the number in each one. Let's see
- 34:12:39what that looks like. And you'll see
- 34:12:41that it puts 118 in the training module
- 34:12:43and it puts 32 in the testing module,
- 34:12:46which lets us know that there was 150
- 34:12:48lines of data in here. So if you went
- 34:12:49and looked at the original data, you
- 34:12:51could see that there's 150 lines and
- 34:12:53that's roughly 75% in one and 25% for us
- 34:12:56to test our model on afterward. So let's
- 34:12:59jump back to our code and see where this
- 34:13:01goes. In the next two steps, we want to
- 34:13:04do one more thing with our data, and
- 34:13:06that's make it readable to humans. Um, I
- 34:13:09don't know about you, but I hate looking
- 34:13:10at zeros and ones. So, let's start with
- 34:13:14the features and let's go ahead and take
- 34:13:17those and make those readable to humans
- 34:13:19and let's put that in our code.
- 34:13:23Let's see. Here we go. Paste it in. And
- 34:13:25you'll see here we've done a couple very
- 34:13:28basic things. We know that the columns
- 34:13:31in our data frame, again, this is a
- 34:13:33panda thing, the DF columns, and we know
- 34:13:37the first four of them, 01, 2, 3, that'd
- 34:13:40be the first four are going to be the
- 34:13:42features or the titles of those columns.
- 34:13:44And so when I run this, you'll see down
- 34:13:47here that it creates an index, sea
- 34:13:49length, sea width, pedal length, and
- 34:13:51pedal width. And this should be familiar
- 34:13:53because if you look up here, here's our
- 34:13:55column titles going across. And here's
- 34:13:57the first four. One thing I want you to
- 34:13:59notice here is that when you're in a
- 34:14:02command line, whether it's Jupyter
- 34:14:03notebook or you're running command line
- 34:14:05in the uh terminal window, if you just
- 34:14:08put the name of it, it'll print it out.
- 34:14:10This is the same as doing print
- 34:14:13features.
- 34:14:15And the shortand is you just put
- 34:14:17features in here. If you're actually
- 34:14:19writing a code and saving the script and
- 34:14:22running it by remote, you really need to
- 34:14:24put the print in there. But for this,
- 34:14:26when I run it, you'll see it gives me
- 34:14:27the same thing.
- 34:14:30But for this, we want to go ahead and
- 34:14:31we'll just leave it as features because
- 34:14:33it doesn't really matter. And this is
- 34:14:34one of the fun thing about Jupyter
- 34:14:36Notebooks is I'm just building the code
- 34:14:37as we go. And then we need to go ahead
- 34:14:39and create the labels for the other
- 34:14:41part. So, let's take a look and see what
- 34:14:42that for. Our final step in prepping our
- 34:14:45data before we actually start running
- 34:14:47the training and the testing is we're
- 34:14:49going to go ahead and convert the
- 34:14:51species on here into something the
- 34:14:53computer understands. So, let's put this
- 34:14:56code into our script and see where that
- 34:14:58takes us.
- 34:15:00All right, here we go. We've set y equal
- 34:15:02to pd.factorize
- 34:15:05train species of zero. So, let's break
- 34:15:09this down just a little bit. We have our
- 34:15:11pandas right here. PD factoriize. What
- 34:15:14is factorized doing? I'm going to come
- 34:15:16back to that in just a second. Let's
- 34:15:18look at what train species is and why
- 34:15:21we're looking at the group zero on
- 34:15:23there. And let's go up here. And here is
- 34:15:26our species.
- 34:15:29Remember this on that? We created this
- 34:15:30whole column here for species. And then
- 34:15:33it has satossa, satossa, satossa,
- 34:15:35satossa. And if you scroll down enough,
- 34:15:37you'd also see virginica and
- 34:15:39veraricolor.
- 34:15:41We need to convert that into something
- 34:15:42the computer understands. Zeros and
- 34:15:44ones. So the train species of zero
- 34:15:48because this is in the format of a of an
- 34:15:51array of arrays. So you have to have the
- 34:15:53zero on the end. And then species is
- 34:15:55just that column. Factoriize goes in
- 34:15:58there and looks at the fact that there's
- 34:16:00only three of them. So when I run this,
- 34:16:02you'll see that Y generates an array
- 34:16:05that's equal to, in this case, it's the
- 34:16:07training set, and it's zeros, ones, and
- 34:16:10twos representing the three different
- 34:16:12kinds of flowers we have. So now we have
- 34:16:14something the computer understands, and
- 34:16:16we have a nice table that we can read
- 34:16:18and understand. And now finally we get
- 34:16:21to actually start doing the predicting.
- 34:16:23So here we go. Uh we have two lines of
- 34:16:27code. Oh my goodness, that was a lot of
- 34:16:29work to get to two lines of code. But
- 34:16:31there is a lot in these two lines of
- 34:16:33code. So let's take a look and see
- 34:16:34what's going on here and put this into
- 34:16:36our full script that we're running. And
- 34:16:39let's paste this in here. And let's take
- 34:16:41a look and see what this is. We have
- 34:16:44we're creating a variable CLF. And we're
- 34:16:46going to set this equal to the random
- 34:16:48forest classifier. And we're passing two
- 34:16:51variables in here. And there's a lot of
- 34:16:52variables you can play with. As far as
- 34:16:55these two are concerned, they're very
- 34:16:56standard. In jobs, all that does is to
- 34:16:59prioritize it. Not something to really
- 34:17:01worry about. Usually when you're doing
- 34:17:03this on your own computer, you do end
- 34:17:05jobs equals 2. If you're working in a
- 34:17:07larger or big data and you need to
- 34:17:09prioritize it differently, this is what
- 34:17:11that number does is it changes your
- 34:17:12priorities and how it's going to run
- 34:17:14across the system and things like that.
- 34:17:16And then the random state is just how it
- 34:17:18starts. Zero is fine for here.
- 34:17:21But uh let's go ahead and run this.
- 34:17:24We also have clf.fit train features, y.
- 34:17:29And before we run it, let's talk about
- 34:17:31this a little bit more. CLF.fit.
- 34:17:35So, we're fitting, we're training it. We
- 34:17:37are actually creating our random forest
- 34:17:40classifier right here. This is the code
- 34:17:43that does everything. And we're going to
- 34:17:44take our training set. Remember, we kept
- 34:17:46our test off to the side. And we're
- 34:17:48going to take our training set with the
- 34:17:50features. And then we're going to go
- 34:17:51ahead and put that in. And here's our
- 34:17:53target, the Y. So, the Y is 0, 1, and
- 34:17:57two that we just created. And the
- 34:17:59features is the actual data going in
- 34:18:02that we put into the training set. And
- 34:18:04let's go ahead and run that.
- 34:18:07And this is kind of an interesting thing
- 34:18:08because it printed out the random force
- 34:18:11classifier
- 34:18:13and everything around it. And so when
- 34:18:16you're running this in your terminal
- 34:18:18window or in a script like this, this
- 34:18:20automatically treats this like just like
- 34:18:22when we were up here and I typed in y
- 34:18:24and it printed out y instead of print y.
- 34:18:27This does the same thing. It treats this
- 34:18:29as a variable and prints it out. But if
- 34:18:32you were actually running your code,
- 34:18:33that wouldn't be the case. And what is
- 34:18:35printed out is it shows us all the
- 34:18:37different variables we can change. And
- 34:18:40if we go down here, you can actually see
- 34:18:41in jobs equals 2. You can see the random
- 34:18:44state equals zero. Those are the two
- 34:18:46that we sent in there. You would really
- 34:18:47have to dig deep to find out all these
- 34:18:50different meanings of all these
- 34:18:51different settings on here. Some of them
- 34:18:53are self-explanatory if you kind of
- 34:18:55think about it a little bit. Like max
- 34:18:56features is auto. So all the features
- 34:18:58that we're putting in there, it's just
- 34:19:00going to automatically take all four of
- 34:19:02them. Whatever we send it, it'll take.
- 34:19:03Some of them might have so many features
- 34:19:05because you're processing words. There
- 34:19:07might be like 1.4 million features in
- 34:19:10there because you're doing legal
- 34:19:11documents and that's how many different
- 34:19:12words are in there. At that point, you
- 34:19:14probably want to limit the maximum
- 34:19:16features that you're going to process.
- 34:19:17And leaf nodes, that's the end nodes.
- 34:19:19Remember, we had the fruit and we're
- 34:19:20talking about the leaf nodes. Like I
- 34:19:22said, there's a lot in this. We're
- 34:19:24looking at a lot of stuff here. So you
- 34:19:25might have uh in this case there's
- 34:19:27probably only think three leaf nodes,
- 34:19:29maybe four. You might have thousands of
- 34:19:31leaf nodes at which point you do need to
- 34:19:32put a cap on that and say, "Okay, you
- 34:19:34can only go so far and then we're going
- 34:19:35to use all of our resources on
- 34:19:37processing this." And that really is
- 34:19:39what most of these are about is limiting
- 34:19:42the process and making sure we don't uh
- 34:19:45overwhelm a system. And there's some
- 34:19:47other settings in here. Again, we're not
- 34:19:48going to go over all of them. Warm start
- 34:19:50equals false. or start as if you're
- 34:19:52programming it one piece at a time
- 34:19:54externally since we're not we're not
- 34:19:56going to have like we're not going to
- 34:19:58continually to train this particular
- 34:19:59learning tree and again like I said
- 34:20:01there's a lot of things in here that
- 34:20:02you'll want to look up more detail from
- 34:20:04the sklearn and if you're digging in
- 34:20:07deep and running a major project on here
- 34:20:09for today though all we need to do is
- 34:20:11fit or train our features and our target
- 34:20:13Y. So now we have our training model.
- 34:20:16What's next? If we're going to create a
- 34:20:18model,
- 34:20:20we now need to test it. Remember, we set
- 34:20:23aside the test features, test group, 25%
- 34:20:27of the data. So let's go ahead and take
- 34:20:28this code and let's put it into our uh
- 34:20:31script and see what that looks like.
- 34:20:33Okay, here we go. And we're going to run
- 34:20:35this.
- 34:20:38And it's going to come out with a bunch
- 34:20:39of zeros, ones, and twos, which
- 34:20:42represents the three type of flowers,
- 34:20:44the satossa, the virginica, and the
- 34:20:45versa color. And what we're putting into
- 34:20:47our predict is the test features. And I
- 34:20:51always kind of like to know what it is I
- 34:20:53am looking at. So, real quick, we're
- 34:20:56going to do test
- 34:20:58features. And remember, features is an
- 34:21:01array
- 34:21:03of sele
- 34:21:05width, pedal length, pedal width. So
- 34:21:07when we put it in this way, it actually
- 34:21:09loads all these different columns that
- 34:21:11we loaded into features. So if we did
- 34:21:13just features, let me just do features
- 34:21:15in here so you can see what features
- 34:21:16looks like. This is just playing with
- 34:21:18the with Panda's data frames. You'll see
- 34:21:21that it's an index. So when you put an
- 34:21:23index in like this
- 34:21:27into test features into test, it then
- 34:21:31takes those columns and creates a Panda
- 34:21:34data frames from those columns. And in
- 34:21:36this case, we're going to go ahead and
- 34:21:39put those into our predict. So, we're
- 34:21:42going to put each one of these lines of
- 34:21:43data, the 5.0, 3.4, 1.5, point2, and
- 34:21:48we're going to put those in, and we're
- 34:21:49going to predict what our new um forest
- 34:21:53classifier is going to come up with. And
- 34:21:55this is what it predicts. It predicts uh
- 34:21:570000121122.
- 34:22:00and and uh again this is the flower type
- 34:22:04satossa vica and versa color. So now
- 34:22:07that we've taken our test features let's
- 34:22:10explore that. Let's see exactly what
- 34:22:12that data means to us. So the first
- 34:22:14thing we can do with our predicts is we
- 34:22:17can actually generate a different
- 34:22:19prediction model. When I say different,
- 34:22:21we're going to view it differently. It's
- 34:22:23not that the data itself is different.
- 34:22:25So let's take this next piece of code
- 34:22:26and put it into our script.
- 34:22:29So we're pasting it in here and you'll
- 34:22:31see that we're doing uh predict and
- 34:22:33we've added underscore proba for
- 34:22:36probability. So there's our clff.predict
- 34:22:39probability. So we're we're running it
- 34:22:41just like we ran it up here, but this
- 34:22:43time with this we're going to get a
- 34:22:45slightly different result and we're only
- 34:22:47going to look at the first 10. So you'll
- 34:22:50see down here instead of looking at all
- 34:22:51of them uh which was uh what 27 you'll
- 34:22:54see right down here that this generates
- 34:22:56a much larger field on the probability
- 34:22:59and let's take a look and see what that
- 34:23:00looks like and what that means. So when
- 34:23:04we do the predict underscore probaba for
- 34:23:07probability it generates three numbers.
- 34:23:10So we had three leaf nodes at the end
- 34:23:12and if you remember from all the theory
- 34:23:14we did this is the predictors. The first
- 34:23:17one is predicting a one for satossa. It
- 34:23:21predicts a zero for virginica. And it
- 34:23:24predicts a zero for versol. And so on
- 34:23:27and so on and so on. And let's um you
- 34:23:29know what? I'm going to change this just
- 34:23:30a little bit. Let's look at 10
- 34:23:33to 20 just because we can.
- 34:23:37And we start to get in a little
- 34:23:38different of data. And you'll see right
- 34:23:40down here it gets to this one. This line
- 34:23:42right here. And this line has zero 0.5
- 34:23:460.5.
- 34:23:48And so if we're going to vote and we
- 34:23:50have two equal votes, it's going to go
- 34:23:51with the first one. So it says uh
- 34:23:53Satossa gets zero votes, virginica
- 34:23:56gets.5 votes, VersaColor gets.5 votes,
- 34:23:59but let's just go with the virginica
- 34:24:02since these two are equal and so on and
- 34:24:04so on down the list. You can see how
- 34:24:05they vary on here. So now we've looked
- 34:24:07at both how to do a basic predict of the
- 34:24:09features and we've looked at the predict
- 34:24:12probability. Let's see what's next on
- 34:24:14here. So now we want to go ahead and
- 34:24:17start mapping names for the plants. We
- 34:24:19want to attach names so that it makes a
- 34:24:21little more sense for us. And that's
- 34:24:23what we're going to do in these next two
- 34:24:24steps. We're going to start by setting
- 34:24:27up our predictions and mapping them to
- 34:24:30the name. So let's see what that looks
- 34:24:32like. And let's go ahead and paste that
- 34:24:35code in here and run it. And this goes
- 34:24:37along with the next piece of code. So
- 34:24:39we'll skip through this quickly and then
- 34:24:40come back to it a little bit. So, here's
- 34:24:42iris.target
- 34:24:44names.
- 34:24:47And uh if you remember correctly, this
- 34:24:49was the the names that we've been
- 34:24:51talking about this whole time, the
- 34:24:52Satossa, Vica, VersaColor. And then
- 34:24:55we're going to go ahead and do the
- 34:24:56prediction again. We've run it. We could
- 34:24:58have just set a variable equal to this
- 34:25:00instead of rerunning it each time, but
- 34:25:01we're going ahead and run it again.
- 34:25:02CLF.predict test features. Remember that
- 34:25:06returns the zeros, the ones, and the
- 34:25:07twos. And then we're going to set that
- 34:25:09equal to predictions. So this time we're
- 34:25:12actually putting it in a variable. And
- 34:25:14when I run this,
- 34:25:16it distributes and it comes out as an
- 34:25:18array. And the array is satossa,
- 34:25:20satossa, satossa, satossa, satossa.
- 34:25:22We're only looking at the first five. We
- 34:25:24could actually do let's do the first 25
- 34:25:27just so we can see a little bit more on
- 34:25:28there. And you'll see that it starts
- 34:25:30mapping it to all the different flower
- 34:25:32types, the versa color and the virginica
- 34:25:34in there. And let's see how this goes
- 34:25:36with the next one. So, let's take a look
- 34:25:38at the top part of our species in here.
- 34:25:41And we'll take this code and put it in
- 34:25:43our script.
- 34:25:45And let's put that down here and paste
- 34:25:47it. There we go. And we'll go ahead and
- 34:25:49run it. And let's talk about both these
- 34:25:51sections of code here and how they go
- 34:25:54together. The first one is our
- 34:25:57predictions. And I went ahead and did uh
- 34:25:59predictions through 25. Let's just do
- 34:26:01five.
- 34:26:03And so we have stosis, satossis, stosis,
- 34:26:05satossis. That's what we're predicting
- 34:26:06from our test model. And then we come
- 34:26:09down here and we look at test species.
- 34:26:12And remember, I could have just done
- 34:26:13test.species.head.
- 34:26:15And you'll see it says Satossa, Satossa,
- 34:26:17Satossa, Satossa. And they match. So the
- 34:26:20first one is what our forest is doing
- 34:26:24and the second one is what the actual
- 34:26:27data is. Now is we need to combine these
- 34:26:30so that we can understand what that
- 34:26:31means. We need to know how good our
- 34:26:33forest is, how good it is at predicting
- 34:26:34the features. So that's where we come up
- 34:26:36to the next step, which is lots of fun.
- 34:26:39We're going to use a single line of code
- 34:26:41to combine our predictions and our
- 34:26:43actuals so we have a nice chart to look
- 34:26:46at. And let's go ahead and put that in
- 34:26:47our script in our Jupyter notebook here.
- 34:26:50Let's see. Let's go ahead and paste that
- 34:26:51in. And then I'm going to because I'm on
- 34:26:54the Jupyter notebook, I can do a control
- 34:26:56minus so we can see the whole line
- 34:26:57there.
- 34:26:59There we go. resize it and let's take a
- 34:27:02look and see what's going on here. We're
- 34:27:04going to create in pandas. Remember PD
- 34:27:06stands for pandas and we're doing a
- 34:27:08cross tab. This function takes two sets
- 34:27:11of data and creates a chart out of them.
- 34:27:13So when I run it, you'll get a nice
- 34:27:14chart down here. And we have the
- 34:27:17predicted species.
- 34:27:19So across the top you'll see the satossa
- 34:27:21versus color virginica and the actual
- 34:27:24species satossa versus color virginica.
- 34:27:27And so the way to read this chart and
- 34:27:29let's go ahead and take a look on how to
- 34:27:31read this chart here. When you read this
- 34:27:33chart, you have satossa where they meet,
- 34:27:35you have versolar where they meet, and
- 34:27:37you have virginica where they meet. And
- 34:27:39they're meeting where the actual and the
- 34:27:41predicted agree. So this is the number
- 34:27:44of accurate predictions. So in this
- 34:27:46case, it equals 30. If you add 13 + 5 +
- 34:27:4912, you get 30. And then we notice here
- 34:27:52where it says virginica, but it was
- 34:27:54supposed to be versol. This is
- 34:27:55inaccurate. So now we have two two
- 34:27:58inaccurate predictions and 30 accurate
- 34:28:01predictions. So we'll say that the model
- 34:28:03accuracy is 93. That's just 30 divided
- 34:28:07by 32. And if we multiply it by 100, we
- 34:28:11can say that it is 93% accurate. So we
- 34:28:14have a 93% accuracy with our model. I
- 34:28:18did want to add one more quick thing in
- 34:28:20here on our scripting before we wrap it
- 34:28:22up. So let's flip back on over to my
- 34:28:24script. in here. We're going to take
- 34:28:26this uh line of code from up above. I
- 34:28:29don't know if you remember it, but
- 34:28:30predicts equals the iris.target_names.
- 34:28:34So, we're going to map it to the names
- 34:28:36and we're going to run the prediction.
- 34:28:38And we read it on test features. But,
- 34:28:40you know, we're not just testing it. We
- 34:28:42want to actually deploy it. So, at this
- 34:28:43point, I would go ahead and change this.
- 34:28:46And this is an array of arrays. This is
- 34:28:48really important when you're running
- 34:28:50these to know that. So, you need the
- 34:28:52double brackets. And I could actually
- 34:28:54create data. Maybe let's let's just do
- 34:28:56two flowers. So maybe I'm processing
- 34:28:58more data coming in. And we'll put two
- 34:29:00flowers in here. And then uh I actually
- 34:29:03want to see what the answer is. So let's
- 34:29:06go ahead and type in PRS and print that
- 34:29:08out. And when I run this, you'll see
- 34:29:11that I've now predicted two flowers that
- 34:29:13maybe I measured in my front yard as
- 34:29:15VersaColor and VersaColor.
- 34:29:18Not surprising since I put the same data
- 34:29:20in for each one. This would be the
- 34:29:22actual uh end product going out to be
- 34:29:25used on data that you don't know the
- 34:29:27answer for.
- 34:29:30So that's going to conclude our
- 34:29:31scripting part of this. Introducing
- 34:29:33naive base classifier. Have you ever
- 34:29:36wondered how your mail provider
- 34:29:38implements spam filtering or how online
- 34:29:40news channels perform news text
- 34:29:42classification or how companies perform
- 34:29:44sentimental analysis of their audience
- 34:29:46on social media? All of this and more is
- 34:29:48done through a machine learning
- 34:29:50algorithm called naive bay classifier.
- 34:29:53Welcome to Naive Bay tutorial. My name
- 34:29:56is Richard Kersner. I'm with the
- 34:29:58SimplyLearn team. That's
- 34:29:59www.simplearn.com.
- 34:30:02Get certified get ahead. What's in it
- 34:30:04for you? We'll start with what is naive
- 34:30:07bays? A basic overview of how it works.
- 34:30:10We'll get into naive bays and machine
- 34:30:12learning where it fits in with our other
- 34:30:14machine learning tools. Why do we need
- 34:30:16naive bays and understanding naive bays
- 34:30:18classifier a much more in-depth of how
- 34:30:20the math works in the background?
- 34:30:22Finally, we'll get into the advantages
- 34:30:24of the naive bay classifier in the
- 34:30:26machine learning setup. And then we'll
- 34:30:28roll up our sleeves and do my favorite
- 34:30:30part. We'll actually do some Python
- 34:30:31coding and do some text classification
- 34:30:34using the naive bays. What is naive
- 34:30:36bays? Let's start with a basic
- 34:30:38introduction to the bay theorem named
- 34:30:40after Thomas Bae from the 1700s who
- 34:30:43first coined this in the western
- 34:30:44literature. Naive bay classifier works
- 34:30:47on the principle of conditional
- 34:30:48probability as given by the bay theorem.
- 34:30:50Before we move ahead, let us go through
- 34:30:52some of the simple concepts in the
- 34:30:54probability that we will be using. Let
- 34:30:56us consider the following example of
- 34:30:57tossing two coins. Here we have two
- 34:31:00quarters and if we look at all the
- 34:31:02different possibilities of what they can
- 34:31:03come up as, we get that they could come
- 34:31:05up as head heads. come up as head, tail,
- 34:31:07tail, head and tell tail. When doing the
- 34:31:09math on probability, we usually denote
- 34:31:12probability as a P, a capital P. So the
- 34:31:15probability of getting two heads equals
- 34:31:161/4. You can see in our data set, we
- 34:31:19have two heads and this occurs once out
- 34:31:21of the four possibilities. And then the
- 34:31:23probability of at least one tail occurs
- 34:31:25three/arters of the time. You'll see on
- 34:31:27three of the coin tosses, we have tails
- 34:31:29in them. And out of four, that's
- 34:31:30three/4s. And then the probability of
- 34:31:32the second coin being a head given the
- 34:31:35first coin is tail is 1/2. And the
- 34:31:38probability of getting two heads given
- 34:31:40the first coin is a head is 1/2. We'll
- 34:31:42demonstrate that in just a minute and
- 34:31:44show you how that math works. Now when
- 34:31:45we're doing it with two coins, it's easy
- 34:31:47to see. But when you have something more
- 34:31:49complex, you can see where these pro
- 34:31:50these formulas really come in and work.
- 34:31:53So the base theorem gives us the
- 34:31:55conditional probability of an event A
- 34:31:57given another event B has occurred. In
- 34:32:00this case, the first coin toss will be B
- 34:32:03and the second coin toss A. This could
- 34:32:05be confusing because we've actually
- 34:32:06reversed the order of them and go from B
- 34:32:09to A instead of A to B. You'll see this
- 34:32:11a lot when you work in probabilities.
- 34:32:13The reason is we're looking for event A,
- 34:32:15we want to know what that is. So, we're
- 34:32:17going to label that A since that's our
- 34:32:18focus. And then given another event B
- 34:32:21has occurred. In the Baze theorem, as
- 34:32:23you can see on the left, the probability
- 34:32:25of A occurring given B has occurred
- 34:32:28equals the probability of B occurring
- 34:32:30given A has occurred times the
- 34:32:32probability of A over the probability of
- 34:32:34B. This simple formula can be moved
- 34:32:36around just like any algebra formula.
- 34:32:38And we could do the probability of A
- 34:32:40after given B times probability of B
- 34:32:43equals the probability of B given A
- 34:32:46times probability of A. You can easily
- 34:32:47move that around and multiply it and
- 34:32:49divide it out. Let us apply B theorem to
- 34:32:52our example. Here we have our two
- 34:32:53quarters and we'll notice that the first
- 34:32:55two probabilities of getting two heads
- 34:32:58and at least one tail we compute
- 34:33:00directly off the data. So you can easily
- 34:33:02see that we have one example hh out of
- 34:33:06four 1/4 and we have three with tails in
- 34:33:09them giving us three quarters or 3/4
- 34:33:1175%. The second condition the second uh
- 34:33:15set three and four we're going to
- 34:33:16explore a little bit more in detail.
- 34:33:18Now, we stick to a simple example with
- 34:33:20two coins because you can easily
- 34:33:21understand the math. The probability of
- 34:33:23throwing a tail doesn't matter what
- 34:33:25comes before it. And the same with the
- 34:33:26heads. So, it's still going to be 50% or
- 34:33:291/2. But when that come when that
- 34:33:31probability gets more complicated, let's
- 34:33:32say you have a d6 dice or some other
- 34:33:35instance, then this formula really comes
- 34:33:37in handy. But let's stick to the simple
- 34:33:38example for now. In this sample space,
- 34:33:41let A be the event that the second coin
- 34:33:43is head and b be the event that the
- 34:33:45first coin is tails. Again, we reversed
- 34:33:47it because we want to know what the
- 34:33:48second event's going to be. So, we're
- 34:33:49going to be focusing on A. And we write
- 34:33:51that out as the probability of A given
- 34:33:54B. And we know this from our formula
- 34:33:56that that equals the probability of B
- 34:33:58given A times the probability of A over
- 34:34:00the probability of B. And when we plug
- 34:34:02that in, we plug in the probability of
- 34:34:04the first coin being tails given the
- 34:34:06second coin is heads and the probability
- 34:34:08of the second coin being heads given the
- 34:34:10first coin being over the probability of
- 34:34:12the first coin being tails. When we plug
- 34:34:14that data in and we have the probability
- 34:34:16of the first coin being tails given the
- 34:34:18second coin is heads times the
- 34:34:20probability of the second coin being
- 34:34:22heads over the probability of the first
- 34:34:24coin being tails. You can see it's a
- 34:34:25simple formula to calculate. We have 1/2
- 34:34:28* 1/2 over 1/2 or 1/2 =.5 or 1/4. So the
- 34:34:34B theorem basically calculates the
- 34:34:36conditional probability of the
- 34:34:38occurrence of an event based on prior
- 34:34:40knowledge of conditions that might be
- 34:34:42related to the event. We will explore
- 34:34:44this in detail when we take up an
- 34:34:45example of online shopping further in
- 34:34:47this tutorial. Understanding naive bays
- 34:34:49and machine learning. Like with any of
- 34:34:51our other machine learning tools, it's
- 34:34:53important to understand where the naive
- 34:34:55bays fits in the hierarchy. So under the
- 34:34:57machine learning, we have supervised
- 34:34:59learning and there is other things like
- 34:35:00unsupervised learning. There's also
- 34:35:02reward system. This falls under the
- 34:35:04supervised learning. And then under the
- 34:35:06supervised learning, there's
- 34:35:07classification. There's also regression.
- 34:35:09But we're going to be in the
- 34:35:10classification side. And then under
- 34:35:12classification is your naive bays. Let's
- 34:35:15go ahead and glance into where is naive
- 34:35:18bays used. Let's look at some of the use
- 34:35:20scenarios for it. As a classifier, we
- 34:35:22use it in face recognition. Is this
- 34:35:24Cindy or is it not Cindy or whoever? Or
- 34:35:27it might be used to identify parts of
- 34:35:29the face that they then feed into
- 34:35:30another part of the face recognition
- 34:35:32program. This is the eye. This is the
- 34:35:34nose. This is the mouth. Weather
- 34:35:36prediction. Is it going to be rainy or
- 34:35:37sunny? Medical recognition. News
- 34:35:39prediction. It's also used in medical
- 34:35:41diagnosis. We might diagnose somebody as
- 34:35:44either as high risk or not as high risk
- 34:35:46for cancer or heart disease or other
- 34:35:49ailments. And news classification you
- 34:35:51look at the Google news and it says well
- 34:35:53is this political or is this world news
- 34:35:56or a lot of that's all done with the
- 34:35:58naive bays. Understanding naive bay
- 34:36:01classifier. Now we already went through
- 34:36:03a basic understanding with the coins and
- 34:36:06the two heads and two tails and head
- 34:36:08tail tail heads etc. We're going to do
- 34:36:10just a quick review on that and remind
- 34:36:12you that the naive bay classifier is
- 34:36:14based on the bay theorem which gives a
- 34:36:16conditional probability of event A given
- 34:36:20event B. And that's where the
- 34:36:21probability of A given B equals the
- 34:36:24probability of B given A times
- 34:36:26probability of A over probability of B.
- 34:36:28Remember this is an algebraic function
- 34:36:30so we can move these different entities
- 34:36:32around. We could multiply by the
- 34:36:34probability of B. So it goes to the left
- 34:36:36hand side and then we could divide by
- 34:36:38the probability of A given B and just as
- 34:36:40easily come up with a new formula for
- 34:36:42the probability of B. To me staring at
- 34:36:44these algebraic functions kind of gives
- 34:36:46me a slight headache. It's a lot better
- 34:36:49to see if we can actually understand how
- 34:36:50this data fits together in a table. And
- 34:36:52let's go ahead and start applying it to
- 34:36:54some actual data so you can see what
- 34:36:55that looks like. So, we're going to
- 34:36:57start with the shopping demo problem
- 34:36:59statement. And remember, we're going to
- 34:37:00solve this first in a table form so you
- 34:37:02can see what the math looks like. And
- 34:37:04then we're going to solve it in Python.
- 34:37:06And in here, we want to predict whether
- 34:37:07the person will purchase a product. Are
- 34:37:09they going to buy or don't buy? Very
- 34:37:11important. If you're running a business,
- 34:37:12you want to know how to maximize your
- 34:37:14profits or at least maximize the
- 34:37:16purchase of the people coming into your
- 34:37:17store. And we're going to look at a
- 34:37:19specific combination of different
- 34:37:21variables. In this case, we're going to
- 34:37:23look at the day, the discount, and the
- 34:37:25free delivery. And you can see here
- 34:37:26under the day we want to know whether
- 34:37:28it's uh on the weekday, you know,
- 34:37:29somebody's working, they come in after
- 34:37:31work or maybe they don't work. Weekend,
- 34:37:33you can see the bright colors coming
- 34:37:34down there celebrating not being in work
- 34:37:36or holiday. And did we offer a discount
- 34:37:39that day? Yes or no. Did we offer free
- 34:37:41delivery that day? Yes or no. And from
- 34:37:43this, we want to know whether the
- 34:37:44person's going to buy based on these
- 34:37:45traits so we can maximize them and find
- 34:37:48out the best system for getting somebody
- 34:37:49to come in and purchase our goods and
- 34:37:51products from our store. Now, having a
- 34:37:53nice visual is great, but we do need to
- 34:37:55dig into the data. So, let's go ahead
- 34:37:57and take a look at the data set. We have
- 34:37:59a small sample data set of 30 rows.
- 34:38:01We're showing you the first 15 of those
- 34:38:03rows for this demo. Now, the actual data
- 34:38:05file you can request. Just type in below
- 34:38:08under the comments on the YouTube video
- 34:38:10and we'll send you some more information
- 34:38:11and send you that file. As you can see
- 34:38:13here, the file is very simple columns
- 34:38:15and rows. We have the day, the discount,
- 34:38:18the free delivery, and did the person
- 34:38:20purchase or not. And then we have under
- 34:38:22the day whether it was a weekday, a
- 34:38:24holiday, was it the weekend? This is a
- 34:38:26pretty simple set of data. And long
- 34:38:29before computers, people used to look at
- 34:38:30this data and calculate this all by
- 34:38:32hand. So let's go ahead and walk through
- 34:38:34this and see what that looks like when
- 34:38:36we put that into tables. Also note in
- 34:38:38today's world, we're not usually looking
- 34:38:40at three different variables and 30
- 34:38:42rows. Nowadays, because we're able to
- 34:38:44collect data so much, we're usually
- 34:38:45looking at 27, 30 variables across
- 34:38:48hundreds of rows. The first thing we
- 34:38:51want to do is we're going to take this
- 34:38:52data and uh based on the data set
- 34:38:55containing our three inputs day,
- 34:38:57discount, and free delivery, we're going
- 34:38:58to go ahead and populate that to
- 34:39:00frequency tables for each attribute. So,
- 34:39:02we want to know if they had a discount,
- 34:39:04how many people buy and did not buy. Uh
- 34:39:07did they have a discount? Yes or no. Do
- 34:39:09we have a free delivery? Yes or no. On
- 34:39:11those days, how many people made a
- 34:39:13purchase and how many people didn't? And
- 34:39:14the same with the three days of the
- 34:39:16week. Was it a weekday, a weekend, a
- 34:39:17holiday? And did they buy? Yes or no? As
- 34:39:20we dig in deeper to this table for our
- 34:39:22bay theorem, let the event buy be a. Now
- 34:39:25remember when we looked at the coins, I
- 34:39:26said we really want to know what the
- 34:39:27outcome is. Did the person buy or not?
- 34:39:30And that's usually event A is what
- 34:39:32you're looking for. And the independent
- 34:39:33variables, discount, free delivery, and
- 34:39:35day be B. So we'll call that probability
- 34:39:38of B. Now let us calculate the
- 34:39:40likelihood table for one of the
- 34:39:41variables. Let's start with day, which
- 34:39:44includes weekday, weekend, and holiday.
- 34:39:46And let us start by summing all of our
- 34:39:48rows. So, we have the uh weekday row,
- 34:39:51and out of the weekdays, there's 9 plus
- 34:39:532, so there's 11 weekdays. There's eight
- 34:39:55weekend days and 11 holidays. Wow,
- 34:39:58that's a lot of holidays. And then we
- 34:39:59want to sum up the total number of days.
- 34:40:01So, we're looking at a total of 30 days.
- 34:40:04Let's start pulling some information
- 34:40:06from our chart and see where that takes
- 34:40:08us. And when we fill in the chart on the
- 34:40:10right, you can see that nine out of 24
- 34:40:13purchases are made on the weekday, 7 out
- 34:40:15of 24 purchases on the weekend, and
- 34:40:18eight out of 24 purchases on a holiday.
- 34:40:20And out of all the people who come in,
- 34:40:2224 out of 30 purchase. You can also see
- 34:40:24how many people do not purchase. On the
- 34:40:26weekday, it's two out of six didn't
- 34:40:28purchase and so on and so on. We can
- 34:40:30also look at the totals and you'll see
- 34:40:31on the right, we put together some of
- 34:40:33the formulas. The probability of making
- 34:40:35a purchase on the weekend comes out 11
- 34:40:38out of 30. So out of the 30 people who
- 34:40:40came into the store throughout the
- 34:40:41weekend, weekday and holiday, 11 of
- 34:40:44those purchases were made on the
- 34:40:45weekday. And then you can also see the
- 34:40:47probability of them not making a
- 34:40:49purchase. And this is done for doesn't
- 34:40:52matter which day of the week. So we call
- 34:40:53that probability of no buy would be 6
- 34:40:56over 30 or 0.2. So there's a 20% chance
- 34:40:59that they're not going to make a
- 34:41:00purchase no matter what day of the week
- 34:41:01it is. And finally, we look at the
- 34:41:03probability of B if A. In this case,
- 34:41:06we're going to look at the probability
- 34:41:07of the weekday and not buying. Two of
- 34:41:10the no buys were done out of the weekend
- 34:41:11out of the six people who did not make
- 34:41:13purchases. So when we look at that,
- 34:41:15probability of the week day without a
- 34:41:17purchase is going to be.33 or 33%. Let's
- 34:41:21take a look at this at different
- 34:41:22probabilities. And uh based on this
- 34:41:25likelihood table, let's go ahead and
- 34:41:27calculate conditional probabilities as
- 34:41:28below. The first three we just did. The
- 34:41:31probability of making a purchase on the
- 34:41:32weekday is 11 out of 30 or roughly 36 or
- 34:41:3637%
- 34:41:37367. The probability of not making a
- 34:41:40purchase at all doesn't matter what day
- 34:41:41of the week is roughly.2 or 20%. And the
- 34:41:45probability of a weekday no purchase is
- 34:41:49roughly two out of six. So two out of
- 34:41:51six of our no purchases were made on the
- 34:41:53weekday. And then finally we take our P
- 34:41:56of A. If you looked we've kept the
- 34:41:58symbols up there. So we got P of
- 34:42:00probability of B, probability of A,
- 34:42:01probability of B if A. We should
- 34:42:04remember that the probability of A if B
- 34:42:07is equal to the first one times the
- 34:42:10probability of no per buys over the
- 34:42:13probability of the weekday. So we could
- 34:42:15calculate it both off the uh table we
- 34:42:17created. We can also calculate this by
- 34:42:19the formula and we get the.367
- 34:42:22which equals or.33
- 34:42:24*2 over.367 which equals.179
- 34:42:28or roughly uh 17 to 18%. And that'd be
- 34:42:32the probability of no purchase done on
- 34:42:34the weekday. And this is important
- 34:42:36because we can look at this and say as
- 34:42:38the probability of buying on the weekday
- 34:42:41is more than the probability of not
- 34:42:43buying on the weekday, we can conclude
- 34:42:45that customers will most likely buy the
- 34:42:47product on a weekday. Now, we've kept
- 34:42:49our chart simple and we're only looking
- 34:42:51at one aspect. So, you should be able to
- 34:42:53look at the table and come up with the
- 34:42:54same information or the same conclusion.
- 34:42:56That should be kind of intuitive at this
- 34:42:58point. Next, we can take the same setup.
- 34:43:01We have the frequency tables of all
- 34:43:03three independent variables. Now we can
- 34:43:05construct the likelihood tables for all
- 34:43:07three of the variables we're working
- 34:43:09with. We can take our day like we did
- 34:43:12before. We have weekday, weekend, and
- 34:43:13holiday. And we filled in this table.
- 34:43:15And then we can come in and also do that
- 34:43:16for the discount. Yes or no. Did they
- 34:43:19buy? Yes or no. And we fill in that full
- 34:43:21table. So now we have our probabilities
- 34:43:24for a discount and whether the discount
- 34:43:27leads to a purchase or not. And the
- 34:43:29probability for free delivery. Does that
- 34:43:31lead to a purchase or not? And this is
- 34:43:33where it starts getting really exciting.
- 34:43:34Let us use these three likelihood tables
- 34:43:37to calculate whether a customer will
- 34:43:38purchase a product on a specific
- 34:43:40combination of day, discount, and free
- 34:43:43delivery or not purchase. Here, let us
- 34:43:46take a combination of these factors. Day
- 34:43:48equals holiday, discount equals yes,
- 34:43:50free delivery equals yes. Let's dig
- 34:43:52deeper into the math and actually see
- 34:43:54what this looks like. And we're going to
- 34:43:56start with looking for the probability
- 34:43:58of them not purchasing on the following
- 34:44:01combinations of days. We are actually
- 34:44:03looking for the probability of A equal
- 34:44:05no buy. No purchase. And our probability
- 34:44:08of B we're going to set equal to is it a
- 34:44:10holiday? Did they get a discount? Yes.
- 34:44:12And was it a free delivery? Yes. Before
- 34:44:14we go further, let's look at the
- 34:44:16original equation. the probability of a
- 34:44:19if b equals the probability of b given
- 34:44:22the condition a and the probability
- 34:44:24times probability of a over the
- 34:44:26probability of b occurring. Now this is
- 34:44:28basic algebra so we can multiply this
- 34:44:31information together. So when you see
- 34:44:33the probability of a given b in this
- 34:44:36case the condition is b c and d or the
- 34:44:39three different variables we're looking
- 34:44:41at. And when you see the probability of
- 34:44:43B, that would be the conditions. We're
- 34:44:45actually going to multiply those three
- 34:44:47separate conditions out. Probability of
- 34:44:50you'll see that in just a second in the
- 34:44:51formula times the full probability of A
- 34:44:54over the full probability of B. So here
- 34:44:57we are back to this and we're going to
- 34:44:59have let A equal no purchase. And we're
- 34:45:01looking for the probability of B on the
- 34:45:03condition A where A sets for three
- 34:45:06different things. Remember that equals
- 34:45:08the probability of A given the condition
- 34:45:10B. And in this case, we just multiply
- 34:45:13those three different variables
- 34:45:15together. So we have the probability of
- 34:45:17the discount times the probability of
- 34:45:21free delivery times the probability is
- 34:45:23the day equal a holiday. Those are our
- 34:45:26three variables of the probability of A
- 34:45:28if B. And then that is going to be
- 34:45:30multiplied by the probability of them
- 34:45:32not making a purchase. And then we want
- 34:45:34to divide that by the total
- 34:45:36probabilities and they're multiplied
- 34:45:38together. So we have the probability of
- 34:45:39a discount, the probability of a free
- 34:45:42delivery, and the probability of it
- 34:45:43being on a holiday. When we plug those
- 34:45:45numbers in, we see that one out of six
- 34:45:48were no purchase on a discounted day,
- 34:45:50two out of six were a no purchase on a
- 34:45:53free delivery day, and three out of six
- 34:45:55were a no purchase on a holiday. Those
- 34:45:58are our three probabilities of A of B
- 34:46:01multiplied out. And then that has to be
- 34:46:03multiplied by the probability of a no
- 34:46:04purchase. And remember the prob
- 34:46:06probability of a noby is across all the
- 34:46:08data. So that's where we get the 6 out
- 34:46:10of 30. We divide that out by the
- 34:46:13probability of each category over the
- 34:46:16total number. So we get the 20 out of 30
- 34:46:19had a discount, 23 out of 30 had a yes
- 34:46:22for free delivery, and 11 out of 30 were
- 34:46:25on a holiday. We plug all those numbers
- 34:46:27in, we get.178.
- 34:46:30So in our probability math, we have
- 34:46:32a.178
- 34:46:34if it's a no-by for a holiday, a
- 34:46:36discount, and a free delivery. Let's
- 34:46:38turn that around and see what that looks
- 34:46:40like if we have a purchase. I promise
- 34:46:42this is the last page of math before we
- 34:46:44dig into the Python script. So here
- 34:46:46we're calculating the probability of the
- 34:46:47purchase using the same math we did to
- 34:46:49find out if they didn't buy. Now we want
- 34:46:51to know if they did buy. And again,
- 34:46:52we're going to go by the day equals a
- 34:46:54holiday, discount equals yes, free
- 34:46:56delivery equals yes, and let a equal
- 34:46:58buy. Now, right about now, you might be
- 34:47:00asking, why are we doing both
- 34:47:02calculations? Why why would we want to
- 34:47:04know the no buys and buys for the same
- 34:47:07data going in? Well, we're going to show
- 34:47:09you that in just a moment, but we have
- 34:47:11to have both of those pieces of
- 34:47:12information so that we can figure it out
- 34:47:14as a percentage as opposed to a
- 34:47:16probability equation. And we'll get to
- 34:47:18that normalization here in just a
- 34:47:20moment. Let's go ahead and walk through
- 34:47:21this calculation. And as you can see
- 34:47:23here, the probability of A on the
- 34:47:25condition of B, B being all three
- 34:47:27categories, did we have a discount with
- 34:47:30a purchase, did we have a free delivery
- 34:47:32with a purchase, and did we is a day
- 34:47:34equal to holiday. And when we plug this
- 34:47:36all into that formula and multiply it
- 34:47:38all out, we get our probability of a
- 34:47:40discount, probability of a free
- 34:47:42delivery, probability of the day being a
- 34:47:44holiday times the overall probability of
- 34:47:47it being a purchase divided by again
- 34:47:50multiplying the three variables out. The
- 34:47:52full probability of there being a
- 34:47:53discount, the full probability of being
- 34:47:55a free delivery, and the full
- 34:47:57probability of there being a day equal
- 34:47:58holiday. And that's where we get this 19
- 34:48:00over 24 * 21 over 24 * 8 over 24 * the p
- 34:48:05of a 24 over 30 divided by the
- 34:48:09probability of the discount the free
- 34:48:11delivery times the day or 20 over 30 23
- 34:48:15over 30 * 11 over 30 and that gives us
- 34:48:17our 986.
- 34:48:20So what are we going to do with these
- 34:48:22two pieces of data we just generated?
- 34:48:24Well, let's go ahead and go over them.
- 34:48:25We have a probability of purchase
- 34:48:27equals.986.
- 34:48:29We have a probability of no purchase
- 34:48:31equals.178.
- 34:48:34So finally we have a conditional
- 34:48:36probabilities of purchase on this day.
- 34:48:37Let us take that we're going to
- 34:48:38normalize it and we're going to take
- 34:48:40these probabilities and turn them into
- 34:48:42percentages. This is simply done by
- 34:48:44taking the sum of probabilities which
- 34:48:46equals 98686 plus.178
- 34:48:50and that equals the 1.164.
- 34:48:53If we divide each probability by the
- 34:48:56sum, we get the percentage. And so the
- 34:48:58likelihood of a purchase is 84.71%.
- 34:49:02And the likelihood of no purchase is
- 34:49:0415.29%
- 34:49:06given these three different variables.
- 34:49:09So it's if it's on a holiday, if it's a
- 34:49:11with a discount and has free delivery,
- 34:49:13then there's an 84.71%
- 34:49:15chance that the customer is going to
- 34:49:16come in and make a purchase. Hooray,
- 34:49:18they purchased our stuff. We're making
- 34:49:20money. If you were owning a shop, that's
- 34:49:21like is the bottom line is you want to
- 34:49:23make some money so you can keep your
- 34:49:24shop open and have a living. Now, I
- 34:49:26promised you that we were going to be
- 34:49:27finishing up the math here with a few
- 34:49:29pages. So, we're going to move on and
- 34:49:31we're going to do two steps. The first
- 34:49:33step is I want you to understand why you
- 34:49:36want to why you want to use the naive
- 34:49:38bays. What are the advantages of naive
- 34:49:40bays? And then once we understand those
- 34:49:42advantages, we just look at that
- 34:49:44briefly. Then we're going to dive in and
- 34:49:46do some Python coding. Advantages of
- 34:49:48naive bay classifier. So let's take a
- 34:49:51look at the six advantages of the naive
- 34:49:53bay classifier. And we're going to walk
- 34:49:55around this lovely wheel. Looks like an
- 34:49:57origami folded paper. The first one is
- 34:49:59very simple and easy to implement.
- 34:50:01Certainly you could walk through the
- 34:50:03tables and do this by hand. You got to
- 34:50:05be a little careful because the
- 34:50:06notations can get confusing. You have
- 34:50:08all these different probabilities and I
- 34:50:10certainly mess those up as I put them
- 34:50:11on, you know, is it on the top or the
- 34:50:13bottom? We got to really pay close
- 34:50:14attention to that. When you put it into
- 34:50:16Python, it's really nice because you
- 34:50:18don't have to worry about any of that.
- 34:50:19You let the Python handle that, the
- 34:50:21Python module. But understanding it, you
- 34:50:23can put it on a table and you can easily
- 34:50:24see how it works. And it's a simple
- 34:50:26algebraic function. It needs less
- 34:50:28training data. So if you have smaller
- 34:50:29amounts of data, this is great powerful
- 34:50:31tool for that. Handles both continuous
- 34:50:34and discrete data. It's highly scalable
- 34:50:36with number of predictors and data
- 34:50:38points. So, as you can see, you can just
- 34:50:40keep multiplying different probabilities
- 34:50:42in there and you can cover not just
- 34:50:43three different variables or sets. You
- 34:50:46can now expand this to even more
- 34:50:47categories. Number five, it's fast. It
- 34:50:50can be used in real time predictions.
- 34:50:52This is so important. This is why it's
- 34:50:54used in a lot of our predictions on
- 34:50:56online shopping carts, uh, referrals,
- 34:50:59spam filters, is because there's no time
- 34:51:01delay as it has to go through and figure
- 34:51:03out a neural network or one of the other
- 34:51:06mini setups where you're doing
- 34:51:07classification. And certainly there's a
- 34:51:09lot of other tools out there in the
- 34:51:10machine learning that can handle these,
- 34:51:13but most of them are not as fast as the
- 34:51:15naive bays. And then finally, it's not
- 34:51:17sensitive to irrelevant features. So it
- 34:51:20picks up on your different
- 34:51:21probabilities. And if you're short on
- 34:51:23data on one probability, you can kind of
- 34:51:25it automatically adjusts for that. Those
- 34:51:27formulas are very automatic. And so you
- 34:51:29can still get a very solid
- 34:51:31predictability even if you're missing
- 34:51:33data or you have overlapping data for
- 34:51:35two completely different areas. We see
- 34:51:36that a lot in doing census and studying
- 34:51:39of people and habits where they might
- 34:51:42have one study that covers one aspect
- 34:51:44and another one that overlaps and
- 34:51:45because the two overlap they can then
- 34:51:47predict the unknowns for the group that
- 34:51:49they haven't done the second study on or
- 34:51:51vice versa. So it's very powerful in
- 34:51:53that it is not sensitive to the
- 34:51:55irrelevant features and in fact you can
- 34:51:57use it to help predict features that
- 34:51:59aren't even in there. So now we're down
- 34:52:01to my favorite part. We're going to roll
- 34:52:03up our sleeves and do some actual
- 34:52:05programming. We're going to do the use
- 34:52:06case text classification. Now, I would
- 34:52:10challenge you to go back and send us a
- 34:52:12note on the notes below underneath the
- 34:52:14video and request the data for the
- 34:52:16shopping cart. So, you can plug that
- 34:52:18into Python code and do that on your own
- 34:52:20time. So, you can walk through it since
- 34:52:22we walk through all the information on
- 34:52:24it. But, we're going to do a Python code
- 34:52:26doing text classification. Very popular
- 34:52:28for doing the naive bays. So, we're
- 34:52:31going to use our new tool to perform a
- 34:52:33text classification of news headlines
- 34:52:35and classify news into different topics
- 34:52:37for a news website. As you can see here,
- 34:52:39we have a nice image of the Google News
- 34:52:42and then related on the right subgroups.
- 34:52:45I'm not sure where they actually pulled
- 34:52:46the actual data we're going to use from.
- 34:52:48It's one of the standard sets, but
- 34:52:50certainly this can be used on any of our
- 34:52:52news headlines in classification. So,
- 34:52:54let's see how it can be done using the
- 34:52:56naive base classifier. Now, we're at my
- 34:52:58favorite part. We're actually going to
- 34:53:00write some Python script, roll up our
- 34:53:02sleeves, and we're going to start by
- 34:53:04doing our imports. These are very basic
- 34:53:06imports, including our news group. And
- 34:53:08we'll take a quick glance at the target
- 34:53:10names. Then we're going to go ahead and
- 34:53:11start training our data set and putting
- 34:53:13it together. We'll put together a nice
- 34:53:15graph because it's always good to have a
- 34:53:17graph to show what's going on. And once
- 34:53:19we've trained it and we've shown you a
- 34:53:20graph of what's going on, then we're
- 34:53:22going to explore how to use it and see
- 34:53:24what that looks like. Now I'm going to
- 34:53:26open up my favorite editor or inline
- 34:53:28editor for Python. You don't have to use
- 34:53:30this. You can use whatever your editor
- 34:53:31that you like, whatever uh interface IDE
- 34:53:34you want. This just happens to be the
- 34:53:36Anaconda Jupiter notebook. And I'm going
- 34:53:39to paste that first piece of code in
- 34:53:40here so we can walk through it. Let's
- 34:53:42make it a little bigger on the screen so
- 34:53:44you have a nice view of what's going on.
- 34:53:45Uh and we're using Python 3, in this
- 34:53:47case 3.5. So this would work in any of
- 34:53:50your 3X if you have it set up correctly.
- 34:53:53should also work in a lot of the 2x. You
- 34:53:55just have to make sure all of the the
- 34:53:56versions of the modules match your
- 34:53:58Python version. And in here, you'll
- 34:54:00notice the first line is your percentage
- 34:54:02mattplot library in line. Now, three of
- 34:54:05these lines of code are all about
- 34:54:07plotting the graph. This one lets the
- 34:54:11notebook know and is inline setup that
- 34:54:14we want the graphs to show up on this
- 34:54:16page. Without it, in a notebook like
- 34:54:18this, which is an explorer interface, it
- 34:54:20won't show up. Now, a lot of IDEs don't
- 34:54:23require that. A lot of them, like on if
- 34:54:25I'm working on one of my other setups,
- 34:54:27it just has a popup and the graph pops
- 34:54:29up on there. So, you have a that setup
- 34:54:31also. But for this, we want the mattplot
- 34:54:33library in line. And then we're going to
- 34:54:35import numpy as np. That's number
- 34:54:39python, which has a lot of different
- 34:54:40formulas in it that we use for both of
- 34:54:43our sklearn module. And we also use it
- 34:54:46for any of the upper math functions in
- 34:54:48python. And it's very common to see that
- 34:54:49as NP numpy as NP. The next two lines
- 34:54:53are all about our graphing. Remember I
- 34:54:55said three of these were about graphing.
- 34:54:56Well, we need our mattplot
- 34:54:58library.pipplot
- 34:54:59as plt. And you'll see that plt is a
- 34:55:02very common setup as is the sns and just
- 34:55:04like the np. And we're going to import
- 34:55:06seabor as sns and we're going to do the
- 34:55:09sns set. Now seabor sits on top of
- 34:55:13pipplot and it just makes a really nice
- 34:55:15heat map. It's really good for heat
- 34:55:17maps. And if you're not familiar with
- 34:55:18heat maps, that just means we give it a
- 34:55:19color scale. The term comes from the
- 34:55:22brighter red it is, the hotter it is in
- 34:55:25some form of data. And you can set it to
- 34:55:26whatever you want. And we'll see that
- 34:55:28later on. So those you'll see that those
- 34:55:30three lines of code here are just
- 34:55:32importing the graph function so we can
- 34:55:33graph it. And as a data scientist, you
- 34:55:36always want to graph your data and have
- 34:55:38some kind of visual. It's really hard
- 34:55:39just to shove numbers in front of people
- 34:55:41and they look at it and it doesn't mean
- 34:55:42anything. And then from the sklearn data
- 34:55:45sets, we're going to import the fetch 20
- 34:55:47news groups. Very common one for
- 34:55:50analyzing tokenizing words and setting
- 34:55:52them up and exploring how the words work
- 34:55:54and how do you categorize different
- 34:55:55things when you're dealing with
- 34:55:56documents. And then we set our data
- 34:55:58equal to fetch 20 news groups. So our
- 34:56:01data variable will have the data in it.
- 34:56:03And we're going to go ahead and just
- 34:56:04print the target names. data.target
- 34:56:07names. And let's see what that looks
- 34:56:08like. And you'll see here we have alt
- 34:56:11atheism comp graphics composs
- 34:56:14windows.mmiscellaneous
- 34:56:16and it goes all the way down to talk
- 34:56:18politics.mmiscellaneous talk
- 34:56:20religion.mmiscellaneous. These are the
- 34:56:22categories they've already assigned to
- 34:56:24this news group and it's called fetch 20
- 34:56:26because you'll see there's I believe
- 34:56:27there's 20 different topics in here or
- 34:56:2920 different categories as we scroll
- 34:56:31down. Now, we've gone through the 20
- 34:56:33different categories and we're going to
- 34:56:35go ahead and start defining all the
- 34:56:37categories and set up our data. So,
- 34:56:39we're actually getting here going to go
- 34:56:41ahead and get it get the data all set up
- 34:56:43and take a look at our data. And let's
- 34:56:44move this over to our Jupyter notebook.
- 34:56:47And let's see what this code does.
- 34:56:49First, we're going to set our
- 34:56:51categories. Now, if you noticed up here,
- 34:56:53I could have just as easily set this
- 34:56:54equal to data.target_names target names
- 34:56:58because it's the same thing, but we want
- 34:57:00to kind of spell it out for you so you
- 34:57:01can see the different categories. It
- 34:57:03kind of makes it more visual so you can
- 34:57:04see what your data is looking like in
- 34:57:06the background. Once we've created the
- 34:57:08categories, we're going to open up a
- 34:57:10train set. So this training set of data
- 34:57:13is going to go into fetch 20 news groups
- 34:57:16and it's a subset in there called train
- 34:57:18and categories equals categories. So
- 34:57:20we're pulling out those categories that
- 34:57:22match. And then if you have a train set,
- 34:57:24you should also have the testing set. We
- 34:57:25have test equals fetch 20 news group
- 34:57:27subset equals test and categories equals
- 34:57:29categories. Let's go down one size so it
- 34:57:32all fits on my screen. There we go. And
- 34:57:34just so we can really see what's going
- 34:57:35on, let's see what happens when we print
- 34:57:37out one part of that data. So it creates
- 34:57:42train and under train, it creates train
- 34:57:44data. And we're just going to look at
- 34:57:46data piece number five. And let's go
- 34:57:48ahead and run that and see what that
- 34:57:50looks like. And you can see when I print
- 34:57:52train.data data number five under train.
- 34:57:55It prints out one of the articles. This
- 34:57:57is article number five. You can go
- 34:57:58through and read it on there. And we can
- 34:58:00also go in here and change this to test,
- 34:58:03which should look identical because it's
- 34:58:05splitting the date up into different
- 34:58:06groups. Train and test. And we'll see
- 34:58:08test number five is a a different
- 34:58:10article, but it's another article in
- 34:58:12here. And maybe you're curious and you
- 34:58:14want to see just how many articles are
- 34:58:16in here. We could do length of train.
- 34:58:21data. And if we run that, you'll see
- 34:58:23that the training data has 11,314
- 34:58:27articles. So, we're not going to go
- 34:58:29through all those articles. That's a lot
- 34:58:30of articles, but um we can look at one
- 34:58:32of them just so you can see what kind of
- 34:58:34information is coming out of it and what
- 34:58:35we're looking at. And we'll just look at
- 34:58:37number five for today. And here we have
- 34:58:39it. Rewarding the Second Amendment IDs,
- 34:58:41VTT, line 58, lines 58 in article, uh
- 34:58:45etc. And you can scroll all the way down
- 34:58:47and see all the different parts to
- 34:58:48there. Now, we've looked at it and
- 34:58:50that's pretty complicated when you look
- 34:58:51at one of these articles to try to
- 34:58:52figure out how do you weight this. If
- 34:58:54you look down here, we have different
- 34:58:56words and maybe the word from. Well,
- 34:58:58from is probably in all the articles.
- 34:59:00So, it's not going to have a lot of
- 34:59:02meaning as far as trying to figure out
- 34:59:03whether this article fits one of the
- 34:59:05categories or not. So, trying to figure
- 34:59:06out which category it fits in based on
- 34:59:09these words is where the challenge comes
- 34:59:10in. Now that we've viewed our data,
- 34:59:12we're going to dive in and do the actual
- 34:59:14predictions. This is the actual naive
- 34:59:17bays. And we're going to throw another
- 34:59:18model at you or another module at you
- 34:59:20here in just a second. We can't go into
- 34:59:22too much detail, but it deals
- 34:59:23specifically working with words and text
- 34:59:26and what they call tokenizing those
- 34:59:28words. So, let's take this code and
- 34:59:31let's uh skip on over to our Jupyter
- 34:59:33notebook and walk through it. And here
- 34:59:35we are in our Jupyter notebook. Let's
- 34:59:36paste that in there. And I can run this
- 34:59:38code right off the bat. It's not
- 34:59:39actually going to display anything yet,
- 34:59:41but it has a lot going on in here. So
- 34:59:44the top we had the print module from the
- 34:59:46earlier one. I didn't know why that was
- 34:59:47in there. So we're going to start by
- 34:59:49importing our necessary packages. And
- 34:59:51from the sklearn features
- 34:59:53extraction.ext,
- 34:59:55we're going to import TF IDF vectorzer.
- 34:59:59I told you we're going to throw a module
- 35:00:00at you. We can't go too much into the
- 35:00:03math behind this or how it works. You
- 35:00:04can look it up. The notation for the
- 35:00:06math is usually TF.idf.
- 35:00:09And that's just a way of weighing the
- 35:00:11words. and it weighs the words based on
- 35:00:14how many times are used in a document,
- 35:00:16how many times or how many documents
- 35:00:18they're used in. And it's a well-used
- 35:00:20formula. It's been around for a while.
- 35:00:21It's a little confusing to put this in
- 35:00:23here. Uh, but let's let them know that
- 35:00:25it just goes in there and weights the
- 35:00:27different words in the document for us.
- 35:00:29That way, we don't have to wait. And if
- 35:00:31you put a weight on it, if you remember,
- 35:00:33I was talking about that up here
- 35:00:34earlier. If these are all emails, they
- 35:00:36probably all have the word from in them.
- 35:00:38From probably has a very low weight. It
- 35:00:40has very little value in telling you
- 35:00:41what this document's about. Same with
- 35:00:43words like in an article in articles in
- 35:00:46cost of un maybe cost might or where
- 35:00:50words like criminal weapons destruction
- 35:00:53these might have a heavier weight
- 35:00:55because they describe a little bit more
- 35:00:56what the article is doing. Well, how do
- 35:00:58you figure out all those weights in the
- 35:00:59different articles? That's what this
- 35:01:01module does. That's what the TF
- 35:01:04vectorizer is going to do for us. And
- 35:01:06then we're going to import our
- 35:01:07sklearn.na naive bays and that's our
- 35:01:10multinnomial NB multinnomial naive bay
- 35:01:14pretty easy to understand that where
- 35:01:15that comes from and then finally we have
- 35:01:17the skyarn pipeline import make pipeline
- 35:01:21now the make pipeline is just a cool
- 35:01:23piece of code because we're going to
- 35:01:25take the information we get from the TF
- 35:01:28vectorizer and we're going to pump that
- 35:01:31into the multinnomial NB. So, a pipeline
- 35:01:35is just a way of organizing how things
- 35:01:38flow. It's used commonly. You probably
- 35:01:40already guessed what it is. If you've
- 35:01:41done any businesses, they talk about the
- 35:01:43sales pipeline. If you're on a work crew
- 35:01:46or project manager, you have your
- 35:01:48pipeline of information that's going
- 35:01:49through or your projects and what has to
- 35:01:51be done in what order. That's all this
- 35:01:52pipeline is. We're going to take the
- 35:01:55TFID vectorzer and then we're going to
- 35:01:57push that into the multinnomial inb. Now
- 35:02:00we've designated that as the variable
- 35:02:03model. We have our pipeline model and
- 35:02:05we're going to take that model and this
- 35:02:08is just so elegant. This is done in just
- 35:02:09a couple lines of code. model.fit and
- 35:02:13we're going to fit the data. And first
- 35:02:15the train data and then the train
- 35:02:17target. Now the train data has the
- 35:02:20different articles in it. You can see
- 35:02:22the one we were just looking at and the
- 35:02:24train.target target is what category
- 35:02:26they already categorized that that
- 35:02:28particular article as. And what's
- 35:02:30happening here is the train data is
- 35:02:33going into the TF ID vectorizer. So when
- 35:02:36you have one of these articles, it goes
- 35:02:38in there, it weights all the words in
- 35:02:40there. So there's thousands of words
- 35:02:42with different weights on them. I
- 35:02:43remember once running a model on this
- 35:02:44and I literally had 2.4 million tokens
- 35:02:48go into this. So when you're dealing
- 35:02:50like large document bases, you can have
- 35:02:52a huge number of different words. It
- 35:02:54then takes those words, gives them a
- 35:02:56weight, and then based on that weight,
- 35:02:59based on the words and the weights, and
- 35:03:00then puts that into the multinnomial NB.
- 35:03:03And once we go into our naive bay, we
- 35:03:06want to put the train target in there.
- 35:03:08So the train data that's been mapped to
- 35:03:10the TFID vectorzer is now going through
- 35:03:14the multinnomial NB. And then we're
- 35:03:16telling it, well, these are the answers.
- 35:03:18These are the answers to the different
- 35:03:19documents. So this document that has all
- 35:03:21these words with these different weights
- 35:03:23from the first part is going to be
- 35:03:25whatever category it comes out of. Maybe
- 35:03:27it's the um talk show or the article on
- 35:03:30religion miscellaneous. Once we fit that
- 35:03:33model, we can then take labels and we're
- 35:03:36going to set that equal to
- 35:03:38model.predict. Most of the sklearn use
- 35:03:41the term.predict to let us know that
- 35:03:43we've now trained the model and now we
- 35:03:45want to get some answers. And we're
- 35:03:46going to put our test data in there
- 35:03:48because our test data is the stuff we
- 35:03:50held off to the side. We didn't train it
- 35:03:52on there and we don't know what's going
- 35:03:54to come up out of it and we just want to
- 35:03:55find out how good our labels are. Do
- 35:03:57they match what they should be? Now,
- 35:03:59I've already run this through. There's
- 35:04:01no actual output to it to show. This is
- 35:04:03just setting it all up. This is just
- 35:04:05training our model, creating the labels
- 35:04:07so we can see how good it is, and then
- 35:04:08we move on to the next step to find out
- 35:04:10what happened. To do this, we're going
- 35:04:13to go ahead and create a confusion
- 35:04:15matrix and a heat map. So, the confusion
- 35:04:18matrix, which is confusing just by its
- 35:04:21very name, is basically going to ask how
- 35:04:23confused is our answer. Did it get it
- 35:04:26correct or did it miss some things in
- 35:04:28there or have some missed labels? And
- 35:04:30then we're going to put that on a heat
- 35:04:31map so we have some nice colors to look
- 35:04:33at to see how that plots out. Let's go
- 35:04:35ahead and take this code and see how
- 35:04:37that uh take a walk through it and see
- 35:04:39what that looks like. So, back to our
- 35:04:41Jupyter notebook. I'm going to put the
- 35:04:42code in there and let's go ahead and run
- 35:04:45that code. Take it just a moment. And
- 35:04:48remember, we had the inline. That way,
- 35:04:50my graph shows up on the inline here.
- 35:04:53And let's walk through the code and then
- 35:04:54we'll look at this and see what that
- 35:04:56means. So, make it a little bit bigger.
- 35:04:58There we go. No reason not to use the
- 35:05:00whole screen. Too big. So, we have here
- 35:05:02from sklearn metrics import confusion
- 35:05:05matrix. And that's just going to
- 35:05:07generate a set of data that says I the
- 35:05:10prediction was such the actual truth was
- 35:05:14either agreed with it or was something
- 35:05:16different. And it's going to add up
- 35:05:17those numbers so we can take a look and
- 35:05:18just see how well it worked. And we're
- 35:05:20going to set a variable Matt equal to
- 35:05:22confusion matrix. We have our test
- 35:05:25target, our test data that was not part
- 35:05:27of the training. Very important in data
- 35:05:29science, we always keep our test data
- 35:05:31separate. Otherwise, it's not a valid
- 35:05:34model if we can't properly test it with
- 35:05:36new data. And this is the labels we
- 35:05:38created from that test data. These are
- 35:05:40the ones that we predict it's going to
- 35:05:42be. So, we go in and we create our SN
- 35:05:44heat map. The SNS is our seaborn which
- 35:05:47sits on top of the piplot. So, we create
- 35:05:50a SNS.heet map. We take our confusion
- 35:05:53matrix and it's going to be uh matt.t.
- 35:05:57And then we have other variables that go
- 35:05:59into the SNS heat map. We're not going
- 35:06:02to go into detail what all the variables
- 35:06:04mean. The annotation equals true. That's
- 35:06:06what tells it to put the numbers here.
- 35:06:07So you have the 166, the one, the 00001.
- 35:06:11Format D and C bar equals false have to
- 35:06:13do with the uh format. If you take those
- 35:06:16out, you'll see that some things
- 35:06:17disappear. And then the X tick labels
- 35:06:19and the Y tick labels. Those are our
- 35:06:21target names. And you can see right
- 35:06:23here, that's the alt atheism comp
- 35:06:26graphics composs windows.mmiscellaneous.
- 35:06:29And then finally we have our plt.xl
- 35:06:32label. Remember the SNS or the seabor
- 35:06:35sits on top of our mattplot library our
- 35:06:37plt. And so we want to just tell it x
- 35:06:39label equals a true is is true. The
- 35:06:41labels are true. And then the y label is
- 35:06:44prediction label. So when we say a true,
- 35:06:47this is what it actually is. And the
- 35:06:49prediction is what we predicted. And
- 35:06:51let's look at this graph because that's
- 35:06:53probably a little confusing the way I
- 35:06:54rattled through it. And what I'm going
- 35:06:56to do is I'm going to go ahead and flip
- 35:06:58back to the slides because they have a
- 35:07:00black background they put in there that
- 35:07:02helps it shine a little bit better so
- 35:07:03you can see the graph a little bit
- 35:07:04easier. So in reading this graph, what
- 35:07:07we want to look at is how the color
- 35:07:09scheme has come out. And you'll see a
- 35:07:11line right down the middle diagonally
- 35:07:14from upper left to bottom right. What
- 35:07:16that is is if you look at the labels, we
- 35:07:18have our predicted label on the left and
- 35:07:21our true label on the right. Those are
- 35:07:23the numbers where the prediction and the
- 35:07:25true come together. And this is what we
- 35:07:27want to see is we want to see those lit
- 35:07:28up. That's what that heat map does. As
- 35:07:30you can see that it did a good job of
- 35:07:32finding those data. And you'll notice
- 35:07:34that there's a couple of red spots on
- 35:07:36there where it missed. You know, it it's
- 35:07:38a little confused when we talk about
- 35:07:40talk religion miscellaneous versus talk
- 35:07:42politics miscellaneous, social religion
- 35:07:45Christian versus alt atheism. It
- 35:07:47mislabeled some of those. And those are
- 35:07:49very similar topics. so you could
- 35:07:51understand why it might mislabel them.
- 35:07:53But overall, it did a pretty good job.
- 35:07:55If we're going to create these models,
- 35:07:56we want to go ahead and be able to use
- 35:07:58them. So, let's see what that looks
- 35:08:00like. To do this, let's go ahead and
- 35:08:02create a definition, a function to run.
- 35:08:05And we're going to call this function.
- 35:08:06Let me just expand that just a notch
- 35:08:08here. There we go. I like mine in big
- 35:08:10letters. Predict category. So, we want
- 35:08:12to predict the category. We're going to
- 35:08:14send it as a string. And then we're
- 35:08:17sending it train equals train. We have
- 35:08:19our training model. And then we had our
- 35:08:21pipeline model equals model. This way we
- 35:08:23don't have to resend these variables
- 35:08:24each time. The definition knows that
- 35:08:27because I said train equals train and I
- 35:08:29put the equal for model. And then we're
- 35:08:31going to set the prediction equal to the
- 35:08:32model.predict s. So it's going to send
- 35:08:35whatever string we send to it. It's
- 35:08:38going to push that string through the
- 35:08:39pipeline, the model pipeline. It's going
- 35:08:41to go through and uh tokenize it and put
- 35:08:44it through the TF IDF, convert that into
- 35:08:48numbers and weights for all the
- 35:08:49different documents and words. And then
- 35:08:51it'll put that through our naive bay.
- 35:08:54And from it, we'll go ahead and get our
- 35:08:56prediction. We're going to predict what
- 35:08:58value it is. And so we're going to
- 35:09:00return train.target names predict of
- 35:09:03zero. And remember that the train.target
- 35:09:06names, that's just categories. I could
- 35:09:08have just as easily put uh categories in
- 35:09:10there.predict of zero. So we're taking
- 35:09:13the prediction which is a number and
- 35:09:15we're converting it to an actual
- 35:09:16category. We're converting it from um I
- 35:09:19don't know what the actual numbers are.
- 35:09:20Let's say zero equals alt atheism. So
- 35:09:23we're going to convert that zero to the
- 35:09:24word or uh one maybe it equals comp
- 35:09:27graphics. So we're going to convert
- 35:09:28number one into comp graphics. That's
- 35:09:30all that is. And then we got to go ahead
- 35:09:32and and then we need to go ahead and run
- 35:09:35this. So I load that up. And then once I
- 35:09:38run that, we can start doing some
- 35:09:40predictions. Let me go ahead and type in
- 35:09:42predict category. And let's just do
- 35:09:45predict category, Jesus Christ. And it
- 35:09:47comes back and says it's social,
- 35:09:49religion, Christian. That's pretty good.
- 35:09:52Now note, I didn't put print on this.
- 35:09:54One of the nice things about the Jupiter
- 35:09:56notebook editor and a lot of inline
- 35:09:58editors is if you just put the name of
- 35:10:00the variable out, it's returning the
- 35:10:02variable train.target_ames, target
- 35:10:04names. It'll automatically print that
- 35:10:06for you. In your own IDE, you might have
- 35:10:08to put in print. Let's see where else we
- 35:10:10can take this. And maybe you're a space
- 35:10:12science buff. So, how about sending load
- 35:10:16to international
- 35:10:19space station.
- 35:10:21And if we run that, we get science
- 35:10:24space. Or maybe you're a uh automobile
- 35:10:28buff. And let's do um Oh, they were
- 35:10:31gonna tell me Audi is better than BMW,
- 35:10:32but I'm going to do BMW is better than
- 35:10:36an Audi. So maybe our car buff. And we
- 35:10:39run that. And you'll see it says
- 35:10:41recreational. I'm assuming that's what
- 35:10:42RECC stands for. Autos. So I did a
- 35:10:45pretty good job labeling that one. How
- 35:10:47about uh if we have something like a
- 35:10:49caption running through there, President
- 35:10:51of India. And if we run that, it comes
- 35:10:54up and says talk politics miscellaneous.
- 35:10:58So when we take our definition or our
- 35:11:00function and we run all these things
- 35:11:02through, kudos, we made it. We were able
- 35:11:04to correctly classify texts into
- 35:11:06different groups based on which category
- 35:11:08they belong to using the naive base
- 35:11:11classifier. Now we did throw in the
- 35:11:13pipeline, the TF IDF vectorzer, we threw
- 35:11:17in the graphs. Those are all things that
- 35:11:19you don't necessarily have to know to
- 35:11:21understand the naive base setup or
- 35:11:24classifier, but they're important to
- 35:11:25know. One of the main uses for the naive
- 35:11:27bays is with the TF IDF tokenizer
- 35:11:31vectorzer where it tokenizes a word and
- 35:11:33has labels and we use the pipeline
- 35:11:36because you need to push all that data
- 35:11:38through and it makes it really easy and
- 35:11:39fast. You don't have to know those to
- 35:11:41understand naive bays but they certainly
- 35:11:43help for understanding the industry and
- 35:11:45data science. And we can see our
- 35:11:47categorizer, our naive base classifier.
- 35:11:50We were able to predict the category
- 35:11:52religion, space, motorcycles, autos,
- 35:11:55politics, and properly classify all
- 35:11:57these different things we pushed into
- 35:11:59our prediction and our trained model.
- 35:12:01Before we dive into the SVM, let's take
- 35:12:04a look at applications of the support
- 35:12:06vector machine, at least some general
- 35:12:08ones that are commonly used with it.
- 35:12:10face detection, text and hypertext
- 35:12:12categorization, classification of
- 35:12:14images, and bioinformatics.
- 35:12:17These are only but a few of those that
- 35:12:19are used with this SVM. As we go through
- 35:12:21this lesson, see if you can figure out
- 35:12:23what other ones you could apply it to,
- 35:12:24and also what you would want to use some
- 35:12:26other tools for. So, in this example,
- 35:12:28last week, my son and I visited a fruit
- 35:12:31shop. Dad, is that an apple or a
- 35:12:33strawberry? So, the question comes up,
- 35:12:35what fruit did I just pick up from the
- 35:12:37fruit stand? After a couple of seconds,
- 35:12:39you can figure out that it was a
- 35:12:40strawberry. So, let's take this model a
- 35:12:42step further and let's uh why not build
- 35:12:45a model which can predict an unknown
- 35:12:46data. And in this, we're going to be
- 35:12:48looking at some sweet strawberries or
- 35:12:50crispy apples. We wanted to be able to
- 35:12:52label those two and decide what the
- 35:12:53fruit is. And we do that by having data
- 35:12:56already put in. So, we already have a
- 35:12:58bunch of strawberries. We know our
- 35:12:59strawberries and they're already labeled
- 35:13:00as such. We already have a bunch of
- 35:13:02apples. We know our apples and are
- 35:13:03labeled as such. Then once we train our
- 35:13:05model, that model then can be given the
- 35:13:07new data and the new data is this image.
- 35:13:09In this case, you can see a question
- 35:13:11mark on it and it comes through and goes
- 35:13:12it's a strawberry. In this case, we're
- 35:13:14using the support vector machine model.
- 35:13:17SVM is a supervised learning method that
- 35:13:20looks at data and sorts it into one of
- 35:13:23two categories. And in this case, we're
- 35:13:25sorting the strawberry into the
- 35:13:27strawberry side. At this point, you
- 35:13:29should be asking the question, how does
- 35:13:31the prediction work? Before we dig into
- 35:13:33an example with numbers, let's apply
- 35:13:35this to our fruit scenario. We have our
- 35:13:37support vector machine. We've taken it
- 35:13:39and we've taken labeled sample of data,
- 35:13:42strawberries and apples, and we draw on
- 35:13:44a line down the middle between the two
- 35:13:46groups. This split now allows us to take
- 35:13:49new data, in this case an apple and a
- 35:13:51strawberry, and place them in the
- 35:13:53appropriate group based on which side of
- 35:13:54the line they fall in. And that way we
- 35:13:56can predict the unknown. As colorful and
- 35:13:58tasty as the fruit example is, let's
- 35:14:00take a look at another example with some
- 35:14:02numbers involved. And we can take a
- 35:14:04closer look at how the math works. In
- 35:14:06this example, we're going to be
- 35:14:07classifying men and women. And we're
- 35:14:09going to start with a set of people with
- 35:14:11a different height and a different
- 35:14:13weight. And to make this work, we'll
- 35:14:15have to have a sample data set of female
- 35:14:17where we have their height and weight
- 35:14:19174, 65, 174, 88, and so on. And we'll
- 35:14:23need a sample data set of the male. They
- 35:14:24have a height 179, 90, 180 to 80 and so
- 35:14:28on. Let's go ahead and put this on a
- 35:14:29graph so we have a nice visual. So you
- 35:14:31can see here we have two groups based on
- 35:14:33the height versus the weight. And on the
- 35:14:36left side we're going to have the women,
- 35:14:37on the right side we're going to have
- 35:14:39the men. Now if we're going to create a
- 35:14:40classifier, let's add a new data point
- 35:14:42and figure out if it's male or female.
- 35:14:44So before we can do that, we need to
- 35:14:47split our data first. We can split our
- 35:14:49data by choosing any of these lines. In
- 35:14:52this case, we draw in two lines through
- 35:14:54the data in the middle that separates
- 35:14:56the men from the women. But to predict
- 35:14:57the gender of a new data point, we
- 35:14:59should split the data in the best
- 35:15:01possible way. And we say the best
- 35:15:03possible way because this line has a
- 35:15:06maximum space that separates the two
- 35:15:08classes. Here you can see there's a
- 35:15:10clear split between the two different
- 35:15:12classes. And in this one, there's not so
- 35:15:15much a clear split. This doesn't have
- 35:15:17the maximum space that separates the
- 35:15:18two. That is why this line best splits
- 35:15:21the data. We don't want to just do this
- 35:15:23by eyeballing it. And before we go
- 35:15:25further, we need to add some technical
- 35:15:27terms to this. We can also say that the
- 35:15:30distance between the points in the line
- 35:15:31should be as far as possible. In
- 35:15:33technical terms, we can say the distance
- 35:15:35between the support vector and the hyper
- 35:15:38plane should be as far as possible. And
- 35:15:40this is where the support vectors are
- 35:15:42the extreme points in the data set. And
- 35:15:44if you look at this data set, they have
- 35:15:46circled two points which seem to be
- 35:15:48right on the outskirts of the women and
- 35:15:50one on the outskirts of the men. And
- 35:15:52hyper plane has a maximum distance to
- 35:15:54the support vectors of any class. Now
- 35:15:56you'll see the line down the middle and
- 35:15:58we call this the hyper plane because
- 35:16:00when you're dealing with multiple
- 35:16:01dimensions, it's really not just a line
- 35:16:03but a plane of intersections. And you
- 35:16:05can see here where the support vectors
- 35:16:07have been drawn in dashed lines. The
- 35:16:10math behind this is very simple. We take
- 35:16:12D+ the shortest distance to the closest
- 35:16:15positive point which would be on the
- 35:16:17men's side and D minus is the shortest
- 35:16:19distance to the closest negative point
- 35:16:21which is on the women's side. The sum of
- 35:16:23D plus and D minus is called the
- 35:16:25distance margin or the distance between
- 35:16:27the two support vectors that are shown
- 35:16:29in the dashed lines. And then by finding
- 35:16:32the largest distance margin, we can get
- 35:16:35the optimal hyper plane. Once we've
- 35:16:37created an optimal hyper plane, we can
- 35:16:39easily see which side the new data fits
- 35:16:41in. And based on the hyper plane, we can
- 35:16:43say the new data point belongs to the
- 35:16:44male gender. Hopefully that's clear how
- 35:16:47that works on a visual level. As a data
- 35:16:49scientist, you should also be asking
- 35:16:51what happens if the hyper plane is not
- 35:16:53optimal. If we select a hyper plane
- 35:16:55having low margin, then there is a high
- 35:16:57chance of mclassification. This
- 35:16:59particular SVM model, the one we
- 35:17:02discussed so far, is also called
- 35:17:04referred to as the LSVM.
- 35:17:06So far so clear, but a question should
- 35:17:09be coming up. We have our sample data
- 35:17:11set. But instead of looking like this,
- 35:17:13what if it looked like this where we
- 35:17:16have two sets of data, but one of them
- 35:17:18occurs in the middle of another set. You
- 35:17:20can see here where we have the blue and
- 35:17:22the yellow and then blue again on the
- 35:17:24other side of our data line. In this
- 35:17:26data set, we can't use a hyper plane. So
- 35:17:29when you see data like this, it's
- 35:17:31necessary to move away from a 1D view of
- 35:17:33the data to a two-dimensional view of
- 35:17:35the data. And for the transformation, we
- 35:17:37use what's called a kernel function. The
- 35:17:40kernel function will take the 1D input
- 35:17:42and transfer it to a two-dimensional
- 35:17:44output. As you can see in this picture
- 35:17:47here, the 1D when transferred to a
- 35:17:49two-dimensional makes it very easy to
- 35:17:51draw a line between the two data sets.
- 35:17:53What if we make it even more
- 35:17:55complicated? How do we perform an SVM
- 35:17:57for this type of data set? Here you can
- 35:17:59see we have a two-dimensional data set
- 35:18:01where the data is in the middle
- 35:18:03surrounded by the green data on the
- 35:18:05outside. In this case, we're going to
- 35:18:07segregate the two classes. We have our
- 35:18:09sample data set and if you draw a line
- 35:18:12through, it's obviously not an optimal
- 35:18:13hyper plane in there. So to do that, we
- 35:18:15need to transfer the 2D to a 3D array.
- 35:18:18And when you translate it into a
- 35:18:20three-dimensional array using the
- 35:18:21kernel, you can see where you can place
- 35:18:22a hyper plane right through it and
- 35:18:24easily split the data. Before we start
- 35:18:26looking at a programming example and
- 35:18:28dive into the script, let's look at the
- 35:18:30advantage of the support vector machine.
- 35:18:32We'll start with highdimensional input
- 35:18:34space or sometimes referred to as the
- 35:18:37curse of dimensionality. We looked at
- 35:18:39earlier one dimension, two dimension,
- 35:18:41three dimension. When you get to a
- 35:18:43thousand dimensions, a lot of problems
- 35:18:45start occurring with most algorithms
- 35:18:46that have to be adjusted for. The SVM
- 35:18:49automatically does that in
- 35:18:50highdimensional space. One of the
- 35:18:52highdimensional space, one
- 35:18:54highdimensional space that we work on is
- 35:18:56sparse document vectors. This is where
- 35:18:58we tokenize the words in documents so we
- 35:19:00can run our machine learning algorithms
- 35:19:02over them. I've seen ones get as high as
- 35:19:042.4 million different tokens. That's a
- 35:19:06lot of vectors to look at. And finally,
- 35:19:09we have regularization parameter. The
- 35:19:11realization parameter or lambda is a
- 35:19:14parameter that helps figure out whether
- 35:19:15we're going to have a bias or
- 35:19:17overfitting of the data. Whether it's
- 35:19:19going to be overfitted to very specific
- 35:19:20instance or it's going to be biased to a
- 35:19:22high or low value. With the SVM, it
- 35:19:25naturally avoids the overfitting and
- 35:19:27bias problems that we see in many other
- 35:19:29algorithms. These three advantages of
- 35:19:31the support vector machine make it a
- 35:19:33very powerful tool to add to your
- 35:19:35repertoire of machine learning tools.
- 35:19:37Now, we did promise you a use case
- 35:19:39study. We're actually going to dive in
- 35:19:40to some Python programming. And so we're
- 35:19:42going to go into a problem statement and
- 35:19:44start off with the zoo. So in the zoo
- 35:19:46example, we have um family members going
- 35:19:49to the zoo and we have the young child
- 35:19:50going, "Dad, is that a group of
- 35:19:52crocodiles or alligators?" Well, that's
- 35:19:54hard to differentiate. And zoos are a
- 35:19:56great place to start looking at science
- 35:19:58and understanding how things work,
- 35:20:00especially as a young child. And so we
- 35:20:02can see the parents sitting here
- 35:20:03thinking, well, what is the difference
- 35:20:04between a crocodile and an alligator?
- 35:20:06Well, one, crocodiles are larger in
- 35:20:08size. Alligators are smaller in size.
- 35:20:10Snout width. The crocodiles have a
- 35:20:12narrow snout and alligators have a wider
- 35:20:14snout. And of course, in the modern day
- 35:20:16and age, the father's sitting here is
- 35:20:17thinking, "How can I turn this into a
- 35:20:19lesson for my son?" And he goes, "Let a
- 35:20:21support vector machine segregate the two
- 35:20:23groups." I don't know if my dad ever
- 35:20:25told me that, but that would be funny.
- 35:20:26Now, in this example, we're not going to
- 35:20:29use actual measurements and data. We're
- 35:20:31just using that for imagery. And that's
- 35:20:33very common in a lot of machine learning
- 35:20:34algorithms and setting them up. But
- 35:20:36let's roll up our sleeves and we'll talk
- 35:20:38about that more in just a moment as we
- 35:20:39break into our Python script. So here we
- 35:20:42arrive in our actual coding and I'm
- 35:20:45going to move this into a Python editor
- 35:20:47in just a moment. But let's talk a
- 35:20:49little bit about what we're going to
- 35:20:50cover. First, we're going to cover in
- 35:20:52the code the setup, how to actually
- 35:20:55create our SVM. And you're going to find
- 35:20:57that there's only two lines of code that
- 35:20:58actually create it. And the rest of it
- 35:21:00is done so quick and fast that it's all
- 35:21:02here in the first page. and we'll show
- 35:21:04you what that looks like as far as our
- 35:21:06data because we're going to create some
- 35:21:07data. I talked about creating data just
- 35:21:09a minute ago. And so we'll get into the
- 35:21:10creating data here and you'll see this
- 35:21:12nice correction of our two blobs and
- 35:21:13we'll go through that in just a second.
- 35:21:15And then the second part is we're going
- 35:21:16to take this and we're going to bump it
- 35:21:18up a notch. We're going to show you what
- 35:21:19it looks like behind the scenes. But
- 35:21:21let's start with actually creating our
- 35:21:23setup. I like to use the Anaconda
- 35:21:25Jupyter notebook because it's very easy
- 35:21:27to use, but you can use any of your
- 35:21:29favorite Python editors or setups and go
- 35:21:32in there. But let's go ahead and switch
- 35:21:33over there and see what that looks like.
- 35:21:35So here we are in the Anaconda Python
- 35:21:38notebook or Anaconda Jupyter notebook
- 35:21:40with Python. We're using Python 3. I
- 35:21:43believe this is 3.5, but it should be
- 35:21:45work in any of your 3x versions. And uh
- 35:21:48you'd have to look at the sklearn and
- 35:21:50make sure if you're using a 2x version,
- 35:21:52an earlier version. Let's go and put our
- 35:21:54code in there. And one of the things I
- 35:21:55like about the Jupyter notebook is I can
- 35:21:57go up to view and I'm going to go ahead
- 35:21:59and toggle the line numbers on to make
- 35:22:00it a little bit easier to talk about.
- 35:22:03And we can even increase the size
- 35:22:04because this is edited in in this case
- 35:22:06I'm using Google Chrome explorer and
- 35:22:08that's how it opens up for the editor.
- 35:22:10Although anyone any like I said any
- 35:22:11editor will work. Now the first step is
- 35:22:13going to be our imports and we're going
- 35:22:15to import four different parts. The
- 35:22:18first two I want you to look at are line
- 35:22:20one and line two are numpy as np and
- 35:22:23mapplot library.pipplot
- 35:22:25as plt. Now these are very standardized
- 35:22:29imports when you're doing work. The
- 35:22:30first one is the numbers python. We need
- 35:22:33that because part of the platform we're
- 35:22:35using uses that for the numpy array. And
- 35:22:38I'll talk about that in a minute so you
- 35:22:39can understand why we want to use a
- 35:22:41numpy array versus a standard python
- 35:22:43array. And normally it's pretty standard
- 35:22:45setup to use NP for numpy. The map plot
- 35:22:48library is how we're going to view our
- 35:22:50data. So this has uh you do need the NP
- 35:22:53for the sklearn module, but the map plot
- 35:22:55library is purely for our use for
- 35:22:57visualization. And so you really don't
- 35:22:59need that for the SVM, but we're going
- 35:23:00to put it there so you have a nice
- 35:23:02visual aid and we can show you what it
- 35:23:03looks like. That's really important at
- 35:23:04the end when you finish everything so
- 35:23:06you have a nice display for everybody to
- 35:23:08look at. And then finally, we're going
- 35:23:09to I'm going to jump one ahead to line
- 35:23:12number four. That's the sklearn.datas
- 35:23:14sets.samples generator import make
- 35:23:18blobs. And I told you that we were going
- 35:23:20to make up data. And this is a tool
- 35:23:21that's in the sklearn to make up data. I
- 35:23:23personally don't want to go to the zoo,
- 35:23:25get in trouble for jumping over the
- 35:23:26fence, and probably get eaten by the
- 35:23:28crocodiles or alligators as I work on
- 35:23:30measuring their snouts and width and
- 35:23:32length. Instead, we're just going to
- 35:23:34make up some data. And that's what that
- 35:23:36make blobs is. It's a wonderful tool. If
- 35:23:38you're ready to test your your uh setup
- 35:23:41and you're not sure about what data
- 35:23:42you're going to put in there, you can
- 35:23:43create this blob and it makes it really
- 35:23:45easy to use. And finally, we have our
- 35:23:47actual SVM, the sklearn import SVM on
- 35:23:50line three. So that covers all our
- 35:23:52imports. We're going to create, remember
- 35:23:54I used the make blobs to create data.
- 35:23:56And we're going to create a capital X
- 35:23:58and a lowercase Y equals make blobs in
- 35:24:01samples equals 40. So we're going to
- 35:24:02make 40 lines of data. It's going to
- 35:24:04have two centers with a random state
- 35:24:07equals 20. So each each each group's
- 35:24:09going to have 20 different pieces of
- 35:24:10data in it. And the way that looks is
- 35:24:13that we'll have under X um an XY plane.
- 35:24:16So I have two numbers under X and Y will
- 35:24:18be 01. That's the two different centers.
- 35:24:21So we have yes or no in this case
- 35:24:23alligator crocodile. That's what that
- 35:24:25represents. And then I told you that the
- 35:24:28actual sklearn or the SVM is in two
- 35:24:31lines of code. And we see it right here
- 35:24:33with CLF equals SVM. SVC kernel equals
- 35:24:37linear. And I set C equal to one.
- 35:24:39Although in this example, since we are
- 35:24:41not uh regularizing the data because we
- 35:24:43want it to be very clear and easy to
- 35:24:44see, I went ahead. You can set it to a
- 35:24:46th00and a lot of times when you're not
- 35:24:48doing that. But for this thing linear,
- 35:24:50because it's a very simple linear
- 35:24:51example, we only have the two dimensions
- 35:24:53and it'll be a nice linear hyper plane.
- 35:24:56It'll be a nice linear line instead of a
- 35:24:58full plane. So we're not dealing with a
- 35:25:00huge amount of data. And then all we
- 35:25:01have to do is do clff.fit
- 35:25:04x, y. And that's it. CLF has been
- 35:25:07created. And then we're going to go
- 35:25:08ahead and display it. And I'm going to
- 35:25:10talk about this display here in just a
- 35:25:12second. But let me go ahead and run this
- 35:25:13code. And this is what we've done is
- 35:25:15we've created two blobs. You'll see the
- 35:25:17blue on the side and then kind of an
- 35:25:19orang-ish uh on the other side. That's
- 35:25:21our two sets of data. They represent one
- 35:25:23represents crocodiles and one represents
- 35:25:25alligators. And then we have our
- 35:25:27measurements. In this case, we have like
- 35:25:28the width and length of the snout. And I
- 35:25:31did say I was going to come up here and
- 35:25:32talk just a little bit about our plot.
- 35:25:34And you'll see plt. That's what we
- 35:25:36imported. We're going to do a scatter
- 35:25:38plot. That means we're just putting dots
- 35:25:40on there. And then look at this
- 35:25:41notation. I have the capital X and then
- 35:25:44in brackets I have a colon, 0ero. That's
- 35:25:47from numpy. If you did that in a regular
- 35:25:49array, you'll get an error in a Python
- 35:25:51array. You have to have that in a numpy
- 35:25:53array. It turns out that our make blobs
- 35:25:55returns a numpy array. And this notation
- 35:25:58is great because what it means is the
- 35:26:00first part is the colon means we're
- 35:26:02going to do all the rows. That's all the
- 35:26:04data in our blob we created under
- 35:26:06capital X. And then the second part has
- 35:26:08a comma 0ero. We're only going to take
- 35:26:10the first value. And then if you notice,
- 35:26:13we do the same thing, but we're going to
- 35:26:14take the second value. Remember, we
- 35:26:16always start with zero and then one. So
- 35:26:18we have column zero and column one. And
- 35:26:20you can look at this as our XY plots.
- 35:26:23The first one is the xplot and the
- 35:26:25second one is the y plot. So the first
- 35:26:27one is on the bottom 0 2 4 6 8 and 10.
- 35:26:31And then the second one x of the one is
- 35:26:34the 4 5 6 7 8 9 10 going up the left
- 35:26:37hand side. S= 30 is just the size of the
- 35:26:39dots. We can see them instead of real
- 35:26:41tiny dots. And then cmap equals
- 35:26:44plt.cm.paired.
- 35:26:46And you'll also see the c equals y.
- 35:26:48That's the color. We're using two colors
- 35:26:5101. And that's why we get the nice blue
- 35:26:54and the two different colors for the
- 35:26:55alligator and the crocodile. Now you can
- 35:26:58see here that we did this the actual fit
- 35:27:00was done in two lines of code. A lot of
- 35:27:02times there'll be a third line where we
- 35:27:04regularize the data. We set it between
- 35:27:06like minus one and one and we reshape
- 35:27:08it. But for this it's not necessary and
- 35:27:10it's also kind of nice because you can
- 35:27:12actually see what's going on. And then
- 35:27:14if we wanted to we wanted to actually
- 35:27:16run a prediction. Let's take a look and
- 35:27:17see what that looks like. And to predict
- 35:27:19some new data and we'll show this again
- 35:27:22as we get towards the end of digging in
- 35:27:23deep. You can simply assign your new
- 35:27:26data. In this case I am giving it a uh
- 35:27:29width and length 34 and a width and
- 35:27:31length 56. And note that I put the data
- 35:27:34as a set of brackets and then I have the
- 35:27:36brackets inside. And the reason I do
- 35:27:38that is because when we're looking at
- 35:27:40data it's designed to process a large
- 35:27:43amount of data coming in. We don't want
- 35:27:45to just process one line at a time. And
- 35:27:47so in this case, I'm processing two
- 35:27:48lines. And then I'm just going to print
- 35:27:50and you'll see clf.predict new data. So
- 35:27:53the CLF and the predict part is going to
- 35:27:56give us an answer. And let's see what
- 35:27:57that looks like. And you'll see 01. So
- 35:28:00predicted the first one, the 34 is going
- 35:28:02to be on the one side and the 56 is
- 35:28:04going to be on the other side. So one
- 35:28:06came out as a alligator and one came out
- 35:28:08as a crocodile. Now that's pretty short
- 35:28:10explanation for the setup, but really we
- 35:28:12want to dug in and see what it's going
- 35:28:14on behind the scenes. and let's see what
- 35:28:16that looks like. So, the next step is to
- 35:28:19dig in deep and find out what's going on
- 35:28:22behind the scenes and also put that in a
- 35:28:24nice pretty graph. We're going to spend
- 35:28:26more work on this than we did actually
- 35:28:28generating the original model. And
- 35:28:30you'll see here that we go through a few
- 35:28:32steps and I'm I'll move this over to our
- 35:28:34editor in just a second. We come in, we
- 35:28:36create our original data. It's exactly
- 35:28:38identical to the first part and I'll
- 35:28:40explain why we redid that and show you
- 35:28:41how not to redo that. And then we're
- 35:28:43going to go in there and add in those
- 35:28:45lines. We're going to see what those
- 35:28:47lines look like and how to set those up.
- 35:28:49And finally, we're going to plot all
- 35:28:51that on here and show it. And you'll get
- 35:28:53a nice graph with the what we saw
- 35:28:55earlier when we were going through the
- 35:28:56theory behind this where it shows the
- 35:28:59support vectors and the hyper plane. And
- 35:29:02those are done where you can see the
- 35:29:03support vectors as the dash lines and
- 35:29:05the solid line which is the hyper plane.
- 35:29:07Let's get that into our Jupyter
- 35:29:09notebook. Before I scroll down to a new
- 35:29:12line, I want you to notice line 13. It
- 35:29:15has plot show. And we're going to talk
- 35:29:17about that here in just a second. But
- 35:29:18let's scroll down to a new line down
- 35:29:20here. And I'm going to paste that code
- 35:29:22in. And you'll see that the plot show
- 35:29:24has moved down below. Let's scroll up a
- 35:29:26little bit. And if you look at the top
- 35:29:28here of our new section, 1 2 3 and four
- 35:29:32is the same code we had before. And
- 35:29:34let's go back up here and take a look at
- 35:29:36that. We're going to fit the values on
- 35:29:38our SVM. And then we're going to plot
- 35:29:40scatter it. And then we're going to do a
- 35:29:42plot show. So you should be asking why
- 35:29:44are we redoing the same code. Well, when
- 35:29:47you do the plot show, that blanks out
- 35:29:49what's in the plot. So once I've done
- 35:29:51this plot show, I have to reload that
- 35:29:53data. Now, we could do this simply by
- 35:29:55removing it up here, rerunning it, and
- 35:29:58then coming down here, and then we
- 35:30:00wouldn't have to rerun these first four
- 35:30:01lines of code. Now, in this, it doesn't
- 35:30:04matter too much. And you'll see the plot
- 35:30:05show is down here and then removed right
- 35:30:07there on line five. I'll go ahead and
- 35:30:10just delete that out of there because we
- 35:30:11don't want to blank out our screen. We
- 35:30:13want to move on to the next setup. So,
- 35:30:15we can go ahead and just skip the first
- 35:30:17four lines because we did that before.
- 35:30:19And let's take a look at the ax=
- 35:30:22plt.gca.
- 35:30:24Now, right now, we're actually spending
- 35:30:25a lot of time just graphing. That's all
- 35:30:27we're doing here. Okay. So, this is how
- 35:30:29we display a nice graph with our results
- 35:30:32and our data. AX is very standard not
- 35:30:35used variable when you're talking about
- 35:30:36PLT and it's just setting it to that
- 35:30:39axis the last axis in the PLT. It can
- 35:30:41get very confusing if you're working
- 35:30:43with many different layers of data on
- 35:30:45the same graph and this makes it very
- 35:30:46easy to reference the ax. So this
- 35:30:49reference is looking at the PLT that we
- 35:30:52created and we already mapped out our
- 35:30:54two blobs on. And then we want to know
- 35:30:56the limits. So we want to know how big
- 35:30:58the graph is. And we can find out the x
- 35:31:00limit and the y limit simply with the
- 35:31:02get x limit and get yimit commands which
- 35:31:04is part of our metplot library. And then
- 35:31:07we're going to create a grid. And you'll
- 35:31:09see down here we have we've set the
- 35:31:11variable xx equal to npines space ximit
- 35:31:150 ximit 1a 30. And we've done the same
- 35:31:18thing for the yspace. And then we're
- 35:31:20going to go in here and we create a mesh
- 35:31:22grid. And this is a numpy command. So
- 35:31:25we're back to our numbers python. Let's
- 35:31:28go through what these numpy commands
- 35:31:30mean with the line space in the mesh
- 35:31:32grid. We've taken xx small xx= np line
- 35:31:36space. And we have our x limit zero and
- 35:31:38our x limit one and we're going to
- 35:31:40create 30 points on it. And we're going
- 35:31:43to do the same thing for the y axis. Now
- 35:31:45this has nothing to do with our
- 35:31:46evaluation. It's uh all we're doing is
- 35:31:48we're creating a grid of data. And so
- 35:31:52we're creating a set of points between
- 35:31:53zero and the x limit. We're creating 30
- 35:31:56points. And the same thing with the y.
- 35:31:58And then the mesh grid loops those all
- 35:32:00together. So it forms a nice grid. So if
- 35:32:02we were going to do this say between the
- 35:32:04limit 0 and 10 and do 10 points, we
- 35:32:06would have a 0 0 1 1 0 1 02 03 04 to 10
- 35:32:12and so on. You can just imagine a point
- 35:32:14at each corner one of those boxes. And
- 35:32:16the mesh grid combines them all. So we
- 35:32:18take the y and the xx we created and
- 35:32:20creates the full grid. And we've set
- 35:32:21that grid into the y coordinates and the
- 35:32:24xx coordinates. Now remember, when we're
- 35:32:26working with Numbi in Python, we like to
- 35:32:29separate those. We like to have instead
- 35:32:30of it being x comma 1, you know, x comma
- 35:32:34y and then x2 comma y2 and in the next
- 35:32:38set of data, it would be a column of x's
- 35:32:40and a column of y's. And that's what we
- 35:32:42have here is we have a column of y's. We
- 35:32:44put it as a capital y y and a column of
- 35:32:46x's, capital xx with all those different
- 35:32:49points being listed. And finally, we get
- 35:32:51down to the numpy vstack. Just as we
- 35:32:54created those in the mesh grid, we're
- 35:32:57now going to put them all into one
- 35:32:59array, XY array. Now that we've created
- 35:33:01the stack of data points, we're going to
- 35:33:04do something interesting here. We're
- 35:33:05going to create a value Z. And the Z
- 35:33:08equals the CLF. That's our uh that's our
- 35:33:11support vector machine we created and
- 35:33:13we've already trained. And we have a
- 35:33:16decision function. And we're going to
- 35:33:17put the XY in there. So here we have all
- 35:33:19this data. We're going to put that XY in
- 35:33:22there, that data, and we're going to
- 35:33:23reshape it. And you'll see that we have
- 35:33:25the xx.shape in here. This literally
- 35:33:28takes the xx, resets it up, connected to
- 35:33:31the y, and the zvalue lets us know
- 35:33:34whether it is the left hand side. It's
- 35:33:37going to generate three different
- 35:33:38values. The zvalue does, and it'll tell
- 35:33:40us whether that data is a support vector
- 35:33:43to the left, the hyper plane in the
- 35:33:45middle, or the support vector to the
- 35:33:47right. So it generates three different
- 35:33:49values for each of those points. And
- 35:33:50those points have been reshaped so
- 35:33:52they're right on a line on those three
- 35:33:54different lines. So we've set all of our
- 35:33:56data up. We've labeled it to three
- 35:33:58different areas and we've reshaped it.
- 35:34:00And we've just taken 30 points in each
- 35:34:02direction. If you do the math, you have
- 35:34:0430 * 30. So that's 900 points of data.
- 35:34:07And we separated it between the three
- 35:34:08lines and reshaped it to fit those three
- 35:34:10lines. We can then go back to our map
- 35:34:12plot library where we've created the AX
- 35:34:15and we're going to create a contour. And
- 35:34:16you'll see here where we have contour,
- 35:34:18capital XX, capital Y, Y. These have
- 35:34:21been reshaped to fit those lines. Z is
- 35:34:23the labels. So now we have the three
- 35:34:25different points with the labels in
- 35:34:26there. And we can set the colors equals
- 35:34:28K. And I told you we had three different
- 35:34:30labels, but we have uh three levels of
- 35:34:33data. The alpha is just makes it kind of
- 35:34:36see-through. So it's only uh 0.5 of the
- 35:34:38value in there. So when we graph it, the
- 35:34:40data will show up from behind it,
- 35:34:41wherever the lines go. And finally, the
- 35:34:43line styles. This is where we set the
- 35:34:46two support vectors to be dash dash
- 35:34:49lines and then a single one is just a
- 35:34:51straight line. That's what all that
- 35:34:53setup does. And then finally, we take
- 35:34:55our ax.scatter. We're going to go ahead
- 35:34:58and plot the support vectors, but we've
- 35:35:00programmed it in there so that they look
- 35:35:02nice like the dash dash line and the
- 35:35:04dash line on that grid. And you can see
- 35:35:06here when we do the CLFS support
- 35:35:09vectors, we are looking at column zero
- 35:35:11and column one. And then again we have
- 35:35:14the S equals 100. So we're going to make
- 35:35:16them larger. And the line width equals
- 35:35:181, face colors equals none. Let's take a
- 35:35:20look and see what that looks like when
- 35:35:21we show it. And you can see when we get
- 35:35:23down to our end result, it creates a
- 35:35:25really nice graph. We have our two
- 35:35:28support vectors and dash lines. And they
- 35:35:30have the near data. So you can see those
- 35:35:32two points or in this case the four
- 35:35:34points where those lines nicely cleave
- 35:35:36the data. And then you have your hyper
- 35:35:38plane down the middle which is as far
- 35:35:40from the two different points as
- 35:35:41possible creating the maximum distance.
- 35:35:43So you can see that we have our nice
- 35:35:45output for the size of the body and the
- 35:35:47width of the snout and we've easily
- 35:35:49separated the two groups of crocodile
- 35:35:51and alligator. Congratulations. You've
- 35:35:54done it. We've made it. Of course, these
- 35:35:55are pretend data for our crocodiles and
- 35:35:58alligators. But this hands-on example
- 35:36:00will help you to encounter any support
- 35:36:02vector machine projects in the future.
- 35:36:04And you can see how easy they are to set
- 35:36:06up and look at in depth. We're going to
- 35:36:08cover the K nearest neighbors a lot
- 35:36:10referred to as KNN. And KNN is really a
- 35:36:14fundamental place to start in the
- 35:36:16machine learning. It's a basis of a lot
- 35:36:18of other things and just the logic
- 35:36:19behind it is easy to understand and
- 35:36:21incorporated in other forms of machine
- 35:36:24learning. So today, what's in it for
- 35:36:26you? Why do we need KNN? What is KN&N?
- 35:36:30How do we choose the factor K? When do
- 35:36:33we use KNN? How does KN&N algorithm
- 35:36:37work? And then we'll dive in to my
- 35:36:40favorite part, the use case. Predict
- 35:36:42whether a person will have diabetes or
- 35:36:44not. That is a very common and popular
- 35:36:46used data set as far as testing out
- 35:36:50models and learning how to use the
- 35:36:52different models in machine learning. By
- 35:36:54now, we all know machine learning models
- 35:36:56make predictions by learning from the
- 35:36:58past data available. So we have our
- 35:37:00input values. Our machine learning model
- 35:37:02builds on those inputs of what we
- 35:37:04already know and then we use that to
- 35:37:06create a predicted output. Is that a
- 35:37:09dog? Little kid looking over there and
- 35:37:11watching the black cat cross their path.
- 35:37:14No, dear. You can differentiate between
- 35:37:16a cat and a dog based on their
- 35:37:18characteristics.
- 35:37:20Cats. Cats have sharp claws, uses to
- 35:37:23climb, smaller length of ears, meows and
- 35:37:25purr. Doesn't love to play around. dogs.
- 35:37:29They have dull claws, bigger length of
- 35:37:30ears, barks, loves to run around. You
- 35:37:33usually don't see a cat running around
- 35:37:35people, although I do have a cat that
- 35:37:36does that where dogs do. And we can look
- 35:37:38at these. We can say uh we can evaluate
- 35:37:40the sharpness of the claws. How sharp
- 35:37:42are their claws? And we can evaluate the
- 35:37:45length of the ears. And we can usually
- 35:37:47sort out cats from dogs based on even
- 35:37:49those two characteristics. Now, tell me
- 35:37:52if it is a cat or a dog. Not question.
- 35:37:54Usually little kids know cats and dogs
- 35:37:56by now. unless you live a place where
- 35:37:58there's not many cats or dogs. So, if we
- 35:38:00look at the sharpness of the claws, the
- 35:38:01length of the ears, and we can see that
- 35:38:03the cat has smaller ears and sharper
- 35:38:06claws than the other animals. Its
- 35:38:09features are more like cats. It must be
- 35:38:11a cat. Sharp claws, length of ears, and
- 35:38:14it goes in the cat group. Because KN&N
- 35:38:17is based on feature similarity, we can
- 35:38:19do classification using KN&N classifier.
- 35:38:22So, we have our input value, the picture
- 35:38:24of the black cat. It goes into our
- 35:38:26trained model and it predicts that this
- 35:38:28is a cat coming out. So what is knn?
- 35:38:31What is the kn&n algorithm? K nearest
- 35:38:35neighbors is what that stands for. Is
- 35:38:37one of the simplest supervised machine
- 35:38:39learning algorithms mostly used for
- 35:38:41classification. So we want to know is
- 35:38:44this a dog or it's not a dog? Is it a
- 35:38:46cat or not a cat? It classifies a data
- 35:38:49point based on how its neighbors are
- 35:38:51classified. KN&N stores all available
- 35:38:53cases and classifies new cases based on
- 35:38:56a similarity measure. And here we've
- 35:38:58gone from cats and dogs right into wine.
- 35:39:01Another favorite of mine. KN&N stores
- 35:39:03all available cases and classifies new
- 35:39:05cases based on a similarity measure. And
- 35:39:07here you see we have a measurement of
- 35:39:09sulfur dioxide versus the chloride level
- 35:39:12and then the different wines they've
- 35:39:13tested and where they fall on that graph
- 35:39:15based on how much sulfur dioxide and how
- 35:39:17much chloride. K and K&N is a perimeter
- 35:39:19that refers to the number of nearest
- 35:39:21neighbors to include in the majority of
- 35:39:23the voting process. And so if we add a
- 35:39:25new glass of wine there, red or white,
- 35:39:27we want to know what the neighbors are.
- 35:39:29In this case, we're going to put K
- 35:39:30equals 5. We'll talk about K in just a
- 35:39:33minute. A data point is classified by
- 35:39:35the majority of votes from its five
- 35:39:36nearest neighbors. Here, the unknown
- 35:39:39point would be classified as red since
- 35:39:41four out of five neighbors are red. So,
- 35:39:43how do we choose K? How do we know K
- 35:39:46equals 5? I mean that's was the value we
- 35:39:48put in there. I said we're going to talk
- 35:39:49about it. How do we choose the factor K?
- 35:39:52KN&N algorithm is based on feature
- 35:39:54similarity. Choosing the right value of
- 35:39:56K is a process called parameter tuning
- 35:39:59and is important for better accuracy. So
- 35:40:02at K equals 3, we can classify we have a
- 35:40:04question mark in the middle as either a
- 35:40:06as a square or not. Is it a square or is
- 35:40:08it in this case a triangle? And so if we
- 35:40:10set K equals to three, we're going to
- 35:40:12look at the three nearest neighbors.
- 35:40:14We're going to say this is a square. And
- 35:40:16if we put k equals a 7, we classify as a
- 35:40:19triangle depending on what the other
- 35:40:21data is around it. And you can see as
- 35:40:22the k changes depending on where that
- 35:40:24point is, that drastically changes your
- 35:40:26answer. And uh we jump here. We go, how
- 35:40:29do we choose the factor of k? You'll
- 35:40:31find this in all machine learning.
- 35:40:33Choosing these factors, that's the face
- 35:40:35you get. It's like, oh my gosh, did I
- 35:40:37choose the right K? Did I set it right
- 35:40:39my values in whatever machine learning
- 35:40:41tool you're looking at? so that you
- 35:40:42don't have a huge bias in one direction
- 35:40:45or the other. And in terms of KNN, the
- 35:40:48number of K, if you choose it too low,
- 35:40:50the bias is based on it's just too
- 35:40:52noisy. It's it's right next to a couple
- 35:40:54things and it's going to pick those
- 35:40:56things and you might get a skewed
- 35:40:57answer. And if your K is too big, then
- 35:41:00it's going to take forever to process.
- 35:41:02So you're going to run into processing
- 35:41:03issues and resource issues. So what we
- 35:41:06do the most common use and there's other
- 35:41:08options for choosing k is to use the
- 35:41:11square root of n. So n is a total number
- 35:41:14of values you have you take the square
- 35:41:16root of it. In most cases you also if
- 35:41:18it's an even number so if you're using
- 35:41:20uh like in this case squares and
- 35:41:22triangles if it's even you want to make
- 35:41:24your k value odd. That helps it select
- 35:41:27better. So in other words you're not
- 35:41:28going to have a balance between two
- 35:41:30different factors that are equal. So
- 35:41:32usually take the square root of n and if
- 35:41:34it's even you add one to it or subtract
- 35:41:36one from it and that's where you get the
- 35:41:37k value from that is the most common use
- 35:41:39and it's pretty solid. It works very
- 35:41:41well. When do we use kn? We can use kn
- 35:41:45when data is labeled. So you need a
- 35:41:47label on it. We know we have a group of
- 35:41:49pictures with dogs cats cats. Data is
- 35:41:52noisefree. And so you can see here when
- 35:41:55we have a class and we have like
- 35:41:57underweight 140 23 Hello kitty normal
- 35:42:00that's pretty confusing. We have a a
- 35:42:02high variety of data coming in. So it's
- 35:42:04very noisy and that would cause an
- 35:42:06issue. Data set is small. So we're
- 35:42:08usually working with smaller data sets
- 35:42:10where you might get into gig of data if
- 35:42:13it's really clean. It doesn't have a lot
- 35:42:14of noise because KN&N is a lazy learner.
- 35:42:17I.e. it doesn't learn a discriminative
- 35:42:19function from the training set. So it's
- 35:42:21very lazy. So if you have very
- 35:42:23complicated data and you have a large
- 35:42:25amount of it, you're not going to use
- 35:42:26the kn. But it's really great to get a
- 35:42:28place to start. Even with large data,
- 35:42:30you can sort out a small sample and get
- 35:42:32an idea of what that looks like using
- 35:42:33the KN&N and also just using for smaller
- 35:42:36data sets. KN&N works really good. How
- 35:42:39does the KN&N algorithm work? Consider a
- 35:42:42data set having two variables, height in
- 35:42:44centimeters and weight in kilograms. And
- 35:42:47each point is classified as normal or
- 35:42:49underweight. So we can see right here we
- 35:42:51have two variables, you know, true
- 35:42:53false. They're either normal or they're
- 35:42:55not. They're underweight. On the basis
- 35:42:56of the given data, we have to classify
- 35:42:59the below set as normal or underweight
- 35:43:01using KN&N. So if we have new data
- 35:43:03coming in that says 57 kg and 177 cm, is
- 35:43:08that going to be normal or underweight?
- 35:43:10To find the nearest neighbors, we'll
- 35:43:12calculate the ukitian distance.
- 35:43:14According to the uklitian distance
- 35:43:16formula, the distance between two points
- 35:43:18in the plane with the coordinates xy and
- 35:43:21ab is given by distance d equals the
- 35:43:24square root of x - a^2 + y - b^2. And
- 35:43:29you can remember that from the two edges
- 35:43:31of a triangle. We're computing the third
- 35:43:33edge since we know the x side and the y
- 35:43:36side. Let's calculate it to understand
- 35:43:38clearly. So we have our unknown point
- 35:43:40and we placed it there in red. And we
- 35:43:42have our other points where the data is
- 35:43:44scattered around. The distance d1 is the
- 35:43:47square<unk> of 170 minus 167^ squar + 57
- 35:43:51- 51^ 2ar which is about 6.7 and
- 35:43:55distance 2 is about 13 and distance 3 is
- 35:43:59about 13.4. Similarly, we will calculate
- 35:44:03the ukitian distance of unknown data
- 35:44:05point from all the points in the data
- 35:44:07set. And because we're dealing with
- 35:44:08small amount of data, that's not that
- 35:44:10hard to do and it's actually pretty
- 35:44:11quick for a computer and it's not a
- 35:44:13really complicated math. You can just
- 35:44:14see how close is the data based on the
- 35:44:16uklidian distance. Hence, we have
- 35:44:18calculated the uklidian distance of
- 35:44:20unknown data point from all the points
- 35:44:22as shown where x1 and y1 equal 57 and
- 35:44:26170 whose class we have to classify. So
- 35:44:29now we're looking at that. We're saying
- 35:44:30well here's the ukitian distance. Who's
- 35:44:32going to be their closest neighbors? Now
- 35:44:34let's calculate the nearest neighbor at
- 35:44:36k equals 3. And we can see the three
- 35:44:39closest neighbors puts them at normal.
- 35:44:41And that's pretty self-evident when you
- 35:44:43look at this graph. It's pretty easy to
- 35:44:44say okay what you know we're just voting
- 35:44:46normal normal normal. Three votes for
- 35:44:48normal. This is going to be a normal
- 35:44:49weight. So majority of neighbors are
- 35:44:51pointing towards normal. Hence as per
- 35:44:53KN&N algorithm the class of 571 170
- 35:44:56should be normal. So a recap of KN&N
- 35:44:59positive integer K is specified along
- 35:45:02with a new sample. We select the K
- 35:45:04entries in our database which are
- 35:45:05closest to the new sample. We find the
- 35:45:07most common classification of these
- 35:45:09entries. This is the classification we
- 35:45:11give to the new sample. So, as you can
- 35:45:13see, it's pretty straightforward. We're
- 35:45:14just looking for the closest things that
- 35:45:16match what we got. So, let's take a look
- 35:45:18and see what that looks like in a use
- 35:45:20case in Python. So, let's dive into the
- 35:45:23predict diabetes use case. So, use case,
- 35:45:26predict diabetes. The objective, predict
- 35:45:29whether a person will be diagnosed with
- 35:45:31diabetes or not. We have a data set of
- 35:45:34768 people who were or were not
- 35:45:37diagnosed with diabetes. And let's go
- 35:45:39ahead and open that file and just take a
- 35:45:41look at that data. And this is in a
- 35:45:43simple spreadsheet format. The data
- 35:45:46itself is commaepparated. Very common
- 35:45:48set of data. And it's also a very common
- 35:45:50way to get the data. And you can see
- 35:45:51here we have columns A through I. That's
- 35:45:54what 1 2 3 4 5 6 7 8. um eight columns
- 35:45:59with a particular attribute and then the
- 35:46:02ninth column which is the outcome is
- 35:46:04whether they have diabetes. As a data
- 35:46:06scientist, the first thing you should be
- 35:46:07looking at is insulin. Well, you know,
- 35:46:09if someone has insulin, they have
- 35:46:11diabetes because that's why they're
- 35:46:12taking it. And that could cause issue in
- 35:46:13some of the machine learning packages,
- 35:46:15but for very basic setup, this works
- 35:46:17fine for doing the KNN. And the next
- 35:46:20thing you notice is it it didn't take
- 35:46:22very much to open it up. Um I can scroll
- 35:46:24down to the bottom of the data. There's
- 35:46:25768.
- 35:46:26It's pretty much a small data set. You
- 35:46:28know, at 769, I can easily fit this into
- 35:46:32my RAM on my computer. I can look at it.
- 35:46:35I can manipulate it. And it's not going
- 35:46:37to really tax just a regular desktop
- 35:46:39computer. You don't even need an
- 35:46:40enterprise version to run a lot of this.
- 35:46:42So, let's start with importing all the
- 35:46:44tools we need. And before that, of
- 35:46:46course, we need to discuss what IDE I'm
- 35:46:48using. Certainly, you can use any uh
- 35:46:50particular editor for Python, but I like
- 35:46:52to use for doing uh very basic visual
- 35:46:55stuff. the Anaconda, which is great for
- 35:46:57doing demos with the Jupyter Notebook.
- 35:47:00And just a quick view of the Anaconda
- 35:47:02Navigator, which is the new release out
- 35:47:04there, which is really nice. You can see
- 35:47:06under home, I can choose my application.
- 35:47:09We're going to be using Python 3.6. I
- 35:47:11have a couple different uh versions on
- 35:47:13this particular machine. If I go under
- 35:47:15environments, I can create a unique
- 35:47:16environment for each one, which is nice.
- 35:47:18And there's even a little button there
- 35:47:20where I can install different packages.
- 35:47:21So, if I click on that button and open
- 35:47:23the terminal, I can then use a simple
- 35:47:24pip install to install different
- 35:47:26packages I'm working with. Let's go
- 35:47:28ahead and go back under home and we're
- 35:47:29going to launch our notebook. And I've
- 35:47:31already, you know, kind of like uh the
- 35:47:33old cooking shows, I've already prepared
- 35:47:34a lot of my stuff. So, we don't have to
- 35:47:36wait for it to launch because it takes a
- 35:47:37few minutes for it to open up a browser
- 35:47:40window. In this case, I'm going to it's
- 35:47:42going to open up Chrome because that's
- 35:47:43my default that I use. And since the
- 35:47:45script is pre-done, you'll see I have a
- 35:47:46number of windows open up at the top,
- 35:47:48the one we're working in. And uh since
- 35:47:50we're working on the KN&N predict
- 35:47:53whether a person will have diabetes or
- 35:47:54not. Let's go and put that title in
- 35:47:56there. And I'm also going to go up here
- 35:47:58and click on cell. Actually, we want to
- 35:48:00go ahead and first insert a cell below.
- 35:48:02And then I'm going to go back up to the
- 35:48:04top cell. And I'm going to change the
- 35:48:06cell type to markdown. That means this
- 35:48:09is not going to run as Python. It's a
- 35:48:10markdown language. So if I run this
- 35:48:12first one, it comes up in nice big
- 35:48:13letters, which is kind of nice. Remind
- 35:48:15us what we're working on. And by now you
- 35:48:18should be familiar with doing all of our
- 35:48:19imports. We're going to import the
- 35:48:21pandas as pd import numpy is np. Pandas
- 35:48:25is the uh pandas data frame and numpy is
- 35:48:28a number array. Very powerful tools to
- 35:48:30use in here. So we have our imports. So
- 35:48:33we've brought in our pandas or numpy our
- 35:48:35two general python tools. And then you
- 35:48:38can see over here we have our train test
- 35:48:40split. By now you should be familiar
- 35:48:42with splitting the data. We want to
- 35:48:44split part of it for training our thing
- 35:48:45and then training our particular model
- 35:48:48and then we want to go ahead and test
- 35:48:49the remaining data to see how good it
- 35:48:51is. Pre-processing a standard scaler
- 35:48:54pre-processor so we don't have a bias of
- 35:48:56really large numbers. Remember in the
- 35:48:58data we had like number of pregnancies
- 35:49:00isn't going to get very large where the
- 35:49:02amount of insulin they take and get up
- 35:49:03to 256. So 256 versus six that will skew
- 35:49:08results. So we want to go ahead and
- 35:49:09change that so they're all uniform
- 35:49:11between minus1 and one. And then the
- 35:49:13actual tool. This is the K neighbors
- 35:49:15classifier we're going to use. And
- 35:49:18finally, the last three are three tools
- 35:49:21to test. All about testing our model.
- 35:49:23How good is it? We just put down test on
- 35:49:25there. And we have our confusion matrix,
- 35:49:27our F1 score, and our accuracy. So we
- 35:49:29have our two general Python modules
- 35:49:32we're importing. And then we have our
- 35:49:34six modules specific from the sklearn
- 35:49:37setup. And then we do need to go ahead
- 35:49:39and run this. So these are actually
- 35:49:42imported. There we go. And then move on
- 35:49:44to the next step. And so in this set,
- 35:49:46we're going to go ahead and load the
- 35:49:47database. We're going to use pandas.
- 35:49:49Remember pandas is pd. And we'll take a
- 35:49:51look at the data in Python. We looked at
- 35:49:53it in a simple spreadsheet, but usually
- 35:49:55I like to also pull it up so that we can
- 35:49:57see what we're doing. So here's our data
- 35:49:59set equals PD read CSV. That's a pandas
- 35:50:03command. And the diabetes folder I just
- 35:50:06put in the same folder where my IPython
- 35:50:08script is. If you put in a different
- 35:50:10folder, you'd need the full length on
- 35:50:12there. We can also do a quick length of
- 35:50:15uh the data set. That is a simple Python
- 35:50:18command. Leen for length. We might even
- 35:50:20let's go ahead and print that. We'll go
- 35:50:21print. And if you do it on its own line,
- 35:50:24length data set in the Jupyter notebook,
- 35:50:26it'll automatically print it. But when
- 35:50:28you're in most of your different setups,
- 35:50:30you want to do the print in front of
- 35:50:31there. And then we want to take a look
- 35:50:33at the actual data set. And since we're
- 35:50:34in pandas, we can simply do data set
- 35:50:37head. And again, let's go ahead and add
- 35:50:40the print in there. If you put a bunch
- 35:50:42of these in a row, you know that data
- 35:50:44set one head, data set two head, it only
- 35:50:46prints out the last one. So, I usually
- 35:50:48always like to keep the print statement
- 35:50:49in there. But because most projects only
- 35:50:52use one data frame, Panda's data frame,
- 35:50:54doing it this way doesn't really matter.
- 35:50:56The other way works just fine. And you
- 35:50:58can see when we hit the run button, we
- 35:51:00have the 768 lines, which we knew, and
- 35:51:02we have our pregnancies. It's
- 35:51:04automatically given a label on the left.
- 35:51:06Remember the head only shows the first
- 35:51:08five lines. So we have zero through
- 35:51:11four. And just a quick look at the data.
- 35:51:13You can see it matches what we looked at
- 35:51:15before. We have pregnancy, glucose,
- 35:51:17blood pressure all the way to age. And
- 35:51:20then the outcome on the end. And we're
- 35:51:22going to do a couple things in this next
- 35:51:24step. We're going to create a list of
- 35:51:26columns where we can't have zero.
- 35:51:28There's no such thing as zero skin
- 35:51:30thickness or zero blood pressure, zero
- 35:51:33glucose. Uh any of those, you'd be dead.
- 35:51:36So, not a really good factor if they
- 35:51:37don't if they have a zero in there
- 35:51:39because they didn't have the data. And
- 35:51:40we'll take a look at that because we're
- 35:51:41going to start replacing that
- 35:51:42information with a couple of different
- 35:51:45things. And let's see what that looks
- 35:51:46like. So, first we create a nice list.
- 35:51:49As you can see, we have the values
- 35:51:51talked about glucose, blood pressure,
- 35:51:52skin thickness. Uh, and this is a nice
- 35:51:54way when you're working with columns is
- 35:51:56to list the columns you need to do some
- 35:51:58kind of transformation on. Uh, very
- 35:51:59common thing to do. And then for this
- 35:52:01particular setup, we certainly could use
- 35:52:03the there's some Panda tools that will
- 35:52:06do a lot of this where we can replace
- 35:52:07the NA, but we're going to go ahead and
- 35:52:10do it as a data set column equals data
- 35:52:13set column.replace. This is this is
- 35:52:15still pandas. You can do a direct.
- 35:52:17There's also one that that you look for
- 35:52:19your nan. A lot of different options in
- 35:52:21here. But the nan numpan is what that
- 35:52:24stands for is non doesn't exist. So the
- 35:52:27first thing we're doing here is we're
- 35:52:29replacing the zero with a numpy none.
- 35:52:33There's no data there. That's what that
- 35:52:34says. That's what this is saying right
- 35:52:36here. So put the zero in and we're going
- 35:52:38to replace zeros with no data. So if
- 35:52:41it's a zero, that means the person's
- 35:52:43well hopefully not dead. Hopefully they
- 35:52:44just didn't get the data. The next thing
- 35:52:45we want to do is we're going to create
- 35:52:47the mean which is the in integer from
- 35:52:49the data set from the column mean where
- 35:52:52we skip NAS. We can do that. That is a
- 35:52:55pandas command there, the skip na. So
- 35:52:57we're going to figure out the mean of
- 35:52:59that data set. And then we're going to
- 35:53:00take that data set column and we're
- 35:53:02going to replace all the npnan
- 35:53:06with the means. Why did we do that? And
- 35:53:08we could have actually just uh taken
- 35:53:10this step and gone right down here and
- 35:53:11just replace zero and skip anything
- 35:53:13where except you could actually there's
- 35:53:15a way to skip zeros and then just
- 35:53:16replace all the zeros. But in this case,
- 35:53:18we want to go ahead and do it this way.
- 35:53:19So you could see that we're switching
- 35:53:21this to a non-existent value. Then we're
- 35:53:23going to create the mean. Well, this is
- 35:53:25the average person. So if we don't know
- 35:53:28what it is, if they did not get the data
- 35:53:30and the data is missing, one of the
- 35:53:32tricks is you replace it with the
- 35:53:34average. What is the most common data
- 35:53:36for that? This way you can still use the
- 35:53:39rest of those values to do your
- 35:53:40computation and it kind of just brings
- 35:53:43that particular value or those missing
- 35:53:44values out of the equation. Let's go
- 35:53:46ahead and take this and we'll go ahead
- 35:53:48and run it. Doesn't actually do
- 35:53:50anything. So we're still preparing our
- 35:53:52data. If you want to see what that looks
- 35:53:54like, we don't have anything in the
- 35:53:55first few lines, so it's not going to
- 35:53:57show up. But we certainly could look at
- 35:53:59a row. Let's do that. Let's go into our
- 35:54:01data set. Let's print a data set. And
- 35:54:04let's pick in this case, let's just do
- 35:54:07glucose. And if I run this, this is
- 35:54:10going to print all the different glucose
- 35:54:11levels going down. And we thankfully
- 35:54:14don't see anything in here that looks
- 35:54:16like missing data, at least on the ones
- 35:54:17it shows. You can see it skipped a bunch
- 35:54:19in the middle because that's what it
- 35:54:20does. If you have too many lines in
- 35:54:21Jupyter notebook, it'll skip a few and
- 35:54:23and go on to the next in a data set. Let
- 35:54:25me go and remove this. And we'll just
- 35:54:27zero out that. And of course, before we
- 35:54:30do any processing, before proceeding any
- 35:54:32further, we need to split the data set
- 35:54:34into our train and testing data. That
- 35:54:36way, we have something to train it with
- 35:54:37and something to test it on. And you're
- 35:54:40going to notice we did a little
- 35:54:40something here with the uh pandas
- 35:54:42database code. There we go. My drawing
- 35:54:45tool. We've added in this right here off
- 35:54:47the data set. And what this says is that
- 35:54:50the first one in pandas, this is from
- 35:54:52the PD pandas. It's going to say within
- 35:54:55the data set, we want to look at the eye
- 35:54:57location and it is all rows. That's what
- 35:54:59that says. So we're going to keep all
- 35:55:01the rows, but we're only looking at
- 35:55:02zero, column 0 to 8. Remember column 9.
- 35:55:06Here it is right up here. We printed it
- 35:55:07in here is outcome. Well, that's not
- 35:55:09part of the training data. That's part
- 35:55:10of the answer. Yeah, it's column 9, but
- 35:55:12it's listed as eight. Number eight. So 0
- 35:55:14to eight is nine columns. So uh eight is
- 35:55:17the value. And when you see it in here,
- 35:55:19zero, this is actually 0 to 7. It
- 35:55:22doesn't include the last one. And then
- 35:55:24we go down here to Y, which is our
- 35:55:25answer. And we want just the last one,
- 35:55:29just column 8. And you can do it this
- 35:55:31way with this particular notation. And
- 35:55:33then if you remember, we imported the
- 35:55:34train test split that's part of the
- 35:55:37sklearn right there. And we simply put
- 35:55:39in our X and our Y. We're going to do
- 35:55:42random state equals zero. You don't have
- 35:55:44to necessarily seed it. That's a seed
- 35:55:45number. I think the default is one when
- 35:55:47you seated it. I'd have to look that up.
- 35:55:49And then the test size. Test size is
- 35:55:510.2. That simply means we're going to
- 35:55:53take 20% of the data and put it aside so
- 35:55:55that we can test it later. That's all
- 35:55:57that is. And again, we're going to run
- 35:55:59it. Not very exciting. So far, we
- 35:56:01haven't had any print out other than to
- 35:56:02look at the data. But that is a lot of
- 35:56:04this is prepping this data. Once you
- 35:56:06prep it, the actual lines of code are
- 35:56:08quick and easy. And we're almost there.
- 35:56:10But the actual writing of our KN&N, we
- 35:56:12need to go ahead and do a scale the
- 35:56:14data. If you remember correctly, we're
- 35:56:16fitting the data in a standard scaler,
- 35:56:18which means instead of the data being
- 35:56:20from, you know, five to 303 in one
- 35:56:23column and the next column is 1 to six,
- 35:56:26we're going to set that all so that all
- 35:56:27the data is between minus1 and one.
- 35:56:30That's what that standard scaler does.
- 35:56:32Keeps it standardized. And we only want
- 35:56:34to fit the scaler with the training set,
- 35:56:37but we want to make sure the testing set
- 35:56:39is the X test going in is also
- 35:56:42transformed. So it's processing it the
- 35:56:45same. So here we go with our standard
- 35:56:47scaler. We're going to call it sc__x for
- 35:56:49the scaler. And we're going to import
- 35:56:51the standard scaler into this variable.
- 35:56:53And then our xrain equals sc_x.fit
- 35:56:58transform. So we're creating the scaler
- 35:57:00on the x-ra variable. And then our x
- 35:57:02test, we're also going to transform it.
- 35:57:04So we've trained and transformed the
- 35:57:06x-ra. And then the x test isn't part of
- 35:57:09that training. It isn't part of that of
- 35:57:11training the transformer. it just gets
- 35:57:13transformed. That's all it does. And
- 35:57:15again, we're going to go and run this.
- 35:57:16And if you look at this, we've now gone
- 35:57:18through these steps, all three of them.
- 35:57:21We've taken care of replacing our zeros
- 35:57:24for key columns that shouldn't be zero,
- 35:57:27and we've replaced that with the means
- 35:57:30of those columns. That way, that they
- 35:57:32fit right in with our data models. We've
- 35:57:34come down here, and we split the data.
- 35:57:36So, now we have our test data and our
- 35:57:39training data. And then we've taken and
- 35:57:41we've scaled the data. So all of our
- 35:57:43data going in. No, no, we don't tra we
- 35:57:45don't train the Y part, the Y train and
- 35:57:48Y test that never has to be trained.
- 35:57:51It's only the data going in. That's what
- 35:57:53we want to train in there. Then define
- 35:57:55the model using K neighbors classifier
- 35:57:57and fit the train data in the model. So
- 35:57:59we do all that data prep. And you can
- 35:58:01see down here we're only going to have a
- 35:58:03couple lines of code where we're
- 35:58:05actually building our model and training
- 35:58:07it. That's one of the cool things about
- 35:58:09Python and how far we've come. It's such
- 35:58:11an exciting time to be in machine
- 35:58:12learning because there's so many
- 35:58:13automated tools. Let's see. Before we do
- 35:58:16this, let's do a quick length of and
- 35:58:18let's do y. We want let's just do length
- 35:58:21of y. And we get 768. And if we import
- 35:58:26math, we do math dot square root. Let's
- 35:58:30do y train. There we go. It's actually
- 35:58:33supposed to be x train. Before we do
- 35:58:36this, let's go ahead and do import math
- 35:58:38and do math square root length of y
- 35:58:41test. And when I run that, we get
- 35:58:4312.409.
- 35:58:45I want to see show you where this number
- 35:58:46comes from. We're about to use 12 is an
- 35:58:48even number. So if you know if you're
- 35:58:50ever voting on things, remember the
- 35:58:52neighbors all vote. Don't want to have
- 35:58:53an even number of neighbors voting. So
- 35:58:55we want to do something odd. And let's
- 35:58:57just take one away. We'll make it 11.
- 35:58:59Let me delete this out of here. That's
- 35:59:00one of the reasons I love Jupyter
- 35:59:01Notebook because you can flip around and
- 35:59:03do all kinds of things on the fly. So,
- 35:59:05we'll go ahead and put in our
- 35:59:06classifier. We're creating our
- 35:59:07classifier now and it's going to be the
- 35:59:09K neighbors classifier. In neighbors
- 35:59:11equal 11. Remember, we did 12 - 1 for
- 35:59:1411. So, we have an odd number of
- 35:59:16neighbors. P= 2 because we're looking
- 35:59:18for is it are they diabetic or not? And
- 35:59:21we're using the ukitian metric. There
- 35:59:23are other means of measuring the
- 35:59:25distance. You could do like square
- 35:59:27square means value. There's all kinds of
- 35:59:29measure this, but the uklidian is the
- 35:59:31most common one and it works quite well.
- 35:59:33It's important to evaluate the model.
- 35:59:35Let's use the confusion matrix to do
- 35:59:37that. And we're going to use the
- 35:59:38confusion matrix. Wonderful tool. And
- 35:59:41then we'll jump into the F1 score. And
- 35:59:44finally, accuracy score, which is
- 35:59:46probably the most commonly used quoted
- 35:59:49number when you go into a meeting or
- 35:59:51something like that. So, let's go ahead
- 35:59:52and paste that in there. And we'll set
- 35:59:54the CM equal to confusion matrix. Y
- 35:59:57test, Y predict. So those are the two
- 35:59:59values we're going to put in there. And
- 36:00:01let me go ahead and run that and print
- 36:00:02it out. And the way you interpret this
- 36:00:05is you have the Y predicted, which would
- 36:00:08be your title up here. You can do uh
- 36:00:10let's just do Predicted
- 36:00:13across the top and actual going down.
- 36:00:17Actual. It's always hard to to write in
- 36:00:20here. Actual. That means that this
- 36:00:22column here down the middle, that's the
- 36:00:24important column. And it means that our
- 36:00:26prediction said 94 and prediction in the
- 36:00:30actual agreed on 94 and 32. This number
- 36:00:34here, the 13 and the 15, those are what
- 36:00:38was wrong. So you could have like three
- 36:00:40different if you're looking at this
- 36:00:41across three different variables instead
- 36:00:43of just two. You'd end up with a third
- 36:00:45row down here in the column going down
- 36:00:47the middle. So in the first case, we
- 36:00:48have the the and I believe the zero is a
- 36:00:5294 people who don't have diabetes. The
- 36:00:54prediction said that 13 of those people
- 36:00:56did have diabetes and were at high risk.
- 36:00:58And the 32 that had diabetes had
- 36:01:00correct, but our prediction said another
- 36:01:0315 out of that 15, it classified as
- 36:01:07incorrect. So you can see where that
- 36:01:09classification comes in and how that
- 36:01:12works on the confusion matrix. Then
- 36:01:14we're going to go ahead and print the F1
- 36:01:16score. Let me just run that. And you see
- 36:01:19we get a 69 in our F1 score. The F1
- 36:01:24takes into account both sides of the
- 36:01:26balance of false positives where if we
- 36:01:29go ahead and just do the accuracy
- 36:01:31account and that's what most people
- 36:01:33think of is it looks at just how many we
- 36:01:36got right out of how many we got wrong.
- 36:01:38So a lot of people when you're a data
- 36:01:39scientist and you're talking to other
- 36:01:41data scientists they're going to ask you
- 36:01:43what the F1 score the Fore is. If you're
- 36:01:45talking to the general public or the uh
- 36:01:48decision makers in the business, they're
- 36:01:50going to ask what the accuracy is. And
- 36:01:51the accuracy is always better than the
- 36:01:54F1 score. But the F1 score is more
- 36:01:56telling. It lets us know that there's
- 36:01:58more false positives than we would like
- 36:02:00on here. But 82% not too bad for a quick
- 36:02:03flash look at people's different
- 36:02:06statistics and running an sklearn and
- 36:02:08running the KNN, the K nearest neighbor
- 36:02:10on it. So we have created a model using
- 36:02:13KN&N which can predict whether a person
- 36:02:16will have diabetes or not or at the very
- 36:02:18least whether they should go get a
- 36:02:19checkup and have their glucose checked
- 36:02:21regularly or not. The print accuracy
- 36:02:24score we got the 0818 was pretty close
- 36:02:26to what we got and we can pretty much
- 36:02:28round that off and just say we have an
- 36:02:29accuracy of 80%. Tells us it is a pretty
- 36:02:32fair fit in the model.
- 36:02:34>> So what is game means clustering? C
- 36:02:37means clustering is an unsupervised
- 36:02:40learning algorithm. In this case, you
- 36:02:43don't have labeled data unlike in
- 36:02:45supervised learning. So you have a set
- 36:02:47of data and you want to group them and
- 36:02:50as the name suggests, you want to put
- 36:02:52them into clusters which means objects
- 36:02:55that are similar in nature, similar in
- 36:02:57characteristics need to be put together.
- 36:03:00So that's what K means clustering is all
- 36:03:03about. The term K is basically is a
- 36:03:07number. So we need to tell the system
- 36:03:09how many clusters we need to perform. So
- 36:03:11if K is equal to two, there will be two
- 36:03:13clusters. If K is equal to three, three
- 36:03:15clusters and so on and so forth. That's
- 36:03:17what the K stands for. And of course
- 36:03:19there is a way of finding out what is
- 36:03:21the best or optimum value of K for a
- 36:03:24given data. We will look at that. So
- 36:03:26that is K means clustering. So let's
- 36:03:30take an example. C means clustering is
- 36:03:32used in many many scenarios but let's
- 36:03:35take an example of cricket the game of
- 36:03:37cricket let's say you received data of a
- 36:03:40lot of players from maybe all over the
- 36:03:43country or all over the world and this
- 36:03:46data has information about the runs
- 36:03:49scored by the people or by the player
- 36:03:52and the wickets taken by the player and
- 36:03:55based on this information we need to
- 36:03:58cluster this data into two clusters
- 36:04:02batsmen and bowlers. So this is an
- 36:04:04interesting example. Let's see how we
- 36:04:06can perform this. So we have the data
- 36:04:09which consists of primarily two
- 36:04:12characteristics which is the runs and
- 36:04:15the wickets. So the bowlers basically
- 36:04:17take wickets and the batsmen score runs.
- 36:04:20There will be of course a few bowlers
- 36:04:22who can score some runs and similarly
- 36:04:25there will be some batsmen who will who
- 36:04:27would have taken a few wickets. But with
- 36:04:29this information, we want to cluster
- 36:04:31this players into batsmen and bowlers.
- 36:04:34So how does this work? Let's say this is
- 36:04:36how the data is. So there are
- 36:04:39information there is information on the
- 36:04:41y-axis about the run scored and on the
- 36:04:44x-axis about the wickets taken by the
- 36:04:46players. So if we do a quick plot, this
- 36:04:49is how it would look. And um when we do
- 36:04:53the clustering, we need to have the
- 36:04:55clusters like shown in the third diagram
- 36:04:59out here. We need to have a cluster
- 36:05:01which consists of people who have scored
- 36:05:04high runs which is basically the
- 36:05:06batsmen. And then we need a cluster with
- 36:05:08people who have taken a lot of wickets
- 36:05:11which is typically the bowlers. There
- 36:05:13may be a certain amount of overlap but
- 36:05:15we will not talk about it right now. So
- 36:05:18with K means clustering we will have
- 36:05:20here that means K is equal to two and we
- 36:05:23will have two clusters which is batsmen
- 36:05:25and bowlers. So how does this work? The
- 36:05:27way it works is the first step in K
- 36:05:30means clustering is the allocation of
- 36:05:33two centroidids randomly. So two points
- 36:05:37are assigned as so-called centrids. So
- 36:05:41in this case we want two clusters which
- 36:05:44means K is equal to two. So two points
- 36:05:47have been randomly assigned as centrids.
- 36:05:51Keep in mind these points can be
- 36:05:54anywhere. There are random points. They
- 36:05:56are not initially they are not really
- 36:05:59the centroidids. Centr means it's a
- 36:06:02central point of a given data set. But
- 36:06:04in this case when it starts off it's not
- 36:06:07really the centroid. Okay. So these
- 36:06:09points though in our presentation here
- 36:06:12we have shown them one point closer to
- 36:06:14these data points and another closer to
- 36:06:16these data points. They can be assigned
- 36:06:18randomly anywhere. Okay. So that's the
- 36:06:21first step. The next step is to
- 36:06:23determine the distance of each of the
- 36:06:26data points from each of the randomly
- 36:06:30assigned centrids. So for example we
- 36:06:32take this point and find the distance
- 36:06:35from this centr and the distance from
- 36:06:38this cent. This point is taken and the
- 36:06:40distance is found from this centroid and
- 36:06:42this c and so on and so forth. So for
- 36:06:45every point the distance is measured
- 36:06:48from both the centroids and then
- 36:06:51whichever distance is less that point is
- 36:06:54assigned to that centroid. So for
- 36:06:56example in this case visually it is very
- 36:06:59obvious that all these data points are
- 36:07:01assigned to this centroid and all these
- 36:07:04data points are assigned to this
- 36:07:05centroid and that's what is represented
- 36:07:07here in blue color and in this yellow
- 36:07:10color. The next step is to actually
- 36:07:12determine the central point or the
- 36:07:15actual centrid for these two clusters.
- 36:07:18So we have this one initial cluster,
- 36:07:21this one initial cluster. But as you can
- 36:07:23see these points are not really the
- 36:07:25centroid. Centroid means it should be
- 36:07:27the central position of this data set.
- 36:07:30Central position of this data set. So
- 36:07:32that is what needs to be determined as
- 36:07:35the next step. So the central point of
- 36:07:38the actual centrid is determined and the
- 36:07:41original randomly allocated centr is
- 36:07:43repositioned to the actual centroid of
- 36:07:46this new clusters and this process is
- 36:07:50actually repeated. Now what might happen
- 36:07:52is some of these points may get
- 36:07:55reallocated. In our example that is not
- 36:07:57happening probably but it may so happen
- 36:07:59that the distance is found between each
- 36:08:02of these data points once again with
- 36:08:04these centroidids. And if there is if it
- 36:08:06is required some points may be
- 36:08:08reallocated. We will see that in a later
- 36:08:10example but for now we will keep it
- 36:08:12simple. So this process is continued
- 36:08:16till the centrid repositioning stops and
- 36:08:20that is our final cluster. So this is
- 36:08:23our so after iteration we come to this
- 36:08:26position this situation where the
- 36:08:28centroid doesn't need any more
- 36:08:30repositioning and that means our
- 36:08:33algorithm has converged convergence has
- 36:08:36occurred and we have the cluster two
- 36:08:38clusters we have the clusters with a
- 36:08:41centroid. So this process is repeated.
- 36:08:45The process of calculating the distance
- 36:08:47and repositioning the centrid is
- 36:08:50repeated till the repositioning stops
- 36:08:53which means that the algorithm has
- 36:08:56converged and we have the final cluster
- 36:09:00with the data points and the
- 36:09:01centroidids. So this is what you're
- 36:09:03going to learn from this session. We
- 36:09:05will talk about the types of clustering.
- 36:09:08What is K means clustering? application
- 36:09:10of K means clustering. C means
- 36:09:12clustering is done using distance
- 36:09:14measure. So we will talk about the
- 36:09:17common distance measures and then we
- 36:09:19will talk about how K means clustering
- 36:09:22works and go into the details of K means
- 36:09:24clustering algorithm and then we will
- 36:09:27end with a demo and a use case for K
- 36:09:29means clustering. So let's begin. First
- 36:09:32of all, what are the types of
- 36:09:33clustering? There are primarily two
- 36:09:35categories of clustering. hierarchical
- 36:09:38clustering and then partitional
- 36:09:40clustering and each of these categories
- 36:09:42are further subdivided into elomerative
- 36:09:45and divisive clustering and K means and
- 36:09:48fuzzy C means clustering. Let's take a
- 36:09:50quick look at what each of these types
- 36:09:52of clustering are. In hierarchical
- 36:09:55clustering, the clusters have a treelike
- 36:09:58structure and hierarchical clustering is
- 36:10:01further divided into elomerative and
- 36:10:04divisive. Elomemerative clustering is a
- 36:10:07bottomup approach. We begin with each
- 36:10:09element as a separate cluster and merge
- 36:10:12them into successively larger clusters.
- 36:10:14So for example, we have A B CDE E F. We
- 36:10:18start by combining B and C form one
- 36:10:20cluster. D and E form one more. Then we
- 36:10:23combine D, E and F one more bigger
- 36:10:25cluster and then add BC to that and then
- 36:10:27finally A to it. Compared to that
- 36:10:30divisive clustering or divisive
- 36:10:32clustering is a top- down approach. We
- 36:10:34begin with the whole set and proceed to
- 36:10:36divide it into successively smaller
- 36:10:39clusters. So we have ABCDE E F. We first
- 36:10:42take that as a single cluster and then
- 36:10:44break it down into A B C D E and F. Then
- 36:10:49we have partitional clustering split
- 36:10:51into two subtypes. K means clustering
- 36:10:54and fuzzy C means. In K means clustering
- 36:10:58the objects are divided into the number
- 36:11:00of clusters mentioned by the number K.
- 36:11:04That's where the K comes from. So if we
- 36:11:06say K is equal to two, the objects are
- 36:11:08divided into two clusters C1 and C2. And
- 36:11:12the way it is done is the features or
- 36:11:14characteristics are compared and all
- 36:11:17objects having similar characteristics
- 36:11:19are clubed together. So that's how K
- 36:11:22means clustering is done. We will see it
- 36:11:24in more detail as we move forward. And
- 36:11:26fuzzy C means is very similar to K means
- 36:11:29in the sense that it clubs objects that
- 36:11:32have similar characteristics together.
- 36:11:34But while in K means clustering two
- 36:11:37objects cannot belong to or any object a
- 36:11:40single object cannot belong to two
- 36:11:41different clusters in C means objects
- 36:11:44can belong to more than one cluster. So
- 36:11:46that is the primary difference between K
- 36:11:49means and fuzzy C means. So what are
- 36:11:51some of the applications of K means
- 36:11:54clustering? C means clustering is used
- 36:11:56in a variety of examples or variety of
- 36:11:59business cases in real life starting
- 36:12:02from academic performance, diagnostic
- 36:12:04systems, search engines and wireless
- 36:12:07sensor networks and many more. So let us
- 36:12:10take a little deeper look at each of
- 36:12:11these examples. Academic performance. So
- 36:12:14based on the scores of the students,
- 36:12:16students are categorized into A, B, C
- 36:12:19and so on. Clustering forms a backbone
- 36:12:21of search engines. When a search is
- 36:12:24performed, the search results need to be
- 36:12:26grouped together. The search engines
- 36:12:28very often use clustering to do this.
- 36:12:31And similarly, in case of wireless
- 36:12:34sensor networks, the clustering
- 36:12:36algorithm plays the role of finding the
- 36:12:38cluster heads which collects all the
- 36:12:41data in its respective cluster. So
- 36:12:44clustering especially K means clustering
- 36:12:46uses distance measure. So let's take a
- 36:12:49look at what is distance measure. So
- 36:12:51while these are the different types of
- 36:12:53clustering in this video we will focus
- 36:12:56on K means clustering. So distance
- 36:12:59measure tells how similar some objects
- 36:13:03are. So the similarity is measured using
- 36:13:07what is known as distance measure and
- 36:13:09what are the various types of distance
- 36:13:11measures. There is ukidian distance.
- 36:13:14There is Manhattan distance. Then we
- 36:13:17have squared ukitian distance measure
- 36:13:19and cosine distance measure. These are
- 36:13:22some of the distance measures supported
- 36:13:24by k means clustering. Let's take a look
- 36:13:26at each of these. What is ukidian
- 36:13:29distance measure? This is nothing but
- 36:13:31the distance between two points. So we
- 36:13:33have learned in high school how to find
- 36:13:35the distance between two points. This is
- 36:13:38a little sophisticated formula for that.
- 36:13:40But we know a simpler one is square
- 36:13:43roo<unk> of y2 - y1 square + x2 - x1
- 36:13:48square. So this is an extension of that
- 36:13:50formula. So that is the ukidian distance
- 36:13:53between two points. What is the squared
- 36:13:56ukidian distance measure? It's nothing
- 36:13:58but the square of the ukidian distance
- 36:14:02as the name suggests. So instead of
- 36:14:04taking the square root, we leave the
- 36:14:08square as it is. And then we have
- 36:14:10Manhattan distance measure. In case of
- 36:14:12Manhattan distance, it is the sum of the
- 36:14:16distances across the x-axis and the
- 36:14:19y-axis. And note that we are taking the
- 36:14:22absolute value so that the negative
- 36:14:24values don't come into play. So that is
- 36:14:25the Manhattan distance measure. Then we
- 36:14:27have cosine distance measure. In this
- 36:14:30case, we take the angle between the two
- 36:14:33vectors formed by joining the points
- 36:14:35from the origin. So that is the cosine
- 36:14:38distance measure. Okay. So that was a
- 36:14:40quick overview about the various
- 36:14:42distance measures that are supported by
- 36:14:45K means. Now let's go and check how
- 36:14:48exactly K means clustering works. Okay.
- 36:14:51So this is how K means clustering works.
- 36:14:55This is like a flowchart of the whole
- 36:14:57process. There is a starting point and
- 36:15:00then we specify the number of clusters
- 36:15:03that we want. Now there are couple of
- 36:15:05ways of doing this. We can do by trial
- 36:15:08and error. So we specify a certain
- 36:15:11number maybe k is equal to 3 or four or
- 36:15:13five to start with and then as we
- 36:15:15progress we keep changing until we get
- 36:15:18the best clusters or there is a
- 36:15:20technique called elbow technique whereby
- 36:15:23we can determine the value of k. What
- 36:15:26should be the best value of k? How many
- 36:15:28clusters should be formed? So once we
- 36:15:30have the value of K we specify that and
- 36:15:33then the system will assign that many
- 36:15:37centrids. So it picks randomly that to
- 36:15:39start with randomly that many points
- 36:15:42that are considered to be the centrids
- 36:15:44of these clusters and then it measures
- 36:15:47the distance of each of the data points
- 36:15:50from these centroidids and assigns those
- 36:15:54points to the corresponding centr from
- 36:15:57which the distance is minimum. So each
- 36:15:59data point will be assigned to the
- 36:16:02centrid which is closest to it and
- 36:16:05thereby we have k number of initial
- 36:16:09clusters. However this is not the final
- 36:16:12clusters. The next step it does is for
- 36:16:16the new groups for the clusters that
- 36:16:19have been formed it calculates the mean
- 36:16:21position thereby calculates the new
- 36:16:25centroid position. the position of the
- 36:16:27centrid moves compared to the randomly
- 36:16:30allocated one. So it's an iterative
- 36:16:31process. Once again the distance of each
- 36:16:34point is measured from this new centroid
- 36:16:36point and if required the data points
- 36:16:40are reallocated to the new centroidids
- 36:16:43and the mean position or the new centrid
- 36:16:46is calculated once again. If the centrid
- 36:16:48moves then the iteration continues which
- 36:16:51means the convergence has not happened.
- 36:16:53The clustering has not converged. So as
- 36:16:56long as there is a movement of the
- 36:16:58centrid this iteration keeps happening.
- 36:17:01But once the centrid stops moving which
- 36:17:04means that the cluster has converged or
- 36:17:07the clustering process has converged
- 36:17:10that will be the end result. So now we
- 36:17:12have the final position of the centroid
- 36:17:15and the data points are allocated
- 36:17:18accordingly to the closest centrid. I
- 36:17:21know it's a little difficult to
- 36:17:22understand from this simple flowchart.
- 36:17:25So let's do a little bit of
- 36:17:27visualization and see if we can explain
- 36:17:30it better. Let's take an example. If we
- 36:17:32have a data set for a grocery shop. So
- 36:17:35let's say we have a data set for a
- 36:17:37grocery shop and now we want to find out
- 36:17:41how many clusters this has to be spread
- 36:17:44across. So how do we find the optimum
- 36:17:47number of clusters? There is a technique
- 36:17:49called the elbow method. So when these
- 36:17:52clusters are formed, there is a
- 36:17:54parameter called within sum of squares.
- 36:17:57And the lower this value is, the better
- 36:18:01the cluster is. That means all these
- 36:18:04points are very close to each other. So
- 36:18:07we use this within sum of squares as a
- 36:18:10measure to find the optimum number of
- 36:18:15clusters that can be formed for a given
- 36:18:17data set. So we create clusters or we
- 36:18:20let the system create clusters of a
- 36:18:23variety of numbers maybe of 10 10
- 36:18:26clusters and for each value of K the
- 36:18:30within SS is measured and the value of K
- 36:18:34which has the least amount of within SS
- 36:18:37or WSS that is taken as the optimum
- 36:18:41value of K. So this is the diagrammatic
- 36:18:44representation. So we have on the y-axis
- 36:18:47the within sum of squares or wss and on
- 36:18:51the x-axis we have the number of
- 36:18:53clusters. So as you can imagine if you
- 36:18:56have k is equal to one which means all
- 36:18:58the data points are in a single cluster
- 36:19:00the within ss value will be very high
- 36:19:03because they are probably scattered all
- 36:19:05over. The moment you split it into two
- 36:19:08there will be a drastic fall in the
- 36:19:10within ss value and that's what is
- 36:19:13represented here. But then as the value
- 36:19:15of K increases the decrease the rate of
- 36:19:19decrease will not be so high. It will
- 36:19:22continue to decrease but probably the
- 36:19:24rate of decrease will not be high. So
- 36:19:26that gives us an idea. So from here we
- 36:19:29get an idea for example the optimum
- 36:19:32value of K should be either two or three
- 36:19:35or at the most four but beyond that
- 36:19:38increasing the number of clusters is not
- 36:19:41dramatically changing the value in WSS
- 36:19:45because that pretty much gets
- 36:19:46stabilized. Okay. Now that we have got
- 36:19:49the value of K and let's assume that
- 36:19:52these are our delivery points. The next
- 36:19:54step is basically to assign two centrids
- 36:19:59randomly. So let's say C1 and C2 are the
- 36:20:02centrids assigned randomly. Now the
- 36:20:04distance of each location from the
- 36:20:07centrid is measured and each point is
- 36:20:10assigned to the centrid which is closest
- 36:20:14to it. So for example these points are
- 36:20:17very obvious that these are closest to
- 36:20:19C1 whereas this point is far away from
- 36:20:22C2. So these points will be assigned
- 36:20:26which are close to C1 will be assigned
- 36:20:27to C1 and these points or locations
- 36:20:30which are close to C2 will be assigned
- 36:20:32to C2. And then so this is the how the
- 36:20:35initial grouping is done. This is part
- 36:20:38of C1 and this is part of C2. Then the
- 36:20:41next step is to calculate the actual
- 36:20:44centrid of this data because remember C1
- 36:20:47and C2 are not the centrids. They've
- 36:20:49been randomly assigned points and only
- 36:20:53thing that has been done was the data
- 36:20:55points which are closest to them have
- 36:20:57been assigned to them. But now in this
- 36:21:00step the actual centroid will be
- 36:21:02calculated which may be for each of
- 36:21:04these data sets somewhere in the middle.
- 36:21:06So that's like the main point that will
- 36:21:08be calculated and the centr will
- 36:21:10actually be positioned or repositioned
- 36:21:13there. Same with C2. So the new centroid
- 36:21:17for this group is C2. this new position
- 36:21:19and C1 is in this new position. Once
- 36:21:22again, the distance of each of the data
- 36:21:24points is calculated from these
- 36:21:26centroids. Now remember, it's not
- 36:21:28necessary that the distance still
- 36:21:30remains the or each of these data points
- 36:21:32still remain in the same group. By
- 36:21:34recalculating the distance, it may be
- 36:21:36possible that some points get
- 36:21:38reallocated like so. You see this? So
- 36:21:41this point earlier was closer to C2
- 36:21:45because C2 was here. But after
- 36:21:48recalculating repositioning it is
- 36:21:51observed that this is closer to C1 than
- 36:21:55C2. So this is the new grouping. So some
- 36:21:58points will be reassigned. And again the
- 36:22:01centrid will be calculated and if the
- 36:22:04centroid doesn't change so that is a
- 36:22:06repetative process, iterative process.
- 36:22:08And if the centroid doesn't change once
- 36:22:10the centroid stops changing that means
- 36:22:13the algorithm has converged and this is
- 36:22:16our final cluster with this as the
- 36:22:18centroid C1 and C2 as the centroids
- 36:22:21these data points as a part of each
- 36:22:23cluster. So I hope this helps in
- 36:22:25understanding the whole process
- 36:22:27iterative process of K means clustering.
- 36:22:30So let's take a look at the K means
- 36:22:33clustering algorithm. Let's say we have
- 36:22:36x1, x2, x3, n number of points as our
- 36:22:40inputs and we want to split this into k
- 36:22:44clusters or we want to create k
- 36:22:46clusters. So the first step is to
- 36:22:48randomly pick k points and call them
- 36:22:52centroidids. They are not real centrids
- 36:22:55because centr is supposed to be a center
- 36:22:57point but they are just called centrids.
- 36:23:00And we calculate the distance of each
- 36:23:03and every input point from each of the
- 36:23:07centroidids. So the distance of X1 from
- 36:23:12C1 from C2 C3 each of the distances we
- 36:23:16calculate and then find out which
- 36:23:19distance is the lowest and assign X1 to
- 36:23:22that particular random centroid. Repeat
- 36:23:24that process for X2. calculate its
- 36:23:28distance from each of the centroid C1,
- 36:23:31C2, C3 up to CK and find which is the
- 36:23:34lowest distance and assign X2 to that
- 36:23:36particular centroid. Same with X3 and so
- 36:23:38on. So that is the first round of
- 36:23:41assignment that is done. Now we have K
- 36:23:45groups because there are we have
- 36:23:47assigned the value of K. So there are K
- 36:23:50centroids and uh so there are K groups.
- 36:23:54All these inputs have been split into K
- 36:23:56groups. However, remember we picked the
- 36:23:59centrids randomly. So they are not real
- 36:24:02centrids. So now what we have to do, we
- 36:24:05have to calculate the actual centroids
- 36:24:08for each of these groups which is like
- 36:24:10the mean position which means that the
- 36:24:13position of the randomly selected
- 36:24:15centrids will now change and they will
- 36:24:18be the main positions of these newly
- 36:24:21formed K groups. And once that is done,
- 36:24:25we once again repeat this process of
- 36:24:28calculating the distance. Right? So this
- 36:24:31is what we are doing as a part of step
- 36:24:33four. We repeat step two and three. So
- 36:24:36we again calculate the distance of X1
- 36:24:40from the centroid C1, C2, C3 and then
- 36:24:44see which is the lowest value and assign
- 36:24:47X1 to that. Calculate the distance of X2
- 36:24:50from C1, C2, C3 or whatever up to CK and
- 36:24:54find whichever is the lowest distance
- 36:24:56and assign X2 to that centroid and so
- 36:24:59on. In this process there may be some
- 36:25:01reassignment. X1 was probably assigned
- 36:25:04to cluster C2 and after doing this
- 36:25:07calculation maybe now X1 is assigned to
- 36:25:09C1. So that kind of reallocation may
- 36:25:12happen. So we repeat the steps two and
- 36:25:14three till the position of the centrids
- 36:25:18don't change or stop changing and that's
- 36:25:21when we have convergence. So let's take
- 36:25:24a detailed look at at each of these
- 36:25:26steps. So we randomly pick K cluster
- 36:25:29centers. We call them centroidids
- 36:25:31because they are not initially they are
- 36:25:34not really the centrids. So we let us
- 36:25:36name them C1 C2 up to CK. And then step
- 36:25:39two, we assign each data point to the
- 36:25:42closest center. So what we do, we
- 36:25:45calculate the distance of each X value
- 36:25:48from each C value. So the distance
- 36:25:51between X1 C1 distance between X1 C2 X1
- 36:25:57C3 and then we find which is the lowest
- 36:26:00value. Right? That's the minimum value
- 36:26:02we find and assign X1 to that particular
- 36:26:06centroid. Then we go next to x2. Find
- 36:26:09the distance of x2 from c1, x2 from c2,
- 36:26:13x2 from c3 and so on up to ck. And then
- 36:26:16assign it to the point or to the
- 36:26:19centroid which has the lowest value and
- 36:26:21so on. So that is step number two. In
- 36:26:24step number three, we now find the
- 36:26:27actual centr for each group. So what has
- 36:26:30happened as a part of step number two?
- 36:26:33We now have all the points, all the data
- 36:26:36points grouped into K groups because we
- 36:26:40we wanted to create K clusters, right?
- 36:26:42So we have K groups. Each one may be
- 36:26:45having a certain number of input values.
- 36:26:47They need not be equally distributed. By
- 36:26:50the way, based on the distance, we will
- 36:26:52have K groups. But remember the initial
- 36:26:56values of the C1 C2 were not really the
- 36:26:59centrids of these groups, right? we
- 36:27:01assign them randomly. So now in step
- 36:27:04three, we actually calculate the centr
- 36:27:08of each group which means the original
- 36:27:11point which we thought was the centrid
- 36:27:13will shift to the new position which is
- 36:27:16the actual centrid for each of these
- 36:27:18groups. Okay? And we again calculate the
- 36:27:22distance. So we go back to step two
- 36:27:25which is what we calculate again the
- 36:27:27distance of each of these points from
- 36:27:30the newly positioned centroidids and if
- 36:27:34required we reassign these points to the
- 36:27:38new centroidids. So as I said earlier
- 36:27:41there may be a reallocation. So we now
- 36:27:43have a new set or a new group. We still
- 36:27:47have K groups but the number of items
- 36:27:50and the actual assignment may be
- 36:27:52different from what was in step two
- 36:27:56here. Okay, so that might change. Then
- 36:27:59we perform step three once again to find
- 36:28:02the new centroid of this new group. So
- 36:28:05we have again a new set of clusters, new
- 36:28:09centroidids and new assignments. We
- 36:28:12repeat this step two again. Once again
- 36:28:14we find and then it is possible that
- 36:28:17after iterating through three or four or
- 36:28:20five times the centrid will stop moving
- 36:28:24in the sense that when you calculate the
- 36:28:26new value of the centrid that will be
- 36:28:29same as the original value or there will
- 36:28:31be very marginal change. So that is when
- 36:28:34we say convergence has occurred and that
- 36:28:37is our final cluster. That's the
- 36:28:41formation of the final cluster. All
- 36:28:43right. So let's see a couple of demos of
- 36:28:47uh K means clustering. We will actually
- 36:28:50see some live demos in uh Python
- 36:28:52notebook using Python notebook. But
- 36:28:54before that let's find out what's the
- 36:28:57problem that we are trying to solve. The
- 36:28:59problem statement is let's say Walmart
- 36:29:01wants to open a chain of stores across
- 36:29:04the state of Florida and uh it wants to
- 36:29:08find the optimal store locations. Now
- 36:29:11the issue here is if they open too many
- 36:29:14stores close to each other obviously the
- 36:29:16they will not make profit but if they if
- 36:29:20the stores are too far apart then they
- 36:29:22will not have enough sales. So how do
- 36:29:24they optimize this? Now for an
- 36:29:27organization like Walmart which is an
- 36:29:30e-commerce giant they already have the
- 36:29:33addresses of their customers in their
- 36:29:36database. So they can actually use this
- 36:29:39information or this data and use K means
- 36:29:42clustering to find the optimal location.
- 36:29:46Now before we go into the Python
- 36:29:49notebook and show you the live code, I
- 36:29:52wanted to take you through very quickly
- 36:29:54a summary of the code in the slides and
- 36:29:57then we will go into the Python
- 36:29:59notebook. So in this block we are
- 36:30:02basically importing all the required
- 36:30:05libraries like numpy, mattplot lib and
- 36:30:09so on and we are loading the data that
- 36:30:13is available in the form of let's say
- 36:30:15the addresses for simplicity sake we
- 36:30:18will just take them as some data points.
- 36:30:21Then the next thing we do is quickly do
- 36:30:23a scatter plot to see how they are
- 36:30:27related to each other with respect to
- 36:30:29each other. So in the scatter plot we
- 36:30:31see that there are a few distinct groups
- 36:30:35already being formed. So you can
- 36:30:37actually get an idea about how the
- 36:30:39cluster would look and how many clusters
- 36:30:41what is the optimal number of clusters
- 36:30:44and then starts the actual K means
- 36:30:46clustering process. So we will assign
- 36:30:50each of these points to the centrids and
- 36:30:53then check whether they are the optimal
- 36:30:56distance which is the shortest distance
- 36:30:58and assign each of the points data
- 36:31:01points to the centroidids and then go
- 36:31:05through this iterative process till the
- 36:31:07whole process converges and finally we
- 36:31:10get an output like this. So we have four
- 36:31:13distinct clusters and um which is we can
- 36:31:17say that this is how the population is
- 36:31:20probably distributed across Florida
- 36:31:22state and uh these centroidids are like
- 36:31:26the location where the store should be
- 36:31:29the optimum location where the store
- 36:31:31should be. So that's the way we
- 36:31:34determine the best locations for the
- 36:31:37store and that's how we can help Walmart
- 36:31:40find the best locations for their stores
- 36:31:42in Florida. So now let's take this into
- 36:31:46Python notebook. Let's see how this
- 36:31:48looks when we are learning running the
- 36:31:50code live. All right. So this is the
- 36:31:52code for K means clustering in Jupyter
- 36:31:56notebook. We have a few examples here
- 36:31:59which we will demonstrate how K means
- 36:32:02clustering is used and even there is a
- 36:32:04small implementation of K means
- 36:32:06clustering as well. Okay. So let's get
- 36:32:08started. Okay. So this block is
- 36:32:11basically importing the various
- 36:32:13libraries that are required like
- 36:32:14mattplot lib and numpy and so on and so
- 36:32:17forth which would be used as a part of
- 36:32:20the code. Then we are going and creating
- 36:32:23blobs which are similar to clusters. Now
- 36:32:26this is a very neat feature which is
- 36:32:28available in scikitlearn. Make blobs is
- 36:32:31a nice feature which creates clusters of
- 36:32:34data sets. So that's a wonderful
- 36:32:37functionality that is readily available
- 36:32:38for us to create some test data kind of
- 36:32:41thing. Okay. So that's exactly what we
- 36:32:44are doing here. We are using make blobs
- 36:32:46and we can specify how many clusters we
- 36:32:50want. So centers we are mentioning here.
- 36:32:53So it will go ahead and so we just
- 36:32:54mentioned four. So it will go ahead and
- 36:32:56create some test data for us. And this
- 36:33:00is how it looks. As you can see visually
- 36:33:02also we can figure out that there are
- 36:33:05four distinct classes or clusters in
- 36:33:08this data set. And that is what make
- 36:33:12blobs actually provides. Now from here
- 36:33:16onwards we will basically run the
- 36:33:18standard K means functionality that is
- 36:33:21readily available. So we really don't
- 36:33:23have to implement K means itself. The C
- 36:33:26means functionality or the the function
- 36:33:28is readily available. You just need to
- 36:33:31feed the data and we'll create the
- 36:33:33clusters. So this is the code for that.
- 36:33:36We import k means and then we create an
- 36:33:39instance of k means and we specify the
- 36:33:42value of k. This n_clusters is the value
- 36:33:46of k. Remember K means in K means K is
- 36:33:49basically the number of clusters that
- 36:33:51you want to create and it is a integer
- 36:33:54value. So this is where we are
- 36:33:55specifying that. So we have K is equal
- 36:33:57to four and so that instance is created.
- 36:34:01We take that instance and as with any
- 36:34:04other machine learning functionality fit
- 36:34:06is what we use the function or the
- 36:34:09method rather fit is what we use to
- 36:34:12train the model. Here there is no real
- 36:34:14training uh kind of thing but that's the
- 36:34:17call. Okay. So we are calling fit and
- 36:34:19what we are doing here we are just
- 36:34:21passing the data. So x has these values
- 36:34:23the data that has been created right. So
- 36:34:26that is what we are passing here and uh
- 36:34:30this will go ahead and create the
- 36:34:33clusters and uh then we are using
- 36:34:38after doing uh fit we run the predict
- 36:34:42which basically assigns for each of
- 36:34:44these observations which cluster it
- 36:34:47belongs to. All right. So it will name
- 36:34:50the clusters. Maybe this is cluster one.
- 36:34:52This is two, three and so on. Or will
- 36:34:55actually start from zero, cluster 0, 1,
- 36:34:582 and 3 maybe. And then for each of the
- 36:35:02observations it will assign based on
- 36:35:04which cluster it belongs to it will
- 36:35:06assign a value. So that is stored in y_k
- 36:35:10means when we call predict that is what
- 36:35:12it does. And we can take a quick look at
- 36:35:15these uh y_k means or the cluster
- 36:35:19numbers that have been assigned for each
- 36:35:21observation. So this is the cluster
- 36:35:23number assigned for observation one.
- 36:35:25Maybe this is for observation two,
- 36:35:27observation three and so on. So we have
- 36:35:30how many about I think 300 samples
- 36:35:32right? So all the 300 samples there are
- 36:35:34300 values here. Each of them the
- 36:35:37cluster number is given and the cluster
- 36:35:39number goes from 0 to three. So there
- 36:35:42are four clusters. So the numbers go
- 36:35:44from 0 1 2 3. So that's what is seen
- 36:35:48here. Okay. Now, so this was a quick
- 36:35:50example of generating some dummy data
- 36:35:53and then clustering that. Okay. And this
- 36:35:56can be applied if you have proper data.
- 36:35:57You can just load it up into X for
- 36:36:00example here and then run the K. So this
- 36:36:03is the central part of the K means
- 36:36:05clustering program example. So you
- 36:36:07basically create an instance and you
- 36:36:10mention how many clusters you want by
- 36:36:12specifying this parameter n_clusters and
- 36:36:15that is also the value of k and then
- 36:36:17pass the data to get the values. Now the
- 36:36:20next section of this code is the
- 36:36:23implementation of a k means. Now this is
- 36:36:26kind of a a rough implementation of the
- 36:36:29k means algorithm. So we will just walk
- 36:36:32you through I will walk you through the
- 36:36:33code uh at each step what it is doing
- 36:36:36and then we will see a couple of more
- 36:36:38examples of how K means clustering can
- 36:36:41be used in maybe some real life examples
- 36:36:44real life use cases. All right. So in
- 36:36:46this case here what we're doing is
- 36:36:48basically implementing K means
- 36:36:51clustering and there is a function or a
- 36:36:55library calculates for a given two pairs
- 36:36:58of points it will calculate the the
- 36:37:01distance between them and see which one
- 36:37:03is the closest and so on. So this is
- 36:37:05like this is pretty much like what K
- 36:37:07means does right. So it calculates the
- 36:37:09distance of each point or each data set
- 36:37:12from predefined centroid and then based
- 36:37:15on whichever is the lowest this
- 36:37:17particular data point is assigned to
- 36:37:19that centroid. So that is basically
- 36:37:21available as a standard function and we
- 36:37:23will be using that here. So as explained
- 36:37:26in the slides the first step that is
- 36:37:28done in case of C means clustering is to
- 36:37:32randomly assign some centrides. So as a
- 36:37:37first step we randomly allocate a couple
- 36:37:40of centrids which we call here we're
- 36:37:42calling as centers
- 36:37:44and then we put this in a loop and we
- 36:37:47take it through an iterative process.
- 36:37:50For each of the data points, we first
- 36:37:53find out using this function pair-wise
- 36:37:55distance argument. For each of the
- 36:37:57points, we find out which one which
- 36:38:00center or which uh randomly selected
- 36:38:02centrid is the closest and accordingly
- 36:38:06we assign that data or the data point to
- 36:38:10that particular centrid or cluster. And
- 36:38:13once that is done for all the data
- 36:38:16points, we calculate the new centr by
- 36:38:20finding out the mean position with the
- 36:38:22the center position. Right? So we
- 36:38:24calculate the new centroid and then we
- 36:38:26check if the new centroid is the
- 36:38:29coordinates or the position is the same
- 36:38:32as the previous centroid. The positions
- 36:38:35we will compare and if it is the same
- 36:38:39that means the process has converged. So
- 36:38:42remember we do this process till the
- 36:38:44centroidids or the centrid doesn't move
- 36:38:46anymore right so the centroid gets
- 36:38:48relocated each time this reallocation is
- 36:38:52done so the moment it doesn't change
- 36:38:55anymore the position of the cent doesn't
- 36:38:57change anymore we know that convergence
- 36:38:59has occurred so till then so you see
- 36:39:01here this is like an infinite loop while
- 36:39:03true is an infinite loop it only breaks
- 36:39:06when the centers are the same the new
- 36:39:09center and old center positions are the
- 36:39:11name and once that is uh done uh we
- 36:39:14return the centers and the labels. Now
- 36:39:16of course as explained this is not a
- 36:39:18very sophisticated and advanced
- 36:39:20implementation very basic implementation
- 36:39:22because one of the flaws in this is that
- 36:39:24sometimes what happens is the centroid
- 36:39:27the position will keep moving but in the
- 36:39:31change will be very minor. So in that
- 36:39:33case also that is actually convergence
- 36:39:36right. So for example the change is 0.1
- 36:39:39we can consider that as convergence
- 36:39:41otherwise what will happen is this will
- 36:39:43either take forever or it will be never
- 36:39:46ending. So that's a small flaw here. So
- 36:39:48that is something additional checks may
- 36:39:50have to be added here. But again as
- 36:39:52mentioned this is not the most
- 36:39:54sophisticated uh implementation. This is
- 36:39:57like a kind of a rough implementation of
- 36:39:59the k means clustering. Okay. So if we
- 36:40:02execute this code this is what we get as
- 36:40:05the output. So this is the definition of
- 36:40:07this particular function and then we
- 36:40:10call that find_clusters and we pass our
- 36:40:13data x and the number of clusters which
- 36:40:15is four and if we run that and plot it
- 36:40:19this is the output that we get. So this
- 36:40:21is of course each cluster is represented
- 36:40:23by a different color. So we have a
- 36:40:25cluster in green color, yellow color and
- 36:40:27so on and so forth. And these big points
- 36:40:30here these are the centroidids is the
- 36:40:32final position of the centroidids. And
- 36:40:33as you can see visually also this
- 36:40:35appears like a kind of a center of all
- 36:40:38these points here. Right? Similarly this
- 36:40:41is like the center of all these points
- 36:40:43here and so on. So this is the example
- 36:40:46or this is an example of a
- 36:40:47implementation of K means clustering and
- 36:40:52uh next we will move on to see a couple
- 36:40:54of examples of how K means clustering is
- 36:40:58used in maybe some real life scenarios
- 36:41:01or use cases. In the next example or
- 36:41:04demo, we are going to see how we can use
- 36:41:07K means clustering to perform color
- 36:41:10compression. We will take a couple of
- 36:41:12images. So there will be two examples
- 36:41:15and uh we will try to use C means
- 36:41:18clustering to compress the colors. This
- 36:41:21is a common situation in image
- 36:41:23processing when you have an image with
- 36:41:26millions of uh colors but then you
- 36:41:29cannot render it on some devices which
- 36:41:32may not have enough memory. Uh so that
- 36:41:34is the scenario where where something
- 36:41:36like this can be used. So before again
- 36:41:40we go into the Python notebook let's
- 36:41:43take a look at quickly the the code. As
- 36:41:46usual we import the libraries and then
- 36:41:49we import the image and uh then we will
- 36:41:53flatten it. So the reshaping is
- 36:41:55basically we have the image information
- 36:41:58is stored in the form of pixels and uh
- 36:42:02if the image is like for example 427x
- 36:42:05640 and it has three colors. So that's
- 36:42:08the overall dimension of the of the
- 36:42:12initial image. we just reshape it and um
- 36:42:16then feed this to our algorithm and this
- 36:42:20will then create clusters of only 16
- 36:42:23clusters. So this this colors there are
- 36:42:26millions of colors and now we need to
- 36:42:29bring it down to 16 colors. So we use k
- 36:42:32is equal to 16 and u this is how when we
- 36:42:36visualize this is how it looks. There
- 36:42:37are these are all about 16 million
- 36:42:40possible colors. The input color space
- 36:42:42has 16 million possible colors and we
- 36:42:46just sub compress it to 16 colors. So
- 36:42:49this is how it would look when we
- 36:42:51compress it to 16 colors. And this is
- 36:42:54how the original image looks. And after
- 36:42:57compression to 16 colors, this is how
- 36:43:00the new image looks. As you can see,
- 36:43:02there is not a lot of information that
- 36:43:06has been lost. though the image quality
- 36:43:10is definitely reduced a little bit. So
- 36:43:14this is an example which we are going to
- 36:43:16now see in Python notebook. Let's go
- 36:43:19into the Python and once again as always
- 36:43:22we will import some libraries and load
- 36:43:25this image called flower.jpg.
- 36:43:29Okay. So let we'll load that and this is
- 36:43:31how it looks. This is the original image
- 36:43:34which has I think 16 million colors and
- 36:43:38uh this is the shape of this image which
- 36:43:40is basically what is the shape is
- 36:43:42nothing but the overall size right so
- 36:43:45this is 427 pixel by 640 pixel and then
- 36:43:49there are three layers which is this
- 36:43:51three basically is for RGB which is red
- 36:43:53green blue so color image will have that
- 36:43:56right so that is the shape of this now
- 36:43:58what we need to do is data let's take a
- 36:44:01look at how data is looking. So let me
- 36:44:03just create a new cell and show you what
- 36:44:07is in data. Basically we have captured
- 36:44:09this information.
- 36:44:12So data is what? Let me just show you
- 36:44:14here.
- 36:44:18All right. So let's take a look at
- 36:44:21China. What are the values in China? And
- 36:44:25uh if you see here, this is how the data
- 36:44:27is stored. This is nothing but the pixel
- 36:44:29values. Okay? So this is like a matrix
- 36:44:32and each one has about for for this 427x
- 36:44:36640 pixels. All right. So this is how it
- 36:44:39looks. Now the issue here is these
- 36:44:40values are large. The numbers are large.
- 36:44:44So we need to normalize them to between
- 36:44:460 and one. Right? So that's why we will
- 36:44:49basically create one more variable which
- 36:44:52is data which will contain the values
- 36:44:54between 0 and one. And the way to do
- 36:44:56that is divide by 255. So we divide
- 36:44:59China by 255 and we get the new values
- 36:45:02in data. So let's just run this uh piece
- 36:45:05of code and this is the shape. So we now
- 36:45:08have also yeah what we have done is we
- 36:45:10changed using reshape we converted into
- 36:45:13the three-dimensional into a
- 36:45:15two-dimensional data set. And let us
- 36:45:18also take a look at how
- 36:45:21let me just insert
- 36:45:24probably a cell here and take a look at
- 36:45:27how data is looking. All right. So this
- 36:45:29is how data is looking and now you see
- 36:45:31this is the values are between 0 and
- 36:45:34one. Right? So if you earlier noticed in
- 36:45:36case of China the values were large
- 36:45:38numbers. Now everything is between 0 and
- 36:45:40one. This is one of the things we need
- 36:45:42to do. All right. So after that the next
- 36:45:45thing that we need to do is to visualize
- 36:45:47this and uh we can take random set of
- 36:45:51maybe 10,000 points and plot it and
- 36:45:54check and see how this looks. So let us
- 36:45:57just plot this and uh so this is how the
- 36:46:00original the color the pixel
- 36:46:02distribution is. These are two plots one
- 36:46:05is red against green and another is red
- 36:46:07against blue and this is the original
- 36:46:09distribution of the color. So then what
- 36:46:11we will do is we will use K means
- 36:46:13clustering to create just 16 clusters
- 36:46:17for the various colors and then apply
- 36:46:20that to the image. Now what will happen
- 36:46:23is since the data is large because there
- 36:46:25are millions of colors using regular K
- 36:46:27means may be a little time consuming. So
- 36:46:30there is another version of K means
- 36:46:32which is called mini batch K means. So
- 36:46:34we will use that which is which
- 36:46:36processes in the overall concept remains
- 36:46:39the same but this basically processes it
- 36:46:41in smaller batches. That's the only
- 36:46:43thing. Okay. So the results will pretty
- 36:46:45much be the same. So let's go ahead and
- 36:46:48execute this piece of code and also
- 36:46:51visualize this so that we can see that
- 36:46:53there are the this is how the 16 colors
- 36:46:56uh would look. So this is red against
- 36:46:58green and this is red against blue.
- 36:47:01there is uh quite a bit of similarity
- 36:47:03between this original color schema and
- 36:47:06the new one. Right? So it doesn't look
- 36:47:08very very completely different or
- 36:47:10anything like that. Now we apply this
- 36:47:12the newly created colors to the image
- 36:47:16and uh we can take a look uh how this is
- 36:47:18uh looking. Now we can compare both the
- 36:47:20images. So this is our original image
- 36:47:23and this is our new image. So as you can
- 36:47:25see there is not a lot of information
- 36:47:28that has been lost. uh it pretty much
- 36:47:31looks like the original image. Yes, we
- 36:47:34can see that for example here there is a
- 36:47:36little bit uh it appears a little
- 36:47:39dullish compared to this one right
- 36:47:41because uh we kind of took off some of
- 36:47:43the finer details of the color but
- 36:47:47overall the highle information has been
- 36:47:50maintained. At the same time, the main
- 36:47:52advantage is that now this can be this
- 36:47:54is an image which can be rendered on a
- 36:47:57device which may not be that very
- 36:47:59sophisticated. Now let's take one more
- 36:48:01example with a different image. In the
- 36:48:04second example, we will take an image of
- 36:48:06the summer palace in China and we repeat
- 36:48:10the same process. This is a high
- 36:48:12definition color image with millions of
- 36:48:15colors and also uh three-dimensional.
- 36:48:19Now we will reduce that to 16 colors
- 36:48:22using K means clustering. And um we do
- 36:48:26the same process like before. We reshape
- 36:48:28it and then we cluster the colors to 16
- 36:48:32and then we render the image once again.
- 36:48:36And we will see that the color the
- 36:48:38quality of the image is slightly
- 36:48:41deteriorates. As you can see here, this
- 36:48:43has much finer details in this which are
- 36:48:46probably missing here. But then that's
- 36:48:48the compromise because there are some
- 36:48:50devices which may not be able to handle
- 36:48:53this kind of a high density images. So
- 36:48:56let's run this code in Python notebook.
- 36:49:00All right. So let's apply the same
- 36:49:02technique for another picture which is
- 36:49:04uh even more intricate and has probably
- 36:49:08much complicated color schema. So this
- 36:49:11is the image. Now once again uh we can
- 36:49:14take a look at the shape which is 427x
- 36:49:17640x3
- 36:49:19and this is the new data would look
- 36:49:22somewhat like this compared to the
- 36:49:24flower image. So we have some new values
- 36:49:27here and we will also bring this as you
- 36:49:31can see the numbers are much big. So we
- 36:49:33will much bigger so we will now have to
- 36:49:36scale them down to values between 0 and
- 36:49:39one. And that is done by dividing by
- 36:49:41255. So let's go ahead and uh do that
- 36:49:46and reshape it. Okay. So we get a
- 36:49:49two-dimensional matrix and uh we will
- 36:49:53then as the next step we will go ahead
- 36:49:55and visualize this how it looks the the
- 36:49:5816 colors and this is basically how it
- 36:50:01would look 16 million colors. And now we
- 36:50:05can create the clusters out of this. The
- 36:50:0916 K means clusters we will create. So
- 36:50:13this is how the distribution of the
- 36:50:15pixels would look with 16 colors. And
- 36:50:19then we go ahead and uh apply this and
- 36:50:23visualize how it is looking for with the
- 36:50:26with the new just the 16 color. So once
- 36:50:29again, as you can see, this looks much
- 36:50:32richer in color, but at the same time,
- 36:50:35and this probably doesn't have, as we
- 36:50:38can see, it doesn't look as rich as this
- 36:50:40one, but nevertheless, the information
- 36:50:42is not lost, the shape and all that
- 36:50:44stuff. And this can be also rendered on
- 36:50:47a slightly a device which is probably
- 36:50:51not that sophisticated. Okay, so that's
- 36:50:54pretty much it. So we have seen two
- 36:50:56examples of how color compression can be
- 36:50:59done uh using K means clustering and we
- 36:51:03have also seen in the previous examples
- 36:51:05of how to implement C means the code to
- 36:51:08roughly how to implement C means
- 36:51:10clustering and we use some sample data
- 36:51:13using blob to just execute the C means
- 36:51:17cluster that takes place after data
- 36:51:20collection and before statistical
- 36:51:21analysis. So before you conduct any
- 36:51:24statistical formulas and analysis on the
- 36:51:26data and squeeze the data to extract
- 36:51:30some valuable insights, the process
- 36:51:32which you perform is called as initial
- 36:51:34data analysis. Like taking the data from
- 36:51:36the source, cleaning the data,
- 36:51:38transforming the data into a readable
- 36:51:40format and using that readable data to
- 36:51:44build some basic charts what exactly is
- 36:51:46happening with this particular company,
- 36:51:49brand or anything. Let's say I give you
- 36:51:51some data from the company. Then you get
- 36:51:53some insights of it. How many number of
- 36:51:55traffic you received? How many number of
- 36:51:57orders you received? What's the sale
- 36:51:59that you made in a specific month,
- 36:52:01specific quarter or specific year? And
- 36:52:04what was the profit? So basic
- 36:52:06information which you convert from the
- 36:52:08data and create a dashboard. Right? That
- 36:52:10is called as initial data analysis. So a
- 36:52:14step beyond initial data analysis is
- 36:52:16known as the exploratory data analysis.
- 36:52:19This is where you perform some
- 36:52:20statistics and probability and predict
- 36:52:23the future. Right? So let's dive deep
- 36:52:25and learn what exactly is exploratory
- 36:52:28data analysis. So a simple definition
- 36:52:30for exploratory data analysis is as
- 36:52:32follows. Exploratory data analysis is a
- 36:52:35key step in data analysis process that
- 36:52:38helps you identify patterns, outliners
- 36:52:40and relationships between variables
- 36:52:43before making assumptions. It is not
- 36:52:46like you just create a dashboard out of
- 36:52:47the initial data analysis and you can
- 36:52:49predict the future. No, it's not like
- 36:52:51that. You might have to last. You might
- 36:52:53have to go through some permutations and
- 36:52:55combinations. You might have to check
- 36:52:57the seasons. You might have to check the
- 36:52:59possibilities, right? During some
- 36:53:02particular seasons in the year, let's
- 36:53:04say it's Christmas, then you can expect
- 36:53:06some good sales. Let's say it's some
- 36:53:09festival, it's some special occasion,
- 36:53:11you can expect some good sales on the
- 36:53:13product, right? and maybe a part of the
- 36:53:16year, maybe a part of 10 years, a
- 36:53:19decade, right? In a certain period of
- 36:53:21time, there might be some reason due to
- 36:53:23which the sales of a certain product
- 36:53:26were high. So to make sure that your
- 36:53:29assumptions to make sure that your
- 36:53:31projection of the sales is 100% accurate
- 36:53:34or at least 90 to 95% accurate then you
- 36:53:37might have to go through the exploratory
- 36:53:39data analysis where you make use of
- 36:53:42statistics in your data analysis. Now
- 36:53:45there are certain steps that you might
- 36:53:46have to follow while going through
- 36:53:48exploratory data analysis. So following
- 36:53:51are the steps. So a first few steps
- 36:53:54might be slightly relevant to initial
- 36:53:57data analysis like connecting data,
- 36:53:59cleaning it, transforming it and loading
- 36:54:01it. After that you will import certain
- 36:54:04libraries from Python and after that you
- 36:54:07read the data what exactly you have in
- 36:54:10your data. The number of columns, the
- 36:54:12number of rows and if there are any null
- 36:54:14values, if there are any uh entries
- 36:54:16which are invalid, you might have to
- 36:54:18check that, read that and you might have
- 36:54:20to check for duplicate entries. It is
- 36:54:22possible that one entry might have been
- 36:54:25entered by two different people, right?
- 36:54:27There might be a duplication. So you
- 36:54:28might have to eliminate those
- 36:54:30duplications. You might have to check
- 36:54:32for some missing values. You might have
- 36:54:34to calculate the total number of missing
- 36:54:35values from the data set and try to
- 36:54:37eliminate them from the calculation
- 36:54:40during your exploratory data analysis.
- 36:54:42Followed by that you have to do some
- 36:54:43model engineering. Followed by that you
- 36:54:45might have to do some feature
- 36:54:46engineering creating features and then
- 36:54:49you will get started with exploratory
- 36:54:51data analysis and the at the end you
- 36:54:53will generate a projection or a
- 36:54:55prediction or give your assumption that
- 36:54:57this might happen in the future and you
- 36:54:59might have to take action to avoid it or
- 36:55:02you might have to take action to
- 36:55:04improvise it right so this is how the
- 36:55:06steps in exploratory data analysis take
- 36:55:08part now let's proceed and start with
- 36:55:12our demo on Python's exploratory data
- 36:55:15analysis and in this session we will be
- 36:55:18using the use case that we discussed
- 36:55:19before which happens to be the students
- 36:55:22performance data set. So in this
- 36:55:23particular data set we will be having
- 36:55:25some columns based on physical activity
- 36:55:28the distance from home parental
- 36:55:30education the subjects the marks they
- 36:55:32have scored in the previous exam. The
- 36:55:34marks that they have scored in the
- 36:55:36previous exam and if they have any
- 36:55:38disabilities if they are having any
- 36:55:40resources that they require to write the
- 36:55:42exams right. So a few parameters the
- 36:55:45important parameters that we will be
- 36:55:46discussing in this session and followed
- 36:55:48by that we will project the future that
- 36:55:52how they will you know improve in their
- 36:55:54exams and if there is a problem and if
- 36:55:58there is a solution to it then we can
- 36:56:00implement that solution and help
- 36:56:01students to gain better marks in their
- 36:56:03exam. So that's the overall use case for
- 36:56:05this demonstration. Now let's get
- 36:56:07started with our Jupyter notebook. Now
- 36:56:10we are on Jupiter notebook. Now let's
- 36:56:13get started. So I would like to have a
- 36:56:15title for my um notebook. So I'll write
- 36:56:18an HTML code for that.
- 36:56:22So HTML code uh
- 36:56:26and I have uh three apostrophes
- 36:56:32and here I would like to write something
- 36:56:34in H1. So I want my title to be in H1.
- 36:56:37So
- 36:56:39style will be
- 36:56:43background color
- 36:56:51name will be dark blue.
- 36:56:56So we let's uh proceed with the simply
- 36:56:58dance background which will be t usually
- 36:57:02and the color
- 36:57:04of text will be orange
- 36:57:09and I also want to have the font let's
- 36:57:13let me give the font size as 30.
- 36:57:24And in the next line, I'd like to have
- 36:57:26border
- 36:57:30radius. I can give the border radius as
- 36:57:3320 pixels
- 36:57:37and padding
- 36:57:39to be 16 pixels.
- 36:57:48Text alignment, I'd like to keep it
- 36:57:50center.
- 36:57:55There you go. Let's code the Let's close
- 36:57:58the H1. And now let's code the uh
- 36:58:02background color or border color. So
- 36:58:06B style.
- 36:58:09So what we can do is we can basically
- 36:58:12have this uh code here and what we will
- 36:58:14do is we will reuse this particular code
- 36:58:17segment because we will be having
- 36:58:19multiple HTML codes in this particular
- 36:58:23uh workbook which will explain the
- 36:58:26results of the analysis that we are
- 36:58:29doing. So the way we just run the code
- 36:58:31and after that we will be getting some
- 36:58:33visualizations and I will be writing
- 36:58:36some textual content in XT in an HTML
- 36:58:39page so that it will be easier for the
- 36:58:42people to understand what's what exactly
- 36:58:44is happening here right so the color
- 36:58:47will be light blue and now comes the
- 36:58:50text we will close this and here we will
- 36:58:53write the text as
- 36:58:56Python
- 36:58:58explorate
- 36:59:00data analysis
- 36:59:04and we will
- 36:59:07break here
- 36:59:10the student performance
- 36:59:19and here we will close the H1
- 36:59:24and lastly we will display this in HTML
- 36:59:27code. So basically we missed this
- 36:59:29library. So we will be importing from
- 36:59:32ipython display import html to display
- 36:59:34this particular code. So without this
- 36:59:36particular library we cannot display any
- 36:59:39HTML codes in our notebook. So we will
- 36:59:41quickly run that and we have a title
- 36:59:43over here. Now let's proceed with the
- 36:59:46next part. Now we will uh use some
- 36:59:49libraries like CAD boost and light bgm.
- 36:59:53So for that we might have to install the
- 36:59:55these libraries. So we will be using pip
- 36:59:58install here
- 37:00:00pip install cat boost
- 37:00:04and control enter to run this particular
- 37:00:06code segment or you can also use run and
- 37:00:09it's installed and after that we will
- 37:00:11also install
- 37:00:15light bgm
- 37:00:19g sorry not bgm
- 37:00:24so it's already installed
- 37:00:26Now we will start importing the
- 37:00:29libraries that we need. So we will be
- 37:00:32needing numpy, panda, seabon, mattplot
- 37:00:35lib and we will also import another
- 37:00:37special library which is for ignoring
- 37:00:40warnings. So we will import warnings and
- 37:00:43after that from we will be importing
- 37:00:45that warnings from IPython display
- 37:00:47import clear output and after that we
- 37:00:50will uh tell the jupyter notebook to
- 37:00:52import if there are any warnings. So
- 37:00:54that code will be warning dot filter
- 37:00:57warnings and in the uh brackets we will
- 37:01:00write ignore. So uh the basic libraries
- 37:01:04which we will be needing are as follows.
- 37:01:07import
- 37:01:09numpy
- 37:01:11as np. Let's quickly copy this and
- 37:01:16proceed. Enter. And now we will be
- 37:01:19needing pandas as pd.
- 37:01:23Enter. And now we will be needing
- 37:01:26seabbone
- 37:01:30SNS.
- 37:01:32And we will also use mattplot lib
- 37:01:36py plot
- 37:01:43as plt.
- 37:01:46And after that we will import warnings
- 37:01:51from
- 37:01:54I Python
- 37:01:56dot display
- 37:01:59import
- 37:02:00clear
- 37:02:02output
- 37:02:05warnings
- 37:02:08dot filter warnings
- 37:02:12ignore.
- 37:02:15Now we'll just quickly run this query.
- 37:02:18Run. And now we will proceed with
- 37:02:21feature engineering.
- 37:02:23So you can use a hashtag to ignore that
- 37:02:26particular line from execution for
- 37:02:29Jupyter notebook.
- 37:02:32And here we will be importing import
- 37:02:37from skarn
- 37:02:40dot impute import
- 37:02:45simple computer
- 37:02:50from skarn
- 37:02:52dot model
- 37:02:54selection
- 37:02:57import kf fold. There you go. Now let's
- 37:03:01quickly run this query.
- 37:03:04There you go. Now we will proceed with
- 37:03:06modeling and model evaluation. Once we
- 37:03:09are done with this then we will directly
- 37:03:12input the data into our notebook. So for
- 37:03:15modeling the data we will again use a
- 37:03:17hash code so that this particular line
- 37:03:20will not execute
- 37:03:24and import some libraries lit
- 37:03:28GBM
- 37:03:31as LGB
- 37:03:34from light GBM library.
- 37:03:41import
- 37:03:42LGBM regressor
- 37:03:47from CAD boost import
- 37:03:51boost regressor. There you go. Let's
- 37:03:54quickly run this command.
- 37:03:57There you go. Now, lastly, we have one
- 37:04:00more task before importing the data that
- 37:04:03is model evaluation. Once that is done,
- 37:04:06we can proceed with importing the data.
- 37:04:09There you go. Let's quickly run it. And
- 37:04:11now so far so good. We have uh done the
- 37:04:14basic library imports and feature
- 37:04:17engineering is done, modeling is done
- 37:04:19and model evaluation is also done. Now
- 37:04:21we can begin with importing the data. So
- 37:04:25we let's also add um the HTML code where
- 37:04:29we have done this importing. So here I
- 37:04:32will try to add another segment and I
- 37:04:35will import the HTML code here which
- 37:04:38displays a similar HTML format which
- 37:04:41explains what exactly is happening here.
- 37:04:43Just a moment I have the code ready.
- 37:04:46I'll just paste it here. There you go.
- 37:04:48Let's quickly run this so that we have a
- 37:04:51HTML code here page here which explains
- 37:04:54what exactly is happening here. So we're
- 37:04:55importing libraries and also student
- 37:04:58data. Now let's import the student data.
- 37:05:01For that let's write the query. So we
- 37:05:04are importing the data as data frame and
- 37:05:07after that we will write pandas read CSV
- 37:05:13right and here we will add the location
- 37:05:17of the file. Right? So the file is
- 37:05:20located in my downloads section. So
- 37:05:24let's quickly copy that location and
- 37:05:25paste it here. So this is the location
- 37:05:28of my file. Let's quickly run it. So
- 37:05:31there might be some error. It's okay. We
- 37:05:33can resolve it. So in such scenarios
- 37:05:37don't have to worry either you can add
- 37:05:38an R but even if that doesn't work you
- 37:05:41might have to change uh in this case it
- 37:05:44worked but in case if it doesn't work
- 37:05:45what you can do is uh you can eliminate
- 37:05:48uh the r and you can just change this
- 37:05:51from uh forward slash to backlash. This
- 37:05:55will also help. So this could be worth
- 37:05:58it. So this can also work. So these are
- 37:06:00the situations where you can use this.
- 37:06:02Now let's proceed with some more uh
- 37:06:06interesting facts. Let's try to
- 37:06:08understand what's going on with our uh
- 37:06:10data set. Right? So what you can do is
- 37:06:13now the data is stored in df variable as
- 37:06:17a data frame. So what you can do is read
- 37:06:19this particular data df.shape shape so
- 37:06:23that you can understand what's the uh
- 37:06:26what's happening with this data. Right?
- 37:06:28So it can tell you that there are uh 6
- 37:06:32sorry 6,67
- 37:06:34rows and 20 columns. Now let's add
- 37:06:36another query part and here you can uh
- 37:06:40try to see the head right what head
- 37:06:42means basically head means uh the column
- 37:06:45headers. So what you can do is just
- 37:06:47write df dot head
- 37:06:51and run. Now you have the column headers
- 37:06:54and couple of sample columns
- 37:06:57sorry couple of sample rows. Now let's
- 37:07:01proceed with uh checking the duplicates
- 37:07:04and identifying the total number of
- 37:07:06duplicate entries in this particular
- 37:07:07data set. So you can write down df dot
- 37:07:11duplicated
- 37:07:14dot
- 37:07:16sum and you will get the number of
- 37:07:18duplicates present in this particular
- 37:07:21file. So we have zero duplicate entries.
- 37:07:23Now let's see if there are any null uh
- 37:07:26elements in this particular data frame.
- 37:07:28So df dot is null
- 37:07:33dot sum.
- 37:07:35So these are all functions. Control
- 37:07:36enter. There you go. So in the column
- 37:07:40teacher quality there are 78 null
- 37:07:43entries and in parental educational
- 37:07:45level there are 90 null entries and
- 37:07:47distance from home there are u 67 null
- 37:07:51entries. Now what we can do is uh from
- 37:07:55the okay so from df shape we can add a
- 37:07:59new cell here and we can write a HTML
- 37:08:02code so that the viewer can understand
- 37:08:04that we are trying to understand our
- 37:08:05data. So let's write a quick code for
- 37:08:08that. So let's not waste much time in
- 37:08:10just writing the HTML code. So I've got
- 37:08:12that HTML code written in a notepad
- 37:08:14already. So I'll just paste it over here
- 37:08:16and let's quickly run it so that we have
- 37:08:18a HTML visibility here. So this was
- 37:08:22supposed to be the result. So we will
- 37:08:23add it here. So what we will do is
- 37:08:26quickly edit this particular content.
- 37:08:28We'll just quickly copy this code and
- 37:08:32paste it over here.
- 37:08:36Change the content from importing
- 37:08:40libraries to
- 37:08:43reading
- 37:08:45student data and we will cut this code
- 37:08:48from here and we will paste it in here
- 37:08:55so that it get give us some information.
- 37:08:58Let's also run this particular code
- 37:09:00segment
- 37:09:02so that we have trading student data.
- 37:09:04There you go. Now the next part of this
- 37:09:06session will be about creating a target
- 37:09:09variable. So overall target of this
- 37:09:11particular data analysis is about exam
- 37:09:14score. Right? So we can name our target
- 37:09:16variable as exam score. And let's
- 37:09:18understand the distribution of this
- 37:09:20particular exam score with uh the
- 37:09:25variables we have. Now let's write down
- 37:09:28plot dot figure.
- 37:09:33Figure size should be around 15 comma 9
- 37:09:41equals to let's add a bracket here 15
- 37:09:44comma 9 or let's keep it as six 9 would
- 37:09:48be a little bigger. Now enter now we
- 37:09:52will use seabbond library here. SNS dot
- 37:09:55count plot
- 37:09:58x is equals to data frame cleaned
- 37:10:01target. Okay. Uh before cleaned target
- 37:10:04we might have to run a few more. Okay.
- 37:10:07We did not perform data cleaning so far,
- 37:10:09right? So let's proceed with data
- 37:10:12cleaning so far. So we found some empty
- 37:10:14entries, right? Null entries and we also
- 37:10:17found some So here we have 299 rows
- 37:10:21which have missing values. So we will
- 37:10:23have to remove that. For that uh we
- 37:10:26might have to create a new column which
- 37:10:28has to be named as not um assigned right
- 37:10:31df data frame not assigned which is
- 37:10:34equals to df dot drop na. So we will be
- 37:10:39dropping the null values here. It's a
- 37:10:41function. And here let's print the
- 37:10:44values. Print df dot df
- 37:10:49na dot shape. So how many number of rows
- 37:10:53and columns we have right and after that
- 37:10:57let's also try to eliminate the null
- 37:10:59values as well dot is null so we don't
- 37:11:02have basically we don't have null values
- 37:11:04but we have u some illegal entries maybe
- 37:11:09some there you go now let's quickly run
- 37:11:11this query so there you go so the new
- 37:11:14data is about 678
- 37:11:1820
- 37:11:19so there is Um, okay. We did uh some
- 37:11:23mistake here. So, we supposed to add it
- 37:11:25as null. N U N L N N N N N N N N N N N N
- 37:11:27N N N N N N L N N N N N N N N N N N N N
- 37:11:27N N N N N N N N N N N N N N N N N N N N
- 37:11:27N N N N N N N U L. Now quickly run this.
- 37:11:30So we should not get any errors this
- 37:11:31time.
- 37:11:34There you go. No errors. So far so good.
- 37:11:37Now let's describe the new data set. df
- 37:11:40na dot describe. So these are the new
- 37:11:45columns and rows that we have. And we
- 37:11:48have uh mean, standard variation,
- 37:11:50standard deviation, minimum, maximum. So
- 37:11:53the scores are split into 25%, 50% and
- 37:11:5775%. Which could be based on hours
- 37:12:00studied which could be based on
- 37:12:01attendance, sleep hours, previous
- 37:12:03course, due training sessions,
- 37:12:04everything. So uh minimum sleep hours,
- 37:12:07maximum sleep hours, 25% of that, 50% of
- 37:12:10that, 75% of that. So that is supposed
- 37:12:12to be the u describe. Now what we will
- 37:12:16do is we have a target variable which is
- 37:12:18exam score. Right? Now we will do some
- 37:12:22changes to it. We already know in an
- 37:12:25exam there will be a threshold value. It
- 37:12:28can be 25 marks per exam. It can be 50
- 37:12:31marks per exam and it can be 100 marks
- 37:12:34per exam. Right? In our situation let's
- 37:12:36consider the threshold value is 100
- 37:12:38marks. Right? If there is a situation
- 37:12:41where marks is entered in a wrong way
- 37:12:45right it can if they add if they wanted
- 37:12:47to add 11 but by mistake if they added
- 37:12:49uh another one right triple one it's not
- 37:12:52a right entry right so we will try to
- 37:12:55eliminate those kind of uh data
- 37:12:59so we will create a new uh data frame
- 37:13:01here which is dataf frame cleaned is
- 37:13:04equals to dataf frame
- 37:13:08not null so We have eliminated the null
- 37:13:10values. BF NA. Now we will add our
- 37:13:15target variable which is exam score
- 37:13:19should be. So let's come out of this and
- 37:13:23here we will add it as should be less
- 37:13:26than or equal to 100 but not more than
- 37:13:28100. Let's use square brackets.
- 37:13:32Here we also the format is square
- 37:13:35brackets. ing action now enter and
- 37:13:41we will describe this particular
- 37:13:44data set instead of the FNA we will copy
- 37:13:47paste this here now let's run this
- 37:13:51okay exam score is not identified let's
- 37:13:55quickly check the error and resolve it
- 37:13:56yeah so we missed out to add colons here
- 37:14:00it's okay not a problem so this was
- 37:14:03supposed to be how it is now let's run
- 37:14:05this and we will have the answer over
- 37:14:07here. So we have the output. Now let's
- 37:14:10check the uh head of this particular
- 37:14:12clean data set. So we can make use of
- 37:14:14the same code here and paste it right
- 37:14:17here and instead of describe let's write
- 37:14:20head so that we have the header of uh
- 37:14:23this particular data set. So we have our
- 37:14:25study and everything normal and we will
- 37:14:28categorize the data right. So we will
- 37:14:31make use of three columns our study
- 37:14:33attendance and previous scores and uh
- 37:14:36after that we will also make use of
- 37:14:39other columns in this particular data
- 37:14:41set which happens to be the parental
- 37:14:43involvement access to resources sleep
- 37:14:45hours ting sessions etc. And now our
- 37:14:49target will be the exam score that we
- 37:14:51created over here. Right? This exam
- 37:14:53score will be our target. And using this
- 37:14:55particular exam score target, we will
- 37:14:58categorize the data. Okay? We will
- 37:15:00categorize the data in terms of uh let's
- 37:15:03say uh first class uh second class and
- 37:15:06uh pass or something like that. Right?
- 37:15:09So if if a student is uh scoring below
- 37:15:1264 and uh that is a separate category.
- 37:15:15If the score student is scoring equals
- 37:15:17to or above 65, that is a different
- 37:15:19category. And if the student is scoring
- 37:15:21beyond 70, that's a different category.
- 37:15:25And uh before we proceed with that,
- 37:15:28let's try to add a HTML code before this
- 37:15:31particular data set so that we have u an
- 37:15:34understanding of what exactly happened
- 37:15:36here. So I let's uh I'll just quickly
- 37:15:38copy paste this particular code here. So
- 37:15:40we will run it and now next we will
- 37:15:43describe check the data type column
- 37:15:44separate and everything and we will you
- 37:15:46know create data type category variables
- 37:15:50right now so far so good. Now we will
- 37:15:54create the categories.
- 37:15:56So num call
- 37:15:58equals to
- 37:16:02hours studied
- 37:16:05comma attendance.
- 37:16:10So let's quickly add the data
- 37:16:14previous exam scores.
- 37:16:17Just a minute. Let's quickly add the
- 37:16:19columns. Let me take a while. There you
- 37:16:22go. I've added the columns. So we are
- 37:16:23considering three different columns. uh
- 37:16:26our study attendance and previous scores
- 37:16:28for num call and cat call. We are
- 37:16:30considering the other columns apart from
- 37:16:32the first three and our target value is
- 37:16:35exam score. Let's quickly run this.
- 37:16:36There you go. And now we will try to
- 37:16:40build some visualizations and before
- 37:16:42that let's try to uh add some data uh
- 37:16:46from in HTML. Let's try to create a
- 37:16:50Okay, what we can do is simply copy this
- 37:16:53particular HTML file here. We can take
- 37:16:57this
- 37:16:59and add it here so that we will
- 37:17:01understand what exactly is happening
- 37:17:03next. And in place of reading student
- 37:17:06data, we will write data visualization
- 37:17:10for student data.
- 37:17:13And we will keep the colors same dark
- 37:17:15blue background and u the color for data
- 37:17:19visualization will be light blue and
- 37:17:20student data will be orange. Let's
- 37:17:23quickly run and there we have it. Now
- 37:17:25our target variable is exam score. Right
- 37:17:28now we will compare this particular
- 37:17:31target variable with three other
- 37:17:34parameters. So our parameters will be
- 37:17:37the following uh as we discussed uh
- 37:17:40creating the segregation in data set
- 37:17:42right. So first will be distribution of
- 37:17:45target variable with other parameters.
- 37:17:48So we will create another HTML file for
- 37:17:51that right here. Just a minute while I
- 37:17:53paste the code for um HTML quickly run
- 37:17:57this. There you go. So distribution of
- 37:18:00target variable exam score against some
- 37:18:03parameters. Now we will write the plot
- 37:18:06for that
- 37:18:08plot dot figure. So we are going to
- 37:18:12consider the size equals to 15 6
- 37:18:17big size
- 37:18:20is equals to 15 6 the same one that we
- 37:18:24considered before. And we will be using
- 37:18:26Cbond SNS dot count plot
- 37:18:32open bracket. This is a function x is
- 37:18:34equals to df
- 37:18:37clean. Okay. Uh what we can do is just
- 37:18:39quickly take the column name so that we
- 37:18:42don't create any mistakes here and we
- 37:18:46will paste it over here instead of df.
- 37:18:49There you go. Or target.
- 37:18:53So our target is exam score,
- 37:18:57and we will use the pellet as green
- 37:19:02and the plot title will be distribution
- 37:19:04of target variable exam score. We can
- 37:19:07copy this. Okay, just a minute before
- 37:19:10that plt do.
- 37:19:14Should be let's use double quotes now.
- 37:19:16Copy this and paste it here.
- 37:19:20I think semicolons went off. Okay, not a
- 37:19:23problem.
- 37:19:24It's right here. Let's add a dot as a
- 37:19:27full stop. And the next line, if you
- 37:19:31need, you can add the full stop. If not
- 37:19:32you can ignore plt dot grid true
- 37:19:38which equals to major
- 37:19:43comma
- 37:19:44access is equals to y
- 37:19:49comma line style equals to so I want
- 37:19:55lines to be hyphen hyphen in this way I
- 37:19:57want the lines and comma line width um
- 37:20:01let's say 0.5 or 0.7. Let's go with 0.7
- 37:20:07line width equals to 0.9 mm. There you
- 37:20:11go. Now let's quickly run this query.
- 37:20:13There you go. Done. And we have the
- 37:20:16first visualization. So here you can see
- 37:20:19there are some students which are
- 37:20:21scoring 58 59 and you can see maximum
- 37:20:24number of students are already scoring
- 37:20:25good marks which is under 65 and uh
- 37:20:28sorry which is under 70 and above 65 and
- 37:20:32there is a good number of students uh
- 37:20:34which are also scoring uh above 65 as
- 37:20:37well right so sorry 70 70 as well. So we
- 37:20:40have now less than or equal to 70 71.
- 37:20:43And highest scorer in some situations
- 37:20:46there is also 100. If you can see there
- 37:20:48is slight growth here. There are a few
- 37:20:50students toppers maybe which have
- 37:20:52already scored 100 as well. Now we have
- 37:20:55the list here. Now we what we need to do
- 37:20:56is we need to segregate that is part one
- 37:20:59which is less than or equal to 64 which
- 37:21:02falls under 65 and another category
- 37:21:05which falls in between 65 to 70 and
- 37:21:08above 70. So we need to categorize these
- 37:21:11three uh datas and segregate them as
- 37:21:13bottom 65, top which is above 70 and mid
- 37:21:17between 70 to 65. Right now before that
- 37:21:21if you want to add an HTML document uh
- 37:21:24sorry segment here which explains about
- 37:21:26the distribution of target you can also
- 37:21:28do that. It's already added here. Now
- 37:21:31let's continue. But in case if you want
- 37:21:33to uh add uh some data which explains
- 37:21:37that we're trying to segregate, you can
- 37:21:39also do that. I would like to do that.
- 37:21:42Let's quickly uh add that HTML code
- 37:21:44here. So what this particular code will
- 37:21:47do is it will tell the percentage of
- 37:21:49students which are scoring less than 65.
- 37:21:52Number of stu uh percentage of students
- 37:21:55uh scoring in between 65 and 70 and the
- 37:21:58percent of students which are scoring
- 37:22:00beyond 70. Right? Let's run this. And
- 37:22:02here we have the result. Bottom 21.81%
- 37:22:05scores under 64 while 24% scores over 70
- 37:22:09and 50% are in between 65 to 69. Right
- 37:22:13now let's uh continue with the
- 37:22:15segregation part of the data. So for
- 37:22:17segregation we will create three
- 37:22:19different variables A, B, C. So first A
- 37:22:22is equals to length of DF claimed. So
- 37:22:25let's copy the column name sorry data
- 37:22:29frame name length of DF cleaned inside
- 37:22:34the square brackets we'll again add DF
- 37:22:36cleaned of target variable which is exam
- 37:22:40score
- 37:22:42let's also add uh single quotes here
- 37:22:46who are scoring in between or um less
- 37:22:50than let's start with less than or equal
- 37:22:53to 64 we'll not consider is 65 we'll
- 37:22:56consider 64 divided by len of df cleaned
- 37:23:02target variable exam score single quotes
- 37:23:06into 100
- 37:23:09which will give us the percentage now
- 37:23:11similarly let's just copy and paste this
- 37:23:14three more times for B and C. So here
- 37:23:18instead of minus we're supposed to add
- 37:23:19equals to and another one. So here
- 37:23:24equals to and instead of A I will write
- 37:23:26B and the last one is C. And instead of
- 37:23:3064 we will add 70 here for top and here
- 37:23:35we will make some changes. It should be
- 37:23:38greater than or equal to
- 37:23:4265. So this is the third category A B C
- 37:23:45and then we will proceed with printing
- 37:23:50the files. So print
- 37:23:54the bottom
- 37:23:57for the first one which is f of
- 37:24:01a
- 37:24:04is to do 2f
- 37:24:07and we will add the percentage symbol
- 37:24:08over here r under
- 37:24:1264.
- 37:24:14Now we can copy paste the same here and
- 37:24:17we can change the variables.
- 37:24:21So here we will be adding under over 70.
- 37:24:27Lastly in between the ones in between
- 37:24:3365 and 70.
- 37:24:36There you go. Here we will change the
- 37:24:39values from A to B and here A to C.
- 37:24:44There you go. And we can quickly run
- 37:24:46this query. So it's not 70. It was
- 37:24:49supposed to be 69.
- 37:24:51There you go. So we forgot to mention
- 37:24:54this particular one. Now let's run this.
- 37:24:58There you go. So we have 21%
- 37:25:01of people who are scoring under 64, 24%
- 37:25:04over 70 and 53% are in between average.
- 37:25:08So I think the school is focusing on
- 37:25:10improving this particular percentage,
- 37:25:13reducing this particular percentage and
- 37:25:15increasing this particular percentage
- 37:25:17and try to eliminate if possible this
- 37:25:19particular one which are under 64. So
- 37:25:22that is the overall moto I guess. Now so
- 37:25:25far so good. Let's now try to remove
- 37:25:28infinite values from HTML, right? So
- 37:25:30before that, let's add uh this
- 37:25:32particular HTML code here. So which
- 37:25:34explains what we are trying to do. So we
- 37:25:37will first implement the code that
- 37:25:39prevents warning about infinite values
- 37:25:41during data visualization. And now let's
- 37:25:43add the code which will try to eliminate
- 37:25:45the uh infinite values. Let's copy this
- 37:25:48particular data frame cleaned uh data
- 37:25:52frame name here. Now df cleaned
- 37:25:56dotreplace
- 37:26:01square brackets np dot info
- 37:26:07comma
- 37:26:08minus np
- 37:26:11dot info out of these square brackets
- 37:26:14dot np
- 37:26:16na
- 37:26:18non na values we're trying to eliminate
- 37:26:20na values in place of those values you
- 37:26:24can write true and after that we will
- 37:26:27try to eliminate the null. So if it is
- 37:26:31null
- 37:26:33dot sum give me the total number of null
- 37:26:36values after this. Right? So let's try
- 37:26:38to execute that. There you go. So all
- 37:26:40the null entries have been removed here.
- 37:26:43Now let's see the distribution of
- 37:26:45numerical values here. So before that
- 37:26:48let's add the HTML code for that. So
- 37:26:51let's quickly run this. So distribution
- 37:26:54of numerical values. So we will be
- 37:26:56considering these three parameters. So
- 37:26:59if you go back here you can see our
- 37:27:02studied attendance and previous scores.
- 37:27:05So we will be considering these three
- 37:27:07values or these three columns and check
- 37:27:10the distribution of these variables
- 37:27:13against the exam score. So uh is it
- 37:27:16making any um you know kind of variation
- 37:27:19if the u attendance is high? If if the
- 37:27:22attendance is high is the exam score
- 37:27:24high and uh apart from that we have if
- 37:27:28our studies is high is the uh mark score
- 37:27:31is high and if the previous scores are
- 37:27:34high is there a chance to get better
- 37:27:36scores in this particular exam. So what
- 37:27:38we are trying to do is we are trying to
- 37:27:40see if there is any direct involvement
- 37:27:42of number of study hours and number of
- 37:27:45days attended and number of uh or the
- 37:27:48number of marks they received in the
- 37:27:49previous course and we'll try to build a
- 37:27:52visualization on that front.
- 37:27:55So let's go and build that. So we'll try
- 37:27:57the try to write the code here.
- 37:28:00Figure axis
- 37:28:05equals tot
- 37:28:08dot
- 37:28:10subplots. So we will be having three
- 37:28:12different plots here since we're
- 37:28:13considering three different u
- 37:28:16categories.
- 37:28:18And the fixed size should be equal to
- 37:28:2212A 4. There you go. Now
- 37:28:28access
- 37:28:30is equals to access dot variable
- 37:28:36per idx
- 37:28:39comma call
- 37:28:41in enumerate
- 37:28:52and we will try to import the seaborn
- 37:28:54library here. We will try to create
- 37:28:57histo plots here. Histograms here
- 37:29:01plot.
- 37:29:22So line style we will be selecting this
- 37:29:24one
- 37:29:25and the comma
- 37:29:28line width will be 0.7.
- 37:29:34There you go. Next will be access
- 37:29:41dot set
- 37:29:44title. So for this we will be uh setting
- 37:29:47the title as distribution of columns. So
- 37:29:50the columns will be the three uh ones
- 37:29:52attendance, hours studied and uh what
- 37:29:55was the third one that we considered
- 37:29:59previous course. Right? So instead of
- 37:30:02mentioning them specifically, what we
- 37:30:03can do is we can just write columns
- 37:30:05here. C L and close.
- 37:30:10There you go. And lastly,
- 37:30:15plt.tight Right.
- 37:30:18And show the plot. There you go. Let's
- 37:30:21quickly run this query. Run. And now we
- 37:30:24will be having the visualizations here.
- 37:30:27So um you can directly see the
- 37:30:29involvement of these three parameters
- 37:30:30here. If uh they are trying to help if
- 37:30:34the number of hours are increased then
- 37:30:36you can see if there is a better
- 37:30:37improvement in scores. If the attendance
- 37:30:39is increased, if there is a betterment
- 37:30:40in scores or if the previous uh scores
- 37:30:44are helping then you can find it out how
- 37:30:46it is. There you go. Now we can write a
- 37:30:49result here in the form of HTML page.
- 37:30:52And if we run this, it will give you the
- 37:30:54result. The breaks or gaps in the hour
- 37:30:56study variable may be due to the
- 37:30:58respondents answering appropriately. The
- 37:31:00variables attendance and previous scores
- 37:31:02which exhibit a uniform distribution
- 37:31:05have a normal impact on exam scores
- 37:31:07variable which is our target variable.
- 37:31:09Now let's proceed with another part of
- 37:31:11this session which will be about the
- 37:31:13relationship between the numerical
- 37:31:15values and the target variables. Now we
- 37:31:17will copy paste the same code and make
- 37:31:19some minute changes to it. So the only
- 37:31:22change that we did to it is we're trying
- 37:31:24to uh build a scatter plot. So we will
- 37:31:27be getting a scatter plot here. But
- 37:31:28before that let's try to add another
- 37:31:33column here and try to add an HTML code
- 37:31:35which explains why we are doing it.
- 37:31:38There you go a scatter plot. So
- 37:31:40basically these two are one and the
- 37:31:41same. Here we use some column graphs. So
- 37:31:44here we did the same using the scatter
- 37:31:46plot which will help for a better
- 37:31:48understanding. Now we will try to build
- 37:31:50some correlations.
- 37:31:52So basically a list is called
- 37:31:54correlation is created containing the
- 37:31:56names of the columns for which you want
- 37:31:58to calculate the correlation. So here in
- 37:32:00our situation it is the df c r which is
- 37:32:04a data frame and it is created by
- 37:32:06selecting only these columns for df
- 37:32:08cleaned data frame effectively created a
- 37:32:11new data frame containing only the
- 37:32:12specified columns. Now the second one
- 37:32:14which is the co r which is a calculate
- 37:32:17correlation. So this method computes the
- 37:32:20correlation matrix for selected columns
- 37:32:22which is n df c r. The one indicates
- 37:32:26perfect positive correlation minus one
- 37:32:28indicates the perfect negative
- 37:32:30correlation and zero indicates no
- 37:32:31correlation. And lastly the setup of
- 37:32:34plot. This line sets up the figure size
- 37:32:37of the plot. In this particular
- 37:32:39situation we are choosing five and four.
- 37:32:42Right now let's quickly try to execute
- 37:32:44this query and see the answer. And we
- 37:32:47will also add the HTML code for this so
- 37:32:50that we have a better understanding for
- 37:32:52this. So we will be adding that HTML
- 37:32:56uh box here which will explain what
- 37:32:58exactly happened here. So this is our
- 37:33:01plot and this is the correlation. Now
- 37:33:03let's try to add that HTML code right
- 37:33:06here. The result of this particular data
- 37:33:09visualization will be maintained here.
- 37:33:12So the hours studying and attendance
- 37:33:14shows a positive correlation with the
- 37:33:16target variable which is exams hour.
- 37:33:18However, previous course appears to have
- 37:33:20no or little relationship with the
- 37:33:22target variable. Right now let's
- 37:33:24continue with our next uh part of this
- 37:33:27session. So now we will try to identify
- 37:33:30the relationship between studies
- 37:33:34hours and attendance and extracurricular
- 37:33:37scores. Right? So we have other u
- 37:33:40columns to consider which is
- 37:33:41extracurricular activities. So there is
- 37:33:43a belief that extracurricular activities
- 37:33:45will also help students to study better.
- 37:33:47So we will find if there is a relation
- 37:33:49between the target variable and this
- 37:33:51extracurricular activities attendance
- 37:33:53and study hours. Now let's quickly add
- 37:33:56the code here. Now let's quickly execute
- 37:33:58the code. Now we have the visualization
- 37:34:02which explains the relationship between
- 37:34:04the number of hours studied
- 37:34:06extracurricular activities etc. So here
- 37:34:08it is and now let's add an HTML code
- 37:34:10which explains about this result
- 37:34:13in this particular code was supposed to
- 37:34:15be added here.
- 37:34:20So this uh is the resultant column here.
- 37:34:24Now here it explains about the
- 37:34:26influences that it performs. So the
- 37:34:28extracal activities, parental income and
- 37:34:30extra things that influence the scores
- 37:34:34and there you go.
- 37:34:37Now let us also consider other columns
- 37:34:40right the other parameters like
- 37:34:43resources are available or not parental
- 37:34:45education and other things which also
- 37:34:47might have influenced the exam scores of
- 37:34:50students. So for that let's add an HTML
- 37:34:52code so that we have an HTML page here
- 37:34:54which explains what is the next
- 37:34:55procedure that we are following. Right
- 37:34:57now let's add the query here. So here we
- 37:35:00are considering the other parameters
- 37:35:02like family income, peer influence,
- 37:35:04motivation level, gender, parental
- 37:35:06involvement, parental educational level
- 37:35:08and extracurricular activities. And we
- 37:35:10are considering them against the target
- 37:35:13variable which is exam score. And now
- 37:35:16let's execute this query. There you go.
- 37:35:19Now we have generated a graph which
- 37:35:20explains about this particular
- 37:35:23parameters against the target variable.
- 37:35:25And now let's add the HTML page here
- 37:35:27which explains about these results.
- 37:35:31Let's quickly run it. And there you go.
- 37:35:34So when certain factors affect Q1 and Q2
- 37:35:36but not Q2, it can be understood that
- 37:35:38individual has overcome challenges
- 37:35:40through personal effort. Right? So if
- 37:35:42government policies and corporate social
- 37:35:45contributors are focused on addressing
- 37:35:47these aspect, it seems that we could
- 37:35:49create dynamic country with greater
- 37:35:52social mobility and open opportunities
- 37:35:53for all. Right? So if extracurricular
- 37:35:56activities can outweigh the influence of
- 37:35:58other variables in academic performance
- 37:35:59then we should foster that kind of
- 37:36:01environment right. So this is how u you
- 37:36:04can get extract some statistical
- 37:36:07analysis on this particular data set.
- 37:36:09Now let's quickly rename this uh python
- 37:36:14eda
- 37:36:16students
- 37:36:18performance
- 37:36:20and you can quickly rename and save it.
- 37:36:23Welcome to math refresher probability
- 37:36:26and statistics.
- 37:36:28In this lesson, we are going to explain
- 37:36:30the concepts of statistics and
- 37:36:33probability.
- 37:36:34Describe conditional probability. Define
- 37:36:37the chain rule of probability. Discuss
- 37:36:40the measure of variance. Identify the
- 37:36:42types of gshian distribution.
- 37:36:45Basic of statistics and probability.
- 37:36:48Probability and statistics. Data science
- 37:36:51relies heavily on estimates and
- 37:36:53predictions. A significant portion of
- 37:36:56data science is made up of evaluations
- 37:36:58and forecast.
- 37:37:00Statistical methods are used to make
- 37:37:02estimates for further analysis.
- 37:37:05Probability theory is helpful for making
- 37:37:07predictions. Statistical methods are
- 37:37:10highly dependent on probability theory
- 37:37:13and all probability and statistics are
- 37:37:16dependent on data. Data is information
- 37:37:19acquired for reference or research via
- 37:37:22observations, facts, and measurements.
- 37:37:26Data is a set of facts structured in the
- 37:37:29form that computers can interpret such
- 37:37:31as numbers, words, estimations, and
- 37:37:34views. Importance of data. Data aids in
- 37:37:38seeing more about the information by
- 37:37:40identifying possible connections between
- 37:37:42two features. Data assists in the
- 37:37:45detection of distortion by uncovering
- 37:37:48hidden patterns based on prior
- 37:37:50information patterns. Data may be
- 37:37:53utilized to anticipate the future or
- 37:37:55predict the current state of affairs.
- 37:37:58Also, data aids in determining whether
- 37:38:00two pieces of information have any
- 37:38:02instance in common or not. Types of
- 37:38:05data. Data might be quantitative. That
- 37:38:09is data that can be measured or counted
- 37:38:11in numbers. Or it may be qualitative
- 37:38:14which is data which is generally divided
- 37:38:16into groups or in simpler words which
- 37:38:19cannot be counted or measured in
- 37:38:21numbers. Let's consider an example. A
- 37:38:24customer information data of a bank may
- 37:38:27contain quantitative and qualitative
- 37:38:29data. Consider this snapshot where we
- 37:38:32have customer ID, surname, geography,
- 37:38:36gender, age, balance, has C or card is
- 37:38:39active member. Amongst these variables
- 37:38:42we can see surname is mostly qualitative
- 37:38:45as it cannot be counted and measured in
- 37:38:47numbers. Geography and gender are also
- 37:38:51qualitative as they cannot be counted in
- 37:38:53numbers and are mostly groups. has C or
- 37:38:57card that is has credit card and is
- 37:39:00active member although are containing
- 37:39:02numerical in form but these are
- 37:39:05categorical that means these have been
- 37:39:07divided into groups of one and zero that
- 37:39:11represent yes and no as an answer hence
- 37:39:14these two variables are also qualitative
- 37:39:18customer ID is again although a
- 37:39:21numerical data however the significance
- 37:39:24or intuition behind Customer ID is
- 37:39:27categorical.
- 37:39:28Hence, it may be kept in the qualitative
- 37:39:31data also. However, age and balance
- 37:39:35these are numerical information which
- 37:39:37have been measured or counted and
- 37:39:39numerical operations can be performed on
- 37:39:42them. Hence, these are under
- 37:39:44quantitative data categories.
- 37:39:46Introduction to descriptive statistics.
- 37:39:49Descriptive statistics. A descriptive
- 37:39:52measurement is summary measure that
- 37:39:54quantitatively portrays the most
- 37:39:56important features of a set of data
- 37:39:59allowing for a better comprehension of
- 37:40:01the information. Data can be measured as
- 37:40:04different levels. The levels of
- 37:40:06measurement describe the nature of
- 37:40:08information stored in the data assigned
- 37:40:10to the variables. Qualitative data can
- 37:40:13be measured as nominal or ordinal.
- 37:40:15Quantitative data can be measured in
- 37:40:17terms of interval and ratio type.
- 37:40:20Nominal data. The data is categorized
- 37:40:23using names, labels or qualities. For
- 37:40:26example, brand name, zip code, and
- 37:40:28gender. Ordinal data can be arranged in
- 37:40:31order or ranked, and can be compared.
- 37:40:34Examples include grades, star reviews,
- 37:40:38position, and race, and date. Interval
- 37:40:41data is the data that is ordered and has
- 37:40:43meaningful differences between the data
- 37:40:46points. Example, temperature in Celsius
- 37:40:49and year of birth. Ratio data is similar
- 37:40:52to the interval level with the added
- 37:40:55property of inherent zero. Mathematical
- 37:40:58calculations can be performed on both
- 37:41:00interval as well as ratio data. For
- 37:41:03example, height, age, and weight.
- 37:41:06Population versus sample. Before
- 37:41:09analyzing the data, it's important to
- 37:41:11figure out if it's from a population or
- 37:41:13a sample. Population is a collection of
- 37:41:17all available items as well as each unit
- 37:41:19in our study. Sample is a subset of the
- 37:41:22population that contains only a few
- 37:41:25units of the population. Population data
- 37:41:28is used for study when the data pool is
- 37:41:31very small and can give all the required
- 37:41:33information. Samples are collected
- 37:41:36randomly and represent the entire
- 37:41:39population in the best possible way.
- 37:41:42Measures of central tendency. The
- 37:41:45central tendency is a single value that
- 37:41:48aids in the description of the data by
- 37:41:50determining its center position.
- 37:41:53Measures of central tendency are
- 37:41:55sometimes known as summary statistics or
- 37:41:58measures of central location. The most
- 37:42:01popular measurements of central tendency
- 37:42:04are mean, median, and mode. The normal
- 37:42:07distribution is a bell-shaped
- 37:42:10symmetrical distribution in which mean,
- 37:42:12median, and mode all are equal. The
- 37:42:15curve over here shows the bell-shaped
- 37:42:17curve or the normal distribution of
- 37:42:19variable X. The point over here that is
- 37:42:22X1 is the point which represents the
- 37:42:26mean, median and mode of this
- 37:42:28distribution. Mean mean is calculated by
- 37:42:31dividing these sum of all data values by
- 37:42:34the total number of data values. It gets
- 37:42:38affected when there are unusual or
- 37:42:40extreme values. It is sensitive to the
- 37:42:43outliers. Mean can be calculated as
- 37:42:46summation over all the values of X in a
- 37:42:49collection divided by the size of the
- 37:42:51collection.
- 37:42:53For example, we have a collection where
- 37:42:55we have values as 7 3 4 1 6 and 7.
- 37:43:01We find out the sum of these values
- 37:43:03which is 28 and there are total of six
- 37:43:07values. So 28 / 6 gives us a mean value
- 37:43:11of 4.66.
- 37:43:14Median,
- 37:43:16it is the middle value in the set of the
- 37:43:18data that has been sorted in ascending
- 37:43:20order.
- 37:43:22It is a better alternative to mean since
- 37:43:24it is less impacted by outliers and
- 37:43:27skewess.
- 37:43:28It is closer to the actual central
- 37:43:31value.
- 37:43:32Median is calculated differently for
- 37:43:35different sizes of data.
- 37:43:37Differentiated as if the total number of
- 37:43:39values is odd or if the total number of
- 37:43:43values is even. If the size of the data
- 37:43:46is odd. For example, in this case we
- 37:43:50have five elements.
- 37:43:54After sorting whatever middle value we
- 37:43:56get
- 37:43:58that means n + 1 by 2 term in this case
- 37:44:035 + 1 / 2
- 37:44:06that is the third term which is four is
- 37:44:09the median value.
- 37:44:12In case when the total number of values
- 37:44:14is even like here there are six values.
- 37:44:18The average or the mean of the two
- 37:44:20central values is considered as the
- 37:44:22median. In this case the median is the
- 37:44:25mean of 6 and four which is five. Mode.
- 37:44:30Mode represents the most common value in
- 37:44:33the data set. It is not at all affected
- 37:44:36by extreme observations.
- 37:44:40It is the best measure of central
- 37:44:42tendency for highly skewed or non-normal
- 37:44:45distribution.
- 37:44:46Mode for categorical data is determined
- 37:44:49by estimating the frequencies for each
- 37:44:51categories
- 37:44:53and then the category with the highest
- 37:44:55frequency is considered to be mode.
- 37:44:58Like in this case 7 has the highest
- 37:45:01frequency. Hence seven becomes the mode
- 37:45:03value. However, in case of continuous
- 37:45:07data or quantitative data, the
- 37:45:09calculation of mode is slightly
- 37:45:11different. The first step in calculation
- 37:45:14of mode is dividing the data into
- 37:45:16classes which are equal with then
- 37:45:18getting the frequency of data points
- 37:45:20lying in within that range of classes
- 37:45:23and finally selecting the class with the
- 37:45:26highest frequency.
- 37:45:29Using the range of that class and the
- 37:45:31frequencies, we can get the final mode
- 37:45:33value.
- 37:45:35Using the formula L+
- 37:45:39minus F_sub_1 multiplied to H divided by
- 37:45:42FM minus F_sub_1 plus FM minus F_sub_2.
- 37:45:47Here L is the lower limit or the lower
- 37:45:50observation of the mode class.
- 37:45:53H is the size of the mode class.
- 37:45:57FM is the frequency of the mode class.
- 37:46:00F_sub_1 is the frequency of the class
- 37:46:03proceeding to mode. And F_sub_2 is the
- 37:46:06frequency of the class succeeding to
- 37:46:09mode. This gives us the final mode
- 37:46:11value.
- 37:46:13Mean versus expectation.
- 37:46:16Now let's talk about mean versus
- 37:46:18expectation.
- 37:46:19So in general we use the expected value
- 37:46:22or expectation when we want to calculate
- 37:46:25the mean of a probability distribution
- 37:46:28that represents the average value we
- 37:46:31expect to occur before collecting any
- 37:46:33data. And mean on the other hand mean is
- 37:46:36basically used when we want to calculate
- 37:46:39the average value of a given sample.
- 37:46:42This represents the average value of raw
- 37:46:45data that we may have already collected.
- 37:46:48We can understand this by using a simple
- 37:46:51example.
- 37:46:53Now to calculate the expected value of
- 37:46:56this probability distribution, we can
- 37:46:59use a specific formula from the previous
- 37:47:01discussion.
- 37:47:03This is going to be the expected value
- 37:47:05where X is going to be the data value
- 37:47:08and this PX is the probability of value.
- 37:47:13For example, we could calculate the
- 37:47:15expected value for this probability
- 37:47:17distribution to be as shown.
- 37:47:22So here it will be 1.45 goals.
- 37:47:27So this represents the expected number
- 37:47:29of goals that the team will score in any
- 37:47:32given game.
- 37:47:34And then if you talk about calculating
- 37:47:36mean, so we typically calculate the mean
- 37:47:39after we have actually collected raw
- 37:47:41data.
- 37:47:44For example, suppose we record the
- 37:47:46number of goals that a soccer team will
- 37:47:48score in 15 different games.
- 37:47:53Now to calculate the mean number of
- 37:47:55goals scored per game,
- 37:47:58we can use the following formula
- 37:48:01where sum of x is basically the sum of
- 37:48:04all the goals divided by n and the
- 37:48:06number of records or we can say the
- 37:48:08sample size.
- 37:48:11It is as shown on the screen.
- 37:48:29So this represents the mean number of
- 37:48:31goals scored per game by the team.
- 37:48:34Measures of asymmetry.
- 37:48:37The difference between the three
- 37:48:38distinct curves can be studied in this
- 37:48:41image.
- 37:48:42The central curve is the normal or no
- 37:48:45skeus curve. Here mean, median and mode
- 37:48:48all lie on the same point. This normal
- 37:48:51curve is symmetrical about its mean,
- 37:48:53median and mode.
- 37:48:57That means the left hand side of the
- 37:48:59curve is a mirror image of the right
- 37:49:01hand side of the curve.
- 37:49:04However, in case of negatively skewed
- 37:49:07data, the tail is elongated on the left
- 37:49:11hand side
- 37:49:12and the mean is smaller than the mode
- 37:49:15and the median values or is on the left
- 37:49:18hand side of the mode.
- 37:49:21Hence indicating that the outliers are
- 37:49:23in the negative direction.
- 37:49:26On the other hand, in case of positively
- 37:49:29skewed, the data is concentrated on the
- 37:49:31left hand side of the curve.
- 37:49:35While the tail is elongated or longer on
- 37:49:37the right hand side of the curve,
- 37:49:41the mean is greater than the mode and
- 37:49:43median
- 37:49:44or is on the right hand side of the mode
- 37:49:47and median indicating that the outliers
- 37:49:49are in the positive direction.
- 37:49:56Let's consider an example.
- 37:49:59The graph here shows the global income
- 37:50:02distribution for the year 2003 2013 and
- 37:50:06a projection for 2035.
- 37:50:09If we see the global income distribution
- 37:50:12statistics for 2003 it is highly right
- 37:50:15skewed.
- 37:50:18We can observe in the previous graph
- 37:50:20that in 2003
- 37:50:24the mean of $3,451
- 37:50:29was higher than the median of $1090.
- 37:50:33The global income is definitely not
- 37:50:36evenly distributed. The majority of
- 37:50:38people make less than $2,000 each year.
- 37:50:44while only a small percentage of the
- 37:50:46population earns more than $14,000.
- 37:50:51Measures of variability.
- 37:50:56Measures of variability.
- 37:50:58Dispersion. The measure of central
- 37:51:01tendencies provide a single value that
- 37:51:03addresses the full worth. However, the
- 37:51:06central tendency cannot depict the
- 37:51:08viewpoint entirely. The metric of
- 37:51:11dispersion helps us focus on the
- 37:51:13inconsistency in the data spread.
- 37:51:16Measures of dispersion describe the
- 37:51:18spread of the data.
- 37:51:21The range, intercortile range, standard
- 37:51:24deviation and variance are examples of
- 37:51:27dispersion measures.
- 37:51:30Range.
- 37:51:32The range of distribution is the
- 37:51:34difference between the largest and the
- 37:51:36smallest amount of data.
- 37:51:39The range, for example, does not include
- 37:51:42all of a series positive aspects.
- 37:51:46It concentrates on the most shocking
- 37:51:48aspects and ignores that aren't
- 37:51:50considered critical. For example, for a
- 37:51:52set 13, 33, 45, 67, 70.
- 37:51:58The range is 57. That is the maximum of
- 37:52:03this which is 70 minus the minimum over
- 37:52:05here which is 13.
- 37:52:10Variance.
- 37:52:12Variance is the average of all squared
- 37:52:14deviations.
- 37:52:17It is defined as the sum of squared
- 37:52:19distance between each point and the mean
- 37:52:22or the dispersion around the mean.
- 37:52:25The standard deviation is used as
- 37:52:28variance suffers from a unit difference.
- 37:52:32Variance can be computed as sigma square
- 37:52:35summation over x - mu^ 2
- 37:52:39divided by n
- 37:52:41where mu is the mean of the data, x is
- 37:52:45the individual data point
- 37:52:48and n is the size of the data.
- 37:52:52This representation is for a population
- 37:52:54data.
- 37:52:56for a sample data variance can be
- 37:52:58computed as X minus
- 37:53:01Xar whole square summation
- 37:53:04over it divided by n minus one.
- 37:53:08Here Xbar is the mean of these sample
- 37:53:11data and n is the sample size.
- 37:53:16The units of values and variance are not
- 37:53:18equal.
- 37:53:20So another variability measure is used.
- 37:53:24Standard deviation.
- 37:53:28Standard deviation is a statistical term
- 37:53:30used to measure the amount of
- 37:53:32variability or dispersion around a mean.
- 37:53:37The standard deviation is calculated as
- 37:53:40the square root of variance. It depicts
- 37:53:43the concentration of the data around the
- 37:53:46mean of the data set.
- 37:53:50Standard deviation as indicated
- 37:53:52previously can be computed as square
- 37:53:54root of variance
- 37:53:56for a population data. Standard
- 37:53:59deviation sigma can be computed as
- 37:54:02square root of summation over x i minus
- 37:54:05mu^ square / n
- 37:54:09where mu is the mean of the data x i are
- 37:54:13the data points and n is the size. Let's
- 37:54:16consider an example.
- 37:54:19Let's find out the mean, variance, and
- 37:54:22standard deviation for this data. The
- 37:54:25data values are 3, 5, 6, 9, and 10. To
- 37:54:30find out the mean, we first find the sum
- 37:54:32of all these data values
- 37:54:36that is 33 and divide it by the count,
- 37:54:39which is five.
- 37:54:42We get the mean of 6.6. To compute the
- 37:54:45variance, we start by computing the
- 37:54:48deviation.
- 37:54:49That is X minus the mean of X. Here 3 is
- 37:54:54one of the values of the data and 6.6 is
- 37:54:57the mean.
- 37:54:59So 3 - 6.6 squared and we do that
- 37:55:05to find out sum of all the deviations
- 37:55:07divided by the count
- 37:55:10which is five.
- 37:55:12We end up getting an overall variance of
- 37:55:146.64.
- 37:55:18Standard deviation as we know is
- 37:55:21measured at square root of variance that
- 37:55:23is square<unk> of 6.64
- 37:55:27which amounts to 2.576.
- 37:55:31Measures of relationship.
- 37:55:34Measures of relationship coariance.
- 37:55:37Covariance is the measure of joint
- 37:55:39variability of two variables.
- 37:55:43It measures the direction of the
- 37:55:44relationship between the variables. It
- 37:55:47determines if one variable will cause
- 37:55:50the other to alter in the same way.
- 37:55:54Coariance between variable X and Y can
- 37:55:57be computed as summation over the
- 37:56:00product of X I - XR
- 37:56:03and Y I - Y bar the whole divided by N
- 37:56:07minus one.
- 37:56:10Here Xar and Y bar are the mean of X and
- 37:56:13Y respectively. The value of covariance
- 37:56:16can range from minus infinity to a plus
- 37:56:19infinity.
- 37:56:22Correlation. Correlation is normalized
- 37:56:25coariance.
- 37:56:28It measures the strength of association
- 37:56:31between two variables. The most common
- 37:56:33measure for correlation is the Pearson
- 37:56:36correlation coefficient.
- 37:56:38Correlation between two variables
- 37:56:42X and Y can be measured with respect to
- 37:56:44coariance as coariance between X
- 37:56:48and Y divided by the standard deviation
- 37:56:51of X and standard deviation of Y.
- 37:56:55The value of correlation ranges from a
- 37:56:58negative 1 to positive 1.
- 37:57:02Types of correlation.
- 37:57:06Correlation can be either a positive
- 37:57:08correlation,
- 37:57:10zero correlation or a negative
- 37:57:12correlation.
- 37:57:16The first picture over here represents a
- 37:57:18perfect positive correlation
- 37:57:22wherein a straight line with a positive
- 37:57:24slope
- 37:57:26is representing the relationship between
- 37:57:28the two variables.
- 37:57:31Zero correlation means that the line
- 37:57:33representing the relationship between
- 37:57:35the two variables is horizontal to the
- 37:57:38xaxis.
- 37:57:41Perfect negative correlation can be
- 37:57:44represented by a straight line with a
- 37:57:46negative slope.
- 37:57:49Correlation equals to 1 implies a
- 37:57:52positive relationship. That is when one
- 37:57:55variable increases the other variable
- 37:57:57also increases. A correlation value of
- 37:58:00negative one implies a negative
- 37:58:02relationship. That is when one variable
- 37:58:05increases the other decreases.
- 37:58:09The correlation coefficient of zero
- 37:58:12shows that the variables are completely
- 37:58:14independent of each other.
- 37:58:17Let's consider an example.
- 37:58:21Here we have two variables height and
- 37:58:24weight.
- 37:58:27To compute the correlation between
- 37:58:29height and weight,
- 37:58:31we use the correlation formula as
- 37:58:33covariance of X
- 37:58:36and Y divided by standard deviation of X
- 37:58:39and standard deviation of Y.
- 37:58:42Here height is the X variable and weight
- 37:58:45is the Y variable.
- 37:58:48First to compute coariance we compute
- 37:58:51the x - xar and y - y bar values and
- 37:58:56then the product of them.
- 37:58:59We then compute x - xr²
- 37:59:04and y - y bar square values to compute
- 37:59:07the standard deviations of height and
- 37:59:09weight respectively. Correlation as we
- 37:59:12know has been defined as covariance of X
- 37:59:16and I and Y divided by standard
- 37:59:18deviations of X and Y.
- 37:59:22This can also be represented as
- 37:59:24summation over x - xr multiplied to y -
- 37:59:29y bar
- 37:59:31divided by square root of summation over
- 37:59:33sum of squared deviations that is x - xr
- 37:59:38square multiplied to square root of
- 37:59:40summation over y - y bar square that is
- 37:59:45sum of square deviations for y.
- 37:59:49Now let's find out values to put into
- 37:59:52this formula.
- 37:59:55First we find out the overall sum of
- 37:59:58height to get the mean of height which
- 38:00:01is 5.14.
- 38:00:03Similarly we get the sum of weight to
- 38:00:05get the mean of weight as 50. We now get
- 38:00:08the summation over x - xr multiplied to
- 38:00:12y - y bar to get the numerator for the
- 38:00:15formula. Then we compute x - xr square
- 38:00:19summation
- 38:00:21and y - y bar square that is sum of
- 38:00:24squared deviation of x and y
- 38:00:27respectively.
- 38:00:29Now we put in the values in this final
- 38:00:31correlation formula to get a correlation
- 38:00:33value of 0.889.
- 38:00:38This indicates that height and weight
- 38:00:40have a positive relationship.
- 38:00:43It is evident that as height grows,
- 38:00:46weight also increases.
- 38:00:50In this module, we will be talking about
- 38:00:52expectation and variance.
- 38:00:56So the expected value or we can say mean
- 38:00:58of a given variable that we can denote
- 38:01:01by X is a discrete random variable where
- 38:01:04it is a weighted average of the possible
- 38:01:06values that X can take and each value is
- 38:01:09going to be according to the probability
- 38:01:11of that specific event occurring.
- 38:01:15So usually the expected value of X is
- 38:01:17denoted by a simple formula where we can
- 38:01:21define the expectation based on the X
- 38:01:23parameter
- 38:01:26which is going to be the sum of each
- 38:01:28possible outcome multiplied by the
- 38:01:30probability of the outcome occurring.
- 38:01:34So in more concrete terms, the
- 38:01:37expectation is what we would expect the
- 38:01:39outcome of an experiment to be on
- 38:01:41average.
- 38:01:45We can take an example for the coin. If
- 38:01:48a coin is being tossed 10 times, then
- 38:01:51one is most likely to get five heads and
- 38:01:54five tails.
- 38:01:57Same logic can be discussed if we talk
- 38:02:00about another example of rolling a
- 38:02:02dieice. So there are six possible
- 38:02:04outcomes when you roll a dieice. 1 2 3 4
- 38:02:085 6. And each of these has a probability
- 38:02:11of 1x 6 of occurring. So we can say that
- 38:02:15the expectation is going to be 1
- 38:02:17multiplied by the probability of that
- 38:02:19happening which is going to be 1x 6 + 2x
- 38:02:236 + 3x 6 + 4x 6 + 5x 6 + 6x 6 and that
- 38:02:31is going to give us 3.5 as an output.
- 38:02:34The expected value is 3.5.
- 38:02:38So if you think about it, 3.5 is halfway
- 38:02:41between the possible values that I can
- 38:02:43take and this is what we should have
- 38:02:46expected.
- 38:02:47Next we talk about the concept of
- 38:02:49variance. So variance of a random
- 38:02:52variable allows us to know something
- 38:02:54about the spread of the possible values
- 38:02:57of the variable.
- 38:02:59So for a discrete random variable X, the
- 38:03:02variances of X is going to be denoted by
- 38:03:04using a simple formula that is going to
- 38:03:06be var=
- 38:03:08E X - M the whole square where M is
- 38:03:12basically the expected value of the
- 38:03:14expectation of X. So this is more like a
- 38:03:17standard deviation of X which can also
- 38:03:20be represented by using this formula. So
- 38:03:23the variance does not behave in the same
- 38:03:25way as expectation when we multiply and
- 38:03:28add constants to random variables.
- 38:03:33So now there are two different type of
- 38:03:35variance that we can have a fair
- 38:03:37understanding on. First of all we have
- 38:03:40low variance and then we have high
- 38:03:43variance.
- 38:03:45So low variance simply means that there
- 38:03:48is a small variation in the production
- 38:03:50of the target function with changes in
- 38:03:53the trading data set and at the same
- 38:03:55time high variance as we can see here
- 38:03:58high variance shows a large variation in
- 38:04:01prediction of the target function with
- 38:04:03changes in the trading data set. So a
- 38:04:06model that shows high variance learns a
- 38:04:08lot and perform well with the training
- 38:04:10data set and it does not generalize well
- 38:04:13with the unseen data set and that's why
- 38:04:16as a result such a model gives good
- 38:04:19results with training data set but shows
- 38:04:21high error rates on the test data set
- 38:04:24and since the high variance a model
- 38:04:26learns too much from the data set it
- 38:04:29leads to an overfitting of the model. So
- 38:04:32model with high variance will be having
- 38:04:34couple of issues like it may lead to
- 38:04:36overfitting or it may also lead to
- 38:04:39increase in model complexities.
- 38:04:43Next we have skewess.
- 38:04:46So skewess in simple terms is basically
- 38:04:49a measure of asymmetry of a
- 38:04:51distribution. So distribution is
- 38:04:54asymmetrical when its left and right
- 38:04:56sides are not the mirror images.
- 38:04:59Right now this is a mirrored image and a
- 38:05:02distribution can have right positive or
- 38:05:04we can say negative or it can have zero
- 38:05:07skewess.
- 38:05:10So right skewed in this scenario is
- 38:05:13basically the distribution is longer on
- 38:05:15the right side of its peak
- 38:05:18and a left skew distribution is going to
- 38:05:20be we can say where it is longer on the
- 38:05:23left side.
- 38:05:25So we can see we have this one as a part
- 38:05:28of right side. It is more elongated
- 38:05:30towards the right side and this one is
- 38:05:33more elongated towards the left side. So
- 38:05:36we can think of skewess in terms of
- 38:05:38tails. A tail is long tampering and the
- 38:05:41end of a distribution. So it simply
- 38:05:44indicates that they are observations at
- 38:05:46one end of the distribution but that
- 38:05:48they are relatively infrequent. So a
- 38:05:51right skew distribution has a long tail
- 38:05:54on the right side as you can see here.
- 38:05:56So the number supports observed. Let's
- 38:05:59say we have a data on a per year basis.
- 38:06:02So again we can have a more skewess
- 38:06:04towards the right side where data is
- 38:06:06being dropping as we continue to
- 38:06:09increase the number of years. For
- 38:06:11example we may have a high sales towards
- 38:06:14the beginning of year suppose in 2022
- 38:06:17but again as we proceed to 2023 second
- 38:06:21half we are seeing the dip in
- 38:06:23performance. So that is rightly skewed
- 38:06:26and same way let's suppose if we started
- 38:06:28with the sales figure it was really less
- 38:06:31in suppose 2002
- 38:06:34but again as we proceeded to 2023 now
- 38:06:37our sales have been gradually
- 38:06:39increasing. So it's more like skew
- 38:06:42towards the left section as a part of
- 38:06:44negative skew. Next we have curtosis.
- 38:06:49So curtosis is basically a measure of
- 38:06:52the tailness of a distribution.
- 38:06:55So taeness is how often the outliers
- 38:06:58occur and act as curtis is the tailness
- 38:07:01of the distribution related to a normal
- 38:07:04distribution. So a distribution with
- 38:07:07medium curttosis is called as meocurtic.
- 38:07:10A distribution with low curtosis like
- 38:07:12this one. This is called as the
- 38:07:15platicurtic and then distribution with
- 38:07:17high curtosis like this one. This is
- 38:07:19called as the leptoccuric.
- 38:07:23So tails here they are tapering ends on
- 38:07:25either side of a distribution like this.
- 38:07:28So they represent the probability or the
- 38:07:31frequency of values that are extremely
- 38:07:33high or extremely low to the mean.
- 38:07:36In other words, tails here represents
- 38:07:39how often the outliers occur.
- 38:07:43So there are three type of curtis. We
- 38:07:46have platocurtic which is negative,
- 38:07:48leptocortic which is a positive towards
- 38:07:50the upper end and then we have messertic
- 38:07:53which is a normal distribution. So
- 38:07:56messertic is the medium tail. So normal
- 38:07:59distributions they have a curtosis of
- 38:08:01three. So any distribution with a curtis
- 38:08:04of a prox value of three is going to be
- 38:08:07messertic. And curtosis is described in
- 38:08:10terms of excess curttosis which is
- 38:08:13curtosis minus3. And since normal
- 38:08:16distribution they have a curtosis of
- 38:08:19three axis curtises makes comparing a
- 38:08:22distribution curtosis to a normal
- 38:08:24distribution even easier. Introduction
- 38:08:27to probability.
- 38:08:30Probability theory. Probability is a
- 38:08:33measure of the likelihood that an event
- 38:08:35will occur.
- 38:08:38Let's consider an example of coin toss
- 38:08:42where the chances of getting heads on a
- 38:08:44coin are 1 by two or 50%.
- 38:08:48The probability of each given event is
- 38:08:50between zero and one both inclusive.
- 38:08:54Sum of an events cumulative probability
- 38:08:57cannot be greater than one.
- 38:09:00Hence the probability of an event x lies
- 38:09:03between zero and one. This means that
- 38:09:06the integral of probability of
- 38:09:08distribution over x equals to 1.
- 38:09:14Conditional probability. Conditional
- 38:09:16probability of any event A is defined as
- 38:09:20the probability of occurrence of A given
- 38:09:23that event B has previously occurred.
- 38:09:28Condition probability of event A given B
- 38:09:31can be estimated as probability of A
- 38:09:34intersection B that is probability of
- 38:09:37both A and B happening together
- 38:09:40divided by the probability of B.
- 38:09:45It is also written as that probability
- 38:09:47of A intersection B equals to
- 38:09:50probability of A given B multiplied to
- 38:09:54probability of B.
- 38:09:59Let's consider an example.
- 38:10:01In a coin, we are doing a two coin flip.
- 38:10:04Coin one gets heads, tails, heads, and
- 38:10:07tails in subsequent flips.
- 38:10:12while coin 2 gets tails, heads, heads,
- 38:10:15and tails in the subsequent flips. Now,
- 38:10:18the probability that coin one will get a
- 38:10:20head is 2 out of four. While the
- 38:10:23probability that coin two will get heads
- 38:10:26is again two out of four.
- 38:10:29The probability that both coin one and
- 38:10:31coin two will have a heads is just one
- 38:10:34out of the four flips.
- 38:10:38Hence the probability that coin one will
- 38:10:40get heads given that coin 2 is already
- 38:10:43heads can be computed as probability of
- 38:10:46coin one edge intersection coin 2 edge
- 38:10:50that is 1x4 divided by probability of
- 38:10:54coin 2 edge
- 38:10:57that's a given that is 2x 4 which is
- 38:11:00going to be 0.5 or 50% based
- 38:11:05base theorem Base theorem calculates the
- 38:11:08conditional probability of an event
- 38:11:10based on its prior probabilities.
- 38:11:14Basically base theorem incorporates the
- 38:11:17prior probability distribution to
- 38:11:19predict the posterior probabilities base
- 38:11:22theorem for conditional probability
- 38:11:25can be expressed as probability of A
- 38:11:28given B equals probability of B given A
- 38:11:32divided by probability of B multiplied
- 38:11:35to probability of A.
- 38:11:38Base theorem allows updating the
- 38:11:40probability values by using new
- 38:11:42information or evidence. Here
- 38:11:45probability of A is known as prior
- 38:11:48probability. That is the probability of
- 38:11:50event that before any new data is
- 38:11:52collected. Probability of A given B is
- 38:11:56known as the posterior probability. It
- 38:11:59is the revised probability of an event
- 38:12:01occurring after taking into
- 38:12:03consideration the new information
- 38:12:05probability of B given A is known as the
- 38:12:09likelihood and probability of B is
- 38:12:11probability of observing an evidence B
- 38:12:14model. An example consider an example
- 38:12:18for calculating the likelihood of having
- 38:12:20diabetes based on frequency of fast food
- 38:12:23consumption. Here is the observed data.
- 38:12:26Let's say the fast food audience is 20%.
- 38:12:30Diabetes prevalence is 10% and 5% is
- 38:12:34fast food and diabetes.
- 38:12:37The chances of diabetes given fast food
- 38:12:39that is the conditional probability of D
- 38:12:42given B can be calculated as probability
- 38:12:45of diabetes and fast food together
- 38:12:48divided by probability of fast food.
- 38:12:51That means 5% divided by 20%. that
- 38:12:55equals 25%.
- 38:12:57Define an analysis can state eating fast
- 38:13:00food increases the chance of having
- 38:13:02diabetes by 25%.
- 38:13:05The multiplication rule of probability
- 38:13:08if events A and B are statistically
- 38:13:11independent and probability of A
- 38:13:14intersection B can be given as
- 38:13:16probability of A given B multiplied to
- 38:13:20probability of B. However, probability
- 38:13:23of A intersection B is also given as
- 38:13:26probability of A multiplied to
- 38:13:29probability of B. Here probability of A
- 38:13:33given B equals to probability of A when
- 38:13:37we assume that probability of B is non
- 38:13:40zero. Similarly, probability of B equals
- 38:13:43probability of B given A assuming
- 38:13:46probability of A is non zero. Chain rule
- 38:13:50of probability joint probability
- 38:13:52distributions over many random variables
- 38:13:55can be reduced into conditional
- 38:13:57distributions over a single variable. It
- 38:14:00can be expressed as probability of X1 X2
- 38:14:04so on until XN equals probability of X1
- 38:14:08intersection probability of X I given
- 38:14:11probability of X1 till X I minus one.
- 38:14:16For example, the joint probability of A,
- 38:14:19B and C can be given as probability of A
- 38:14:23given B. C multiplied to probability of
- 38:14:27B given C multiply to probability of C.
- 38:14:32Logistic sigmoid.
- 38:14:35The logistics function is a type of
- 38:14:37sigmoid function that aims to predict
- 38:14:39the class to which a particular sample
- 38:14:42belongs. Its outcome is discrete binary
- 38:14:45value. a probability between zero and
- 38:14:48one. The logistic sigmoid is a useful
- 38:14:51function that follows the yes curve. It
- 38:14:54saturates when the input is very large
- 38:14:56or very small. Logistic sigmoid is
- 38:15:00expressed as sigma of x= 1 upon 1 + e to
- 38:15:04the power minus x.
- 38:15:07The logistic sigmoid can be expressed as
- 38:15:10sigmoid function of x is given as 1 upon
- 38:15:131 + e ^ minus x where e is the ooler's
- 38:15:17number.
- 38:15:19Gshian distribution.
- 38:15:22The gossian distribution is a type of
- 38:15:24distribution in which data tends to
- 38:15:26cluster around a central value with
- 38:15:29little or no bias to the left or right.
- 38:15:33It is often referred to as normal
- 38:15:35distribution.
- 38:15:37In absence of prior information, the
- 38:15:40normal distribution is frequently a fair
- 38:15:42assumption in machine learning
- 38:15:45equation.
- 38:15:47The formula for calculating Gaussian
- 38:15:49distribution is described as the normal
- 38:15:52distribution of X.
- 38:15:55That is the function of x given mean as
- 38:15:57mu and variance is sigma square can be
- 38:16:00calculated as 1 upon sigma square
- 38:16:03roo<unk> of 2 pi e to the power -/ x -
- 38:16:07mood / sigma square
- 38:16:11where mu is the mean or peak value which
- 38:16:14also is the expected value of x.
- 38:16:18Sigma is the standard deviation. Sigma
- 38:16:21square is the variance.
- 38:16:23A standard normal distribution has a
- 38:16:26mean of zero and a standard deviation of
- 38:16:28one.
- 38:16:31Goshan distribution can be univariate
- 38:16:35which describes the distribution of a
- 38:16:37single variable X.
- 38:16:39It can also be multivariat where it can
- 38:16:42just use to describe the distribution of
- 38:16:44several variables.
- 38:16:47It is represented in 3D of ND formats.
- 38:16:53Law of large numbers.
- 38:16:57Now let's talk about law of large
- 38:16:59numbers. The law of large numbers states
- 38:17:02that an observed sample average from a
- 38:17:05large sample will be close to the true
- 38:17:07population average and that it will get
- 38:17:10closer in the larger sample. So the law
- 38:17:12of large number does not guarantee that
- 38:17:15a given sample spatially a small sample
- 38:17:17will reflect the true population
- 38:17:19characteristics or that a sample does
- 38:17:22not reflect the true population will be
- 38:17:24balanced by a subsequent sample. This is
- 38:17:27for the law of large numbers to express
- 38:17:30the relationship between scale and
- 38:17:32growth rate.
- 38:17:35So there are multiple examples through
- 38:17:37which we can understand
- 38:17:41and it is widely used in statistical
- 38:17:43analysis in working with the central
- 38:17:45limit theorem in terms of the business
- 38:17:47growth. So there are multiple real time
- 38:17:50setup in which these are going to be
- 38:17:52used. So if you talk about tossing a
- 38:17:55coin so tossing a coin in a number of
- 38:17:58times will give us two different type of
- 38:18:00outcomes.
- 38:18:03the result will spread evenly between
- 38:18:05head and tails and the expected average
- 38:18:08value is going to be half.
- 38:18:10That means 50 times tails and 30 times
- 38:18:13heads. But again, if you toss a coin
- 38:18:161,000 times, then the result can be in
- 38:18:19different manners because out of 1,000,
- 38:18:22let's say 850 times it has been head and
- 38:18:26only 150 times it has been tails and so
- 38:18:30on. So that's why the possibility of one
- 38:18:32event occurring is going to be changed
- 38:18:35in large sample sets as compared to a
- 38:18:37small sample sets as in let's say 10
- 38:18:40times. So the number of heads and tails
- 38:18:43unbalanced for lower number of trials.
- 38:18:45So we can see it is unbalanced.
- 38:18:49But again as soon as we toss more number
- 38:18:51of coins more leans towards the balance
- 38:18:54value or we can see the observed
- 38:18:56averages.
- 38:18:58Next we have p value.
- 38:19:01So p value is basically a number
- 38:19:04calculated from the statistical test
- 38:19:07that describes how likely we are to have
- 38:19:09found a particular set of observations
- 38:19:11if the null hypothesis were true. So p
- 38:19:15values are used in hypothesis testing to
- 38:19:18help decide whether to reject the null
- 38:19:20hypothesis.
- 38:19:22And the smaller the p value, the more
- 38:19:24likely we are to reject the null
- 38:19:26hypothesis.
- 38:19:27So we have a term called as null
- 38:19:30hypothesis. So all statistical tests
- 38:19:33they have null hypothesis. So for most
- 38:19:36tests the null hypothesis is that there
- 38:19:38is no relationship between our variables
- 38:19:41of in first or that there is no
- 38:19:43difference among groups. For example in
- 38:19:46a two-taile t test the non-hypothesis is
- 38:19:49that the difference between two groups
- 38:19:51is going to be zero.
- 38:19:54So p value is going to tell us how
- 38:19:56likely it is that our data could have
- 38:19:58occurred under the null hypothesis.
- 38:20:02It is done by calculating the likelihood
- 38:20:04of a test statistic
- 38:20:06which is the number calculated by a
- 38:20:08statistical test using our data. So p
- 38:20:12value tell us how often we would expect
- 38:20:14to see a test statistic as extreme or
- 38:20:17more extreme
- 38:20:18than one calculated by a statistical
- 38:20:21test. if the null hypothesis of the test
- 38:20:24was true.
- 38:20:26So there are multiple limitations as
- 38:20:28well. So first one is the results can be
- 38:20:31significant but again they are they may
- 38:20:34not be practical as we have compared it
- 38:20:37can be based on multiple hypothesis for
- 38:20:39a game for the healthcare test. If the
- 38:20:42test is going to be positive or not it
- 38:20:45may show even values of the effect of a
- 38:20:47variable but not the magnitude in real
- 38:20:50life. What exactly is going to be the
- 38:20:52application of a drug test being failed
- 38:20:55in pharma company? Therefore, it is
- 38:20:58recommended to use confidence and levels
- 38:21:00in addition to the p values to quantify
- 38:21:03or we can say to give a solid figure to
- 38:21:05the reserve which we are going to get.
- 38:21:08The p values they are interpreted as
- 38:21:11supporting or we can say refuting the
- 38:21:13alternative hypothesis.
- 38:21:15So p value can only tell you whether or
- 38:21:18not the null hypothesis is supported. It
- 38:21:21cannot tell us whether our alternative
- 38:21:23hypothesis is true or why. So the risk
- 38:21:27of rejecting the null hypothesis is
- 38:21:30often higher than the p value. So
- 38:21:33especially when we are looking at a
- 38:21:34single study or when using small sample
- 38:21:37sizes. So this is because the smaller
- 38:21:40frame of reference, the greater are the
- 38:21:42chance that as we stumble across a
- 38:21:45statistically significant pattern
- 38:21:47completely by accident.
- 38:21:49Key takeaways.
- 38:21:51Key takeaways. Probability and
- 38:21:54statistics structure the premise of the
- 38:21:56data. The data helps in anticipating the
- 38:22:00future or gauging in view of the past
- 38:22:03patterns of information.
- 38:22:05The central tendency is a single value
- 38:22:08that helps to describe the data by
- 38:22:10identifying these central positions. The
- 38:22:12mean, median and mode are the measures
- 38:22:15of central tendencies.
- 38:22:18The distribution where the data tends to
- 38:22:20be around a central value with a lack of
- 38:22:23bias or minimal bias towards the left or
- 38:22:26right is called as gshian distribution.
- 38:22:30So now let's dive into the definition of
- 38:22:32the probability distribution function.
- 38:22:36What is probability distribution
- 38:22:38function? A function which defines the
- 38:22:41relationship between a random variable
- 38:22:43and its probability such that you can
- 38:22:45find the probability of the variable
- 38:22:47using the function is called a
- 38:22:49probability density function.
- 38:22:52In simple words, probability density is
- 38:22:55the relationship between an observation
- 38:22:58and the probability. Some outcomes of a
- 38:23:00random variable will have low
- 38:23:02probability density and other outcomes
- 38:23:04will have a very high probability
- 38:23:06density. Basically, the probability of a
- 38:23:09variable X happening or occurring will
- 38:23:12vary and it can sometimes take on a
- 38:23:15lower value or it can take on a way
- 38:23:16higher value.
- 38:23:18The overall shape of the probability
- 38:23:20density is referred to as probability
- 38:23:22distribution. And the calculation of
- 38:23:24probabilities for specific outcomes of a
- 38:23:27random variable is performed by a
- 38:23:29probability density function or PDF for
- 38:23:32short. Now consider a variable X with a
- 38:23:35PDF of f ofx.
- 38:23:39This is what your probability density
- 38:23:41function will look like. There might be
- 38:23:42a point where the probability of X
- 38:23:45occurring is very high. Hence your
- 38:23:49probability distribution function or f
- 38:23:51ofx will also be very high. At other
- 38:23:54points the distribution or the
- 38:23:55probability of X happening or occurring
- 38:23:58is going to be very low. Hence your f
- 38:24:00ofx is also going to have a very small
- 38:24:03value. Basically given the random sample
- 38:24:06of a variable we might want to know
- 38:24:08things like the shape of the probability
- 38:24:10distribution. This here is something
- 38:24:12called a normal distribution where a
- 38:24:14probability distribution function takes
- 38:24:16on a bell shape.
- 38:24:19However, this is not the probability
- 38:24:22density function that might always
- 38:24:23occur. There are different probability
- 38:24:25distribution functions and all of their
- 38:24:27graphs look very different from each
- 38:24:29other. Knowing the probability
- 38:24:31distribution for a random variable can
- 38:24:33help you calculate movements of the
- 38:24:35distribution like the mean and variance.
- 38:24:37But it can also be useful for other more
- 38:24:40general considerations like determining
- 38:24:42whether an observation is unlikely or
- 38:24:44very unlikely and might be an outlier or
- 38:24:47an anomaly like consider this graph
- 38:24:50itself. In this graph, these points over
- 38:24:54here which have very less probability
- 38:24:57distribution
- 38:24:58are outliers which means that the chance
- 38:25:01of them occurring is very low. And
- 38:25:04basically this is not something that
- 38:25:06you're going to see in your regular
- 38:25:08scenario for your variable X. Now let's
- 38:25:11consider two points A and B which are
- 38:25:13values that a variable X can take. P of
- 38:25:17A and P of B just represent the
- 38:25:19probability of A and the probability of
- 38:25:22B which can be found out by drawing a
- 38:25:25straight line and coinciding it with our
- 38:25:27graphs. The area under the graph over
- 38:25:30here which is going to give you your
- 38:25:33probability of this region occurring can
- 38:25:36be written as probability of A less than
- 38:25:40equal to X which is a probability that
- 38:25:42we're searching for here less than equal
- 38:25:44to probability of B. What does this mean
- 38:25:47exactly? This means that this area is
- 38:25:51always going to be greater than or equal
- 38:25:53to the probability of A but less than or
- 38:25:56equal to the probability of B. This
- 38:25:58gives us the narrow region
- 38:26:01of the probability which is present over
- 38:26:03here. And doing this we can find the
- 38:26:06probability of occurrence for any value
- 38:26:09of X. Suppose you want to find the
- 38:26:12probability of B happening. For a
- 38:26:15probability distribution function, the
- 38:26:18probability of B happening is not simply
- 38:26:21this point here, but the entire area of
- 38:26:24the graph which is taking place before
- 38:26:27this point itself. So if you want to
- 38:26:30find the probability between these
- 38:26:32regions, you're going to have to find
- 38:26:33the entire area and not simply the
- 38:26:36probability at one point.
- 38:26:39Now, so far we've been talking about
- 38:26:41different types of variables which is
- 38:26:42discrete random variables and continuous
- 38:26:44variables. What exactly do these mean? A
- 38:26:48variable which can only take a value
- 38:26:50within a certain range is called a
- 38:26:52discrete random variable. The value is
- 38:26:55usually within a certain distance of
- 38:26:57another finite value. An example of this
- 38:27:00would be the sum of two dices. Basically
- 38:27:03values which are well defined are called
- 38:27:05discrete values or and a variable which
- 38:27:08has well- definfined values will be
- 38:27:10called a discrete random variable.
- 38:27:13This variable can only take values which
- 38:27:15fall within a certain set of values.
- 38:27:19Let's say you roll a dice. The dice can
- 38:27:21only give you specific outcomes which
- 38:27:23range from 1 to six. This is what you
- 38:27:26would call a discrete output.
- 38:27:29On the other hand, a continuous random
- 38:27:31variable can take on infinite different
- 38:27:33values within a range of values. For
- 38:27:36example, the height of a student. The
- 38:27:38height of a student is not fixed. Even
- 38:27:40if the height is 1.7 m, in reality, the
- 38:27:46height can be 1.77 or 1.765
- 38:27:50or 1.789.
- 38:27:52The exact height is very hard to
- 38:27:54determine because it's not easy for us
- 38:27:56to find the precise value of the height
- 38:28:00of a student. So basically the height
- 38:28:02can take on an infinite different range
- 38:28:04of values. When we're trying to define
- 38:28:06the values that a continuous random
- 38:28:08variable can take, we usually say it in
- 38:28:11the form of a range of values which
- 38:28:13means that the value can fall in that
- 38:28:15range and can take on any value in that
- 38:28:18range. It's not like a discrete random
- 38:28:20variable where you can define definitive
- 38:28:23values.
- 38:28:26Now let's understand a probability
- 38:28:27density function with the help of a
- 38:28:29graph. Consider the graph below which
- 38:28:31shows the rainfall distribution in an
- 38:28:33year in a city. The x-axis has the
- 38:28:36rainfall in inches or the amount of rain
- 38:28:38that we're getting and the y-axis has
- 38:28:40the probability density function of
- 38:28:42getting that amount of rain.
- 38:28:45The probability of some amount of
- 38:28:47rainfall is obtained by finding the area
- 38:28:50of the curve to the left of it. So let's
- 38:28:53say we have a 3.
- 38:28:56If you want to find the probability of 3
- 38:28:59in of rainfall occurring, we would have
- 38:29:01to find the area of the curve
- 38:29:05which falls to the left of three. When
- 38:29:08we draw a line from three which
- 38:29:10intercepts the graph and further extend
- 38:29:12it onto the yaxis, we get a value of
- 38:29:160.5.
- 38:29:18Simply put, this means that the
- 38:29:20probability of 3 in of rainfall
- 38:29:23occurring is going to be lesser than or
- 38:29:25equal to 0.5. The exact probability can
- 38:29:29be found out by finding the area of the
- 38:29:31curve
- 38:29:34which falls to the left of three.
- 38:29:38How do we find the probability
- 38:29:40distribution function?
- 38:29:43The first step is to summarize your
- 38:29:45density with the help of a histogram.
- 38:29:47The first step in a density estimation
- 38:29:49is to create a histogram of the
- 38:29:51observations
- 38:29:53in the random sample.
- 38:29:56Now what is a histogram? A histogram is
- 38:29:58a plot which involves first grouping the
- 38:30:00observation into bins and counting the
- 38:30:03number of events that fall in each bin.
- 38:30:06The counts or frequency of observation
- 38:30:08in each bins are then plotted as a bar
- 38:30:10graph with the bins on the x-axis and
- 38:30:13the frequency on the y-axis. The choice
- 38:30:16of the number of bins is important as it
- 38:30:18controls the coarseness of the
- 38:30:19distribution and in turn how well the
- 38:30:22density of the observation is plotted.
- 38:30:24It is a good idea to experiment with
- 38:30:26different bin sizes for a given data
- 38:30:28sample to get multiple perspectives or
- 38:30:31views on the same data.
- 38:30:35At the same time, the number of bins is
- 38:30:37important as it determines how many bars
- 38:30:40the histogram will have and their
- 38:30:41widths. This will change not only the
- 38:30:44shape of the graph but also how the
- 38:30:46graph is read. This will also determine
- 38:30:49how our density is plotted. Now let's
- 38:30:51see how we can summarize our density
- 38:30:53with histograms using Python. First
- 38:30:56let's import all of our necessary
- 38:30:58modules which we're going to require.
- 38:31:00We're going to require Mattplot lib to
- 38:31:02plot graphs. We're going to need the
- 38:31:05normal random function so that we can
- 38:31:07get a normal distribution. We're going
- 38:31:09to import mean and standard deviation
- 38:31:12from numpy to use on our graphs and also
- 38:31:15going to normalize our uh data. So we're
- 38:31:18going to import the nom function from
- 38:31:20sci.
- 38:31:25We finished importing all of our
- 38:31:27necessary modules. Now let's generate a
- 38:31:30sample
- 38:31:32which has a size of thousand and it's
- 38:31:34going to be a normal distribution. And
- 38:31:36we're going to also plot this with the
- 38:31:38help of a histogram in bins of 10.
- 38:31:44So as you can see here you get a normal
- 38:31:46distribution which is nothing but a
- 38:31:48almost bell-shaped curve and we have 10
- 38:31:52bins here which are centered at zero and
- 38:31:54which extend from minus3 to 3.
- 38:31:59How will our graph look if we change the
- 38:32:02number of bins though?
- 38:32:04Let's run it and see. So you still have
- 38:32:07a normal distribution but it's not as
- 38:32:09well defined because of how less the
- 38:32:11number of bins are. you lose a majority
- 38:32:13of the data which will contribute to
- 38:32:15your normal distribution. It doesn't
- 38:32:18look like a proper normal distribution
- 38:32:19but it looks more like a discrete data
- 38:32:21at this point. Now let's take a look at
- 38:32:24the next step of finding a probability
- 38:32:27distribution function.
- 38:32:29The next step is called parametric
- 38:32:31density estimation. What exactly is
- 38:32:34parametric density estimation? The
- 38:32:37probability density function is of many
- 38:32:39types. The shape of your histogram will
- 38:32:42help you determine what type of a
- 38:32:43function it is. We can also calculate
- 38:32:46the parameters associated with the
- 38:32:48function to get our density. Now
- 38:32:51different probability distribution
- 38:32:54functions will have
- 38:32:57different graphs which will have
- 38:33:00different shapes and which will also
- 38:33:01have different parameters like mean,
- 38:33:03standard deviation etc associated with
- 38:33:06them. Using these parameters, we can
- 38:33:09find important points of our data.
- 38:33:13Hence, it's very important for us to
- 38:33:15recognize what type of a distribution it
- 38:33:17is. Common distributions will occur
- 38:33:20again and again in different and
- 38:33:21sometimes unexpected domains.
- 38:33:24Getting familiar with common probability
- 38:33:26distributions will help you identify a
- 38:33:29distribution from a histogram. And once
- 38:33:31identified, you can attempt to estimate
- 38:33:33the density of the random variable with
- 38:33:36a chosen probability distribution. This
- 38:33:38can be achieved by estimating the
- 38:33:40parameters of the distribution from a
- 38:33:42random sample of data. Now, an example
- 38:33:45of this would be a normal distribution
- 38:33:46which has two main parameters, the mean
- 38:33:49and standard deviation. Given these two
- 38:33:51parameters, we will now know the
- 38:33:53probability distribution function. These
- 38:33:56parameters can be estimated from data by
- 38:33:58calculating the sample mean and sample
- 38:34:01standard deviation. This entire process
- 38:34:03is known as parametric density
- 38:34:05estimation and it includes identifying
- 38:34:08your probability distribution function
- 38:34:11and getting the parameters which are
- 38:34:12associated with it.
- 38:34:16Now once we have estimated the density,
- 38:34:18we can check if it's a good fit.
- 38:34:24This can be done in three different
- 38:34:26ways. One is plotting the density of the
- 38:34:28function and comparing the shape to the
- 38:34:30histogram. The next is sampling the
- 38:34:33density function and comparing the
- 38:34:35generated sample to the real sample. And
- 38:34:38the last one is using a statistical test
- 38:34:40to confirm if the data fits the
- 38:34:42distribution. Now over here as you can
- 38:34:44see all we've done is taken our data and
- 38:34:48plotted the density function on top of
- 38:34:51our histogram and we've compared the
- 38:34:53shape. So the distribution so the
- 38:34:55density function that we're actually
- 38:34:56considering here is a normal
- 38:34:58distribution and from this graph we can
- 38:35:01see that it's almost an exact fit to our
- 38:35:03histogram. Now let's see how we can
- 38:35:06perform parametric density estimation
- 38:35:08using Python. To begin with,
- 38:35:12let's generate a random sample of
- 38:35:14thousand observations from a normal
- 38:35:16distribution with a mean of 50 which is
- 38:35:19determined by the LOC parameter and a
- 38:35:22standard deviation of five which is
- 38:35:24determined by the scale parameter.
- 38:35:29Now just to show you what the
- 38:35:31distribution looks like, we're going to
- 38:35:33plot it in the form of a histogram. So
- 38:35:35this is what the histogram looks like.
- 38:35:38But this is just to give you a basic
- 38:35:40idea of our data and what it looks like
- 38:35:42once plotted. But let's assume that we
- 38:35:45don't know the probability distribution
- 38:35:49and and we don't know what it looks like
- 38:35:51as a histogram and we don't know that
- 38:35:53that it's normal. So now if we just
- 38:35:56assume that it's normal, we can
- 38:35:58calculate the parameters of the
- 38:35:59distribution specifically the mean and
- 38:36:02the standard deviation.
- 38:36:04We would not expect the mean and
- 38:36:05standard deviation to be 50 and five.
- 38:36:07Exactly given the small sample size and
- 38:36:10the noise in the sampling data.
- 38:36:12So because of this noise and the small
- 38:36:15sample size, we have a mean of almost 50
- 38:36:18and a standard deviation of a little
- 38:36:20more than five.
- 38:36:22Now let's define the distribution as
- 38:36:25normal. So now using this we've defined
- 38:36:28a normal distribution. We've used the
- 38:36:30norm method of the sci-fi uh library and
- 38:36:34uh we're doing this with the mean and
- 38:36:37the standard deviation that we've
- 38:36:39obtained from our samples. So up until
- 38:36:42now we're just assuming that it's a
- 38:36:44normal distribution and because of that
- 38:36:46the parameters that we've calculated is
- 38:36:48the mean and standard deviation
- 38:36:51and using the calculated mean and
- 38:36:53standard deviation we've gotten a normal
- 38:36:55distribution.
- 38:36:57And up until now again keep in mind we
- 38:36:59do not know for certain that it is a
- 38:37:02normal distribution. So far all we have
- 38:37:04is this data.
- 38:37:06So the next thing that we're going to do
- 38:37:08is fit the distribution with these
- 38:37:10parameters
- 38:37:12and then sample the probabilities for a
- 38:37:14distribution for a range of values in
- 38:37:17our domain which in this case is 30 and
- 38:37:1970. So all we're doing is we're
- 38:37:22calculating probabilities for a range of
- 38:37:24outcomes. And in this case we've taken
- 38:37:2730 and 70 as our domain.
- 38:37:32So these are the probability
- 38:37:33distribution values for the normal
- 38:37:36distribution that we've defined over
- 38:37:38here. And this is going to uh this is
- 38:37:41basically going to give you the outline
- 38:37:42of your normal distribution.
- 38:37:45These are the points at which your
- 38:37:46normal distribution will be plotted. Uh
- 38:37:48now what we're basically going to do is
- 38:37:50we're going to plot our histograms using
- 38:37:52the samples that we've already generated
- 38:37:55along with the values and probabilities
- 38:37:58of the normal function that we defined
- 38:38:00over here.
- 38:38:05So as you can see it's an all it's
- 38:38:07almost a complete fit. The normal
- 38:38:09distribution that we have here is made
- 38:38:13using the mean and the standard
- 38:38:14deviation of our actual samples.
- 38:38:18The reason we took mean and standard
- 38:38:20deviation was because we assumed it was
- 38:38:22a normal distribution and the parameters
- 38:38:25associated with the normal distribution
- 38:38:28are mean and standard deviation. Using
- 38:38:30the mean and standard deviation, we got
- 38:38:32the normal distribution. We calculated
- 38:38:36probabilities for this normal
- 38:38:38distribution using a random domain of 30
- 38:38:40and 70
- 38:38:42and we plotted the probabilities and the
- 38:38:45values on top of our histogram to see if
- 38:38:48the normal distribution was a fit to our
- 38:38:50histogram. If it was not a fit, you
- 38:38:53would have to go and do the same
- 38:38:55procedure with other common probability
- 38:38:58density functions
- 38:39:01until you found a function which was a
- 38:39:03proper fit to your histogram. Now let's
- 38:39:06move on to the final step which is used
- 38:39:08in the calculation of a PDF. This final
- 38:39:11step is called nonparametric density
- 38:39:13estimation and it's only used when the
- 38:39:16shape of a histogram doesn't match a
- 38:39:18common probability density function or
- 38:39:21it cannot be made to fit one. In this
- 38:39:23case, we will calculate the density
- 38:39:25using all samples in our data using
- 38:39:28certain algorithms.
- 38:39:30This is only done when a data sample
- 38:39:31does not resemble a common probability
- 38:39:34distribution or it cannot be easily made
- 38:39:36to fit the distribution. And this is
- 38:39:38often the case when the data has two
- 38:39:40peaks. This is also called a biodal
- 38:39:43distribution or it has many peaks which
- 38:39:46is also called a multimodal
- 38:39:48distribution. In this case, the
- 38:39:50parametric density estimation will not
- 38:39:52be feasible and alternative methods can
- 38:39:55be used that do not use a common
- 38:39:57distribution. Instead, you will use an
- 38:39:59algorithm which is used to approximate
- 38:40:01the probability distribution of the data
- 38:40:04without a predefined distribution which
- 38:40:06is also referred to as a non-parametric
- 38:40:09method because we're not using any
- 38:40:12predefined parameters. The distribution
- 38:40:14will still have parameters but these are
- 38:40:17not controllable in the same way as a
- 38:40:19simple probability distribution. For
- 38:40:22example, a non-parametric method might
- 38:40:24estimate the density using all
- 38:40:26observations in a random sample in
- 38:40:29effect making all observations in the
- 38:40:31sample parameters.
- 38:40:37Now consider this graph which has two
- 38:40:39peaks. You this is not a normal
- 38:40:42distribution or any other sort of
- 38:40:44distribution that we are familiar with.
- 38:40:46So for this we're not going to use a
- 38:40:48parametric estimation method but we're
- 38:40:52just going to calculate the parameters
- 38:40:54for every single sample point in this.
- 38:40:57Perhaps the most common nonparametric
- 38:41:00approach for estimating the probability
- 38:41:02density function of a continuous random
- 38:41:04variable is called kernel smoothing or
- 38:41:07kernel density estimation or KDE for
- 38:41:09short. Kernal density estimation is a
- 38:41:13nonparametric method for using a data
- 38:41:15set to estimate probabilities for new
- 38:41:17points.
- 38:41:19It uses a mathematical function and
- 38:41:21smoothing probabilities. So the so the
- 38:41:23sum of the resultant probabilities is
- 38:41:26always one. Now in this case a kernel is
- 38:41:28a mathematical function that returns a
- 38:41:31probability for a given value of a
- 38:41:33random variable. The kernel effectively
- 38:41:35smooths or interpolates the
- 38:41:37probabilities across a range of outcomes
- 38:41:40for a random variable such that the sum
- 38:41:42of probabilities always equals one. A
- 38:41:45requirement of well- behaved
- 38:41:47probabilities. You also have a parameter
- 38:41:50called the smoothing parameter which
- 38:41:52controls the scope or the window of
- 38:41:54observations from the data samples that
- 38:41:57contributes to estimating the
- 38:41:59probability for a given sample. As such
- 38:42:02the kernel density estimation is s is
- 38:42:05sometimes referred to as your parsen
- 38:42:07rosenbalt window. Now at the end you
- 38:42:10also have a basis function which is a
- 38:42:12function which is chosen to control the
- 38:42:14contribution of samples in the data set
- 38:42:16towards estimating the probability of a
- 38:42:18new point. This is only done to make
- 38:42:21sure that you're not learning from a lot
- 38:42:23of noise and that you're not using a lot
- 38:42:26of the outliers. Again let's see how we
- 38:42:29can perform non-parametric density
- 38:42:31estimation with the help of Python.
- 38:42:35So first we'll start by importing all
- 38:42:37the necessary modules along with the
- 38:42:39kernel density estimation which can be
- 38:42:42imported from skarn.
- 38:42:47Now let's create a biodial distribution
- 38:42:51by combining two different samples.
- 38:42:53Sample one and sample two. Sample one
- 38:42:56has 300 examples with a mean of 20 and a
- 38:43:00standard deviation of five. While sample
- 38:43:02two has 700 examples with a mean of 40
- 38:43:06and a standard deviation of five.
- 38:43:10We're then going to use it stack to com
- 38:43:13to merge both of them together to get a
- 38:43:15final sample.
- 38:43:18The means that we've chosen which is 20
- 38:43:20and 40 are chosen close together to
- 38:43:23ensure that the distributions overlap in
- 38:43:25the combined sample.
- 38:43:28So this is what our distribution is.
- 38:43:30Let's just plot it so you get a basic
- 38:43:32idea of what our graph looks like.
- 38:43:35So this is what our graph looks like.
- 38:43:39Now we already know that none of the
- 38:43:41various different uh uh probability
- 38:43:44distribution functions fit these graphs.
- 38:43:47So now we're going to perform
- 38:43:48nonparametric estimations.
- 38:43:52To perform nonparametric estimations,
- 38:43:57we're going to use the scikitlearn
- 38:43:59machine learning library which provides
- 38:44:01the kernel density class that implements
- 38:44:04kernel density sorry that implements
- 38:44:07kernel density estimation. First the
- 38:44:10class is constructed with the desired
- 38:44:12bandwidth or window size of two
- 38:44:17and your basis function
- 38:44:20which in this case is a gshian function.
- 38:44:24It's a good idea to at least test
- 38:44:25different configurations to your data.
- 38:44:28And in this case we're only going to try
- 38:44:29a bandwidth of two and a gshian kernel.
- 38:44:32Uh but usually there are multiple
- 38:44:34different kernels that you can uh you
- 38:44:37know like uh that you can play around
- 38:44:39with and you can also tweak your
- 38:44:40bandwidth to exactly fit the
- 38:44:43distribution that you have.
- 38:44:47Now let's run this.
- 38:44:50Uh so now we've gotten our kernel
- 38:44:52density estimation. Uh now we can
- 38:44:54evaluate how well the density estimates
- 38:44:57matches our data by calculating
- 38:45:00probabilities
- 38:45:01for a range of observations and
- 38:45:03comparing shapes to the histogram just
- 38:45:05like we did for the parametric case
- 38:45:07before. So again we're going to just
- 38:45:10calculate different probabilities using
- 38:45:12the kernel density function
- 38:45:15and we're just going to plot it on top
- 38:45:16of a histogram to see how well this the
- 38:45:19kernel density function is estimating
- 38:45:22for our data.
- 38:45:24So these are the probabilities that
- 38:45:25we've gotten finally with the con uh
- 38:45:27with the kernel density estimation. And
- 38:45:30now we're going to plot it on top of our
- 38:45:32histogram.
- 38:45:34So as you can see it's almost a complete
- 38:45:37fit. It's just left out some of these
- 38:45:39outlier values which again are ranging
- 38:45:42very high. But overall we have a pretty
- 38:45:44good fit.
- 38:45:47Uh the only problem is it's not very
- 38:45:49smooth and you can uh again try tweaking
- 38:45:52the bandwidth to different values. Uh so
- 38:45:55let's just in this case try tweaking it
- 38:45:57to three and see how well it runs. Okay.
- 38:46:02So now we've got a new probabilities and
- 38:46:04let's run it on top of our bandwidth. Uh
- 38:46:08so again we using a bandwidth of three.
- 38:46:10You can see that we're fitting our data
- 38:46:12even better and we're again
- 38:46:16ignoring a lot of the outliers which are
- 38:46:18out there. So this is going to give us a
- 38:46:20better estimation. The first question
- 38:46:22that is probably in your mind is what's
- 38:46:24in it for you? What can you expect from
- 38:46:27this video?
- 38:46:29First we will explain the concept of
- 38:46:31regression a machine learning algorithm
- 38:46:34to you.
- 38:46:36Next we will take a look at the R squar
- 38:46:38error which can be used to calculate the
- 38:46:41error in regression models.
- 38:46:43Next
- 38:46:45we will teach you how to calculate the R
- 38:46:47squar error and finally we will
- 38:46:50implement the R squared error with the
- 38:46:52help of Python.
- 38:46:54So what is regression?
- 38:46:58Regression is nothing but a machine
- 38:47:00learning algorithm that helps us
- 38:47:02determine the relationship between two
- 38:47:04or more variables. It uses input or
- 38:47:07independent variables to find the value
- 38:47:09of the output or dependent variables.
- 38:47:13Regression is a prediction algorithm
- 38:47:16which means that given some variables we
- 38:47:19can predict the value of an output
- 38:47:21variable.
- 38:47:23The predicted value is not going to be
- 38:47:25from a set of values but it's going to
- 38:47:27be a unique value in itself.
- 38:47:31Now let's understand what exactly
- 38:47:33regression is with the help of a few
- 38:47:35independent input variables. In this
- 38:47:38case, the variables that we'll be
- 38:47:40looking at is rainwater, fertilizer, and
- 38:47:43seeds. When we pass these independent
- 38:47:46input variables through regression
- 38:47:48model, we're going to get a predicted
- 38:47:50output.
- 38:47:52The output predicted is that a crop will
- 38:47:55germinate when all three of these
- 38:47:56components are put together in certain
- 38:47:59quantities. So with the help of
- 38:48:01regression given raw input data we can
- 38:48:04find out the dependent output variable
- 38:48:07that we'll get
- 38:48:09when all of these input variables are
- 38:48:12more or less combined. Regression is
- 38:48:14nothing but a statistical method which
- 38:48:17is used to determine the strength and
- 38:48:19character of the relationship between a
- 38:48:21dependent variable and a series of other
- 38:48:24variables.
- 38:48:26Now using regression if you have a set
- 38:48:28of data points we can use a regression
- 38:48:31model to fit a line which passes through
- 38:48:33most of these data points and use it to
- 38:48:35predict the outcome for new data points.
- 38:48:38The line is fit using equation of a
- 38:48:40straight line or a polomial equation.
- 38:48:43Now in this graph consider that we have
- 38:48:46our dependent variable or the output y
- 38:48:48and our independent variable or the
- 38:48:50output x. for a certain value of our
- 38:48:54independent variable. We are going to
- 38:48:56get a certain value of our dependent
- 38:48:58variable or our output. Using the data
- 38:49:01given here, we can see how our output
- 38:49:03varies when our input varies. To predict
- 38:49:06the value of our outcome Y, we're going
- 38:49:09to need to find a relationship between
- 38:49:11all of these data points. To do this,
- 38:49:14we're going to plot a straight line
- 38:49:15through it. Because, as you can see, all
- 38:49:17the data points lie more or less along a
- 38:49:20given straight line. Now using the
- 38:49:22straight line for any value of x we can
- 38:49:25find the approximate value of y. So
- 38:49:29suppose you want to find the value of y
- 38:49:32at a point x which is given here say
- 38:49:35then all you have to do is extend a line
- 38:49:37from this point onto a predicted model
- 38:49:40which is this line here and then we can
- 38:49:43see where this point on the line
- 38:49:46coincides with the y-axis and get the
- 38:49:48approximate value of the outcome.
- 38:49:52The equation of a straight line is given
- 38:49:55as y = b + b1 x + e. Where b is a
- 38:50:01constant given by the y intercept of our
- 38:50:04line or basically where the line
- 38:50:06intersects on the y-axis.
- 38:50:11B1 is the slope of our line and x is the
- 38:50:14point for which we want to find the
- 38:50:16output. E in this case is nothing but an
- 38:50:19error correcting term.
- 38:50:22So this is basically how regression
- 38:50:24works and this is how prediction takes
- 38:50:27place in a regression model. Next we
- 38:50:30will explain the concept of the R squar
- 38:50:32error to you. So what is R squar error?
- 38:50:36R squared error is nothing but an error
- 38:50:39measurement term which calculates how
- 38:50:41well a regression model fits the data.
- 38:50:44It determines the amount of variance in
- 38:50:46a model caused by the input variables.
- 38:50:49Now, R squar is a statistical measure
- 38:50:52that represents the portion of variance
- 38:50:54for a dependent variable that's
- 38:50:56explained by an independent variable or
- 38:50:58variables in a regression model. R 2 is
- 38:51:01used to explain to what extent the
- 38:51:04variance of one variable affects the
- 38:51:06variance of a second variable.
- 38:51:09So if the R square of a model is 0.5
- 38:51:12then approximately half of the observed
- 38:51:14variation can be explained by the
- 38:51:16model's inputs.
- 38:51:20In other words, an R squar of 60%
- 38:51:23reveals that 60% of our data fits our
- 38:51:26regression model exactly.
- 38:51:30Now in this case in this graph if we
- 38:51:34have a variance of 60% it means that 60%
- 38:51:38of our data points fall exactly on our
- 38:51:41regression line. However it is not
- 38:51:44always the case that a high R squar is
- 38:51:46good for a regression model. The quality
- 38:51:49of the statistical measure depends on
- 38:51:51many factors such as the nature of the
- 38:51:53variables employed in the model, the
- 38:51:56unit of measure of variables and the
- 38:51:58applied data transformation.
- 38:52:00Thus, sometimes a high R squar can
- 38:52:03indicate the problems with the
- 38:52:05regression model. A low R squar figure
- 38:52:08is generally a bad sign for predictive
- 38:52:10models. Now, all this time we've been
- 38:52:13talking about variance in our model.
- 38:52:15What exactly is variance? Variance is
- 38:52:18nothing but a statistical term which
- 38:52:20determines how spread out our data is
- 38:52:22and tells us how many outliers are
- 38:52:25present in it. Basically, it's a measure
- 38:52:27of how far a set of numbers is spread
- 38:52:30out from their average value.
- 38:52:33Using variance, you can basically figure
- 38:52:35out where your data is centered
- 38:52:38and how spread it is from the mean and
- 38:52:41also you can find out how many outliers
- 38:52:44it has.
- 38:52:46Now how can you calculate the R squar
- 38:52:49error?
- 38:52:52Let's start by considering the data that
- 38:52:55we have been given as shown below. Now
- 38:52:57we can find the relationship between the
- 38:52:59input and the output variables by
- 38:53:01plotting a straight line or a regression
- 38:53:03model that passes through most of the
- 38:53:06data. To get the perfect fit for a
- 38:53:08model, we don't necessarily need to have
- 38:53:11the line passing through as many data
- 38:53:13points as possible.
- 38:53:15A true measure of a good model is that
- 38:53:17we reduce the error which is present in
- 38:53:19our model. Now how do you find this
- 38:53:22error? The error present in our model is
- 38:53:25given by nothing but the distance
- 38:53:27between our predicted line and the data
- 38:53:29points which do not fall on the line.
- 38:53:32This is what we have to minimize.
- 38:53:35The distance between our data points and
- 38:53:37our line can be calculated by
- 38:53:39subtracting
- 38:53:41our data point from the point at which
- 38:53:44it coincides on our regression line. We
- 38:53:47square this just to get rid of any
- 38:53:49negative coefficients that may occur due
- 38:53:51to finding the difference between the
- 38:53:52two points. Now to find the variance in
- 38:53:55our data, we're going to find the mean
- 38:53:57and subtract the data points from the
- 38:53:59mean. This will basically tell us how
- 38:54:02spread out our data point is from the
- 38:54:04average value. We can then square these
- 38:54:08differences and add up the result to get
- 38:54:10our total variance. This is also known
- 38:54:12as the sum of squares total. And using
- 38:54:14this we can find the total variance in
- 38:54:16our data. The mean is nothing but the
- 38:54:19average of our data. And using the mean
- 38:54:22we can find the center of our data.
- 38:54:26So this is exactly where our data is
- 38:54:28centered. This is the average value that
- 38:54:30occurs in our data. Now, the variance is
- 38:54:33nothing but the distance of all of our
- 38:54:35data points from the mean. If we do
- 38:54:39this, we're basically going to find out
- 38:54:40how spread apart our data is from the
- 38:54:43average value or how far all of our data
- 38:54:46points lie from each other.
- 38:54:48When we subtract the position of our
- 38:54:50data points from our mean and square it
- 38:54:53and add all of that up, we get something
- 38:54:55known as the sum of squares total.
- 38:54:59Now the R squ error is the total
- 38:55:01variance in our input data. It can be
- 38:55:03obtained by dividing the SSR by the SST
- 38:55:06and subtracting the results from one.
- 38:55:10So the R squar error totally becomes 1
- 38:55:13minus the sum of our squared errors
- 38:55:16divided by the sum of the squared
- 38:55:19difference between our data points and
- 38:55:22the mean. Now this value gives us the
- 38:55:25variance and this is why we can say that
- 38:55:29R squ is used to find the portion of
- 38:55:32variation in our data.
- 38:55:35Now how can you implement R squ error
- 38:55:37with Python?
- 38:55:39To calculate R squared error with
- 38:55:41Python, we're going to look at the data
- 38:55:44which depicts the weather conditions
- 38:55:46which were present during World War II.
- 38:55:48And using the variables which are
- 38:55:50present in our data, we're going to
- 38:55:51create a model which predicts the daily
- 38:55:54weather forecast during World War II.
- 38:55:57And then we're going to use R squared
- 38:55:58error to find the accuracy of our model.
- 38:56:02So this is our R square uh model. So
- 38:56:05we're going to start off by importing
- 38:56:07all of our necessary modules. We're
- 38:56:09going to use the model numpy to perform
- 38:56:12numerical calculations on our database
- 38:56:15and arrays and we're going to use
- 38:56:17seaborn and mattplot lib to plot our
- 38:56:20data.
- 38:56:22So now we've managed to import all of
- 38:56:24our data sets.
- 38:56:27Let's also load our data by reading in
- 38:56:30the CSV file that it is stored as in the
- 38:56:33form of a data frame.
- 38:56:36So over here as you can see we've read
- 38:56:38in the CSV file into uh a variable
- 38:56:42called weather and then after that we're
- 38:56:44changing weather into a panda's data
- 38:56:46frame called climate. This is what our
- 38:56:48data frame finally looks like.
- 38:56:53So as you can see in our data frame we
- 38:56:55have five rows because we're only
- 38:56:58looking at the top five rows and we have
- 38:57:0031 columns. So these are values which
- 38:57:03are not required and which are basically
- 38:57:05going to increase our error value. So
- 38:57:08let's drop them and get rid of them.
- 38:57:13Now let's also drop any
- 38:57:17empty values which may occur in the
- 38:57:19remaining columns of our data set and
- 38:57:21see what the final data set looks like.
- 38:57:24So this is our final data set. We have
- 38:57:27at the end we're only left with max
- 38:57:29temperature, minimum temperature, and
- 38:57:30mean temperature.
- 38:57:33Now let's plot a count plot of a max
- 38:57:36temperature.
- 38:57:41A count plot is basically going to go
- 38:57:44through the entire max temperature
- 38:57:46column and figure out how many times
- 38:57:49every temperature value occurs. So it's
- 38:57:53going to figure out how many it's going
- 38:57:54to count how many times 29.44444
- 38:57:58has occurred and it's going to plot that
- 38:58:00on this graph and it is going to do that
- 38:58:03for every unique temperature value which
- 38:58:05is present in our column. So finally
- 38:58:08this is the value that we get. Uh so
- 38:58:11over here as you can see the majority of
- 38:58:14our temperature values are concentrated
- 38:58:16within this range. This means that these
- 38:58:19temperature values are the ones which
- 38:58:22occur most frequently. The other ones
- 38:58:24can be considered as outliers because
- 38:58:27they rarely are seen in our data
- 38:58:31and they can further skew the output
- 38:58:33that we're going to get. Now let's do
- 38:58:35the same with minimum temperature. Let's
- 38:58:38plot a count plot for minimum
- 38:58:40temperature.
- 38:58:42So for the minimum temperature we can
- 38:58:44see a very similar plot to the one that
- 38:58:46we got for a maximum temperature. Most
- 38:58:49of the values are concentrated around
- 38:58:51this region but the outlier values here
- 38:58:54are way fewer.
- 38:58:59Now let's plot a regression plot between
- 38:59:01our maximum and our minimum temperature.
- 38:59:04So using the regression plot we can plot
- 38:59:07a regression line for the two variables
- 38:59:10in our x and y axis. So this is our
- 38:59:13x-axis and this is our y-axis. This is
- 38:59:15basically going to plot a straight line
- 38:59:19which best fits the data
- 38:59:24that we are getting here. So over here
- 38:59:27as you can see this is a regression
- 38:59:28line. This thin blue line is a
- 38:59:30regression line which intersects our
- 38:59:33minimum temperature at a value which is
- 38:59:37between -30 and -40 and it passes
- 38:59:40through our entire data. So uh for a
- 38:59:44value of maximum temperature which is
- 38:59:47zero using this we can predict the
- 38:59:50minimum temperature that would have
- 38:59:53occurred on the same day.
- 38:59:56So for zero it'll be somewhere around -
- 38:59:5910°C. So if we saw maximum temperature
- 39:00:02of 0° on that day, we would have seen a
- 39:00:04minimum temperature of minus 10 on the
- 39:00:06same day. Now let's plot a heat map to
- 39:00:09see how these values are correlated with
- 39:00:12each other. So the correlation is
- 39:00:14basically
- 39:00:15used to find which values affect each
- 39:00:19other linearly
- 39:00:22or which values
- 39:00:25when changed will also affect the change
- 39:00:27in other values. So over here as you can
- 39:00:30see minimum temperature
- 39:00:39and maximum temperature have a
- 39:00:41correlation of 0.8. 88 which means if
- 39:00:44minimum temperature changes then the
- 39:00:47maximum temperature will also change to
- 39:00:510.88.
- 39:00:55Now the best correlation is obviously
- 39:00:57going to be between the mean
- 39:00:58temperatures and the minimum and maximum
- 39:01:00temperatures. The mean temperature is
- 39:01:02nothing but the average temperature
- 39:01:04value that we have. So this this is
- 39:01:07basically going to lie in the middle of
- 39:01:10all of our temperature values which is
- 39:01:12why we going to have a better
- 39:01:14correlation for mean temperature. But
- 39:01:16minimum temperature and maximum
- 39:01:17temperature are also pretty well
- 39:01:19correlated with a correlation value of
- 39:01:210.88.
- 39:01:22This means that if our minimum
- 39:01:25temperature fluctuates, our maximum
- 39:01:28temperature will also fluctuate
- 39:01:29proportionately.
- 39:01:31Now let's separate our input and output
- 39:01:33values. We're going to predict the
- 39:01:36maximum temperature given our minimum
- 39:01:38temperature. Here x is our input
- 39:01:40variable and y is our output variable.
- 39:01:46So now we're basically just going to get
- 39:01:48all the important values in our x and y
- 39:01:51data sets. So after that this is what
- 39:01:53our x and y data sets are going to look
- 39:01:55like. Now let's split our data set into
- 39:01:58training and testing sets.
- 39:02:01The training set will be used to train
- 39:02:03our regression model and the testing set
- 39:02:06will be used to predict how well our
- 39:02:08regression model is performing.
- 39:02:11The training data is the data which will
- 39:02:14be visible to our model or the data
- 39:02:17which a model is allowed to have access
- 39:02:19to. Testing data will be data which the
- 39:02:21model has never seen before or which it
- 39:02:24doesn't have access to and hence it will
- 39:02:27be made to work on completely new data
- 39:02:29to better test how well we've fitted to
- 39:02:32our data set. Now we can split our data
- 39:02:35set into training and testing sets by
- 39:02:36using the train test split functionality
- 39:02:39from our skarn.mmodel selection library.
- 39:02:44Now finally from our scikitlearn library
- 39:02:47let's import a linear regression model.
- 39:02:51We're going to initialize a linear
- 39:02:53regression model to a variable called
- 39:02:55regressor and then we're going to fit a
- 39:02:57linear regression model to our training
- 39:03:00data set.
- 39:03:02So we finally got our trained linear
- 39:03:05regression model. Now let's use this
- 39:03:07model to perform predictions on our
- 39:03:09testing data set. So these are the
- 39:03:12values that we've gotten after running
- 39:03:14our linear regression model on our
- 39:03:17testing data set. Let's see how well
- 39:03:20we've performed.
- 39:03:22We are going to import the R2 score from
- 39:03:25our skarn metric.
- 39:03:28Now the R2 score will directly perform R
- 39:03:31squared error on our prediction and
- 39:03:34testing data set and see how well our
- 39:03:36testing data set matches a prediction
- 39:03:39data set.
- 39:03:42So now we've gotten an R squared error
- 39:03:44of 0.9345
- 39:03:46which basically means that 93% of our
- 39:03:53output values are influenced by our
- 39:03:55input values. This also means that a
- 39:03:58model is 93% accurate
- 39:04:02and that approximately
- 39:04:0493% of our observed variation can be
- 39:04:08explained by the model's inputs. Ever
- 39:04:10wondered how to build an AI project that
- 39:04:13actually gets noticed by Google, OpenAI,
- 39:04:15or top startups, not just a chartboard
- 39:04:18or recycled homework. Today I'm going to
- 39:04:21walk you through 10 AI project ideas for
- 39:04:2426 that are practical, futuristic and
- 39:04:28portfolio ready. I'll tell you exactly
- 39:04:30which models, framework and data sets to
- 39:04:33you so that you can start coding
- 39:04:35immediately. Now before we jump into hit
- 39:04:38that like button, share and subscribe
- 39:04:40because keeping up with future proof AI
- 39:04:43projects is going to give you a massive
- 39:04:45edge. Let's start with the AI shopping
- 39:04:48buddy. This project acts like a personal
- 39:04:51stylist and interior designer. Users
- 39:04:54upload photos of the room, outfit, or
- 39:04:56even face, and the AI suggest products
- 39:04:59that match color, style, and
- 39:05:01preferences. This isn't just about
- 39:05:04throwing recommendations at someone.
- 39:05:06It's about computer vision to understand
- 39:05:08images, generative AI to create style
- 39:05:11suggestions, and recommendation
- 39:05:13algorithms to find the perfect products.
- 39:05:15Personalized recommendation systems
- 39:05:17drive massive engagement and conversions
- 39:05:20which is why companies like Amazon,
- 39:05:22Flipkart, Myntra or Urban Ladder would
- 39:05:26be thrilled to hit someone who can build
- 39:05:28this. Completing a project like this
- 39:05:30demonstrates skills in deep learning,
- 39:05:33computer vision, generative AI and full
- 39:05:35stack deployment for web or mobile app.
- 39:05:38While shopping and lifestyle AI is
- 39:05:40exciting, the next project takes up to
- 39:05:43our health and wellness. The smart
- 39:05:46health analyzer predicts stress burnout
- 39:05:48or sleep issues by analyzing voice,
- 39:05:51facial microp expressions and variable
- 39:05:54data. It uses multimodel AI that
- 39:05:57integrates time series analysis for
- 39:05:59variable data. NLP for voice and text
- 39:06:02and computer vision for micro
- 39:06:04expressions. Health tech startups in
- 39:06:06India and globally like healthy, cure
- 39:06:08fit, Fitbit and Apple Health are looking
- 39:06:11for engineers who can make predictive
- 39:06:14wellness tools. Building this project
- 39:06:17demonstrates your ability to work with
- 39:06:19multimodel AI, pre-process complex data
- 39:06:22sets, train models, and visualize
- 39:06:24result. Moving from personal health to
- 39:06:27professional efficiency, the AI
- 39:06:29productivity agent automates your daily
- 39:06:32workflow. It reads emails, scans your
- 39:06:34calendar, understands priorities, and
- 39:06:37builds an optimized schedule. It uses
- 39:06:39NLP to parse emails, API integration
- 39:06:43with other tools like Gmail, Slack, and
- 39:06:45Notion, and optimization algorithms to
- 39:06:48prioritize task efficiently.
- 39:06:50Productivity loss is a major issue for
- 39:06:52companies which is why tech giants like
- 39:06:55Google Workspace, Microsoft 365 and
- 39:06:58startups in workflow automation would be
- 39:07:00very interested in this project. It is a
- 39:07:03great way to demonstrate automation, NLP
- 39:07:06API integration and practical problem
- 39:07:08solving skills. Taking automation to the
- 39:07:11next level, the voice toaction system
- 39:07:13allows users to speak commands and have
- 39:07:16the AI perform multi-step action such as
- 39:07:19booking flights, organizing files or
- 39:07:21generating reports. It relies on
- 39:07:24speechtoext models, intent
- 39:07:26classification using NLP and task
- 39:07:29automation pipelines. You can train
- 39:07:32intent classification models using data
- 39:07:34set into the snips NLU data set. This
- 39:07:38project builds directly on productivity
- 39:07:40AI and is exactly the kind of work that
- 39:07:43Amazon openai or Apple would notice for
- 39:07:46voiced driven automation solutions. Once
- 39:07:49we have automated task, why not explore
- 39:07:52creativity? Generative AI story maker
- 39:07:54allows you to create full stories
- 39:07:56including scripts, characters and
- 39:07:59visuals based on just a few keywords.
- 39:08:02Now it uses large language models for
- 39:08:05text generation and image generation
- 39:08:07models like stable diffusion deli3 for
- 39:08:10visuals and you can also train
- 39:08:12fine-tuned models on data sets like CMU
- 39:08:15book summary corpus on writing prompts
- 39:08:18text to speech library such as scope TTS
- 39:08:21or GTTS can add narration media
- 39:08:25companies like Netflix, Ubisoft and
- 39:08:27Adobe are actively investing in
- 39:08:30generative AI and a project like this
- 39:08:32would definitely stand out. Building on
- 39:08:34the idea of multiple AI capabilities
- 39:08:37working together, the multi- aent AI
- 39:08:40team project introduces collaboration
- 39:08:42between AI agents. Multiple AI agents
- 39:08:45are assigned specialized roles such as
- 39:08:48researching, writing, criticing, and
- 39:08:50summarizing. They communicate and
- 39:08:53coordinate to complete complex task
- 39:08:55using multi- aent reinforcement,
- 39:08:57learning and communication protocols.
- 39:09:00Enterprise AI and automation platforms
- 39:09:03are investing heavily in this approach
- 39:09:05and companies like Enthropic, OpenAI, AI
- 39:09:08workflow startups are actively seeking
- 39:09:11engineers who can build collaborative AI
- 39:09:14systems from collaboration to
- 39:09:16observation. The AI body language reader
- 39:09:18analyzes micro expressions, tone of
- 39:09:21voice and posture to provide feedback on
- 39:09:24communication skills. It can be applied
- 39:09:26in interviews, public speaking or remote
- 39:09:29coaching. Computer vision tools such as
- 39:09:32open pose or media pipe pose track
- 39:09:34gestures and posture. While audio
- 39:09:37processing libraries such as librosa
- 39:09:40analyze tone models can be trained using
- 39:09:43data sets like Raves for audio and CK
- 39:09:46plus for facial expressions. HR tech
- 39:09:49companies like High View, Pytrics and AI
- 39:09:52coaching startups would highly value
- 39:09:54this type of project as it help bridge
- 39:09:57human behavior and AI analysis. Nucation
- 39:10:01is also another area being transformed
- 39:10:03by AI. The personalized tutor with
- 39:10:06adaptive difficulty creates an AI tutor
- 39:10:09that adjusts lessons in real time based
- 39:10:12on student performance. Knowledge
- 39:10:14tracing models like deep knowledge
- 39:10:16tracing combined with transformers for
- 39:10:18content generation allow the AI to adapt
- 39:10:21to each learner. Data sets such as
- 39:10:24assessments or edn nets can be used for
- 39:10:27training. Now this AI can generate new
- 39:10:29questions, explanations and motivational
- 39:10:32feedback based on learning pace.
- 39:10:35Companies like Baiju, Vidanto, Corsera
- 39:10:38and Udemy are constantly looking for
- 39:10:40talent that can build adaptive learning
- 39:10:43platforms. Next, we move into research
- 39:10:45augmentation with the autonomous
- 39:10:48research agent. This AI can answer
- 39:10:50research questions by reading academic
- 39:10:52papers, extracting insights, summarizing
- 39:10:55information, and citing sources
- 39:10:57automatically. It uses the S2 or data
- 39:11:01set for academic papers. Cybboard for
- 39:11:04scientific text embeddings and hugging
- 39:11:06face transformers or lang chain for
- 39:11:09reasoning and summarization. Citation
- 39:11:12extraction can be done with NLP passers
- 39:11:15or rejects academic platforms. AI labs
- 39:11:18and companies like Google research, open
- 39:11:20AI, research gate and LCV would hire
- 39:11:23engineers who can build this system.
- 39:11:26This project connects perfectly with
- 39:11:28education focused AI extending learning
- 39:11:31into automated research capabilities.
- 39:11:35Finally, we arrive at realworld robotics
- 39:11:38control with the AI. The ultimate
- 39:11:40demonstration of cuttingedge skill. This
- 39:11:43project trains AI to control robotic
- 39:11:46arms or humanoids based on goals rather
- 39:11:49than just rigid instructions. It uses pi
- 39:11:52bullet or vbots for simulation. Stable
- 39:11:55baseline 3 or R lib for reinforcement
- 39:11:58learning and open CV or media pipe for
- 39:12:01vision input. Sim to real transfer
- 39:12:04techniques bring simulations into realw
- 39:12:06world scenarios. Robotics companies like
- 39:12:09Boston Dynamics, Appronic, Amazon
- 39:12:12Robotics and Agibot are seeking
- 39:12:15engineers capable of endtoend AIdriven
- 39:12:18robotic systems. After exploring
- 39:12:20softwarebased AI projects, robotics is
- 39:12:23the next step to show mastery of AI
- 39:12:26applied in the physical world. These 10
- 39:12:28AI projects are more than just ideas.
- 39:12:31>> We will learn about some of the machine
- 39:12:33learning and deep learning interview
- 39:12:34questions.
- 39:12:36So let's begin with our first question.
- 39:12:39The first question is how to detect
- 39:12:41outliers in data. So in data analytics
- 39:12:45and machine learning, you often find
- 39:12:47data points that lie at an abnormal
- 39:12:49distance from other points in a random
- 39:12:51sample from a population. Those are
- 39:12:54called outliers. Now outliers in data
- 39:12:56can significantly impact any prediction
- 39:12:58analysis. There are majorly three
- 39:13:00different methods to treat outliers.
- 39:13:03First we have the univariate method. It
- 39:13:06is one of the simplest methods for
- 39:13:07detecting outliers. The univariate
- 39:13:10method uses box plots. A box plot is a
- 39:13:13graphical display for describing the
- 39:13:14distributions of the data. Box plots use
- 39:13:18the median and the lower and upper
- 39:13:19quartiles.
- 39:13:21This method looks for data points with
- 39:13:23extreme values on one variable. Next, we
- 39:13:26have the multivariate method. So, the
- 39:13:28multivariate outliers can be found in an
- 39:13:30n- dimensional space having n features.
- 39:13:33We look for unusual combinations of all
- 39:13:35the variables in this method. Finally,
- 39:13:37we have Minowski error. This method
- 39:13:40reduces the contribution of potential
- 39:13:42outliers in the training process. The
- 39:13:44Minowski error is a loss index that is
- 39:13:46more insensitive to outliers than the
- 39:13:48standard mean squared error. Now moving
- 39:13:50on to the second question. What is a
- 39:13:53confusion matrix? So a confusion matrix
- 39:13:56is a table that is used to describe the
- 39:13:58performance of a classification model on
- 39:14:00a set of test data for which the true
- 39:14:02values are already known. The target
- 39:14:05variable has two values positive or
- 39:14:07negative. The columns represent the
- 39:14:10actual values of the target variable
- 39:14:12which you can see here. The rows
- 39:14:14represent the predicted values of the
- 39:14:15target variable which you can see here.
- 39:14:18Now there are four important terms that
- 39:14:20are related to confusion matrix. First
- 39:14:23we have true positive which is this one.
- 39:14:27So in true positive the predicted value
- 39:14:29matches the actual value. So the actual
- 39:14:31value was positive and the model also
- 39:14:33predicted a positive value. Then we have
- 39:14:36true negative which is also represented
- 39:14:39as tn. The true negative depicts the
- 39:14:42predicted value matches the actual
- 39:14:44value. Now the actual value was negative
- 39:14:46and the model predicted a negative
- 39:14:48value. Next we have false positive. Now
- 39:14:52false positive is also known as a type
- 39:14:54one error. In false positive the
- 39:14:57predicted value was falsely predicted.
- 39:14:59The actual value was negative but the
- 39:15:01model predicted a positive value.
- 39:15:03Finally we have false negative. A false
- 39:15:06negative is also known as type two
- 39:15:08error. So in false negative the
- 39:15:10predicted value was falsely predicted.
- 39:15:13The actual value was positive but the
- 39:15:15model predicted a negative value. Now
- 39:15:17moving to our third question which is
- 39:15:20explain the ROC curve. Now the ROC curve
- 39:15:23is one of the most important evaluation
- 39:15:25metrics for checking the performance of
- 39:15:27any classification model. ROC stands for
- 39:15:30receiver operating characteristic.
- 39:15:33Receiver operating characteristic or ROC
- 39:15:35curve is a method to compare the
- 39:15:37diagnostic tests. The ROC curve is
- 39:15:39created by plotting the true positive
- 39:15:41rate against the false positive rate at
- 39:15:43various threshold settings. So here on
- 39:15:46the y-axis you have the true positive
- 39:15:47rate. On the x-axis we have the false
- 39:15:50positive rate. The true positive rate
- 39:15:52indicates the proportion of observations
- 39:15:54that were correctly predicted to be
- 39:15:55positive out of all positive
- 39:15:57observations. Similarly, the false
- 39:16:00positive rate is the proportion of
- 39:16:02observations that are incorrectly
- 39:16:03predicted to be positive out of all
- 39:16:05negative observations.
- 39:16:08You can take an example. Suppose in
- 39:16:10medical testing, the true positive rate
- 39:16:12is the rate in which people are
- 39:16:14correctly identified to test positive
- 39:16:16for the disease in question. Let's say
- 39:16:18the corona virus testing. ROC does not
- 39:16:21depend on any class distribution. This
- 39:16:23makes it useful for evaluating
- 39:16:25classifiers predicting rare events such
- 39:16:27as diseases or disasters. Now moving to
- 39:16:29the fourth question we have what are the
- 39:16:32assumptions for linear regression. So
- 39:16:34linear regression analysis is used for
- 39:16:36modeling the relationship between a
- 39:16:38single dependent variable Y and one or
- 39:16:40more feature or predictor variables.
- 39:16:43Some of the important assumptions for
- 39:16:45linear regression are so first they
- 39:16:48should have linearity. So linear
- 39:16:50regression needs the relationship
- 39:16:52between the independent and the
- 39:16:53dependent variables to be linear. It is
- 39:16:56also crucial to check for outliers since
- 39:16:58linear regression is sensitive to
- 39:17:00outlier effects. Next we have
- 39:17:02homocyasticity.
- 39:17:04Homoscadasticity
- 39:17:05illustrates a situation in which the
- 39:17:07error term that is the noise or random
- 39:17:10disturbance in the relationship between
- 39:17:11the features and the target variable is
- 39:17:13the same across all levels of the
- 39:17:15dependent variables. Third we have
- 39:17:17independence. So observations should be
- 39:17:19independent of each other. Finally, we
- 39:17:22have no multi-olinearity.
- 39:17:25So there should be little or no
- 39:17:26multi-olinearity.
- 39:17:28Independent variables should not be too
- 39:17:29highly correlated. Now moving to our
- 39:17:32fifth question in our list of interview
- 39:17:33questions.
- 39:17:35The question is what is regularization
- 39:17:38in machine learning? Explain the L2
- 39:17:40regularization.
- 39:17:43So regularization is a machine learning
- 39:17:45technique that is used to reduce the
- 39:17:46errors by fitting the function
- 39:17:48appropriately on the training set in
- 39:17:50order to avoid overfitting of data. So
- 39:17:52overfitting happens when a model learns
- 39:17:54the detail and noise in the training
- 39:17:56data to the extent that it negatively
- 39:17:58impacts the performance of the model on
- 39:18:00new data. So here you can see we have a
- 39:18:03nice plot which shows how overfitting of
- 39:18:06data can be visualized
- 39:18:08and here we have a good fit line over
- 39:18:13the same data points. So this is also
- 39:18:15known as the regression line. Now L2
- 39:18:18regularization is also known as ridge
- 39:18:20regression. So ridge regression modifies
- 39:18:23the overfitted model by adding the
- 39:18:24squared magnitude of coefficient as a
- 39:18:26penalty term to the loss function. So on
- 39:18:29the right you can see a set of data
- 39:18:32points plotted and we have our linear
- 39:18:34regression line and here we are
- 39:18:37calculating the cost function for the
- 39:18:39ridge regression line. So our cost
- 39:18:41function is actually loss plus lambda
- 39:18:45into summation of w ^ 2 where loss is
- 39:18:48actually the sum of squared errors or
- 39:18:50squared residuals. Lambda stands for
- 39:18:53penalty for the errors. W is called the
- 39:18:55slope of the curve or line. Okay. Now
- 39:18:59consider a case where there are two
- 39:19:02points passing through the linear
- 39:19:04regression line. Now if you calculate
- 39:19:06the cost function, we get the value as
- 39:19:091.69. So here we have assumed that loss
- 39:19:12is zero since the two points lie
- 39:19:14directly on the line. We have taken
- 39:19:17lambda to be 1 and w is 1.3. So if you
- 39:19:20use this function or this formula, you
- 39:19:24get the cost function as 1.69.
- 39:19:28Now moving ahead, let's consider another
- 39:19:31situation where we'll calculate the same
- 39:19:33cost function for the ridge regression
- 39:19:35line. There is some loss for both the
- 39:19:38points as they are not on the same line.
- 39:19:41So here you can see the sum of squared
- 39:19:43residuals is 0.05 which actually is the
- 39:19:47sum of 2 squared. I'm assuming this as 2
- 39:19:50and 0.1 for this one.
- 39:19:53So if you square both and add it the
- 39:19:56value is 05 a lambda is again 1 and w we
- 39:20:00have assumed to be 6. Now if you find
- 39:20:03the cost function the value is 41. Let's
- 39:20:06draw the linear regression line and the
- 39:20:08ridge regression line with all the
- 39:20:10points we find that the ridge regression
- 39:20:12line as the best fit since its cost
- 39:20:14function is less. Now coming to the
- 39:20:17sixth question.
- 39:20:19What are the different methods to split
- 39:20:21a tree in a decision tree algorithm?
- 39:20:25So there are three methods to split a
- 39:20:27decision tree. First we have variance.
- 39:20:30So reduction in variance is an algorithm
- 39:20:33that is used for continuous target
- 39:20:34variables. This algorithm uses the
- 39:20:37standard formula variance to choose the
- 39:20:39best split. So here you can see the
- 39:20:41standard formula variance which is
- 39:20:42summation of x that is all the
- 39:20:45individual points minus xar which is the
- 39:20:48mean squared divided by the total number
- 39:20:51of observations. Now the split with
- 39:20:53lower variance is selected as the
- 39:20:55criteria to split the population.
- 39:20:58Now the steps to calculate variance is
- 39:21:00you need to calculate variance for each
- 39:21:02node and then you need to calculate for
- 39:21:05each split as the weighted average of
- 39:21:07each node variance.
- 39:21:10Moving ahead, the second method we have
- 39:21:12is information gain. So information gain
- 39:21:15is used for splitting the nodes when the
- 39:21:17target variable is categorical.
- 39:21:20It works on the concept of entropy. Now
- 39:21:23the degree of disorganization in a
- 39:21:25system is known as entropy. So here you
- 39:21:27can see the formula for information gain
- 39:21:29which is 1 minus entropy.
- 39:21:32Finally we have genie impurity. So genie
- 39:21:36impurity is the probability of
- 39:21:37incorrectly classifying a randomly
- 39:21:39chosen element in the data set if it
- 39:21:42were randomly labeled according to the
- 39:21:43class distribution in the data set. So
- 39:21:45below you can see the formula for gen
- 39:21:48impurity. So we have 1 minus summation
- 39:21:51of pi whole square where n represents
- 39:21:54the number of classes and p of i
- 39:21:56represents the probability of randomly
- 39:21:58picking an element of class i.
- 39:22:02Now moving to the seventh question.
- 39:22:06So the question is how do we find the
- 39:22:08optimum cluster value in K means
- 39:22:09clustering algorithm. Now there are two
- 39:22:12methods to find the optimum cluster
- 39:22:14value. So first we have the elbow method
- 39:22:17which is one of the most wellknown for
- 39:22:18finding the optimum number of clusters.
- 39:22:22So in this method you need to calculate
- 39:22:23the within cluster sum of squared errors
- 39:22:26for different values of K and choose the
- 39:22:28K for which within cluster sum of
- 39:22:30squared errors first starts to diminish.
- 39:22:34So in the below plot of squared errors
- 39:22:36versus the number of clusters K you can
- 39:22:39see at K is equal to 4 the squared error
- 39:22:42starts to diminish. So hence our optimum
- 39:22:46K value is four. Next we have the siloid
- 39:22:50method.
- 39:22:51So the celloid method measures how
- 39:22:53similar a point is to its own cluster
- 39:22:56compared to other clusters. The average
- 39:22:59seloid method computes the average of
- 39:23:01observations for different values of K.
- 39:23:04The optimum number of clusters K is the
- 39:23:06one that maximizes the average over a
- 39:23:09range of possible values for K. The
- 39:23:12seloid score reaches its global maximum
- 39:23:14at the optimal K. So in our case the
- 39:23:17average seloid reaches maximum at k is
- 39:23:20equal to two which you can see here. So
- 39:23:22our optimum cluster value will be two
- 39:23:25here. Moving ahead the eighth question
- 39:23:28in our list is how does the pooling
- 39:23:31layer work in a convolutional neural
- 39:23:32network. So the pooling layer performs a
- 39:23:35downsampling operation in order to
- 39:23:37reduce the dimensionality of the feature
- 39:23:39map. So in the pooling operation, you
- 39:23:42slide a two-dimensional filter over each
- 39:23:45channel of feature map and summarize the
- 39:23:47features lying within the region covered
- 39:23:49by the filter.
- 39:23:51It is a common practice to periodically
- 39:23:53insert a pooling layer in between
- 39:23:54successive convolutional layers in a
- 39:23:56convolutional neural network
- 39:23:58architecture.
- 39:23:59So the pooling layer operates
- 39:24:01independently on every depth slice in
- 39:24:04the input and resizes it specially using
- 39:24:07the max operation. So in the diagram
- 39:24:10shown here you can see we have a
- 39:24:12rectified feature map. We are using a 2
- 39:24:16+2 filter and performing a max pooling
- 39:24:19operation. So consider this as the
- 39:24:21filter. If you perform the max operation
- 39:24:23over the values
- 39:24:27let's say 0 5 3 and 1. So considering
- 39:24:30this one our pool feature map maximum
- 39:24:33value will be five. Similarly for this
- 39:24:35chunk of data it is going to be seven.
- 39:24:38Next, if you slide the filter over this
- 39:24:40square frame, you get eight. And
- 39:24:42similarly here you get six. So this is
- 39:24:45also known as a pulled feature map.
- 39:24:48Moving ahead, the ninth question in our
- 39:24:50list is how does LSTM network work? So
- 39:24:53long short-term memory networks are a
- 39:24:55type of recurrent neural networks that
- 39:24:57are capable of learning order dependence
- 39:24:59and sequence prediction problems. So
- 39:25:01remembering information for long periods
- 39:25:03of time is practically their default
- 39:25:05behavior. Now, LSTMs also have this
- 39:25:08chain-like structure which you can see
- 39:25:10here.
- 39:25:12But the repeating module has a different
- 39:25:14structure. So, instead of having a
- 39:25:16single neural network layer, there are
- 39:25:18four interacting in a very special way.
- 39:25:21Now, you can see these are called as
- 39:25:23gates. These gates contain sigmoid
- 39:25:25activations. A sigmoid activation is
- 39:25:28similar to the tanh activation. Instead
- 39:25:30of squishing values between minus1 and +
- 39:25:34one, it squishes values between 0 and
- 39:25:36one. An LSTDM has four gates.
- 39:25:41Now these are called forget, remember,
- 39:25:44learn and use or output. So if you see
- 39:25:47this in the first step, we use the
- 39:25:49forget gate that decides what
- 39:25:51information should be thrown away or
- 39:25:53kept. the information from the previous
- 39:25:55hidden state and the information from
- 39:25:58the current input is passed through the
- 39:25:59sigmoid function. Values come out
- 39:26:02between zero and one.
- 39:26:04So if the value is closer to zero, it
- 39:26:06means you need to forget that
- 39:26:07information and if the value is closer
- 39:26:09to one, it means you need to keep that
- 39:26:11information.
- 39:26:13Next we have the input gate. So the
- 39:26:16input gate is used to update the cell
- 39:26:18state. First, we pass the previous
- 39:26:20hidden state and the current input into
- 39:26:22a sigmoid function
- 39:26:25that decides which values will be
- 39:26:26updated by transforming the values to be
- 39:26:29between 0 and 1. Zero means not
- 39:26:31important and one means important. You
- 39:26:34also pass the hidden state and current
- 39:26:36input into the tan function to flatten
- 39:26:38the values between minus1 and + one.
- 39:26:41This helps to regulate the network.
- 39:26:44Then you multiply the tan output with
- 39:26:46the sigmoid output. The sigmoid output
- 39:26:49will decide which information is
- 39:26:50important to keep from the tanh output.
- 39:26:54And finally in step three we have the
- 39:26:56output gate. This output gate is used to
- 39:26:59decide what the next hidden state should
- 39:27:02be. First we pass the previous hidden
- 39:27:04state and the current input into a
- 39:27:06sigmoid function. Then we pass the newly
- 39:27:09modified cell state into the tanage
- 39:27:11function. We then multiply the tanage
- 39:27:14output with the sigmoid output to decide
- 39:27:16what information the hidden state should
- 39:27:18carry. The output is the hidden state.
- 39:27:21The new cell state and the new hidden
- 39:27:23state is then carried over to the next
- 39:27:25time step. Finally, talking about the
- 39:27:29last question in our list of interview
- 39:27:30questions, we have explained the concept
- 39:27:33of gradient descent in deep learning.
- 39:27:36Now gradient descent is an optimization
- 39:27:38algorithm which is mainly used to find
- 39:27:40the minimum of a function in machine
- 39:27:42learning. Gradient descent is used to
- 39:27:44update the parameters in a model.
- 39:27:47Parameters can vary according to the
- 39:27:49algorithms such as coefficients in
- 39:27:51linear regression and weights in neural
- 39:27:53networks.
- 39:27:55You can see we have these maps and on
- 39:27:59the y-axis we have the loss. On the
- 39:28:02x-axis we have the weight and here we
- 39:28:05are trying to find the local minimum or
- 39:28:08the global minimum. Now this gradient
- 39:28:11descent method is used to minimize the
- 39:28:13cost function and update the parameters
- 39:28:14of the learning model. The gradient
- 39:28:17always points in the direction of the
- 39:28:19steepest increase in the loss function.
- 39:28:22The gradient descent algorithm takes a
- 39:28:24step in the direction of the negative
- 39:28:25gradient in order to reduce the loss as
- 39:28:28quickly as possible. To determine the
- 39:28:30next point along the loss function
- 39:28:32curve, the gradient descent algorithm
- 39:28:34adds some fraction of the gradient's
- 39:28:36magnitude to the starting point. Now
- 39:28:38this process is repeated to find the
- 39:28:40global minimum.
- 39:28:42>> Welcome to math refresher probability
- 39:28:45and statistics.
- 39:28:47In this lesson, we are going to explain
- 39:28:49the concepts of statistics and
- 39:28:52probability.
- 39:28:53Describe conditional probability. Define
- 39:28:56the chain rule of probability. Discuss
- 39:28:59the measure of variance. Identify the
- 39:29:01types of gshian distribution.
- 39:29:04Basic of statistics and probability.
- 39:29:07Probability and statistics. Data science
- 39:29:10relies heavily on estimates and
- 39:29:12predictions. A significant portion of
- 39:29:15data science is made up of evaluations
- 39:29:17and forecast.
- 39:29:19Statistical methods are used to make
- 39:29:21estimates for further analysis.
- 39:29:24Probability theory is helpful for making
- 39:29:26predictions. Statistical methods are
- 39:29:29highly dependent on probability theory
- 39:29:32and all probability and statistics are
- 39:29:35dependent on data.
- 39:29:37Data is information acquired for
- 39:29:39reference or research via observations,
- 39:29:43facts, and measurements. Data is a set
- 39:29:46of facts structured in the form that
- 39:29:48computers can interpret such as numbers,
- 39:29:51words, estimations, and views.
- 39:29:54Importance of data. Data aids in seeing
- 39:29:57more about the information by
- 39:29:59identifying possible connections between
- 39:30:01two features. Data assists in the
- 39:30:04detection of distortion by uncovering
- 39:30:06hidden patterns based on prior
- 39:30:09information patterns. Data may be
- 39:30:12utilized to anticipate the future or
- 39:30:14predict the current state of affairs.
- 39:30:17Also, data aids in determining whether
- 39:30:19two pieces of information have any
- 39:30:21instance in common or not. Types of
- 39:30:24data. Data might be quantitative. That
- 39:30:28is data that can be measured or counted
- 39:30:30in numbers or it may be qualitative
- 39:30:33which is data which is generally divided
- 39:30:35into groups or in simpler words which
- 39:30:38cannot be counted or measured in
- 39:30:40numbers. Let's consider an example a
- 39:30:43customer information data of a bank may
- 39:30:46contain quantitative and qualitative
- 39:30:48data. Consider this snapshot where we
- 39:30:51have customer ID, surname, geography,
- 39:30:55gender, age, balance, has C or card is
- 39:30:58active member. Amongst these variables
- 39:31:01we can see surname is mostly qualitative
- 39:31:04as it cannot be counted and measured in
- 39:31:06numbers. Geography and gender are also
- 39:31:10qualitative as they cannot be counted in
- 39:31:12numbers and are mostly groups. has C or
- 39:31:16card that is has credit card and is
- 39:31:19active member although are containing
- 39:31:21numerical in form but these are
- 39:31:24categorical that means these have been
- 39:31:26divided into groups of one and zero that
- 39:31:30represent yes and no as an answer hence
- 39:31:33these two variables are also qualitative
- 39:31:37customer ID is again although a
- 39:31:40numerical data however the significance
- 39:31:43or intuition behind Customer ID is
- 39:31:46categorical.
- 39:31:47Hence, it may be kept in the qualitative
- 39:31:50data also. However, age and balance
- 39:31:54these are numerical information which
- 39:31:56have been measured or counted and
- 39:31:58numerical operations can be performed on
- 39:32:01them. Hence, these are under
- 39:32:03quantitative data categories.
- 39:32:05Introduction to descriptive statistics.
- 39:32:08Descriptive statistics. A descriptive
- 39:32:11measurement is summary measure that
- 39:32:13quantitatively portrays the most
- 39:32:15important features of a set of data
- 39:32:18allowing for a better comprehension of
- 39:32:20the information. Data can be measured as
- 39:32:22different levels. The levels of
- 39:32:25measurement describe the nature of
- 39:32:26information stored in the data assigned
- 39:32:29to the variables. Qualitative data can
- 39:32:32be measured as nominal or ordinal.
- 39:32:34Quantitative data can be measured in
- 39:32:36terms of interval and ratio type.
- 39:32:39Nominal data. The data is categorized
- 39:32:42using names, labels or qualities. For
- 39:32:45example, brand name, zip code, and
- 39:32:47gender. Ordinal data can be arranged in
- 39:32:50order or ranked and can be compared.
- 39:32:53Examples include grades, star reviews,
- 39:32:57position, and race, and date. Interval
- 39:33:00data is the data that is ordered and has
- 39:33:02meaningful differences between the data
- 39:33:05points. Example temperature in Celsius
- 39:33:08and year of birth. Ratio data is similar
- 39:33:11to the interval level with the added
- 39:33:14property of inherent zero. Mathematical
- 39:33:17calculations can be performed on both
- 39:33:19interval as well as ratio data. For
- 39:33:22example, height, age, and weight.
- 39:33:25Population versus sample. Before
- 39:33:28analyzing the data, it's important to
- 39:33:30figure out if it's from a population or
- 39:33:32a sample. Population is a collection of
- 39:33:36all available items as well as each unit
- 39:33:38in our study. Sample is a subset of the
- 39:33:41population that contains only a few
- 39:33:44units of the population. Population data
- 39:33:47is used for study when the data pool is
- 39:33:50very small and can give all the required
- 39:33:52information. Samples are collected
- 39:33:55randomly and represent the entire
- 39:33:58population in the best possible way.
- 39:34:01Measures of central tendency.
- 39:34:04The central tendency is a single value
- 39:34:06that aids in the description of the data
- 39:34:09by determining its center position.
- 39:34:12Measures of central tendency are
- 39:34:14sometimes known as summary statistics or
- 39:34:17measures of central location.
- 39:34:19The most popular measurements of central
- 39:34:22tendency are mean, median, and mode. The
- 39:34:26normal distribution is a bell-shaped
- 39:34:28symmetrical distribution in which mean,
- 39:34:31median, and mode all are equal. The
- 39:34:34curve over here shows the bell-shaped
- 39:34:36curve or the normal distribution of
- 39:34:38variable X. The point over here that is
- 39:34:41X1 is the point which represents the
- 39:34:45mean, median and mode of this
- 39:34:47distribution. Mean mean is calculated by
- 39:34:50dividing these sum of all data values by
- 39:34:53the total number of data values. It gets
- 39:34:57affected when there are unusual or
- 39:34:59extreme values. It is sensitive to the
- 39:35:02outliers. Mean can be calculated as
- 39:35:05summation over all the values of X in a
- 39:35:08collection divided by the size of the
- 39:35:10collection.
- 39:35:12For example, we have a collection where
- 39:35:14we have values as 7 3 4 1 6 and 7.
- 39:35:20We find out the sum of these values
- 39:35:22which is 28 and there are total of six
- 39:35:26values. So 28 / 6 gives us a mean value
- 39:35:30of 4.66.
- 39:35:33Median,
- 39:35:35it is the middle value in the set of the
- 39:35:37data that has been sorted in ascending
- 39:35:39order.
- 39:35:41It is a better alternative to mean since
- 39:35:43it is less impacted by outliers and
- 39:35:46skewess.
- 39:35:47It is closer to the actual central
- 39:35:50value.
- 39:35:51Median is calculated differently for
- 39:35:54different sizes of data.
- 39:35:56Differentiated as if the total number of
- 39:35:58values is odd or if the total number of
- 39:36:02values is even. If the size of the data
- 39:36:05is odd. For example, in this case we
- 39:36:09have five elements.
- 39:36:12After sorting whatever middle value we
- 39:36:15get
- 39:36:17that means n + 1 by 2 term in this case
- 39:36:225 + 1 / 2
- 39:36:25that is the third term which is four is
- 39:36:28the median value.
- 39:36:31In case when the total number of values
- 39:36:33is even like here there are six values.
- 39:36:37The average or the mean of the two
- 39:36:39central values is considered as the
- 39:36:41median. In this case the median is the
- 39:36:44mean of 6 and four which is five. Mode.
- 39:36:50Mode represents the most common value in
- 39:36:52the data set. It is not at all affected
- 39:36:55by extreme observations.
- 39:36:59It is the best measure of central
- 39:37:01tendency for highly skewed or non-normal
- 39:37:04distribution.
- 39:37:05Mode for categorical data is determined
- 39:37:08by estimating the frequencies for each
- 39:37:10categories
- 39:37:12and then the category with the highest
- 39:37:14frequency is considered to be mode.
- 39:37:17Like in this case 7 has the highest
- 39:37:20frequency. Hence seven becomes the mode
- 39:37:22value. However, in case of continuous
- 39:37:26data or quantitative data, the
- 39:37:28calculation of mode is slightly
- 39:37:30different. The first step in calculation
- 39:37:33of mode is dividing the data into
- 39:37:35classes which are equal with then
- 39:37:37getting the frequency of data points
- 39:37:39lying in within that range of classes
- 39:37:42and finally selecting the class with the
- 39:37:45highest frequency.
- 39:37:48Using the range of that class and the
- 39:37:50frequencies, we can get the final mode
- 39:37:52value.
- 39:37:54Using the formula L+
- 39:37:58minus F_sub_1 multiplied to H / FM minus
- 39:38:02F_sub_1 plus FM minus F_sub_2.
- 39:38:06Here L is the lower limit or the lower
- 39:38:09observation of the mode class.
- 39:38:12H is the size of the mode class.
- 39:38:16FM is the frequency of the mode class.
- 39:38:19F_sub_1 is the frequency of the class
- 39:38:22proceeding to mode and F_sub_2 is the
- 39:38:25frequency of the class succeeding to
- 39:38:28mode. This gives us the final mode
- 39:38:30value,
- 39:38:32mean versus expectation.
- 39:38:35Now let's talk about mean versus
- 39:38:37expectation.
- 39:38:38So in general we use the expected value
- 39:38:41or expectation when we want to calculate
- 39:38:44the mean of a probability distribution
- 39:38:47that represents the average value we
- 39:38:50expect to occur before collecting any
- 39:38:52data. And mean on the other hand mean is
- 39:38:55basically used when we want to calculate
- 39:38:58the average value of a given sample.
- 39:39:01This represents the average value of raw
- 39:39:04data that we may have already collected.
- 39:39:07We can understand this by using a simple
- 39:39:10example.
- 39:39:12Now to calculate the expected value of
- 39:39:15this probability distribution, we can
- 39:39:18use a specific formula from the previous
- 39:39:20discussion.
- 39:39:22This is going to be the expected value
- 39:39:24where X is going to be the data value
- 39:39:27and this PX is the probability of value.
- 39:39:32For example, we could calculate the
- 39:39:34expected value for this probability
- 39:39:36distribution to be as shown.
- 39:39:41So here it will be 1.45 goals.
- 39:39:46So this represents the expected number
- 39:39:48of goals that the team will score in any
- 39:39:50given game.
- 39:39:53And then if you talk about calculating
- 39:39:55mean, so we typically calculate the mean
- 39:39:58after we have actually collected raw
- 39:40:00data.
- 39:40:03For example, suppose we record the
- 39:40:05number of goals that a soccer team will
- 39:40:07score in 15 different games.
- 39:40:12Now to calculate the mean number of
- 39:40:14goals scored per game,
- 39:40:17we can use the following formula
- 39:40:20where sum of x is basically the sum of
- 39:40:23all the goals divided by n and the
- 39:40:26number of records or we can say the
- 39:40:27sample size.
- 39:40:30It is as shown on the screen.
- 39:40:48So this represents the mean number of
- 39:40:50goals scored per game by the team.
- 39:40:53Measures of asymmetry.
- 39:40:56The difference between the three
- 39:40:57distinct curves can be studied in this
- 39:41:00image.
- 39:41:01The central curve is the normal or no
- 39:41:04skewess curve. Here mean, median and
- 39:41:07mode all lie on the same point. This
- 39:41:10normal curve is symmetrical about its
- 39:41:12mean, median and mode.
- 39:41:16That means the left hand side of the
- 39:41:18curve is a mirror image of the right
- 39:41:20hand side of the curve.
- 39:41:23However, in case of negatively skewed
- 39:41:26data, the tail is elongated on the left
- 39:41:30hand side
- 39:41:32and the mean is smaller than the mode
- 39:41:34and the median values or is on the left
- 39:41:37hand side of the mode.
- 39:41:40Hence indicating that the outliers are
- 39:41:42in the negative direction.
- 39:41:45On the other hand, in case of positively
- 39:41:48skewed, the data is concentrated on the
- 39:41:50left hand side of the curve.
- 39:41:54While the tail is elongated or longer on
- 39:41:56the right hand side of the curve,
- 39:42:00the mean is greater than the mode and
- 39:42:02median
- 39:42:03or is on the right hand side of the mode
- 39:42:06and median indicating that the outliers
- 39:42:08are in the positive direction.
- 39:42:15Let's consider an example.
- 39:42:18The graph here shows the global income
- 39:42:20distribution for the year 2003 2013 and
- 39:42:25a projection for 2035.
- 39:42:28If we see the global income distribution
- 39:42:30statistics for 2003 it is highly right
- 39:42:34skewed.
- 39:42:37We can observe in the previous graph
- 39:42:39that in 2003
- 39:42:43the mean of $3,451
- 39:42:48was higher than the median of $1090.
- 39:42:52The global income is definitely not
- 39:42:55evenly distributed. The majority of
- 39:42:57people make less than $2,000 each year.
- 39:43:03while only a small percentage of the
- 39:43:05population earns more than $14,000.
- 39:43:10Measures of variability.
- 39:43:15Measures of variability.
- 39:43:17Dispersion. The measure of central
- 39:43:20tendencies provide a single value that
- 39:43:22addresses the full worth. However, the
- 39:43:25central tendency cannot depict the
- 39:43:27viewpoint entirely. The metric of
- 39:43:30dispersion helps us focus on the
- 39:43:32inconsistency in the data spread.
- 39:43:35Measures of dispersion describe the
- 39:43:37spread of the data.
- 39:43:40The range, intercortile range, standard
- 39:43:43deviation and variance are examples of
- 39:43:46dispersion measures.
- 39:43:49Range.
- 39:43:51The range of distribution is the
- 39:43:53difference between the largest and the
- 39:43:55smallest amount of data.
- 39:43:58The range, for example, does not include
- 39:44:01all of a series positive aspects.
- 39:44:04It concentrates on the most shocking
- 39:44:07aspects and ignores that aren't
- 39:44:09considered critical. For example, for a
- 39:44:11set 13, 33, 45, 67, 70.
- 39:44:17The range is 57. That is the maximum of
- 39:44:22this which is 70 minus the minimum over
- 39:44:24here which is 13.
- 39:44:28Variance.
- 39:44:31Variance is the average of all squared
- 39:44:33deviations.
- 39:44:36It is defined as the sum of squared
- 39:44:38distance between each point and the mean
- 39:44:41or the dispersion around the mean.
- 39:44:44The standard deviation is used as
- 39:44:46variance suffers from a unit difference.
- 39:44:51Variance can be computed as sigma square
- 39:44:54summation over x - mu^ 2
- 39:44:58divided by n
- 39:45:00where mu is the mean of the data, x is
- 39:45:03the individual data point
- 39:45:06and n is the size of the data.
- 39:45:11This representation is for a population
- 39:45:13data.
- 39:45:15For a sample data variance can be
- 39:45:18computed as X minus
- 39:45:20Xar whole square summation
- 39:45:23over it divided by n minus one.
- 39:45:27Here Xbar is the mean of these sample
- 39:45:30data and n is the sample size.
- 39:45:34The units of values and variance are not
- 39:45:37equal.
- 39:45:39So another variability measure is used.
- 39:45:43Standard deviation.
- 39:45:47Standard deviation is a statistical term
- 39:45:49used to measure the amount of
- 39:45:51variability or dispersion around a mean.
- 39:45:56The standard deviation is calculated as
- 39:45:59the square root of variance. It depicts
- 39:46:02the concentration of the data around the
- 39:46:05mean of the data set.
- 39:46:08Standard deviation as indicated
- 39:46:11previously can be computed as square
- 39:46:13root of variance
- 39:46:15for a population data. Standard
- 39:46:18deviation sigma can be computed as
- 39:46:21square root of summation over x i minus
- 39:46:24mu^ square / n
- 39:46:28where mu is the mean of the data x i are
- 39:46:32the data points and n is the size. Let's
- 39:46:35consider an example.
- 39:46:38Let's find out the mean, variance, and
- 39:46:41standard deviation for this data. The
- 39:46:44data values are three, 5, 6, 9, and 10.
- 39:46:49To find out the mean, we first find the
- 39:46:51sum of all these data values
- 39:46:55that is 33 and divide it by the count,
- 39:46:58which is five.
- 39:47:01We get the mean of 6.6. To compute the
- 39:47:04variance, we start by computing the
- 39:47:07deviation.
- 39:47:08That is X minus the mean of X. Here
- 39:47:12three is one of the values of the data
- 39:47:14and 6.6 is the mean.
- 39:47:18So 3 - 6.6 squared and we do that
- 39:47:24to find out sum of all the deviations
- 39:47:26divided by the count
- 39:47:29which is five.
- 39:47:31we end up getting an overall variance of
- 39:47:336.64.
- 39:47:37Standard deviation as we know is
- 39:47:39measured at square root of variance that
- 39:47:42is square<unk> of 6.64
- 39:47:46which amounts to 2.576.
- 39:47:50Measures of relationship.
- 39:47:53Measures of relationship. Coariance.
- 39:47:56Coariance is the measure of joint
- 39:47:58variability of two variables.
- 39:48:01It measures the direction of the
- 39:48:03relationship between the variables. It
- 39:48:06determines if one variable will cause
- 39:48:08the other to alter in the same way.
- 39:48:13Coariance between variable X and Y can
- 39:48:16be computed as summation over the
- 39:48:19product of X I - XR
- 39:48:22and Y I - Y bar the whole divided by N
- 39:48:26minus one.
- 39:48:29Here Xar and Y bar are the mean of X and
- 39:48:32Y respectively. The value of covariance
- 39:48:36can range from minus infinity to a plus
- 39:48:38infinity.
- 39:48:41Correlation.
- 39:48:43Correlation is normalized coariance.
- 39:48:47It measures the strength of association
- 39:48:50between two variables. The most common
- 39:48:52measure for correlation is the Pearson
- 39:48:55correlation coefficient.
- 39:48:57Correlation between two variables
- 39:49:01X and Y can be measured with respect to
- 39:49:03coariance as coariance between X
- 39:49:07and Y divided by the standard deviation
- 39:49:10of X and standard deviation of Y.
- 39:49:14The value of correlation ranges from a
- 39:49:17negative 1 to positive 1.
- 39:49:21Types of correlation.
- 39:49:24Correlation can be either a positive
- 39:49:27correlation,
- 39:49:28zero correlation or a negative
- 39:49:31correlation.
- 39:49:35The first picture over here represents a
- 39:49:37perfect positive correlation
- 39:49:41wherein a straight line with a positive
- 39:49:43slope
- 39:49:45is representing the relationship between
- 39:49:47the two variables.
- 39:49:50Zero correlation means that the line
- 39:49:52representing the relationship between
- 39:49:54the two variables is horizontal to the
- 39:49:57xaxis.
- 39:50:00Perfect negative correlation can be
- 39:50:02represented by a straight line with a
- 39:50:05negative slope.
- 39:50:08Correlation equals to 1 implies a
- 39:50:11positive relationship. That is when one
- 39:50:14variable increases the other variable
- 39:50:16also increases. A correlation value of
- 39:50:19negative 1 implies a negative
- 39:50:21relationship. That is when one variable
- 39:50:24increases the other decreases.
- 39:50:28The correlation coefficient of zero
- 39:50:31shows that the variables are completely
- 39:50:33independent of each other.
- 39:50:36Let's consider an example.
- 39:50:40Here we have two variables height and
- 39:50:43weight.
- 39:50:46To compute the correlation between
- 39:50:48height and weight,
- 39:50:50we use the correlation formula as
- 39:50:52covariance of X
- 39:50:55and Y divided by standard deviation of X
- 39:50:58and standard deviation of Y.
- 39:51:01Here height is the X variable and weight
- 39:51:04is the Y variable.
- 39:51:07First to compute coariance we compute
- 39:51:10the x - xar and y - y bar values and
- 39:51:14then the product of them.
- 39:51:18We then compute x - xr²
- 39:51:23and y - y bar square values to compute
- 39:51:26the standard deviations of height and
- 39:51:28weight respectively. Correlation as we
- 39:51:31know has been defined as covariance of x
- 39:51:34and i and y divided by standard
- 39:51:37deviations of x and y.
- 39:51:41This can also be represented as
- 39:51:43summation over x - xr multiplied to y -
- 39:51:48y bar
- 39:51:50divided by square root of summation over
- 39:51:52sum of squared deviations
- 39:51:55that is x - xr square multiplied to
- 39:51:58square root of summation over y - yar
- 39:52:02whole square that is sum of square
- 39:52:05deviations for y.
- 39:52:08Now let's find out values to put into
- 39:52:11this formula.
- 39:52:14First we find out the overall sum of
- 39:52:17height to get the mean of height which
- 39:52:20is 5.14.
- 39:52:22Similarly we get the sum of weight to
- 39:52:24get the mean of weight as 50. We now get
- 39:52:27the summation over x - xr multiplied to
- 39:52:31y - y bar to get the numerator for the
- 39:52:34formula. Then we compute x - xr square
- 39:52:38summation
- 39:52:40and y - y bar square that is sum of
- 39:52:43squared deviation of x and y
- 39:52:46respectively.
- 39:52:48Now we put in the values in this final
- 39:52:50correlation formula to get a correlation
- 39:52:52value of 0.889.
- 39:52:57This indicates that height and weight
- 39:52:59have a positive relationship.
- 39:53:02It is evident that as height grows,
- 39:53:05weight also increases.
- 39:53:09In this module, we will be talking about
- 39:53:11expectation and variance.
- 39:53:14So the expected value or we can say mean
- 39:53:17of a given variable that we can denote
- 39:53:20by X is a discrete random variable where
- 39:53:22it is a weighted average of the possible
- 39:53:25values that X can take and each value is
- 39:53:28going to be according to the probability
- 39:53:30of that specific event occurring.
- 39:53:34So usually the expected value of X is
- 39:53:36denoted by a simple formula where we can
- 39:53:40define the expectation based on the X
- 39:53:42parameter.
- 39:53:45which is going to be the sum of each
- 39:53:47possible outcome multiplied by the
- 39:53:49probability of the outcome occurring.
- 39:53:53So in more concrete terms, the
- 39:53:56expectation is what we would expect the
- 39:53:58outcome of an experiment to be on
- 39:54:00average.
- 39:54:04We can take an example for the coin. If
- 39:54:07a coin is being tossed 10 times, then
- 39:54:10one is most likely to get five heads and
- 39:54:13five tails.
- 39:54:16Same logic can be discussed if we talk
- 39:54:19about another example of rolling a die.
- 39:54:22So there are six possible outcomes when
- 39:54:24you roll a dieice 1 2 3 4 5 6. And each
- 39:54:28of these has a probability of 1 by 6 of
- 39:54:31occurring. So we can say that the
- 39:54:34expectation is going to be 1 multiplied
- 39:54:36by the probability of that happening
- 39:54:39which is going to be 1x 6 + 2x 6 + 3x 6
- 39:54:44+ 4x 6 + 5x 6 + 6x 6 and that is going
- 39:54:50to give us 3.5 as an output. The
- 39:54:53expected value is 3.5.
- 39:54:57So if you think about it, 3.5 is halfway
- 39:55:00between the possible values that I can
- 39:55:02take and this is what we should have
- 39:55:05expected.
- 39:55:06Next we talk about the concept of
- 39:55:08variance. So variance of a random
- 39:55:11variable allows us to know something
- 39:55:13about the spread of the possible values
- 39:55:16of the variable. So for a discrete
- 39:55:19random variable X the variances of X is
- 39:55:22going to be denoted by using a simple
- 39:55:24formula that is going to be var=
- 39:55:27E X - M the whole square where M is
- 39:55:31basically the expected value of the
- 39:55:33expectation of X. So this is more like a
- 39:55:36standard deviation of X which can also
- 39:55:39be represented by using this formula. So
- 39:55:42the variance does not behave in the same
- 39:55:44way as expectation when we multiply and
- 39:55:47add constants to random variables.
- 39:55:52So now there are two different type of
- 39:55:54variance that we can have a fair
- 39:55:56understanding on. First of all we have
- 39:55:59low variance and then we have high
- 39:56:02variance.
- 39:56:04So low variance simply means that there
- 39:56:06is a small variation in the production
- 39:56:09of the target function with changes in
- 39:56:12the trading data set and at the same
- 39:56:14time high variance as we can see here
- 39:56:17high variance shows a large variation in
- 39:56:20prediction of the target function with
- 39:56:22changes in the trading data set. So a
- 39:56:25model that shows high variance learns a
- 39:56:27lot and perform well with the training
- 39:56:29data set and it does not generalize well
- 39:56:32with the unseen data set and that's why
- 39:56:35as a result such a model gives good
- 39:56:38results with training data set but shows
- 39:56:40high error rates on the test data set
- 39:56:43and since the high variance a model
- 39:56:45learns too much from the data set it
- 39:56:48leads to an overfitting of the model. So
- 39:56:51model with high variance will be having
- 39:56:53couple of issues like it may lead to
- 39:56:55overfitting or it may also lead to
- 39:56:57increase in model complexities.
- 39:57:02Next we have skewess.
- 39:57:05So skewess in simple terms is basically
- 39:57:08a measure of asymmetry of a
- 39:57:10distribution. So distribution is
- 39:57:13asymmetrical when its left and right
- 39:57:15sides are not the mirror images.
- 39:57:18Right now this is a mirrored image and a
- 39:57:21distribution can have right positive or
- 39:57:23we can say negative or it can have zero
- 39:57:26skewess.
- 39:57:29So right skewed in this scenario is
- 39:57:32basically the distribution is longer on
- 39:57:34the right side of its peak
- 39:57:37and a left skew distribution is going to
- 39:57:39be we can say where it is longer on the
- 39:57:42left side.
- 39:57:44So we can see we have this one as a part
- 39:57:46of right side. It is more elongated
- 39:57:49towards the right side and this one is
- 39:57:52more elongated towards the left side. So
- 39:57:55we can think of skewess in terms of
- 39:57:57tails. A tail is long tampering and the
- 39:58:00end of a distribution. So it simply
- 39:58:03indicates that they are observations at
- 39:58:05one end of the distribution but that
- 39:58:07they are relatively infrequent. So a
- 39:58:10right skew distribution has a long tail
- 39:58:13on the right side as you can see here.
- 39:58:15So the number supports observed. Let's
- 39:58:18say we have a data on a per year basis.
- 39:58:21So again we can have a more skewess
- 39:58:23towards the right side where data is
- 39:58:25being dropping as we continue to
- 39:58:28increase the number of years. For
- 39:58:30example we may have a high sales towards
- 39:58:32the beginning of year suppose in 2022
- 39:58:36but again as we proceed to 2023 second
- 39:58:40half we are seeing the dip in
- 39:58:42performance. So that is rightly skewed
- 39:58:45and same way let's suppose if we started
- 39:58:47with the sales figure it was really less
- 39:58:50in suppose 2002
- 39:58:53but again as we proceeded to 2023 now
- 39:58:56our sales have been gradually
- 39:58:58increasing. So it's more like skew
- 39:59:01towards the left section as a part of
- 39:59:03negative skew. Next we have curtosis.
- 39:59:08So curtosis is basically a measure of
- 39:59:11the tailness of a distribution.
- 39:59:14So taeness is how often the outliers
- 39:59:17occur and act as curtis is the tailness
- 39:59:20of the distribution related to a normal
- 39:59:23distribution. So a distribution with
- 39:59:26medium curttosis is called as messortic.
- 39:59:29A distribution with low curtosis like
- 39:59:31this one. This is called as the
- 39:59:33platicurtic and then distribution with
- 39:59:36high curtosis like this one. This is
- 39:59:38called as the leptocortic.
- 39:59:42So tails here they are tapering ends on
- 39:59:44either side of a distribution like this.
- 39:59:47So they represent the probability or the
- 39:59:49frequency of values that are extremely
- 39:59:52high or extremely low to the mean.
- 39:59:55In other words, tails here represents
- 39:59:58how often the outliers occur.
- 40:00:02So there are three type of curtis. We
- 40:00:05have platicurtic which is negative,
- 40:00:07leptocortic which is a positive towards
- 40:00:09the upper end and then we have messertic
- 40:00:12which is a normal distribution. So
- 40:00:15messertic is the medium tail. So normal
- 40:00:18distributions they have a curtosis of
- 40:00:20three. So any distribution with a
- 40:00:23kurtosis of approx value of three is
- 40:00:26going to be messertic. And curtosis is
- 40:00:29described in terms of excess curttosis
- 40:00:32which is curtosis minus3. And since
- 40:00:35normal distribution they have a curtosis
- 40:00:38of three axis curtises makes comparing a
- 40:00:41distribution curtosis to a normal
- 40:00:43distribution even easier. Introduction
- 40:00:46to probability.
- 40:00:49Probability theory. Probability is a
- 40:00:52measure of the likelihood that an event
- 40:00:54will occur.
- 40:00:57Let's consider an example of coin toss
- 40:01:01where the chances of getting heads on a
- 40:01:03coin are 1 by two or 50%.
- 40:01:07The probability of each given event is
- 40:01:09between zero and one both inclusive. Sum
- 40:01:13of an events cumulative probability
- 40:01:16cannot be greater than one.
- 40:01:19Hence the probability of an event X lies
- 40:01:22between zero and one. This means that
- 40:01:25the integral of probability of
- 40:01:27distribution over x equals to 1.
- 40:01:33Conditional probability. Conditional
- 40:01:35probability of any event A is defined as
- 40:01:38the probability of occurrence of A given
- 40:01:42that event B has previously occurred.
- 40:01:47Condition probability of event A given B
- 40:01:50can be estimated as probability of A
- 40:01:53intersection B that is probability of
- 40:01:56both A and B happening together
- 40:01:59divided by the probability of B.
- 40:02:04It is also written as that probability
- 40:02:06of A intersection B equals to
- 40:02:09probability of A given B multiplied to
- 40:02:13probability of B.
- 40:02:18Let's consider an example.
- 40:02:20In a coin, we are doing a two coin flip.
- 40:02:23Coin one gets heads, tails, heads, and
- 40:02:26tails in subsequent flips.
- 40:02:30while coin two gets tails, heads, heads,
- 40:02:33and tails in the subsequent flips. Now,
- 40:02:37the probability that coin one will get a
- 40:02:39head is 2 out of four. While the
- 40:02:42probability that coin two will get heads
- 40:02:45is again two out of four.
- 40:02:48The probability that both coin one and
- 40:02:50coin two will have a heads is just one
- 40:02:53out of the four flips.
- 40:02:57Hence the probability that coin one will
- 40:02:59get heads given that coin 2 is already
- 40:03:02heads can be computed as probability of
- 40:03:05coin one edge intersection coin 2 edge
- 40:03:09that is 1x4 divided by probability of
- 40:03:12coin 2 edge
- 40:03:16that's a given that is 2x 4 which is
- 40:03:19going to be 0.5 or 50% based
- 40:03:24base theorem Base theorem calculates the
- 40:03:27conditional probability of an event
- 40:03:29based on its prior probabilities.
- 40:03:33Basically base theorem incorporates the
- 40:03:36prior probability distribution to
- 40:03:38predict the posterior probabilities.
- 40:03:40Base theorem for conditional probability
- 40:03:44can be expressed as probability of A
- 40:03:47given B equals probability of B given A
- 40:03:51divided by probability of B multiplied
- 40:03:54to probability of A.
- 40:03:57Base theorem allows updating the
- 40:03:59probability values by using new
- 40:04:01information or evidence. Here
- 40:04:04probability of A is known as prior
- 40:04:06probability. That is the probability of
- 40:04:09event before any new data is collected.
- 40:04:12Probability of A given B is known as the
- 40:04:16posterior probability. It is the revised
- 40:04:19probability of an event occurring after
- 40:04:21taking into consideration the new
- 40:04:24information probability of B given A is
- 40:04:27known as the likelihood and probability
- 40:04:29of B is probability of observing an
- 40:04:32evidence B model. An example consider an
- 40:04:36example for calculating the likelihood
- 40:04:38of having diabetes based on frequency of
- 40:04:41fast food consumption. Here is the
- 40:04:44observed data. Let's say the fast food
- 40:04:47audience is 20%. Diabetes prevalence is
- 40:04:5110% and 5% is fast food and diabetes.
- 40:04:56The chances of diabetes given fast food
- 40:04:58that is the conditional probability of D
- 40:05:01given B can be calculated as probability
- 40:05:04of diabetes and fast food together
- 40:05:07divided by probability of fast food.
- 40:05:10That means 5% divided by 20%. that
- 40:05:14equals 25%.
- 40:05:16Define an analysis can state eating fast
- 40:05:19food increases the chance of having
- 40:05:21diabetes by 25%.
- 40:05:24The multiplication rule of probability
- 40:05:27if events A and B are statistically
- 40:05:30independent and probability of A
- 40:05:33intersection B can be given as
- 40:05:35probability of A given B multiplied to
- 40:05:39probability of B. However, probability
- 40:05:42of A intersection B is also given as
- 40:05:45probability of A multiplied to
- 40:05:48probability of B. Here probability of A
- 40:05:52given B equals to probability of A when
- 40:05:56we assume that probability of B is non
- 40:05:59zero. Similarly, probability of B equals
- 40:06:02probability of B given A assuming
- 40:06:05probability of A is non zero.
- 40:06:08Chain rule of probability joint
- 40:06:10probability distributions over many
- 40:06:13random variables can be reduced into
- 40:06:16conditional distributions over a single
- 40:06:18variable. It can be expressed as
- 40:06:21probability of X1 X2 so on until Xn
- 40:06:25equals probability of X1 intersection
- 40:06:28probability of X I given probability of
- 40:06:31X1 till X I minus one.
- 40:06:36For example, the joint probability of A,
- 40:06:38B and C can be given as probability of A
- 40:06:42given B. C multiplied to probability of
- 40:06:46B given C multiply to probability of C.
- 40:06:50Logistic sigmoid.
- 40:06:54The logistics function is a type of
- 40:06:56sigmoid function that aims to predict
- 40:06:58the class to which a particular sample
- 40:07:01belongs. Its outcome is discrete binary
- 40:07:04value. a probability between zero and
- 40:07:07one. The logistic sigmoid is a useful
- 40:07:10function that follows the yes curve. It
- 40:07:13saturates when the input is very large
- 40:07:15or very small. Logistic sigmoid is
- 40:07:19expressed as sigma of x= 1 upon 1 + e to
- 40:07:23the power minus x.
- 40:07:26The logistic sigmoid can be expressed as
- 40:07:29sigmoid function of x is given as 1 upon
- 40:07:321 + e ^ minus x where e is the ooler's
- 40:07:36number.
- 40:07:38Gshian distribution.
- 40:07:41The gossian distribution is a type of
- 40:07:43distribution in which data tends to
- 40:07:45cluster around a central value with
- 40:07:48little or no bias to the left or right.
- 40:07:52It is often referred to as normal
- 40:07:54distribution.
- 40:07:56In absence of prior information, the
- 40:07:59normal distribution is frequently a fair
- 40:08:01assumption in machine learning
- 40:08:04equation.
- 40:08:06The formula for calculating Gaussian
- 40:08:08distribution is described as the normal
- 40:08:11distribution of X.
- 40:08:14That is the function of x given mean as
- 40:08:16mu and variance is sigma square can be
- 40:08:19calculated as 1 upon sigma square
- 40:08:22roo<unk> of 2 pi e to the power -/ x -
- 40:08:26mood / sigma square
- 40:08:30where mu is the mean or peak value which
- 40:08:33also is the expected value of x.
- 40:08:37Sigma is the standard deviation. Sigma
- 40:08:40square is the variance.
- 40:08:42A standard normal distribution has a
- 40:08:44mean of zero and a standard deviation of
- 40:08:47one.
- 40:08:50Gshian distribution can be univariate
- 40:08:54which describes the distribution of a
- 40:08:56single variable X.
- 40:08:58It can also be multivariate where it can
- 40:09:01just use to describe the distribution of
- 40:09:03several variables.
- 40:09:06It is represented in 3D of ND formats.
- 40:09:12Law of large numbers.
- 40:09:16Now let's talk about law of large
- 40:09:18numbers. The law of large numbers states
- 40:09:21that an observed sample average from a
- 40:09:24large sample will be close to the true
- 40:09:26population average and that it will get
- 40:09:28closer in the larger sample. So the law
- 40:09:32of large number does not guarantee that
- 40:09:34a given sample spatially a small sample
- 40:09:36will reflect the true population
- 40:09:38characteristics or that a sample does
- 40:09:41not reflect the true population will be
- 40:09:43balanced by a subsequent sample. This is
- 40:09:46for the law of large numbers to express
- 40:09:49the relationship between scale and
- 40:09:51growth rate.
- 40:09:54So there are multiple examples through
- 40:09:56which we can understand
- 40:10:00and it is widely used in statistical
- 40:10:02analysis in working with the central
- 40:10:04limit theorem in terms of the business
- 40:10:06growth. So there are multiple real time
- 40:10:09setup in which these are going to be
- 40:10:11used. So if you talk about tossing a
- 40:10:14coin so tossing a coin in a number of
- 40:10:17times will give us two different type of
- 40:10:19outcomes.
- 40:10:22the result will spread evenly between
- 40:10:24head and tails and the expected average
- 40:10:27value is going to be half.
- 40:10:29That means 50 times tails and 30 times
- 40:10:32heads. But again, if you toss a coin
- 40:10:351,000 times, then the result can be in
- 40:10:38different manners because out of 1,000,
- 40:10:41let's say 850 times it has been head and
- 40:10:45only 150 times it has been tails and so
- 40:10:49on. So that's why the possibility of one
- 40:10:51event occurring is going to be changed
- 40:10:54in large sample sets as compared to a
- 40:10:56small sample sets as in let's say 10
- 40:10:59times. So the number of heads and tails
- 40:11:02unbalanced for lower number of trials.
- 40:11:04So we can see it is unbalanced.
- 40:11:08But again as soon as we toss more number
- 40:11:10of coins more leans towards the balance
- 40:11:13value or we can see the observed
- 40:11:15averages.
- 40:11:17Next we have p value.
- 40:11:20So p value is basically a number
- 40:11:23calculated from the statistical test
- 40:11:26that describes how likely we are to have
- 40:11:28found a particular set of observations
- 40:11:30if the null hypothesis were true. So p
- 40:11:34values are used in hypothesis testing to
- 40:11:37help decide whether to reject the null
- 40:11:39hypothesis. And the smaller the p value,
- 40:11:42the more likely we are to reject the
- 40:11:45null hypothesis.
- 40:11:46So we have a term called as null
- 40:11:48hypothesis. So all statistical tests
- 40:11:52they have null hypothesis. So for most
- 40:11:55tests the null hypothesis is that there
- 40:11:57is no relationship between our variables
- 40:12:00of in first or that there is no
- 40:12:02difference among groups. For example in
- 40:12:05a two-tail t test the non-hypothesis is
- 40:12:08that the difference between two groups
- 40:12:10is going to be zero.
- 40:12:13So p value is going to tell us how
- 40:12:15likely it is that our data could have
- 40:12:17occurred under the null hypothesis.
- 40:12:21It is done by calculating the likelihood
- 40:12:23of a test statistic
- 40:12:25which is the number calculated by a
- 40:12:27statistical test using our data. So p
- 40:12:30value tell us how often we would expect
- 40:12:33to see a test statistic as extreme or
- 40:12:36more extreme
- 40:12:37than one calculated by a statistical
- 40:12:40test. if the null hypothesis of the test
- 40:12:43was true.
- 40:12:45So there are multiple limitations as
- 40:12:47well. So first one is the results can be
- 40:12:50significant but again they are they may
- 40:12:53not be practical as we have compared it
- 40:12:56can be based on multiple hypothesis for
- 40:12:58a game for the healthcare test. If the
- 40:13:01test is going to be positive or not it
- 40:13:04may show even values of the effect of a
- 40:13:06variable but not the magnitude in real
- 40:13:09life. What exactly is going to be the
- 40:13:11application of a drug test being failed
- 40:13:14in pharma company? Therefore, it is
- 40:13:17recommended to use confidence and levels
- 40:13:19in addition to the p values to quantify
- 40:13:22or we can say to give a solid figure to
- 40:13:24the reserve which we are going to get.
- 40:13:27The p values they are interpreted as
- 40:13:30supporting or we can say refuting the
- 40:13:32alternative hypothesis.
- 40:13:34So p value can only tell you whether or
- 40:13:37not the null hypothesis is supported. It
- 40:13:40cannot tell us whether our alternative
- 40:13:42hypothesis is true or why. So the risk
- 40:13:46of rejecting the null hypothesis is
- 40:13:49often higher than the p value. So
- 40:13:52especially when we are looking at a
- 40:13:53single study or when using small sample
- 40:13:56sizes. So this is because the smaller
- 40:13:59frame of reference, the greater are the
- 40:14:01chance that as we stumble across a
- 40:14:03statistically significant pattern
- 40:14:06completely by accident.
- 40:14:08Key takeaways.
- 40:14:10Key takeaways. Probability and
- 40:14:13statistics structure the premise of the
- 40:14:15data. The data helps in anticipating the
- 40:14:18future or gauging in view of the past
- 40:14:21patterns of information.
- 40:14:24The central tendency is a single value
- 40:14:27that helps to describe the data by
- 40:14:29identifying these central positions. The
- 40:14:31mean, median, and mode are the measures
- 40:14:34of central tendencies.
- 40:14:37The distribution where the data tends to
- 40:14:39be around a central value with a lack of
- 40:14:42bias or minimal bias towards the left or
- 40:14:45right is called as gshian distribution.
- 40:14:49>> My name is Richard Kersner with the
- 40:14:50simply learn team. That's get certified,
- 40:14:53get ahead. We're going to cover
- 40:14:54mathematics for machine learning. So
- 40:14:57today's agenda is going to cover data
- 40:14:59and its types. Then we're going to dive
- 40:15:01into linear algebra and its concepts,
- 40:15:04calculus, statistics for machine
- 40:15:06learning, probability for machine
- 40:15:08learning, hands-on demos, and of course
- 40:15:12throwing in there in the middle is going
- 40:15:13to be your matrixes and a few other
- 40:15:15things to go along with all this.
- 40:15:18Data and its types. Data denotes the
- 40:15:20individual pieces of factual information
- 40:15:22collected from various sources. It is
- 40:15:25stored, processed and later used for
- 40:15:26analysis.
- 40:15:28And so we see here uh just a huge
- 40:15:30grouping of information, a lot of tech
- 40:15:32stuff, money, dollar signs, numbers
- 40:15:36uh and then you have your performing
- 40:15:38analytics to drive insights and
- 40:15:40hopefully you have a nice share your
- 40:15:41shareholders gathered at the meeting and
- 40:15:43you're able to explain it in something
- 40:15:44they can understand. So we talk about
- 40:15:47datas types of data we have in our types
- 40:15:50of data we have a qualitative
- 40:15:52categorical
- 40:15:54you think nominal or ordinal and then
- 40:15:56you have your quantitative or numerical
- 40:15:58which is discrete or continuous
- 40:16:01and let's look a little closer at those
- 40:16:03data type vocabulary always people's
- 40:16:06favorite is the vocabulary words okay
- 40:16:09not mine uh but let's dive into this
- 40:16:11what we mean by nominal nominal they are
- 40:16:14used to label various just uh label our
- 40:16:17variables without providing any
- 40:16:19measurable value. Uh country, gender,
- 40:16:22race, hair, color, etc. It's something
- 40:16:26that you either mark true or false. This
- 40:16:28is a label. It's on or off. Either they
- 40:16:30have a red hat on or they do not. Uh so
- 40:16:33a lot of times when you're thinking
- 40:16:34nominal data labels, uh think of it as a
- 40:16:38true false kind of setup. And we look at
- 40:16:40ordinal. This is categorical data with a
- 40:16:42set order or a scale to it. Uh and you
- 40:16:45can think of salary range is a great
- 40:16:47one. Uh movie ratings etc. You see here
- 40:16:50the salary range if you have 10,000 to
- 40:16:5220,000 number of employees earning that
- 40:16:55rate is 150 20,000 to 30,000 100 and so
- 40:16:59forth. Some of the terms you'll hear is
- 40:17:02bucket. Uh this is where you have 10
- 40:17:04different buckets and you want to
- 40:17:05separate it into something that makes
- 40:17:07sense into those 10 buckets. And so when
- 40:17:10we start talking about ordinal, a lot of
- 40:17:12times when you get down to the brass
- 40:17:14bones, again, we're talking true false.
- 40:17:17Uh so if you're a member of the 10 to
- 40:17:1820k range, uh so forth, those would each
- 40:17:22be either part of that group or you're
- 40:17:24not. But now we're talking about buckets
- 40:17:26and we want to count how many people are
- 40:17:27in that bucket. Quantitative numerical
- 40:17:30data uh falls into two classes, discrete
- 40:17:34or continuous. And so data with a final
- 40:17:37set of values which can be categorized
- 40:17:39class strength questions answered
- 40:17:42correctly and runs hit in cricket. A lot
- 40:17:45of times when you see this you can think
- 40:17:47integer uh and a very restricted integer
- 40:17:50i.e. you can only have 100 questions um
- 40:17:53on a test. So you can it's very
- 40:17:55discreet. I only have a 100 different
- 40:17:56values that it can attain. So think
- 40:17:59usually you're talking about integers
- 40:18:01but within a very small range. They
- 40:18:03don't have an open end or anything like
- 40:18:04that.
- 40:18:05Uh so discrete is very solid, simple to
- 40:18:08count, set number. Continuous on the
- 40:18:12other hand uh continuous data can take
- 40:18:14any numerical value within a range. So
- 40:18:17water pressure, weight of a person etc.
- 40:18:20Usually we start thinking about float
- 40:18:21values where they can get phenomenally
- 40:18:24small in their in what they're worth.
- 40:18:26And there's a whole series of values
- 40:18:28that falls right between discrete and
- 40:18:30continuous. Um you can think of the
- 40:18:32stock market. You have dollar amounts.
- 40:18:34It's still discreet, but it starts to
- 40:18:37get complicated enough when you have
- 40:18:39like, you know, jump in the stock market
- 40:18:40from $525.33
- 40:18:44to $580.67.
- 40:18:48There's a lot of point values in there.
- 40:18:49It'd still be called discreet, but you
- 40:18:52start looking at it as almost continuous
- 40:18:54because it does have such a variance in
- 40:18:56it. Now uh we talk about n we did we
- 40:18:59went over nominal and ordinal uh almost
- 40:19:01true false charts and we looked at
- 40:19:04quantitative and numerical data which
- 40:19:06we're starting to get into numbers.
- 40:19:08Discrete you can usually a lot of times
- 40:19:10discreet will be put into it could be
- 40:19:12put into true false but usually it's
- 40:19:14not. Uh so we want to address this stuff
- 40:19:15and the first thing we want to look at
- 40:19:17is the very basic which is your algebra.
- 40:19:19So we're going to take a look at linear
- 40:19:21algebra. You can remember back when your
- 40:19:24uklidian geometry uh we have a line.
- 40:19:27Well, let's go through this. We have
- 40:19:28linear algebra is the domain of
- 40:19:30mathematics concerning linear equations
- 40:19:33and their representations in vector
- 40:19:35spaces and through matrices. I told you
- 40:19:37we're going to talk about matrices. Uh
- 40:19:40so a linear equation is simply um uh 2x
- 40:19:44+ 4 y - 3 z = 10. Very linear. 10 x +
- 40:19:5012.4 4 y = z. And now you can actually
- 40:19:53solve these two equations by combining
- 40:19:55them. Uh, and that's we're talking about
- 40:19:57a linear equation.
- 40:19:59In the vectors, we have a + b= c. Now,
- 40:20:03we're starting to look at a direction.
- 40:20:05And these values usually think of an xyz
- 40:20:08plot. Um, so each one is a direction.
- 40:20:11And the actual distance of like a
- 40:20:14triangle A is C. And then your matrix
- 40:20:17can describe all kinds of things. Um, I
- 40:20:20find matrixes uh confuse a lot of
- 40:20:22people, not because they're particularly
- 40:20:25difficult, but because of the magnitude
- 40:20:28and the different things are used for.
- 40:20:31And a matrix is a chart or a um, you
- 40:20:35know, think of a spreadsheet, but you
- 40:20:36have your rows and your columns. And
- 40:20:39you'll see here we have a * b= c. Very
- 40:20:43important to know your counts. Uh, so
- 40:20:47depending on how the math is being done,
- 40:20:48what you're using it for, making sure
- 40:20:50you have the same rows and the number of
- 40:20:52columns or a single number, there's all
- 40:20:54kinds of things that play in that that
- 40:20:56can make matrixes confusing. Uh, but
- 40:20:58really it has a lot more to do with what
- 40:21:00domain you're working in. Uh, are you
- 40:21:02adding in multiple polomials where you
- 40:21:05have like uh uh ax^2 plus b y plus, you
- 40:21:10know, you start to see that can be very
- 40:21:12confusing versus a very straightforward
- 40:21:14matrix. And let's just go a little
- 40:21:16deeper into these because these are such
- 40:21:18primary this is what we're here to talk
- 40:21:20about is these different math uh
- 40:21:22mathematical computations that come up.
- 40:21:25So we're looking at linear equations.
- 40:21:26Let's dig deeper into that one. An
- 40:21:28equation having a maximum order of one
- 40:21:30is called a linear equation. Uh so it's
- 40:21:33linear because when you look at this we
- 40:21:35have uh ax plus b= c which is a one
- 40:21:38variable. We have two variable ax plus b
- 40:21:42y = c ax plus b y plus z c cz z= d and
- 40:21:46so forth. But all of these are to the
- 40:21:50power of one. You don't see x squar. You
- 40:21:52don't see x cubed. So we're talking
- 40:21:54about linear equations. That's what
- 40:21:55we're talking about in their addition.
- 40:21:57If you have already dived into say
- 40:22:00neural networks, you should recognize
- 40:22:02this ax plus b y plus cz um setup plus
- 40:22:06the intercept. uh which is basically
- 40:22:08your your neural network each node
- 40:22:10adding up all the different inputs and
- 40:22:13we can drill down into that most common
- 40:22:15formula is your y = mx + c.
- 40:22:20So you have your uh y equals the m which
- 40:22:24is your slope, your x value plus c which
- 40:22:28is your um y intercept. They kind of
- 40:22:31labeled it wrong here
- 40:22:33threw me for a loop but the the c would
- 40:22:35be your y intercept. So when you set x
- 40:22:37equal to zero, y equals c. And that's
- 40:22:40that's your y intercept right there. Uh
- 40:22:43and that's they they just had reversed
- 40:22:45value of y. When x equals 0, it equals
- 40:22:48the y intercept, which is c. And your
- 40:22:50slope gradient line, which is your m. So
- 40:22:52you get your y = 2x + 3. And there's
- 40:22:56lots of easy ways to compute this. This
- 40:22:58why this is why we always start with the
- 40:23:00most basic one when we're solving one of
- 40:23:01these problems. And then of course the
- 40:23:03um one of the most important takeaways
- 40:23:05is the slope gradient of the line. Uh so
- 40:23:08the slope is very important that m
- 40:23:10value. Uh in this case we went ahead and
- 40:23:12solved this. If you have y = 2x + 3 you
- 40:23:16can see how it has a nice line graph
- 40:23:18here on the right.
- 40:23:20So matrixes a matrix refers to a
- 40:23:23rectangular representation of an array
- 40:23:25of numbers arranged in columns and rows.
- 40:23:29So we're talking m rows by n columns
- 40:23:31here. A11 is denotes the element of the
- 40:23:34first row in the first column. Similarly
- 40:23:37a12 and it's really pronounced a11 in
- 40:23:40this particular setup. So it's row one
- 40:23:43column one. A12 is a of row one column 2
- 40:23:48uh first row and second column and so
- 40:23:50on.
- 40:23:52And there's a lot of ways to denote
- 40:23:53this. I've seen these as like a capital
- 40:23:56letter A, smaller case A for the top row
- 40:23:58or I mean you can see where they can go
- 40:24:01all kinds of different directions as far
- 40:24:02as the value. You just take a moment to
- 40:24:05realize there's need to be some
- 40:24:06designation as far as what row it's in
- 40:24:09and what column it's in. And we have our
- 40:24:12uh basic operations. We have addition.
- 40:24:14So when you think about addition, you
- 40:24:16have uh two matrices of 2x two and you
- 40:24:20just add each individual number in that
- 40:24:23matrix and then when you get to the
- 40:24:25bottom you have uh in this case the
- 40:24:27solution is 12, 10 + 2 is 12, 5 + 3 is 8
- 40:24:30and so on. And the same thing with
- 40:24:32subtraction.
- 40:24:34Now again you're counting matrices you
- 40:24:37want to check your um dimensions of the
- 40:24:39matrix the shape you'll see shape come
- 40:24:42up a lot in programming. So we're
- 40:24:44talking about dimensions we're talking
- 40:24:46about the shape. If the two shapes are
- 40:24:48equal this is what happens when you add
- 40:24:51them together or subtract them. And we
- 40:24:54have multiplication. When you look at
- 40:24:56the multiplication you end up with a
- 40:24:57very slightly different setup going.
- 40:25:00Now, if we look at our last one, we're
- 40:25:03um uh we're like, why? This always gets
- 40:25:06to me when we get to matrices. They
- 40:25:08don't really say why you multiply
- 40:25:10matrices. Um you know, my first thought
- 40:25:12is 1 * 2, 4 * 3. But if you look at
- 40:25:15this, we get 1 * 2 + 4 * 3, 1 * 3 + 4 *
- 40:25:205,
- 40:25:22uh 6 * 2 + 3 * 3, 6 * 3 + 3 * 5. If
- 40:25:27you're looking at these matrices, uh,
- 40:25:29think of this more as an equation. And
- 40:25:32so we have, uh, if you remember when we
- 40:25:33back up here for our multiple line
- 40:25:35equations, let's just go back up a
- 40:25:37couple slides where we were looking at,
- 40:25:39uh, two variable. So this is a two
- 40:25:41variable equation. ax plus b y= c.
- 40:25:45Um, and this is a way to make it very
- 40:25:48quick to solve these variables. And
- 40:25:50that's why you have the matrix, and
- 40:25:51that's why you do
- 40:25:53the multiplication the way they do. And
- 40:25:56this is the dotproduct of uh 1 * 2 + 4 *
- 40:26:013
- 40:26:031 * 3 + 4 * 5
- 40:26:07uh 6 * 2 + 3 * 3 6 * 3 + 3 * 5. And it
- 40:26:13gives us a nice little 14, 23, 21, and
- 40:26:1633 over here, which then can be used and
- 40:26:19reduced down to a simple um formula as
- 40:26:23far as solving the variables as you have
- 40:26:25enough inputs. Uh and then in matrix
- 40:26:27operations, when you're dealing with a
- 40:26:29lot of matrices, uh now keep in mind
- 40:26:32multiplying matrices is different than
- 40:26:34finding the product of two matrices.
- 40:26:36Okay? So we're talking about
- 40:26:37multiplication, we're talking about
- 40:26:39solving uh for equations. When you're
- 40:26:42finding the product, you are just
- 40:26:43finding one time two. Keep that in mind
- 40:26:45because that does come up. I've had that
- 40:26:47come up a number of times where I am
- 40:26:49altering data and I get confused as to
- 40:26:51what I'm doing with it. Uh transpose
- 40:26:54flipping the matrix over it's diagonal.
- 40:26:56Comes up all the time where you have you
- 40:26:58still have 12, but instead of it being
- 40:27:00uh 128, it's now 1214 821. You're just
- 40:27:05flipping the columns and the rows. Uh
- 40:27:07and then of course you can do an inverse
- 40:27:09um changing the signs of the values
- 40:27:11across this main diagonal. And you can
- 40:27:13see here we have the inverse a to the
- 40:27:15minus1 and ends up with uh instead of 12
- 40:27:188 14 12 it's now -22 -12 vectors uh
- 40:27:23vector just means we have
- 40:27:26a value and a direction and we have down
- 40:27:30four numbers here on our vector.
- 40:27:33uh in mathematics a one-dimensional
- 40:27:35matrix is called a vector. Uh so if you
- 40:27:39have your xplot and you have a single
- 40:27:41value that values along the x- axis and
- 40:27:44it's a single dimension. If you have two
- 40:27:46dimensions you can think about putting
- 40:27:48them on a graph. You might have x and
- 40:27:50you might have y and each value denotes
- 40:27:53a direction. And then of course the
- 40:27:55actual distance is going to be the
- 40:27:57hypothesis of that triangle. Uh and you
- 40:27:59can do that with three dimensionals x y
- 40:28:01and z. uh and you can do it all the way
- 40:28:03to nth dimensions. So when they talk
- 40:28:06about the k means uh for categorizing
- 40:28:09and how close data is together, they
- 40:28:12will compute that based on the
- 40:28:14Pythagorean theorem. So you would take
- 40:28:16uh the square of each value, add them
- 40:28:18all together and find the square root
- 40:28:20and that gives you a distance as far as
- 40:28:22where that point is, where that vector
- 40:28:24exists or an actual point value. And
- 40:28:26then you can compare that point value to
- 40:28:29another one and it makes a very easy
- 40:28:31comparison versus comparing uh 50 or 60
- 40:28:34different numbers. And that brings us up
- 40:28:36to gene vectors and I gene values. Uh I
- 40:28:41gene vectors the vectors that don't
- 40:28:43change their span while transformation
- 40:28:46and I gene values the scalar values that
- 40:28:49are associated to the vectors.
- 40:28:52Conceptually you can think of the vector
- 40:28:54as your picture. you have a picture.
- 40:28:56It's um uh two dimensions x and y. And
- 40:29:00so when you do those two dimensions and
- 40:29:02those two values or whatever that value
- 40:29:04is um that is that point but the values
- 40:29:09change when you skew it and so if we
- 40:29:12take and we have a vector a and that's a
- 40:29:16set value uh b is um your is your you
- 40:29:20have a and b which is your hygiene
- 40:29:21vector. Two is the i gene value. So,
- 40:29:25we're altering all the values by two.
- 40:29:28That means we're u maybe we're
- 40:29:30stretching it out one direction, making
- 40:29:31it tall if you're doing picture editing.
- 40:29:34Um that that's one of the places this
- 40:29:36comes in. But you can see when you're
- 40:29:38transforming uh your different
- 40:29:40information, how you transform it is
- 40:29:43then your hygiene value. And you can see
- 40:29:45here uh vector after line transition
- 40:29:49uh we have 3 a is the hygiene vector.
- 40:29:52Three is the hygiene value. So A doesn't
- 40:29:55change. That's whatever we started with.
- 40:29:57That's your original picture. And three
- 40:29:59uh is skewing it one direction and maybe
- 40:30:02uh B is being skewed another direction.
- 40:30:05And so you have a nice tilted picture
- 40:30:06because you've altered it by those by
- 40:30:08the hygiene values.
- 40:30:10So let's go ahead and pull up a demo on
- 40:30:13linear algebra. And to do this, I'm
- 40:30:16going to go through my trusted Anaconda
- 40:30:19into my Jupiter notebook. and we'll
- 40:30:22create a new uh notebook called linear
- 40:30:25algebra. Since we are working in Python,
- 40:30:28uh we're going to use our numpy. I
- 40:30:30always import that as np or numpy array.
- 40:30:33Probably the most popular um module for
- 40:30:36doing matrixes and things in
- 40:30:39given that this is part of a series. I'm
- 40:30:41not going to go too much into numpy. Uh
- 40:30:43we are going to go ahead and create two
- 40:30:45different variables. A for a numpy array
- 40:30:4710 15 and b 29.
- 40:30:51We'll go ahead and run this. And you can
- 40:30:52see there's our two arrays 105 29. And I
- 40:30:55went ahead and added a space there in
- 40:30:57between so it's easier to read. And
- 40:31:00since it's the last line, we don't have
- 40:31:02to put the print statement on it unless
- 40:31:04you want. We can simp but we can simply
- 40:31:06do a plus b. So when I run this, uh, we
- 40:31:10have 10 15 29 and we get 30 24, which is
- 40:31:15what you expect. 10 + 20 15 + 9. You
- 40:31:19could almost look at this addition as
- 40:31:21being um
- 40:31:24just adding up the columns on here
- 40:31:26coming down. And if we wanted to do it a
- 40:31:28different way, we could also do a t plus
- 40:31:32b dot t. Remember that t flips them. And
- 40:31:35so if we do that, we now get them uh we
- 40:31:39now have 304 going the other way. We
- 40:31:42could also do something kind of fun.
- 40:31:44There's a lot of different ways to do
- 40:31:45this. Uh, as far as a plus b, I can also
- 40:31:49do a plus b. T and you're going to see
- 40:31:53that that will come out the same. The 30
- 40:31:5524 whether I transpose a and b or
- 40:31:57transpose them both at the end.
- 40:32:01And likewise, we can very easily
- 40:32:03subtract two vectors. I can go a minus
- 40:32:06b. And we run that and we get - 106. Now
- 40:32:11remember, this is the last line in this
- 40:32:13particular section. That's why I don't
- 40:32:14have to put the print around it. Um, and
- 40:32:17just like we did before, we can
- 40:32:20transpose either the individual or we
- 40:32:22can transpose the main setup and then we
- 40:32:25get a minus 106 going the other way.
- 40:32:30Now, we didn't mention this in our
- 40:32:32notes, but you can also do a scalar
- 40:32:35multiplication.
- 40:32:37Let me just put down scaler so you can
- 40:32:38remember that. Uh what we're talking
- 40:32:41about here is I have uh this array here
- 40:32:45u and if I go a time u uh we'll take the
- 40:32:51value two we'll multiply it by every
- 40:32:52value in here. So 2 * 30 is 60 2 * 15
- 40:32:58and just like we did before
- 40:33:01um this happens a lot because when
- 40:33:03you're doing matrices you do need to
- 40:33:04flip them you get 6030 coming this way.
- 40:33:08So in numpy uh we have what they call
- 40:33:11dotproduct
- 40:33:14and uh what this this in a
- 40:33:16twodimensional vectors it is the
- 40:33:18equivalent of two matrix multiplication
- 40:33:21and remember we were talking about
- 40:33:22matrix multiplication
- 40:33:24uh where it is the well let's walk
- 40:33:27through it
- 40:33:30we'll go ahead and start by defining two
- 40:33:32um numpy arrays we'll have uh 10 20 256
- 40:33:37or our u and our E uh and then we're
- 40:33:39going to go ahead and do if we take
- 40:33:43the values uh and if you remember
- 40:33:45correctly
- 40:33:47an array like this would be 10 * 25 + 20
- 40:33:52* 6. We'll go ahead and uh print that.
- 40:34:00There we go.
- 40:34:02And then we'll go ahead and do the uh np
- 40:34:05dot of u comma
- 40:34:10v.
- 40:34:12And we'll find when we do this, we go
- 40:34:14and run this uh we're going to get uh
- 40:34:17370
- 40:34:18370.
- 40:34:20So this is a strain multiplication where
- 40:34:22they use it to solve uh linear algebra
- 40:34:26uh when you have multiple numbers going
- 40:34:28across. And so this could be very
- 40:34:30complicated. We could have a whole
- 40:34:31string of different variables going in
- 40:34:33here. But for this we get a nice uh
- 40:34:35value for our dot multiplication
- 40:34:39and we did um addition earlier which was
- 40:34:42just your basic addition. Uh and of
- 40:34:44course a matrix you can get very
- 40:34:46complicated on these or in this case
- 40:34:48we'll go ahead and do um let's create
- 40:34:51two complex matrixes.
- 40:34:55This one is a matrix of um you know 1210
- 40:34:5946 431. We'll just print out A so you
- 40:35:02can see what that looks like. Here's
- 40:35:04print A.
- 40:35:06We print A out. You can see that we have
- 40:35:09a um 2x3
- 40:35:13layer matrix for A. And we can also put
- 40:35:16together always kind of fun when you're
- 40:35:18playing with print values. Uh we could
- 40:35:20do something like this. We could go in
- 40:35:22here. There we go. Uh, we could print a.
- 40:35:25We have it end with uh equals a run. And
- 40:35:29this kind of gives it a nice look. Uh,
- 40:35:31here's your matrix. That's all this is.
- 40:35:33Comma, n means it just tags it on the
- 40:35:35end. That's all all that is doing on
- 40:35:37there. And then we can simply add in
- 40:35:39what is a plus b. And you should already
- 40:35:42guess because this is the same as what
- 40:35:43we did before. There's no difference.
- 40:35:45Uh, we do a simple vector addition. We
- 40:35:47have 12 + 2 is 14, 10 + 8 is 18. And so
- 40:35:51on. And just like we did the uh matrix
- 40:35:54addition, we can also do a minus b and
- 40:35:58do our matrix subtraction.
- 40:36:01And we look at this uh we have what? 12
- 40:36:03- 2 is 10. 10 - 8 um where are we?
- 40:36:11Oh, there we go. 8 min
- 40:36:15confusing what I'm looking at. I should
- 40:36:16have reprinted out the original numbers.
- 40:36:18Uh but we can see here 12 - 2 is of
- 40:36:21course 10. 10 - 8 is 2. Uh 4 - 46 is -
- 40:36:2542 and so forth. So same as a
- 40:36:28subtraction as before, we just call it
- 40:36:30matrix subtraction. It's identical.
- 40:36:33Now if you remember up here, we had
- 40:36:35scalar addition where we're adding just
- 40:36:37one number to a matrix. You can also do
- 40:36:41scalar multiplication. Uh and so simply
- 40:36:44if you have a single value A and you
- 40:36:46have B which is your array, we can also
- 40:36:48do A * B. When we run that, uh, you can
- 40:36:53see here we have 2 * 4 is 8. Uh, 5 * 4
- 40:36:57is 20 and so forth. You're just
- 40:36:59multiplying the four across each one of
- 40:37:01these values. And this is an interesting
- 40:37:03one that comes up. A little bit of a
- 40:37:06brain teaser is matrix and vector
- 40:37:09multiplication.
- 40:37:11And so when we're looking at this,
- 40:37:14uh, we are just do a regular arrays. It
- 40:37:17doesn't necessarily have to be a numpy
- 40:37:18array. We have a
- 40:37:21which has our um array of arrays and b
- 40:37:25which is a single array and so we can
- 40:37:28from here
- 40:37:30do the dot
- 40:37:33a b and this is going to return two
- 40:37:36values and the first value is that it's
- 40:37:39you could say it's like uh um we're
- 40:37:41doing the this array b array first with
- 40:37:45a and then with a second one and so it
- 40:37:47splits it up so you have a matrix of
- 40:37:49vector multiplication and you mix and
- 40:37:50match. When you get into really
- 40:37:52complicated uh backend stuff, this
- 40:37:54becomes more common because you're now
- 40:37:56you got layers upon layers of data and
- 40:37:59so you you'll end up with a matrix and a
- 40:38:02set of uh vector matrices. Do you want
- 40:38:04to multiply?
- 40:38:06Now, keep in mind that if you're doing
- 40:38:09data science, a lot of times you're not
- 40:38:11looking at this. This is what's going on
- 40:38:12behind the scenes. So if you're in um
- 40:38:15the scikit looking at sklearn where
- 40:38:17you're doing linear regression models,
- 40:38:20this is some of the math that's hidden
- 40:38:21behind the scenes that's going on. Other
- 40:38:24times you might find yourself having to
- 40:38:26do part of this and manipulate the data
- 40:38:28around so it fits right and then you go
- 40:38:30back in and you run it through the
- 40:38:31scikit. And if we can do um up here
- 40:38:36where we did a uh matrix and vector
- 40:38:39multiplication, we can also do matrix to
- 40:38:41matrix multiplication. And if we run
- 40:38:43this where we have the two matrices, uh
- 40:38:45you can see we have a very complicated
- 40:38:47array that of course comes out on there
- 40:38:48for our dot. And just to reiterate it,
- 40:38:52we have our transpose a matrix which is
- 40:38:54your T. And so if we create a matrix A
- 40:38:57and then we do transpose it, you can see
- 40:38:59how it flips it from 5 10 15 20 25 30 to
- 40:39:045 15 25 10 20 30 uh rows and columns.
- 40:39:10And certainly with the math, uh, this
- 40:39:12comes up a lot. Um, it also comes up a
- 40:39:15lot with XY plotting. When you put it
- 40:39:17into piplot, you have one format where
- 40:39:19they're looking at pairs of numbers and
- 40:39:21then they want all of X's and all Y's.
- 40:39:24So, you know, the transpose is an
- 40:39:25important tool both for your math and
- 40:39:27for plotting and all kinds of things.
- 40:39:30Another tool that we didn't discuss uh
- 40:39:32is your identity matrix. Uh and this one
- 40:39:37is more definition.
- 40:39:40Uh the identity matrix. Um we have here
- 40:39:43one where we just did uh two. So it
- 40:39:46comes down as one 0 0 1 uh 1 0 0 1 0. It
- 40:39:51creates a diagonal of one. And what that
- 40:39:53is is when you're doing your identities,
- 40:39:55you could be comparing all your
- 40:39:58different features to the different
- 40:40:00features and how they correlate. And of
- 40:40:02course when you have uh feature one
- 40:40:04compared to feature one to itself it is
- 40:40:06always one uh where usually it's between
- 40:40:10zero one depending on how well
- 40:40:12correlates. So when we're talking about
- 40:40:14identity matrix that's what we're
- 40:40:16talking about right here is that you
- 40:40:18create this preset matrix and then you
- 40:40:20might adjust these numbers depending on
- 40:40:22what you're working with and what the
- 40:40:23domain is. And then another thing we can
- 40:40:26do uh to kind of wrap this up. We'll hit
- 40:40:28you with the most complicated uh um
- 40:40:30piece of this puzzle here is an inverse
- 40:40:34um a matrix. And let's just go ahead and
- 40:40:37put the um it's a lengthy description.
- 40:40:41Let's go and put the description. This
- 40:40:43is straight out of the uh the website
- 40:40:46for um numpy. Uh so given a square
- 40:40:50matrix A, here's our square matrix A,
- 40:40:53which is 2 1 0 0 1 0 1 2 1. Keep in mind
- 40:40:573x3, it's square. It's got to be equal.
- 40:40:59It's going to return the matrix A
- 40:41:02inverse satisfying dot A um A inverse.
- 40:41:06So here's our matrix multiplication.
- 40:41:10Um and then of course it equals the dot
- 40:41:13uh yeah a inverse of a um with an
- 40:41:17identity shape of uh a dotshaped zero.
- 40:41:20This is just reshaping the identity.
- 40:41:23That's a little complicated there. Uh so
- 40:41:25we go and have our here's our array. Uh
- 40:41:27we'll go ahead and run this. And you can
- 40:41:30see what we end up with is we end up
- 40:41:32with uh an array 0.5 minus 0.5 and so
- 40:41:36forth with our 211 going down to 1 0 0 1
- 40:41:400 1 2 1. Um getting into a little deep
- 40:41:44on the math understanding when you need
- 40:41:47this is probably really is is what's
- 40:41:49really important when you're doing data
- 40:41:50science versus uh handwriting this out
- 40:41:54and looking up the math and handwriting
- 40:41:55all the pieces out. you do need to know
- 40:41:57about the linear algorithm inverse of a.
- 40:42:00Uh so if it comes up, you can easily
- 40:42:02pull it up or at least remember where to
- 40:42:03look it up. You took a look at the
- 40:42:06algebra side of it. Let's go ahead and
- 40:42:07take a look at the calculus side of uh
- 40:42:10what's going on here with the machine
- 40:42:11learning. So calculus, oh my goodness,
- 40:42:14and differential equations, you got to
- 40:42:16throw that in there because that's all
- 40:42:18part of the bag of tricks, especially
- 40:42:21when you're doing large neural networks,
- 40:42:23but also comes up in many other areas.
- 40:42:25The good news is most of it's already
- 40:42:27done for you in the back end. Uh so when
- 40:42:29it comes up, you really do need to
- 40:42:30understand from the data science, not
- 40:42:32data analytics. Data analytics means
- 40:42:34you're digging deep into actually
- 40:42:36solving these math equations. U and a
- 40:42:39neural network is just a giant
- 40:42:41differential equation. Uh so we talk
- 40:42:43about calculus uh we're going to go
- 40:42:45ahead and understand it by talking about
- 40:42:49cars versus time and speed. uh so helps
- 40:42:53to calculate the spontaneous rate of
- 40:42:56change.
- 40:42:58Uh so suppose we plot a graph of the
- 40:43:00speed of a car with respect to time. So
- 40:43:02as you can see here going down the
- 40:43:04highway probably merged into the highway
- 40:43:06from an on-ramp. So I had to accelerate
- 40:43:09so my speed went way up uh stuck in
- 40:43:12traffic merged into the traffic. Traffic
- 40:43:15opens up and I accelerate again up to
- 40:43:16the speed limit and u maybe it peters
- 40:43:19off up there. So you can look at this as
- 40:43:22as um the speed versus time. I'm getting
- 40:43:25faster and faster because I'm
- 40:43:26continually accelerating. And if I hit
- 40:43:29the brakes, it go the other way. So the
- 40:43:31rate of change of speed with respect of
- 40:43:33time is nothing but acceleration. How
- 40:43:36fast are we accelerating? The
- 40:43:38acceleration is the area between the
- 40:43:40start point of x and the end point of
- 40:43:42delta x. Uh so we can calculate a simple
- 40:43:46if you had x and delta x we could put a
- 40:43:48line there and that slope of the line is
- 40:43:51our acceleration.
- 40:43:53Now that's pretty easy when you're doing
- 40:43:55linear algebra but I don't want to know
- 40:43:58it just for that line and those two
- 40:44:00points. I want to know it across the
- 40:44:02whole of what I'm working with. That's
- 40:44:04where we get into calculus. So when we
- 40:44:06talk about the distance between x and
- 40:44:08delta x it has to be the smallest
- 40:44:10possible near to zero in order to
- 40:44:12approximate the acceleration.
- 40:44:15Uh so the idea is that instead of I mean
- 40:44:17if you ever did took a basic calculus
- 40:44:19class they would draw bars down here and
- 40:44:22you would divide this area up um let's
- 40:44:25go back up a screen. you divide this
- 40:44:27area of this time period up into maybe
- 40:44:3010 sections and you'd use that and you
- 40:44:32could calculate the acceleration between
- 40:44:34each one of those 10 sections kind of
- 40:44:35thing. Uh and then we just keep making
- 40:44:38that space smaller and smaller until
- 40:44:40delta x is almost uh infantismally
- 40:44:44small. And so we get a function of a uh
- 40:44:48equals a limit as h goes to zero of a
- 40:44:51function of a plus h minus a function of
- 40:44:54a over h. And that is you're computing
- 40:44:57the slope of the line.
- 40:45:00We're just computing that slope under
- 40:45:01smaller and smaller and smaller samples.
- 40:45:04Uh and that's what calculus is. Calculus
- 40:45:06is the integral. You can see down here
- 40:45:08we have our nice uh integral sign. Looks
- 40:45:11like a giant s. And that's what that
- 40:45:13means is that we've taken this down to
- 40:45:16as small as we can for that sampling. Uh
- 40:45:20so we're talking about calculus. Finding
- 40:45:22the area under the slope is the main
- 40:45:24process in the integration. Similar
- 40:45:27small intervals are made of the smallest
- 40:45:29possible length of x plus delta x where
- 40:45:32delta x approaches almost an infantismly
- 40:45:35small space. And then it helps to find
- 40:45:37the overall acceleration by summing up
- 40:45:39all the lengths together. Uh so we're
- 40:45:42summing up all the accelerations from
- 40:45:44the beginning to the end. And so here's
- 40:45:46our integral. we sum of a of x * d ofx =
- 40:45:50a + c. Uh that is our basic calculus
- 40:45:55here. So when we talk about
- 40:45:57multivvariant calculus, uh multivariate
- 40:46:01calculus deals with functions that have
- 40:46:02multiple variables and you can see here
- 40:46:05we start getting into some very
- 40:46:06complicated equations. Um uh change in w
- 40:46:10over change of time equals change of w
- 40:46:13over change of z. the differential of z
- 40:46:16to dx differential of x to dt. It gets
- 40:46:19pretty complicated. Uh and it really
- 40:46:21translates into the multivariate
- 40:46:23integration using double integrals. And
- 40:46:25so you have the the sum of the sum of f
- 40:46:28ofxy of d of a equals the sum from c to
- 40:46:31d and a to b of f ofxy dx dy equals uh
- 40:46:36the sum of a to b sum of c to d of fxy
- 40:46:40dy dx.
- 40:46:42understanding the very specifics of
- 40:46:44everything going on in here and actually
- 40:46:46doing the math is usually calculus one,
- 40:46:50calculus 2, and differential equations.
- 40:46:52Uh so you're talking about three
- 40:46:54fulllength courses to dig into and solve
- 40:46:57these math equations. What we want to
- 40:47:00take from here is we're talking about
- 40:47:01calculus. Uh we're talking about summing
- 40:47:04of all these different slopes. And so
- 40:47:07we're still solving a linear uh
- 40:47:09expression. We're still solving y = mx +
- 40:47:13b, but we're doing this for
- 40:47:15infantismally small x's. And then we
- 40:47:17want to sum them up. That's what this
- 40:47:19integral sign means. The the sum of a of
- 40:47:21x d of x= a plus c.
- 40:47:25And when you see these very complicated
- 40:47:27uh multivariate differentiation using
- 40:47:29the chain rule uh when we come in here
- 40:47:31and we have the change of w to the
- 40:47:34change of t equals the change of w dz uh
- 40:47:37and so forth. That's what's going on
- 40:47:40here. That's what these means. We're
- 40:47:41basically looking for the area under the
- 40:47:43curve which really comes to how is the
- 40:47:46change changing and speed's going up.
- 40:47:49How is that changing? And then you end
- 40:47:51up with a multiple layer. So if I have
- 40:47:53three layers of neural networks, how is
- 40:47:55the third layer changing based on the
- 40:47:57second layer changing which is based on
- 40:47:59the first layer changing? And you get
- 40:48:01the picture here that now we have a very
- 40:48:03complicated uh multivariate integration
- 40:48:06um with integrals.
- 40:48:08The good news is we can solve this uh
- 40:48:11mathematically and that's what we do
- 40:48:13when you do neural networks and reverse
- 40:48:14propagation. Uh so the nice thing is
- 40:48:17that you don't have to solve this on
- 40:48:18paper unless you're a data analysis and
- 40:48:20you're working on the back end of
- 40:48:22integrating these formulas and building
- 40:48:24the script to actually build them. So we
- 40:48:26talk about applications of calculus. Uh
- 40:48:28it provides us the tools to build an
- 40:48:30accurate predictive model. Um so it's
- 40:48:32really behind the scenes we want to
- 40:48:34guess at what the change of the change
- 40:48:36of the change is.
- 40:48:38That's a little goofy. I I know I just
- 40:48:40threw that out there. It's kind of a
- 40:48:41meta term. But if you can guess how
- 40:48:44things are going to change, then you can
- 40:48:46guess what the new numbers are.
- 40:48:48Multivariate calculus explains the
- 40:48:51change in our target variable in
- 40:48:52relation to the rate of change in the
- 40:48:54input variables. So there's our multiple
- 40:48:57variables going in there. If uh one
- 40:49:00variable is changing, how does it affect
- 40:49:01the other variable? And then in gradient
- 40:49:04descent, calculus is used to find the
- 40:49:07local and global maxima. And this is
- 40:49:10really big. Uh we're actually going to
- 40:49:12have a whole section here on gradient
- 40:49:14descent because it is really I mean I
- 40:49:18talked about neural networks and how you
- 40:49:19can see how the different layers go in
- 40:49:21there, but gradient descent is one of
- 40:49:23the most key things for trying to guess
- 40:49:26the best answer to something. So let's
- 40:49:30take a look at the code behind gradient
- 40:49:33descent. And uh before we open up the
- 40:49:36code, let's just do real quick uh
- 40:49:39gradient descent.
- 40:49:42Let's say we have a curve like this. And
- 40:49:44most common is that this is going to
- 40:49:47represent your error. Oops.
- 40:49:51Error. There we go. Error. Ah, hard to
- 40:49:55read there. And I want to make the error
- 40:49:57as low as possible. And so what I'm
- 40:50:00looking at it is I want to find this
- 40:50:02line here which is the minimum value. So
- 40:50:07we're looking for the minimum and it
- 40:50:10does that by uh sampling there and then
- 40:50:14it based on this it guesses it might be
- 40:50:16someplace here and it goes hey this is
- 40:50:18still going down. It goes here and then
- 40:50:21goes back over here and then goes a
- 40:50:23little bit closer and it's just playing
- 40:50:25a high low until it gets to that spot,
- 40:50:28that bottom spot. And so we want to
- 40:50:30minimize the error in uh on the flip
- 40:50:34note, you could also want to be
- 40:50:37maximizing something. You want to get
- 40:50:38the best output of it. Uh that's simply
- 40:50:41uh minus the value. Uh so if you're
- 40:50:44looking for where the peak is, this is
- 40:50:46the same as a negative for where the
- 40:50:49valley is and looking for that valley.
- 40:50:52Uh that's all that is and this is a way
- 40:50:53of finding it. So the cool thing is um
- 40:50:57all the heavy lifting's done. Um I
- 40:50:59actually ended up putting together one
- 40:51:01of these a while back as uh when I
- 40:51:03didn't know about sidekick and I was
- 40:51:05just starting. Boy, it's a long while
- 40:51:08back and uh is playing high low. How do
- 40:51:11you play high low? not get stuck in the
- 40:51:13valleys, uh, figure out these curves and
- 40:51:16things like that. Well, you do that and
- 40:51:18the back end is all the calculus and
- 40:51:19differential equations to calculate this
- 40:51:21out. The good news is you don't have to
- 40:51:24do those. Uh, so instead, we're going to
- 40:51:27put together the code and let's go ahead
- 40:51:31and see what we can do with that.
- 40:51:36So, uh, guys in the back put together a
- 40:51:38nice little piece of code here, which is
- 40:51:40kind of fun.
- 40:51:41uh some things we're going to note and
- 40:51:43this is this is really important stuff
- 40:51:45because when you start doing your data
- 40:51:47science and digging into your machine
- 40:51:49learning models uh you're going to find
- 40:51:52these things are stumbling blocks. Uh
- 40:51:54the first one is current x. Where do we
- 40:51:56start at? Uh keep in mind your model
- 40:52:00that you're working with is very
- 40:52:02generic. So whatever you use to minimize
- 40:52:04it the first question is where do we
- 40:52:06start? Um, and we started at this cuz
- 40:52:09the algorithm starts at x= 3. So, we
- 40:52:12arbitrarily picked five. Learning rate
- 40:52:15is uh how many bars to skip going one
- 40:52:17way or the other. Uh, I'm in fact, I'm
- 40:52:19going to separate that a little bit
- 40:52:20because these two are really important.
- 40:52:22Um, if we're dealing with something like
- 40:52:24this where we're talking about um uh
- 40:52:26well, here's our here's the function
- 40:52:28we're going to use our um gradient of
- 40:52:31our function um 2 * x + 5. Keep it
- 40:52:34simple. So that's a function we're going
- 40:52:36to work with. So if I'm dealing with
- 40:52:38increments of a th00and 0.1 is going to
- 40:52:41be a very long time. And if I'm dealing
- 40:52:44with increments of 0.001,
- 40:52:47uh 0.1 is going to skip over my answer.
- 40:52:50So I won't get a very good answer. Um
- 40:52:52and then we look at precision. This
- 40:52:54tells us when to stop the algorithm. So
- 40:52:56again, very specific to what you're
- 40:52:59working on. uh if you're working with
- 40:53:01money and you don't convert it into a
- 40:53:05float value uh you might be dealing with
- 40:53:080.01 which is a penny that might be your
- 40:53:11precision you're working with. Um and
- 40:53:14then of course the previous step size
- 40:53:16max iterations uh we want something to
- 40:53:18cut out at a certain point. Usually
- 40:53:20that's built into a lot of minimization
- 40:53:22functions. And then here's our actual uh
- 40:53:25formula we're going to be working with.
- 40:53:27And then we come in, we go while
- 40:53:29previous step size is greater than
- 40:53:31precision and its is less than max its
- 40:53:36say that 10 times fast. Um
- 40:53:40we're just saying if it's uh if we're if
- 40:53:42we're still greater than our precision
- 40:53:43level, we still got to keep digging
- 40:53:45deeper. Um and then we also don't want
- 40:53:47to go past a thou or whatever this is, a
- 40:53:50million or 10,000 uh running. That's
- 40:53:52actually pretty high. um almost never do
- 40:53:54max iterations more than like 100 or
- 40:53:57200. Rare occasions you might go up to
- 40:53:59four or 500 if it's depending on the
- 40:54:01problem you're working with. Uh so we
- 40:54:04have our previous equals our current.
- 40:54:06That way we can track timewise.
- 40:54:08Uh the current now equals the current
- 40:54:10minus the rate times the formula of our
- 40:54:13previous x. So now we've generated our
- 40:54:16new version. Uh previous step size
- 40:54:19equals the absolute current previous.
- 40:54:22Uh, so we're looking for the change in x
- 40:54:25itters equals iterations + one. That's
- 40:54:27so we know to stop if we get too far.
- 40:54:30And then we're just going to print the
- 40:54:31local minimum occurs at x on here. And
- 40:54:35if we go ahead and run this,
- 40:54:37uh, you can see right here it gets down
- 40:54:39to this point and it says, hey, um,
- 40:54:43local minimum is minus 3.3222
- 40:54:46for this particular series we created.
- 40:54:49Uh, and this is created off of our
- 40:54:50formula here. lambda x2 * x + 5. Now,
- 40:54:55when I'm running this stuff, uh you'll
- 40:54:57see this come up a lot
- 40:55:00and uh with the sklearn kit and and one
- 40:55:04of the nice reasons of breaking this
- 40:55:05down the way we did is I could go over
- 40:55:08those top pieces. Uh those top pieces
- 40:55:10are everything when you start looking at
- 40:55:12these minimization toolkits in built-in
- 40:55:15code. And so from um we'll just do it's
- 40:55:19actually docs.cipi.org
- 40:55:26and we're looking at the scikit. There
- 40:55:30we go. Um optimize minimize. You can
- 40:55:34only minimize one value. You have the
- 40:55:36function that's going in. This function
- 40:55:38can be very complicated. Uh so we used a
- 40:55:41very simple function up here. It could
- 40:55:43be there's all kinds of things that
- 40:55:46could be on there. And there's a number
- 40:55:47of methods to solve this as far as how
- 40:55:49they shrink down. Uh and your x knot.
- 40:55:52There's your there's your start value.
- 40:55:54So your function, your start value. Um
- 40:55:57there's all kinds of things that come in
- 40:55:58here that we can look at which we're not
- 40:56:00going to. Um optimization automatically
- 40:56:03creates constraints bounds. Some of this
- 40:56:06it does automatically, but you really
- 40:56:08the big thing I want to point out here
- 40:56:10is you need to have a starting point.
- 40:56:12You want to start with something that
- 40:56:13you already know is mostly the answer.
- 40:56:15Uh if you don't, then it's going to have
- 40:56:17a heck of a time trying to calculate it
- 40:56:18out.
- 40:56:20Or you can write your own little script
- 40:56:21that does this and and does a high low
- 40:56:23guessing and tries to find the max
- 40:56:25value. That brings us to statistics.
- 40:56:29What this is kind of all about is
- 40:56:30figuring things out. Lot of vocabulary
- 40:56:33and statistics. Uh so statistics, well,
- 40:56:36I guess it's all relative. It's
- 40:56:38definitely not an ed class. Uh so a
- 40:56:40bunch of stuff going on. Statistics.
- 40:56:42Statistics concerns with the collection,
- 40:56:45organization, analysis, interpretation,
- 40:56:48and presentation of data. That is a
- 40:56:52mouthful. Um so we have from end to end
- 40:56:56we're
- 40:56:58valid, what does it mean? How do we
- 40:57:00organize it? Um how do we analyze it?
- 40:57:03Then you got to take those analysis and
- 40:57:04interpret it into something that uh
- 40:57:06people can use. kind of reduce it to
- 40:57:09understandable. Um, and nowadays you
- 40:57:11have to be able to present it. If you
- 40:57:13can't present it, then no one else is
- 40:57:14going to understand what the heck you
- 40:57:15did.
- 40:57:17So, we look at the terminologies. Uh,
- 40:57:20there is a lot of terminologies
- 40:57:22depending on what domain you're working
- 40:57:24in. So clearly if you're working in um a
- 40:57:28domain that deals with
- 40:57:31viruses and tea cells and and how does
- 40:57:36you know where does that come from and
- 40:57:37you're studying the different people
- 40:57:38then you're going to have a population.
- 40:57:40if you are working with um mechanical
- 40:57:44gear um you know a little bit different
- 40:57:46if you're looking for the wobbling
- 40:57:48statistics uh to know when to replace a
- 40:57:51rotor on a machine or something like
- 40:57:52that uh that can be a big deal. You
- 40:57:54know, we have these huge fans that turn
- 40:57:57in our sewage processing systems. And so
- 40:58:01those fans, they start to wobble and hum
- 40:58:03and do different things that the sensors
- 40:58:05pick up. At one point, do you replace
- 40:58:07them? Instead of waiting for it to
- 40:58:08break, in which case it cost a lot of
- 40:58:10money. Instead of replacing a bushing,
- 40:58:11you're replacing the whole fan unit. Uh
- 40:58:14an interesting project that came up for
- 40:58:16our city a while back. Uh so population,
- 40:58:19all objects are measurements whose
- 40:58:21properties are being observed. Uh so
- 40:58:24that's your population all the objects.
- 40:58:27It's easy to see it with people because
- 40:58:28we have our population and large. Um but
- 40:58:32in the case of the sewer fans we're
- 40:58:34talking about how the fan units. That's
- 40:58:36the population of fans that we're
- 40:58:37working with.
- 40:58:39You have a parameter a matrix uh that is
- 40:58:42used to represent a population or
- 40:58:44characteristic.
- 40:58:45You have your sample a subset of the
- 40:58:48population studied. You don't want to do
- 40:58:50them all because then you don't have a
- 40:58:51if you come up with a conclusion for
- 40:58:53everyone, you don't have a way of
- 40:58:55testing it. So you take a sample. Uh
- 40:58:57sometimes you don't have a choice. You
- 40:58:58can only take a sample of what's going
- 40:58:59on. You can't u study the whole
- 40:59:02population. And a variable, a metric of
- 40:59:05interest for each person or object in a
- 40:59:07population.
- 40:59:09Types of sampling. We have probabilistic
- 40:59:12approach. uh selecting samples from a
- 40:59:14larger population using a method based
- 40:59:17on the theory of probability
- 40:59:20and we'll go into a little bit more
- 40:59:21deeper on these. We have random
- 40:59:23systematic stratified and then you have
- 40:59:26nonprobabilistic approach selecting
- 40:59:28samples based on the subjective judgment
- 40:59:30of the researcher rather than random
- 40:59:33selection. Uh it has to do with
- 40:59:35convenience trying to reach a quota um
- 40:59:37or snowball. Uh and they're very biased.
- 40:59:41That's one of the reasons you'll see
- 40:59:42this big stamp on it says biased. Uh so
- 40:59:44you got to be very careful on that. So
- 40:59:47probabilistic sampling uh when we talk
- 40:59:50about a random sampling, we select
- 40:59:52random size samples from each group or
- 40:59:54category. So we it's as random as you
- 40:59:56can get. Uh we talk about systematic
- 40:59:59sampling. We're selecting randomsiz
- 41:00:02samples from each group or category with
- 41:00:04a fixed periodic interval. Uh so we kind
- 41:00:08of split it up. This would be like a
- 41:00:09time setup or different categories. And
- 41:00:12you might ask your question, what is a
- 41:00:13category or a group? Uh if you look at
- 41:00:17I'm going to go back a window. Let's say
- 41:00:19we're studying um economics of different
- 41:00:21of an area. Um we know pretty much that
- 41:00:25based on their culture, where they came
- 41:00:28from, they might need to be separated.
- 41:00:31And so uh and when I say separated, I
- 41:00:33don't mean separated from their their uh
- 41:00:35place where they live. I mean, as far as
- 41:00:38the analysis, we want to look at the
- 41:00:39different groups and make sure they're
- 41:00:41all represented. So, if we had like an
- 41:00:4480% uh of a group that is uh say
- 41:00:47Hispanic and or Indian and also in that
- 41:00:51same area, we have 20% 20% who are let's
- 41:00:55call our expatriots. They left America
- 41:00:57and they're nice and uh your Caucasian
- 41:00:59group. We might want to sample a group
- 41:01:02that is representative of both. Uh, so
- 41:01:05we're talking about stratified sampling
- 41:01:08and we're talking about groups. Those
- 41:01:09are the groups we're talking about. And
- 41:01:10it brings us to stratified sampling,
- 41:01:12selecting approximately equalized
- 41:01:14samples from each group or category. Uh,
- 41:01:17this way we can actually separate the
- 41:01:19categories and give us an insight into
- 41:01:22the different cultures and how that
- 41:01:24might affect them in that area. Uh so
- 41:01:26you can see these are very very
- 41:01:28different kind of depends on what you're
- 41:01:30working with um as far as your data and
- 41:01:33what you're studying. And so we can see
- 41:01:35here just to go a little bit more we'd
- 41:01:37have selecting 25 employees from a
- 41:01:39company of 250 employees randomly. Don't
- 41:01:41care anything about them. What groups
- 41:01:42they're in, which office are in,
- 41:01:44nothing. Um and we might be selecting
- 41:01:47one employee from every 50 unique
- 41:01:49employees in a company of 250 employees.
- 41:01:52And then we have selecting one employee
- 41:01:54from every branch in the company office.
- 41:01:56So we have all the different branches.
- 41:01:58There's our group or our categories by
- 41:02:00the branch. And the category could
- 41:02:01depend on what you're studying. So it
- 41:02:03has a lot of variation on there. You see
- 41:02:06this kind of grouping and categorizing
- 41:02:08is also used to generate a lot of
- 41:02:10misinformation.
- 41:02:12Uh so if you only study one group and
- 41:02:14you say this is what it is, then
- 41:02:16everybody assumes that's what it is for
- 41:02:18everybody. And so you got to be very
- 41:02:20careful of that. and it's very unethical
- 41:02:21thing to kind of do. So, types of
- 41:02:24statistics. Uh we talk about statistics,
- 41:02:27we're going to talk about descriptive
- 41:02:29and inferential statistics. There are so
- 41:02:33many different terms in statistics to
- 41:02:35break it up. Uh so we so we're talking
- 41:02:37about a particular setup. So we're
- 41:02:41talking about descriptive and
- 41:02:42inferential uh statistics. You the base
- 41:02:45of the word describe is pretty solid.
- 41:02:48you're describing the data. What does it
- 41:02:51look like? With inferial statistics,
- 41:02:54we're going to take that from the small
- 41:02:55population to a large population. So, if
- 41:02:58you're working with a drug company, uh
- 41:02:59you might look at the data and say,
- 41:03:01"These people were helped by this drug.
- 41:03:04They did uh 80% better as far as their
- 41:03:07health or 80% better survival rate than
- 41:03:09the people um who did not have the drug.
- 41:03:13So, we can infer that that drug will
- 41:03:15work in the greater populace and will
- 41:03:16help people. So that's where you get
- 41:03:18your inferential. Uh so we are
- 41:03:20predicting how it's going to affect the
- 41:03:22greater population.
- 41:03:24So descriptive statistics it is used to
- 41:03:26describe the basic features of data and
- 41:03:28form the basis of quantitative analysis
- 41:03:31of data. So we have a measure of central
- 41:03:34tendencies. We have your mean, median
- 41:03:36and mode. And then we have a measure of
- 41:03:39spread like your range, your
- 41:03:40interquartile range, your variance and
- 41:03:43your standard deviation. And we're going
- 41:03:45to look at all these a little deeper
- 41:03:46here in a second. Uh but one of them you
- 41:03:49can think of is um how the data
- 41:03:53difference differences you know what's
- 41:03:55the max men range all that stuff is your
- 41:03:58spread and anything that's just a single
- 41:04:00number is usually your central uh
- 41:04:02tendencies measure of central
- 41:04:04tendencies. So we talk about the mean it
- 41:04:07is the average of the set of values
- 41:04:08considered. what is the average outcome
- 41:04:11of whatever's going on? And then your
- 41:04:13median separates the higher half and the
- 41:04:16lower half of data.
- 41:04:19Uh so where's the center point of all
- 41:04:21your different data points? So your mean
- 41:04:24might have some a couple really big
- 41:04:26numbers that skew it uh so that the
- 41:04:29average is much higher than if you took
- 41:04:31those outliers out where the median
- 41:04:34would by separating the high from the
- 41:04:36low might give you a much lower number.
- 41:04:39you might look at and say, "Oh, that's
- 41:04:40that's odd. Why is the average so much
- 41:04:42higher than the median?" Well, it's
- 41:04:44because you have some outliers, or why
- 41:04:45is it so much lower? And then the mode
- 41:04:47is the most frequent appearing value.
- 41:04:50Uh, this is really interesting. If
- 41:04:51you're studying economics and how people
- 41:04:53are doing, you might find that the most
- 41:04:55common um income like in the US was at
- 41:04:59one point 24,000 a year where the
- 41:05:02average was closer to 80,000. And it's
- 41:05:05like, wow, what a difference. Well,
- 41:05:07there's some people have a lot of money
- 41:05:09and so that skews that way up. So the
- 41:05:11average person is not making that kind
- 41:05:13of money. And then you look at the
- 41:05:14median income and you're like, well, the
- 41:05:16median income is a little bit closer to
- 41:05:18the average. Uh so it does create a very
- 41:05:20interesting way of looking at the data.
- 41:05:22Again, these are all uh central
- 41:05:25tendencies, single numbers you can look
- 41:05:26at for the whole spread of the data.
- 41:05:29And we look at the measure of central
- 41:05:32tendencies. The mean is the average
- 41:05:33marks of a students in a classroom. So
- 41:05:35here we have the mean sum of the marks
- 41:05:37of the students total number of students
- 41:05:40and as we talked about the median uh if
- 41:05:42we have 0 through 10 and we take half
- 41:05:46the numbers and put them on one side of
- 41:05:47the line half the numbers on the other
- 41:05:49side of the line uh we end up with five
- 41:05:51in the middle and then the mode what
- 41:05:53mark was scored by most of the students
- 41:05:55in a test in a simple case where most
- 41:05:58people scored like an 82% and got
- 41:06:01certain problems wrong easy to figure
- 41:06:03out. uh not so easy when you have
- 41:06:06different areas where like you have like
- 41:06:08the um oh let's go back to economy a
- 41:06:11little bit more difficult to calculate
- 41:06:12if you have a large group that scores
- 41:06:14that makes 30,000 and a slightly bigger
- 41:06:17group that makes 26,000. So what do you
- 41:06:19put down for the mode? Uh certainly
- 41:06:21there's a number of ways to calculate
- 41:06:22that and there's actually a different
- 41:06:24variations depending on what you're
- 41:06:25doing. So now we're looking at a measure
- 41:06:28of spread uh range. What's the
- 41:06:30difference between the highest and the
- 41:06:31lowest value? First thing you want to
- 41:06:33look at, you know, it's we had everybody
- 41:06:35in the test scored between 60 and 100%,
- 41:06:37somebody got 100% or maybe 60 to 90%. It
- 41:06:41was so hard that a lot of people could
- 41:06:42not get 100%. Um, and you have your
- 41:06:46interquartile range. Quartortiles divide
- 41:06:49a rankorder data set into four equal
- 41:06:52parts.
- 41:06:53very common thing to do as part of all
- 41:06:55the basic packages whether you're
- 41:06:57working in uh dataf frames with pandas
- 41:07:01whether you're working in scala whether
- 41:07:03you're working in R um you'll see this
- 41:07:05come up where they have range your min
- 41:07:07your max and then it'll have your
- 41:07:09interquartile range how does it look
- 41:07:11like in each quarter of data variance
- 41:07:13measures how far each number in the set
- 41:07:16is from the mean and therefore from
- 41:07:18every other number in the set uh so you
- 41:07:21have like a how much turbulence is going
- 41:07:23on in this data. And then the standard
- 41:07:26deviation, it is the measure of the
- 41:07:28variance or the dispersion of a set of
- 41:07:30values from the mean. And you'll usually
- 41:07:33see uh if I'm doing a graph, I might
- 41:07:35have the value graphed. Um and then
- 41:07:37based on the the error, I might graph
- 41:07:40graph the standard deviation and the
- 41:07:42error on the graph as a background so
- 41:07:44you can see how far off it is. Uh so
- 41:07:47standard deviation is used a lot. So
- 41:07:49measurement of spread uh marks of a
- 41:07:51student out of a 100 uh we have here
- 41:07:54from 50 to 63 or 50 to 90 uh so the
- 41:07:58range maximum marks minimum marks we
- 41:08:00have 90 to 45 and the spread of that is
- 41:08:0245 90 - 45 and then we have the
- 41:08:05interquartile range using the same marks
- 41:08:08over there you can see here where the
- 41:08:10median is and then there's the first
- 41:08:13quarter the second quarter and the third
- 41:08:15quarter based on splitting it apart by
- 41:08:17those values
- 41:08:19And to understand the variance and
- 41:08:20standard deviation, we first need to
- 41:08:22find out the mean. Uh so here's our our
- 41:08:25you know calculating the average there.
- 41:08:27We end up at approximately 66 for the
- 41:08:29average. And then we look at that the
- 41:08:31variance once we know the means we can
- 41:08:33do equals the marks minus the mean
- 41:08:35squared. Why is it squared? Uh because
- 41:08:39one, you want to make sure it's you
- 41:08:41don't have like if you if you're putting
- 41:08:43all this stuff together, you end up with
- 41:08:45an error as far as one's negative, one's
- 41:08:47positive, one's a little higher, one's a
- 41:08:49little lower. Uh so you always see the
- 41:08:52squared value and over the total
- 41:08:54observations. And so the standard
- 41:08:56deviation equals the square root of the
- 41:08:58variance, which is approximately 16. And
- 41:09:02if you were looking at um a predictable
- 41:09:04model, you would be looking at the
- 41:09:06deviation based on the error. How much
- 41:09:09error does it have? Uh that's again
- 41:09:12really important to know if you're if
- 41:09:14your prediction is predicting something,
- 41:09:16what's a chance of it being way off or
- 41:09:18just a little bit off.
- 41:09:21Now that we've looked at the um tools as
- 41:09:24far as some of the basics for doing your
- 41:09:26statistics and what we're talking about,
- 41:09:28let's go ahead and pull up a little demo
- 41:09:30and show you what that looks like in
- 41:09:31Python code. Uh so you can get some
- 41:09:33little hands-on here. For that, let's go
- 41:09:35back into our Jupyter notebook in
- 41:09:37Python. Now, almost all of this you can
- 41:09:40do in numpy. Last time we worked um in
- 41:09:42numpy. This time we're going to go ahead
- 41:09:44and use pandas. And if you remember from
- 41:09:47pandas on here, uh this is basically a
- 41:09:50data frame, rows, columns. Let's just go
- 41:09:53ahead and do a print df. head
- 41:09:58and run that.
- 41:10:00And you can see we have uh the name
- 41:10:02Jane, Michael, William, Rosie, Hannah,
- 41:10:04and their salaries on here. And of
- 41:10:06course, instead of having to do all
- 41:10:08those hand calculations and add
- 41:10:09everything together and divide by the
- 41:10:11total, we can do something very simple
- 41:10:13on this uh like use the command mean in
- 41:10:17pandas. And so if I go ahead and do this
- 41:10:19print df, pick our column salary because
- 41:10:22we want to find the means of that
- 41:10:24colery.
- 41:10:25We want to find the means of that
- 41:10:27column. Uh and we go and print this out.
- 41:10:29And you can see that the uh average
- 41:10:32income on here is 71,000.
- 41:10:35Uh, and let's just go ahead and do this.
- 41:10:37We'll go ahead and put in uh means.
- 41:10:42And if we're going to do that, we also
- 41:10:44might want to find the median.
- 41:10:48And the median is uh very similar except
- 41:10:52it actually is just median. Uh we're
- 41:10:54used to means and average. It's kind of
- 41:10:56interesting that those are they use the
- 41:10:58two different words. Uh there can be in
- 41:11:01some computations slight differences but
- 41:11:03for the most part the means is the
- 41:11:05average. Uh and then the median oops
- 41:11:08let's put a
- 41:11:12median here. DF salary that way it
- 41:11:14displays a little better. We can see the
- 41:11:16median is 54 um000. So the halfway mark
- 41:11:20is significantly below the average. Why?
- 41:11:23Because we have somebody in here who
- 41:11:24makes 189,000.
- 41:11:26Darn you Rosie for throwing off our
- 41:11:28numbers. Uh but that's something you'd
- 41:11:30want to notice. This is this is the
- 41:11:32difference between these is huge and so
- 41:11:34is what is the meaning behind that when
- 41:11:36you're studying a populace and looking
- 41:11:37at uh the different data coming in. And
- 41:11:40of course we also want to find out hey
- 41:11:42what's the most uh common income that
- 41:11:46people make in this little tiny sample.
- 41:11:48And so we'll go ahead and do the mode.
- 41:11:51And you can see here with the mode uh
- 41:11:53it's at 50,000.
- 41:11:55So this is this is very telling that
- 41:11:57most people are making 50,000. The
- 41:12:00middle point is at 54,000. So half the
- 41:12:03people are making more than that. What
- 41:12:05that tells me is that if the most common
- 41:12:08income is way is below the median, then
- 41:12:13there's a few there's a SK, you know,
- 41:12:14there's a a lot of high salaries going
- 41:12:16up, but there's some really low salaries
- 41:12:18in there. And so this trend which is
- 41:12:21very common in statistic you when you're
- 41:12:23analyzing the economy and different
- 41:12:26people's income is pretty common and the
- 41:12:28bigger difference between these is also
- 41:12:31very important when we're studying
- 41:12:32statistics. Uh and when you hear someone
- 41:12:35just say hey the average income was you
- 41:12:37might start asking questions at that
- 41:12:39point. Why aren't you talking about the
- 41:12:40median income? Why aren't you talking
- 41:12:42about the mode the most common income?
- 41:12:44What are you hiding? Uh and if you're
- 41:12:47doing these analysis, you should be
- 41:12:48looking at these saying, "Hey, why why
- 41:12:49are this discrepancies? Why are these so
- 41:12:51different?" And of course, with any uh
- 41:12:53analysis, it's important to find out the
- 41:12:55minimum
- 41:12:57and the maximum. So, we'll go ahead.
- 41:12:59It's just simply uh um min'll pull up
- 41:13:04your minimum and then do max pulls up
- 41:13:06the maximum. pretty straightforward on
- 41:13:09as far as um translating it and knowing
- 41:13:13what your you know what the your lowest
- 41:13:15value and what your highest value is
- 41:13:16here. Um which you'll use to generate
- 41:13:19like a spread later on. And real quick
- 41:13:22on no mode mode, uh note that it puts
- 41:13:24mode zero. Like I said, there's a couple
- 41:13:26different ways you can compute the mode.
- 41:13:28Um although, you know, standard one's
- 41:13:30pretty good. We can of course do the
- 41:13:32range, which is your max minus your min.
- 41:13:35So now we have a range of 149,000
- 41:13:38between the upper end and the lower end.
- 41:13:40And you might want to be looking up the
- 41:13:42individual values on all of these. But
- 41:13:45it turns out there is a describe
- 41:13:49feature in pandas.
- 41:13:51And so in pandas we can actually do df
- 41:13:54salary describe. And if we do this you
- 41:13:56can see we have that there's seven uh
- 41:13:58setups. Here's our mean. Um, our
- 41:14:01standard deviation, which we didn't
- 41:14:03compute yet, which would just be a STD.
- 41:14:06And you got to be a little careful
- 41:14:07because when it computes it, it looks
- 41:14:08for axes and things like that. Uh, we
- 41:14:10have our minimum value, and here's our
- 41:14:12cortiles,
- 41:14:14uh, our maximum value, and then of
- 41:14:16course the name salary. Uh, so these are
- 41:14:18the these are the basic statistics. You
- 41:14:20can pull them up and just describe. This
- 41:14:22is a dictionary. So I could actually do
- 41:14:25something like um in here I could
- 41:14:28actually go uh count and run. And now it
- 41:14:32just prints the count. Uh so because
- 41:14:34this is a dictionary, you can pull any
- 41:14:36one of these values out of here. It's
- 41:14:38kind of a quick and dirty way to pull
- 41:14:40all the different information and then
- 41:14:42split it up and depending on what you
- 41:14:43need. Now if I just walked in and gave
- 41:14:46you this information um in a meeting, at
- 41:14:49some point you would just kind of fall
- 41:14:51asleep. That's what I would do anyway.
- 41:14:55Um, so we want to go ahead and and see
- 41:14:57about graphing it here. And we'll go
- 41:14:59ahead and put it into a histogram and
- 41:15:01plot that graph on it of the salaries.
- 41:15:04And let's just go ahead and put that in
- 41:15:06here. So we do our map plot inline.
- 41:15:09Remember that's a Jupiter's notebook
- 41:15:10thing. Uh, a lot of the new version of
- 41:15:13the mapplot library does it
- 41:15:14automatically, but just in case I always
- 41:15:17put it in there. Uh, import mattplot
- 41:15:18library piplot as plt. That's my
- 41:15:21plotting.
- 41:15:23And then we have our data frame. Uh I
- 41:15:25don't I guess I really don't need to
- 41:15:26respell the data frame. Maybe we could
- 41:15:28just remind oursel what's in it. So
- 41:15:30we'll go ahead and just uh print
- 41:15:32DF. That way we still have it. And then
- 41:15:35we have our salary. DF salary
- 41:15:38salary.plot history title salary
- 41:15:40distribution color gray. Uh plot AXV
- 41:15:44line salary the mean value. So, we're
- 41:15:47going to take the mean value um color
- 41:15:50violet line style dash. This is just all
- 41:15:53making it pretty. Uh what color dash
- 41:15:56line width of two that kind of thing.
- 41:15:59And the median. And let's go ahead and
- 41:16:00run this just so you can see what we're
- 41:16:02talking about.
- 41:16:04And so up here we are taking on our
- 41:16:07plot. Um so here's the data. Here's our
- 41:16:10our data frame printed out so you can
- 41:16:12see it with the salaries. We're looking
- 41:16:14at the salary distribution and just look
- 41:16:16at this the way they're the salary is
- 41:16:18distributed. Um you have our in this
- 41:16:22case we did let's see we had red for the
- 41:16:25median we have violet
- 41:16:28for our average or mean and you can just
- 41:16:32see how it really here's our outlier.
- 41:16:35Here's our person who makes a lot of
- 41:16:36money. Here's the um average and here's
- 41:16:40the median. Um, and so as you look at
- 41:16:42this, you can say, "Wow." Um, based on
- 41:16:44the average, it really doesn't tell you
- 41:16:46much about what people are really taking
- 41:16:47home. All it does is tell you how much
- 41:16:50money is in this, you know, what the
- 41:16:52average salary is. So, some of the
- 41:16:55things you want to take away in addition
- 41:16:57to this is that it's very easy to plot
- 41:17:01um an AXV line. These are these up and
- 41:17:04down lines for your markers. Um, and as
- 41:17:08you display display the data, I mean,
- 41:17:09you can add all kinds of things to this
- 41:17:11and get really complicated. Keeping it
- 41:17:13simple is pretty straightforward. I look
- 41:17:14at this and I can see we have a major
- 41:17:16outlier out here. We can definitely do a
- 41:17:18histogram and stuff like that. Um, but
- 41:17:20you know, picture's worth a thousand
- 41:17:22words. What you really want to make sure
- 41:17:24you take away is that we can do a basic
- 41:17:26describe which pulls all this
- 41:17:29information out and we can print any of
- 41:17:31the individual information from the
- 41:17:33describe uh because this is a
- 41:17:35dictionary.
- 41:17:38And so if we want to go ahead and look
- 41:17:40up um the mean value, we can also do
- 41:17:43describe mean. So if you're doing a lot
- 41:17:45of statistics, uh being able to
- 41:17:49doesn't have the print on there, so it's
- 41:17:50only going to print um the last one,
- 41:17:52which happens to be the mean. Uh you can
- 41:17:54very easily reference any one of these.
- 41:17:56And then you can also, if you're doing
- 41:17:58something a little bit more complicated
- 41:17:59and you don't need just the basics, you
- 41:18:01can come through and pull any one of the
- 41:18:04individual um
- 41:18:07references from the from the pandas on
- 41:18:10here. So now we've had a chance to
- 41:18:13describe our data. Uh let's get into
- 41:18:16inferential statistics. Inferial
- 41:18:19statistics allows you to make
- 41:18:20predictions or inferences from data. And
- 41:18:24you can see here we have a nice little
- 41:18:25picture movie ratings and um if we took
- 41:18:29this group of people and said hey how
- 41:18:31many people like the movie dislike it
- 41:18:33can't say and then you ask just a random
- 41:18:36person who comes out of the movie who
- 41:18:37hasn't been in this study uh you can
- 41:18:39infer that 55% chance of saying liked
- 41:18:4335% chance of saying disliked or a 10 or
- 41:18:4611% chance of can't say. So that that's
- 41:18:49real basics of what we're talking about
- 41:18:51is you're going to infer that the next
- 41:18:53person is going to follow these
- 41:18:54statistics.
- 41:18:57Uh so let's look at point estimation. Uh
- 41:19:00it is a process of finding an
- 41:19:01approximate value for a population's
- 41:19:03parameter like mean or average from
- 41:19:06random samples of the population. Let's
- 41:19:09take an example of testing vaccines for
- 41:19:11COVID 19. Uh vaccines and flu bugs, all
- 41:19:14that. It's a pretty big thing of how do
- 41:19:16you test these out and make sure they're
- 41:19:17going to work on the populace. A group
- 41:19:20of people are chosen from the
- 41:19:21population. Medical trials are
- 41:19:23performed. Results are generalized for
- 41:19:26the whole population. So here's a
- 41:19:28protected here's our small group up here
- 41:19:30where we've selected them. We run
- 41:19:32medical trials on them and then the
- 41:19:33results work for the population. You
- 41:19:35nice diagram with the arrows going back
- 41:19:37and forth and the very scary co virus in
- 41:19:40the middle of one. And let's take a look
- 41:19:42at the applications of inferial
- 41:19:44statistics.
- 41:19:46Very central is what they call
- 41:19:47hypothesis testing uh and the confidence
- 41:19:51interval which go with that. And then as
- 41:19:54we get into
- 41:19:56probability, we get into our binomial
- 41:19:59theorem, our normal distribution and
- 41:20:01central limit theorem. Hypothesis
- 41:20:04testing. Hypothesis testing is used to
- 41:20:06measure the plausibility of a hypothesis
- 41:20:09assumption by using sample data. Now
- 41:20:13when we talk about theorems, theory,
- 41:20:17hypothesis,
- 41:20:19uh keep in mind that if you are in a
- 41:20:21philosophy class, theory is the same as
- 41:20:25hypothesis where theorem is a scientific
- 41:20:29uh statement that is something that has
- 41:20:30been proven although it is always up for
- 41:20:33debate because in science we always want
- 41:20:35to make sure things are up to debate. So
- 41:20:37a hypothesis is the same as a phil
- 41:20:39philosophical class calling a theory
- 41:20:41where theory in science is not the same.
- 41:20:44Theory in science says this has been
- 41:20:45well proven. Gravity is a theory. Uh so
- 41:20:48if you want to debate the theory of
- 41:20:50gravity try jumping up and down. If you
- 41:20:52want to have a theory about why the
- 41:20:54economy is collap collapsing in your
- 41:20:56area that is a philosophical debate.
- 41:21:00Very important. I've heard people mix
- 41:21:01those up and it is a pet peeve of mine.
- 41:21:04When we talk about hypothesis testing,
- 41:21:06the steps involved in hypothesis testing
- 41:21:08is first we formulate a hypothesis. We
- 41:21:11figure out the right test to test our
- 41:21:13hypothesis. We execute the test and we
- 41:21:16make a decision. And so when you're
- 41:21:18talking about hypothesis, you're usually
- 41:21:20trying to disprove it. If you can't
- 41:21:22disprove it and it works for all the
- 41:21:25facts, then you might call that a
- 41:21:27theorem at some point. So in a use case,
- 41:21:30uh let's consider an example. We have
- 41:21:32four students. were given a task to
- 41:21:33clean a room every day. Sounds like
- 41:21:35working with my kids. They decided to
- 41:21:37distribute the job of cleaning the room
- 41:21:39among themselves. They did so by making
- 41:21:41four chits which has their names on it
- 41:21:44and the name that gets picked up has to
- 41:21:46do the cleaning for that day. Rob took
- 41:21:48the opportunity to make chits and wrote
- 41:21:50everyone's name on it. So here's our
- 41:21:52four people, Nick, Rob, Imlia, Imlia,
- 41:21:55and Summer.
- 41:21:58Now Rick, Imlia and Summer are asking us
- 41:22:00to decide whether Rob has done some
- 41:22:03mischief in preparing the chits i.e
- 41:22:05whether Rob has written his name on one
- 41:22:07of the chit. For that we will find out
- 41:22:09the probability of Rob getting the
- 41:22:11cleaning job on first day, second day,
- 41:22:13third day and so on till 12 days. The
- 41:22:16probability of Rob getting the job
- 41:22:18decreases every day. I.e. his turn never
- 41:22:21comes up. Then definitely he has done
- 41:22:23some mischief while making the chits. So
- 41:22:26the probability of Rob not doing work on
- 41:22:28day one is uh three out of four. There's
- 41:22:30a 75 chance that he didn't do work. Uh
- 41:22:33two days 34s * 34s equals.56.
- 41:22:383 days you have 3/4 34 34 which
- 41:22:41equals42.
- 41:22:43Uh when you get to day 12 it's 0032
- 41:22:46which is less than 0.05.
- 41:22:48Remember this 0.05 uh that comes up a
- 41:22:51lot when we're talking about um certain
- 41:22:54values when we're looking at statistics.
- 41:22:57Rob is cheating as he wasn't chosen for
- 41:22:5912 consecutive days. That's a very high
- 41:23:01probability when on day 12 he still
- 41:23:05hasn't gotten the job cleaning the room.
- 41:23:08So we come up to our important important
- 41:23:10terminologies.
- 41:23:12We have null hypothesis.
- 41:23:15a general statement that states that
- 41:23:16there is no relationship between two
- 41:23:18measured phenomenon or no assoc
- 41:23:20association among the groups.
- 41:23:23Alternative hypothesis contrary to the
- 41:23:26null hypothesis it states whenever
- 41:23:28something is happening a new theory is
- 41:23:30preferred instead of an old one. And so
- 41:23:33the two hypothesis go hand in hand. Uh
- 41:23:36so your null this is always interesting
- 41:23:38in in we're talking about data science
- 41:23:40and the math behind it. It's about
- 41:23:42proving that the things have no
- 41:23:44correlation. Null hypothesis says these
- 41:23:47two have zero relation to each other.
- 41:23:49Where the alternative hypothesis says,
- 41:23:51hey, we found a relation. This is what
- 41:23:53it is. We have p value. The p value is
- 41:23:56the probability of finding the observed
- 41:23:58or more extreme results when the null
- 41:24:00hypothesis of a study question is true.
- 41:24:04And the t value, it is simply the
- 41:24:06calculated difference represented in
- 41:24:08units of standard error. The greater the
- 41:24:10magnitude of t, the greater the evidence
- 41:24:12against the null hypothesis. And you can
- 41:24:14look at the t value as being specific to
- 41:24:17the test you're doing where the p value
- 41:24:20is derived from your t value and you're
- 41:24:22looking for what they call the 5% or the
- 41:24:250.05
- 41:24:26showing that it has a high correlation.
- 41:24:28So digging in deeper, let's assume that
- 41:24:31a new drug is developed with the goal of
- 41:24:33lowering the blood pressure more than
- 41:24:35the existing drug. And this is a good
- 41:24:37one because uh the null value here isn't
- 41:24:40that you don't have any drug. The null
- 41:24:41value here is that it's better than the
- 41:24:43existing drug. The new drug doesn't
- 41:24:45lower the blood pressure more than the
- 41:24:47existing drug. Now if we get that uh
- 41:24:50that says our null hypothesis is
- 41:24:52correct. There is no correlation and the
- 41:24:54new drug is not doing its job. The
- 41:24:57alternative hypothesis the new drug does
- 41:25:00significantly lower the blood pressure
- 41:25:02more than the existing drug. Uh, yay, we
- 41:25:05got a new drug out there. And that's our
- 41:25:06alternative hypothesis or the H1 or HA.
- 41:25:10And we look at the p value results from
- 41:25:13the evidence like medical trials showing
- 41:25:15positive results which will reject the
- 41:25:17null hypothesis. And again, they're
- 41:25:19looking for um a 0.05 or 5%. And the t
- 41:25:23value comparing all the positive test
- 41:25:25results and finding means of different
- 41:25:27samples in order to test hypothesis. So
- 41:25:30this is specific to the test. how uh
- 41:25:32what percentage of increase did they
- 41:25:34have and this leads us to the confidence
- 41:25:37intervals. Uh a confidence interval is a
- 41:25:40range of values we are sure our true
- 41:25:42values of observations lie in. Let's say
- 41:25:45you asked a dog owner around you and
- 41:25:47asked them how many cans of food do you
- 41:25:50buy for your uh per year for your dog.
- 41:25:53Through calculations you got to know
- 41:25:55that the on an average around 95% of the
- 41:25:58people bought around 200 to 300 cans of
- 41:26:00food. Hence we can say that we have a
- 41:26:03confidence interval of 230 where 95% of
- 41:26:06our values lie in that spread data
- 41:26:09spread. Uh and this the graph really
- 41:26:12helps a lot. So you can start seeing
- 41:26:13what you're looking at here where you
- 41:26:15have the 95%. You have your peak in this
- 41:26:18case it's a normal distribution. So you
- 41:26:19have the nice bell curve equal on both
- 41:26:21sides. It's not asymmetrical. And 95% of
- 41:26:24all the values lie within a very small
- 41:26:26range. And then you have your outliers
- 41:26:28the 2.5% going each way.
- 41:26:31So we touched upon hypothesis uh and
- 41:26:34we're going to move into probability. Uh
- 41:26:36so you have your hypothesis. Once you've
- 41:26:38generated your hypothesis, we want to
- 41:26:39know the probability of something
- 41:26:41occurring. Probability is a measure of
- 41:26:43the likelihood of an event to occur. Any
- 41:26:46event can be predicted with total
- 41:26:47certainty and can only be predicted as a
- 41:26:50likelihood of its occurrence. So any
- 41:26:53event cannot be predicted with total
- 41:26:54certainty. It can only be predicted as a
- 41:26:56likelihood of its occurrence. Uh score
- 41:26:59prediction. how good you're going to do
- 41:27:01in whatever sport you're in, weather
- 41:27:04prediction, stock prediction, if you've
- 41:27:07studied physics and chaos theory, even
- 41:27:09the location of the chair you're sitting
- 41:27:11on has a probability that it might move
- 41:27:133 ft over. Granted, that probability is
- 41:27:16one in like uh I think we calculated as
- 41:27:18under one in trillions upon trillions.
- 41:27:21So, it's the better the probability, the
- 41:27:24more likely it's going to happen. There
- 41:27:25are some things that have such a low
- 41:27:26probability that we don't see them. So
- 41:27:29we talk about a random variable. Uh
- 41:27:31random variable is a variable whose
- 41:27:33possible values are numerical outcomes
- 41:27:35of a random phenomena. So uh we have the
- 41:27:38coin toss. How many heads will occur in
- 41:27:41the series of 20 coin flips? Probably
- 41:27:44you know the on average there are 10,
- 41:27:45but you really can't know because it's
- 41:27:47very random. How many times a red ball
- 41:27:49is picked from a bag of balls if there's
- 41:27:51equal number of of red balls and blue
- 41:27:54balls and green balls in there. How many
- 41:27:56times the sum of digits on two dice uh
- 41:27:59result or five each? Um so you know
- 41:28:02there's how often you're going to roll
- 41:28:04two fives on your pair of dice. So in a
- 41:28:07use case uh let's consider the example
- 41:28:09of rolling two dice. We have a random
- 41:28:11variable outcome equals y. You can take
- 41:28:12values 2 3 4 5 6 7 8 9 10 11 12. So we
- 41:28:17have a random variable and a combination
- 41:28:19of dice and instead of looking at how
- 41:28:22many times um both dice were roll five
- 41:28:25let's go ahead and look at a total sum
- 41:28:27of five and you have in as far as your
- 41:28:29random variables you can have a one four
- 41:28:31equals 5 4 1 2 3 32 so four of those
- 41:28:35roles can be four if you look at all the
- 41:28:38different options you have four of those
- 41:28:40random rolls can be a five and if we
- 41:28:43look at the total number
- 41:28:46which happens to be 36 different
- 41:28:48options. Uh you can see that we have
- 41:28:51four out of 36 chance every time you
- 41:28:53roll the dice that you're going to roll
- 41:28:54a total of five. You're going to have an
- 41:28:56outcome of five. And uh we'll look a
- 41:28:59little deeper as to what that means. Uh
- 41:29:01but you could think of that at what
- 41:29:03point if someone never rolls a five or
- 41:29:05they always roll a five, can you say,
- 41:29:07"Hey, that person's probably cheating."
- 41:29:09uh we'll look a little closer at the
- 41:29:11math behind that but let's just consider
- 41:29:13this as one of the cases is rolling two
- 41:29:15dice and gambling. There's also a
- 41:29:17binomial distribution. It is the
- 41:29:19probability of getting success or
- 41:29:21failure as an outcome in an experiment
- 41:29:23or trial that is repeated multiple
- 41:29:25times. And the key is is by meaning two
- 41:29:29binomial. Uh so passing or failing an
- 41:29:31exam, winning or losing a game and
- 41:29:34getting either head or tails. So if you
- 41:29:36ever see binomial distribution, it's
- 41:29:38based on a um true false kind of setup.
- 41:29:41You win or lose. Let's consider a uh use
- 41:29:45case and let's consider the game of
- 41:29:47football between two clubs Barcelona and
- 41:29:51Dortmund. The teams will have to play a
- 41:29:53total of four matches and we have to
- 41:29:55find out the chances of Barcelona
- 41:29:58winning the series. So we look at the
- 41:30:00total games and we're looking at five
- 41:30:02different games or matches. Let's say
- 41:30:04that the winning chance for Barcelona is
- 41:30:0675% or 75. That means at each game they
- 41:30:10have a 75% chance that they're going to
- 41:30:12win that game and losing chances are 25%
- 41:30:15or 0.25. Clearly 75 plus 0.25 equals 1.
- 41:30:20So that accounts for 100% of the game.
- 41:30:22Probability for getting K wins in n
- 41:30:25matches is calculated.
- 41:30:27And we we're talking like so if you have
- 41:30:29five games uh and you want to know if I
- 41:30:31play um how many wins in those five
- 41:30:34games should I get? What's a percentage
- 41:30:36on those? And the probability for
- 41:30:38getting k wins in n matches is
- 41:30:41calculated by px= k= n k p the k q to
- 41:30:47the n minus k. Here p is the probability
- 41:30:50of success and q is the probability of
- 41:30:53failure. And so we can do total games of
- 41:30:55n equals 5 where k equals 012345.
- 41:31:00P which is the chance of winning is 75.
- 41:31:04Q the chance of losing equals 1 minus p
- 41:31:07which equals 1 - 0075 which equals 0.25.
- 41:31:11The probability that Barcelona will lose
- 41:31:13all of the matches can then just plug in
- 41:31:15the numbers and we end up with a
- 41:31:1809765625.
- 41:31:23So very small chance they're going to
- 41:31:25lose all their matches.
- 41:31:27And we can plug in uh the value for two
- 41:31:29matches. Probability that Barcelona will
- 41:31:32win at least two matches is 00878. And
- 41:31:36of course we can go on to probability
- 41:31:38that Barcelona will win three matches
- 41:31:40the 26 and of course four matches and so
- 41:31:44on. And it's always nice to take this
- 41:31:46information um and let let's find the
- 41:31:48cumulative discrete probabilities for
- 41:31:50each of the outcomes where Barcelona has
- 41:31:53won three or more matches x= 3 x= 4 x= 5
- 41:31:58and we end up with the p =264 plus 395 +
- 41:32:02237 which equals89.
- 41:32:05In reality the probability of Barcelona
- 41:32:08winning the series is much higher than
- 41:32:1075. And it's always nice to uh put out a
- 41:32:15nice graph so you can actually see the
- 41:32:17number of wins to the probability and
- 41:32:19how that pans out with our binomial
- 41:32:22case. Continuing in our important
- 41:32:24terminology, location, the location of
- 41:32:27the center of the graph depends on the
- 41:32:29mean value. And uh this is some very
- 41:32:32important things. So much of the data we
- 41:32:35look at and when you start looking at
- 41:32:36probabilities almost always has a
- 41:32:38normalized look like the graph in the
- 41:32:40middle.
- 41:32:42uh but you do have left skewed where the
- 41:32:44data is skewed off to the left and you
- 41:32:45have more stuff happening off to the
- 41:32:47left and you have right skewed data and
- 41:32:49so when this comes up and these
- 41:32:50probabilities come up where they're
- 41:32:52skewed it's really important to take a
- 41:32:54closer look at that uh mostly you end up
- 41:32:56with a normalized set of data but you
- 41:32:58got to also be aware that sometimes it's
- 41:33:00a skewed data and then the height height
- 41:33:03of the slope inversely depends upon the
- 41:33:05standard deviation
- 41:33:07so you can see down here the standard
- 41:33:09deviation is really large it kind of
- 41:33:10squishes it out. And if the standard
- 41:33:12deviation is small, then most of your
- 41:33:15data is going to hit right there in the
- 41:33:16middle. You're going to have a nice
- 41:33:17peak. Um, and so being aware of this
- 41:33:19that you might have a probability that
- 41:33:22fits certain data, but it has a lot of
- 41:33:24outliers. So you're if you have a really
- 41:33:26high standard deviation, um, if you're
- 41:33:28doing stock market analysis,
- 41:33:31this means your predictions are probably
- 41:33:33not going to make you much money. uh
- 41:33:35where if you have a very small
- 41:33:36deviation, you might be right on target
- 41:33:38and set to become a millionaire. Which
- 41:33:40leads us to the zcore. Zcore tells you
- 41:33:43how far from the mean a data point is.
- 41:33:46It is measured in terms of standard
- 41:33:48deviations from the mean. Around 68% of
- 41:33:51the results are found between one
- 41:33:53standard deviation. Around 95% of the
- 41:33:56results are found between two standard
- 41:33:58deviations.
- 41:33:59And you read the symbols. Of course,
- 41:34:01they love to throw some Greek letters in
- 41:34:02there. we have mu minus 2 sigma. Mu is
- 41:34:07just a quick way. It's that kind of
- 41:34:09funky u. It just means the mean. Uh and
- 41:34:12then the sigma is the standard
- 41:34:14deviation. And that's the o with a
- 41:34:16little arrow off to the right or the
- 41:34:18little waggly tail going up. The o with
- 41:34:20a with a line on it. Uh so mu minus 2
- 41:34:23sigma is your uh 95% of the results are
- 41:34:27found between two standard deviations.
- 41:34:30The central limit theorem. This goes
- 41:34:33back to the skew. If you remember, we
- 41:34:35were looking at the skew values on this
- 41:34:37previous slide. Have left skewed,
- 41:34:40normalized, and right skewed. When we're
- 41:34:42talking about it being skewed or not
- 41:34:44skewed, the distribution of the sample
- 41:34:46means will be approximately normally
- 41:34:49distributed, evenly distributed, not
- 41:34:51skewed. If you take large random samples
- 41:34:54from the population with the mean mu and
- 41:34:57the standard deviation sigma with
- 41:35:00replacement
- 41:35:02and you can see here um uh of course we
- 41:35:04have our uh mu minus 2 sigma and the
- 41:35:07spread down here the mean the median and
- 41:35:09the mode and so when you're talking
- 41:35:11about very large populations
- 41:35:14these numbers should come together and
- 41:35:15you shouldn't have a skewed value. If
- 41:35:17you do that's a flag that something's
- 41:35:19wrong. That's why this is so important
- 41:35:22to be aware of what's going on with your
- 41:35:24data, where your samples are coming
- 41:35:26from, and the math behind it. And if
- 41:35:29you're going to do all this, we got to
- 41:35:30jump into conditional probability. The
- 41:35:34conditional probability of an event A is
- 41:35:36a probability that the event will occur
- 41:35:39given the knowledge that an event B has
- 41:35:41already occurred. And you'll see this as
- 41:35:44Baze theorem. B A Y S bay. Uh, and this
- 41:35:48is read. I mean, you have these funky
- 41:35:50looking little P brackets. A B. This is
- 41:35:54the probability of A being true while B
- 41:35:58is already true. And you have the
- 41:36:00probability of B being true when A is
- 41:36:02already true. So, P B of A probability
- 41:36:06of A being true divided by the
- 41:36:08probability of B being true. And we talk
- 41:36:11about BA's theorem which occurred back
- 41:36:13in the 1800s when he discovered this.
- 41:36:16This is such an important formula and
- 41:36:18it's really it's not if you actually do
- 41:36:20the math you could just kind of do um um
- 41:36:23XY equals J K and then you divide them
- 41:36:27out and you're going to see the same
- 41:36:28math but it works with probabilities
- 41:36:30which makes it really nice. And so if
- 41:36:33you have a s you might have uh eight or
- 41:36:35nine different studies going on in
- 41:36:38different areas different people have
- 41:36:39done the studies they brought them
- 41:36:41together. Um if we look at today's co
- 41:36:44virus the virus spread uh certainly the
- 41:36:47studies done in China versus the studies
- 41:36:50the way they're done in the US that data
- 41:36:52is different in each of those studies
- 41:36:54but if you can find a place where it
- 41:36:56overlaps where they're studying the same
- 41:36:58thing together you can then compute the
- 41:37:01changes that you need to make in one
- 41:37:02study to make them equal and this is
- 41:37:05also true if you have a study of uh um
- 41:37:09one group and you want to find out more
- 41:37:10about it. So this formula is very
- 41:37:13powerful. Uh it really has to do with
- 41:37:15the data collection part of the math and
- 41:37:17data science and understanding where
- 41:37:19your data is coming from and how you're
- 41:37:21going to combine different studies in
- 41:37:23different groups. And we'll go ahead and
- 41:37:25go into a use case. Uh let's find out
- 41:37:27the chance of a person getting lung
- 41:37:29disease due to smoking. Uh and this is
- 41:37:32kind of interesting the way they word
- 41:37:33this. Um let's say that according to
- 41:37:35medical report provided by the hospital
- 41:37:38states that around 10% of all patients
- 41:37:41they treated suffered lung lung disease.
- 41:37:44Uh so we have kind of a generic medical
- 41:37:46report. They further found out uh by a
- 41:37:50survey that 15% of the patients that
- 41:37:52visit them smoke. So we have 10% that
- 41:37:55are lung disease and um 15% of the
- 41:37:58patients smoke. And finally, 5% of the
- 41:38:02people continued smoke even when they
- 41:38:04had lung disease. Uh not the brightest
- 41:38:07choice um but you know it is an
- 41:38:09addiction so it can be really difficult
- 41:38:10to kick. And so we can look at the
- 41:38:13probability of a uh prior probability of
- 41:38:1510% people having lung disease. And then
- 41:38:18probability b probability that a patient
- 41:38:21smokes is 15%.
- 41:38:24Uh and the probability of B um if B then
- 41:38:28A. The probability of a patient smokes
- 41:38:30even though they have lung disease is
- 41:38:335%. And probability of A is B.
- 41:38:36Probability that the patient will have
- 41:38:38lung disease if they smoke. And then
- 41:38:40when you put the formulas together, uh
- 41:38:42you get a nice solution here. You get
- 41:38:43the probability of A of B, probability
- 41:38:45that the patient will have lung disease
- 41:38:47if they smoke. And you can just plug the
- 41:38:49numbers right in and we get a 3.33%
- 41:38:53chance. Hence, there is a 3.33% chance
- 41:38:56that a person who smokes will get a lung
- 41:38:58disease. So, we're going to pull up a
- 41:39:01little Python code, always my favorite,
- 41:39:03roll up the sleeves. Keep in mind, we're
- 41:39:06going to be doing this um kind of like
- 41:39:08the backend way so that you can see
- 41:39:11what's going on. And then later on we're
- 41:39:14going to create um we'll get into
- 41:39:17another demo which shows you some of the
- 41:39:19tools that are already pre-built for
- 41:39:20this. Let's start by creating a set. So
- 41:39:25we're going to create a set with curly
- 41:39:26braces. This means that our set has um
- 41:39:30only unique values. So you have a list
- 41:39:34uh you have your tupils which can never
- 41:39:36change and then you have um in this case
- 41:39:39the the set. So 47, you can't create a
- 41:39:4347, 4. It'll delete the four out. So
- 41:39:46it's only unique values. And if you use
- 41:39:49dictionaries,
- 41:39:51quick reminder, this should look
- 41:39:53familiar because it is a dictionary uh
- 41:39:56where you have a value and that value is
- 41:39:58assigned to or that key is assigned to a
- 41:40:01value. Uh so you could have a key value
- 41:40:03set up as a dictionary. So it's like a
- 41:40:05dictionary without the value. It's just
- 41:40:07the keys and they all have to be unique.
- 41:40:11And if we run this, we have a set of 47.
- 41:40:17We can also take a list, a regular um
- 41:40:20setup. And I'm going to go ahead and
- 41:40:21just throw in another number in here,
- 41:40:22four, and run it. Uh, and you can see
- 41:40:25here if I take my list 1 2 3 4, and I
- 41:40:29convert it to a set, and here it is. My
- 41:40:32set from list equals set my list.
- 41:40:36The result is 1 2 3 4. So, it just
- 41:40:38deletes that last four right out of
- 41:40:40there.
- 41:40:42And with the sets, you can also go in
- 41:40:44there and um print here is my set. My
- 41:40:48set uh three is in the set. And then if
- 41:40:51you do three in my set,
- 41:40:54that's going to be a logic function. Uh
- 41:40:57and one in my set, six is not in the
- 41:41:00set, and so forth. If we run this,
- 41:41:04we get three is in the set true one is
- 41:41:06in the set false because 357 is another
- 41:41:09one. Six is in the set uh six is not in
- 41:41:12the set. So not in my set. You can also
- 41:41:16use this with a list. We could have just
- 41:41:18used 357 and it would have um the same
- 41:41:22response on there is three and usually
- 41:41:25you do if three is in but three in my
- 41:41:27set is still works on a just a regular
- 41:41:29list. And we'll go ahead and do a little
- 41:41:32iteration. We're going to do kind of the
- 41:41:34dice one. Remember um uh 1 2 3 4 5 6.
- 41:41:38And so we're going to bring in an
- 41:41:39iteration tool and import product as
- 41:41:42product.
- 41:41:44And uh I'll show you what that means in
- 41:41:46just a second. So we have our two dice.
- 41:41:48We have dice A and it's going to be a
- 41:41:51set of values. Um they can only have one
- 41:41:53value for each one. That's why they put
- 41:41:55it in a set. And if you remember from
- 41:41:57range, it is up to seven. So this is
- 41:42:00going to be 1 2 3 4 5 6. It will not
- 41:42:03include the seven. And the same thing
- 41:42:05for our dice B.
- 41:42:08And then we're going to do is we're
- 41:42:09going to create a list which is the
- 41:42:12product of A and B. So what's um a + b?
- 41:42:17And if we go ahead and run this uh it'll
- 41:42:19print that out. And you'll see um in
- 41:42:22this case when they say product because
- 41:42:23it's an iteration tool,
- 41:42:27we're talking about creating a tupole of
- 41:42:29the two. So we've now created a tupole
- 41:42:31of all possible outcomes of the dice
- 41:42:34where dice A is one to three one to six
- 41:42:37and dice B is 1 to six. And you can see
- 41:42:38one to one, one to two, one to three and
- 41:42:40so forth. You remember we had a slide on
- 41:42:42this earlier where we talked about um
- 41:42:45the different all the different outcomes
- 41:42:47of a dice. We can play around with this
- 41:42:50a little bit. Uh we can do in dice
- 41:42:52equals two divi dice faces 1 2 3 4 5 6.
- 41:42:57Uh another way of doing what we did
- 41:42:59before and then we can create an event
- 41:43:00space where we have a set which is the
- 41:43:03product of the dice faces repeat equals
- 41:43:06end dice. And we'll go ahead and just
- 41:43:07run this. And you can see here it just
- 41:43:10again puts it through all the different
- 41:43:12possible variables we can have. And then
- 41:43:15if we wanted to take the same uh set on
- 41:43:17here and print them all out like we had
- 41:43:20before uh we can just go through for
- 41:43:22outcome and event space. Outcome end
- 41:43:25equals. So the event space is creating
- 41:43:30a sequence and as you can see here when
- 41:43:32we print it out it stacks them versus
- 41:43:34going through and putting them in a nice
- 41:43:36line.
- 41:43:38and we'll go ahead and do something. Um,
- 41:43:40let's go print. Since we have the end
- 41:43:43printing with a comma, that just means
- 41:43:45it's just going to it's not going to hit
- 41:43:47the return going down to the next line.
- 41:43:49Uh, and we'll go ahead and do the length
- 41:43:54of our event space. Uh, that'll be an
- 41:43:57important variable we're going to want
- 41:43:58to know in a minute.
- 41:44:01And of course, if I get carried away
- 41:44:02with my typing of length, uh, we'll
- 41:44:04print it twice and it'll give me an
- 41:44:06error. Uh so we have 36 different
- 41:44:09possible variations here
- 41:44:12and we might want to calculate something
- 41:44:14like um what about the multiple of
- 41:44:16three? What if we want to have
- 41:44:19uh the probability of the multiple of
- 41:44:21three in our setup?
- 41:44:25And so uh we can put together the code
- 41:44:27for the outcome in event space of xy
- 41:44:30equals outcome if x + y
- 41:44:34remainder 3. So, we're going to divide
- 41:44:36by three and look at the remainder and
- 41:44:37it equals zero.
- 41:44:40Then it's a favorable outcome and we're
- 41:44:42going to pop that outcome on the end
- 41:44:43there.
- 41:44:45And we'll turn it into a set. So, the
- 41:44:47favor outcome equals a set. Not
- 41:44:50necessary uh because we know it's not
- 41:44:52going to be repeating itself, but just
- 41:44:54in case, we'll go ahead and do that.
- 41:44:58And if we want to print out the outcome,
- 41:45:01we can go ahead and see what that looks
- 41:45:03like. And you can see here these are all
- 41:45:05uh multiples of three. Uh 1 plus 2 is 3,
- 41:45:085 + 4 is 9, which divided by 3 is 3, and
- 41:45:11so forth.
- 41:45:15And just like we looked up the length uh
- 41:45:17of the one before, let's go ahead and
- 41:45:18print the length of our f outcome so we
- 41:45:24can see what that looks like.
- 41:45:30There we go.
- 41:45:32And of course, I did forget to add the
- 41:45:34print in the middle because we're
- 41:45:35looping through and putting an end on
- 41:45:37the on the setup on there. So, we're
- 41:45:38going to put the print in there. And if
- 41:45:40I run this, you can see um
- 41:45:46we end up with 12. So, we have 36 total
- 41:45:49options. Uh we have 12 that are multiple
- 41:45:53that um add up to a multiple of three.
- 41:45:57And we can easily conver compute the
- 41:45:59probability of this uh by simply taking
- 41:46:02the length of our favorable outcome over
- 41:46:05the length of the event space.
- 41:46:09And if we print it out, let me put that
- 41:46:10in there. Probability
- 41:46:13last line. So we just type it in. We end
- 41:46:15up with a 3333 chance. And it's roughly
- 41:46:19a third.
- 41:46:21And we might want to make this look
- 41:46:23nice. So let's go ahead and put in
- 41:46:24another line there. The probability of
- 41:46:26getting the sum which is a multiple of
- 41:46:27three is
- 41:46:303333.
- 41:46:34We can compute the same thing for five
- 41:46:36dice.
- 41:46:38And if we do this for five dice and go
- 41:46:40ahead and run it, you can see we just
- 41:46:42have a huge amount of choices. So it
- 41:46:45just goes on and on down here. And we
- 41:46:47can look at the uh length of the event
- 41:46:51space.
- 41:46:59And we have over 7,776
- 41:47:02choices. That's a lot of choices.
- 41:47:05And if we want to ask the question like
- 41:47:07we did above, uh what is the sum where
- 41:47:10the sum is a multiple of five but not a
- 41:47:12multiple of three? We can go through all
- 41:47:15of these different options. And then uh
- 41:47:18you can see here uh d1 d2 d3 d4 d5
- 41:47:22equals the outcome. And if uh you add
- 41:47:24these all together and the
- 41:47:27division by five does not have a
- 41:47:29remainder of zero but the remainder is
- 41:47:32also of a division by three is not equal
- 41:47:35to zero. So the multiple of five is
- 41:47:38equal to zero but the multiple of three
- 41:47:39is not. We can just appin that on here
- 41:47:42and then we can look at that uh
- 41:47:44favorable outcome. We'll go ahead and
- 41:47:47set that and we'll just take a look at
- 41:47:48this. What's our length of our favorable
- 41:47:52outcome?
- 41:47:58It's always good to see what we're
- 41:47:59working with. And so we have 94 out of
- 41:48:02776.
- 41:48:06And then of course we can just do a
- 41:48:08simple division to get the probability
- 41:48:10on here. What's the probability that
- 41:48:11we're going to roll a multiple of five
- 41:48:14when you add them together?
- 41:48:16but not a multiple of three. And so
- 41:48:19we're just going to divide those two
- 41:48:20numbers. And you can see here we get
- 41:48:22uh.16255
- 41:48:24or 11.62%.
- 41:48:30And so you can really have a nice visual
- 41:48:32that this is not really complicated math
- 41:48:35right here on probabilities. uh it's
- 41:48:37just how many options do you have and
- 41:48:39how many of those are you possibly going
- 41:48:41to be able to um come up with with the
- 41:48:44solution you're looking for. And this
- 41:48:46leads us to a confusion matrix. A
- 41:48:49confusion matrix is a table which is
- 41:48:51used to describe the performance of a
- 41:48:53classification model on a set of test
- 41:48:55data for which the true values are
- 41:48:57known. And so you'll see on the left we
- 41:49:00have the predicted and the actual and we
- 41:49:03have a negative uh false negative
- 41:49:06positive true positive
- 41:49:08um and then we have false positive and
- 41:49:11true negative. And you can think of this
- 41:49:14as your predicted model. What does that
- 41:49:17mean? That means if you divided your
- 41:49:19data and you use twothird of it to
- 41:49:21create the model, you might then test it
- 41:49:24against an actual case for the last
- 41:49:25third to see how well it comes out. How
- 41:49:27many times was it uh true positive
- 41:49:30versus uh false positive? It gave a
- 41:49:33false positive response. And you can
- 41:49:35imagine in medical uh situations, this
- 41:49:38is a pretty big deal. You don't want to
- 41:49:40give a false positive. So you might
- 41:49:42adjust your model accordingly so you
- 41:49:44don't have a false positive. Say with a
- 41:49:46co virus test, it'd be better to have a
- 41:49:48false negative and then go back and get
- 41:49:50retested than to have 30% false
- 41:49:53positives where then the test is pretty
- 41:49:55much invalid. So in a use case uh like
- 41:49:58cancer prediction, let's consider an
- 41:50:00example where a cancer prediction model
- 41:50:02is put to the test for its accuracy and
- 41:50:04precision. Actual result of a person's
- 41:50:07medical report is compared with the
- 41:50:09prediction made by the machine learning
- 41:50:11model. And so you can see here here's
- 41:50:13our actual predicted uh whether they
- 41:50:15have cancer or not. You know cancer a
- 41:50:17big one. You don't want to have a uh
- 41:50:19false positive. I mean a false negative.
- 41:50:22In other words, you don't want to have
- 41:50:23it tell you that you don't have cancer
- 41:50:25when you do. So that would be something
- 41:50:27you'd really be looking for in this
- 41:50:29particular domain. You don't want a
- 41:50:31false negative. Uh and this is again,
- 41:50:34you know, you've created a model, you
- 41:50:35have hundreds of people or thousands of
- 41:50:38pieces of data that come in. There's a
- 41:50:40real famous case study where they have
- 41:50:42the imagery and all the measurements
- 41:50:43they take and there's about 36 different
- 41:50:45measurements they take. And then if you
- 41:50:48run the a basic model, you want to know
- 41:50:50just how accurate it is. How many um
- 41:50:52negative results do you have that are
- 41:50:54either telling people they have cancer
- 41:50:56that don't or telling people that don't
- 41:50:57have cancer that they do? And then we
- 41:50:59can take these numbers and we can feed
- 41:51:02them into our accuracy, our precision,
- 41:51:04and our recall. Uh so accuracy,
- 41:51:07precision, and recall, accuracy metric
- 41:51:09to measure how accurately the results
- 41:51:11are predicted. And this is your um total
- 41:51:15um true where you got the right results.
- 41:51:17you add them together, the true
- 41:51:18positive, the true negative over all the
- 41:51:21results. So what percentage of them were
- 41:51:23accurate versus what were wrong. We talk
- 41:51:26about precision is a metric to measure
- 41:51:28how many of the correctly predicted
- 41:51:30cases are actually turned out to be
- 41:51:32positive. Uh so we have a precision on
- 41:51:36true positive. Again, if you're talking
- 41:51:38about like uh COVID testing with the
- 41:51:42viruses, uh you really want this to be a
- 41:51:44a high number. you want this true um
- 41:51:47that to be the center point where you
- 41:51:49might have the opposite if you're
- 41:51:51dealing with cancer where you want no
- 41:51:53false negatives. Uh so this is your
- 41:51:56metric on here. Precision is your test
- 41:51:58positive uh true positive plus uh false
- 41:52:02positive. And then your recall how many
- 41:52:05of the actual positive cases we were
- 41:52:07able to predict quickly with our model.
- 41:52:09Uh so test positive is the test positive
- 41:52:12plus the false negative on there. And
- 41:52:15we'll want to go ahead and do a demo on
- 41:52:17the naive bay classifier. Before I get
- 41:52:21too far into uh naive baze classifier
- 41:52:24because we're going to pull it from the
- 41:52:25sklearn or the scikit. Um let's go ahead
- 41:52:30kind of an interesting page here for
- 41:52:31classifiers. When you go into the
- 41:52:33sklearn kit, there's a lot of ways to do
- 41:52:35classification. I'll just zoom up in
- 41:52:37here so you can see some of the titles.
- 41:52:40Uh there's everything from the nearest
- 41:52:42neighbor linear
- 41:52:44uh but we're going to be focusing on the
- 41:52:46naive bays over here. And this is just
- 41:52:50um a sample data set that they put
- 41:52:52together. And you can see how some of
- 41:52:54these have a very different output. The
- 41:52:57naive bay remember is set up as probably
- 41:53:00the most simplified uh calculator or um
- 41:53:03set of predictions out there. And so
- 41:53:05what we've been talking about with the
- 41:53:07true false and stuff like that where
- 41:53:08there's a uh
- 41:53:11an belief that there is a independent
- 41:53:13assumption between the features where
- 41:53:14the features are very assumed to have
- 41:53:17some kind of connection uh then we can
- 41:53:19go ahead and use that for the
- 41:53:21prediction. And so that's what we're
- 41:53:23using as a naive bay classifier versus
- 41:53:26many of the other classifiers that are
- 41:53:27out there.
- 41:53:30For this we're going to use uh the
- 41:53:32social network ads. It's a little data
- 41:53:35set on here and let me go and just open
- 41:53:38that up the file. Uh here we go. It has
- 41:53:42user ID, gender, age, estimated salary,
- 41:53:45uh purchased. And so we have you can see
- 41:53:48the user ID, male 19, uh estimated
- 41:53:52salary 19,000 and purchased zero. Uh so
- 41:53:56it's either going to make a purchase or
- 41:53:57not. So look at that last one. 01. We
- 41:54:01should be thinking of binomials. we
- 41:54:03should be thinking of simple naive base
- 41:54:05classifier kind of setup.
- 41:54:09So if we close this out, we're going to
- 41:54:10go ahead and import our numpy as np.
- 41:54:14We're nice to have a a good visual of
- 41:54:16our data. So we'll put in our mattplot
- 41:54:18library. Here's our pandas, our data
- 41:54:21frame.
- 41:54:23Uh and then we're going to go ahead and
- 41:54:24import the data set. And the data set's
- 41:54:26going to be we're going to read it from
- 41:54:28the social network ads.csv. Then we're
- 41:54:31going to print the head just so you can
- 41:54:32see it again uh even though I showed you
- 41:54:34it in the file. And X equals the data
- 41:54:37set I location uh two three values and Y
- 41:54:40is going to be the four uh column 4. Let
- 41:54:43me just run this so it's a little easier
- 41:54:45to go over that. Um you can see right
- 41:54:47here we're going to be looking at uh 012
- 41:54:50is age and estimated salary. So 2 three
- 41:54:54and that's what I location just means um
- 41:54:57that we're looking at the number versus
- 41:55:00a regular location. Uh regular location
- 41:55:02you'd actually say age and estimated
- 41:55:04salary.
- 41:55:06And then column four is did they make a
- 41:55:08purchase? They purchased something. Uh
- 41:55:10so those are the three columns we're
- 41:55:12going to be looking at when we do this.
- 41:55:13And we've gone ahead and imported these
- 41:55:15and imported the data. So now our data
- 41:55:17set is all set with this information in
- 41:55:19it.
- 41:55:23And we'll need to go ahead and split the
- 41:55:24data up. Uh so we need our from the
- 41:55:26sklearn model selection we can import
- 41:55:29train test split. Uh this does a nice
- 41:55:32job. We can set the random state so it
- 41:55:34randomly picks the data. And we're just
- 41:55:36going to take uh 25% of it is going to
- 41:55:39go into the test our x test and our y
- 41:55:41test and the 75% will go to x train and
- 41:55:44y train. That way once we create our
- 41:55:48model, we can then have data to see just
- 41:55:50how accurate or how well it has
- 41:55:52performed with our um prediction.
- 41:55:56The next step in pre-processing our data
- 41:55:59is to go ahead and do feature scaling.
- 41:56:02Now, a lot of this is start to look
- 41:56:04familiar. If you've done a number of the
- 41:56:06other modules and setup, you should
- 41:56:08start noticing that we bring in our
- 41:56:10data. We take a look at what we're
- 41:56:12working with. uh we go ahead and split
- 41:56:14it up into training and testing. Uh in
- 41:56:17this case, we're going to go ahead and
- 41:56:18scale it. Scale it means we're putting
- 41:56:20it between a value of minus1 and one uh
- 41:56:24or someplace in that middle ground
- 41:56:26there. This way, if you have any huge
- 41:56:28set, you don't have this huge um setup.
- 41:56:31If we go back up to here where salary uh
- 41:56:33salary is 20,000 versus age 35, well,
- 41:56:39there's a good chance with a lot of the
- 41:56:40back-end math that 20,000 will skew the
- 41:56:43results and the estimated salary will
- 41:56:45have a higher impact than the age
- 41:56:47instead of balancing them out and
- 41:56:48letting the calculations weigh them
- 41:56:50properly.
- 41:56:52And finally, we get to actually create
- 41:56:55our naive bay model.
- 41:56:58Um, and then we're going to go ahead and
- 41:57:00import the Gazian naive bays.
- 41:57:04And the Gazian is is uh the most basic
- 41:57:07one. That's what we're looking at now.
- 41:57:09It turns out though, if you go to the SK
- 41:57:12um learn kit, uh they have a number of
- 41:57:15different ones you can pull in there.
- 41:57:17There's a um Bernoli. I I've never used
- 41:57:20that one. Categorical
- 41:57:22um compliment. And here's our Gazian. Uh
- 41:57:25so there's a number of different options
- 41:57:27you can look at. Gazian when you come to
- 41:57:30the naive bays is the most commonly
- 41:57:32used. Uh so we're talking about the
- 41:57:34naive bays that's usually what people
- 41:57:36are talking about when they when they're
- 41:57:37pulling this in. And one of the nice
- 41:57:39things about the gazian if you go to
- 41:57:41their website um to sklearn the naive
- 41:57:44bay gazian there's a lot of cool
- 41:57:46features. One of them is you can do
- 41:57:47partial fit on here. Um that means if
- 41:57:50you have a huge amount of data, you
- 41:57:51don't have to process it all at on you
- 41:57:53once. You can batch it into the Gausian
- 41:57:57uh NB model. And there's many other
- 41:57:59different things you can do with it as
- 41:58:01far as fitting the data and how you um
- 41:58:03manipulate it. We're just doing the
- 41:58:06basics. So we're going to go ahead and
- 41:58:07create our classifier. We're going to
- 41:58:09equal the Gausian NB.
- 41:58:12And then we're going to do a fit. We're
- 41:58:13going to fit our training data and our
- 41:58:15training solution. So, X-Rain, Y train,
- 41:58:20and we'll go ahead and run this. Uh,
- 41:58:22it's going to tell us that it it ran the
- 41:58:24code right there.
- 41:58:26And now we have our trained classifier
- 41:58:29model. So, the next step is we need to
- 41:58:32go ahead and run a prediction. We're
- 41:58:33going to do our Y predict equals the
- 41:58:35classifier.predict
- 41:58:37X test. So, here we fit the data and now
- 41:58:40we're going to go ahead and predict.
- 41:58:45And now we get to our confusion matrix.
- 41:58:49Uh so from the sklearn matrix metrics
- 41:58:52you can import your confusion matrix
- 41:58:54just as saves you from doing all the
- 41:58:56simple math. It does it all for you. And
- 41:58:59then we'll go ahead and create our
- 41:59:00confusion metrics with the y test and
- 41:59:02the y predict. So we have our actual and
- 41:59:05we have our predicted value.
- 41:59:08And you can see from here this is the
- 41:59:10chart we looked at. Here's predicted.
- 41:59:11So, true positive, false positive, false
- 41:59:14negative, true negative.
- 41:59:18And if we go ahead and run this, there
- 41:59:20we have it. 653725.
- 41:59:24And in this particular uh prediction, we
- 41:59:27had 65 uh or predicted the truth as far
- 41:59:30as a a purchase. They're going to make a
- 41:59:32purchase, and we guessed three wrong.
- 41:59:35And then we had 25 we predicted would
- 41:59:37not purchase, and seven of them did. So,
- 41:59:40there's our our confusion matrix.
- 41:59:44At this point, if you were uh with your
- 41:59:46shareholders or a board meeting, um you
- 41:59:49would start to hear some snoozing if
- 41:59:50they were looking at the numbers and you
- 41:59:52say, "Hey, here's my confusion mat uh
- 41:59:54matrix." So, let's go ahead and
- 41:59:56visualize the results.
- 41:59:58We're going to pull from the map plot
- 42:00:00library colors import listed color map.
- 42:00:04Um, and this is actually my machine's
- 42:00:06going to throw an error because this is
- 42:00:09being um because of the way the setup
- 42:00:12is. I have a newer version on here than
- 42:00:14when they put together the demo. And we
- 42:00:17need our um X set and our Y set, which
- 42:00:19is our X train and Y train. And then
- 42:00:22we'll create our X1, X2. And we'll put
- 42:00:25that into a grid. Uh, and we set our X
- 42:00:28set minimum stop and our X set max stop.
- 42:00:31And if you come all the way over here,
- 42:00:33we're going to step 0. 001. This is
- 42:00:35going to give us a nice line, uh, is
- 42:00:37what that's doing. And then we're going
- 42:00:38to plot the contour, uh, plot the x
- 42:00:41limit, plot the y limit, and put the
- 42:00:44scatter plot in there. And let's go
- 42:00:46ahead and run this. Uh, to be honest,
- 42:00:48when I'm doing these graphs, there's so
- 42:00:50many different ways to do that. There's
- 42:00:52so many different ways to put this code
- 42:00:53together to show you what we're doing.
- 42:00:56it's uh a lot easier to pull up the
- 42:00:59graph and then go back up and explain
- 42:01:00it. So the first thing we want to note
- 42:01:04here when we're looking at the data
- 42:01:07is this is the training set.
- 42:01:10And so we have those who didn't make a
- 42:01:12purchase. We've drawn a nice area for
- 42:01:14that that's defined by the naive bay
- 42:01:17setup. And then we have those who did
- 42:01:19make a purchase, the green. And you can
- 42:01:21see that some of the green dots fall
- 42:01:23into the red area and some of the red
- 42:01:25dots fall into the green. So even our
- 42:01:27training set isn't going to be 100%. Uh
- 42:01:30we couldn't do that. And so we're
- 42:01:32looking at our different data coming
- 42:01:34down. Uh we can kind of arrange our x1
- 42:01:37x2 so we have a nice plot going on. And
- 42:01:40we're going to create the um contour.
- 42:01:43That's that nice line that's drawn down
- 42:01:45the middle on here with the red green.
- 42:01:47Um that's what that's what this is doing
- 42:01:49right here with the reshape and notice
- 42:01:51that we had to uh do the t if you
- 42:01:55remember from numpy um if you did the
- 42:01:57numpy module um you end up with pairs
- 42:02:00you know x uh x1 x2 x1 x2 next row and
- 42:02:05so forth you have to flip it so it's all
- 42:02:07one row you have all your x1's and all
- 42:02:09your x2s. Um so this what we're kind of
- 42:02:11looking for right here on this setup.
- 42:02:15Uh, and then the scatter plot is of
- 42:02:17course um your scattered data across
- 42:02:19there. We're just going through all the
- 42:02:20points that puts these nice little dots
- 42:02:22onto our setup on here. And we have our
- 42:02:25estimated salary and our H. And then of
- 42:02:27course the dots are did they make a
- 42:02:29purchase or not. And just a quick note,
- 42:02:31this is kind of funny. You can see up
- 42:02:33here where it says X set Y set equals uh
- 42:02:36X train Y train, which seems kind of a
- 42:02:39little weird to do. Um, this is because
- 42:02:42this is probably originally a
- 42:02:43definition. Uh, so it's its own module
- 42:02:46that could be called over and over
- 42:02:47again. And which is really a good way to
- 42:02:50do it because the next thing we're going
- 42:02:51to want to do is do the exact same
- 42:02:52thing, but we're going to visualize the
- 42:02:55test set results. Uh, that way we can
- 42:02:57see what happened with our test group,
- 42:02:59our 25%.
- 42:03:02And you can see down here we have um the
- 42:03:04test set. Uh, and it, if you look at the
- 42:03:07two graphs next to each other, this one
- 42:03:09obviously has um 75% of the data, so
- 42:03:12it's going to show a lot more. This is
- 42:03:14only 25% of the data. You can see that
- 42:03:17there's a number that are kind of on the
- 42:03:19edge as to whether they could guess by
- 42:03:21age and income they're going to make a
- 42:03:22purchase or not. U, but that said, it
- 42:03:25still is pretty clear. It's pretty good
- 42:03:27as far as how much the estimate is and
- 42:03:28how good it does.
- 42:03:31Now, graphs are really effective for
- 42:03:34showing people what's going on, but you
- 42:03:37also need to have the numbers. And so,
- 42:03:39we're going to do from sklearn, we're
- 42:03:41going to import metrics, and then we're
- 42:03:43going to print our metrics
- 42:03:44classification port from the Y test and
- 42:03:46the Y predict.
- 42:03:49And you can see here we have precision
- 42:03:52uh precision of zeros is 90. There's our
- 42:03:54recall
- 42:03:5696. We have an F1 score and a support.
- 42:04:00And we have our precision, the recall on
- 42:04:03getting it right. Uh, and then we can do
- 42:04:05our accuracy, the macro average, and the
- 42:04:07weighted average. Uh, so you can see it
- 42:04:10pulls in pretty good as far as um how
- 42:04:13accurate it is. You could say it's going
- 42:04:15to be about 90% is going to guess
- 42:04:18correctly um that it that they're not
- 42:04:21going to purchase. And we had an 89%
- 42:04:23chance that they are going to purchase.
- 42:04:25Um, and then the other numbers as you
- 42:04:27get down have a little bit different
- 42:04:29meaning, but it's pretty straightforward
- 42:04:31on here. Here's our accuracy, and here's
- 42:04:33our micro average, and the weighted
- 42:04:35average, and everything else you might
- 42:04:36need. And if you forgot the exact
- 42:04:38definition of accuracy, it is the true
- 42:04:42positive, true negative over all of the
- 42:04:44different setups. Precision is your true
- 42:04:47positive over all positives, true and
- 42:04:50false. And recall is a true positive
- 42:04:53over true positive plus false negative.
- 42:04:56And we can just real quick flip back
- 42:04:58there so you can see those numbers on
- 42:05:00here. Uh here's our precision, here's
- 42:05:03our recall, and here's our accuracy on
- 42:05:06this.
- 42:05:07>> Welcome to this exciting journey into
- 42:05:09the world of statistics for data
- 42:05:11science. Have you ever wondered how data
- 42:05:13transforms from raw numbers into
- 42:05:15powerful insights that drive decisions?
- 42:05:18Well, statistics is the magic behind it
- 42:05:21all. Today, we will uncover how
- 42:05:23statistical methods help us summarize
- 42:05:25data, model uncertaintity, test
- 42:05:28hypothesis, and find relationships that
- 42:05:31can predict the future. So, buckle up.
- 42:05:33We are about to turn numbers into
- 42:05:35knowledge. Without any further ado,
- 42:05:37let's get started. Now, to start off,
- 42:05:40here's a key question. What are
- 42:05:42statistics in data science? Now,
- 42:05:44statistics is the science of collecting,
- 42:05:46analyzing and interpreting data. By
- 42:05:49applying statistical methods, we can
- 42:05:51uncover patterns in the data and make
- 42:05:54informed decisions. Now, as we continue,
- 42:05:57you notice how these essential concepts
- 42:05:59will form the backbone of many data
- 42:06:02science practices. Let's explore the key
- 42:06:04functions of statistics in data science.
- 42:06:07First, statistic helps summarize data
- 42:06:10using measures like mean, median and
- 42:06:13variance. Next, it models uncertaintity
- 42:06:16with probability and distributions. So,
- 42:06:19we can better understand risk and
- 42:06:21variability in our data. It also tests
- 42:06:25hypothesis such as when we use AB
- 42:06:28testing to compare different outcomes.
- 42:06:30Statistics finds relationship through
- 42:06:33methods like regression and correlation
- 42:06:36revealing how variables impact each
- 42:06:38other. And finally, all these tools
- 42:06:41enable datadriven decision-m turning raw
- 42:06:44numbers into actionable insights. Now,
- 42:06:47let's talk about why does statistics
- 42:06:49matter in data science. Statistics form
- 42:06:51the backbone of data science, providing
- 42:06:54the mathematical framework needed to
- 42:06:56make sense of data and draw reliable
- 42:06:59conclusions. Without statistics, it
- 42:07:01would be impossible to turn raw data
- 42:07:03into meaningful insights or make
- 42:07:05confident evidence-based decisions. Now
- 42:07:09let's have a look at the main branches
- 42:07:11of statistics. Descriptive statistics
- 42:07:14and inferential statistic. So first
- 42:07:16let's talk about the definition. Then we
- 42:07:19have got methods and measures. Now in
- 42:07:21the case of descriptive statistics, now
- 42:07:24let's have a look at the main branches
- 42:07:26of statistics. So basically there are
- 42:07:28two core branches. Descriptive
- 42:07:30statistics and inferial statistics.
- 42:07:33Descriptive statistics summarizes and
- 42:07:36describes data using measures like mean,
- 42:07:39median, mode, range and standard
- 42:07:41deviation. Its purpose is to organize
- 42:07:43and present data typically with charts,
- 42:07:46graph or summary tables. The scope of
- 42:07:49descriptive statistics is limited to the
- 42:07:52sample data itself. Now on the other
- 42:07:54hand, inferential statistics makes
- 42:07:56inferences about populations based on
- 42:07:59samples. It uses methods like hypothesis
- 42:08:03testing, confidence intervals and
- 42:08:05regression. The purpose here is to draw
- 42:08:08conclusion and make predictions with
- 42:08:11common examples including AB testing and
- 42:08:14survey analysis. The scope of inferial
- 42:08:17statistics extends beyond the sample to
- 42:08:20the larger population. Understanding
- 42:08:22both these branches is essential for
- 42:08:25analyzing and interpreting data in any
- 42:08:28data science project. Now let's take a
- 42:08:30closer look at the descriptive
- 42:08:32statistics starting with measures of
- 42:08:35central tendency. The first measure here
- 42:08:37is mean which is the average of all the
- 42:08:40values simply calculated as the sum of
- 42:08:43all the data points divided by the total
- 42:08:46count. Next is the median which
- 42:08:48represents the middle value when the
- 42:08:50data is sorted from lowest to highest.
- 42:08:53This is especially useful when dealing
- 42:08:55with skewed distributions. And finally,
- 42:08:58the mode is the most frequently
- 42:09:00occurring value in the data set, helping
- 42:09:03us identify common patterns or repeated
- 42:09:06outcomes. Let's continue our deep dive
- 42:09:09into descriptive statistics by looking
- 42:09:11at the measures of variability. First,
- 42:09:14we've got the range. This is simply the
- 42:09:16difference between the maximum and
- 42:09:18minimum value in a data set showing us
- 42:09:21the spread of our data which measures
- 42:09:23the average of the squared differences
- 42:09:26from the mean. This tells us how much
- 42:09:29the values in our data set differ from
- 42:09:31the average. Closely related is the
- 42:09:34standard deviation which is the square
- 42:09:36root of the variance. It gives us a more
- 42:09:39intuitive sense of how much the values
- 42:09:43typically deviate from the mean. And
- 42:09:45finally, we've got the interquartile
- 42:09:47range or we say IQR. This shows us the
- 42:09:50range of the middle 50% of our data,
- 42:09:54helping us understand how data is
- 42:09:56distributed across the center and avoid
- 42:09:59the effect of outliers. Understanding
- 42:10:02these four measures allows us to
- 42:10:04summarize not just the center of our
- 42:10:06data, but how spread out and varied our
- 42:10:10data set is. Now let's look at some
- 42:10:12practical applications of descriptive
- 42:10:15statistics. One major use is the data
- 42:10:18exploration and summarization where we
- 42:10:20quickly get an overview and basic
- 42:10:23understanding of complex data sets
- 42:10:25helping track performance and detect
- 42:10:27problems early in fields like
- 42:10:29manufacturing or operations. And
- 42:10:32finally, they are central to business
- 42:10:34reporting and dashboards where concise
- 42:10:37summaries are essential for managers to
- 42:10:40review trends and make datadriven
- 42:10:42decisions. So in short, descriptive
- 42:10:45statistics help transform raw data into
- 42:10:48clear actionable information across many
- 42:10:51business, scientific and operational
- 42:10:53context. Let's explore the first type of
- 42:10:56data in statistics, qualitative or
- 42:10:59categorical data. This type of data
- 42:11:01includes descriptive information that
- 42:11:04cannot be measured numerically such as
- 42:11:06categories or labels and yes or no
- 42:11:09responses. These are all about qualities
- 42:11:12or characteristics rather than
- 42:11:14quantities. Qualitative data can be
- 42:11:17further divided into two types. Nominal
- 42:11:20where the categories have no specific
- 42:11:22order and ordinal where the categories
- 42:11:25do have an order or ranking. Recognizing
- 42:11:28and classifying qualitative data is very
- 42:11:31important as it affects how information
- 42:11:33is analyzed and interpreted in
- 42:11:36statistics. Now let's discuss the second
- 42:11:38main type of data in statistics which is
- 42:11:41quantitative or numerical data. This
- 42:11:44type of data includes information that
- 42:11:46can be measured and expressed with
- 42:11:48numbers making it ideal for mathematical
- 42:11:51analysis. Quantitative data is further
- 42:11:54divided into two main categories. First,
- 42:11:57there is discrete data. These are
- 42:12:00countable values with specific fixed
- 42:12:02points such as the number of students,
- 42:12:05cars sold or website clicks. Second,
- 42:12:08there is continuous data which includes
- 42:12:10infinite possible values within a given
- 42:12:13range. Examples of continuous data
- 42:12:16include height, weight, temperature or
- 42:12:18time. Understanding the distinction
- 42:12:21between discrete and continuous data is
- 42:12:23very important as it determines which
- 42:12:26statistical methods and visualizations
- 42:12:29will be most appropriate. Now let's dive
- 42:12:31into the fundamentals of probability.
- 42:12:34Probability measures the likelihood of
- 42:12:36an event occurring and it's always
- 42:12:38expressed as a value between 0 and 1.
- 42:12:42Here are some key concepts. If the
- 42:12:44probability or P equals to zero, that
- 42:12:47means the event will never occur. If P
- 42:12:50is equals to 1, the event will always
- 42:12:53occur. And if P is equals to 0.5, the
- 42:12:56event has an equal chance of occurring
- 42:12:58or not occurring. It's truly a 50/50%
- 42:13:02scenario. Now, understanding these basic
- 42:13:04principle help us quantify uncertaintity
- 42:13:07and make informed predictions about
- 42:13:10future outcomes. Now let's look at the
- 42:13:13different types of probability. First
- 42:13:15there's classical probability. This is
- 42:13:18based on equally likely outcomes such as
- 42:13:20flipping a fair coin or rolling a
- 42:13:23balanced die. Next empirical probability
- 42:13:26which relies on observed frequency. It's
- 42:13:30calculated from actual data such as the
- 42:13:32proportion of rainy days over the past
- 42:13:35month. Let's say for example the
- 42:13:37probability of it's raining today given
- 42:13:40that it's cloudy. Understanding these
- 42:13:43three type help us choose the right
- 42:13:45approach for different situations.
- 42:13:47Whether we are predicting outcomes,
- 42:13:50analyzing data or making decisions under
- 42:13:53uncertaintity. Now probability has a
- 42:13:55wide range of powerful applications in
- 42:13:58data science. First of all, it is used
- 42:14:00in predictive modeling and machine
- 42:14:02learning where algorithms estimate
- 42:14:05future outcomes based on existing data.
- 42:14:08Probability also plays a key role in
- 42:14:11risk management and decision making
- 42:14:13helping businesses and researchers
- 42:14:15evaluate the likelihood of different
- 42:14:18scenarios and plan accordingly.
- 42:14:21Conditional probability calculating the
- 42:14:23chance of one event given that the
- 42:14:26another has occurred. This is crucial in
- 42:14:29fields like healthcare, fraud detection
- 42:14:31and marketing analytics.
- 42:14:35And finally, probability is foundational
- 42:14:37in AB testing and experimental design,
- 42:14:41allowing us to measure the effectiveness
- 42:14:43of new strategies or products. These
- 42:14:46applications show how probability
- 42:14:48enables smarter evidence-driven progress
- 42:14:51in modern data science. Let's look at
- 42:14:54one of the most common probability
- 42:14:56distributions, the normal distribution.
- 42:14:59This distribution is famous for its
- 42:15:01bell-shaped symmetric curve, which shows
- 42:15:04that most values cluster around the mean
- 42:15:07with fewer and fewer values appearing as
- 42:15:10you move away from the center. Now many
- 42:15:13natural phenomena like heights, test
- 42:15:16scores and measurement errors tend to
- 42:15:18follow this pattern making the normal
- 42:15:21distribution a key concept in
- 42:15:23statistics. It's characterized by two
- 42:15:26main parameters. The mean which
- 42:15:28determines the center of the curve and
- 42:15:30the standard deviation which controls
- 42:15:33its spread.
- 42:15:34Recognizing the distribution help
- 42:15:36analysts make predictions, calculate
- 42:15:39probabilities, and apply statistical
- 42:15:41techniques to real world data. Next,
- 42:15:44let's explore the binomial distribution,
- 42:15:47which is another common probability
- 42:15:49distribution. The binomial distribution
- 42:15:51is discrete and is used for situations
- 42:15:54with binary outcomes like success or
- 42:15:57failure. It is based on a fixed number
- 42:16:00of trials where each trial has a
- 42:16:03constant probability of success such as
- 42:16:05flipping a coin a certain number of
- 42:16:07times or tracking pass fail rates.
- 42:16:10Examples of binomial experiment includes
- 42:16:13coin flips and measuring how many
- 42:16:15students pass or fail a test. This
- 42:16:18distribution help us model and analyze
- 42:16:20outcomes when only two possibilities
- 42:16:23exist in each trial. Now let's focus on
- 42:16:26the poison distribution. Another key
- 42:16:29type of probability distribution. The
- 42:16:31poion distribution is a discrete
- 42:16:33distribution specifically used to model
- 42:16:36rare events. It's particularly helpful
- 42:16:39for modeling events that occur
- 42:16:41independently over a fixed interval of
- 42:16:44time or space. For example, it predicts
- 42:16:47how many times an event like a customer
- 42:16:50arriving on a website, receiving a visit
- 42:16:53might happen in a certain period.
- 42:16:55Typical examples include customer
- 42:16:57arrivals at a store, defect rates, and
- 42:17:00manufacturing or counts of website
- 42:17:02visits over a set period. This
- 42:17:05distribution is great tool for
- 42:17:07understanding and predicting random
- 42:17:09independent events that don't happen
- 42:17:12very often, but are important to track.
- 42:17:14It's a branch that allows us to make
- 42:17:17inferences, predictions, or
- 42:17:19generalizations about a larger
- 42:17:21population using sample data. Here are
- 42:17:24some key concepts. First is the
- 42:17:26population versus sample. The population
- 42:17:29represents the entire group we want to
- 42:17:31know about while the sample is the
- 42:17:33subset we actually collect data from.
- 42:17:36Next concept is sampling distribution
- 42:17:39which refers to the distribution of a
- 42:17:41statistic across multiple samples from
- 42:17:44the same population. Standard error
- 42:17:47measures how much the sample statistic
- 42:17:49is expected to vary due to random
- 42:17:52sampling. And finally, margin of error
- 42:17:55tells us how much we can expect our
- 42:17:57estimates to differ from the true
- 42:18:00population value. Together these concept
- 42:18:03form the foundation for drawing reliable
- 42:18:05insights from sample data in inferential
- 42:18:08statistics.
- 42:18:10Let's review some of the main techniques
- 42:18:12used in inferial statistic. First, we've
- 42:18:15got hypothesis testing. This method
- 42:18:17allows us to test claims or ideas about
- 42:18:21population parameters based on sample
- 42:18:24data helping us determine if observed
- 42:18:27results are statistically significant
- 42:18:29and reliability of our estimate. Next
- 42:18:32are confidence intervals. These provide
- 42:18:34a range of likely values for a
- 42:18:36population parameter giving us a sense
- 42:18:39of possible variation reliability of our
- 42:18:42estimate. And lastly, regression
- 42:18:44analysis is used to model and analyze
- 42:18:47the relationships between variables,
- 42:18:49allowing us to make predictions, uncover
- 42:18:52trends, and understand how changes in
- 42:18:54one factor might affect another. Now,
- 42:18:57let's walk through the hypothesis
- 42:18:59testing process. The first step is to
- 42:19:02formulate hypothesis. Start with null
- 42:19:04hypothesis represented as Hnot, which
- 42:19:08states that there is no effect or
- 42:19:10difference. Then there's the alternative
- 42:19:13hypothesis represented as H1 which
- 42:19:16suggests that an effect or difference
- 42:19:19does exist. Now after setting up the
- 42:19:21hypothesis the next step is to choose a
- 42:19:24significance level noted by alpha.
- 42:19:27Common choices for significance levels
- 42:19:29include 0.05 5% 01 1% 010 which is 10%.
- 42:19:36The significance level is important
- 42:19:38because it controls the likelihood of
- 42:19:40making a type one error also known as
- 42:19:42false positive. And now each step in the
- 42:19:46process is critical for ensuring that
- 42:19:48statistical results are both meaningful
- 42:19:50and reliable. The next step is to
- 42:19:53collect and analyze sample data. Ensure
- 42:19:55the data is represented by choosing a
- 42:19:58good sample. Then calculate the
- 42:20:00appropriate and test statistic for your
- 42:20:03hypothesis test. Once that's done, it's
- 42:20:06time to make a decision. Compare the p
- 42:20:08value to the chosen significance level
- 42:20:11alpha. Now, if the p value is less than
- 42:20:13or equals to alpha, you reject the null
- 42:20:16hypothesis. If the p value is greater
- 42:20:19than alpha, you fail to reject the null
- 42:20:21hypothesis. And finally, interpret your
- 42:20:24results. Draw your conclusions with
- 42:20:26respect to the context and problem at
- 42:20:29hand. Always keeping the bigger picture
- 42:20:31in mind. Following these step helps
- 42:20:34ensure your hypothesis test is robust,
- 42:20:37clear and meaningful. Let's review the
- 42:20:40common types of hypothesis test used in
- 42:20:42statistic. The one sample test compares
- 42:20:45the mean of a sample to a known value
- 42:20:48and the one sample zed test is used when
- 42:20:51the population standard deviation is
- 42:20:53known. Next, we have two sample test.
- 42:20:56The independent samples t test compares
- 42:20:59to the means of two different groups
- 42:21:02while paired samples t test compares
- 42:21:04before and after measurements for the
- 42:21:07same subject. For categorical data test,
- 42:21:10the shear test checks for the
- 42:21:12independence or goodness of fit. And
- 42:21:16fiser's exact test is useful for small
- 42:21:19sample sizes. And lastly, non-parametric
- 42:21:22tests such as the man Whitney U test and
- 42:21:25Wil Coxson signed rank test serve as
- 42:21:28alternatives to the t test when data
- 42:21:31doesn't meet certain parametric
- 42:21:33assumptions. Choosing the right
- 42:21:35hypothesis test depends on your data
- 42:21:37type and the specific question you want
- 42:21:40to answer. Let's explore the central
- 42:21:42limit theorem or CLT of foundation for
- 42:21:45inferial statistics. The central limit
- 42:21:48theorem states that as the sample size
- 42:21:51increases, the distribution of sample
- 42:21:53means approaches a normal distribution
- 42:21:56even if the original data is a normally
- 42:21:59distributed. This sample works for any
- 42:22:01population distribution making it
- 42:22:04incredibly powerful. Now for good
- 42:22:06results, the sample size should
- 42:22:08typically be 30 or more. What's
- 42:22:11interesting is the sample mean
- 42:22:12distribution which will have the same
- 42:22:15mean as the population. The standard
- 42:22:18error which measures variability by the
- 42:22:20sample mean equals sigma / the square
- 42:22:23roo<unk> of n. And this gets smaller as
- 42:22:26the sample size grows. And finally the
- 42:22:29formula shown here which is z is equ= to
- 42:22:32x -
- 42:22:35sigma divided by sigma over the square
- 42:22:38root of n. Let's standardize and compare
- 42:22:41sample means. The CLT makes most
- 42:22:44parametric statistics possible and is
- 42:22:47the backbone for many statistical test.
- 42:22:49Let's review the main types of
- 42:22:51regression analysis which are used to
- 42:22:54model and understand relationships
- 42:22:56between variables. First up is linear
- 42:22:59regression. This technique analyzes the
- 42:23:02relationship between a continuous
- 42:23:04variable and another variable resulting
- 42:23:07in a straight line. It's commonly used
- 42:23:09in predicting sales or prices. Next is
- 42:23:13logistic regression. Unlike linear
- 42:23:15regression, this method is used for
- 42:23:18binary outcomes, helping estimate
- 42:23:20probabilities such as whether an email
- 42:23:23is spam or a patient has a disease.
- 42:23:26Moving to multiple regression, this
- 42:23:28allows us to account for the effect of
- 42:23:31several variables at once, modeling more
- 42:23:34complex relationships like determining
- 42:23:36house prices. Lastly, polomial
- 42:23:39regression which is used for nonlinear
- 42:23:42relationships. The resulting curve
- 42:23:44rather than a straight line lets us
- 42:23:46capture growth trends and other patterns
- 42:23:48that aren't linear. Understanding these
- 42:23:51types of regression help analysts choose
- 42:23:54the right model for the data and
- 42:23:56business's problem at hand. Let's
- 42:23:59clarify the important differences
- 42:24:00between correlation and causation.
- 42:24:03Correlation is statistical measure of
- 42:24:06how two variables move together.
- 42:24:08Correlation values range from minus 10
- 42:24:10to + one. But remember correlation does
- 42:24:14not imply causation. On the other hand,
- 42:24:16causation means that one variable
- 42:24:18actually causes changes in another.
- 42:24:21Establishing causation requires
- 42:24:24controlled experimentation and is much
- 42:24:27stronger relationship than simple
- 42:24:29correlation. Always be careful when
- 42:24:32interpreting results. Just because two
- 42:24:34variables move together doesn't mean one
- 42:24:36cause the other. Let's look at some
- 42:24:39common problems with interpreting
- 42:24:41correlation and causation. First is the
- 42:24:43third variable problem which happens
- 42:24:46when a hidden variable affects both
- 42:24:48variables in a question leading to a
- 42:24:50misleading connection. There's also this
- 42:24:52directionality problem where it's
- 42:24:54unclear which variable is causing others
- 42:24:57to change but both are actually caused
- 42:25:00by a third variable which is the hot
- 42:25:02weather. This demonstrate that
- 42:25:03correlation does not mean one variable
- 42:25:06causes the other. So always look for
- 42:25:08hidden factors before assuming
- 42:25:10causation. Let's talk about statistical
- 42:25:13errors specifically type one and type
- 42:25:15two errors. Type one error also called
- 42:25:18as false positive occurs when we reject
- 42:25:21the true null hypothesis. In other
- 42:25:23words, we wrongly conclude there's an
- 42:25:26effect when there's actually isn't. The
- 42:25:28probability of making this error is
- 42:25:31equal to the significance level alpha.
- 42:25:34For example, concluding a drug works
- 42:25:36when it actually doesn't. On the other
- 42:25:38hand, type two error is false negative
- 42:25:42means failing to reject a false null
- 42:25:44hypothesis. That's when we miss a real
- 42:25:47effect or difference. The probability of
- 42:25:50type two error is beta. For example,
- 42:25:53missing the real effect of a drug and
- 42:25:55seeing it doesn't work when it actually
- 42:25:57does. Understanding these errors is key
- 42:26:00for designing good experiments and
- 42:26:03interpreting statistical results
- 42:26:05properly. Let's look at three common
- 42:26:07sampling methods using statistics. The
- 42:26:10first one is random sampling. Here every
- 42:26:12individual in the population has a equal
- 42:26:15chance of being selected where every nth
- 42:26:18individual is chosen. Random sampling is
- 42:26:21crucial for ensuring representative
- 42:26:23sample. Next is stratified sampling. The
- 42:26:27population is divided onto homogeneous
- 42:26:29subgroups or strata and then a random
- 42:26:33sample is drawn from each group. This
- 42:26:35approach makes sure all the groups are
- 42:26:38represented in the sample. And finally,
- 42:26:40cluster sampling divides the population
- 42:26:43into clusters, often based on geography.
- 42:26:46Entire clusters are then randomly
- 42:26:48selected. It's cost effective and useful
- 42:26:51for large spread out populations.
- 42:26:54Choosing the right method ensures the
- 42:26:56data truly represents the whole
- 42:26:59population and strengthens the study's
- 42:27:01conclusions. So guys, let's wrap up with
- 42:27:04some real world applications of
- 42:27:06statistics. In business analytics,
- 42:27:08statistics are used for AB testing to
- 42:27:11optimize websites, customer segmentation
- 42:27:14and targeting, sales forecasting and
- 42:27:17demand planning and quality control and
- 42:27:20also process improvement. These
- 42:27:22techniques help businesses make smarter
- 42:27:24datadriven decisions every day.
- 42:27:27Statistics play a vital role in
- 42:27:29healthcare and medicine as well. They
- 42:27:32are key for analyzing clinical trial
- 42:27:34results, conducting epidemological
- 42:27:37studies, evaluating treatment
- 42:27:39effectiveness, and identifying risk
- 42:27:42factors. By using these approaches,
- 42:27:44healthcare researchers and practitioners
- 42:27:47can improve patient outcomes in public
- 42:27:49health. From businesses to medicine,
- 42:27:52statistics transform raw information
- 42:27:54into actionable insights that create
- 42:27:57real impact. Statistics has a huge
- 42:27:59impact in technology, data science and
- 42:28:02finance and power recommener systems
- 42:28:04that personalize what users see. In
- 42:28:07finance, statistic help with risk
- 42:28:10assessment and management, optimizing
- 42:28:12investment portfolios, determining
- 42:28:15credit scores, and supporting market
- 42:28:17research and analysis. Now, these
- 42:28:19applications show how statistical
- 42:28:22techniques help make smarter decisions
- 42:28:24and solve complex challenges across high
- 42:28:27impact industry. Are you one of the many
- 42:28:30who dreams of becoming a data scientist?
- 42:28:32Keep watching this video if you're
- 42:28:34passionate about data science because we
- 42:28:36will tell you how does it really work
- 42:28:38under the hood. Emma is a data
- 42:28:39scientist. Let's see how a day in her
- 42:28:41life goes while she's working on a data
- 42:28:44science project. Well, it is very
- 42:28:46important to understand the business
- 42:28:47problem first. In our meeting with the
- 42:28:49clients, Emma asks relevant questions,
- 42:28:52understands and defines objectives for
- 42:28:55the problem that needs to be tackled.
- 42:28:57She's a curious soul who asks a lot of
- 42:29:00wise, one of the many traits of a good
- 42:29:02data scientist. Now, she ges up for data
- 42:29:05acquisition to gather and scrape data
- 42:29:07from multiple sources like web servers,
- 42:29:10logs, databases, APIs, and online
- 42:29:12repositories. Oh, it seems like finding
- 42:29:15the right data takes both time and
- 42:29:17effort. After the data is gathered comes
- 42:29:20data preparation. This step involves
- 42:29:22data cleaning and data transformation.
- 42:29:25Data cleaning is the most time consuming
- 42:29:27process as it involves handling many
- 42:29:29complex scenarios. Here Emma deals with
- 42:29:32inconsistent data types, misspelled
- 42:29:34attributes, missing values, duplicate
- 42:29:37values and whatnot. Then in data
- 42:29:39transformation she modifies the data
- 42:29:42based on defined mapping rules. In a
- 42:29:44project ETL tools like talent and
- 42:29:46Informatica are used to perform complex
- 42:29:49transformations that helps the team to
- 42:29:51understand the data structure better.
- 42:29:53Then understanding what you actually can
- 42:29:55do with your data is very crucial. For
- 42:29:57that Emma does exploratory data analysis
- 42:30:00with the help of EDA. She defines and
- 42:30:03refineses the selection of feature
- 42:30:04variables that will be used in the model
- 42:30:07development. But what if Emma skips this
- 42:30:09step? She might end up choosing the
- 42:30:11wrong variables which will produce an
- 42:30:13inaccurate model. Thus, exploratory data
- 42:30:15analysis becomes the most important
- 42:30:18step. Now, she proceeds to the core
- 42:30:20activity of a data science project which
- 42:30:22is data modeling. She repetitively
- 42:30:24applies diverse machine learning
- 42:30:26techniques like KN&N, decision tree,
- 42:30:29knives based to the data to identify the
- 42:30:32model that best fits the business
- 42:30:34requirements. She trains the models on
- 42:30:36the training data set and tests them to
- 42:30:38select the best performing model. Emma
- 42:30:41prefers Python for modeling the data.
- 42:30:43However, it can also be done using R and
- 42:30:46SAS. Well, the trickiest part is not yet
- 42:30:49over. Visualization and communication.
- 42:30:51Emma meets the clients again to
- 42:30:53communicate the business findings in a
- 42:30:55simple and effective manner to convince
- 42:30:58the stakeholders. She uses tools like
- 42:31:00Tableau, PowerBI and ClickView that can
- 42:31:03help her in creating powerful reports
- 42:31:05and dashboards. And then finally, she
- 42:31:07deploys and maintains the model. She
- 42:31:10tests the selected model in a
- 42:31:11pre-production environment before
- 42:31:13deploying it in the production
- 42:31:15environment which is the best practice.
- 42:31:17Right? After successfully deploying it,
- 42:31:19she uses reports and dashboards to get
- 42:31:22realtime analytics. Further, she also
- 42:31:25monitors and maintains the project's
- 42:31:27performance. Well, that's how Emma
- 42:31:29completes the data science project. We
- 42:31:31have seen the daily routine of a data
- 42:31:33scientist is a whole lot of fun, has a
- 42:31:35lot of interesting aspects and comes
- 42:31:37with its own share of challenges. Now,
- 42:31:40let's see how data science is changing
- 42:31:42the world. Data science techniques along
- 42:31:44with genomic data provides a deeper
- 42:31:47understanding of genetic issues and
- 42:31:49reaction to particular drugs and
- 42:31:50diseases. Logistic companies like DHL,
- 42:31:54FedEx have discovered the best routes to
- 42:31:56ship, the best suited time to deliver,
- 42:31:58the best mode of transport to choose,
- 42:32:00thus leading to cost efficiency. With
- 42:32:03data science, it is possible to not only
- 42:32:06predict employee attrition, but to also
- 42:32:08understand the key variables that
- 42:32:10influence employee turnover. Also, the
- 42:32:13airline companies can now easily predict
- 42:32:15flight delay and notify the passengers
- 42:32:18beforehand to enhance their travel
- 42:32:20experience. Well, if you're wondering,
- 42:32:22there are various roles offered to a
- 42:32:24data scientist like data analyst,
- 42:32:27machine learning engineer, deep learning
- 42:32:29engineer, data engineer, and of course,
- 42:32:32data scientist. The median base salaries
- 42:32:35of a data scientist can range from
- 42:32:37$95,000 to $165,000.
- 42:32:41So that was about the data science. Are
- 42:32:43you ready to be a data scientist? If
- 42:32:45yes, then start today. The world of
- 42:32:48data. Picture this. You're shopping
- 42:32:50online and suddenly you see a product
- 42:32:52that feels like it was made just for
- 42:32:55you. How did they know? It's not by
- 42:32:58chance. It's data science. Data science
- 42:33:00help businesses understand what you
- 42:33:02like, predict what you'll need next, and
- 42:33:05improve the way we shop and use
- 42:33:07technology. And here's the best part.
- 42:33:10Data science isn't just about watching
- 42:33:12Netflix. It's one of the fastest growing
- 42:33:15careers in the world right now. In fact,
- 42:33:17the US Bureau of Labor Statistic says
- 42:33:20that data science jobs are expected to
- 42:33:23grow 36% by 2033, way faster than most
- 42:33:28of the other jobs. Companies everywhere
- 42:33:30are using data to make smarter
- 42:33:32decisions. That means the demand for
- 42:33:35data scientists is huge. And let's talk
- 42:33:38about the salary. You're probably
- 42:33:39wondering how much can I earn actually?
- 42:33:42Well, for entry- level position, data
- 42:33:44scientists in the US are earning around
- 42:33:47$152,000
- 42:33:49per year right now. And by 2025, some
- 42:33:52can make as much as $230,000.
- 42:33:55And in India, starting salaries range
- 42:33:57from 50,000 rupees to 1 lakh per month.
- 42:34:01And experienced professionals can earn
- 42:34:03more than 5 lakh rupees per month.
- 42:34:05That's impressive, right? But the best
- 42:34:07part is as a data scientist, you won't
- 42:34:09just stop here. The skills you develop
- 42:34:12in this role like machine learning, data
- 42:34:14visualization, and statistics are highly
- 42:34:17transferable and crucial for moving into
- 42:34:20AI roles. So whether it's becoming an AI
- 42:34:22engineer or an AI specialist, the
- 42:34:24foundation you build in data science
- 42:34:27will help you level up and pursue
- 42:34:29exciting hyping opportunities in the AI
- 42:34:32field. Now if you're thinking this
- 42:34:34sounds great but where do I even start?
- 42:34:37Well that's exactly what the
- 42:34:39professional certificate course in data
- 42:34:41science from IIT Kpur and Simply Learn
- 42:34:44is designed to do. It will get you
- 42:34:46started and make sure you're ready for
- 42:34:48this booming industry. In this 11 month
- 42:34:50live online interactive program, you
- 42:34:53will learn the skills you need to become
- 42:34:54a data science professional. No more
- 42:34:57theory, no more fluff. You get hands-on
- 42:34:59projects, life classes and mentorship
- 42:35:01from IT Kpur faculty plus real world
- 42:35:04industry expert who will help you build
- 42:35:06your skills. So this course comes with
- 42:35:09exciting amazing features that makes
- 42:35:11learning even more impactful. Eight
- 42:35:13times high interaction in live online
- 42:35:16classes with industry expert. Regular
- 42:35:18live online classes conducted by
- 42:35:20experienced professionals who bring real
- 42:35:22world knowledge into every session.
- 42:35:25You'll also get to master 13 plus key
- 42:35:27skills including generative AI, prompt
- 42:35:29engineering, charge, expendable AI,
- 42:35:32conversional AI, NLP, and many more.
- 42:35:34These are the skills that top companies
- 42:35:36use every day. And you'll gain hands-on
- 42:35:38experience with 14 plus industry tools
- 42:35:41like Python, SQL, Tableau, Dali 2,
- 42:35:43Midjourney, TensorFlow, and more. And by
- 42:35:46the end of this course, you will be
- 42:35:48ready to tackle real world challenges
- 42:35:50using these powerful tools and
- 42:35:52techniques. But we are not just talking
- 42:35:54about textbook and theory. You'll work
- 42:35:57on 25 plus real world projects giving
- 42:36:00you hands-on experience with the tools
- 42:36:02and skills you'll use in the industry.
- 42:36:05For example, our first project would be
- 42:36:07about sales analysis. You'll use Python
- 42:36:10to analyze a clothing company fourth
- 42:36:12quarter sales data across Australian
- 42:36:15state, helping the company make informed
- 42:36:17decisions. The next project would be
- 42:36:19about employee performance analysis
- 42:36:21where you learn how to build machine
- 42:36:23learning models to understand the
- 42:36:24factors influencing employees turnover.
- 42:36:27Our third project would be about
- 42:36:29e-commerce which will help Amazon
- 42:36:31improve its recommendation engine to
- 42:36:33offer better recommendation to
- 42:36:35customers. Upon successful completion,
- 42:36:37you'll receive a program certificate
- 42:36:39directly issued by the ENI city academy
- 42:36:41IT Kpool within 45 days of completing
- 42:36:44your cohort. This prestigious
- 42:36:46certificate will help you boost your
- 42:36:48resume and show potential employees that
- 42:36:50you have mastered the skills needed to
- 42:36:53succeed. You'll also benefit from master
- 42:36:55classes delivered by distinguished IT
- 42:36:56Kur faculty who bring their deep
- 42:36:58expertise in this course. Plus, you will
- 42:37:01be exposed to trending tools like
- 42:37:02chargeb2 geni and prompt engineering.
- 42:37:05Now, if you're wondering who will be
- 42:37:06teaching all of this, then the IT
- 42:37:08carpool faculty is here for you. These
- 42:37:10experts have been in this field for
- 42:37:12years and have worked with the top
- 42:37:14companies. You'll also get master
- 42:37:16classes from them and they will guide
- 42:37:17you through the learning process. You're
- 42:37:19not just learning from a textbook. You
- 42:37:21are learning from the people who have
- 42:37:23been there and done that. And the best
- 42:37:25part is once you have learned the
- 42:37:26skills, simply learns career assistance
- 42:37:28team will help you take to the next step
- 42:37:31which will help you to build a killer
- 42:37:32resume, show you how to stand out top
- 42:37:35recruiters and even give you access to
- 42:37:37mock interviews. Plus, you will also
- 42:37:39gain access to exclusive networking
- 42:37:41events and hackathons to connect with
- 42:37:43industry professionals. Along with
- 42:37:45Simply Learn's job assistant, you'll
- 42:37:47also get access to ID Kpur's career
- 42:37:49services helping you connect with top
- 42:37:51recruiters and land interviews with
- 42:37:53leading tech companies. So, upon
- 42:37:55finishing this course, you'll be ready
- 42:37:56for top data science roles like data
- 42:37:58scientists, machine learning engineer,
- 42:38:00AI specialist, and business analyst. The
- 42:38:02good news is top companies like Amazon,
- 42:38:05EY, Fidelity Investment, Johnson and
- 42:38:07Johnson, Borafhone, Accenture, Infosys
- 42:38:10and Nvidia is looking for professionals
- 42:38:13just like you. And with salaries in the
- 42:38:15US hitting around $230,000 plus and in
- 42:38:18India reaching up to five lakh per
- 42:38:20month, your career outlook looks great.
- 42:38:23So what will you be actually learning in
- 42:38:25this course? So here's a sneak peek of
- 42:38:27the syllabus which you'll be learning in
- 42:38:28this course which is the foundation in
- 42:38:30Python, SQL and mathematics, core data
- 42:38:33science like machine learning, data
- 42:38:34visualization, NLP, special topics like
- 42:38:37GNI, charge GBT, prompt engineering and
- 42:38:40you'll also work on industry projects
- 42:38:41which we have already mentioned before.
- 42:38:43You'll also have the option to choose
- 42:38:45electives like data storytelling with
- 42:38:47PowerBI and business analytics with
- 42:38:49Excel to tailor your learning experience
- 42:38:51and focus on the areas that interest you
- 42:38:53the most. So don't wait, hurry up and
- 42:38:56enroll now and find the course link in
- 42:38:57the description box below and in the pin
- 42:38:59comments. Have you ever wondered how
- 42:39:01your favorite online store seems to know
- 42:39:03exactly what you are looking for? Every
- 42:39:06time you browse, add to cart or wish
- 42:39:08list an item, you are leaving clues
- 42:39:10about your style, favorite colors,
- 42:39:12brands, and even shopping times. Data
- 42:39:15scientists jump in, analyze these
- 42:39:16patterns, and create a super
- 42:39:18personalized shopping experience.
- 42:39:20Suddenly, the store is showing you just
- 42:39:22the right pieces at just the right time.
- 42:39:24Almost like it's reading your mind.
- 42:39:27That's data science. Turning your clicks
- 42:39:29into a shopping spree crafted just for
- 42:39:31you. Hello everyone. Welcome back to
- 42:39:33Simply Learn's YouTube channel. If
- 42:39:35you're already a data science enthusiast
- 42:39:37or just got curious about this exciting
- 42:39:39field, you're in the right place. Today
- 42:39:41in this video, I'm diving into 10
- 42:39:44essential steps to help you become the
- 42:39:46next in- demand data scientist and land
- 42:39:48that dream job. No more waiting. Let's
- 42:39:51dive right in and get you on the path to
- 42:39:53your future in data science. So, let's
- 42:39:55see the 10 essential steps to become the
- 42:39:58next data scientist in demand. Step
- 42:40:00number one is programming languages.
- 42:40:02Starting with Python is a beginner is a
- 42:40:05great move because it's simple,
- 42:40:06versatile, and widely used in data
- 42:40:08science. Python straightforward syntax
- 42:40:10makes it beginner friendly, helping you
- 42:40:12grasp programming basics quickly and
- 42:40:15dive into data science libraries like
- 42:40:17pandas, numpy and mattplotive with ease.
- 42:40:20Adding R to your skill set is valuable
- 42:40:22because it excels at statistical
- 42:40:24analysis and data visualization two
- 42:40:26essential parts of data science. You can
- 42:40:29be comfortable with Python and R within
- 42:40:31a month or two. So moving on to the next
- 42:40:33step that is version control system.
- 42:40:35Learning a version control system like
- 42:40:38Git is essential because it allows you
- 42:40:40to track, manage, and collaborate and
- 42:40:42code effectively. With Git, you can save
- 42:40:44different versions of your work, making
- 42:40:46it easy to backtrack if something goes
- 42:40:48wrong or to experiment without losing
- 42:40:50progress. This is especially useful when
- 42:40:53working with complex data science
- 42:40:55projects where you might try out
- 42:40:56different models of analysis techniques.
- 42:40:59One or two weeks for practice along with
- 42:41:01Python and R is good to get start. Now
- 42:41:03moving on to the third step that is data
- 42:41:06structures and algorithms. Learning data
- 42:41:08structures and algorithms is crucial for
- 42:41:10becoming a data scientist because they
- 42:41:12provide the foundation for efficient
- 42:41:13data handling and problem solving. Data
- 42:41:16structures like arrays, stacks, cues and
- 42:41:18trees help you store and organize data
- 42:41:20in ways that make it easier and faster
- 42:41:22to access, process and analyze.
- 42:41:24Algorithms on the other hand give you
- 42:41:26strategies to perform tasks like
- 42:41:28searching, sorting and optimizing data
- 42:41:30operations which are essential for
- 42:41:32handling large data sets. While many
- 42:41:34candidates struggle with the essay,
- 42:41:36mastering it gives you an age helping
- 42:41:38you stand out in the interviews and
- 42:41:40shine as a skilled data scientist
- 42:41:42capable of tackling the toughest data
- 42:41:44problems. Spend about two months in
- 42:41:46this, you will get in the shape for
- 42:41:48sure. Now moving on to the step number
- 42:41:50four that is SQL. Learning SQL is
- 42:41:53essential for data scientists because it
- 42:41:55enables you to access, manage, and
- 42:41:57manipulate data directly within
- 42:41:58databases where most real world data
- 42:42:02resides. With SQL, you can create new
- 42:42:04tables, alter existing ones, delete
- 42:42:07unnecessary records, and run queries to
- 42:42:10filter, sort, and aggregate data. These
- 42:42:13abilities allow you to retrieve, clean,
- 42:42:15and organize data effectively. Core
- 42:42:18skills needed for any data science role.
- 42:42:20It's easy and you don't have to spend
- 42:42:22more than a month to have a deep
- 42:42:24understanding of it. Now moving on to
- 42:42:26the fifth step that is mathematics and
- 42:42:28statistics. Mathematics and statistics
- 42:42:30are essential for data science because
- 42:42:32they form the backbone of data analysis,
- 42:42:34model building and interpretation.
- 42:42:36Topics like linear algebra, calculus,
- 42:42:39probability and statistics gives data
- 42:42:42scientists the tools to understand data
- 42:42:44patterns, perform accurate analysis and
- 42:42:46make datadriven decisions. Mastering
- 42:42:48these areas enables you to build robust
- 42:42:50models, validate results and tackle
- 42:42:53complex problems confidently making you
- 42:42:56a well-rounded and skilled data
- 42:42:58scientist. Make sure you spend two
- 42:43:00months to grasp this topics. Now moving
- 42:43:02on to the step number six that is data
- 42:43:04prep-processing and visualization.
- 42:43:06Learning data prep-processing and
- 42:43:08visualization is essential for a data
- 42:43:10scientist because these skills make you
- 42:43:13data accurate, insightful and easy to
- 42:43:16understand. Python libraries like NumPy
- 42:43:18and Panders are crucial for manipulating
- 42:43:21and creating data, enabling you to
- 42:43:23handle missing values, filter out noise,
- 42:43:26and prepare data for analysis. Once the
- 42:43:28data is ready, visualization lets you
- 42:43:31uncover patterns and communicate results
- 42:43:33effectively. Libraries like Mattplot tip
- 42:43:36and Seaborn help create clear, impactful
- 42:43:39visuals, allowing you to interpret
- 42:43:41trends and convey insights in a way
- 42:43:43that's easily understood by others.
- 42:43:45Together with these tools make data
- 42:43:47prep-processing and visualization
- 42:43:49fundamentals for effective data science.
- 42:43:51If you have a solid foundation on Python
- 42:43:53and mathematics, you will get a good
- 42:43:55understanding of data prep-processing
- 42:43:56and visualization in a month or two. Now
- 42:43:59moving on to the seventh step that is
- 42:44:01machine learning fundamentals. Machine
- 42:44:03learning fundamentals involve
- 42:44:04understanding how algorithms enable
- 42:44:06computers to learn from data and make
- 42:44:08predictions on decisions without
- 42:44:10explicit programming. The two main
- 42:44:12categories are supervised learning and
- 42:44:14unsupervised learning. In supervised
- 42:44:16learning, models are trained on labelled
- 42:44:18data to make predictions while in
- 42:44:19unsupervised learning models find
- 42:44:22patterns in unlabelled data. Popular
- 42:44:25tools like TensorFlow, PyTorch help
- 42:44:28build and train complex models
- 42:44:30especially for deep learning. While
- 42:44:32skyit learn is essential used for
- 42:44:35simpler machine learning algorithms and
- 42:44:37data prep-processing. These tools make
- 42:44:39it easier to implement machine learning
- 42:44:41fundamentals effectively and build
- 42:44:44intelligent datadriven decisions.
- 42:44:46Dedicate about three months to
- 42:44:48understand the core of machine learning.
- 42:44:50Now coming to the next step that is deep
- 42:44:52learning. Deep learning is a subset of
- 42:44:54machine learning that focuses on
- 42:44:56algorithms inspired by the structures of
- 42:44:58the human brain called neural networks.
- 42:45:00Deep learning uses neural networks with
- 42:45:02multiple layers often dozens or hundreds
- 42:45:05to learn complex patterns from large
- 42:45:07data sets. Specialized types like
- 42:45:10convolutional neural networks that is
- 42:45:12CNN's are great for image processing
- 42:45:15while recurrent neural networks RNNs are
- 42:45:18used for sequence data like text or time
- 42:45:20series. Essential tools like TensorFlow,
- 42:45:23PyTorch make building, training and
- 42:45:25deploying deep learning models more
- 42:45:27accessible, allowing you to create
- 42:45:29powerful AI solutions across various
- 42:45:32domains. I think it will take about 2
- 42:45:34months to have a good hold on deep
- 42:45:36learning concepts and how to implement
- 42:45:38them. Now moving on to the ninth step
- 42:45:40that is specializations. Once you have
- 42:45:42grasped the deep learning, it's like
- 42:45:44reaching a new level as a data
- 42:45:46scientist. Just as doctors specialize in
- 42:45:48areas in nephrology and cardiology, data
- 42:45:52scientists often choose to specialize in
- 42:45:54fields like natural language processing
- 42:45:56or computer vision. Natural language
- 42:45:59processing focuses on teaching machines
- 42:46:02to understand and generate human
- 42:46:03language enabling applications like
- 42:46:06chatbot, sentiment analysis, and
- 42:46:08language transition. It's about making
- 42:46:11computers read, write, and even
- 42:46:12interpret human emotions through text or
- 42:46:15speech. Computer vision on the other
- 42:46:17hand is all about enabling machines to
- 42:46:19see and interpret images or videos. This
- 42:46:22field powers innovations like facial
- 42:46:24recognition, object detection and
- 42:46:26autonomous driving. Now you don't need
- 42:46:28to learn both. You can choose what
- 42:46:30interests you the most. Now spend one to
- 42:46:33two months diving deep into one of these
- 42:46:35areas. Now moving on to the last but not
- 42:46:38the least step that is big data. Big
- 42:46:40data refers to extremely large volumes
- 42:46:42of data generated rapidly from sources
- 42:46:44like social media and sensors. For data
- 42:46:47scientists, learning to handle big data
- 42:46:49is crucial as it requires specialized
- 42:46:52tools like Hadoop and Spark to analyze
- 42:46:54and extract insights effectively. With
- 42:46:57companies relying on datadriven
- 42:46:58decisions, big data skills make you a
- 42:47:00highly in- demand professional in the
- 42:47:02field. Focus for about 2 months and you
- 42:47:05will be able to spot trends and patterns
- 42:47:06from data sets very easily. Once you're
- 42:47:09ready, it's time to build a killer
- 42:47:10resume packed with projects that
- 42:47:12showcase your new skills. Start applying
- 42:47:14to jobs on platforms like Noy and Indate
- 42:47:17and supercharge your LinkedIn. Connect
- 42:47:19with data scientists. See what skills
- 42:47:21they are mastering and learn from their
- 42:47:23journeys as well. Keep sharpening your
- 42:47:25own skills and when the time comes, you
- 42:47:27will be ready to crush those interviews
- 42:47:29and land your dream data scientist role
- 42:47:32in 2025.
- 42:47:33>> Yeah. step by step we will go through
- 42:47:35all of this and uh we'll make sure that
- 42:47:38we learn everything and we bring
- 42:47:39everything together towards the end
- 42:47:41right without further ado let me just
- 42:47:44straight away deep dive to business
- 42:47:47right to learn data science
- 42:47:54right
- 42:47:56and with this data science there is also
- 42:47:58something which is prefixed which is
- 42:48:01applied data science
- 42:48:06and suffix for this is with Python
- 42:48:13right apply data science with Python
- 42:48:15right so there are there are two key
- 42:48:17concepts which are going to be a part of
- 42:48:19this course the first one is the
- 42:48:21knowledge about data science that what
- 42:48:23data science is and then because we are
- 42:48:26doing an applied course right we are
- 42:48:27doing an applied course I will try to
- 42:48:29tie up these concepts which we will
- 42:48:32understand in data science with a tool
- 42:48:35right which is Python for you right we
- 42:48:37already know about uh 60 65% of Python
- 42:48:42right which is the fundamental Python
- 42:48:44and now we will be moving to the next
- 42:48:46step to advanced Python
- 42:48:52right and using Python right leveraging
- 42:48:55Python we will be solving a lot of
- 42:48:59problems of data science right using
- 42:49:01this.
- 42:49:03Okay. So, the first few sessions, right?
- 42:49:06The first few sessions will be about
- 42:49:09making you a breast with Python. What
- 42:49:11Python is, right? What how and what
- 42:49:14packages do we have? How do they work in
- 42:49:16reality, right? And all those things.
- 42:49:18And then we will be coupling it up with
- 42:49:20data science concepts. And then finally
- 42:49:22towards the end of the session in the
- 42:49:24last few classes, we will be doing uh we
- 42:49:27will be taking a real data set. And on
- 42:49:29that data set we will be applying all
- 42:49:31these concepts right to understand the
- 42:49:33data better and we will be drawing
- 42:49:35inferences from that to convert that
- 42:49:38into information to take actionable
- 42:49:40insights or using that actionable
- 42:49:42insights taking a better decision.
- 42:49:44Right? We'll do all of that in in the
- 42:49:47actual way. Okay. So now guys if you
- 42:49:51understand this then the next point of
- 42:49:54contention is data sets
- 42:49:58right one of the most famous keywords on
- 42:50:02the planet right now right one of the
- 42:50:04most famous keywords in the planet right
- 42:50:05now do you think that these two things
- 42:50:10okay let me put it different way what do
- 42:50:12you think that can be the possible
- 42:50:15explanation about this term data science
- 42:50:18you know data You know science what do
- 42:50:21you think is going to follow in these
- 42:50:23sessions? What is data science to you as
- 42:50:26per these two words? Okay. So this is
- 42:50:29people made up of two words right data
- 42:50:31and science right. So what we are trying
- 42:50:35to do is we are
- 42:50:40trying to understand data
- 42:50:46right? We are trying to understand data
- 42:50:48right and then do something to it right
- 42:50:52understanding its science understanding
- 42:50:54the uh nature the behavior of this data
- 42:50:58and converting it in something called as
- 42:51:01information
- 42:51:04right do we know difference between data
- 42:51:07and information data is something which
- 42:51:10is completely raw okay it is completely
- 42:51:14raw it has no meaning
- 42:51:19right it has no meaning isn't it for
- 42:51:21example
- 42:51:23I give you these stats of some player
- 42:51:27like suppose Sid Dhoni I give you stats
- 42:51:29of Mahindra Singh Dhoni right that what
- 42:51:31what what were his scores uh what is his
- 42:51:34name what is his age and you know all
- 42:51:37those things now everything is there
- 42:51:40right but we don't know what to do about
- 42:51:42it right do you think the score of Dhoni
- 42:51:44has any context people it has any
- 42:51:47context no right but when I deep down
- 42:51:50but I when I go and deep dive about it
- 42:51:52right what is the first thing you find
- 42:51:54out of scores what is the first thing
- 42:51:57you find out of score scores you try to
- 42:52:00find the average of score isn't it that
- 42:52:03in last 10 innings
- 42:52:05right in last 10 innings before this
- 42:52:07also you need something which is called
- 42:52:09as a problem statement isn't it now for
- 42:52:12example the problem statement is select
- 42:52:15Selectors want to understand selectors
- 42:52:17wants to understand that whether
- 42:52:19Mahindra Singh Dhoni should be picked
- 42:52:21up. So there's a problem now right
- 42:52:24selectors want to see that whether Dhoni
- 42:52:27is fit for the next tournament or not.
- 42:52:31So what we will do we will now try to
- 42:52:34take the mean of the scores right for
- 42:52:40last 10 innings. And if this score is
- 42:52:43suppose X, we will try to compare this
- 42:52:45with Y. What is Y? Y is a reference,
- 42:52:49right? Y is a reference that we want to
- 42:52:52compare it against. Now when you are
- 42:52:54doing this comparisons, when you are
- 42:52:56applying these techniques to this score,
- 42:52:59this is now slowly becoming information,
- 42:53:03right? And at the end of the day once
- 42:53:06you have the strike rate once you have
- 42:53:09the mean score of Dhoni once you have
- 42:53:12his age once you have his fitness score
- 42:53:15all those things will now help you to
- 42:53:18take this particular decision because
- 42:53:21now what you have is called as
- 42:53:23information right because this has
- 42:53:25context
- 42:53:28right this has meaning
- 42:53:31and this is usually processed
- 42:53:35Right? This is usually processed. Right?
- 42:53:37This is usually processed. Now what did
- 42:53:39we do? Now what did we do here? If you
- 42:53:41will go and read about data science,
- 42:53:43data science says,
- 42:53:47data science says
- 42:53:49it is the art of collecting,
- 42:53:54right? cleaning,
- 42:53:58analyzing,
- 42:54:03modeling,
- 42:54:08improving,
- 42:54:12right? And visualizing,
- 42:54:19right? Visualizing
- 42:54:21the day, right? If a person is adept in
- 42:54:25doing all these things, this person
- 42:54:28people is cumulatively called a data
- 42:54:30scientist. Right?
- 42:54:33That person is called a data scientist.
- 42:54:35Right? So before going to the definition
- 42:54:37of data scientist, now I will give you
- 42:54:39some more examples, right? I'll give you
- 42:54:40some more examples. Data science people
- 42:54:43as I said is a combination of these
- 42:54:45things, right? You have to collect the
- 42:54:47data,
- 42:54:49right? You have to collect the data,
- 42:54:52right? Right. And this has a lot of
- 42:54:54things. Data can be connected from two
- 42:54:56types in two types. One is primary
- 42:55:01and the second one is secondary.
- 42:55:05Right? What is the primary way of
- 42:55:07collecting data? From your IoT devices,
- 42:55:09right? From sensors,
- 42:55:12from your inbuilt machines,
- 42:55:16right? Then from your surveys
- 42:55:20which you float, right? Questionnaires,
- 42:55:25right? All these things are primary
- 42:55:26ways. What is the secondary way of
- 42:55:27collecting data?
- 42:55:29Purchasing data,
- 42:55:33right? Using internet data
- 42:55:38because you have not generated it. You
- 42:55:40are just using someone else's data.
- 42:55:42Right? Something like uh transfer
- 42:55:44learning.
- 42:55:48What is transfer learning?
- 42:55:50Transfer learning is a technique where
- 42:55:52suppose I am bank A and you are bank B
- 42:55:56right so bank A has created some model
- 42:56:00right trained on their data now you are
- 42:56:04going to use the exact same model right
- 42:56:07you're going to use the exact same model
- 42:56:09right maybe you're not seeing the data
- 42:56:11but you are just using the property of
- 42:56:13data like mean median mode and a lot of
- 42:56:16modeling things which will come we will
- 42:56:18learn about them and You use this model
- 42:56:20on your particular data right so in a
- 42:56:22way you did not have enough data to
- 42:56:25create the model yourself but you are
- 42:56:26now using someone else's model to run
- 42:56:29your data on it right so this is called
- 42:56:31as transfer learning so this kind of
- 42:56:33collection is basically secondary data
- 42:56:36collection so you can collect the data
- 42:56:39right then you can perform data analysis
- 42:56:46right you can perform data analysis
- 42:56:48right How will you perform this data
- 42:56:50analysis? Using complex
- 42:56:53algorithms,
- 42:56:57right? Using complex algorithms, right?
- 42:56:59Some statistics,
- 42:57:03right? You can use artificial
- 42:57:06intelligence,
- 42:57:12right? Artificial intelligence. You can
- 42:57:14use machine learning.
- 42:57:19Right.
- 42:57:22Right. You can use all these things for
- 42:57:24data analysis. Then you can transform
- 42:57:31transform
- 42:57:33the patterns
- 42:57:39into predictions.
- 42:57:43Right? You can transfer these patterns
- 42:57:45into predictions, right? Which can be
- 42:57:48used for business
- 42:57:50decision making,
- 42:57:55right? For business decision making.
- 42:57:57Then you can validate the results,
- 42:58:03right? And present the results,
- 42:58:09right? So this is like a complete life
- 42:58:11cycle of a data scientist, right? So
- 42:58:14before going further, let me give you
- 42:58:15what combinations do you need to have to
- 42:58:18become a data scientist. The first one
- 42:58:20is
- 42:58:23domain knowledge,
- 42:58:29right? So what is domain knowledge?
- 42:58:32First of all, I told you right there
- 42:58:33will be a problem, right? You'll be
- 42:58:35solving a problem in any project of data
- 42:58:38science. You'll be trying to solve a
- 42:58:39problem right and the problem will be
- 42:58:41belonging to a particular domain even if
- 42:58:44you're working for yourself right even
- 42:58:45if you're an entrepreneur then also
- 42:58:47you'll be solving a problem. So this
- 42:58:49domain knowledge part includes things
- 42:58:51like understanding
- 42:58:55it's a very important diagram
- 42:58:58understanding
- 42:59:00client requirement
- 42:59:03right understanding the client
- 42:59:05requirement right
- 42:59:07important criterians
- 42:59:12right important
- 42:59:14criteria knowledge
- 42:59:18Right? For example, to give an example,
- 42:59:21suppose we have created a machine
- 42:59:23learning model. Okay? Understanding the
- 42:59:25data, we have created a machine learning
- 42:59:27model whose accuracy is 90%. Right? Is
- 42:59:3190% a good accuracy?
- 42:59:34Yeah, fairly decent accuracy. Yes.
- 42:59:37Suppose you have to predict sales,
- 42:59:39right? You are selling something.
- 42:59:40Suppose you are selling clothes and you
- 42:59:42want to predict what will be the sales
- 42:59:44for the next week. When you use this
- 42:59:46model, whatever the output model gives
- 42:59:48you, what is going to be the accuracy of
- 42:59:50your output using this model?
- 42:59:54How much accuracy?
- 42:59:5790%.
- 42:59:58But my question is model is 90%. But my
- 43:00:02question to you is that if 90% accuracy
- 43:00:05is on sales data, a person like me will
- 43:00:08be very very happy. Okay? Very very
- 43:00:10happy. I'll be probably dancing, right?
- 43:00:13But if you try to apply the same model
- 43:00:16right for a medical diagnosis case, will
- 43:00:20you be interested in getting operated in
- 43:00:23such a hospital or an institution where
- 43:00:26the accuracy is coming as 90%. Domain
- 43:00:29knowledge, right? Domain knowledge. We
- 43:00:31need to understand what are the exact
- 43:00:33requirements. We need to understand what
- 43:00:36are the exact expectations,
- 43:00:38right? And we need to know how much do
- 43:00:41we need to pivot right? So the first
- 43:00:43thing in data science is these
- 43:00:45accuracies and everything are subjective
- 43:00:47right they are subjective. So for that
- 43:00:50you need domain knowledge. Domain
- 43:00:52knowledge part very important guys very
- 43:00:55important these three things which I'm
- 43:00:56going to tell. Second part is people
- 43:01:00the game changer right the second part
- 43:01:02is
- 43:01:04computer science.
- 43:01:07Now if I take you back in history okay
- 43:01:10if I take you back in history in 1980s
- 43:01:13or somewhere then do you okay how many
- 43:01:16of you think that data science is a new
- 43:01:17concept
- 43:01:19how many of you think that data science
- 43:01:20is a new concept I hope you all it's not
- 43:01:23a new concept everyone knows that yeah
- 43:01:26it has been happening for ages just like
- 43:01:30you guys will be shocked if you already
- 43:01:32don't know AI was coined in the year
- 43:01:341956
- 43:01:361956 6 at the University of Dharma. AI
- 43:01:40was coined by Paul McCarthy, right? And
- 43:01:43we saw the boom of AI in the year 2010,
- 43:01:48right? Such a long journey. Same case
- 43:01:50with data science because people back in
- 43:01:53the day data science was called as data
- 43:01:56mining. Everyone heard about it data
- 43:01:59mining, knowledge databases. Yeah, we
- 43:02:02need we used to mine the data. Now, what
- 43:02:05were the problems? What were the hiccups
- 43:02:07of data mining? The hiccups for data
- 43:02:09mining was that we were doing everything
- 43:02:14everything manually
- 43:02:18right now if I give you 100 points can
- 43:02:21you calculate the mean
- 43:02:23or let's say if I give you two points to
- 43:02:26multiply 2 * 3 how much time will you
- 43:02:30take?
- 43:02:322 seconds.
- 43:02:34Yep. How much time a computer will take?
- 43:02:372 seconds. If I give you to multiply 2
- 43:02:43489
- 43:02:45multiplied by 200, how much time will
- 43:02:47you take to calculate this? Say 5
- 43:02:50seconds. How much time computer will
- 43:02:52take? 2 seconds. Now if I give you to
- 43:02:56multiply 2 48 9 into 15 1 95 4386
- 43:03:02how much time will you take to calculate
- 43:03:03this manually? Maybe say 1 minute
- 43:03:0780 seconds 1 minute. How much time a
- 43:03:09computer will take? Still 2 seconds
- 43:03:12right? still 2 seconds right and if I
- 43:03:15give you to calculate this over 200
- 43:03:17times you will take 200 minutes right
- 43:03:21using parallel computing computer will
- 43:03:23still take about 3 to 5 seconds right so
- 43:03:27are you understanding the power of
- 43:03:28computer do you understand this concept
- 43:03:31in this relationship what was happening
- 43:03:33back then what computer science did
- 43:03:36people was it revolutionized the way
- 43:03:39data mining was happening and That thing
- 43:03:42now is called as that again that thing
- 43:03:46now is called as data science in which
- 43:03:48computer science is one of the most
- 43:03:50important contributors. So this is just
- 43:03:52one reason. Now manually right manually
- 43:03:55if I give you say 1 million rows of data
- 43:03:59right 1 million rows of data right so
- 43:04:02how many pages will you pages will you
- 43:04:04need to store this data suppose your
- 43:04:07notebook is like this
- 43:04:09right these boxes and here you are
- 43:04:11storing the data 1 million so maybe you
- 43:04:13can buy n number of notebooks but now do
- 43:04:17you think it's as easy in storing
- 43:04:19something in computer because back in
- 43:04:21the day people we had memory issues,
- 43:04:25isn't it? Memory constraints.
- 43:04:30So there is something called as Murray's
- 43:04:32law, right? Which says as the
- 43:04:35advancement in microprocessors will
- 43:04:38increase, the price of microprocessor
- 43:04:41will decrease, right? So this is what is
- 43:04:42happening right now. Back in 1980s, if I
- 43:04:46show you right guys, right? Yeah. 2.5
- 43:04:50kg. Exactly. Right. It was size of a
- 43:04:51fridge hard disk but now it fits in your
- 43:04:54palm right. So this was enabled the
- 43:04:57storage techniques right the processing
- 43:04:59techniques the infrastructure right
- 43:05:02things like big data what kind of data
- 43:05:04do you think we will be dealing with
- 43:05:06people in data science you all know the
- 43:05:08term very famous term the kind of data
- 43:05:12big data right everyone knows about big
- 43:05:14data what is big data yes a data which
- 43:05:17is fast right it has velocity veracity
- 43:05:23variety right so This kind of data needs
- 43:05:26to be stored. This kind of data needs to
- 43:05:28be processed. So which thing brought all
- 43:05:31these things into data science? It was
- 43:05:33given to us by computer science, right?
- 43:05:36So computer science people included
- 43:05:38things like database management,
- 43:05:44right? Data validation, right? Data
- 43:05:48infrastructure,
- 43:05:51right? Data infrastructure,
- 43:05:53right? Then we had languages, computer
- 43:05:57languages
- 43:06:00which is Python right now for us. Right?
- 43:06:04Again do you think when you do this
- 43:06:07thing manually right suppose you do this
- 43:06:09thing manually how easy do you think it
- 43:06:11will become using something like Python
- 43:06:14or any other computer language to create
- 43:06:16complex models. How easy it will be to
- 43:06:19do that to create the complexity in
- 43:06:22models right where you can capture the
- 43:06:24nonlinear nature isn't it people isn't
- 43:06:28it
- 43:06:30for example for example let me tell you
- 43:06:33this
- 43:06:342 4 6 8 10 dash what do you think is the
- 43:06:40next number guys 12 if I tell you to
- 43:06:43define this to me in okay let leave
- 43:06:47leave
- 43:06:48What do you think is going to be the
- 43:06:49next number here?
- 43:06:57What is the next number? 11.
- 43:07:04Next number
- 43:07:0825.
- 43:07:09Perfect. Now guys, if I ask you to write
- 43:07:12these numbers, right, the way you
- 43:07:14predicted them, can you give me a
- 43:07:15function f ofx is equal to what is the f
- 43:07:19of x here?
- 43:07:21It's 2x, right? It's 2x. f ofx is equal
- 43:07:24to 2x. If I tell you to create a
- 43:07:27function, it will be f of x is equal to
- 43:07:292x. What will be the function here,
- 43:07:32guys?
- 43:07:34f ofx
- 43:07:36will be equal to
- 43:07:39x + 1. Yeah. x + 1. Yeah. And here f ofx
- 43:07:46will be equal to
- 43:07:48x². Yeah. Now the last example, right?
- 43:07:52Last example.
- 43:08:00What is the next number here? I don't
- 43:08:02want the number. I want the function. I
- 43:08:04want this so that I can generalize.
- 43:08:07Isn't it? How did you reach this figure?
- 43:08:11How many of you think it's not possible
- 43:08:13to determine this? How many
- 43:08:17of you
- 43:08:20think
- 43:08:23it is not
- 43:08:26possible
- 43:08:28to determine this?
- 43:08:33Yeah. How many of you think what if I
- 43:08:36just change this question and ask you
- 43:08:39how many of you think it is not possible
- 43:08:41to determine this
- 43:08:44manually?
- 43:08:47Same response. But now I say how many of
- 43:08:50you think it is not possible to
- 43:08:52determine this
- 43:08:54with computers?
- 43:08:56Will your answer still remain no? Do you
- 43:08:59think I cannot approximate this function
- 43:09:01using computers?
- 43:09:03We have something called as deep
- 43:09:07neural networks
- 43:09:10and they are called as universal
- 43:09:12function approximators.
- 43:09:14Right? So this is the problem people.
- 43:09:17This is the problem. Right? I will show
- 43:09:18this to you when the time comes. Right?
- 43:09:20I will remember this example and I will
- 43:09:22show this to you. But now what I'm
- 43:09:24trying to tell you is the things which
- 43:09:26seemed impossible manually was solved by
- 43:09:30what? It was solved by computers. The
- 43:09:33distribution of this is like this. Can
- 43:09:36you figure it out yourself? No. Right?
- 43:09:40We cannot. Isn't it? We cannot do that.
- 43:09:43So this kind of approximation will be
- 43:09:45given by what? It will be only given by
- 43:09:48machines. Right? And this is people what
- 43:09:50data science is all about. Right? It is
- 43:09:53what computer science did inside data
- 43:09:56science. Right? I hope this is clear.
- 43:09:59I'm assuming a lot of you will be going
- 43:10:01for interviews and everything after
- 43:10:03these course. Right? So this will be a
- 43:10:05very very important thing for you to
- 43:10:07know. Right? Often it is asked why data
- 43:10:09science is having computer science in
- 43:10:11it. Right? The reason is this. Okay? So
- 43:10:14this is the role of computer science
- 43:10:16inside data science. Now people the
- 43:10:18third thing the third circle which is
- 43:10:21one of the most
- 43:10:23parts is
- 43:10:26maths
- 43:10:28and stats right mathematics statistics
- 43:10:33which was optimization
- 43:10:37right optimization of your models right
- 43:10:42design
- 43:10:45of model. Right? Now guys, if you look
- 43:10:49carefully, if you look carefully, in
- 43:10:52order to approximate this, what you what
- 43:10:54will you be playing with? You will be
- 43:10:56playing with a lot of data. You'll be
- 43:11:00playing with a lot of mathematical to
- 43:11:02mathematical concepts and statistical
- 43:11:04concepts, isn't it? How did you what do
- 43:11:07you call this? This is math, right? This
- 43:11:09is statistics and mathematics. finding
- 43:11:12mean, median, mode, standard deviations,
- 43:11:15probability, statistics, all these will
- 43:11:18lead to this kind of result, isn't it?
- 43:11:20So, this becomes the third wheel of this
- 43:11:23particular uh of of this particular
- 43:11:25diagram. And this point of intersection,
- 43:11:30right? The sweet point of intersection
- 43:11:33is basically data science.
- 43:11:39Yeah. is particularly data science.
- 43:11:44Right? So now people this point okay
- 43:11:47this point
- 43:11:49is basically representing data
- 43:11:53engineering right data engineering right
- 43:11:56data engineering is people the part of
- 43:11:59data science which enables us to capture
- 43:12:02the correct data right how the data will
- 43:12:05flow how the data will be stored how the
- 43:12:08data will be cleaned right all this is
- 43:12:10done by home it is done by data
- 43:12:13engineering Just because so am I right
- 43:12:15there you'll go for interview right
- 43:12:16after this and try to fetch yourself
- 43:12:18jobs in this domain data science AI ML
- 43:12:23if my understanding is correct is that
- 43:12:25the aim yes so now guys there will be
- 43:12:27three types of companies
- 43:12:29or let's say to simplify let's say two
- 43:12:31types one is small and the other one is
- 43:12:35big right so in a small organization if
- 43:12:39you become a part of small or
- 43:12:40organization and you are the data
- 43:12:42scientist there You can be involved in
- 43:12:45all of these things, right? All of these
- 43:12:48things possible, right? Your bosses and
- 43:12:51your management will expect you to
- 43:12:53construct all these flows, right? Know
- 43:12:56computer science, you should know maths
- 43:12:58and stats and you should have the domain
- 43:12:59knowledge and you will be asked to do
- 43:13:01all of this. But if you are going to
- 43:13:04become a part of a big organization,
- 43:13:06usually all these roles are fragmented.
- 43:13:09All of these roles are fragmented,
- 43:13:12right? There's a separate data
- 43:13:13infrastructure team. There's a separate
- 43:13:15data governance team. Now guys, when you
- 43:13:17go on to collect the data, can you
- 43:13:19collect any sensitive data about is it
- 43:13:21possible ethically it's not right? And
- 43:13:24legally also it's not right. My question
- 43:13:26to you is who will look after this
- 43:13:28compliance? Whose responsibility indeed
- 43:13:31it is to look after this compliance?
- 43:13:33Data scientist. So this is about
- 43:13:36fragmentation. If you are part of a big
- 43:13:38organization, this thing will be taken
- 43:13:40by someone else, right? But if you are a
- 43:13:42part of a small organization, you will
- 43:13:44be know you'll be expected to do this
- 43:13:46all by yourself. In if if you are part
- 43:13:48of a big organization then do you think
- 43:13:50you need to have this domain knowledge?
- 43:13:53The answer is no. Why? Because there
- 43:13:56will be separate set of people who are
- 43:13:57called as what? Who are called as
- 43:14:00business analyst. Have you heard about
- 43:14:01this position people? Business analyst.
- 43:14:05What is a business analyst role? It is
- 43:14:07basically a technical translator, right?
- 43:14:10who knows technical, who knows domain
- 43:14:12and that person goes and talks to the
- 43:14:14client, talks to the client in a layman
- 43:14:16language, convert it into technical
- 43:14:18requirement coupled with the domain
- 43:14:20knowledge and give you the document.
- 43:14:22This is basically a medical engineering
- 43:14:24problem. So we need this this this this
- 43:14:27and you need to fulfill this this this
- 43:14:29this criteria. But again, if you're part
- 43:14:31of a small organization, who needs to
- 43:14:33take care of that? Who needs to make
- 43:14:35sure that you know everything about a
- 43:14:36domain? You yourself, right? You
- 43:14:40yourself, right? Then people, this area
- 43:14:45usually represents whom? This area
- 43:14:49represents research
- 43:14:52and analysis.
- 43:14:55Why?
- 43:14:56Because they we have people who have
- 43:14:58domain knowledge and we have people who
- 43:15:00are knowledge of math, stats. Have you
- 43:15:02heard about a position called as actury
- 43:15:05in the world? Acturial science. Acturies
- 43:15:08are people who are basically dealing
- 43:15:11with uh domains which are very very
- 43:15:14heavily data intrinsic. Right? For
- 43:15:16example, finance domain, right? Finance
- 43:15:18is all about numbers. So there we go
- 43:15:20actal science and we have math, stats,
- 43:15:23optimization, model development, all of
- 43:15:25those things happening there. We are
- 43:15:26also at at this point of time people
- 43:15:28belonging to which section research and
- 43:15:30data analysis right and then people
- 43:15:34there's the third intersection right
- 43:15:35there's a third the third intersection
- 43:15:39which is this part and this people is
- 43:15:42called as machine learning
- 43:15:48right machine learning why machine
- 43:15:50learning if you can combine the power of
- 43:15:53maths stats right and you Combine the
- 43:15:56power of computer science, you will find
- 43:15:59yourself to be in a position where you
- 43:16:01can call yourself a machine learning
- 43:16:03engineer. Right? How does a machine
- 43:16:05learning engineer becomes a data
- 43:16:06scientist? When they couple it up with
- 43:16:09the domain expertise, right? So in this
- 43:16:12course people in this course we will
- 43:16:14teach you computer science. We will
- 43:16:16teach you a little bit about math stats.
- 43:16:19But what we cannot teach you is domain
- 43:16:21knowledge. Yeah.
- 43:16:24and making the base of what we are about
- 43:16:26to do. Very very important to
- 43:16:28understand. Right? If you understand
- 43:16:29this then half the battle is won. Right?
- 43:16:33So now to answer the question which was
- 43:16:36posted earlier uh answer to the question
- 43:16:38which was posted earlier. There are
- 43:16:40different things in the world of data
- 43:16:41science. Right? So you can pick and
- 43:16:43choose anything or you can do everything
- 43:16:46by yourself. If you guys are engineers
- 43:16:48then I think you can be at the sweet
- 43:16:50spot going forward in life if you choose
- 43:16:52a domain for yourself. For example, you
- 43:16:54choose to be a part of automobile
- 43:16:56industry, you choose to be a part of
- 43:16:57medical industry, you choose to be a
- 43:16:59part of say retail industry, you choose
- 43:17:03to be a part of aeronautics industry,
- 43:17:06right? You choose to be a part of
- 43:17:08finance industry, right? So whatever you
- 43:17:10will choose, this thing will get
- 43:17:12developed over time, right? This is the
- 43:17:14most difficult out of these three. I try
- 43:17:16to give you the example right my domain
- 43:17:19was agricultural industry right the agri
- 43:17:22products I have worked extensively in
- 43:17:25agricultural industry right so again for
- 43:17:28now in my current role this is something
- 43:17:30which I don't have I have this expertise
- 43:17:32I have this expertise so same will be
- 43:17:34with you and you guys will develop this
- 43:17:36knowledge over the time now I will try
- 43:17:39to give you an example right elections
- 43:17:42to make you understand how data science
- 43:17:45can be used in one particular use case.
- 43:17:48Okay, maybe we can extend that to a lot
- 43:17:50of other use cases and examples, right?
- 43:17:53Talking about election season, right?
- 43:17:55Talking about the election season,
- 43:17:57right? We will we will try to understand
- 43:18:00how do we use data science because it is
- 43:18:03very extensively used in this data
- 43:18:05science. Right? Now, let me talk about
- 43:18:09the first phase. Let's say this is
- 43:18:12pre-election
- 43:18:13phase,
- 43:18:15right?
- 43:18:17Right. This is pre-election phase. In
- 43:18:20this pre-election phase, what do you
- 43:18:22think will be the tasks with which an
- 43:18:24agency like Election Commission of India
- 43:18:26will be doing? The first task can be
- 43:18:30that they will be
- 43:18:32doing the voter
- 43:18:34registration,
- 43:18:37right? Voter registration and data
- 43:18:40management, isn't it?
- 43:18:44Yeah, it will start with that.
- 43:18:47And what will be the things inside this?
- 43:18:49The first thing will be data collection,
- 43:18:54right? First thing will be data
- 43:18:55collection. So, you will collect the
- 43:18:57data from all the registered voters,
- 43:19:00right? Maybe it can be their demographic
- 43:19:02information, where they live, what is
- 43:19:04their age, what is their gender, right?
- 43:19:07What is their past polling behavior?
- 43:19:09Have they turned out previously or not?
- 43:19:11Right? All these details we can collect.
- 43:19:14Then can we also do data cleaning?
- 43:19:19Because I've told you, right? That there
- 43:19:21can be a lot of redundancies. Some
- 43:19:23person's name can appear twice, right?
- 43:19:26Some people can be a mismatch. Suppose
- 43:19:28for example, we have learned this in
- 43:19:30Python. Raghav.
- 43:19:36Raghav.
- 43:19:39Raghav.
- 43:19:42Radha. Right.
- 43:19:45Right. All these are what people?
- 43:19:48This is belonging to the same name.
- 43:19:50Right. This is me. But for a computer,
- 43:19:53for a computer, how many ragavves are
- 43:19:55there? Different ones. All are
- 43:19:56different. Right? So example like these,
- 43:19:59right? Some people might have died. They
- 43:20:01might not be existing anymore. Right? So
- 43:20:03all this part will be taken care where
- 43:20:05people in the data cleaning process.
- 43:20:08Right? Removing the duplicates, updating
- 43:20:10the new addresses, right? Correcting the
- 43:20:13information about every voter, all those
- 43:20:15things, right? And now lastly, we can
- 43:20:18also include a flavor of data analytics,
- 43:20:24right? What will data analytics include
- 43:20:25in this? We can analyze the demographic
- 43:20:28data to identify eligible voters. Isn't
- 43:20:32it? Yeah. People till now we haven't
- 43:20:35understood this why it is not automated.
- 43:20:38How will you pick up all the how how
- 43:20:40will you pick up all the details and
- 43:20:42nuances? Suppose you're filling a form
- 43:20:44right by mistake you have. Suppose you
- 43:20:46are 25 years of age. Suppose you have
- 43:20:48written 250. So does that mean I remove
- 43:20:51this? I remove this entry of yours
- 43:20:54because by mistake you have written your
- 43:20:56age as 250
- 43:20:58is age 250 possible in our current world
- 43:21:01never right so I will have to tell the
- 43:21:04machine that because this is a mistake
- 43:21:07please convert it to 25 isn't it this is
- 43:21:11called as imputation so this is mostly a
- 43:21:14manual task not manual but you have to
- 43:21:16understand the problem manually and then
- 43:21:20code it on
- 43:21:21Suppose someone has written their state
- 43:21:26as
- 43:21:28E D L H I right and country
- 43:21:35as I D I N right so what is this state
- 43:21:41referring to what is this country
- 43:21:42referring to the humans are very smart
- 43:21:46India right we can say this is India and
- 43:21:48if this is India then what is this
- 43:21:50pointing out to
- 43:21:52telly just because you said and it's a
- 43:21:55it's a leading question I want you to
- 43:21:56explain this tell me is it possible to
- 43:21:59do this automatically no right we have
- 43:22:02to employ manual rules right we have to
- 43:22:05tell because we who is more intelligent
- 43:22:08humans or machines humans right machines
- 43:22:11are just more optimized right so we know
- 43:22:14through human intelligence that this is
- 43:22:16pointing to Delhi and this is not EDLHI
- 43:22:19so this is data cleaning Right. Lastly,
- 43:22:21we have data analytics. So, do you think
- 43:22:23people based on these data points, we
- 43:22:26can understand that who are the eligible
- 43:22:30voters
- 43:22:32and maybe who are not registered yet,
- 43:22:34maybe who have not voted in the past.
- 43:22:37Can we do all those analytics
- 43:22:40and reach out to those peoples and
- 43:22:42persons? That's the first part. This
- 43:22:45just the first part pre-election phase.
- 43:22:47Now moving on to the second part right
- 43:22:51moving on to the second part let's say
- 43:22:54uh we say public
- 43:22:58opinion
- 43:23:02regarding the polling right the polling
- 43:23:04which is about to happen the first thing
- 43:23:06will be you want to collect information
- 43:23:08about people right so can can you go out
- 43:23:12and reach all 1.8 8 billion people in
- 43:23:16this country that what is their likable
- 43:23:20vote for which party is it possible
- 43:23:241.8 8 billion do you think it's possible
- 43:23:26for 1 billion
- 43:23:28do you think it's possible for 500
- 43:23:30million do you think it's possible for
- 43:23:32100 million no right so basically I'm
- 43:23:36talking about what I have something
- 43:23:38which is called as a population
- 43:23:42right and if I have to study about this
- 43:23:44population which is 1.8 8 billion people
- 43:23:48which is impossible which you just said
- 43:23:51what do I need to do should I stop my
- 43:23:52process no right I will go and collect
- 43:23:56something which is called a sample
- 43:24:01right we always work in samples right
- 43:24:05suppose someone says that a Coca-Cola
- 43:24:08bottle does not contain 500 ml of liquid
- 43:24:13which it claims right suppose someone
- 43:24:15has put this allegation possible that
- 43:24:17Coca-Cola bottles do not have 500 ml
- 43:24:21liquid which they promise. Now there are
- 43:24:23two ways to deal with this. Right? There
- 43:24:25are two ways to deal with this. Either I
- 43:24:27go and collect all the bottles of
- 43:24:30Coca-Cola in the world. Possible
- 43:24:33never right. So what will I do? I will
- 43:24:36go and pick up handful of bottles.
- 43:24:39Right? Handful of bottles. So what is
- 43:24:41that handful of bottles? Those are
- 43:24:43called as samples. One last thing.
- 43:24:45Suppose someone says that because of an
- 43:24:48industry
- 43:24:49all the fishes of the lake are dying or
- 43:24:53they are infected. Is it possible to go
- 43:24:55and collect and check all the fishes in
- 43:24:57the pond or a lake? No. Right. What will
- 43:25:00we do? We will collect again handful of
- 43:25:03fishes and we will test them. Right?
- 43:25:06Again samples. Now how does the raw data
- 43:25:10collected? Raw data as in I hope you
- 43:25:13understand this. This is no more about
- 43:25:16population
- 43:25:17with this step being told. Now you're
- 43:25:19dealing with samples. So now do you want
- 43:25:22to ask me how is sample created? Yeah,
- 43:25:24now I'm coming to that. Now guys, there
- 43:25:26are a lot of second point is how to
- 43:25:30sample right? How to sample
- 43:25:34right? So we have sampling techniques
- 43:25:36people. One is called as probabilistic
- 43:25:38and one is called as nonrobabilistic.
- 43:25:42Right? I will not go in detail right
- 43:25:44now. I just want to tell you an overview
- 43:25:46probabilistic is suppose uh you are
- 43:25:48manufacturing t-shirts right you are
- 43:25:50manufacturing t-shirts right and suppose
- 43:25:53you created
- 43:25:55100 lots
- 43:25:58of
- 43:26:00thousand t-shirts right so this is box
- 43:26:03one box two box three box four up to up
- 43:26:06to 100 right 100 boxes and in each box
- 43:26:09how many t-shirts are there 1 th00and
- 43:26:12right now suppose you are a Quality
- 43:26:14inspector. You're a quality inspector.
- 43:26:17Is it possible you for you to go over
- 43:26:20all the 100 lots with all checking all
- 43:26:23the thousand t-shirts one by one? No.
- 43:26:25Right. What will you do? You will sample
- 43:26:27again. You will sample. Now the most
- 43:26:31common way of sampling these kind of
- 43:26:34problems is probabilistic sampling. What
- 43:26:37is probability? What is the probability
- 43:26:38of getting heads or a tails when you
- 43:26:40spin the when you flip the coin? equally
- 43:26:42likely 1x2 and 1x2. What is the
- 43:26:45probability of getting 1 2 3 4 5 6 on a
- 43:26:48roll of a dice? 1x 6. Now what is the
- 43:26:51probability of picking any t-shirt from
- 43:26:54this first slot out of thousand
- 43:26:56t-shirts?
- 43:26:571 by,000.
- 43:27:00Yes. So do you think all the t-shirts
- 43:27:02have equally probable equal probability
- 43:27:04of being picked up without any bias? If
- 43:27:08you decide to draw five t-shirts, right,
- 43:27:11from each of this box, right, and
- 43:27:15suppose say two are defective and three
- 43:27:19are not defective, what will you do?
- 43:27:21Will you accept the lot or reject the
- 43:27:23lot? We have majority of t-shirts of
- 43:27:25non-deective
- 43:27:27or let's say we will reject the lot. We
- 43:27:29will reject the lot. We will reject the
- 43:27:30lot. Let's say we will reject the lot.
- 43:27:32Okay. Though this is basically
- 43:27:34subjective as per the company policies
- 43:27:37but let's say we rejected. Now people my
- 43:27:41question to you is what if this entire
- 43:27:44batch had only two defected t-shirts but
- 43:27:47now what will happen? The entire batch
- 43:27:49will be rejected.
- 43:27:52Yes. Let's say you sampled one t-shirt.
- 43:27:55Let's let's change the use case. I say
- 43:27:58you sampled only one t-shirt and that
- 43:28:00t-shirt was defected. Now you will
- 43:28:02reject the batch and that is equally
- 43:28:04likely case. So this is called as
- 43:28:07probabilistic sampling people and there
- 43:28:09is no way you can go back. There is no
- 43:28:12way you cannot say that hey sir please
- 43:28:15uh allow this batch to pass because
- 43:28:17there is a chance that rest of the
- 43:28:19t-shirts are not defective. No it is not
- 43:28:21the way that happens. It happens
- 43:28:24randomly. So right this is called as
- 43:28:26random sampling.
- 43:28:30Right? random sampling. Now suppose you
- 43:28:34are doing a cancer research, right?
- 43:28:37You're doing a cancer research, right?
- 43:28:40So for your cancer research people, what
- 43:28:42kind of people will you need? People who
- 43:28:45had had cancer in the past, isn't it? To
- 43:28:48know more about their problem, to know
- 43:28:50more about their medical condition. So
- 43:28:52is it possible people that in this use
- 43:28:54case you can go and pick up any person
- 43:28:57from the population and ask them
- 43:28:59questions? No. Right? That is not
- 43:29:02possible. So now is the probability
- 43:29:05equally likely or it has changed when
- 43:29:07you pick the sample? It has changed. Now
- 43:29:10there is a bias which is introduced that
- 43:29:13you only want people who had cancer.
- 43:29:16Right? So that kind of sampling people
- 43:29:19is called as nonprobabilistic sampling.
- 43:29:22Right? Non-robabilistic sampling. Clear?
- 43:29:26Now sampling technique. Right? Now third
- 43:29:30thing in this same scheme can be people
- 43:29:32what? It can be the data collection
- 43:29:37right data collection mode right that
- 43:29:40how do you collect the data? You can
- 43:29:42float a survey
- 43:29:45on say internet.
- 43:29:49You can go and stand outside a mall
- 43:29:53or office,
- 43:29:56isn't it? How will you how will you can
- 43:29:58probably interview someone,
- 43:30:01right? Interview someone, right? You can
- 43:30:03have a group discussion.
- 43:30:06Yeah. All these techniques.
- 43:30:09Yes. No, maybe. Right. In the same part,
- 43:30:11public opinion polling, right? Now,
- 43:30:14guys, uh so this was a brief
- 43:30:16introduction, right? And this can be
- 43:30:18extended to any industry. Right? As of
- 43:30:22now, you can have example in the medical
- 43:30:25science,
- 43:30:29right? You can have an example in
- 43:30:31automobile,
- 43:30:34right? You can have an example in
- 43:30:37retail,
- 43:30:39right? Right. Then you can have example
- 43:30:43in manufacturing.
- 43:30:47Right? You can have example in
- 43:30:49education,
- 43:30:52right? You can have example in sports,
- 43:30:56right? IPL analysis, cricket analysis,
- 43:30:59all these sports analysis, right? These
- 43:31:01are the most famous domains, right? They
- 43:31:04are the most famous domains. They are
- 43:31:06not topics, they are domains, right? In
- 43:31:08which data science is used extensively.
- 43:31:10So, I'm just going to check. Guys, in
- 43:31:12automobiles, there's a biggest example,
- 43:31:15self-driving cars.
- 43:31:18Yeah, autonomous driving. How do you
- 43:31:21think that's possible? Data science
- 43:31:24again
- 43:31:26like Tesla. Absolutely. Tesla is level
- 43:31:29three. We have five levels.
- 43:31:32Level three is narrow AI. Level four is
- 43:31:36AGI and level five is super AI. We are
- 43:31:39going to move first right with the
- 43:31:43technical aspect right with the
- 43:31:45technical aspect in our data science
- 43:31:47course right
- 43:31:57right which is based on Python
- 43:32:00right because that is our base language
- 43:32:02which we have learned so far right so in
- 43:32:06Python people we will start and cover
- 43:32:08four of the packages
- 43:32:12right now and as we move on to machine
- 43:32:15learning and other uh deep learning and
- 43:32:17everything you will explore more and
- 43:32:19more packages. The first package we have
- 43:32:21to cover will be numpy.
- 43:32:24Right? I'll explain you in detail what
- 43:32:26numpy is. Then we will cover pandas.
- 43:32:32Then we will cover mattplot lip
- 43:32:36and then finally we will cover something
- 43:32:38called as cbond.
- 43:32:41Right? We will cover something called as
- 43:32:43cbond. So these four packages inherently
- 43:32:45we have to cover in Python to make sure
- 43:32:51we are able to reduce
- 43:32:55the time
- 43:32:57in coding right we are able to reduce
- 43:32:59the time in coding and using these
- 43:33:02packages immediately help us in getting
- 43:33:05the desired results right I hope you all
- 43:33:07remember the concept of modules
- 43:33:10we have covered in Python do we all
- 43:33:12remember functions and modules. You can
- 43:33:14use Jupyter notebook. If your Jupyter
- 43:33:16notebook is not installed, you can use
- 43:33:18something which is called as Google
- 43:33:20Collab, right? Go to Google, type
- 43:33:24Collab,
- 43:33:26right? Let's say you write Collab,
- 43:33:30right? And then you will see this
- 43:33:32option. Click on Google Collab and it
- 43:33:36will allow you to code in Python, right?
- 43:33:40So we are now going to discuss about the
- 43:33:43numpy package in python. Okay, numpy
- 43:33:46package in python.
- 43:33:49So numpy is a
- 43:33:55fundamental
- 43:34:00package
- 43:34:02for data science in Python. Right? It is
- 43:34:08one of the most fundamental packages for
- 43:34:12practicing data science in Python.
- 43:34:14Right? Why is that? Why is so why is
- 43:34:17numpy so fundamental? What's is so
- 43:34:19special? NumPy package
- 43:34:25gives us a new data type
- 43:34:29for handling
- 43:34:32data in Python
- 43:34:36called as
- 43:34:39N D arrays, right? ND arrays which
- 43:34:44stands for
- 43:34:48this stands for
- 43:34:51N dimensional
- 43:34:54arrays right n dimensional arrays right
- 43:34:57this stands for n dimensional arrays
- 43:35:00so till now
- 43:35:03till now we have studied
- 43:35:07about
- 43:35:09list integer
- 43:35:12tpples,
- 43:35:14strings,
- 43:35:16right? Out of which
- 43:35:20out of which
- 43:35:25the data types
- 43:35:29such as list
- 43:35:32pupils have been used to store data,
- 43:35:38right? store data, right?
- 43:35:43And range
- 43:35:46used for generating
- 43:35:51new data which is primarily sequential.
- 43:35:55Right?
- 43:35:57Now there is one now there is one
- 43:35:58problem right? There is one problem and
- 43:36:01there should be a question that there
- 43:36:04should be a question.
- 43:36:09Why do we need a new data type
- 43:36:15to work with data science,
- 43:36:20right? Why do we need this? Yep. So the
- 43:36:24answer to this question people the
- 43:36:26answer to this question uh lies in a
- 43:36:29small explanation right? Yeah lies in a
- 43:36:33small explanation which is that
- 43:36:36Python
- 43:36:39is a
- 43:36:42high
- 43:36:44level
- 43:36:47language right? Python is a highlevel
- 43:36:51language, right? And
- 43:36:56a highle language
- 43:36:59is usually
- 43:37:03very
- 43:37:05distant
- 43:37:08from
- 43:37:10hardware.
- 43:37:13A highle language is close to hardware
- 43:37:15or distant from the hardware. Did you
- 43:37:17not attend the Python programming
- 43:37:19essentials?
- 43:37:21What is the type of programming language
- 43:37:24which is closest to hardware? A
- 43:37:26low-level language.
- 43:37:29If this is my hardware,
- 43:37:33yeah, this is my OS.
- 43:37:37This is my application layer. Right? So,
- 43:37:40hard level langu language is here and
- 43:37:43low-level languages here. Right? Which
- 43:37:45is closest. So binary languages,
- 43:37:48assembly languages,
- 43:37:51right? All these are closest to the
- 43:37:53hardware because where is the processing
- 43:37:57happening? Where is the processing
- 43:37:58happening of the data? At the hardware,
- 43:38:02isn't it? Processing of data
- 43:38:08is happening
- 43:38:11at hardware.
- 43:38:14No worry. Actually it is happening at
- 43:38:16hardware right? What is processing?
- 43:38:18Processing is signals of zeros and ones
- 43:38:20right? What are zeros and ones? These
- 43:38:23are electric signals.
- 43:38:25These are the electric signals right?
- 43:38:28Which is basically on and off. Right?
- 43:38:31And it is communicated to the hardware
- 43:38:33through the help of resistors
- 43:38:36and microprocessors.
- 43:38:40Right? Isn't it right? Why a computer
- 43:38:43only knows zeros and ones? Because zero
- 43:38:45is off and one is on which is the
- 43:38:48electric current
- 43:38:51right electric current to activate or
- 43:38:54deactivate certain things right true and
- 43:38:56false gates right so it is happening at
- 43:38:59hardware so now people if you understand
- 43:39:02this part then try to logically connect
- 43:39:05it to what I'm going to say when you are
- 43:39:09studying data science
- 43:39:13what kind of data you'll be dealing with
- 43:39:15big data,
- 43:39:19right? And as the name suggests, it will
- 43:39:21have a lot of volume,
- 43:39:23right? It will have a lot of volume
- 43:39:25other than a lot of other things, right?
- 43:39:28It will be very very big. And on this
- 43:39:31large volume of data, you'll be doing
- 43:39:34processing.
- 43:39:35You'll be doing processing. What does
- 43:39:37processing means?
- 43:39:40What does processing means? Operations.
- 43:39:43So where is this operation happening?
- 43:39:45This is happening in hardware
- 43:39:48and for hardware which is the closest
- 43:39:50language to hardware a low-level
- 43:39:53language.
- 43:39:54But now people but now we have a
- 43:39:57situation in front of us. What is the
- 43:39:59situation that what are we trying to do
- 43:40:01data science with? What are we trying to
- 43:40:03do data science with?
- 43:40:06Python, right?
- 43:40:09And Python is what?
- 43:40:11A high level language,
- 43:40:14isn't it? Yeah. So, there's a
- 43:40:17discrepancy. Yeah. A big one
- 43:40:20because hard level, high level language,
- 43:40:24these
- 43:40:26are slow
- 43:40:29in processing,
- 43:40:33right? These are very slow in
- 43:40:35processing, right? So for these kind of
- 43:40:38languages to handle this kind of data
- 43:40:41and these kind of operations yeah is
- 43:40:44very difficult right so let's let's keep
- 43:40:48let's keep this part aside if you
- 43:40:50understand this now let's go to the
- 43:40:52second point right
- 43:40:56when you learned Python
- 43:41:00on a scale of 1 to 10 how easy was it
- 43:41:05the ease of use of Python
- 43:41:07It's relatively a higher number, right?
- 43:41:10Relatively a higher number. So now guys,
- 43:41:13when I talk about data science, right?
- 43:41:16When I talk about
- 43:41:18data science, okay?
- 43:41:22Right? When I talk about data science,
- 43:41:25my thing is that this will be used by
- 43:41:30masses,
- 43:41:32right? Will be used by masses. A lot of
- 43:41:34people managers, programmers, business
- 43:41:37analysts, data analysts, possible right
- 43:41:41who are from nontechnical background who
- 43:41:43don't know coding they also can do data
- 43:41:45science because data science is a
- 43:41:47general thing isn't it? Understanding
- 43:41:49the data it should not be limited by
- 43:41:52your capability to understand the uh
- 43:41:55technicalities of a very complex
- 43:41:57language. So for these people which
- 43:42:00language is suitable which is Python
- 43:42:03right? It is easiest to understand. It's
- 43:42:05a high level language almost like
- 43:42:06English. Yeah. So, Python is a simple
- 43:42:09language. So, in this part people,
- 43:42:11Python fits the bill, right? Which is a
- 43:42:13bigger thing. In the second part, when
- 43:42:16we talk about the operations,
- 43:42:20we talk about the operations. In this
- 43:42:23part, there is a problem, right? Python
- 43:42:26fails,
- 43:42:30right? Python fails, right? because it's
- 43:42:33a highle language. We said that okay
- 43:42:36there is a language called as C right
- 43:42:39which is a middle level language
- 43:42:45right and C language people is used to
- 43:42:49create OS operating systems. It is used
- 43:42:52to create networks
- 43:42:55networking applications.
- 43:43:00It is used to create games,
- 43:43:03right? All the things which are close to
- 43:43:07hardware,
- 43:43:10C is used, right? So C fits this bill,
- 43:43:15right? C fits this bill.
- 43:43:18Python
- 43:43:20said that okay, if C fits the bill and
- 43:43:22Python is written
- 43:43:26in C, written on C, right? It's written
- 43:43:28on top of C language. Now what happened
- 43:43:30was Python said okay if my intrinsic
- 43:43:34data types my intrinsic processing is
- 43:43:37not suitable for data science but my
- 43:43:40highlevel nature is let's do one thing
- 43:43:43let's take C language and use its power
- 43:43:47right use its power that it is very
- 43:43:51close to hardware and let's create a new
- 43:43:55data type right let's create a new data
- 43:43:59type which is written on top of C and
- 43:44:04which can integrate with Python
- 43:44:07seamlessly. And people this new data
- 43:44:10type was called as array
- 43:44:16and this was given to you by something
- 43:44:18called as num py package. Right? Nump py
- 43:44:23package. It defined a new data type
- 43:44:26which was array. And along with defining
- 43:44:28the array, it gave various operations.
- 43:44:35Right? It gave various operations
- 43:44:38one could
- 43:44:40perform
- 43:44:42on arrays,
- 43:44:45right? One could perform on arrays.
- 43:44:49It's simple, right? program the the
- 43:44:51power of processing lied with C right it
- 43:44:55was lying with C. So we developed a new
- 43:44:59data type using C on top of Python and
- 43:45:03that new data type was called as array
- 43:45:06and this array was defined in a new
- 43:45:09module which was called as numpy module
- 43:45:13which told you how to create the arrays
- 43:45:15and then how to manipulate those arrays
- 43:45:19for doing data science. We had options
- 43:45:22like Java, we had options like C, right?
- 43:45:25We had options like forotron to be used
- 43:45:28for data science but we chose Python
- 43:45:30because of its simplicity and the simple
- 43:45:33syntaxes that people from
- 43:45:35non-programming background could also
- 43:45:38use Python to do data science.
- 43:45:42Right now the limitation was that
- 43:45:44because it is slow because of being high
- 43:45:46level we needed something which could
- 43:45:49make it fast and that was using an
- 43:45:52external data type which is not internal
- 43:45:54to Python and that was array and this
- 43:45:57array is defined inside a new uh module
- 43:46:01or a library called as num py which is
- 43:46:05numerical python right numerical python.
- 43:46:09Back to the programming right where we
- 43:46:12have understood right about this
- 43:46:15question right.
- 43:46:17So
- 43:46:19numpy
- 43:46:22essentially is
- 43:46:25built on top of
- 43:46:30C
- 43:46:31language which is
- 43:46:34compatible
- 43:46:37with Python.
- 43:46:40It leverages the power of closeness of C
- 43:46:48with hardware,
- 43:46:51right? C with hardware
- 43:46:53which eventually
- 43:46:56makes the processing
- 43:47:00faster in
- 43:47:05Python. Right?
- 43:47:08So this is what the first part is right
- 43:47:11this is what a first part is right now
- 43:47:14the second thing is people so this is
- 43:47:16about the performance bit right these
- 43:47:18are the performance bit the above
- 43:47:22is about the performance
- 43:47:28of
- 43:47:31right now coming to the next part which
- 43:47:34is the memory efficiency Y right memory
- 43:47:39efficiency right so
- 43:47:44arrays created
- 43:47:46by nump py in
- 43:47:51python
- 43:47:53are less memory
- 43:47:57exhaustive
- 43:48:00than lists in Python right and I will
- 43:48:04prove these points later on to you
- 43:48:06through code, right?
- 43:48:08In list, right? In list,
- 43:48:11each item is an object, right? Each item
- 43:48:16is an object, right? I hope you remember
- 43:48:18this, guys. Each item is an object,
- 43:48:21right?
- 43:48:23And it holds, right? It holds
- 43:48:28meta information
- 43:48:32like
- 43:48:34references
- 43:48:36and types,
- 43:48:38right? Etc., right? A lot of information
- 43:48:40it holds, right? This makes
- 43:48:46list consume
- 43:48:49more memory, right? But but
- 43:48:54arrays in
- 43:48:57numpy
- 43:48:59are contiguous
- 43:49:03which means
- 43:49:06that they do not
- 43:49:10create objects
- 43:49:12but rather
- 43:49:14directly store
- 43:49:18the data
- 43:49:20in
- 43:49:22continuous
- 43:49:24memory
- 43:49:27blocks
- 43:49:29one after another. Right? Also the
- 43:49:33arrays are homogeneous in nature. Right?
- 43:49:39You can only store the homogeneous data
- 43:49:42in array unlike lists. In list you could
- 43:49:45store different different data types.
- 43:49:47Right? But in arrays you cannot right.
- 43:49:49you have to store the same kind of data
- 43:49:52in the array homogeneous. So it's a
- 43:49:55contiguous memory block which is meaning
- 43:49:58that you can store data in continuity
- 43:50:02right one after another in the memory
- 43:50:04block. So the access is faster the
- 43:50:06memory location and the memory
- 43:50:08efficiency is very higher right as
- 43:50:11compared to the native data type like
- 43:50:13lists or tpples in Python. Arrays
- 43:50:17are basically
- 43:50:20vectorzed operations
- 43:50:23right I'll talk to about talk about to
- 43:50:25you with vectors what are vectors right
- 43:50:27they are the vectorzed operations and
- 43:50:29they are way more convenient
- 43:50:34to deal with as compared to list
- 43:50:41in Python right in list you have to go
- 43:50:44through a lot of loops, right? We saw
- 43:50:47that we have to go through the list
- 43:50:49comprehension. But you will see in
- 43:50:52Python pandas, sorry, in Python numpy,
- 43:50:55the vectorzed operations are very very
- 43:50:58simple, right? They are very very
- 43:51:00simple, right? Again, for all this, I
- 43:51:02will give you examples, but uh it will
- 43:51:05take some time because you have to
- 43:51:07understand what arrays are first, right?
- 43:51:13Okay. Then
- 43:51:16the most important
- 43:51:20other packages
- 43:51:22which are pandas,
- 43:51:26mattplot lib,
- 43:51:28cb bond,
- 43:51:30sklearn,
- 43:51:32cypy
- 43:51:34all are written
- 43:51:36on top
- 43:51:39of
- 43:51:40numpy
- 43:51:42package.
- 43:51:44That's the reason for this reason
- 43:51:47it's called as
- 43:51:50fundamental
- 43:51:52package.
- 43:51:54All the other packages which make your
- 43:51:56life easier as a data scientist where
- 43:51:58you don't have to worry about code.
- 43:52:02All these packages
- 43:52:05make our data science
- 43:52:09journey smooth
- 43:52:12because we have to worry
- 43:52:16less about code and more about
- 43:52:22logic.
- 43:52:23Yes. So for these for the understanding
- 43:52:26of these packages it's very important
- 43:52:29that we understand numpy first and then
- 43:52:31we move forward right
- 43:52:35now
- 43:52:38then let's get started. So the first
- 43:52:40step right the first step which you have
- 43:52:42to uh see right the first step which you
- 43:52:46have to see while using numpy package is
- 43:52:49basically from where will you import the
- 43:52:52package right from where will you import
- 43:52:54the package.
- 43:52:56So to import the package num py we write
- 43:53:03import nump py as np right where np
- 43:53:12is an alias right it's an alias
- 43:53:17right so please do this import nump py
- 43:53:20as np
- 43:53:23right if you don't get any result for
- 43:53:27this then you can write pip install
- 43:53:30nump py right pip install nump py and
- 43:53:34just execute this right when you will
- 43:53:37execute this it will give you this
- 43:53:39message or it will give you it will
- 43:53:41download this package for you right pip
- 43:53:45stands for
- 43:53:47python
- 43:53:49index package
- 43:53:52right
- 43:53:55it is basically like play store,
- 43:53:59app store,
- 43:54:01right? Or Windows store
- 43:54:06for Python,
- 43:54:08right?
- 43:54:10So, you're just going to these Play
- 43:54:12Store,
- 43:54:16Windows Store, App Store of your Python
- 43:54:19and asking them to download this for
- 43:54:21you, right? It is also called as the
- 43:54:25package manager
- 43:54:28right pip
- 43:54:30if it is done you can also check the
- 43:54:32version you can say np dot
- 43:54:35version
- 43:54:38and it will give you the version of
- 43:54:40numpy
- 43:54:42right
- 43:54:47numpy is an opensource
- 43:54:52package,
- 43:54:54right? Yeah, that's about it. Yeah, it's
- 43:54:58an open-source package, people,
- 43:55:01right? Open source package. And if you
- 43:55:05want to see the code, you can go to
- 43:55:10GitHub.
- 43:55:12Go to Google, write the uh code for
- 43:55:15numpy. It will show you it on GitHub.
- 43:55:18Right. done. So let's say I okay so I
- 43:55:23say
- 43:55:25we have when we so okay before this let
- 43:55:28me come to a little bit of theory before
- 43:55:30I do this with you. So now guys I said
- 43:55:33that num py
- 43:55:36has arrays
- 43:55:40as
- 43:55:43data type
- 43:55:45right and this is written on C which
- 43:55:49runs directly
- 43:55:52on hardware
- 43:55:57also this numpy array
- 43:56:01is basically basically a vector,
- 43:56:05right? It's basically a vector. Now,
- 43:56:07what is a vector, people? What is a
- 43:56:09vector? A vector is a quantity which has
- 43:56:14sign
- 43:56:16plus magnitude,
- 43:56:20right? It has a sign and magnitude. If I
- 43:56:23say this is a cartition space, this is
- 43:56:25I, this is J. And I say this this is 3 I
- 43:56:32and 4 J right 3 I cap 4 Jcap. So this is
- 43:56:37a vector right and this is the direction
- 43:56:41guys. If you have studied elementary
- 43:56:43maths you would know this.
- 43:56:46Yes people this is a vector. If I draw
- 43:56:49another like this
- 43:56:52then this is another vector. So I will
- 43:56:55call this say
- 43:56:57uh 2 I and 5 J. Yeah, this is another
- 43:57:01vector and this is the theta right. This
- 43:57:05is the direction.
- 43:57:06This is the direction. And what is the
- 43:57:09magnitude?
- 43:57:113 I 3² + 4² which is 9 + 16 which is 25
- 43:57:18under root which is 5. So the magnitude
- 43:57:22of vector is five and direction is equal
- 43:57:25to theta. This is a vector quantity.
- 43:57:29Now what is this? What is this? This is
- 43:57:33a scalar.
- 43:57:36This is a scalar. Yeah. Only magnitude
- 43:57:43isn't it? This is scalar only magnitude.
- 43:57:46And then I have 2a 3. Right? This is
- 43:57:50what? This is a vector.
- 43:57:55It has two dimensions
- 43:57:58or one dimension. Only one dimension,
- 43:58:02right?
- 43:58:04This is one dimension vector.
- 43:58:08Yep. One dimension vector. Now if I say
- 43:58:11this 2 3 4 5, what is this called? This
- 43:58:16is called a matrix,
- 43:58:18right? Right? This is called a matrix
- 43:58:20which is what collection of vectors
- 43:58:27and this collection of m vectors matrix
- 43:58:29is called as two-dimensional. Right? It
- 43:58:32is called as two-dimensional.
- 43:58:35Now if you have this
- 43:58:43so these are stacked behind each other.
- 43:58:47This is one. This is two. So this is
- 43:58:48three right? So we have three layers
- 43:58:55in matrix.
- 43:58:58So how many dimensions will be this
- 43:58:59people?
- 43:59:01One dimension, two dimension and three
- 43:59:03dimension. This is a threedimension
- 43:59:06matrix,
- 43:59:09right? Or threedimension vector
- 43:59:13or three-dimension array.
- 43:59:16I'll repeat once again. What is a single
- 43:59:19value? A single value is called as a
- 43:59:22scalar. Right? It only has magnitude.
- 43:59:24Now when you have multiple values,
- 43:59:27right? This is called as a vector. It
- 43:59:29has a direction. It has a magnitude. And
- 43:59:32this is single dimension. Now multiple
- 43:59:36vectors right like this or maybe you can
- 43:59:40say like this, right? Are you
- 43:59:43understanding why I'm calling it one
- 43:59:44dimension? It can be either this
- 43:59:46dimension or it can be this dimension.
- 43:59:48In any dimension you stack two vectors,
- 43:59:51you will get yourself a matrix. Right?
- 43:59:54You will get yourself a matrix which is
- 43:59:56now two dimensions. It has rows and it
- 44:00:00has columns.
- 44:00:03Right?
- 44:00:04Now if you stack multiple such matrix
- 44:00:09one after another, right? This becomes a
- 44:00:13threedimensional matrix. And this can go
- 44:00:16up to how many dimensions people? How
- 44:00:19many dimensions this can go up to? It
- 44:00:21can go up to this
- 44:00:25can
- 44:00:27go up to n dimensions,
- 44:00:32right? N dimensions,
- 44:00:35right? So now do we get it? Why do we
- 44:00:37call it ND
- 44:00:40arrays?
- 44:00:42Yeah, n dimension arrays,
- 44:00:48right? That why are we calling something
- 44:00:50what?
- 44:00:52Right. N dimension arrays. We are only
- 44:00:55capable of viewing three dimensions
- 44:00:57people. It can go up to 100 dimensions,
- 44:00:59500 dimensions, 1,000 dimensions, any
- 44:01:03dimensions.
- 44:01:04Scalas are least important.
- 44:01:07vectors are more important. So now we
- 44:01:10have
- 44:01:14multiple
- 44:01:15dimensions in
- 44:01:19arrays
- 44:01:23namely 0D,
- 44:01:261D,
- 44:01:292D, 3D and so on till the N D, right? NV
- 44:01:36arrays. Right now let's create our first
- 44:01:40array. Right? Let's create our first
- 44:01:42array. And this will be a zero
- 44:01:47dimension
- 44:01:52array, right? Zero dimension array. How
- 44:01:55will you create this? You will say a r0
- 44:02:00is equal to np dot array and you will
- 44:02:05mention what a scalar value is. What is
- 44:02:07a scalar value? It is a simple value. I
- 44:02:11say two, right? np dot array equal to
- 44:02:13two. And when you will now check or
- 44:02:17print the type of ar r0, it will tell
- 44:02:22you class num py nd array. Right? It is
- 44:02:27a zero dimension array. And if you want
- 44:02:30to check
- 44:02:33the dimension,
- 44:02:35you just have to write a r0 dot end div,
- 44:02:43right? And it shows you that there is
- 44:02:46zero dimensions present. So what is this
- 44:02:48in short? This is a scalar. I have
- 44:02:54entered a single value people 2 200 500
- 44:02:58whatever you want to enter. And the
- 44:02:59syntax is np dot array np dot array.
- 44:03:04You're instructing nump py package to
- 44:03:06fetch the function array on method array
- 44:03:10and convert this into that particular
- 44:03:14data type right array zero
- 44:03:18and when you check the type type is nd
- 44:03:20array but what is the dimension of this
- 44:03:22nd array this is zero which is nothing
- 44:03:24but a scalar right we have created this
- 44:03:29this is what we have created
- 44:03:32Right
- 44:03:35now guys,
- 44:03:38if I created this list, right?
- 44:03:45Say Lis is equal to
- 44:03:50Yeah. So this was this was a list,
- 44:03:53right? This was a list and this was this
- 44:03:55is what people what is this that we have
- 44:03:59just studied according to that? What
- 44:04:01dimension is this list?
- 44:04:03So what I'm trying to do is I will
- 44:04:06create a
- 44:04:09one deal
- 44:04:14array with a or let's say from a list
- 44:04:20right let's let me show you how do we do
- 44:04:22that okay
- 44:04:29so I say
- 44:04:31hurry read the error name ar r not
- 44:04:34defined. Why? Because you have created
- 44:04:36array from a a r0. Come on hurry.
- 44:04:41Right.
- 44:04:43I will create a onedimensional array.
- 44:04:44How will I do that people? I will say a
- 44:04:46ar r1 is equal to np dot array. And can
- 44:04:51I pass 1D list inside this? Or can I say
- 44:04:55I can pass lis inside this? When I do
- 44:04:58this people now see what will happen. A
- 44:05:01R R1 will be equal to this right and if
- 44:05:05you say let me say print a ar a ar a ar
- 44:05:07a ar a ar a ar a ar a ar a ar a ar a ar
- 44:05:08r r r r r r r r r r r r r r r r r r r r1
- 44:05:10it will be like this okay this is your
- 44:05:15a ar r r1 dot nim
- 44:05:20you will see that it gives you one right
- 44:05:22this is a onedimension array right on
- 44:05:26dimension array
- 44:05:43Yes. So for example
- 44:05:45when I say scalar right when I say
- 44:05:49scalar I say
- 44:05:5335s right? when I say vector
- 44:05:581D I say
- 44:06:02uh
- 44:06:03so these are suppose my marks
- 44:06:06right now I say 35
- 44:06:0940
- 44:06:1150 right so now what are these my marks
- 44:06:16in three subjects
- 44:06:19yeah marks in three subjects this is my
- 44:06:22say Hindi this is English and this is
- 44:06:26maths right or let's say science because
- 44:06:29not everyone has Hindi science English
- 44:06:32and maths right guys now if I have to
- 44:06:34create a matrix
- 44:06:37of two dimension what will I write so
- 44:06:40that means this is one student this is
- 44:06:43one student people isn't it guys yes no
- 44:06:47maybe so now in matrix we will have what
- 44:06:50we will have multiple students
- 44:06:55Yes. No. Maybe in a matrix people we
- 44:06:58will have multiple students. Suppose
- 44:07:00this was S1. Now you will have S_sub_1,
- 44:07:03S_UB_2, S3, S4 like this.
- 44:07:08And each student will have their own
- 44:07:10individual list of marks. So can I say
- 44:07:14that I'm making a nested list?
- 44:07:19Can I say that people? I'm making a
- 44:07:21nested list. So now let's make it okay
- 44:07:28from a
- 44:07:31nested list. Okay, a nested list.
- 44:07:36So I'll say ar r2 is equal to np array
- 44:07:39list. I will have to create a list
- 44:07:41first. Lis2 is equal to. So this is my
- 44:07:45first bracket. What is this bracket
- 44:07:47representing? this bigger bracket. Now I
- 44:07:49will put another bracket inside this and
- 44:07:51I will write 1 1 22 33 3. I'll put a
- 44:07:54comma again. Write a comma. Then I will
- 44:07:58say
- 44:07:594455 666 comma 778899
- 44:08:05right I'll do this right now I will say
- 44:08:09list to
- 44:08:13two
- 44:08:20when you will do this you will see that
- 44:08:22an array like this has been created
- 44:08:25Right?
- 44:08:28like this AR R2
- 44:08:30right this has been created right when
- 44:08:34we check the dimension it is two
- 44:08:36dimension array right
- 44:08:40yeah marks of three different students
- 44:08:43in three different subjects
- 44:08:46and this is the same technique you can
- 44:08:49create a three-dimension array how will
- 44:08:51you create a three-dimension array
- 44:08:52people
- 44:08:54if I go here how will you create a
- 44:08:56threedimension array
- 44:08:59Now suppose I have data in this and this
- 44:09:03is my master list. Okay, this is my
- 44:09:04master list. In this I have data
- 44:09:08and this can be represented like this.
- 44:09:14This is my
- 44:09:16first matrix isn't it? And this is the
- 44:09:21vector inside this
- 44:09:26V_sub_1, V_sub_2, V3.
- 44:09:29Then this can be called as M1. And now
- 44:09:34to create a three-dimension setup, how
- 44:09:37many M1s do you need? You need multiple
- 44:09:40M1s, isn't it? You need M1, M2, M3,
- 44:09:44multiple matrix like this people.
- 44:09:48like this matrix 1, matrix 2, matrix 3.
- 44:09:54So what will you do? You will have you
- 44:09:57will have what people?
- 44:10:00You will have another yellow, right?
- 44:10:04And you will have inside this yellow
- 44:10:07multiple purples.
- 44:10:12Isn't it
- 44:10:13right? Again like this.
- 44:10:27Yeah. Like this you will have it people.
- 44:10:31So can I say people can I say that as I
- 44:10:35am increasing the dimensions as I am
- 44:10:41increasing
- 44:10:46the dimensions
- 44:10:49I am putting 1D sorry 0D right and okay
- 44:10:54again in this V_sub_1 in this V_sub1 do
- 44:10:58you think you will have multiple scalers
- 44:11:00people can I say that
- 44:11:05can I say multiple scalers create a
- 44:11:08vector multiple vectors create a matrix
- 44:11:13And multiple matrix create one
- 44:11:15three-dimensional matrix.
- 44:11:18Can I say that? Let me talk to you about
- 44:11:21an image. Right?
- 44:11:24Image, right? What is an image made up
- 44:11:30of people?
- 44:11:33What is an image made up of?
- 44:11:38H
- 44:11:41pixels.
- 44:11:44Yes or no? No,
- 44:11:47not frames. Frames is basically videos.
- 44:11:52Pixels are creating an image, right? So,
- 44:11:54we have how many pixels here? 1 2 1 2 3
- 44:11:584 5 6 7 8 9 10 11 12 13 14 15 16. Right?
- 44:12:04Suppose
- 44:12:06this is one vector, right? This is one
- 44:12:08vector and this is one scalar
- 44:12:13right? Scalar 1, scalar 2, scalar 3,
- 44:12:15scalar 4. And this will make vector
- 44:12:18v_sub1.
- 44:12:20This is v_sub_2, v_ub3, v4. And together
- 44:12:24together can I call this m_sub_1 and
- 44:12:28call this blue?
- 44:12:30So for a colored image, how many
- 44:12:33channels are there people? How many
- 44:12:35channels are there? What do we call it?
- 44:12:37We call it the image as
- 44:12:41RGB
- 44:12:43RGB image
- 44:12:45that means red,
- 44:12:48green,
- 44:12:50blue.
- 44:12:52So this is blue part. So similarly you
- 44:12:55will have a red part in front of it.
- 44:12:58Then you will have a green part and then
- 44:13:01finally you will have a blue part. Yeah.
- 44:13:04Are you understanding guys? Why do we
- 44:13:07require three-dimensional arrays?
- 44:13:09Yes. Suppose now you want to make a
- 44:13:12change at this this pixel. So you will
- 44:13:15go to the third layer which is the blue
- 44:13:18layer. Then you will go to the third
- 44:13:20column. You will go to the third column
- 44:13:22and third row. And this is how you will
- 44:13:24reach this pixel. Everyone? Yes. No.
- 44:13:28Maybe.
- 44:13:29Yes. So this is the reason why we need
- 44:13:32to create a 3D array.
- 44:13:36Right. To ingest information like this,
- 44:13:39right? To ingest information like this.
- 44:13:41So we can create a 3D array also. Right.
- 44:13:50Right. And how did I tell you? How many
- 44:13:52brackets will I have? Squared brackets.
- 44:13:54I will have three squared brackets. So
- 44:13:56now this is just one student.
- 44:14:00Right. Now I'll put a comma here.
- 44:14:04Right? Right, I'll put a comma here and
- 44:14:06I will start.
- 44:14:14Right, I'll do this. I'll say this list
- 44:14:17three
- 44:14:19a r3
- 44:14:21list three ar r r3
- 44:14:24a r3
- 44:14:26r
- 44:14:29and now you will see that there's a
- 44:14:31threedimensional array
- 44:14:33right guys
- 44:14:42right this is a threedimensional
- 44:14:44array
- 44:14:51Yep. Moving on. There are multiple ways
- 44:14:57create arrays, right? The first one we
- 44:15:01have done.
- 44:15:03So we have done
- 44:15:05from lists.
- 44:15:08From list we have done.
- 44:15:11Then second will be
- 44:15:13from uh we can create a zero array
- 44:15:19right we can create
- 44:15:22on's array
- 44:15:24right then we can create custom array
- 44:15:31right I will show this all to you
- 44:15:36right I'll show this all to you so let's
- 44:15:38start with the on's array sorry zero
- 44:15:41those array
- 44:15:44right what do you have to do you have to
- 44:15:47write so let's say zero
- 44:15:50ar r r0 okay a zero dimension zero array
- 44:15:52so you say np dot zeros
- 44:15:56right and you create
- 44:15:58a two
- 44:16:03right
- 44:16:04a two right and if I say
- 44:16:10uh 0
- 44:16:12dot end
- 44:16:14right you will see it is a onedimension
- 44:16:17array
- 44:16:20right it is a onedimension array
- 44:16:24right so two is by default taken as so
- 44:16:28let me just show this to you
- 44:16:31it will look like this right it is taken
- 44:16:33as horizontal what is the dimension of
- 44:16:34this guys a vector what is the dimension
- 44:16:37of this vector
- 44:16:39no no it's 1. It is basically 2 + 1,
- 44:16:43right? 2a 1.
- 44:16:46Sorry, 1 comma 2. My bad.
- 44:16:501 comma 2. Isn't it? Now, what if you
- 44:16:53had to create a 2 + 1? What? What if you
- 44:16:57had to create a 2 + 1? Right? So, let me
- 44:17:00just show that to you.
- 44:17:03So, I say wait
- 44:17:08like this. Okay.
- 44:17:13Now when I do this, I say a ar r r once
- 44:17:17and I say here
- 44:17:202, 1, right? I say 2a 1. Now you will
- 44:17:26see people what will happen to this.
- 44:17:31Now
- 44:17:33because you have created right specified
- 44:17:37two dimensions, right? What will this be
- 44:17:39converted to now?
- 44:17:41H what will this be converted to? This
- 44:17:45will be converted to
- 44:17:48a
- 44:17:49two-dimension vector. By default, it was
- 44:17:52this, right? Which was this vector,
- 44:17:56right? By default, it was this vector,
- 44:18:00right? What is the dimension of this
- 44:18:01vector? 1 + how many values you put
- 44:18:04here, right? 1 + 2 like this. So we call
- 44:18:08this only one dimension. We call this
- 44:18:10only one dimension. But now when I will
- 44:18:13run this one, you will see yes 2 + 1.
- 44:18:18And now you will see this is
- 44:18:22two dimension, right? You see this is
- 44:18:24two dimension. Now clear people just the
- 44:18:27orientation has changed. But now you see
- 44:18:29the brackets there are two brackets now
- 44:18:31because what have you now instructed?
- 44:18:33You have now instructed Python and
- 44:18:36rather numpy to create the vector as 2 +
- 44:18:401. So you have said give me this zero
- 44:18:42and give me this zero here. So the
- 44:18:45moment you do this you are now
- 44:18:47specifying the rows and columns.
- 44:18:52Now this 21 right let's say this is 21.
- 44:18:58Now let me create a another two cross
- 44:19:00two dimension matrix for you. Let me
- 44:19:02call this 10
- 44:19:05comma 10. Right? So how many rows and
- 44:19:07columns will it have people?
- 44:19:10How many rows and columns will it have?
- 44:19:13I will say 21.
- 44:19:19It will have 10 rows and 10 columns like
- 44:19:22this. You saw this? Yeah. 10 rows and 10
- 44:19:27columns. Right?
- 44:19:31By default, it is 1 +2 like this. It is
- 44:19:33a one-dimensional vector and this is a
- 44:19:35two-dimensional vector. What about a
- 44:19:37three-dimensional vector? 0 dot 0
- 44:19:42uh a ar r3
- 44:19:45is equal to np dot zeros, right? np
- 44:19:48do.zer,
- 44:19:50right?
- 44:19:54H should I just write 3a 3a 3? And I
- 44:19:59should check for this.
- 44:20:02Yeah, you will have a 3 +3 vector 3 + 3
- 44:20:06matrix with three matrix stacked behind
- 44:20:08each other. Right? If I say four, this
- 44:20:13will be four. So how do we read this?
- 44:20:20number of layers,
- 44:20:24number of rows and number of columns.
- 44:20:30So, can you help me with a syntax? Can
- 44:20:32you help me with a syntax which can
- 44:20:35create me a 3D matrix of five layers,
- 44:20:40three rows and three columns? What will
- 44:20:42I write?
- 44:20:44Five layers, three rows and three
- 44:20:46columns. What will I write?
- 44:20:535 33. Yeah, you'll get five layers. 1 2
- 44:20:573 4 5 right like this. Suppose suppose
- 44:21:02you have to represent
- 44:21:06an image
- 44:21:08with
- 44:21:10with say
- 44:21:13RGB channel
- 44:21:15and
- 44:21:19uh say 256 and 256 pixels. How will you
- 44:21:24create this? So I'll say img right image
- 44:21:29is equal to np dot
- 44:21:32zeros and I will say inside this
- 44:21:37three channel 256 cross 256
- 44:21:42right and when you will just run this
- 44:21:44image it will be like this right it will
- 44:21:47be like this this is one channel this is
- 44:21:51two channel and this is three channel
- 44:21:55And uh we understood that numpy is a
- 44:21:58fundamental package for data science in
- 44:21:59Python. Right? This package gives us a
- 44:22:02new data type to work with which is
- 44:22:05called as n- dimensional arrays. Right?
- 44:22:08Now what are arrays? What are arrays?
- 44:22:11Arrays are nothing but vectors right
- 44:22:15which are stored in a contiguous block
- 44:22:17of memory which means they are stored
- 44:22:20continuously one after another and they
- 44:22:22do not get converted into the object
- 44:22:24unlike the list and they are way faster
- 44:22:27they are more memory efficient than list
- 44:22:29I will prove this fact to you today with
- 44:22:31the help of example through the help of
- 44:22:32code right so why numpy arrays because
- 44:22:36numpy is a package which is built on top
- 44:22:38of C language which is compatible with
- 44:22:40python and C being a middle level
- 44:22:43language interacts directly with the
- 44:22:45hardware. So whatever operation you run
- 44:22:49in a fact that is getting directly
- 44:22:50executed on the hardware itself. Right?
- 44:22:53That is the reason why the uh
- 44:22:56performance is way better when we try to
- 44:22:58use the n dimensional arrays. Along with
- 44:23:00this these syntaxes the type of syntaxes
- 44:23:03we used to type uh in list I will show
- 44:23:05that to you also today with the help of
- 44:23:07example are way simpler when you try to
- 44:23:10do them with nd arrays right so arrays
- 44:23:15are basically vectorized operations
- 44:23:17right and they're also very convenient
- 44:23:20to deal with as compared to lists and
- 44:23:21other native data types in python right
- 44:23:24and the most important part is that in
- 44:23:26our data science journey whatever other
- 44:23:29packages packages we will use. Right?
- 44:23:30Again, what are packages? They have
- 44:23:32predefined things stored for you so that
- 44:23:34you can leverage them and focus less on
- 44:23:37code and more on logic. Right? You need
- 44:23:39to be aware about the logic more than
- 44:23:42the knowledge of the code. Right? So, we
- 44:23:44have packages for that which contains
- 44:23:46methods inside them which you can use
- 44:23:49and uh without any further calculations
- 44:23:52you can work with them directly. Right?
- 44:23:56So that is how we started with numpy
- 44:23:59right and the syntax to import numpy was
- 44:24:02import numpy as np where np was an alias
- 44:24:06right. Uh you could do pip install numpy
- 44:24:08if someone did not have access to numpy
- 44:24:10if numpy was not coming by default. You
- 44:24:13can use pip install numpy which is
- 44:24:15python index package right and uh this
- 44:24:19is like play store app store window for
- 44:24:21python. All the packages are stored in
- 44:24:23pip and you can call pip you can ask pip
- 44:24:26to download that package for you so that
- 44:24:28you can use that right it's basically
- 44:24:29the package manager so numpy is an open
- 44:24:32source and if you want to see the code
- 44:24:34you can go to github and check the code
- 44:24:35out for yourself right in numpy we have
- 44:24:38n dimensional arrays now the question
- 44:24:40arises people that why do we need numpy
- 44:24:43right so my answer to this particular
- 44:24:45question is that when you will deal in
- 44:24:48data science right when you will become
- 44:24:49a data scientist you will be dealing
- 44:24:51with data Right now my question is how
- 44:24:53will you ingest how will you make the
- 44:24:56machine ingest the data right there has
- 44:24:59to be a way right for you to input the
- 44:25:01data to the machine right to make
- 44:25:03manipulations to the data to read the
- 44:25:05data so all this is started as the base
- 44:25:08package of numpy arrays right other than
- 44:25:11that it becomes very difficult and
- 44:25:13cumbersome for us to deal with that and
- 44:25:15this provides us a lot of ease and
- 44:25:17flexibility to deal with such massive
- 44:25:20amounts of data which you will along
- 44:25:21with me as we will move forward in this
- 44:25:23particular course. Yeah, perfect. Now
- 44:25:27people, let me just uh pull up the PBTs.
- 44:25:31This is what we're discussing people. We
- 44:25:34start with something which is called as
- 44:25:36scalar, right? We start with something
- 44:25:38which is called as scalar which is a
- 44:25:41quantity which only has magnitude.
- 44:25:43Right? In numpy language, this is also
- 44:25:46called as 0D, right? It is called as
- 44:25:48zero dimension. Then people we have
- 44:25:50vector right and vector has a constant
- 44:25:53dimension. It has only one dimension
- 44:25:55right you can interpret it as a row or
- 44:25:57you can interpret it as a column it
- 44:25:59doesn't really matter because this is
- 44:26:01only one single dimension right so
- 44:26:03usually it will be written as five comma
- 44:26:06blank right there will be nothing
- 44:26:07written in front of it so this in numpy
- 44:26:10terminology and nomenclature is called
- 44:26:12as one dimension right when you move on
- 44:26:16then you get combine couple of vectors
- 44:26:19you get a shape and now that is called
- 44:26:21as a matrix and In numpy terminology it
- 44:26:24is called as two-dimension right and
- 44:26:27when you try to stack multiple
- 44:26:29two-dimension matrices one before with
- 44:26:31one after each other or one before each
- 44:26:32other then they become something called
- 44:26:35as three-dimensional and now in my
- 44:26:37capability I don't know what four
- 44:26:38dimension looks like but there is a high
- 44:26:40possibility that you have n dimensional
- 44:26:43data right it has it has n we are
- 44:26:45dealing with n dimension data right
- 44:26:48suppose with this I also add time right
- 44:26:52at t equal to 1 at t=2 that will serve
- 44:26:55as the fourth dimension for this data
- 44:26:57but how do how does it look like I don't
- 44:26:59really know that right because humans
- 44:27:00are only capable of visualizing 3D three
- 44:27:04dimensions at max right so you can go to
- 44:27:06n dimensions and hence the name n
- 44:27:09dimensional arrays right nd arrays post
- 44:27:12this right post this we moved on to
- 44:27:15create certain things and I tried to
- 44:27:17explain you the data right so the data
- 44:27:20will look to you like this right You
- 44:27:22might have a scalar quantity which is
- 44:27:24marks right one marks right now if I go
- 44:27:28on to vectors in one day it can be marks
- 44:27:31of one student
- 44:27:34right marks of one student in science
- 44:27:38English
- 44:27:40and maths right 35
- 44:27:4340 and 50 out of say 50 right three
- 44:27:47subjects so this will be characterized
- 44:27:49as a vector right what will be the
- 44:27:51dimension written for For this it will
- 44:27:52be 3 comma nothing. This will be zero
- 44:27:56right shape will be zero. For this
- 44:27:58matrix suppose we have 1 2 3 four
- 44:28:02students and each student will have
- 44:28:05three marks.
- 44:28:12Right? Each student will have three
- 44:28:13marks. So what will be the shape of
- 44:28:15this? We have four rows and three
- 44:28:19columns. Right? So this will be the
- 44:28:21shape right of this 2D matrix right and
- 44:28:25now if you stack images right one behind
- 44:28:28each other then it will be like this
- 44:28:30right image I gave you an example so
- 44:28:33this has 4 + 4 pixels so the shape will
- 44:28:36be 3 + 4 + 4 right this will be the
- 44:28:40shape for this particular 3D matrix
- 44:28:43right guys so this is how you input the
- 44:28:46data just to tell you a little bit more
- 44:28:48since generative AI is very popular
- 44:28:51these days. Right? So what if I tell you
- 44:28:54the fact that the Chad GPT
- 44:28:58which you use or you might have used
- 44:29:02right has
- 44:29:05never seen
- 44:29:07a single
- 44:29:10word
- 44:29:13in its lifetime.
- 44:29:17Right? All it sees
- 44:29:21is numbers, right? Only numbers. How do
- 44:29:25we see numbers? Suppose I say my
- 44:29:29name
- 44:29:30is
- 44:29:32Raghav.
- 44:29:35Raghav
- 44:29:37is a
- 44:29:40nice
- 44:29:41name. Right? So these are two data,
- 44:29:44right? These are two data points. Now we
- 44:29:47all know that computers do not
- 44:29:49understand these right computers do not
- 44:29:51understand these right there is nothing
- 44:29:53no understanding for computers to know
- 44:29:55what text is right it only knows 0 and
- 44:29:58one yes dhika right it only knows zeros
- 44:30:00and ones so see how we will convert this
- 44:30:02so there is something called as
- 44:30:04vocabulary
- 44:30:07right so vocabulary are nothing but the
- 44:30:10unique words
- 44:30:12right how many unique words do I have in
- 44:30:14this my name is Raga four. This is not
- 44:30:18unique. This is repeating. This is
- 44:30:19repeating. Five, six. And this is
- 44:30:22repeating. So I have six words. So now
- 44:30:24guys, I will do something called as word
- 44:30:27to
- 44:30:29right where I will convert these words
- 44:30:30into vectors. How will I convert them?
- 44:30:32Look at this. So I will have suppose
- 44:30:35this is S_sub_1,
- 44:30:37this is S_sub_1 and this is S_sub_2,
- 44:30:39right? So I will represent
- 44:30:42S1 as
- 44:30:47right. I will have a vector
- 44:30:51of size six. How? I will say 1 0 0 0.
- 44:30:59Right? 1 0 0. How many elements does it
- 44:31:02have? Six elements. Right? Name will be
- 44:31:050 1 0 0 0.
- 44:31:08is will be 0 0 0 1 0 0 0
- 44:31:13and ra will be 0 0 0 1 0 0 right this is
- 44:31:18s1 my name is raghub now when it comes
- 44:31:22to s_ub_2 right when it comes to s_ub_2
- 44:31:26how will I enter this s2
- 44:31:280 1 0 0 ragh what is is here
- 44:31:3400 0 1 0 0 0
- 44:31:37what is
- 44:31:380 0 0 1 0 or nice is 0 0 0 1 and name
- 44:31:46will be 0 1 0 0 0 0 right now this will
- 44:31:51be the vector representation of these
- 44:31:54two sentences just to tell you a fact
- 44:31:58GBD3 right GPD3 model right GPD3 or
- 44:32:03GPD3.5
- 44:32:05they have vocabul vabulary
- 44:32:09of 30,000 words, right? 30,000 words.
- 44:32:13And each word, right? Each word
- 44:32:19has a
- 44:32:22dimension
- 44:32:25of
- 44:32:2612,500
- 44:32:29numbers. Right? What do I mean? Suppose
- 44:32:32I say Raghub.
- 44:32:34So it will be one word and it will be
- 44:32:37represented by five 11,500
- 44:32:41different numbers
- 44:32:45right and like raghub there will be
- 44:32:4730,000 words in this GPD model right
- 44:32:5230,000 words in this GPD model right and
- 44:32:56this total number of parameters which
- 44:32:58get trained in the neural network which
- 44:33:00we will learn later on neural networks
- 44:33:02they are almost close to 1
- 44:33:0775
- 44:33:09billion
- 44:33:11parameters right 1.75 billion parameters
- 44:33:15so why I'm telling you all this because
- 44:33:18to showcase to you that what is the
- 44:33:20importance of vectors right in the
- 44:33:23entire machine learning and data science
- 44:33:26yeah people so this is the reason people
- 44:33:29now my question is how will you create
- 44:33:30these vectors how will you read these
- 44:33:32vectors Right? The answer is through
- 44:33:35numpy package because it is the base
- 44:33:38package. Clear people? Yeah. I hope
- 44:33:41today's class will be uh interesting for
- 44:33:43you because you will know the context.
- 44:33:44Why are we doing it? Yeah. So I'll try
- 44:33:46to show that to you how we convert
- 44:33:48things to vectors.
- 44:33:50Right? Okay. Let me go here now. Right.
- 44:33:53Let me go here.
- 44:33:55So people uh we started using numpy. So
- 44:33:59I started with the zero dimension
- 44:34:01arrays. Right? Zero dimension arrays. So
- 44:34:03zero dimension is nothing but a scalar.
- 44:34:05So I created a ar r0 which was np dot
- 44:34:08array and I entered a single word single
- 44:34:11uh element inside this which is nothing
- 44:34:12but a scalar and then I checked the type
- 44:34:15of uh ar0 also right and then I check
- 44:34:19the dimension also. So the answer was
- 44:34:21two class was numpy nd array and the
- 44:34:24dimension was zero right exactly what I
- 44:34:27had mentioned in my pb
- 44:34:31right it will be having a zero dimension
- 44:34:35like this right same thing has been
- 44:34:38proven
- 44:34:39right because it's a scalar now coming
- 44:34:42on to one dimension right coming on to
- 44:34:44one dimension I create a list which is
- 44:34:46nothing but a one-dimension data type
- 44:34:49right now I create an array which is a
- 44:34:51ar r1 from array from this particular
- 44:34:54list lis and then I check the type of
- 44:34:58print ar1 check the type of ar1 and the
- 44:35:01dimension right so when I execute
- 44:35:07right then you will see that it was this
- 44:35:10numpy array and the dimension was one
- 44:35:13right exactly like this so if I show you
- 44:35:15something else say print a ar r r1 one
- 44:35:19dot shape
- 44:35:26you will see it's 4 comma empty right
- 44:35:28and I show it to you here
- 44:35:35it will be empty right empty and this
- 44:35:39will be comma 1 right so this means that
- 44:35:42it has only four elements right if I
- 44:35:45increase these elements to say 55 5 66
- 44:35:4977
- 44:35:51then it will become 7, blank, right?
- 44:35:54Which means it has seven elements as a
- 44:35:56vector. Now we create something with a
- 44:35:59nested list right which is like this. So
- 44:36:02with a nest I want one bracket which is
- 44:36:05running outside right then inside this I
- 44:36:08have one two and three lists inside one
- 44:36:11list. Right? So this is a nested list.
- 44:36:13This is marks of first student, second
- 44:36:15student and the third student. Right? So
- 44:36:19I do this and you see it is this right?
- 44:36:22And I will show you the shape also
- 44:36:30right. It will be 3 + 3 rows and three
- 44:36:33columns. Three rows and three columns.
- 44:36:36Right? Now similarly we can also create
- 44:36:39a
- 44:36:50the 3D matrix
- 44:36:52right with
- 44:36:56two levels right level one and level two
- 44:37:00and 3 + 3 so the shape will be what
- 44:37:03people can someone guess the shape
- 44:37:06what will be the shape of this I've
- 44:37:09shown you
- 44:37:11Right? If this is 3 4 then what will be
- 44:37:14this?
- 44:37:16Three rows and three columns. Right? So
- 44:37:19when you will execute you will get 2 3 3
- 44:37:22right 2 3 3
- 44:37:24right
- 44:37:26right now people there are multiple ways
- 44:37:28right there are multiple ways to create
- 44:37:30arrays and we should know them because
- 44:37:32all of these comes very very handy. Not
- 44:37:35right now. I don't have enough context
- 44:37:37to give you right now. But later on you
- 44:37:39will see with me or with some other
- 44:37:41trainer that how these will be used in
- 44:37:43deep learning specifically, right? They
- 44:37:45are the key of deep learning algorithms,
- 44:37:49right? Where we initialize some weights,
- 44:37:51we initialize some biases and those
- 44:37:53initializations are nothing but
- 44:37:55multi-dimensional numpy arrays, right?
- 44:37:58Numpy arrays.
- 44:38:01Okay.
- 44:38:05Like for example, suppose I have this
- 44:38:12I have to multiply this with some random
- 44:38:15numbers, right? So how will you generate
- 44:38:17these random numbers? You will generate
- 44:38:19them through numpy. And you can generate
- 44:38:21them in a specific kind of uh shape,
- 44:38:25right? Which is 2 + 3. And then you can
- 44:38:28multiply them. You can multiply the
- 44:38:30matrices and you can get your output for
- 44:38:33yourself. Right? So this is the way they
- 44:38:36are used.
- 44:38:38So we saw the first thing from list we
- 44:38:41have already covered. Then now we are
- 44:38:42moving on to creating zero arrays right.
- 44:38:45So I create a zero array of one
- 44:38:49dimension right of one dimension which
- 44:38:51is 0 0. They are represented in floats
- 44:38:54right. They are represented in floats
- 44:38:560.0 zero. Right? Now you can create a
- 44:39:00two-dimensional zero array. Right? You
- 44:39:02can create a two-dimensional zero array
- 44:39:04which is you have to mention just the
- 44:39:05shape inside 2 + 1. So it will have two
- 44:39:08rows and one columns, right? Two rows
- 44:39:10and one columns. The difference here is
- 44:39:12the difference here is that these are
- 44:39:15one dimension and these are two
- 44:39:17dimensions. Right? You have explicitly
- 44:39:18mentioned the rows and columns. So you
- 44:39:20can expand this to 10 + 10 also.
- 44:39:27Right? You can expand in 10 + 10 or 10 +
- 44:39:296 whatever you feel like yourself.
- 44:39:31Right? It will have 10 rows and six
- 44:39:35columns. Right?
- 44:39:38Now you can also create
- 44:39:41the 3D arrays 3D zero arrays.
- 44:39:47Right?
- 44:39:50Which is 5a 3a 3. What does five means?
- 44:39:53What does five means? First element
- 44:39:55represents the number of layers. So you
- 44:39:57have five layers, right? It's five layer
- 44:40:00deep. Then you have three rows and three
- 44:40:03columns, right? So it will look
- 44:40:05something like this.
- 44:40:151
- 44:40:182
- 44:40:241 2 3 4 5 right like this something like
- 44:40:27this right it will look like this tab 1
- 44:40:302 3 1 2 3 1 2 3 right 5 33 okay yes the
- 44:40:35number of matrices hurry what I
- 44:40:37represented
- 44:40:39right layers
- 44:40:41rows
- 44:40:43columns right layer rows and columns.
- 44:40:46Clear?
- 44:40:49So now when you execute this, you will
- 44:40:50get an arrangement like this. Okay. I
- 44:40:52try to show you this thing with another
- 44:40:54example, right? Which is I created an
- 44:40:57image of an RGB image of 256 cross 256
- 44:41:01pixels, right? Which have all zeros
- 44:41:03inside them. And this is how it was
- 44:41:06created, right? Three layers RGB 256
- 44:41:10256. So this is how the image will look
- 44:41:12like
- 44:41:14right.
- 44:41:17This is how it will look like.
- 44:41:20Now you can also create
- 44:41:27arrays with ones. Right? Exactly the
- 44:41:31same way you created it with zeros. I'll
- 44:41:33give it to you. I'll give you 5 minutes
- 44:41:34time to create them. I'll show you one.
- 44:41:36So I say
- 44:41:39a1 is equal to np dot
- 44:41:44once
- 44:41:46and inside I pass
- 44:41:49two
- 44:41:52I check a1. So this is a array like this
- 44:41:55right? It is an array like this. Now you
- 44:41:58create create two dimension
- 44:42:02and
- 44:42:04three dimension
- 44:42:06arrays of one. It is basically
- 44:42:11as a float. Hurry. It's represented as a
- 44:42:13float. Right. It's represented as a
- 44:42:15float. Okay.
- 44:42:18Right.
- 44:42:20Perfect. Right. Also guys uh with this
- 44:42:24right also with this you can create the
- 44:42:27custom arrays. Right. You can create the
- 44:42:29custom arrays. Right. How do we create
- 44:42:32custom arrays people? How do we create
- 44:42:34the custom arrays?
- 44:42:37You have created now zeros. You have
- 44:42:39created now ones. Now what is left that
- 44:42:43you create the custom arrays. Uh forget
- 44:42:47about this. I will come to this later
- 44:42:48on.
- 44:42:50Right? Let's create
- 44:42:53custom arrays. Right? So the syntax
- 44:42:56remains the same. Right? I'll say cus
- 44:42:59arr is equal to np.
- 44:43:04Right? This is the syntax np.
- 44:43:07Right? And you will say 6 + 6 and
- 44:43:10suppose you want an array of all fours.
- 44:43:13Right? This is the dimension 6 + 6. And
- 44:43:16this value after comma is basically the
- 44:43:18value which you want. You execute this
- 44:43:21and you copy this paste this and you
- 44:43:24will get the arrays of fours for
- 44:43:27yourself. Right? If you want of 10, you
- 44:43:30will get of 10. If you want 10.3,
- 44:43:34you will get 10.3. Right? anything which
- 44:43:37you want. If you want case, you will get
- 44:43:40case, right? All the examples. So, let
- 44:43:43me just show that to you.
- 44:43:55Yep. Like this. Now guys, how did we
- 44:43:58create
- 44:44:01or how did we use
- 44:44:04range in Python?
- 44:44:07Can you use
- 44:44:09range to generate
- 44:44:13numbers between 20 to 50,
- 44:44:19right? 20 to 50. Can you give me the
- 44:44:22syntax quickly? How did you do that in
- 44:44:24range?
- 44:44:26How do we do that? We said R is equal to
- 44:44:30range
- 44:44:3220 to 51. Right? And then I said
- 44:44:39I in R
- 44:44:41print I,
- 44:44:44right? And this is how I got the
- 44:44:45numbers, right? So similar
- 44:44:49to range in Python,
- 44:44:52we have
- 44:44:54a range in num py. Right? How do we use
- 44:44:58a range? I say
- 44:45:01uh a range
- 44:45:05ar r is equal to np dot arange. Right?
- 44:45:10And then same syntax I will say 20 to
- 44:45:1351. Right? 20 to 51. And that's it. And
- 44:45:17when I will check my AR range error, you
- 44:45:21will see I have generated myself numbers
- 44:45:23between 20 to 50 and a range in numpy.
- 44:45:28Right? A range in numpy. Yes, if you
- 44:45:32want a interval so you can use this say
- 44:45:37a range one and after comma you pass the
- 44:45:40third argument. Suppose it's three. So
- 44:45:42now it will jump three times, right? 20
- 44:45:4623 26 29 32 35 like this up till 50.
- 44:45:52Now guys there is something which is
- 44:45:54called as lind space
- 44:46:00right. What is lindspace?
- 44:46:02It stands for
- 44:46:06linear spacing
- 44:46:08which means
- 44:46:10between two given numbers.
- 44:46:15This function will fit the required
- 44:46:21number of
- 44:46:23numbers. Right? For example, suppose for
- 44:46:27example,
- 44:46:29we need to create an interval
- 44:46:36from 0 to 1. People, there are infinite
- 44:46:41numbers I can have between 0 to 1. Isn't
- 44:46:44it?
- 44:46:45Infinite numbers I can have between 0 to
- 44:46:481. 0.0000000000001
- 44:46:520 0 1 0 1 01 right I can go in the
- 44:46:56infinite manner right now for example
- 44:47:00you need to create numbers between 0 to
- 44:47:0410 right and you want to create and want
- 44:47:08to have
- 44:47:1010 numbers in it right so how will you
- 44:47:13do this it's not float it's about the
- 44:47:16number theory right between 0 and one
- 44:47:18you have infinite finite numbers, right?
- 44:47:20So you say lindspace is equal to np dot
- 44:47:25lindspace, right? np.tlind space. You
- 44:47:28mention from 0 to 10, you want to have
- 44:47:3110 numbers, right? And when you will
- 44:47:34create lindspace,
- 44:47:35you will see that these are the numbers
- 44:47:38are there which have been created,
- 44:47:40right? These are the numbers which have
- 44:47:42been created,
- 44:47:44right? Nine numbers. Now I say 100
- 44:47:47numbers. I want evenly spaced 100
- 44:47:50numbers, right? Evenly spaced 100
- 44:47:53numbers. How are they even? You can
- 44:47:55simply subtract one number from another
- 44:47:57and the difference for all the numbers
- 44:47:58will be exactly the same. 0.01 0 1 01.
- 44:48:03Subtract any two numbers. It will be
- 44:48:040.01 0 1 01. Right? Where do we need
- 44:48:07this? We need this to plot the axises.
- 44:48:10Right? When you will plot graphs, you
- 44:48:12will need access between this interval.
- 44:48:14You need five values that works like
- 44:48:16this. Okay? Suppose you want from 0 to
- 44:48:1910 five different values, right? You
- 44:48:21will have five different values like
- 44:48:23this, right? Between the gap of 0.5,
- 44:48:27right? If I say 1 to 10, you will have
- 44:48:30values like this,
- 44:48:33right? Like this. Clear? 100 values like
- 44:48:36this. Yeah. Clear guys. How do we use
- 44:48:38lin space? Suppose you want from 0 to
- 44:48:4210. Interval from 0 to 10 and 10 will be
- 44:48:44included. Zero will not be included.
- 44:48:47Right? We'll start from one. So I go to
- 44:48:50one it will be from
- 44:48:55one. Why is 0 not included then? Yeah. 0
- 44:48:59is included. Right? 0 is also included
- 44:49:01and 10 is also included. 100 numbers
- 44:49:04between them. Right?
- 44:49:06Clear? This is what lin space is. Now
- 44:49:09guys, now
- 44:49:12suppose you
- 44:49:15lohan l space is basically used if you
- 44:49:20want to create n numbers between the
- 44:49:23range of numbers right between 0 to 1
- 44:49:27right suppose between 0 to 1 you are
- 44:49:30trying to plot a graph okay and your
- 44:49:32values are 0.2 0.3 0.6 six right and you
- 44:49:37want to draw a graph so you will have to
- 44:49:39mark the axis right the x axis and the
- 44:49:41y- axis so you can use lindspace there
- 44:49:44and what will it do it will take the
- 44:49:46range it will take the interval in
- 44:49:48between you want to add the equal space
- 44:49:51numbers and then the third argument here
- 44:49:54will be that how many numbers do you
- 44:49:57want between them so this syntax tells
- 44:50:00you that
- 44:50:02from
- 44:50:040 to 1
- 44:50:06give me 100 numbers. How are these
- 44:50:09numbers? Equally
- 44:50:12spaced
- 44:50:13numbers. Equally spaced numbers, right?
- 44:50:16So when you will execute this, you will
- 44:50:19see that all there are 100 numbers which
- 44:50:21have been generated, right? 100 numbers.
- 44:50:23And all the numbers are equidistant from
- 44:50:26each other because difference of every
- 44:50:28single number from the next number is
- 44:50:300.01
- 44:50:3201.
- 44:50:38Yep, that's what it does. Right
- 44:50:43now guys, now suppose we want to
- 44:50:48generate
- 44:50:50random numbers, right? We want to
- 44:50:52generate random numbers, right? Now
- 44:50:54we're interested in generating random
- 44:50:55numbers. So we have something called as
- 44:51:00random
- 44:51:02dot random right. What will it do?
- 44:51:05Random.random
- 44:51:06will generate
- 44:51:10random
- 44:51:11float numbers
- 44:51:15between 0 to 1. Random float numbers
- 44:51:18between 0 to 1. How will this happen?
- 44:51:21You will say
- 44:51:23rand rand is equal to np do. random dot
- 44:51:29random and inside you will mention what
- 44:51:32is the dimension that you seek. Suppose
- 44:51:34I want 6 + 6. So when you will check
- 44:51:38this you will have all numbers for 6 + 6
- 44:51:42dimension right 6 + 6 matrix right now
- 44:51:47every time you rerun this the numbers
- 44:51:48will change because all of these are
- 44:51:50random numbers
- 44:51:52right all of these are random numbers
- 44:51:56now I say 100 multiplied by rand rand
- 44:52:02you will see all of them all these
- 44:52:04numbers will be multiplied by 00 right
- 44:52:07all of them
- 44:52:09in one shopping.
- 44:52:12Now just like this we can also create
- 44:52:17random integers right how will we create
- 44:52:20random integers guys
- 44:52:23I say rand intore
- 44:52:26rand is equal to np dot random dot rand
- 44:52:33right and here you will specify that
- 44:52:37what is the range of numbers you want
- 44:52:39from so I say between 20 to 25 I need
- 44:52:44random numbers and then I want it from
- 44:52:48in a 3 + 3 format. Right? And now when
- 44:52:51you will check your random you will get
- 44:52:54random numbers generated like this.
- 44:52:56Okay? Random numbers generated like
- 44:52:58this.
- 44:53:04Right?
- 44:53:06If you say 3 + 3 + 3 you will get a 3 +
- 44:53:103 + 3 matrix. Even if you will only say
- 44:53:123, you will get a 1D.
- 44:53:17Right guys? You can change the
- 44:53:19dimension. So this is the range from
- 44:53:22which you want to choose the random
- 44:53:24numbers and this is the dimension you
- 44:53:26want this matrix or vector to be in.
- 44:53:30Now guys, we'll move on to the next part
- 44:53:33which is basically properties and again
- 44:53:37there are a lot of operations you'll
- 44:53:38have to see it yourself right
- 44:53:42properties and
- 44:53:46attributes
- 44:53:49of numpy
- 44:53:53arrays right property and attributes of
- 44:53:56numpy arrays.
- 44:53:59Okay. Now guys, the first one in this
- 44:54:02scheme of things is shape of array,
- 44:54:06right? Shape of array. What is shape of
- 44:54:08array? It tells you
- 44:54:12the
- 44:54:14dimensions of the array
- 44:54:19stored in a
- 44:54:23tle. Right?
- 44:54:25For example, I say a ar r r r r r r r r
- 44:54:29r r r r r r r r r r r r0
- 44:54:30right a r r r r r r r r r r r r r r r r
- 44:54:32r r r r r 1 a ar a ar a ar a ar a ar a
- 44:54:34ar a ar a ar a ar a ar a r r r r r r r r
- 44:54:34r r r r r r r r r r r r r 2 a ar r r3
- 44:54:38right and then I say
- 44:54:44print this
- 44:54:47dot shape
- 44:54:50right like this and you will see that it
- 44:54:52will give you the shape of each array
- 44:54:55Right? 0D, 1D, 2D and 3D. Right people?
- 44:55:00Shape of the array.
- 44:55:04Please try it out. We have used it one
- 44:55:07or two times. But this is what shape of
- 44:55:09array actually means.
- 44:55:11Second is people
- 44:55:14end
- 44:55:16right is end
- 44:55:20right end of array right it tells you
- 44:55:26the rank of the array whether it's one
- 44:55:30dimensional two dimensional zero
- 44:55:31dimensional threedimensional four
- 44:55:32dimensional
- 44:55:34so again I will do the same and
- 44:55:46right I'll say end and you will see it
- 44:55:48will give you 0 1 2 3 zero dimension
- 44:55:51zero rank one rank two rank and three
- 44:55:53rank and it can go all the way up to end
- 44:55:56rank
- 44:55:58before this I should have also
- 44:56:02printed these arrays
- 44:56:09Right. These are the arrays.
- 44:56:22Yep.
- 44:56:24These are the arrays which we have and
- 44:56:26these are the subsequent things, right?
- 44:56:30Rank copy array. Then guys, the third
- 44:56:32thing is the size of
- 44:56:37array, right? It tells you
- 44:56:42the number of elements inside. How many
- 44:56:47elements do we have inside this array in
- 44:56:49a ar r1? How many elements do we have? 1
- 44:56:522 3 4 5 6 7. How many elements do we
- 44:56:56have in this 2 + 2 ar2? 1 2 3 4 5 6 7 8
- 44:57:009 which is rows multiplied by columns 3
- 44:57:02* 3 right and how many elements do we
- 44:57:05have in this 3D which is 2 * 3 * 3 which
- 44:57:10is 18 right so now when you copy this
- 44:57:14right you can use this
- 44:57:17and say
- 44:57:20size right it will say 17 918
- 44:57:48Yep.
- 44:57:57Fourth is people.
- 44:58:00The D type
- 44:58:03of array tells you the data type of the
- 44:58:10array. Right? And I've told you we place
- 44:58:14only
- 44:58:16homogeneous
- 44:58:18data in array right what will happen if
- 44:58:22we don't do this I will show that to you
- 44:58:24also right so we do this right and we
- 44:58:28say
- 44:58:32retype
- 44:58:34and you will see in 64 all of them are
- 44:58:37integers right all of them are integers
- 44:58:40that's the reason we are getting in 64
- 44:58:42right suppose I create a new array a ar
- 44:58:45r new right let me say head
- 44:58:50right hetro
- 44:58:53heterogenous
- 44:58:54and I say it is like n dot array
- 44:59:01let's say like this
- 44:59:07right like this now people when you will
- 44:59:10Check
- 44:59:12ar r
- 44:59:16dot d type you will see it will give you
- 44:59:18float just because of one floating point
- 44:59:21number inside this entire array it gets
- 44:59:24converted to float right between all the
- 44:59:27integers if you put one float then it
- 44:59:29will be taking float directly right now
- 44:59:34let me just copy this and let's say
- 44:59:36heterogenous one and let me add another
- 44:59:38value which is string and say rather
- 44:59:42right and when you will execute this it
- 44:59:44will give you U32 U32 here is
- 44:59:46representing strings right it is all
- 44:59:48called as objects right these are all
- 44:59:50string values right so precedences
- 44:59:53strings greatest then float and then
- 44:59:56your uh integers right if you place the
- 45:00:00heterogenous data inside the numpy array
- 45:00:04right you only need to put homogeneous
- 45:00:06data in the array
- 45:00:08now fifth is the item size
- 45:00:12of array. Right? What is item size?
- 45:00:17It gives you
- 45:00:20the bite
- 45:00:22occupied
- 45:00:24by each element of an array. Right?
- 45:00:29Because we assume that elements will be
- 45:00:31homogeneous. It will give you the bite
- 45:00:34occupied by each element of the array.
- 45:00:37Only one element. Okay? So how will it
- 45:00:39happen? So let's say a ar r r0
- 45:00:43or let's say ar r r1 dot item size
- 45:00:48right and you will get eight right. So
- 45:00:51why eight? Because
- 45:00:55because each data point right each data
- 45:00:58point is occupying
- 45:01:01the result is 8 bytes. Let me put this
- 45:01:07here.
- 45:01:13Right. Eight bytes
- 45:01:15because each element in ARR1 is
- 45:01:22occupying
- 45:01:2464 bits which are
- 45:01:30equivalent to which are equivalent to 8
- 45:01:34bytes. Right? one bite is equal to 8
- 45:01:37bits. So 64 bits will be equal to 8
- 45:01:41bytes. Right? That is how it is giving
- 45:01:43you the result. Now if you are
- 45:01:45interested in knowing the entire bytes
- 45:01:49right entire bytes then you say n bytes
- 45:01:56will give you the
- 45:02:00total bytes
- 45:02:02occupied
- 45:02:04by the elements of the array. Right? You
- 45:02:08say print
- 45:02:11a ar r1 dot
- 45:02:13n bytes right and write
- 45:02:18bytes it will be 56 bytes
- 45:02:24right why because how many elements do
- 45:02:26we have in our ar r1 1 2 3 4 5 6 7 right
- 45:02:327 8 are 56 right 56 total bytes are
- 45:02:36being occupied with by ar r1 one. Now
- 45:02:39guys, the seventh one
- 45:02:42is
- 45:02:44as type right
- 45:02:48in array.
- 45:02:51This will help you change the data type
- 45:02:57of the array. Right? Change the data
- 45:03:00type of the array. Suppose I have a arr
- 45:03:04type which is uh so I'll say print
- 45:03:10d type right this is end 64 right and
- 45:03:13now what I do is I say print
- 45:03:20uh wait let me give you a structured way
- 45:03:22print a r1 right let me say
- 45:03:28array
- 45:03:33R1
- 45:03:39right D type of array.
- 45:03:43Now
- 45:03:46a ar r2 sorry a ar a ar a ar a ar a ar a
- 45:03:47ar a ar a ar a ar a ar a r r r r r r r r
- 45:03:48r r r r r r r r r r r r r1 is equal to a
- 45:03:50ar r1
- 45:03:52dot as type right dot as type and let's
- 45:03:55say I want to convert this in np dot
- 45:04:00right np dot
- 45:04:02uh
- 45:04:04int 32 right I want to create convert
- 45:04:07this in uh ar r r int 32 right when I do
- 45:04:11this and now when I will copy these same
- 45:04:15things you will see for yourself. Right?
- 45:04:18Now the D type was int 64 and now the DT
- 45:04:20type is int 32. Right?
- 45:04:24If I want I can do this conversion in
- 45:04:27float also
- 45:04:30float 64. Right? And then I will just
- 45:04:33copy this
- 45:04:35and I will paste it here.
- 45:04:39Right? And now you will see now the
- 45:04:42floating point has been activated.
- 45:04:45Right? It has been now activated. We can
- 45:04:48go till int. We can go till int 8.
- 45:04:54Right?
- 45:04:56I can go to 16.
- 45:04:59Right? And I can do this.
- 45:05:04I can go to int
- 45:05:078 also.
- 45:05:12Yeah, like this in date also 64 32.
- 45:05:18So first one has to be
- 45:05:22Yeah.
- 45:05:2832
- 45:05:38whatever right like this okay you can
- 45:05:42convert this
- 45:05:45also people this is later on conversion
- 45:05:48you can define the data type of the
- 45:05:54array while creation time also. How
- 45:05:59would you do that? Suppose you are
- 45:06:00creating a ar r11 and you say np dot
- 45:06:04array right and suppose you take it from
- 45:06:06a list and then you just put a comma and
- 45:06:09say d type. So what will be the default
- 45:06:12data type here people? If I just do this
- 45:06:14if I just execute this what will be the
- 45:06:17default data type? Int 64 is the default
- 45:06:21isn't it? But now suppose I want to
- 45:06:24change it right here. I say data type is
- 45:06:26equal to float 32 right sorry float 64
- 45:06:36np dot
- 45:06:38sorry my bad
- 45:06:42float 64 right and I say a ar r r11 and
- 45:06:46this will be float 64 if you want float
- 45:06:4932 it will also become float 32 right
- 45:06:52right here while you define
- 45:06:55Instead of using as type, you can do it
- 45:06:57right here. Right? These things will
- 45:06:58come in very handy people because you
- 45:07:00will have to save memory because when
- 45:07:02your data becomes very very big, you
- 45:07:03will be always in a crunch for memory
- 45:07:06like this. You want integers, then
- 45:07:08integers will be like this
- 45:07:11random.randent
- 45:07:12like this. Suppose you want to generate
- 45:07:160 to six, right? And suppose you want to
- 45:07:19generate
- 45:07:21say 100 numbers like this 0 to 6 the
- 45:07:27scores
- 45:07:281 to six like this randomly
- 45:07:36right suppose you want to generate 100
- 45:07:39scores for five different batsmen
- 45:07:42randomly it will be like this batsman
- 45:07:45number one batsman number to bat number
- 45:07:48three, fourth and fifth. Right.
- 45:07:54Yep.
- 45:07:59Understand the data guys. Now it's the
- 45:08:01time to understand the data.
- 45:08:03Yep. Now guys, we have methods in numpy
- 45:08:10arrays, right? Methods in numpy arrays.
- 45:08:13So what are these methods? Now the first
- 45:08:16method we have to learn is called as
- 45:08:19reshape right. Reshape.
- 45:08:29Yeah. So reshape is you can use this to
- 45:08:31create
- 45:08:34you can use this to create
- 45:08:37a new shape of the array. Very very
- 45:08:41powerful guys. Very powerful. One of the
- 45:08:43most powerful methods in numpy is uh the
- 45:08:47reshape right and how do we use reshape
- 45:08:50suppose
- 45:08:54we have a 1D array of 20 elements right
- 45:09:01now to reshape this
- 45:09:05reshape it we need to find the factors
- 45:09:12Right. Factors of 20. They are what? 1
- 45:09:1720 4 5
- 45:09:212 10.
- 45:09:23Right. The other factors.
- 45:09:27The other factors. Now see what will I
- 45:09:29do. Right. Now see what will I do. Let
- 45:09:31me create a say random array. Right?
- 45:09:34Random array. I say random
- 45:09:39arr is equal to entprandom
- 45:09:43dot rand right and let me say I want to
- 45:09:47create it from 1 to 50 right and I want
- 45:09:52it to be having 20 elements right so I
- 45:09:57say random arrand
- 45:10:00values inside this right random 20
- 45:10:02values now see Now reshape
- 45:10:08first
- 45:10:11I will reshape in 1 + 20 right 1A 20 how
- 45:10:18will I do that you just have to write
- 45:10:22uh
- 45:10:26print
- 45:10:28a ar r r sorry sorry random dot ar r
- 45:10:31random ar
- 45:10:33dot reshape dot reshape and you just
- 45:10:37pass in the dimension I say 1 20
- 45:10:41right and when you will do this you will
- 45:10:43see it is coming now in 1 20 format
- 45:10:47right so let me just also write print
- 45:11:02right 2D
- 45:11:041A 20
- 45:11:07right and now what I'll do is I'll copy
- 45:11:10this and I will paste this and say 20
- 45:11:15comma 1 right you will see it will be
- 45:11:18like this 20 comma 1 immediately with
- 45:11:21reshape right let me copy this
- 45:11:25let me say 2D I'm still at 2D let me say
- 45:11:292 10 right and you to see this is 2A 10.
- 45:11:35Now I can reshape it in
- 45:11:3910 2
- 45:11:43right 10 2 right then I can reshape the
- 45:11:48same thing
- 45:11:52in
- 45:11:544A 5 and I can reshape this in 5A 4
- 45:11:58right like this guys are you able to see
- 45:12:01the power one dimension I'm able to
- 45:12:04create two dimensions
- 45:12:05And now I will take it a step further
- 45:12:08and I will write it in three dimensions.
- 45:12:11Right? How will I write it in three
- 45:12:12dimension? Let me say this 1 comma
- 45:12:172a 10. Right? This is will also be three
- 45:12:20dimensions. Let's sorry 2a 2a 5.
- 45:12:24Let me say this. And now you will see I
- 45:12:26can have this in three dimensions.
- 45:12:29Right?
- 45:12:32Yep. I can also say in three dimensions
- 45:12:36like this. I want to have five layers
- 45:12:40with two rows and two columns. Right? So
- 45:12:43you will have it like this also. Right?
- 45:12:46So this is how people we can reshape the
- 45:12:49array. Very powerful. Very very
- 45:12:51powerful.
- 45:12:53Right? Very very powerful.
- 45:12:57Right? And you can take this
- 45:13:04and save print
- 45:13:09random dot this and you can say print
- 45:13:14dot shape.
- 45:13:16Right? So this was the first one.
- 45:13:27Yep. like this.
- 45:13:29Also people if you want to visualize we
- 45:13:32can also go this route.
- 45:13:35We can have 10 comma 2 comma 1, right?
- 45:13:40It will look like this, right? 10
- 45:13:43layers. 10 layers you can have,
- 45:13:47right? 10 layers you can have.
- 45:13:52Great. Now, second method which we have
- 45:13:54to learn is called as
- 45:13:58transpose,
- 45:14:00right? Transpose method. Right? What
- 45:14:03does that do? It interchanges the
- 45:14:07dimensions
- 45:14:10like rows and columns, right? Yes.
- 45:14:15Absolutely. Right. Suppose you have a
- 45:14:18matrix, right? Which is
- 45:14:2222, 33, 44, 55, 66, 77, right? This is
- 45:14:29a. So now when you will a transpose it,
- 45:14:32right? The dimension right now the shape
- 45:14:35right now is 3A 2. Now this will become
- 45:14:392a 3. And how this will happen? Rows
- 45:14:42will now become columns and columns will
- 45:14:44now become rows. Right? So let's make
- 45:14:47column the rows. Right? Sorry columns
- 45:14:49the rows. It will be 22 44 66
- 45:14:5533 55 77. Right? So people in transpose
- 45:15:00no information is lost. It is just a
- 45:15:03change in the view right which is
- 45:15:07happening right?
- 45:15:1022 44 66 33 55 77.
- 45:15:19Why do we need transposition? Suppose we
- 45:15:23have two matrix.
- 45:15:25This is 11th class mathematics. Right?
- 45:15:28one has
- 45:15:30a dimension of n cross m and the second
- 45:15:33has dimension of a cross b. If you
- 45:15:40want to multiply
- 45:15:46these two matrix say
- 45:15:50M_sub_1 and M_sub_2,
- 45:15:54right? There needs to be a satisfaction
- 45:15:56of condition. M should be equal to A.
- 45:16:00Right? M should be equal to A. Right? M
- 45:16:04should be equal to A. This should be
- 45:16:05equal to this and the resultant vector
- 45:16:08the resultant matrix which you will get
- 45:16:10will be of n crossb dimension right. So
- 45:16:14often times suppose this is n cross m
- 45:16:17this is n cross m right m cross n and
- 45:16:21you know that m is equal to a. So what
- 45:16:24will you do? You will transpose this
- 45:16:26matrix right? You will transpose this
- 45:16:28matrix then it will become n cross m and
- 45:16:31then m can be equivalent to a. Right?
- 45:16:34For example, what I'm saying, we have
- 45:16:36one matrix which is 2 + 3 and this
- 45:16:38matrix is 2 + 5, right? So, can you
- 45:16:42multiply these matrix people? Is 3 equal
- 45:16:45to 2? The answer is no. Right? The
- 45:16:48answer is no. So, what will you do? You
- 45:16:50will just transpose this and this will
- 45:16:52become 3 + 2 and this is 2 + 5. And now
- 45:16:57you can multiply this and the resultant
- 45:16:59will become 3 + 5 matrix. Right? So for
- 45:17:04operations like these we need
- 45:17:06transposition. Right? So how do we
- 45:17:08transpose it?
- 45:17:11How do we transpose it? So let's say
- 45:17:14again
- 45:17:15uh a ar r2 right this is a ar r2 and I
- 45:17:20want to transpose it. So I say a r2 t is
- 45:17:24equal to uh np.transpose transpose
- 45:17:29sorry a r2 dot
- 45:17:33transpose
- 45:17:35right
- 45:17:37and now when you will see ar r2 ts you
- 45:17:41will see rows and columns have
- 45:17:43interchanged right rows and columns have
- 45:17:46interchanged with each other
- 45:17:52or let me give you one more example
- 45:17:56Uh if this is not clear, let me pick up
- 45:18:02this again
- 45:18:14right now. I say
- 45:18:17dot reshape
- 45:18:20into say
- 45:18:232 + 10. Right? So this is 2 + 10. And
- 45:18:26now when you want to transpose this so
- 45:18:28I'll say this t is equal to this dot
- 45:18:34transpose
- 45:18:38t or a this and we can check this now
- 45:18:44and it will be this. Sorry guys. So this
- 45:18:48is going to be it will be like this
- 45:18:51right transposed.
- 45:18:53And now if you want to see this,
- 45:19:03this was the original shape, right? Rows
- 45:19:05and columns have now been interchanged,
- 45:19:08transposed with each other.
- 45:19:12Now guys, the third method,
- 45:19:16the third method which is there with us
- 45:19:18is called as flatten, right? It is
- 45:19:21called as flatten.
- 45:19:25Right? What does flatten do? It reduces
- 45:19:29the dimension to one dimension. Right?
- 45:19:33Any dimension you have, it reduces it to
- 45:19:36one dimension. For example, I have this,
- 45:19:40right? And now when I say
- 45:19:44this
- 45:19:46dot_f,
- 45:19:48this will become this dot platin.
- 45:19:53And when you will check this up, you
- 45:19:56will see that it has now become one
- 45:19:57dimension. No matter how many dimensions
- 45:19:59you have, it will become one dimension.
- 45:20:03Right? Let me take this again to show
- 45:20:05you one more example.
- 45:20:11Right?
- 45:20:13And here I say dot reshape into
- 45:20:18uh
- 45:20:21uh 5 + 2 + 2 right I do this my random
- 45:20:27ar r is this right it has five layers
- 45:20:30two rows and two columns right so now I
- 45:20:34say this
- 45:20:36flatten is equal to this dot flatten
- 45:20:41and If you will check it now again, you
- 45:20:44will see it has now flattened it out.
- 45:20:46Yep.
- 45:20:49Now why do we need this? We need this
- 45:20:51for a lot of statistical operations. We
- 45:20:53need this to feed the data into the
- 45:20:56algorithms. Right? As we will move
- 45:20:59forward, you will understand the use of
- 45:21:00flattening.
- 45:21:02Right? Now guys, moving on and uh as
- 45:21:06discussed, let me now show you
- 45:21:10the power of
- 45:21:13numpy,
- 45:21:16right? Numpy
- 45:21:18over
- 45:21:20lists and other data types, right? I
- 45:21:25will not take a lot of examples. Just a
- 45:21:27second, guys.
- 45:21:29Yeah. Okay.
- 45:21:32Now guys, I told you that
- 45:21:38less
- 45:21:41take up
- 45:21:44much more memory
- 45:21:48as compared
- 45:21:50to numpy arrays. Right? And I'm going to
- 45:21:53prove this to you. Right? Now let me use
- 45:21:58let me create a random sequence of
- 45:22:00random numbers using range in Python
- 45:22:04right and say I create range of 10,000
- 45:22:08numbers right range of 10,000 numbers so
- 45:22:11what will this give me this will give me
- 45:22:13numbers from 0 to 99999 right continuous
- 45:22:16numbers right so this is range I will
- 45:22:20use
- 45:22:22a range in
- 45:22:24numpy Y to create
- 45:22:28a similar
- 45:22:32series of numbers
- 45:22:35right so let's say array is equal to np
- 45:22:39dot arange
- 45:22:42right same thing same done by both right
- 45:22:45I've shown you above also now let me
- 45:22:48import sis
- 45:22:51library right sis package and I will use
- 45:22:55something called as get size of right
- 45:22:58get size of. What does this do? Get size
- 45:23:01of it's a method
- 45:23:05which calculates
- 45:23:07the bytes
- 45:23:10occupied
- 45:23:12by a single
- 45:23:15element in
- 45:23:18vanilla Python. What is vanilla Python?
- 45:23:20It is the traditional Python,
- 45:23:22right? vanilla Python.
- 45:23:25So let me just show that to you. I'll
- 45:23:26use this and I will say print. Now guys,
- 45:23:30if I get
- 45:23:33size of any random number from this
- 45:23:36range, right? Any random number. Say I
- 45:23:39get size of five, right? And I then
- 45:23:43multiply that byes with the length of
- 45:23:46random, right? With the length of random
- 45:23:50this rand, right?
- 45:23:53Right? With the length of random, do you
- 45:23:55think I will get the bytes for the
- 45:23:58entire
- 45:23:59data structure? What am I saying is
- 45:24:02suppose
- 45:24:04uh I used range
- 45:24:08five. So what will this give me? 0 1 2 3
- 45:24:124. Right? This will be the output. So
- 45:24:15now I say get
- 45:24:18size of say I say two. Right? So suppose
- 45:24:212 is x and then I multiply this with the
- 45:24:25length of this series which is five. So
- 45:24:28do you think I will get 5x which will
- 45:24:31represent the number of bytes occupied
- 45:24:33by the entire data type. Anything
- 45:24:35randomly any random number this can be
- 45:24:38three right? Why not hurry?
- 45:24:42Why not?
- 45:24:47All of these are integers. So integers
- 45:24:50all of 64 bits assuming. So if you
- 45:24:54calculate the side of size of this and
- 45:24:56if you multiply with the total number of
- 45:24:57numbers you will get the total size
- 45:25:00isn't it?
- 45:25:04Huh? Index is in
- 45:25:10no no it's not about that it's about the
- 45:25:12element right homogeneous elements
- 45:25:13inside this.
- 45:25:16I am saying when you use range five what
- 45:25:19is going to be the output? 0 1 2 3 4
- 45:25:22right now all these are elements
- 45:25:27elements of range.
- 45:25:31Right? All of them are elements of
- 45:25:32range. Right? Now I'm saying if I fetch
- 45:25:36the size of one element and multiply it
- 45:25:41with the length of the entire range,
- 45:25:45will I get the bytes occupied by the
- 45:25:48entire range? For example, if I do this,
- 45:25:52right? If I do this,
- 45:25:56this is 28, right? 28 bytes people. 28
- 45:26:00bytes
- 45:26:02bytes are occupied
- 45:26:06by one element of range right
- 45:26:13one element of
- 45:26:16range right now if I just I'm saying I'm
- 45:26:20just saying if I multiply to find how
- 45:26:24many numbers range has
- 45:26:27how many numbers
- 45:26:29range has
- 45:26:31equal to 10,000
- 45:26:35right so total
- 45:26:39memory occupied
- 45:26:41will be will be how much it will be 28
- 45:26:48ult*lied by 10,000 which will be equal
- 45:26:51to 28,000
- 45:26:53yeah and how will you find this you will
- 45:26:55say Print
- 45:27:00this multiplied by length of RAM,
- 45:27:07right? 28,000 bytes. Clear? Now, yes.
- 45:27:11Now, this is for the range. Now, let me
- 45:27:13use another thing. So, how will you
- 45:27:15calculate the length of this array? What
- 45:27:18what property and attribute will you use
- 45:27:21people?
- 45:27:22N bytes, right? N bytes will give you
- 45:27:26total bytes occupied by the elements of
- 45:27:27the array. Right? We will use n bytes
- 45:27:29here. So I come back down and I say
- 45:27:33using n bytes for arrays. Right? And you
- 45:27:38will see what the result comes. Print
- 45:27:43uh array dot n bytes. Right? And I say
- 45:27:50bytes. Are you ready to see the result?
- 45:27:52Do you see what has happened?
- 45:27:55How many bytes this was taking? It was
- 45:27:57taking 28,000 bytes. How many bytes this
- 45:28:00is taking? This is taking 80,000 bytes.
- 45:28:05This was taking 2 lakh 80,000. This is
- 45:28:07taking 80,000. Two lakh extra bytes of
- 45:28:10memory is taken by range.
- 45:28:15Ran is range.
- 45:28:24Yep. And if I just go to million
- 45:28:28numbers,
- 45:28:29right? Million numbers in both.
- 45:28:34See the difference it becomes,
- 45:28:37right? This is now 3 three 28 million
- 45:28:42bytes it is taking and it is taking 8
- 45:28:46million bytes. 20 million extra bytes
- 45:28:49are occupied right now people do you
- 45:28:53believe me? Yeah, that numpy wy are way
- 45:28:56more efficient in memory management as
- 45:28:59compared to the traditional data types
- 45:29:00of Python. Yes. Okay, that's the first
- 45:29:03part. Now second is people performance,
- 45:29:07right? Performance. So what I'm going to
- 45:29:09do is what I'm going to do is I am going
- 45:29:12to
- 45:29:14import
- 45:29:16time, right? It's a it's a module in
- 45:29:19Python, right? Suppose I say x is equal
- 45:29:22to range
- 45:29:26this much right. Okay. And then I have y
- 45:29:31is equal to range say
- 45:29:37this
- 45:29:38to
- 45:29:40this. Right? Both of them will have
- 45:29:43equal amount of numbers. Same numbers
- 45:29:45both of them will have. Right?
- 45:29:47This will have say
- 45:29:50uh 1 2 3 1 2 3 10 million values. 10
- 45:29:56million values. This will also have 10
- 45:29:59million values,
- 45:30:03right? Both of them will have 10 million
- 45:30:04values. Now what I'm trying to do is I
- 45:30:06want to add them up right by bit by bit.
- 45:30:09I want to add them up right. I want to
- 45:30:11add first element of this to first
- 45:30:13element of this. second of this to
- 45:30:15second of this, third of this to third
- 45:30:16of this like this. Okay, I want to do
- 45:30:18this. Now what I'll do is I will run a
- 45:30:21counter, right? I will run a counter
- 45:30:24which is the start time,
- 45:30:28right? And this is given by time dot
- 45:30:31time, right? Which will give you the
- 45:30:35this will give you the
- 45:30:40current time, right? After this I will
- 45:30:43run the operation. I will say C is equal
- 45:30:45to X + Y
- 45:30:49for X Y
- 45:30:51in zip
- 45:30:54X Y right in zip X Y right add X + Y bit
- 45:31:02by bit element by element for X and Y in
- 45:31:05zip zip is a function right which allows
- 45:31:07you to do this operation sequentially
- 45:31:10right sequentially right add the
- 45:31:15elements of X and Y
- 45:31:20element by element right element by
- 45:31:24element
- 45:31:28right element by element right and then
- 45:31:31I'm going to print right so start time
- 45:31:34will start and now I will say time dot
- 45:31:37time which is now the end time minus
- 45:31:40start time so this will give me the Time
- 45:31:43taken for execution isn't it guys
- 45:31:47will give me
- 45:31:49delta of time which is equal to time
- 45:31:53taken for operation
- 45:31:56seconds
- 45:31:58right these many seconds will be taken
- 45:32:01right so let me run this and it takes
- 45:32:04around say
- 45:32:074.3 seconds right to do this right 4.3
- 45:32:11seconds now Guys, see what happens. You
- 45:32:14had to write this complex syntax in the
- 45:32:18traditional Python. Now let me show this
- 45:32:20on arrays. Right? What will happen on
- 45:32:23arrays? I will say a is equal to np dot
- 45:32:27a range.
- 45:32:30Right? And inside a range I will pass
- 45:32:33the same values what I have taken above.
- 45:32:36Right? And I will say b is equal to np
- 45:32:39dot
- 45:32:43a range and I will pass the same values
- 45:32:46inside
- 45:32:48right exactly the same now what I'll do
- 45:32:51is I will say same thing
- 45:32:58right just I will change the execution
- 45:33:01of C will now simply become people A + B
- 45:33:06what is simple this or this
- 45:33:09this or this
- 45:33:13two right do you see the power if not I
- 45:33:15will show this to you again right later
- 45:33:17on and let me run this and you see the
- 45:33:21difference now let me just increase a
- 45:33:23couple of zeros right a couple of zeros
- 45:33:28two zeros I'm increasing in both the use
- 45:33:31cases
- 45:33:36it is going on and on right let's see
- 45:33:39See how much time it will take
- 45:33:42to add say 2 million 1 billion numbers.
- 45:33:481 billion numbers I have asked my system
- 45:33:50to add and I want to see how much time
- 45:33:53it takes.
- 45:33:55Running running running.
- 45:34:00Yep. Colonel has died. Kernel has died.
- 45:34:03People,
- 45:34:06I'll have to restart.
- 45:34:09Right. I will have to import
- 45:34:12numpy
- 45:34:14as np. So let me just remove one zero
- 45:34:19from both.
- 45:34:26It is taking 4 seconds. Removing one
- 45:34:29zero from here also. Right?
- 45:34:33And when I do this it takes 1 second. Do
- 45:34:37you see guys what is the difference in
- 45:34:40performance also right for both of
- 45:34:44these? Yeah.
- 45:34:46And if you didn't understand this, let
- 45:34:49me give you an example.
- 45:34:52Range five. This is 5 to 10, right?
- 45:35:00And this is basically
- 45:35:02adding elements,
- 45:35:06right? 5 7 9 11 13 Right. So this will
- 45:35:13be what will the output of this? This
- 45:35:15will be 0 1 2 3 4 and this will be
- 45:35:19output what 5
- 45:35:226 7 8 9 right so 0 + 5 5 6 + 1 7 7 + 2 9
- 45:35:318 + 3 11 9 + 4 13 and the same thing if
- 45:35:35I do here
- 45:35:37then what will happen I say 5 I say 5
- 45:35:42and 10 right and I say C is equal to a +
- 45:35:46b and I say c. Same thing you get here.
- 45:35:50Right? We can move on. Right? The next
- 45:35:53bit guys which we have to understand the
- 45:35:56next bit which we have to understand is
- 45:35:58called as the indexing in numpy arrays.
- 45:36:02Right? Indexing in numpy arrays. Right?
- 45:36:05How do we index the elements? Right? How
- 45:36:08do we index the elements?
- 45:36:11indexing in
- 45:36:13nump py
- 45:36:15arrays right indexing in numpy arrays
- 45:36:20so again you know indexing from basic
- 45:36:23python so let's start with 1d for 1
- 45:36:28arrays right I will use ar r r1
- 45:36:35yeah this is a ar r1 now right this is a
- 45:36:38ar r1 Okay.
- 45:36:42Now people what I want to do is what I
- 45:36:45want to do is I want to fetch right I
- 45:36:50want to fetch right you can slice and
- 45:36:53dice let's say dice
- 45:36:5533 right so again as per our normal
- 45:36:58indexing of list what is 33
- 45:37:03what is the index of 33 people
- 45:37:08two so you will say the same thing print
- 45:37:11right A R R1 squared bracket 2 and you
- 45:37:15will get 33 for yourself. Right? If you
- 45:37:18wish to slice same things, right? 33 to
- 45:37:23say 66. What is the index?
- 45:37:2933 is 2. 2. Which one? 3 4 5 and 6.
- 45:37:35Right? We will write 3 to six. Not five.
- 45:37:37Hurry. Five is not included. Remember?
- 45:37:41We print
- 45:37:43a ar r r1
- 45:37:452 is to 6 and you will get 33 44 55 66.
- 45:37:52Right? Simple indexing. Please try it
- 45:37:55out. Please try it out. And if you want
- 45:37:58you can have this code also. You can
- 45:38:00write this code. You will always have
- 45:38:02clarity that why do we use it.
- 45:38:06Moving on people. Moving on. Let's see
- 45:38:09indexing.
- 45:38:11in a 2D array. Right? And before I
- 45:38:14explain this to you, let me take you
- 45:38:16here. Right? So a 2D array will be what?
- 45:38:22Right? This is a 2D array. It has three
- 45:38:25rows and three columns. Right? Rows
- 45:38:28columns. So now for rows indexing will
- 45:38:32start from zero. So if you have to pitch
- 45:38:36this particular row, right? So what will
- 45:38:40be the index? It will be row 0. If you
- 45:38:43have to fetch this particular row, the
- 45:38:46index will be one. And if you have to
- 45:38:48fetch this particular row, this the
- 45:38:50index will be two. Similarly, for
- 45:38:53column, if you have to fetch this
- 45:38:55particular column, right, you will have
- 45:38:58column is equal to zero. This particular
- 45:39:01column, column equal to 1. This
- 45:39:04particular column, column equal to two.
- 45:39:07Right? indexing will start from minus
- 45:39:09one again right 012
- 45:39:12so let me just show that to you right
- 45:39:14let's say I call ar r r2 this is my a r2
- 45:39:18now
- 45:39:20indexing
- 45:39:23first
- 45:39:25row right how will I do that I will say
- 45:39:28print a ar r r2 and I will write how how
- 45:39:33will I write this
- 45:39:36h I will Write row 0, right? Row 0. So
- 45:39:41what will this give me? This will give
- 45:39:42me this, right? The syntax is
- 45:39:49row space column. Right? So now if you
- 45:39:52just pass one, it will give you only
- 45:39:55rows, right?
- 45:40:11row one, row two, row three. Right?
- 45:40:13Similarly,
- 45:40:16if you want to create it for columns,
- 45:40:19right? What will you say?
- 45:40:24Sorry,
- 45:40:27uh
- 45:40:28columns.
- 45:40:30Uh
- 45:40:32uh it was
- 45:40:40zero
- 45:40:45column 1 column 2
- 45:40:49column 3. Right?
- 45:40:54Right.
- 45:40:58So what is my column 1? 76 89 98 76 89
- 45:41:0398 right and if you want column 2
- 45:41:10uh sorry
- 45:41:13if you want column two this is the
- 45:41:15column two and this is column 3 right
- 45:41:18guys 90 999 99 independently
- 45:41:23now if I ask you to fetch me a
- 45:41:25particular element
- 45:41:29element, right? Say I want you to fetch
- 45:41:32me 78, right? How will you fetch 78? You
- 45:41:36will say print a ar r2. What is the row
- 45:41:39for 78? Which row does it belong to?
- 45:41:42012. To which row 78 belongs? One row.
- 45:41:46Which column it belongs to? 012.
- 45:41:501. Hurry. Check again whether it belongs
- 45:41:54to zero column, first column or second
- 45:41:56column.
- 45:42:00Right? And you will get uh sorry
- 45:42:0501. So this is two. Right? 78. Right?
- 45:42:10Second row first column. Right? Second
- 45:42:13row first column.
- 45:42:16How to define it as a
- 45:42:27right? This is your matrix. Right? Now,
- 45:42:30what are the index positions for this?
- 45:42:33This row is zero. Row, first row, second
- 45:42:37row. This is your zeroth column, first
- 45:42:41column, second column. Yeah.
- 45:42:47Yes. No. Maybe. Are we understanding
- 45:42:49this? This much is clear. The indexing
- 45:42:52of rows and columns. Now, now if you
- 45:42:55have two, so the syntax is
- 45:43:01syntax is say this is matrix A. So you
- 45:43:05will say A in this A matrix you will
- 45:43:07write row,
- 45:43:10column. Right? Suppose I want to access
- 45:43:13only zeroth row. So there will be no
- 45:43:16column. So what will you get? Row number
- 45:43:18zero.
- 45:43:20Right? Row number zero. And you will
- 45:43:22just put it like this. Or you can put a
- 45:43:25comma and put colon. Colon means what?
- 45:43:28Take everything. So I want all three
- 45:43:30columns together. Zero row and all three
- 45:43:32columns. So this will be your
- 45:43:35show. Let me show that to you.
- 45:43:41Right? See this zero row and all the
- 45:43:44columns. So what will be your answer? 76
- 45:43:4788 90. Right? Then if you want to access
- 45:43:51the second row 89 90 99 89 90 99 one and
- 45:43:57like this. Clear? Now similarly for
- 45:44:00columns what will happen? You will take
- 45:44:02all the rows
- 45:44:04comma which column do you want? If you
- 45:44:06say two what will be the result for
- 45:44:08this? What numbers will you get for
- 45:44:11this? Colon, 2, you will get 33 66 99.
- 45:44:20Now coming to the element, right?
- 45:44:22Suppose now you want to fetch 55s,
- 45:44:26right? So what will you write? You will
- 45:44:29write which row does it belong? Follow
- 45:44:32the syntax. It belongs to the first row.
- 45:44:34Which column does it belong to?
- 45:44:37First column. So what will you get? 55.
- 45:44:40Come to come come to this example. Now
- 45:44:42you want 78 in this particular array or
- 45:44:46matrix. Right? Where is 78? Which row
- 45:44:49does it belong to?
- 45:44:51Is it first row? Check carefully.
- 45:44:55Zero row. First row. Second row.
- 45:45:00Yeah. So I put two here. Now which
- 45:45:03column does it belong to? First column.
- 45:45:06Second column. Sorry. zero column, first
- 45:45:08column, second column belongs to the
- 45:45:10first column. So I put one here. So when
- 45:45:13you put this syntax, you will get 78.
- 45:45:17Clear? Now the third thing is slicing
- 45:45:22through the array. Right? Now for this I
- 45:45:26want I want 90 99 78 99. That means
- 45:45:34what? I want this, this, this, and this.
- 45:45:38Let me take you to the PPT first. Now,
- 45:45:40what I'm asking you to fetch me? I'm
- 45:45:41asking you to fetch me these four
- 45:45:43numbers. Right? These four numbers. So,
- 45:45:47here what will you write? Which rows are
- 45:45:50included in this people? Which rows are
- 45:45:52included in this?
- 45:45:54Row one to all.
- 45:45:58Right. One to all.
- 45:46:00Yeah. Not two. One to all.
- 45:46:04Right? And you leave everything like
- 45:46:06this. If you have the last row, you
- 45:46:08leave it empty after the colon, comma.
- 45:46:12Which columns do you want for this? One
- 45:46:15and two. Right? And when these will
- 45:46:18intersect, when these will intersect,
- 45:46:20you will get this area. You will get
- 45:46:22this shaded area, green shaded area. So
- 45:46:24I will say I need from column one to all
- 45:46:29the columns. Right? So what will this
- 45:46:31fetch you? What will this fetch you?
- 45:46:33This will fetch you 44
- 45:46:3655 66
- 45:46:40and 77 88 99. Right? What will this
- 45:46:46fetch you? This will fetch you 22 55 88
- 45:46:5233 66 99. What are the commonalities
- 45:46:56between both of these
- 45:46:58access
- 45:47:00both of these slices? What are the
- 45:47:01commonalities? It is only this much
- 45:47:05right
- 45:47:0955 66 88 99 right so do you think you
- 45:47:12will get your result yeah let's check it
- 45:47:14here right I say print a ar r2 right and
- 45:47:20this I say
- 45:47:22I need row zero sorry row one to empty
- 45:47:28and then column one empty and you will
- 45:47:30get 90 99 78 90 indexing.
- 45:47:35Yeah.
- 45:47:37Now guys, let's check it for
- 45:47:42three dimensions, right? 3D.
- 45:47:45I say AR R3, right? This is my AR r3,
- 45:47:48right? So, let me just put an example
- 45:47:52here. Suppose my arrays are 1 1 22 2 3 4
- 45:47:594 5 6 7 88 91. Right? This is my first.
- 45:48:07Then behind this I have another matrix
- 45:48:10which is uh
- 45:48:13111 222 333
- 45:48:17444 555 666
- 45:48:21777 888 9999.
- 45:48:25Right. And in my
- 45:48:30third one, I have 11 1 1 222 33 33 3 4
- 45:48:3744 44 44 44 44 44 44 44 44 44 44 44 44 4
- 45:48:3855 555 66 66 66 66 66 6 7 8
- 45:48:459
- 45:48:46is that is it qualifying for a 3D array?
- 45:48:50Can you access anything on this
- 45:48:52particular array if given a chance
- 45:48:55separately?
- 45:48:56like this one. Now people, if you want
- 45:48:59to move between layers, right? If you
- 45:49:02want to move between layers, do you
- 45:49:04think I have told you one particular
- 45:49:07access which is what? Which is the
- 45:49:10layers.
- 45:49:12So what number will be given to this
- 45:49:15layer? Layer number zero,
- 45:49:18layer number one
- 45:49:21and layer number two. So now if you have
- 45:49:24to access this 555
- 45:49:28which layer you will go to first you
- 45:49:31will go to first layer. Then which row
- 45:49:33will you go to? You will go to first row
- 45:49:38and first column. So if you pass this
- 45:49:42syntax what will you get? You will get 5
- 45:49:45five5.
- 45:49:47Right? Problem solved. The only
- 45:49:49bottleneck was the layers part and you
- 45:49:52have additional parameter or argument
- 45:49:54for this layer.
- 45:49:57So if I come back to my example and if
- 45:50:00you have to access this 555 right how
- 45:50:03will you do that? I say print a ar r2 a
- 45:50:08r3 right and this I say and in this I
- 45:50:11say what I have to access this 555 which
- 45:50:14which layer number is this? This is
- 45:50:17layer number zero. Right? This is layer
- 45:50:20number zero. And this is layer number
- 45:50:21one. So I say layer number one. Which
- 45:50:25row is this
- 45:50:27in this particular layer? It is layer
- 45:50:29number. Sorry, it is row number one and
- 45:50:32column number one. And when you do this,
- 45:50:34you will get 5x5.
- 45:50:36Right? If you want 999, what you will
- 45:50:39do? You will change this to 2, 2, right?
- 45:50:46If you want 777 or let's say if you want
- 45:50:4998 what will you do for 98
- 45:50:53layer number zero row number two column
- 45:50:56number zero then you will get this 98
- 45:51:02clear guys? Yep. Layer, row, column. In
- 45:51:06the same way, you will slice it. Right.
- 45:51:10You will slice it.
- 45:51:12Yep.
- 45:51:14Right. I'll write the syntax
- 45:51:17so that you don't get confused. It is
- 45:51:20array square bracket. Layer row column.
- 45:51:27Right? Layer
- 45:51:30row
- 45:51:32column.
- 45:51:34Right. This is the syntax.
- 45:51:37Perfect.
- 45:51:39Right. Great. We can also perform some
- 45:51:44operations, right? Plus, minus,
- 45:51:47multiplication, division between two
- 45:51:50arrays very very easily, right? It
- 45:51:53should not pose any problem to us,
- 45:51:55right? We can do all the operations
- 45:51:57which we want to, right? Between two
- 45:51:59arrays, right? For example, right? You
- 45:52:03had you have
- 45:52:07right operations on arrays.
- 45:52:10So let's say uh
- 45:52:14let's see as a list right list. So we
- 45:52:18have list is equal to 1 2 3 4 5 right
- 45:52:23now suppose you want to square
- 45:52:27each element of the list. Right? What
- 45:52:31will you do? You will say s sq l is
- 45:52:36equal to x to the power of 2 for x in
- 45:52:44l right and then you will say print xq l
- 45:52:50and you will get the squared of list
- 45:52:52Right?
- 45:53:01Original list squared list. Now
- 45:53:05in array what will happen? Suppose I say
- 45:53:10a 1 is equal to np dot array and l I
- 45:53:15create the same array out of this list.
- 45:53:18I say print
- 45:53:21original
- 45:53:23array
- 45:53:24right I say A1 right this is my original
- 45:53:27array same as this now I want to square
- 45:53:30it
- 45:53:33square the elements of arrays very very
- 45:53:36simple nothing you require you just say
- 45:53:40you just say
- 45:53:43sq
- 45:53:44a1 is equal to sq a1 1 is equal to a1 to
- 45:53:51the power of 2, right? A1 to the power
- 45:53:54of 2, right? And then you print
- 45:53:59the
- 45:54:05right you get the squared r. Suppose
- 45:54:10you want to find the mean of
- 45:54:21the mean of numbers
- 45:54:24using list. Right? So what will you do?
- 45:54:28You will say
- 45:54:31mean is equal to sum of
- 45:54:36l right list divided by length of list
- 45:54:42right and you will get the means
- 45:54:45right which is three for this one right
- 45:54:48this original list. Now let me show you
- 45:54:51in arrays
- 45:54:54right. How will you do this? You will
- 45:54:56say mean is equal to np dot mean and you
- 45:55:01will just pass a1 right and when you
- 45:55:04will check mean you will get 3.2
- 45:55:08Right? Nothing like this direct right
- 45:55:11direct like this right you can you have
- 45:55:13I've already showed you add I've already
- 45:55:15showed you uh square and then I believe
- 45:55:19you can understand that what all
- 45:55:21operations are possible using the arrays
- 45:55:25right leveraging the power of arrays
- 45:55:28right also guys in the arrays right in
- 45:55:32the arrays what you can do is you can
- 45:55:34perform
- 45:55:36you can perform some string operations
- 45:55:42right very powerful string operations
- 45:55:44right so for example let's say I have a
- 45:55:49I I have a array right I have an array
- 45:55:55of names of people right so I say names
- 45:56:01is equal to np dot array right and I say
- 45:56:08Radha
- 45:56:12uh then I say D
- 45:56:15right and then I say
- 45:56:19Maduk right these three names I have
- 45:56:22right so you can check the names they
- 45:56:24will be like in the array right and the
- 45:56:26data type will be U6 which is a
- 45:56:28representation of strings right now guys
- 45:56:31suppose you want to capitalize you wish
- 45:56:35to capitalize the names right of people
- 45:56:40what will you do you will say print me
- 45:56:44np docare right npcare dot capitalize
- 45:56:50right capitalize and inside this you
- 45:56:53will pass names
- 45:56:56and you will see all the names have been
- 45:56:58capitalized
- 45:57:00right you see this R has been
- 45:57:03capitalized D has been capitalized. M
- 45:57:05has been capitalized.
- 45:57:10Right?
- 45:57:13You can convert them into upper if you
- 45:57:15want. Print
- 45:57:17np.care dot upper
- 45:57:21names and you will have all of them in
- 45:57:23caps lock. You can say print np.care
- 45:57:28dot lower.
- 45:57:31You will have them in lower.
- 45:57:34Right? You can put the title. Right?
- 45:57:36Suppose I say uh
- 45:57:41title e title is equal to np dot array
- 45:57:46and I say
- 45:57:48rahov
- 45:57:53go
- 45:57:55right then I say
- 45:57:59dhapati
- 45:58:04right and I say mad
- 45:58:10warm
- 45:58:12right I say these three things now if I
- 45:58:15say print
- 45:58:18npcare
- 45:58:20dot
- 45:58:21title right and I say
- 45:58:25titles you will see that all the words
- 45:58:29will be in capitalized mode rael d and s
- 45:58:33of dh sinapati m and V of MaduMa are now
- 45:58:37capitalized. I can also
- 45:58:41replace something if I wish to suppose I
- 45:58:45want to replace Madhu with say suri
- 45:58:50right I will say print
- 45:58:53right np do.care care dot replace
- 45:59:02right and you will say where you want to
- 45:59:05replace I say title
- 45:59:09and in this I want to replace
- 45:59:12madu
- 45:59:14with
- 45:59:19suri
- 45:59:20right and you will say it will be suri_1
- 45:59:24suri one Right. Madu has been replaced
- 45:59:26with
- 45:59:29Right.
- 45:59:33Right. If you want to calculate
- 45:59:36the characters of strings,
- 45:59:40right, you can do that. Print np.car
- 45:59:46dot str length of
- 45:59:50titles and you will get 11 characters
- 45:59:52are there in here.
- 45:59:5515 are there here and 11 are here in
- 45:59:59this particular thing.
- 46:00:01Right? You can do much more powerful
- 46:00:04things also. Let me show you one complex
- 46:00:06function. Right? Suppose I have f name
- 46:00:12is equal to np dot array.
- 46:00:16Right? And we have Ra
- 46:00:25Madu
- 46:00:27right and we have L name
- 46:00:33arapati
- 46:00:42worma. Right, we have these two things.
- 46:00:45Now I can create a new array full name
- 46:00:50by simply right by simply saying np.car
- 46:00:54car dot add right and I say uh f name
- 46:01:02right
- 46:01:06comma
- 46:01:11l
- 46:01:13name right
- 46:01:16Yes.
- 46:01:26Yep.
- 46:01:32Full name.
- 46:01:34Yeah. Ra. I just was trying to add a
- 46:01:38space
- 46:01:40in between.
- 46:01:43Anyway,
- 46:01:54right. We can do that,
- 46:01:58right? You can also people search in
- 46:02:02arrays, right? Very powerful. Again,
- 46:02:04search in arrays using where,
- 46:02:08right?
- 46:02:10Right. You can say suppose a 2 is equal
- 46:02:14to
- 46:02:16np dot array right and I'll say 1 1 22
- 46:02:2033 3 4 4 5 66 right and now you have to
- 46:02:24say a is equal to np dot where right and
- 46:02:30in this you say a2 greater than 20
- 46:02:36right and when you will check A
- 46:02:43uh it's giving me the index. Why is it
- 46:02:45giving me the index
- 46:02:48or is it giving me the index?
- 46:02:52Does it always return index?
- 46:02:56One more thing is you can find a you can
- 46:03:01find an
- 46:03:03a letter
- 46:03:06right through a letter. You can find an
- 46:03:08element
- 46:03:10through
- 46:03:12a letter. Guys, these are all some
- 46:03:14tricks which you should know because you
- 46:03:16will be dealing with data and you need
- 46:03:18to pull data, right? You need to pull
- 46:03:19data a lot, right? Based on conditions
- 46:03:21and based on things. Suppose you want to
- 46:03:24find out the names which have G in them
- 46:03:29right or R A in them. So how will you do
- 46:03:31this? I will say print right and I will
- 46:03:34say uh np do.care care dot find right
- 46:03:41and I will say find this inside full
- 46:03:46names right and find me ra a right ra a
- 46:04:03two p
- 46:04:07it uh right it returns true because it
- 46:04:10has found it here right so I don't want
- 46:04:14to tell you indexing through this but
- 46:04:16anyway you should know this just just
- 46:04:18assume this that I'm telling you to
- 46:04:19write this okay because this is much
- 46:04:21easier when we will go to pandas right
- 46:04:26just uh write it as a syntax okay
- 46:04:29greater than equal to zero I hope this
- 46:04:32is clear
- 46:04:33>> so let's start with the data science
- 46:04:34interview questions and answers and The
- 46:04:36number one problem we would be facing is
- 46:04:38real world problem solving. And the
- 46:04:40question one is handling missing data in
- 46:04:43predictive modeling. So imagine you have
- 46:04:45given a data set where 30% of the data
- 46:04:48for key predictive variable is missing.
- 46:04:50This variable is crucial for a
- 46:04:52predictive model. How would you handle
- 46:04:54this situation to ensure the integrity
- 46:04:56and performance of your model? And
- 46:04:58please describe your approach step by
- 46:05:00step. So starting with the answer you
- 46:05:02can start with handling missing data set
- 46:05:04is a common challenge in data science
- 46:05:06and it's important to address it
- 46:05:08carefully to maintain the accuracy of
- 46:05:10your model and here's how you could
- 46:05:13approach this situation. The number one
- 46:05:14point could be identify the missing
- 46:05:16data. So first you need to understand
- 46:05:18where the missing values are in your
- 46:05:20data set. You can do this by using a
- 46:05:23simple code in Python with libraries
- 46:05:24like mandas. For example, you can use
- 46:05:27the data dot isnull dot sum function
- 46:05:31that will show you the count of missing
- 46:05:33values in each column. Then you can
- 46:05:35analyze the pattern. Determine if
- 46:05:37there's a pattern to the missing data.
- 46:05:39Is it random or is it missing for a
- 46:05:41reason? This can affect your approach.
- 46:05:43If the data is missing at random, the
- 46:05:45methods you use might be different than
- 46:05:47if the data is missing systematically.
- 46:05:49So choosing a method for imputation.
- 46:05:51Let's see the next method that is
- 46:05:54choosing a method for imputation. So if
- 46:05:56the missing data is numeric, you might
- 46:05:58replace missing values with the mean or
- 46:06:00median of that column. This is simple
- 46:06:02and effective but can be used primarily
- 46:06:05when the data is missing completely at
- 46:06:06random. Then comes model based
- 46:06:09imputation. Sometimes you can use other
- 46:06:11variables in the data to predict missing
- 46:06:13values using a regression model. This
- 46:06:15can be more accurate but is also more
- 46:06:17complex. Then we'll use the k nearest
- 46:06:20neighbors can algorithm. But before that
- 46:06:23we have a code snippet here that could
- 46:06:26be used for the implementation of
- 46:06:27imputation. You could use Python or R.
- 46:06:30And now moving on we'll see the K
- 46:06:32nearest neighbors algorithm. So this
- 46:06:34method predicts the missing values based
- 46:06:36on how closely related the data points
- 46:06:38are to each other. So after imputation
- 46:06:41it's crucial to check how your changes
- 46:06:43have affected the overall data set and
- 46:06:44model performance. Sometimes filling in
- 46:06:47too many missing values can introduce
- 46:06:49bias. And then we have visualization. To
- 46:06:52help understand before and after the
- 46:06:53imputation, you could visualize the
- 46:06:55distribution of the variable using
- 46:06:57histograms or box plots. This helps in
- 46:07:00seeing how the imputation has changed
- 46:07:02the statistical properties of the data.
- 46:07:04And by following these steps, you can
- 46:07:05handle missing data thoughtfully and
- 46:07:07maintain the integrity of your
- 46:07:09predictive model. Now moving to the
- 46:07:11question number two that is based on
- 46:07:13evaluating model overfitting. So the
- 46:07:16question is you have developed a
- 46:07:18predictive model but you suspect it
- 46:07:20might be overfitting the training data.
- 46:07:22How would you test and address the
- 46:07:24issue? Please explain your steps and the
- 46:07:26techniques you would use. So you could
- 46:07:28start the answer by explaining what is
- 46:07:31overfitting. So overfitting is a common
- 46:07:33problem where model performs well on
- 46:07:35training data but poorly on unseen data
- 46:07:37indicating it's too closely fitted to
- 46:07:39the training data specific details and
- 46:07:41noise. So now we'll see a step-by-step
- 46:07:44guide on how to address this. The number
- 46:07:46one step is cross validation. So one
- 46:07:48effective way to test for overfitting is
- 46:07:51by using cross validation technique.
- 46:07:53Cross validation involves splitting your
- 46:07:55training data into multiple smaller sets
- 46:07:57that is false and then training a model
- 46:07:59on some of these set and validating it
- 46:08:02on the others. So this helps you
- 46:08:04understand if the model's good
- 46:08:05performance is consistent across
- 46:08:07different subsets of data. For example,
- 46:08:09in Python you can use the cross value
- 46:08:12score function from skarn.mmodel
- 46:08:15selection. So this is the code and this
- 46:08:19is the code snippet of Python that you
- 46:08:21can use for the cross validation and
- 46:08:23here we are importing from skarn that is
- 46:08:26the module and we're importing
- 46:08:29cross_well
- 46:08:30score and here we have used the cross
- 46:08:33val score function and then we have
- 46:08:36printed the average cross validation
- 46:08:38score and the next step we will do is
- 46:08:40running cross validation model. So this
- 46:08:42is your predictive model that you have
- 46:08:44already built using scikit learn and
- 46:08:46here's the x train these are the x input
- 46:08:49features of your training data and y
- 46:08:51train these are the output labels of
- 46:08:53training data. So we are running gross
- 46:08:55validation model here this is your
- 46:08:57predictive model that you have already
- 46:08:59built using scikitlearn. So x train here
- 46:09:02that means these are the input features
- 46:09:03of your training data and y train here
- 46:09:06means these are the output labels of
- 46:09:07training data and cv equal to 5. This
- 46:09:10parameter tests the function to split
- 46:09:11the data into five parts that is false.
- 46:09:14And the model is trained on four of
- 46:09:16these parts and the remaining part is
- 46:09:18used for testing. So this process
- 46:09:20rotates until each part has been used
- 46:09:22for testing once and the printing
- 46:09:24results that is score dot mean. So this
- 46:09:27calculates the average of the scores
- 46:09:29obtained from each gross validation for
- 46:09:31this average score gives you an idea of
- 46:09:34how well your model is likely to perform
- 46:09:36on unseen data. A consistent score
- 46:09:38across different polls suggests your
- 46:09:40model is generalizing well rather than
- 46:09:42overfitting to the training data. So now
- 46:09:44moving to the next point that is
- 46:09:46training versus validation error. So
- 46:09:48plot the training and validation errors
- 46:09:50as a function of training epochs or
- 46:09:52complexity of the model. A model that
- 46:09:54overfits will show a low error on
- 46:09:56training data and a high error on
- 46:09:58validation data as it trains further.
- 46:10:00Then we have pruning the model. If you
- 46:10:03confirm that the model is overfitting,
- 46:10:05consider simplifying it. This might mean
- 46:10:07reducing the number of parameters by
- 46:10:09selecting fewer features using
- 46:10:12regularization techniques like lasso or
- 46:10:14ridge or choosing a less complex model.
- 46:10:17After this step, we will move to
- 46:10:18regularization technique step. So these
- 46:10:20techniques add a penalty to the loss
- 46:10:22function used to train the model which
- 46:10:25can discourage complex models that
- 46:10:27overfeit. Then we have common methods
- 46:10:28that include L1 that is lasso and L2
- 46:10:31ridge regularization. And here's how you
- 46:10:34can add L2 regularization in Python. So
- 46:10:37this is the code snippet here. And what
- 46:10:38we have done here is we are creating the
- 46:10:40ridge model and we have applied alpha
- 46:10:43equal to 1.0. So this parameter controls
- 46:10:45the strength of the regularization. A
- 46:10:47higher alpha value increases the
- 46:10:50regularization effect which helps reduce
- 46:10:52model complexity and combat overfitting.
- 46:10:55The alpha value can be tuned to find the
- 46:10:58optimal balance between bias and
- 46:11:00variance. And now coming for the fitting
- 46:11:02the model. So model do fit and in that
- 46:11:05we have X train and Y train that trains
- 46:11:08the ridge model on the training data. It
- 46:11:10adjusts the weight of the feature in X
- 46:11:12train to predict the Y train while also
- 46:11:15considering the regularization term.
- 46:11:17This helps prevent the model from
- 46:11:19fitting too closely to the noisy aspects
- 46:11:21of the training data. And then we are
- 46:11:23re-evaluating the model. After making
- 46:11:25adjustments, it's important to
- 46:11:27re-evaluate the model again using the
- 46:11:29same cross validation technique to see
- 46:11:32if the issue of overfitting has
- 46:11:33improved. And then we have
- 46:11:34visualization. To help illustrate or
- 46:11:36ffitting, you could create a plot
- 46:11:38showing the training and validation
- 46:11:40errors or the number of epochs or model
- 46:11:42complexity. So by using these
- 46:11:44techniques, you can identify if your
- 46:11:45model is all fitting and take steps to
- 46:11:47correct it ensuring it performs well not
- 46:11:50only on the training data but also on
- 46:11:51new unseen data. So now moving to the
- 46:11:53next question that is question number
- 46:11:55three and it is based on realtime data
- 46:11:57stream processing and the question is
- 46:11:59you are tasked with building a model to
- 46:12:01predict stock prices in real time. The
- 46:12:04data comes in every second and you need
- 46:12:06to update your predictions accordingly.
- 46:12:08Describe how you would set up your
- 46:12:09system to handle this type of data
- 46:12:11effectively and what tools and
- 46:12:13techniques would you use and why. So you
- 46:12:15could start answering this question with
- 46:12:17handling real-time data. So handling
- 46:12:19real-time data especially for something
- 46:12:21as volatile and fastpaced as stock
- 46:12:23prices requires a robust system that can
- 46:12:26process and analyze data quickly and
- 46:12:28accurately. So here's how you could
- 46:12:30approach this. We will set up such a
- 46:12:32system and we'll have some steps. So
- 46:12:35starting with the steps. So the first
- 46:12:36step is choosing the right tools. The
- 46:12:38right tool would be Apache Kafka. So
- 46:12:40this is a popular tool for handling
- 46:12:42real-time data that streams because it
- 46:12:45allows you to publish and subscribe to
- 46:12:46streams of records that is data and it
- 46:12:49can handle high throughput with low
- 46:12:50latency. Kafka acts as a buffer and
- 46:12:53manages the flow of data ensuring that
- 46:12:54your system doesn't get overwhelmed and
- 46:12:57you can also use Apache Spark especially
- 46:12:59Spark streaming is excellent for
- 46:13:01processing the data. It can process data
- 46:13:03in real time and perform complex
- 46:13:05operations like windowing, grouping data
- 46:13:07into chunks of a specified time period
- 46:13:10and aggregating that is summarizing
- 46:13:12data. So you can modify it and perform
- 46:13:14the predicting of stock prices. And then
- 46:13:17the step is data processing pipeline.
- 46:13:19And the first step comes here is
- 46:13:21injection. Data first enters the system
- 46:13:23typically through Kafka which collects
- 46:13:25data sent from the stock market and then
- 46:13:27we do the processing. So spark streaming
- 46:13:30takes over here. Here you can apply
- 46:13:32transformations and run your predictive
- 46:13:34models on the data. For example, you
- 46:13:36might calculate moving averages or other
- 46:13:38indicators that feed into your stock
- 46:13:40price prediction model. And then comes
- 46:13:42the output. Finally, the predictions are
- 46:13:44outputed. This could be to a dashboard
- 46:13:47for traders, an automated trading system
- 46:13:49or even stored for further analysis. And
- 46:13:52then we develop the model. Now comes the
- 46:13:54model development. You would likely use
- 46:13:56a machine learning model that can update
- 46:13:57quickly and incorporate new data as it
- 46:14:00arrives. models such as aim for time
- 46:14:03series forecasting or more complex
- 46:14:05machine learning models like rect neural
- 46:14:07networks RNNs can be suitable. The model
- 46:14:10should be retrained or fine-tuned
- 46:14:12periodically with new data to ensure it
- 46:14:14stays accurate. Now we'll come to
- 46:14:16scalability and reliability. So ensure
- 46:14:19your system can scale as data volume
- 46:14:21increases. This might mean adding more
- 46:14:23servers or optimizing your data
- 46:14:24processing code. Implement monitoring to
- 46:14:27catch any issues early like delays in
- 46:14:29data processing or model performance
- 46:14:31drops. And now we'll see the step that
- 46:14:33is visualization and monitoring.
- 46:14:35Consider setting up a real-time
- 46:14:36dashboard that shows key metrics like
- 46:14:39prediction accuracy and processing time.
- 46:14:41This helps in quickly spotting when
- 46:14:43something goes wrong. By setting up your
- 46:14:45system with these tools and strategies,
- 46:14:47you can effectively handle the challenge
- 46:14:49of predicting stock prices in real time.
- 46:14:51So now we'll move to the next question
- 46:14:52that is question number four and this
- 46:14:55will based on scalable data analytics.
- 46:14:57So we have covered two questions that
- 46:14:59were a bit code based questions and now
- 46:15:02we'll see other questions that would be
- 46:15:04based on scalable data analytics or they
- 46:15:07might be on different areas and with the
- 46:15:1013th question we'll start again with the
- 46:15:12coding ones. So moving with the question
- 46:15:14four that is based on scalable data
- 46:15:16analytics and the question is given a
- 46:15:18scenario where your organization
- 46:15:20suddenly needs to scale its data
- 46:15:22analysis capabilities due to an influx
- 46:15:24of data that would be 10 times the
- 46:15:27normal volume. How would you handle this
- 46:15:29situation to ensure your data analytics
- 46:15:31processes remain efficient and accurate?
- 46:15:33What technologies would you consider and
- 46:15:35what steps would you take? So you can
- 46:15:37start answering this question with
- 46:15:39handling a sudden increase in data
- 46:15:41volume requires a strategic approach to
- 46:15:43scaling your analytics infrastructure
- 46:15:45without compromising on efficiency or
- 46:15:47accuracy. So we'll see some steps from
- 46:15:50that you could effectively manage this
- 46:15:52scenario that you would start answering
- 46:15:54the interviewer that we can start by
- 46:15:56evaluating the current infrastructure's
- 46:15:58ability to handle increased loads. This
- 46:16:01includes assessing your databases,
- 46:16:02servers and analytical tools to identify
- 46:16:05potential bottlenecks or limitations.
- 46:16:08Then you could move to next step that
- 46:16:10would be choosing scalable technologies
- 46:16:12to manage the increased data volume.
- 46:16:14Consider leveraging cloud-based
- 46:16:15solutions such as Amazon web services,
- 46:16:17Google cloud platform or Microsoft
- 46:16:19Azure. These platforms offer scalable
- 46:16:21resources which can be adjusted
- 46:16:23accordingly to the data load ensuring
- 46:16:25you only pay for what you use. integrate
- 46:16:27big data technologies like Apache Hadoop
- 46:16:29for distributed storage and Apache Spark
- 46:16:31for fast data processing. These tools
- 46:16:33are designed to handle massive volumes
- 46:16:35of data efficiently and can scale up to
- 46:16:38meet standard increased demands. Now we
- 46:16:41move to the next step that would be
- 46:16:42optimizing data processing. So implement
- 46:16:44data partitioning and indexing
- 46:16:47strategies to improve the efficiency of
- 46:16:49data queries. This will help in managing
- 46:16:51large data sets by breaking them into
- 46:16:53smaller manageable chunks and speeding
- 46:16:55up search operations and use real-time
- 46:16:58data processing frameworks like Apache
- 46:17:00Kafka or Apache Flink which can handle
- 46:17:02high throughput and provide timely
- 46:17:04insights from large data streams. And
- 46:17:06the next step would be automation and
- 46:17:08monitoring. Automate routine data
- 46:17:09processing task to reduce the manual
- 46:17:11effort and speed up the analysis. This
- 46:17:13can be done through scripting or using
- 46:17:15workflow automation tools. Set up
- 46:17:17comprehensive monitoring systems to
- 46:17:19track the performance of your data
- 46:17:21processes. Tools like Prometheus for
- 46:17:23system monitoring and Graphana for
- 46:17:25analytics and monitoring dashboards are
- 46:17:27useful here. They help ensure that the
- 46:17:29system is running smoothly and alert you
- 46:17:32to potential issues before they become
- 46:17:34critical. And the next step will be
- 46:17:36regular evaluation and scaling.
- 46:17:39Continuously evaluate the performance of
- 46:17:40analytics infrastructure. As your data
- 46:17:43grows, keep adjusting and scaling your
- 46:17:45resources to maintain optimal
- 46:17:46performance. Plan for periodic reviews
- 46:17:48of your technology stack and
- 46:17:50infrastructure to ensure they remain
- 46:17:52aligned with your data needs and
- 46:17:54organizational goals. By following these
- 46:17:56steps, you can ensure that your data
- 46:17:57analytics processes are prepared to
- 46:17:59handle sudden surges in data volume
- 46:18:01effectively maintaining the integrity
- 46:18:03and speed of insights. So this was all
- 46:18:06for the question four. Now moving to the
- 46:18:08question five and this is based on
- 46:18:10integrating machine learning models into
- 46:18:12production and the question is you have
- 46:18:14developed a machine learning model that
- 46:18:15performs well in testing environment.
- 46:18:17Now you need to integrate it into your
- 46:18:19production environment where it will be
- 46:18:21used in realtime applications. What
- 46:18:23steps would you take to ensure the
- 46:18:25successful deployment and operations of
- 46:18:27the model in production? So we'll start
- 46:18:29answering this by successfully deploying
- 46:18:31a machine learning model into production
- 46:18:33involves several critical steps to
- 46:18:35ensure it performs as well in real time
- 46:18:38operations as it does in testing. So you
- 46:18:40would have a clear pathway to make the
- 46:18:43interviewer understand. We will start
- 46:18:45with the pathway with the first step
- 46:18:46that would be model validation. So
- 46:18:49before moving anything into production
- 46:18:51revalidate your model's performance
- 46:18:52using a separate validation data set.
- 46:18:55This helps confirm that the model
- 46:18:57generalizes well to new unseen data. The
- 46:19:00next step will be preparing the
- 46:19:02production environment. Ensure that the
- 46:19:03production environment is ready to
- 46:19:05handle the model. This includes setting
- 46:19:07up the necessary hardware and software
- 46:19:09ensuring that it can handle the expected
- 46:19:11load and that all dependencies are
- 46:19:14correctly installed and configured. Then
- 46:19:16the next step comes that is model
- 46:19:18wrapping. Wrap your model in an API that
- 46:19:20is application programming interface
- 46:19:22making it accessible to other parts of
- 46:19:24your software infrastructure. Frameworks
- 46:19:26like flask for Python can be used to
- 46:19:28create a simple web server that listens
- 46:19:30for data inputs and provides model
- 46:19:32outputs. Then comes the next step that
- 46:19:34is deployment strategies. Consider using
- 46:19:37containerization tools like doer which
- 46:19:39can help encapsulate your model and its
- 46:19:41environment ensuring that it works
- 46:19:43uniformly across different development
- 46:19:46and production settings. And then we'll
- 46:19:48use deployment strategies like blue
- 46:19:50green deployment or canary releases to
- 46:19:53minimize downtime and reduce the risk of
- 46:19:55introducing a faulty model into
- 46:19:57production. And then comes the next step
- 46:19:59that is monitoring and logging.
- 46:20:01Implement logging and monitoring to
- 46:20:03track the model's performance and health
- 46:20:05in real time. Tools like prompts for
- 46:20:07monitoring and ELK elastic search log
- 46:20:11statch kibbana for logging help in
- 46:20:13quickly identifying and diagnosing
- 46:20:15issues in production. And then comes the
- 46:20:17next step that is performance tuning.
- 46:20:19Monitor the model's performance over
- 46:20:21time. If the model's performance
- 46:20:23degrades or if new data shows different
- 46:20:25patterns, you may need to retrain or
- 46:20:28fine-tune the model to maintain
- 46:20:29accuracy. And after this step, there's a
- 46:20:32step for feedback loop. Set a feedback
- 46:20:34loop where predictions and outcomes can
- 46:20:36be compared. This feedback is crucial
- 46:20:39for continuously improving the model and
- 46:20:41catching any drift in data or changes in
- 46:20:44external conditions that affect the
- 46:20:45model. And after this comes a last step
- 46:20:48that is legal and compliance checks.
- 46:20:50Ensure all the data used by the model in
- 46:20:52production complies with privacy laws
- 46:20:54and regulations. This is crucial for
- 46:20:56maintaining trust and legality
- 46:20:58especially when handling sensitive
- 46:21:00information. So by carefully planning
- 46:21:02and executing these steps you can
- 46:21:04smoothly transition your machine
- 46:21:05learning model from a testing
- 46:21:07environment to a fully functional
- 46:21:09component of a production system. So
- 46:21:11this was all about the question number
- 46:21:12five. Now moving to the question number
- 46:21:14six that would be based on datadriven
- 46:21:16decision making. And the question is
- 46:21:18your company wants to shift towards more
- 46:21:20datadriven decision making. You have
- 46:21:23been tasked with developing a strategy
- 46:21:25to implement this. What steps would you
- 46:21:27take to ensure that the data at all
- 46:21:29levels of the organization is utilized
- 46:21:31effectively to make informed decisions
- 46:21:34and what challenges might you face and
- 46:21:36how would you address them? So you can
- 46:21:37start answering this by implementing a
- 46:21:39datadriven decision-m strategy that will
- 46:21:42require a comprehensive approach to
- 46:21:44ensure that reliable data is accessible
- 46:21:47and effectively used across all levels
- 46:21:49of the organization. And now we can
- 46:21:51develop and deploy this strategy. And
- 46:21:53similarly you could tell this strategy
- 46:21:55to the interviewer. So the number one
- 46:21:57step will be assessing current data
- 46:21:59infrastructure. Start by evaluating the
- 46:22:01existing data infrastructure to
- 46:22:03understand what data is available, how
- 46:22:05it is stored and how it is currently
- 46:22:06used. This assessment will help identify
- 46:22:09gaps in data collection, storage and
- 46:22:11access that need to be addressed. Now we
- 46:22:14move to the next step that is developing
- 46:22:16a data governance framework. Implement a
- 46:22:19data governance framework that defines
- 46:22:21who can access data, how it can be used
- 46:22:23and who is responsible for maintaining
- 46:22:25its quality. This framework ensures data
- 46:22:28integrity and security which are
- 46:22:29critical for making reliable decisions.
- 46:22:32Now we'll move to the next step that is
- 46:22:33training and empowerment. So train
- 46:22:35employees at all levels on the
- 46:22:37importance of datadriven decision making
- 46:22:40and provide them with the tools and
- 46:22:42knowledge necessary to analyze and
- 46:22:44interpret data. This might include
- 46:22:46training sessions, workshops and ongoing
- 46:22:48support to ensure everyone can use data
- 46:22:50effectively. Now move to the next step
- 46:22:52that is implementing analytical tools.
- 46:22:54So deploy user-friendly analytical tools
- 46:22:56that can integrate seamlessly into the
- 46:22:58daily workflows of employees. Tools like
- 46:23:01Tableau, Microsoft PowerBI or even
- 46:23:03advanced Excel techniques can provide
- 46:23:05powerful data analysis capabilities
- 46:23:07without requiring extensive technical
- 46:23:09knowledge. After this we'll move to the
- 46:23:12step that would be creating a
- 46:23:13centralized data platform. Developer
- 46:23:16centralized data platform where all
- 46:23:18organizational data can be accessed and
- 46:23:20analyzed. This platform should be
- 46:23:22scalable and secure providing a single
- 46:23:24source of truth for the organization.
- 46:23:27And then we have the promoting a
- 46:23:28datadriven culture. So foster culture
- 46:23:31that values datadriven decision-m
- 46:23:33encourage experimentation and learning
- 46:23:35from datadriven initiatives. celebrate
- 46:23:37successes and learn from failures to
- 46:23:39continually improve the use of
- 46:23:40datadriven in decision making and there
- 46:23:43would be some challenges and solutions
- 46:23:45for that. So one major challenge we know
- 46:23:47here is resistance to change as some
- 46:23:49employees may prefer traditional
- 46:23:50decision-m methods. So address this by
- 46:23:53demonstrating the tangible benefits of
- 46:23:55datadriven decisions through pilot
- 46:23:57projects and success stories. So data
- 46:24:00silos can also hinder effective data use
- 46:24:02promote cross department collaboration
- 46:24:04and integrate disparate data sources to
- 46:24:07overcome this challenge. After that you
- 46:24:09can monitor and do continuous
- 46:24:10improvement. So by systematically
- 46:24:13implementing these steps you can
- 46:24:14transform your organization into one
- 46:24:16that leverages data at all levels to
- 46:24:19make informed and effective decisions.
- 46:24:21And after answering in these steps you
- 46:24:23could make the interviewer have a truth
- 46:24:26and a faith in you that you could make
- 46:24:28these models. Now move to the next
- 46:24:30question that is question number seven
- 46:24:32and that is based on handling large data
- 46:24:34set and the question is your project
- 46:24:36involves analyzing extremely large data
- 46:24:39sets potentially exceeding terabytes in
- 46:24:41size. What strategies would you use to
- 46:24:43manage and analyze such large data sets
- 46:24:46effectively? Describe the tools and
- 46:24:48techniques you might employ and you
- 46:24:50could start this with answering that
- 46:24:52working with large data sets especially
- 46:24:54those in terabyte range presents unique
- 46:24:56challenges in terms of storage
- 46:24:58processing and analysis. So we'll have a
- 46:25:00structured approach to handle these
- 46:25:02challenges effectively. We'll start with
- 46:25:04the data storage that would be use
- 46:25:06distributed file systems. Consider using
- 46:25:09a distributed file systems like Hadoop
- 46:25:11distributed file system HDFS or Amazon
- 46:25:14S3. These systems are designed to store
- 46:25:16vast amounts of data across many servers
- 46:25:18offering high availability and port
- 46:25:21tolerance. And then comes the next step
- 46:25:23that is data processing. Leverage big
- 46:25:25data processing frameworks. Tools like
- 46:25:27Apache Spark are ideal for processing
- 46:25:29large data sets because they handle
- 46:25:31distributed computing effectively. Spark
- 46:25:34can perform data processing task much
- 46:25:36faster than traditional disk based
- 46:25:38processing due to its in-memory
- 46:25:40computing capabilities. And next we
- 46:25:42could start with efficient data
- 46:25:44sampling. So there are many sampling
- 46:25:46techniques that we can use. So when the
- 46:25:48data set is too large to handle even
- 46:25:50with powerful tools consider using data
- 46:25:53sampling techniques to reduce the size
- 46:25:55to a manageable level without losing
- 46:25:57significant insights. Ensure that the
- 46:25:59sample represents the whole data set
- 46:26:00accurately. And then comes optimization
- 46:26:03of data queries. Indexing and
- 46:26:05partitioning. Optimize your data queries
- 46:26:07by implementing indexing and
- 46:26:08partitioning. This can drastically
- 46:26:10reduce the time it takes to perform
- 46:26:12queries by limiting the amounts of data
- 46:26:14scan. And then we can do scalable
- 46:26:16analytics. And then we'll move to the
- 46:26:18next step that is scalable analytics.
- 46:26:20And in that we could start with the
- 46:26:21parallel computing. Use parallel
- 46:26:23computing capabilities of frameworks
- 46:26:25like spark or dask to analyze data
- 46:26:28across multiple nodes. This helps in
- 46:26:30scaling up your analytics operations to
- 46:26:32handle large data sets effectively. And
- 46:26:34now we'll move to the cloud-based
- 46:26:36analytical tools. So consider using
- 46:26:38cloud services like Google BigQuery or
- 46:26:40AWS Red Shift which are designed to
- 46:26:42handle massive data sets and complex
- 46:26:45analytics with ease. And after this step
- 46:26:47we'll move to data cleaning and
- 46:26:48pre-processing. Here we will automate
- 46:26:50pre-processing task. We'll use automated
- 46:26:52tools to clean and pre-process data.
- 46:26:55This includes handling missing values,
- 46:26:57normalizing data and removing duplicates
- 46:27:00which can be particularly challenging
- 46:27:01with large data set. And after this
- 46:27:03step, we'll move to the step that will
- 46:27:05visualize large data set. So we'll use
- 46:27:08specialized tools. That tools could be
- 46:27:10Tableau or PowerBI that can handle large
- 46:27:12data set by aggregating data and using
- 46:27:15efficient backend technologies. For more
- 46:27:18detailed exploration, tools like plotly
- 46:27:20or bouquet can be used as they offer
- 46:27:22capabilities to interactively visualize
- 46:27:24large volumes of data. And after that,
- 46:27:26there would be step for regular
- 46:27:28maintenance and updates. That could be
- 46:27:30continuously monitoring the data
- 46:27:32quality. As new data comes in, you can
- 46:27:35continuously monitor its quality. And
- 46:27:37after this step, you could integrate all
- 46:27:39these strategies and tools into your
- 46:27:41workflow. And you can effectively manage
- 46:27:43and extract valuable insights from
- 46:27:45extremely large data sets thereby
- 46:27:48supporting robust datadriven decision
- 46:27:50making. And you could answer the whole
- 46:27:52strategy to the interviewer. Now moving
- 46:27:55to the question number eight that is
- 46:27:56based on optimizing machine learning
- 46:27:58models and the question is during model
- 46:28:00development you have noticed that your
- 46:28:01machine learning model is
- 46:28:02underperforming. What steps would you
- 46:28:04take to diagnose the problem and
- 46:28:06optimize the model's performance? What
- 46:28:08techniques and tools would you use? So
- 46:28:10you can answer this by starting with the
- 46:28:12optimizing and optimizing a machine
- 46:28:14learning model that is underperforming
- 46:28:17involves several steps to diagnose and
- 46:28:19improve its accuracy and efficiency. And
- 46:28:21here we will have structured approach to
- 46:28:23tackle this issue and you could start
- 46:28:25this with the number one step that is
- 46:28:27diagnosing the problem. Evaluate model
- 46:28:29metrics. Start by thoroughly evaluating
- 46:28:31the performance metrics of your model.
- 46:28:33For classification task, for
- 46:28:35classification task, look at accuracy,
- 46:28:37precision, recall and the F1 score. For
- 46:28:40regression task, consider R squ, mean
- 46:28:42squared error that is MSE and mean
- 46:28:45absolute error that is MA. And then you
- 46:28:48can move to the next step that is use
- 46:28:50plots like ROC curves for classification
- 46:28:52models and residual plots for regression
- 46:28:54to visually assess with the model is
- 46:28:57going wrong. After that, we'll move to
- 46:28:58the next step that is data quality and
- 46:29:00quantity check. Inspect the data that is
- 46:29:03sometimes the quality and quantity of
- 46:29:05data can be the root cause of poor model
- 46:29:08performance. Ensure the data is clean,
- 46:29:10well pre-processed and sufficient. Look
- 46:29:12for issues like missing values, outliers
- 46:29:14or imbalanced classes. And after this
- 46:29:17we'll move to the feature engineering
- 46:29:19step that would be experiment with
- 46:29:21creating new features or transforming
- 46:29:23existing ones to provide better
- 46:29:25predictive power. And then we have the
- 46:29:27next step that is model tuning and
- 46:29:28configuration. After feature
- 46:29:30engineering, we'll move to the next step
- 46:29:32that is model tuning and configuration.
- 46:29:34So, hyperparameter tuning. Use
- 46:29:36techniques like grid search or random
- 46:29:38search to find the optimal settings for
- 46:29:40your model's parameters. Tools like
- 46:29:41scikit learns, grid search CV or
- 46:29:44randomized search CV can automate this
- 46:29:46process. And there's a cross validation
- 46:29:49that would implement cross validation to
- 46:29:51ensure that the model's performance is
- 46:29:53consistent across different subsets of
- 46:29:55the data set. And then we have the next
- 46:29:57step that is trying different models. So
- 46:29:59experiment with algorithms here. If
- 46:30:01initial models are underperforming, try
- 46:30:03different algorithms that might be
- 46:30:05better suited for the problem. For
- 46:30:07instance, if you started with linear
- 46:30:09regression and it's not performing well,
- 46:30:11consider more complex models like random
- 46:30:13forest or gradient boosting machines.
- 46:30:16And after this, we have nseml methods
- 46:30:18that we can use techniques like bagging,
- 46:30:21boosting or stacking to combine the
- 46:30:23predictions of multiple models to
- 46:30:25improve overall performance. After this
- 46:30:27step, we have feature selection that
- 46:30:29includes reduce dimensionality. Use
- 46:30:32techniques like principal component
- 46:30:34analysis that is PCA to reduce the
- 46:30:36number of features which might help in
- 46:30:38improving model performance by removing
- 46:30:39noise and redundancy. And then we have
- 46:30:42select important features. So use model
- 46:30:44based technique to identify and keep
- 46:30:46only the most important features that
- 46:30:48impact the outcome. And then comes the
- 46:30:50last step that is regular updates and
- 46:30:52retraining. So here you can monitor and
- 46:30:54update that could be continuously
- 46:30:56monitoring the model's performance over
- 46:30:58time as new data becomes available
- 46:31:00update and retrain the model to adapt to
- 46:31:03any changes in underlying patterns and
- 46:31:06after that you could have a consultation
- 46:31:07and collaboration work with the other
- 46:31:09teams and by methodically addressing
- 46:31:11each of these areas you can diagnose why
- 46:31:13your machine learning model is
- 46:31:15underperforming and can take steps to
- 46:31:17optimize its accuracy and efficiency. So
- 46:31:19this was all about question number
- 46:31:21eight. So let's start with the question
- 46:31:22number nine and this is based on
- 46:31:24handling unstructured data. So the
- 46:31:26question is you are given a large amount
- 46:31:28of unstructured data including text,
- 46:31:30images and videos. What strategies would
- 46:31:33you use to manage and analyze this type
- 46:31:35of data effectively? Describe the tools
- 46:31:38and techniques you might employ. So you
- 46:31:40can start answering this question by
- 46:31:42describing that dealing with
- 46:31:44unstructured data can be challenging due
- 46:31:46to its lack of predefined format or
- 46:31:48structure. However, with the right
- 46:31:49strategies and tools, you can
- 46:31:51effectively manage and analyze it to
- 46:31:54extract valuable insights and there will
- 46:31:56be a approach how you can do that. So,
- 46:31:58we will discuss the approach here and
- 46:32:00starting with the steps. So, the number
- 46:32:02one step will be data categorization and
- 46:32:04organization. So, the number one step in
- 46:32:08this step will be sorting and tagging.
- 46:32:11We will begin by categorizing the data
- 46:32:13into types that will be text, images or
- 46:32:16videos. Use tagging to add metadata
- 46:32:19which helps in organizing the data and
- 46:32:21makes it easier to access and analyze
- 46:32:23later. Then and after that particularly
- 46:32:25for text data we'll use natural language
- 46:32:28processing NLP. We will employ NLP
- 46:32:30techniques to extract useful information
- 46:32:32from text. Tools like NLTK, spacy or
- 46:32:36even more advanced models like BERT can
- 46:32:38help you perform tasks such as sentiment
- 46:32:40analysis, entity recognition and topic
- 46:32:43modeling. After that we will do text
- 46:32:46indexing. We can use elastic search or
- 46:32:48Apache sle to index large volumes of
- 46:32:51text. These tools provide powerful
- 46:32:53search capabilities and can handle
- 46:32:55complex queries efficiently. And after
- 46:32:57that we'll move to image data. And to
- 46:32:59structure image data we'll use image
- 46:33:01processing. We'll use libraries like
- 46:33:03OpenCV for basic image processing tasks
- 46:33:05such as filtering and transformations.
- 46:33:07For more advanced image analysis,
- 46:33:09consider deep learning models using
- 46:33:11frameworks like TensorFlow or PyTorch.
- 46:33:14And then we'll feature extraction. Apply
- 46:33:16techniques to extract features from
- 46:33:18images such as edges, textures or key
- 46:33:20points which can be used for further
- 46:33:22analysis or machine learning. And then
- 46:33:24we'll come to video data. And here we'll
- 46:33:26do video processing. and we'll use the
- 46:33:28tools like fmpg that can be used for
- 46:33:31basic video processing tasks such as
- 46:33:33format conversion or extracting frames
- 46:33:35for analyzing video content look at
- 46:33:37machine learning models that can
- 46:33:39classify or recognize activities in the
- 46:33:41video and after this we'll move to
- 46:33:43temporal analysis for videos temporal
- 46:33:46components are important techniques like
- 46:33:48sequence modeling or recurrent neural
- 46:33:50networks RNNs can be useful to analyze
- 46:33:53sequences of frames for activities or
- 46:33:56events and then we'll move to data
- 46:33:58storage and management. Here we'll use
- 46:34:00the given volume and complexity of
- 46:34:02unstructured data and use big data
- 46:34:04platforms like Hadoop or cloud services
- 46:34:06like AWS S3 for storage. These platforms
- 46:34:09can scale up to handle large data sizes
- 46:34:11and provide the necessary infrastructure
- 46:34:13to store and retrieve unstructured data
- 46:34:14efficiently. And then we have
- 46:34:16visualization and reporting custom
- 46:34:18dashboards that we'll create here. We
- 46:34:20will develop custom dashboards using
- 46:34:22tools like Tableau or PowerBI which can
- 46:34:24integrate different data types and
- 46:34:25provide a unified view of the analyzed
- 46:34:27data. And after that we will do data
- 46:34:29summarization. Tools that provide
- 46:34:31summarization capabilities can help in
- 46:34:33considering large volumes of
- 46:34:34unstructured data into more manageable
- 46:34:36and interpretable forms. And after that
- 46:34:39we'll leverage these strategies and
- 46:34:41tools and can effectively manage,
- 46:34:43analyze and derive insights from
- 46:34:45unstructured data which can be crucial
- 46:34:47for making informed decisions in various
- 46:34:49applications. And this is the path that
- 46:34:51you can explore and explain to the
- 46:34:53interviewer if this question has been
- 46:34:55asked. Now moving to the question number
- 46:34:5710 and that will be based on scaling AI
- 46:34:59solutions in enterprise and the question
- 46:35:01is your company wants to scale its AI
- 46:35:04operations from a few initial pilot
- 46:35:06projects to enterprisewide
- 46:35:07implementation. What are the key
- 46:35:09considerations and steps you would take
- 46:35:11to ensure the successful scaling of AI
- 46:35:13solutions across the organization and
- 46:35:15what challenges might you face and how
- 46:35:17would you address them? So you can start
- 46:35:19answering this question with the scaling
- 46:35:21AI solutions. You could answer him that
- 46:35:23scaling AI solutions across an
- 46:35:25enterprise requires careful planning and
- 46:35:27strategic implementation to ensure
- 46:35:29success and alignment with business
- 46:35:32objectives and there should be a
- 46:35:33strategic approach to implement this. So
- 46:35:36starting with the approach and the
- 46:35:38number one step will be that will be
- 46:35:40strategic alignment. So identify
- 46:35:43business objectives. Start by
- 46:35:45identifying the business objectives that
- 46:35:46the AI solutions are intended to
- 46:35:48support. This ensures that the AI
- 46:35:50initiatives are aligned with the company
- 46:35:52strategic goals and can demonstrate
- 46:35:54clear business value. And then comes the
- 46:35:57stakeholder engagement. So engage
- 46:35:59stakeholders from various departments
- 46:36:01early in the process to gather input and
- 46:36:03build support. This helps in
- 46:36:05understanding diverse needs and ensures
- 46:36:08broader acceptance of the AI solutions.
- 46:36:10And after that comes the infrastructure
- 46:36:12and technology. So there's an option
- 46:36:14that is assess and upgrade
- 46:36:16infrastructure. Evaluate whether your
- 46:36:18current IT infrastructure can support
- 46:36:20the expanded use of AI. You might need
- 46:36:22to upgrade hardware, invest in cloud
- 46:36:24solutions or adopt technologies that
- 46:36:26facilitate AI processing and data
- 46:36:28handling. And after that we have
- 46:36:30standardization of tools. Standardize
- 46:36:32the tools and platforms used for AI
- 46:36:34development to ensure compatibility and
- 46:36:37ease of maintenance across the
- 46:36:38organizations. And after that we'll move
- 46:36:40to data management. So robust data
- 46:36:42governance that is to implement a strong
- 46:36:45data governance framework to manage
- 46:36:46enterprise data effectively. This
- 46:36:48includes policies for data quality,
- 46:36:50security and compliance especially
- 46:36:52important when scaling AI solutions that
- 46:36:54rely on vast amounts of data. And after
- 46:36:56that we will come to data accessibility.
- 46:36:58So ensure that data is accessible across
- 46:37:00the organization but also secure against
- 46:37:02unauthorized access. This involves
- 46:37:04setting up secure data leaks or
- 46:37:06warehouses that centralize data while
- 46:37:08allowing controlled access. And then we
- 46:37:11come to the next step that is talent and
- 46:37:13training. So build AI competency that is
- 46:37:16develop in-house AI expertise through
- 46:37:18training programs and hiring. So this
- 46:37:20build the necessary skills within the
- 46:37:22organization to develop, manage and
- 46:37:24scale AI solutions and after that you
- 46:37:26can also perform cross functional AI
- 46:37:28teams that could be forming cross
- 46:37:30functional teams that include data
- 46:37:32scientists, IT professionals and domain
- 46:37:34experts. So this fosters collaboration
- 46:37:36and ensure that AI solutions are
- 46:37:38developed with a comprehensive
- 46:37:39understanding. And after forming these
- 46:37:41collaborative teams, we move to scalable
- 46:37:43deployment models. So pilot test and
- 46:37:46phase roll out. Before a full-scale
- 46:37:48rollout, conduct pilot test to go the AI
- 46:37:51solution effectiveness and integration
- 46:37:53capabilities based on feedback, adjust
- 46:37:56and then gradually deploy the solutions
- 46:37:58across the organization. And then we
- 46:38:00have modular and flexible design. So
- 46:38:02design AI systems to be modular and
- 46:38:05scalable allowing for adjustments and
- 46:38:07expansions as needs and then we'll
- 46:38:09monitor and do the continuous
- 46:38:11improvement. So there will be
- 46:38:13performance metrics that would establish
- 46:38:14metrics to regularly assess the
- 46:38:16performance of AI systems. We will
- 46:38:18monitor these systems to ensure they met
- 46:38:21expected outcomes and adapt as
- 46:38:23necessary. And after that we have next
- 46:38:25step that is addressing challenges. So
- 46:38:28there could be cultural resistance that
- 46:38:30there could be employees that would be
- 46:38:32resisting to the changes but we have to
- 46:38:34address this through continuous
- 46:38:36education and by showcasing successful
- 46:38:38AI use cases within the organizations
- 46:38:41and by carefully considering these
- 46:38:42aspects and methodically implementing
- 46:38:45steps you can successfully scale AI
- 46:38:47solutions across your enterprise driving
- 46:38:49significant business value and
- 46:38:50innovation. And that's all for question
- 46:38:53number 10. Now we'll move to question
- 46:38:55number 11 and that is based on ethical
- 46:38:57considerations in data science. So the
- 46:38:59question is in your data science
- 46:39:01projects how do you ensure that ethical
- 46:39:03considerations are addressed? Describe
- 46:39:05the steps you take to identify and
- 46:39:07mitigate ethical risk in your projects.
- 46:39:09What frameworks or guidelines do you
- 46:39:11follow? So you could start answering
- 46:39:13this question with ethical
- 46:39:14considerations that they're crucial in
- 46:39:16data science to ensure that the
- 46:39:17solutions and analyzes do not
- 46:39:20advertently cause harm or bias. Here's
- 46:39:23how you can ensure that. So there are
- 46:39:25some steps and we will discuss those
- 46:39:27steps. Starting with the number one that
- 46:39:30is educate on ethical standards. So stay
- 46:39:32informed about the ethical standards in
- 46:39:34data science such as fairness,
- 46:39:36accountability, transparency and
- 46:39:38privacy. Organizations like the data
- 46:39:40science association and the ACM have
- 46:39:42codes of ethics that we refer to as
- 46:39:44guidelines. And then we have ethical
- 46:39:46risk assessment. Identify potential
- 46:39:48ethical issues. That would be at the
- 46:39:50beginning of each project. Conduct a
- 46:39:52thorough assessment to identify any
- 46:39:54potential ethical risk such as biases in
- 46:39:57data or impact on vulnerable groups.
- 46:40:00This involve reviewing the source of
- 46:40:02data, the methodologies used for data
- 46:40:04collection and the intended use of the
- 46:40:06data analytics results. And then we have
- 46:40:09stakeholder analysis. Engage with
- 46:40:11stakeholders to understand the diverse
- 46:40:13perspectives and potential impact of the
- 46:40:16project. This helps in identifying
- 46:40:18ethical issues that may not be apparent
- 46:40:20from a purely technical standpoint. And
- 46:40:23then we'll move to mitigation
- 46:40:24strategies. Implementing bias mitigation
- 46:40:27techniques. We will use statistical and
- 46:40:29machine learning techniques to detect
- 46:40:31and mitigate biases in data. This might
- 46:40:34involve techniques like resampling,
- 46:40:35reeing or using algorithms designed to
- 46:40:38be fair. And then we have privacy
- 46:40:40preserving methods. Employ methods such
- 46:40:42as data anonymization, encryption or
- 46:40:44differential privacy to protect
- 46:40:46individual privacy when analyzing
- 46:40:48sensitive data. Then we have other
- 46:40:50methods that is transparency and
- 46:40:52explanability. There we have model
- 46:40:54explanability and after that coming to
- 46:40:56documentation and reporting. So we have
- 46:40:59to maintain thorough documentation of
- 46:41:00data sources, model decisions and
- 46:41:03methodologies. And then we have
- 46:41:04continuous monitoring and feedback.
- 46:41:06There you have to monitor outcomes and
- 46:41:08the feedback mechanisms should be
- 46:41:10applied. And then we have the panels
- 46:41:12that is collaboration and advisory
- 46:41:14panels. Then we have ethical review
- 46:41:16boards. So for complex projects setting
- 46:41:19up or consulting with an ethical review
- 46:41:21board can provide oversight and diverse
- 46:41:24perspectives on the ethical implications
- 46:41:26of project methodologies. So by
- 46:41:28proactively addressing ethical
- 46:41:29considerations through these steps you
- 46:41:32can ensure that your data science
- 46:41:33projects uphold high ethical standards
- 46:41:35and positively contribute to society
- 46:41:38while minimizing harm. So this was all
- 46:41:40about question 11. Now moving to
- 46:41:42question number 12 that is based on time
- 46:41:44series forecasting for business
- 46:41:46decisions. So the question number 12 is
- 46:41:48you are tasked with forecasting monthly
- 46:41:50sales for a retail company using time
- 46:41:53series data from the past 5 years. What
- 46:41:56steps would you take to prepare and
- 46:41:58analyze this data to make accurate
- 46:42:00forecast? What specific tools or
- 46:42:02techniques would you use and why? So we
- 46:42:04can start answering this by time series
- 46:42:06forecasting and we could address them
- 46:42:08that it's a powerful tool for predicting
- 46:42:10future events based on past data
- 46:42:12especially in business context like
- 46:42:14retail sales. So we will have a
- 46:42:16structured approach here and we'll start
- 46:42:18with data collection and cleaning.
- 46:42:20First, you will gather data and ensure
- 46:42:22that you have collected all relevant
- 46:42:23data including monthly sale figures from
- 46:42:25the past five years and also considering
- 46:42:27including external factors that might
- 46:42:29affect sales such as economic
- 46:42:31indicators, holidays and promotional
- 46:42:33activities. And then we'll proceed to
- 46:42:35clean data. We will check for and handle
- 46:42:37any inconsistencies or missing values.
- 46:42:39And then we have data visualization.
- 46:42:41Here we will plot the data. We'll use
- 46:42:43plotting libraries like Matt Lib or
- 46:42:45Seabbone in Python to visualize the
- 46:42:47data. This will help in identifying
- 46:42:49patterns, trends and seasonality. And
- 46:42:51then we have decomposition of data. So
- 46:42:53there's a seasonal decomposition and
- 46:42:55we'll use statistical techniques to
- 46:42:57decompose the data into trend seasonally
- 46:43:00and residuals. So this can be
- 46:43:02accomplished with tools like the
- 46:43:04seasonal decompose function from the
- 46:43:06stats models library in Python. And
- 46:43:08we'll understand these components
- 46:43:09separately and can improve the accuracy
- 46:43:12of our forecast. And then the next step
- 46:43:14is model selection and forecasting. So
- 46:43:17there are two models that is a ara and s
- 46:43:20IMA models. So we have to choose
- 46:43:22appropriate forecasting models based on
- 46:43:24data's characteristics. For instance,
- 46:43:26AMA that is auto reggressive integrated
- 46:43:29moving average. It is effective for
- 46:43:31non-season data while SMA that is
- 46:43:34seasonal AMA that is suitable for data
- 46:43:37with seasonal patterns. And after
- 46:43:39choosing the model we'll move to cross
- 46:43:41validation. We will implement time
- 46:43:43series specific cross validation
- 46:43:45techniques like timebased splitting to
- 46:43:47evaluate model performance and this will
- 46:43:49ensure your model generalizes well on
- 46:43:51unseen data and then we have model
- 46:43:53fitting and diagnostics. We will fit the
- 46:43:55model that is by using the cinemax class
- 46:43:58from stat models that will fit your
- 46:44:00model to the data and then we will
- 46:44:02carefully select parameters based on AIC
- 46:44:04that is a cake information criterion
- 46:44:07that scores or thorough grid research
- 46:44:09technique and then we can do the
- 46:44:11diagnostics and forecast and validation
- 46:44:14and after forecast validation we'll move
- 46:44:16to iterative improvement. So there's a
- 46:44:19feedback loop that should be mandatory
- 46:44:21and there should be a regular update for
- 46:44:24the model with new sales data and
- 46:44:25refining the model as needed. So this
- 46:44:28continuous improvement cycle helps adapt
- 46:44:30to changing patterns in sales data. And
- 46:44:33by following these steps and using these
- 46:44:35tools, you can create robust forecast
- 46:44:37that help the retail company plan better
- 46:44:39and make informed decisions. So this was
- 46:44:42all about question number 12. Now move
- 46:44:43to question number 13 that is based on
- 46:44:46customer segmentation using machine
- 46:44:48learning. So the question is you are
- 46:44:50given a data set containing demographic
- 46:44:52and purchasing behavior data for a group
- 46:44:54of customers. Your task is to segment
- 46:44:56these customers into distinct groups
- 46:44:58based on similarities in the purchasing
- 46:45:00behavior and demographics. So what steps
- 46:45:02would you take to perform this
- 46:45:04segmentation and can you provide a
- 46:45:05sample Python code snippet to illustrate
- 46:45:08the initial stages of data handling and
- 46:45:10model application. So we can start this
- 46:45:12by explaining customer segmentation that
- 46:45:14it's a powerful approach to tailor
- 46:45:16marketing strategies and improve
- 46:45:18customer service by identifying distinct
- 46:45:20groups based on their behavior and
- 46:45:22characteristics. And here also we have a
- 46:45:25detailed approach for this task. So
- 46:45:27we'll start with number one step that
- 46:45:29would be data exploration and
- 46:45:31pre-processing. So there will be initial
- 46:45:33exploration that is beginning by
- 46:45:35examining the data set to understand the
- 46:45:37features available such as age, income,
- 46:45:40purchase frequency etc. Then we'll look
- 46:45:42for missing values or anomalies and
- 46:45:45decide how to handle them. That could be
- 46:45:47using imputation and then we'll move to
- 46:45:49feature engineering. We'll create new
- 46:45:51features that might be useful for
- 46:45:52segmentation such as customer lifetime
- 46:45:55value or average transaction amount.
- 46:45:57We'll also use normalization that is
- 46:45:59normalize the data to ensure that one
- 46:46:01feature doesn't disproportionately
- 46:46:03influence the model due to its scale.
- 46:46:05We'll use standard scaling or minmax
- 46:46:07scaling as appropriate. So then we'll
- 46:46:10come to the next step that is choosing
- 46:46:11the segmentation technique. And here we
- 46:46:13have k means clustering. So this is a
- 46:46:15popular method for customer
- 46:46:16segmentation. Here we will decide on the
- 46:46:18number of clusters by using techniques
- 46:46:20like the elbow method or analysis to
- 46:46:23determine the optimal cluster count. And
- 46:46:26then we have model implementation and in
- 46:46:29that we will use data preparation and
- 46:46:31we'll prepare the data by selecting the
- 46:46:33relevant features and applying any final
- 46:46:35transformations and then we have model
- 46:46:37fitting. We fit the C means clustering
- 46:46:39model to the data and evaluate and
- 46:46:42interpret analyzing clusters and after
- 46:46:45analyzing clusters we'll move to the
- 46:46:47next step that is strategic insights. We
- 46:46:50will provide actionable insights based
- 46:46:52on cluster characteristics such as
- 46:46:54targeted marketing strategies for each
- 46:46:56segment. And then we have iterative
- 46:46:59refinement that is feedback
- 46:47:01incorporation and we'll use business
- 46:47:03feedback to refine the segmentation. If
- 46:47:06additional data becomes available
- 46:47:07incorporated to enhance the model and
- 46:47:09now we'll see the sample Python code. So
- 46:47:12for this first we'll import the
- 46:47:14libraries and modules. As you can see on
- 46:47:16the screen we have imported pandas
- 46:47:19random forest classifier train test
- 46:47:21split standard scaler classification
- 46:47:24report and after that we will load the
- 46:47:26data and and for that we have used the
- 46:47:28pandas to read the data that is read ssv
- 46:47:33and after that we are processing the
- 46:47:34data that is data prep-processing we are
- 46:47:37handling missing values and using the
- 46:47:39forward fill or fil to fill missing
- 46:47:42values in the data set and then we are
- 46:47:44featuring scaling that is normalizing
- 46:47:46the selected features that is feature
- 46:47:48one, feature two and feature three using
- 46:47:50standard scaler and then we'll move to
- 46:47:52the next step that is data splitting.
- 46:47:54We'll split the data set into training
- 46:47:56and testing sets. So that test size
- 46:47:58equal to 0.2 parameters specifies that
- 46:48:0120% of the data will be used for testing
- 46:48:03and then we'll train the model. We'll
- 46:48:05initialize and train a random forest
- 46:48:07classifier with 100 trees and a random
- 46:48:10state for reproductibility and then
- 46:48:12we'll evaluate the model. will make
- 46:48:14predictions on the test set that is X
- 46:48:18test using the train model and print a
- 46:48:20classification report showing precision
- 46:48:23recall F1 score and support for each
- 46:48:26class. So this code demonstrates the
- 46:48:28process of loading, pre-processing,
- 46:48:31training and evaluating a machine
- 46:48:33learning model that is random forest
- 46:48:34classifier for predicting equipment
- 46:48:37failures in a manufacturing plant. The
- 46:48:39use of techniques such as data
- 46:48:40prep-processing and splitting along with
- 46:48:42the random forest classifier highlights
- 46:48:44a standard flow for building predictive
- 46:48:46maintenance models. So this was all
- 46:48:48about the question number 13. So now
- 46:48:50moving to the question number 14 that is
- 46:48:52based on predictive customer churn and
- 46:48:54the question is you are tasked with
- 46:48:56developing a model to predict which
- 46:48:58customers are likely to churn from a
- 46:49:00subscription service. So what steps
- 46:49:02would you take to build this model and
- 46:49:04can you provide a sample Python code to
- 46:49:06illustrate the data preparation and
- 46:49:08model training process? So we'll start
- 46:49:09answering this question about depicting
- 46:49:12what is predicting customer churn. So
- 46:49:14predicting customer churn is crucial for
- 46:49:16businesses to implement detention
- 46:49:18strategies proactively and we'll have a
- 46:49:21detailed approach for building a
- 46:49:22predictive model for this purpose.
- 46:49:25Starting with data collection and
- 46:49:26exploration and in this we will collect
- 46:49:29data and after that we'll perform the
- 46:49:31exploratory data analysis that is EDA.
- 46:49:34We'll perform an initial analysis to
- 46:49:36understand patterns and trends and then
- 46:49:38we have feature engineering. We will
- 46:49:40create new features and derive new
- 46:49:42feature that might influence churn such
- 46:49:44as change in usage pattern or service
- 46:49:47upgrades. And then we'll handle the
- 46:49:48missing values if we found any. And then
- 46:49:50we'll encode categoral variables. We'll
- 46:49:53use techniques like one hot encoding or
- 46:49:55label encoding for categorial variables.
- 46:49:58And then we have scale features to
- 46:50:00normalize or standardize numerical
- 46:50:02features to ensure they contribute
- 46:50:04equally to the model's performance. And
- 46:50:06then we'll select the model that is
- 46:50:09we'll choose the appropriate model and
- 46:50:11start with for the knowing handling
- 46:50:13binary classification task that could be
- 46:50:16with logistic regression, random forest
- 46:50:18or gradient boosting machines. And after
- 46:50:21selecting the model, we'll train the
- 46:50:22model and evaluate it. So fit your model
- 46:50:25on the training data and after that
- 46:50:27evaluate the model using appropriate
- 46:50:29metrics like accuracy, precision,
- 46:50:31recall, F1 score and ROC to go its
- 46:50:35performance. And then we'll optimize the
- 46:50:37model using hyperparameter tuning. We'll
- 46:50:40optimize the model parameter using grid
- 46:50:42search or random search to improve
- 46:50:44performance. And then we have feature
- 46:50:46importance that is analyze and rank
- 46:50:48features by their importance in
- 46:50:50predicting churn to refine the model
- 46:50:52further. And then and then the last step
- 46:50:54is deployment and monitoring. We'll
- 46:50:56deploy the model once validated deploy
- 46:50:58the model into a production environment
- 46:51:00where it can predict real-time churn. So
- 46:51:02after deploying the model regularly
- 46:51:04monitor the model to ensure it remains
- 46:51:06effective over time as new data comes
- 46:51:08in. So now we'll see the sample Python
- 46:51:11code for this example. So starting with
- 46:51:13the importing of libraries we will
- 46:51:16import pandas numpy scikitlearn skarn
- 46:51:19tensorflow and the tensorflow kas and
- 46:51:23callbacks and after importing the
- 46:51:25modules we'll start with data loading
- 46:51:27we'll load the data set from a CSV file
- 46:51:30named equipment data dot csv and that
- 46:51:33with the pandas data frame and after
- 46:51:36that we'll do the data prep-processing
- 46:51:37we'll handle missing values and for that
- 46:51:40we'll use forward fill to fill missing
- 46:51:42values in the data set and then we have
- 46:51:44feature scaling that will normalize the
- 46:51:46selected features that is feature one,
- 46:51:48feature two, feature three using
- 46:51:49standard scaler and after that we'll use
- 46:51:52the data splitting. We'll split the data
- 46:51:54set into training and testing sets and
- 46:51:56the test size will be equal to 0.2 and
- 46:51:59this parameter specifies that 20% of the
- 46:52:01data will be used for testing and after
- 46:52:03that we'll start with building the
- 46:52:05model. First we'll see sequential model
- 46:52:07that initializes a sequential model
- 46:52:10technique. And then we have dense layers
- 46:52:12that adds two dense layers with 64 units
- 46:52:15and value activation function. Then we
- 46:52:17have dropout layers that adds two
- 46:52:19dropout layers with a dropout rate of
- 46:52:210.5 to reduce overfitting. After that
- 46:52:24we'll do the model compilation. We'll
- 46:52:26compile the model using the atom
- 46:52:27optimizer and binary cross entropy loss
- 46:52:30function for binary classification. And
- 46:52:33there will be an early stopping that
- 46:52:34will define an early stopping call back
- 46:52:36to stop training when the validation
- 46:52:38loss metric has stopped improving after
- 46:52:40three blocks. And after training the
- 46:52:43model, we will evaluate the model. And
- 46:52:45evaluating the model on the test data
- 46:52:47and print the loss and accuracy metrics.
- 46:52:50So this code demonstrates the process of
- 46:52:52loading, pre-processing, building,
- 46:52:53compiling, training and evaluating a
- 46:52:55deep learning model using TensorFlow and
- 46:52:58KAS for predicting equipment failures in
- 46:53:00a manufacturing plant. So the use of
- 46:53:02techniques such as data prep-processing,
- 46:53:04dropout regularization and early
- 46:53:07stopping helps in building a robust deep
- 46:53:09learning model for predictive
- 46:53:10maintenance. So that's all with question
- 46:53:12number 14. Now we'll start with question
- 46:53:14number 15 that is based on deep learning
- 46:53:16and NLP. And your question is you are
- 46:53:18tasked with developing a sentiment
- 46:53:20analysis model using deep learning to
- 46:53:22understand customer opinions from
- 46:53:24reviews. So what steps would you take to
- 46:53:26build this model and can you provide a
- 46:53:28sample Python code snippet to illustrate
- 46:53:30how you would pre-process data and train
- 46:53:32a simple deep learning model? So we'll
- 46:53:34start answering this with sentiment
- 46:53:36analysis that sentiment analysis using
- 46:53:38deep learning allows businesses to coach
- 46:53:41customer sentiment from text data like
- 46:53:43reviews or comments effectively. And
- 46:53:45we'll have a detailed approach for
- 46:53:47building a sentiment analysis model.
- 46:53:49We'll start with data collection and
- 46:53:51cleaning. We will collect the data,
- 46:53:52gather a substantial data set of text
- 46:53:54reviews and their associated sentiments
- 46:53:57typically labeled as positive, negative,
- 46:53:59or neural. And then we'll clean the
- 46:54:01data, pre-process the data by removing
- 46:54:04noise such as HTML tags, special
- 46:54:06characters, and so words. And we'll
- 46:54:08normalize the text by converting it to
- 46:54:10lower case. And then we have text
- 46:54:12prep-processing. We'll convert text into
- 46:54:14tokens, words, or phrases. And then we
- 46:54:17have vectorization that transforms
- 46:54:19tokens into numerical format using
- 46:54:21techniques like word embeddings or TF
- 46:54:23that is term frequency in document
- 46:54:26frequency and then we'll use the padding
- 46:54:29and then we have the option of model
- 46:54:30selection. will choose a model
- 46:54:32architecture based on a basic approach
- 46:54:34and use a RNN or more advanced
- 46:54:37architecture like LSTM that is long
- 46:54:40short-term memory or GRU that is gated
- 46:54:43recurrent units which are effective for
- 46:54:45sequence data like text and then we have
- 46:54:48model training we'll compile the model
- 46:54:51define the model architecture and
- 46:54:52compile it with a loss function suited
- 46:54:55for classification like categoral cross
- 46:54:57entropy and an optimizer like Adam and
- 46:55:00then we'll train the model. We'll fit
- 46:55:02the model on our pre-processed data.
- 46:55:04We'll evaluate and optimize it.
- 46:55:06Evaluating model performance. Here use
- 46:55:08the metrics such as accuracy, precision,
- 46:55:11recall, and F1 score to assess the
- 46:55:13model. And then we have hyperparameter
- 46:55:15tuning. We'll optimize the model by
- 46:55:17adjusting parameters like learning rate,
- 46:55:19number of layers and units per layer.
- 46:55:22And then coming to deployment. We'll
- 46:55:24deploy the model and integrate the model
- 46:55:26into the existing review processing
- 46:55:28pipeline. So it can automatically
- 46:55:30classify new reviews. So let's see the
- 46:55:32sample Python code and we'll have a
- 46:55:34basic approach for that. Here we'll
- 46:55:36import numpy tensorflow sequential
- 46:55:39embedding LSTM dense stroke out. So
- 46:55:42embedding converts positive integers
- 46:55:44that is indexes into dense vectors of
- 46:55:46fixed size and LSTM that is long
- 46:55:48short-term memory layer that is used for
- 46:55:51learning dependencies in sequence data.
- 46:55:54And then we have dense that is a
- 46:55:55regularly densed connected NN layer. And
- 46:55:59then we will import pad sequences. And
- 46:56:02after that we have the data set and the
- 46:56:04sample text data representing customer
- 46:56:06reviews that will store in variable
- 46:56:09text. And then we have labels that has
- 46:56:11binary labels indicating sentiment one
- 46:56:13for positive zero for negative. And now
- 46:56:16we'll start with the pre-processing of
- 46:56:17data. Here we have declared that
- 46:56:19tokenizer. We will initialize a
- 46:56:21tokenizer that will help only the top
- 46:56:24thousand most frequent words. And then
- 46:56:26we have fit_on
- 46:56:29text that is update the internal
- 46:56:31vocabulary based on the list of text. It
- 46:56:34essentially creates a dictionary of word
- 46:56:36to index pairs. And then we have text to
- 46:56:38sequences that will transform each text
- 46:56:41in text to a sequence of integers. And
- 46:56:44then we have pad sequences that will
- 46:56:46ensure all sequences have the same
- 46:56:48length by padding shorter sequences with
- 46:56:50zeros up to the maximum length. And then
- 46:56:53we'll start building the model. Here we
- 46:56:55have sequential model that will set up a
- 46:56:57linear stack of layers. And then we have
- 46:56:59embedding layer that will map each word
- 46:57:01index to an embedding vector of size 64.
- 46:57:04So the input length is set to 10 that is
- 46:57:06the length of the input sequences. Then
- 46:57:09we'll start with LSTM layers. So two
- 46:57:11LSTM layers are added. The first one
- 46:57:13returns sequences to allow the next LSTM
- 46:57:16layer to process these sequences. And
- 46:57:18after that we have the dropout layer
- 46:57:20that applies dropout with a rate of 0.5
- 46:57:23of the first LSTM layer to reduce
- 46:57:25overfitting. And after that we'll come
- 46:57:27to dense layer that has output of a
- 46:57:30single scalar that represents the
- 46:57:32predicted setment and using sigmoid
- 46:57:35activation to output a probability. And
- 46:57:38now we'll start with model compilation
- 46:57:40and training. So we'll configure the
- 46:57:42model for training and we'll use binary
- 46:57:44cross entropy as the loss function that
- 46:57:46is suitable for binary classification
- 46:57:48and the atom optimizer and tracks
- 46:57:51constantly accuracy as a metric and then
- 46:57:54we have the fit that trains the model
- 46:57:56for a specified number of epochs that is
- 46:57:58iterations or the entire data set and
- 46:58:01then we'll predict the model that is
- 46:58:03after training the model can predict the
- 46:58:05sentiment of the reviews in the data
- 46:58:06set. This is useful for checking how the
- 46:58:08model performs on the training data
- 46:58:10itself. So this breakdown explains each
- 46:58:12step of the coding process detailing how
- 46:58:14the data is prepared and how the model
- 46:58:16is configured and then we'll compile it
- 46:58:19and use for training and prediction. So
- 46:58:21it's detailed explanation should help in
- 46:58:24understanding how to implement a simple
- 46:58:26LSTM model for sentiment analysis in
- 46:58:28TensorFlow. Now moving to the question
- 46:58:30number 16. So let's start with question
- 46:58:33number 16 that is based on anomly
- 46:58:35detection in transaction data. So the
- 46:58:37question is you are tasked with
- 46:58:39identifying unusual transactions in a
- 46:58:41company's financial data that might
- 46:58:43suggest fraudulent activity. So what
- 46:58:45steps would you take to develop an
- 46:58:46anomaly detection model and can you
- 46:58:48provide a sample Python code snippet to
- 46:58:51illustrate how you would pre-process the
- 46:58:52data and apply an anomaly detection
- 46:58:54technique. So we'll start answering this
- 46:58:56with anomaly detection technique that is
- 46:58:59anomaly detection is essential for
- 46:59:00preventing fraud by identifying
- 46:59:02transactions that deviate significantly
- 46:59:05from typical patterns. And now we'll see
- 46:59:07the structured approach to building an
- 46:59:09anomaly detection model for transaction
- 46:59:11data. We'll start with data collection
- 46:59:14and cleaning and we'll collect all the
- 46:59:16compiling transaction data which should
- 46:59:18include details like transaction amount,
- 46:59:20time, user ID and transaction type. Then
- 46:59:22we'll move to feature engineering and
- 46:59:24develop features that capture the
- 46:59:25essence of transaction such as time of
- 46:59:27day and the day of the week. And then we
- 46:59:30have data normalization. We'll use
- 46:59:31scaling techniques such as minmax
- 46:59:33scaling or standardization to ensure
- 46:59:35that the model is perfectly normalized.
- 46:59:38And then we have choosing the anomaly
- 46:59:40detection technique. So here we have to
- 46:59:43choose the technique which is effective
- 46:59:45for highdimensional data sets and works
- 46:59:47for isolating anomalies instead of
- 46:59:50profiling normal data points. After
- 46:59:52choosing the anomaly technique will
- 46:59:54train anomaly identification. We'll fit
- 46:59:57the chosen model to the data and the
- 46:59:59anomalies that would have been chosen
- 47:00:01will be those transactions that the
- 47:00:03model identifies and after this we come
- 47:00:05to the last step that is review and
- 47:00:07action. Here we have manual review that
- 47:00:09is transactions flagged as potential
- 47:00:11anomalies should be reviewed manually to
- 47:00:13confirm fraudent activity and then we
- 47:00:16have continuous improvement that is we
- 47:00:18can regularly update the model with the
- 47:00:19new data and feedback from the review
- 47:00:21process to improve accuracy. And now
- 47:00:23moving to the prediction that is after
- 47:00:25training the model we can predict the
- 47:00:27sentiment of the reviews in the data set
- 47:00:29and this is useful for checking how the
- 47:00:31model performs on the training data
- 47:00:33itself. Now we'll see the Python code to
- 47:00:36see how you can set up this model for
- 47:00:39anomaly detection. We'll start by
- 47:00:41importing the libraries and module and
- 47:00:43after that we'll load and prepare data.
- 47:00:45That is we'll load transaction data from
- 47:00:47a CSV file into the pandas data frame.
- 47:00:50And after that we'll convert the
- 47:00:51transaction time column to date time
- 47:00:53format which allows the extraction of
- 47:00:55additional time based features. And
- 47:00:58after that we'll perform feature
- 47:00:59engineering that will extract the hour
- 47:01:01of the day from the transaction time
- 47:01:03column. This feature can be important as
- 47:01:05transactions occurring at unusual hours
- 47:01:07may be indicative of fraud. And then
- 47:01:09we'll move to the normalization of data.
- 47:01:12This will apply standard scaling to the
- 47:01:14amount n of the day feature. This
- 47:01:16normalization process involves
- 47:01:17subtracting the mean and dividing by the
- 47:01:19standard deviation for each feature
- 47:01:22ensuring that the feature contribute
- 47:01:23equally to the analysis and improving
- 47:01:25the performance of many machine learning
- 47:01:27algorithms. And after that we'll start
- 47:01:29with anomaly detection with isolation
- 47:01:31forest. That's a technique. We'll
- 47:01:34initialize an isolation forest model
- 47:01:36with 100 trees that is n estimators
- 47:01:39equal to 100. Setting the proportions of
- 47:01:41outliers that is contamination to 1% of
- 47:01:44the data. So this parameter is crucial
- 47:01:46as it influences the threshold of
- 47:01:48marking an observation as an anomaly.
- 47:01:51Then we'll fit the model to the scaled
- 47:01:52amount n of the data and predict the
- 47:01:55anomaly status for each transaction. And
- 47:01:58then we'll start with filter and display
- 47:02:00anomalies. We'll filter out transactions
- 47:02:02identified as anomalies that is anomaly
- 47:02:05equal equal to minus one. We'll display
- 47:02:07these transactions which can be reviewed
- 47:02:09manually to determine if they represent
- 47:02:12actual fraud net activity. So this code
- 47:02:14snippet provides a systematic approach
- 47:02:16to detecting anomalies in transaction
- 47:02:18data leveraging the isolation forest
- 47:02:20algorithms ability to handle complex and
- 47:02:23highdimensional data set effectively. So
- 47:02:25the pre-processing steps ensured that
- 47:02:27the data is appropriately formatted and
- 47:02:29normalized for optional model
- 47:02:30performance. So this was all about
- 47:02:32question number 16. Now moving to
- 47:02:34question number 17 and that is based on
- 47:02:36integrating machine learning models into
- 47:02:38web applications. And your question is,
- 47:02:40you have developed a machine learning
- 47:02:42model to predict real estate prices
- 47:02:44based on various features like location,
- 47:02:46size, and amenities. How would you
- 47:02:48integrate this model into a web
- 47:02:50application to allow users to get
- 47:02:52real-time price predictions? Can you
- 47:02:54provide a sample Python code snippet to
- 47:02:56illustrate how you would prepare the
- 47:02:57model for integration and handle user
- 47:03:00request? So, starting with the approach
- 47:03:03that is integrating a machine learning
- 47:03:04model into a web application. This will
- 47:03:07involve several steps to ensure the
- 47:03:09model is accessible and perform well in
- 47:03:11a live environment. So here's how you
- 47:03:13can approach this task. We could divide
- 47:03:15into steps and we'll start with number
- 47:03:17one step that is model preparation.
- 47:03:19We'll finalize and save the model. So
- 47:03:22once your model is trained and
- 47:03:23validated, save it using a format that
- 47:03:26can be easily loaded into a web
- 47:03:28application. So Python's pickle module
- 47:03:30or TensorFlow's save model format are
- 47:03:32commonly used for this purpose. Then we
- 47:03:35can use web application backend setup.
- 47:03:37For this, select a suitable web
- 47:03:39framework. So, Flask is popularly known
- 47:03:42for its simplicity and effectiveness in
- 47:03:44integrating Python based machine
- 47:03:46learning models. And after that, we'll
- 47:03:48develop the API. After developing the
- 47:03:51API within your Flask app that you can
- 47:03:53receive user inputs for model features,
- 47:03:56load the model, make prediction, and
- 47:03:58return the result. And after this, we'll
- 47:04:00develop the UI. We'll design a
- 47:04:02user-friendly interface. We'll create a
- 47:04:04simple and intuitive UI that lets users
- 47:04:07input the feature like location, size
- 47:04:09and submit them for prediction. And
- 47:04:12after that we'll move to the deployment
- 47:04:13phase. We'll use a cloud platform like
- 47:04:15Heroku, AWS or Google Cloud to deploy
- 47:04:18your Flask application. And then we have
- 47:04:20the maintenance and updates. We'll
- 47:04:22monitor and update regularly for the
- 47:04:24performance and use the model as needed
- 47:04:27based on user feedback. So now moving to
- 47:04:30the Python code and see how this model
- 47:04:33can be created. So here we'll start
- 47:04:35importing the libraries and module and
- 47:04:37we are using flask ple and jsonify and
- 47:04:41we will start with app initialization.
- 47:04:43We'll initialize a new flask web
- 47:04:45application that would be a special
- 47:04:47variable which gives python files a
- 47:04:50unique name to differentiate between
- 47:04:52them when they are important into other
- 47:04:54scripts. And after that we'll load the
- 47:04:56model. So loading a pretend machine
- 47:04:58learning model from the file system. So
- 47:05:00this model is assumed to be saved in the
- 47:05:02same directory as this script. So the
- 47:05:04model is loaded in RB mode which stands
- 47:05:07for read binary. And after that we'll
- 47:05:09move to API route and prediction
- 47:05:11function. So we will define an API
- 47:05:14endpoint at predict that listens for
- 47:05:17post request. This is the URL that the
- 47:05:20front end of the web application will
- 47:05:21call to send data to the back end. And
- 47:05:24after that we'll start with predicting
- 47:05:25the function. And here we have extract
- 47:05:27features that retrieves data sent into
- 47:05:30the JSON format from the post request
- 47:05:32that is request get_json and the force
- 47:05:35we have set it as true here and
- 47:05:38forcefully formats the request data into
- 47:05:40JSON ensuring compatibility and then
- 47:05:43we'll extract the relevant features that
- 47:05:44is location size and amenities from the
- 47:05:47JSON object and store them in a list as
- 47:05:49expected by the model and after
- 47:05:51preparing the features we'll make the
- 47:05:53prediction we'll use the loaded model to
- 47:05:55make a prediction based bas on the
- 47:05:56provided feature and then we have the
- 47:05:58return prediction method. Here we will
- 47:06:00convert the prediction result into JSON
- 47:06:02format using JSON and send it back to
- 47:06:05the client and this will ensure that the
- 47:06:06response can be easily handled by the
- 47:06:08client application. So this was all
- 47:06:10about the question number 17. Now moving
- 47:06:12to the question number 18 that is based
- 47:06:14on analyzing. And now we move to the
- 47:06:16question number 18 that is based on
- 47:06:18analyzing geospatial data. And your
- 47:06:20question is you are tasked with
- 47:06:21analyzing geospatial data to help a city
- 47:06:24improve its public transportation
- 47:06:25system. The data includes GPS
- 47:06:27coordinates of bus stops, ridership
- 47:06:30numbers and traffic patterns. What steps
- 47:06:32would you take to analyze this data? And
- 47:06:34can you provide a sample Python code
- 47:06:35snippet to illustrate how you might
- 47:06:37visualize bus stop location and
- 47:06:39ridership? So you can start answering
- 47:06:41this question that juice better data
- 47:06:44analysis can provide critical insights
- 47:06:46into how effectively a public
- 47:06:47transportation system serves its city
- 47:06:49and guide improvements and there's an
- 47:06:52detailed approach for this and we can
- 47:06:54start with data preparation and in this
- 47:06:56we'll do data collection and data
- 47:06:58cleaning and after this step we'll move
- 47:07:00to the next step that is explorative
- 47:07:02data analysis and in this we'll have
- 47:07:04statistical summary we'll generate
- 47:07:06descriptive statistics and then we have
- 47:07:08correlation analysis
- 47:07:10And after moving that we have geospatial
- 47:07:13visualization that is mapping bus stop.
- 47:07:16We'll plot the locations of bus stop on
- 47:07:17a map to visually assess their
- 47:07:19distribution across the city. And after
- 47:07:22that we have heat maps that will create
- 47:07:24ridership data to identify hot sports
- 47:07:26and areas with potential service gaps.
- 47:07:29And after geospatial visualization we'll
- 47:07:31move with spatial analysis. We have
- 47:07:33proximity analysis that will analyze the
- 47:07:36proximity of bus stop to key areas like
- 47:07:38commercial centers or residential areas.
- 47:07:41And now moving to the fifth step that is
- 47:07:43optimization and recommendation. So
- 47:07:45we'll have a route optimization that
- 47:07:47will suggest modifications to route
- 47:07:50based on traffic patterns and ridership
- 47:07:52demand and the policy recommendations
- 47:07:54that will provide actionable
- 47:07:55recommendations for improving bus
- 47:07:57frequencies. Now move to the sample
- 47:07:59Python code where we can define this
- 47:08:02model and use it accordingly. And here
- 47:08:05we will start importing the libraries
- 47:08:06and modules. And here we'll start with
- 47:08:09importing geopandas and m lib dotpipo.
- 47:08:14And after importing we'll start with
- 47:08:16data loading. So we will declare a
- 47:08:19variable bus stops and load the bus
- 47:08:22stops data from a shape file. So shape
- 47:08:25files are popular geospatial vector data
- 47:08:27formats for geographic information
- 47:08:29system software and then we have the
- 47:08:31wrership that will load wrership data
- 47:08:33from a CSV file which includes columns
- 47:08:35for longitude latitude and ridership
- 47:08:38levels and after that we'll create geo
- 47:08:40data frame that will convert the
- 47:08:42wrership data frame into a geo data
- 47:08:44frame and this step involves creating a
- 47:08:46geometry column from the longitude and
- 47:08:48latitude columns and then we have the
- 47:08:51plotting one here we will plot the
- 47:08:53graphs that would figures and axis and
- 47:08:55create a figure for the single subplot
- 47:08:58with a specified size that is 10 + 10
- 47:09:00in. And then we have city map.plot. It
- 47:09:04is assumed that there is a base map of
- 47:09:06the city loaded as a geo data frame
- 47:09:09named city map. This is plotted first
- 47:09:12with a light gray color to serve as a
- 47:09:14background for the other layers. So this
- 47:09:16was all about question number 18. Now
- 47:09:18moving to question number 19 that is
- 47:09:20based on predictive maintenance using
- 47:09:23machine learning and the question is you
- 47:09:25are tasked with developing a predictive
- 47:09:27maintenance system for a manufacturing
- 47:09:29plant that relies heavily on automated
- 47:09:31machinery. So the data available
- 47:09:33includes machine operational parameters,
- 47:09:35maintenance history and failure
- 47:09:36incidents. What steps would you take to
- 47:09:38develop a predictive model and can you
- 47:09:40provide a sample Python code? So you can
- 47:09:42start with predictive maintenance that
- 47:09:44is essential in manufacturing as it
- 47:09:46helps prevent equipment failures
- 47:09:48reducing downtime and maintenance cost.
- 47:09:51And here you would have a detail
- 47:09:53approach or predictive model for this
- 47:09:55starting with data collection and
- 47:09:57integration. Then you can do EDA that is
- 47:09:59exploratory data analysis and then we
- 47:10:02can perform feature engineering and then
- 47:10:04move to data prep-processing task and
- 47:10:07then the selection model and training
- 47:10:09and after that we have model evaluation
- 47:10:11and deployment technique that we can do
- 47:10:13for the model using appropriate metrics
- 47:10:16such as precision, recall and F1 score.
- 47:10:18So this was all about question number
- 47:10:2019. So now move to question number 20
- 47:10:22that is based on personalization using
- 47:10:24machine learning and your question is
- 47:10:26you are tasked with developing a machine
- 47:10:28learning model to personalize content
- 47:10:30recommendations for users on a media
- 47:10:32streaming platform. The data available
- 47:10:34includes user demographic retails
- 47:10:36viewing history and ratings. So what
- 47:10:38steps would you take to build a model
- 47:10:40for personalized recommendations and can
- 47:10:42you provide a sample Python code for
- 47:10:44that? So you can start answering this
- 47:10:46with creating a personalized
- 47:10:47recommendation systems. This would be
- 47:10:49essential for engaging users by
- 47:10:51providing content that is relevant to
- 47:10:54their interest. And there will be a
- 47:10:55systematic approach or personalized
- 47:10:57content recommendation. We'll start with
- 47:11:00data collection and integration. And
- 47:11:02after that, we'll perform EDA that is
- 47:11:04explorative data analysis. And then we
- 47:11:06have feature engineering. In this we'll
- 47:11:08interact features and the temporal
- 47:11:11features. We'll include time based
- 47:11:13features to capture trends and
- 47:11:15seasonality in viewing behavior. And
- 47:11:17then we'll select the model that is by
- 47:11:19collaborative filtering and hybrid
- 47:11:21models. And then we'll train the model
- 47:11:23and validation and implement and monitor
- 47:11:26them. And after that we'll deploy the
- 47:11:28model. So let's start with beginner
- 47:11:30level questions. And number one is what
- 47:11:32is machine learning? So machine learning
- 47:11:34is a subset of artificial intelligence
- 47:11:37that involves the use of algorithms and
- 47:11:39statistical models to enable computers
- 47:11:42to perform task without explicit
- 47:11:44instructions. that is by relying on
- 47:11:46patterns and interference. And now
- 47:11:49moving to number second question that is
- 47:11:51what are the different types of machine
- 47:11:52learning. So the three main types of
- 47:11:55machine learning are number one is
- 47:11:57supervised learning and then comes
- 47:11:59unsupervised learning and then there is
- 47:12:01reinforcement learning. Now moving to
- 47:12:04next question that is third that is what
- 47:12:06is supervised learning. So supervised
- 47:12:09learning involves training a model on a
- 47:12:11label data set which means each training
- 47:12:13example is paired with an output label.
- 47:12:16The model learns to predict the output
- 47:12:18from the input data. Now moving to the
- 47:12:21fourth question that is what is
- 47:12:22unsupervised learning. So unsupervised
- 47:12:25involve training a model on data that
- 47:12:27does not have labeled responses. The
- 47:12:29model tries to learn the patterns and
- 47:12:31the structure from the input data. So
- 47:12:34guys, these are the beginner level
- 47:12:35questions and now we'll move to the
- 47:12:37fifth question that is what is
- 47:12:38reinforcement learning. So reinforcement
- 47:12:41learning is a type of machine learning
- 47:12:43where an agent learns to make decisions
- 47:12:45by performing actions and receiving
- 47:12:47rewards or penalties. The goal is to
- 47:12:50maximize the cumulative reward. So now
- 47:12:52moving to the sixth question that is
- 47:12:54what is a model in machine learning. So
- 47:12:57a model in machine learning is a
- 47:12:59mathematical representation of a real
- 47:13:01world process. It is trained on data to
- 47:13:04recognize patterns and make predictions
- 47:13:06or decisions based on new data. So now
- 47:13:09moving to seventh question that is what
- 47:13:11is overfitting? So overfitting occurs
- 47:13:13when a machine learning model performs
- 47:13:15well on the training data but poorly on
- 47:13:17new unseen data. It indicates that the
- 47:13:20model has learned the noise and details
- 47:13:22in the training data instead of the
- 47:13:24actual patterns. So now coming to
- 47:13:26question number eight that is what is
- 47:13:28underfitting? So underfitting occurs
- 47:13:30when a machine learning model is too
- 47:13:32simple to capture the underlying
- 47:13:34patterns in the data. It performs poorly
- 47:13:37on both the training data and new data.
- 47:13:39Now move to the next question that is
- 47:13:41ninth question and the question is what
- 47:13:43is a confusion matrix? So confusion
- 47:13:45matrix is a table used to evaluate the
- 47:13:48performance of a classification model.
- 47:13:50It summarizes the number of correct and
- 47:13:52incorrect predictions made by the model
- 47:13:55and that is categorized by each class.
- 47:13:58Now moving to the 10th question that is
- 47:13:59what is cross validation? So cross
- 47:14:02validation is a technique for assessing
- 47:14:04how the results of a statistical
- 47:14:06analysis will generalize to an
- 47:14:08independent data set. It involves
- 47:14:10partitioning the data into subsets.
- 47:14:13Training the model on some subsets and
- 47:14:15validating it on the remaining subsets.
- 47:14:18This was all about that is the 10th
- 47:14:21question or the overall 1 to 10
- 47:14:23questions for beginner level. Now we'll
- 47:14:25move to intermediate level and here
- 47:14:27we'll cover 10 questions. So we'll start
- 47:14:30with 11th question that is what is a ROC
- 47:14:33curve. So ROC that is receiver operating
- 47:14:37characteristic curve. It is a graphical
- 47:14:39representation of a classifier's
- 47:14:41performance across different thresholds.
- 47:14:44It plots the true positive rate that is
- 47:14:46TPR against a false positive rate that
- 47:14:49is FPR. Now moving to 12th question that
- 47:14:52is what is precision and recall. So
- 47:14:54precision is the ratio of correctly
- 47:14:56predicted positive observations to the
- 47:14:58total predicted positives and recall is
- 47:15:01the ratio of correctly predicted
- 47:15:03positive observations to all actual
- 47:15:06positives. So the formula is precision
- 47:15:08equal to TP/TP
- 47:15:11plus FP and the recall is TP/TP
- 47:15:15+ F_sub_1. So now we'll move to the 13th
- 47:15:19question that is what is the F1 score?
- 47:15:22So the F1 score is the harmonic mean of
- 47:15:25precision and recall. It provides a
- 47:15:27balance between the two metrics and is
- 47:15:29useful when you need to balance
- 47:15:31precision and recall. F1 score is equal
- 47:15:34to twice into precision into recall and
- 47:15:38that is divided by precision plus
- 47:15:40recall. Now we'll move to 14th question
- 47:15:43and here we will cover regularization.
- 47:15:46So the question is what is
- 47:15:47regularization? So it is a technique
- 47:15:50used to prevent overfitting by adding a
- 47:15:52penalty to the model's complexity and
- 47:15:55the common types of regularization
- 47:15:57include L1 that is lasso and L2 ridge
- 47:16:00regularization. Now we'll move to the
- 47:16:0215th question that is what is the bias
- 47:16:05variance tradeoff. So the bias variance
- 47:16:07trade-off is a fundamental issue in
- 47:16:10machine learning that involves balancing
- 47:16:12the error introduced by the model's
- 47:16:14assumptions and the error due to model
- 47:16:17complexity. So a good model should have
- 47:16:19low bias and low variance. Now we'll
- 47:16:22move to the question number 16 that is
- 47:16:24what is feature engineering. So feature
- 47:16:26engineering is the process of creating
- 47:16:28new features or modifying existing ones
- 47:16:31to improve the performance of a machine
- 47:16:33learning model. It involves techniques
- 47:16:35like normalization and coding
- 47:16:37categorical variables and creating
- 47:16:40interaction terms. So now we'll move to
- 47:16:42question number 17 and that is about
- 47:16:45gradient descent. So the question is
- 47:16:47what is gradient descent and your answer
- 47:16:49is gradient descent is an optimization
- 47:16:52algorithm used to minimize the cost
- 47:16:54function in machine learning models and
- 47:16:56it iteratively adjust the model
- 47:16:58parameters in the direction of the
- 47:17:00steepest descent of the coast function.
- 47:17:04So with this we'll move to the 18th
- 47:17:06question and that will cover with the
- 47:17:08difference between bagging and boosting.
- 47:17:11So the question is what is difference
- 47:17:13between bagging and boosting and you
- 47:17:14could answer this with starting with
- 47:17:16bagging that is bootstrap aggregating
- 47:17:20that involves training multiple models
- 47:17:22on different subsets of the data and
- 47:17:24averaging their predictions. Then comes
- 47:17:26boosting that involves training models
- 47:17:28sequentially with each new model
- 47:17:30focusing on correcting the errors of the
- 47:17:32previous ones. And then we have the
- 47:17:35question number 19 that is what is a
- 47:17:37decision tree? So a decision tree is a
- 47:17:39nonparametric supervised learning
- 47:17:41algorithm used for classification and
- 47:17:44regression. It splits the data into
- 47:17:46subsets based on the value of input
- 47:17:48features resulting in a treel like
- 47:17:50structure of decisions. Now we'll move
- 47:17:52to question number 20 that is what is a
- 47:17:54random forest. So random forest is an
- 47:17:56ansemble learning method that combines
- 47:17:59multiple decision trees to improve the
- 47:18:01accuracy and robustness of the model. It
- 47:18:04builds each tree using a random subset
- 47:18:06of features and data points and then
- 47:18:09averages their predictions. So these
- 47:18:11were the questions that are for the
- 47:18:13intermediate level and these are just
- 47:18:15the basic questions or I will just say
- 47:18:18the theoretical questions that can be
- 47:18:20asked in an interview. So be prepared
- 47:18:22for that. Now we'll move to the advanced
- 47:18:24level interview questions and we'll
- 47:18:26start with question number 21. And here
- 47:18:28also we'll cover the 10 questions. So
- 47:18:30number one question or that is 21
- 47:18:33question and the question is what is a
- 47:18:35support vector machine? So support
- 47:18:37vector machine is a supervised learning
- 47:18:39algorithm used for classification and
- 47:18:42regression. It finds the optimal hyper
- 47:18:44plane that maximizes the margin between
- 47:18:46different classes in the feature space.
- 47:18:49And then comes question number 22 that
- 47:18:51is what is principal component analysis.
- 47:18:54So principal component analysis is a
- 47:18:56dimensionality reduction technique that
- 47:18:58transforms highdimensional data into a
- 47:19:01lower dimensional space by finding the
- 47:19:03directions that is principal components
- 47:19:06that maximize the variance in the data.
- 47:19:08And then comes the question number 23
- 47:19:10that is what is a neural network? So a
- 47:19:12neural network is a series of algorithms
- 47:19:14that attempt to recognize underlying
- 47:19:17relationships in a set of data through a
- 47:19:19process that mimics the way the human
- 47:19:21brain operates. It consists of layer of
- 47:19:24interconnected nodes or neurons. And
- 47:19:27then comes the question number 24 that
- 47:19:29is what is deep learning? So deep
- 47:19:31learning is a subset of machine learning
- 47:19:33that involves neural networks with many
- 47:19:35layers that is deep neural networks and
- 47:19:37it is particularly effective for task
- 47:19:39like image and speech recognition. Now I
- 47:19:42move to question number 25 that is what
- 47:19:44is convolutional neural network that is
- 47:19:48CNN. So we will start the answer by
- 47:19:50answering the interviewer that a
- 47:19:52convolutional neural network is a type
- 47:19:54of deep learning model specifically
- 47:19:56designed for processing structured grid
- 47:19:58data like images. It uses convolutional
- 47:20:01layers to extract special features or
- 47:20:04the spatial features and patterns from
- 47:20:06the input data. Now we move to the
- 47:20:08question number 26 that is what is a
- 47:20:10recurrent neural network or RNN. So a
- 47:20:14recurrent neural network is a type of
- 47:20:16neural network designed for sequential
- 47:20:18data and it has connections that form
- 47:20:20directed cycles allowing it to maintain
- 47:20:23a memory of previous inputs and process
- 47:20:26sequences of data. So this was all about
- 47:20:28question number 26 and now we will cover
- 47:20:30the question number 27 that is what is
- 47:20:32the difference between batch gradient
- 47:20:34descent and stoastic gradient descent.
- 47:20:38So batch gradient descent computes the
- 47:20:40gradient of the coast function using the
- 47:20:43entire training data set while
- 47:20:45stochastic gradient descent that is SGD
- 47:20:48computes the gradient using only one
- 47:20:50training example at a time. So SGD is
- 47:20:54faster but noisier. Now we move to
- 47:20:56question number 28 that is what is
- 47:20:58dropout in neural networks. So dropout
- 47:21:00is a regularization technique used in
- 47:21:03neural networks to prevent overfitting
- 47:21:05and it involves randomly setting a
- 47:21:07fraction of the neurons to zero during
- 47:21:10training forcing the network to learn
- 47:21:12more robust features. And now we'll move
- 47:21:15to question number 29 and that will be
- 47:21:17about transfer learning. And your
- 47:21:19question is what is transfer learning?
- 47:21:22So we'll answer this to the interviewer
- 47:21:23by starting that transfer learning is a
- 47:21:26technique in machine learning where a
- 47:21:28model developed for one task is reused
- 47:21:30as the starting point for a model on a
- 47:21:33second related task. It is particularly
- 47:21:36useful when there is limited data
- 47:21:37available for the second task. Now we'll
- 47:21:40move to the last question and the 30th
- 47:21:42question. So that is what is a
- 47:21:44generative adversial network that is GN.
- 47:21:48So you can start answering this. So
- 47:21:50generative adversial network is a type
- 47:21:52of deep learning model consisting of two
- 47:21:55neural networks a generator and a
- 47:21:57discriminator that are trained
- 47:21:59simultaneously. The generator creates
- 47:22:01fake data while the discriminator tries
- 47:22:04to distinguish between real and fake
- 47:22:06data leading to the generator producing
- 47:22:08increasingly realistic data. And these
- 47:22:11questions and answers are over and these
- 47:22:14covers a wide range of topics in machine
- 47:22:16learning and should help prepare for
- 47:22:18interviews at various levels.
- 47:22:19>> And with that we have reached the end of
- 47:22:21the session on the AI and machine
- 47:22:23learning engineer full course for
- 47:22:25beginners. If you have any doubts or
- 47:22:27questions about this video let us know
- 47:22:29in the comment section below and a team
- 47:22:30of experts will be happy to help you.
- 47:22:32Until next time thank you and keep
- 47:22:34learning. Stay tuned for more from
- 47:22:35SimplyLearn.
About this transcript
This page contains the full transcript of AI And Machine Learning Full Course [FREE] | Learn AI And Machine Learning In 24 Hours | Simplilearn by Simplilearn, generated from the public captions YouTube serves with the video. The transcript has 423,981 words across 63,517 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.