Harvard CS50’s Artificial Intelligence with Python – Full University Course — Transcript
Full transcript
- 0:00This course from Harvard University explores the concepts and algorithms at the foundation of modern
- 0:07artificial intelligence, diving into the ideas that give rise to technologies like game-playing
- 0:13engines, handwriting recognition, and machine translation. You'll gain exposure to the theory
- 0:20behind graph search algorithms, classification, optimization, reinforcement learning,
- 0:26and other topics in artificial intelligence and machine learning. Brian Yu teaches this course.
- 0:33Hello, world. This is CS50, and this is an introduction to artificial intelligence with
- 0:56Python with CS50's own Brian Yu. This course picks up where CS50 itself leaves off and explores the
- 1:02concepts and algorithms at the foundation of modern AI.
- 1:06We'll start with a look at how AI can search for solutions to problems,
- 1:10whether those problems are learning how to play a game or trying
- 1:12to find driving directions to a destination.
- 1:15We'll then look at how AI can represent information, both knowledge that our AI
- 1:19is certain about, but also information and events about which our AI might be uncertain,
- 1:23learning how to represent that information, but more importantly,
- 1:26how to use that information to draw inferences and new conclusions as well.
- 1:30We'll explore how AI can solve various types of optimization problems,
- 1:33trying to maximize profits or minimize costs or satisfy some other constraints
- 1:38before turning our attention to the fast-growing field of machine learning,
- 1:41where we won't tell our AI exactly how to solve a problem, but instead,
- 1:45give our AI access to data and experiences
- 1:48so that our AI can learn on its own how to perform these tasks.
- 1:52In particular, we'll look at neural networks, one of the most popular tools
- 1:55in modern machine learning, inspired by the way that human brains learn and reason as well
- 1:59before finally taking a look at the world of natural language processing
- 2:03so that it's not just us humans learning to learn how artificial intelligence is
- 2:06able to speak, but also AI learning how to understand and interpret human language as well.
- 2:11We'll explore these ideas and algorithms, and along the way,
- 2:14give you the opportunity to build your own AI programs to implement all of this and more.
- 2:19This is CS50.
- 2:44All right.
- 2:44Welcome, everyone, to an introduction to artificial intelligence with Python.
- 2:48My name is Brian Yu, and in this class, we'll
- 2:50explore some of the ideas and techniques and algorithms
- 2:53that are at the foundation of artificial intelligence.
- 2:56Now, artificial intelligence covers a wide variety of types of techniques.
- 3:00Anytime you see a computer do something that
- 3:01appears to be intelligent or rational in some way,
- 3:05like recognizing someone's face in a photo,
- 3:07or being able to play a game better than people can,
- 3:09or being able to understand human language when we talk to our phones
- 3:12and they understand what we mean and are able to respond back to us,
- 3:16these are all examples of AI, or artificial intelligence.
- 3:19And in this class, we'll explore some of the ideas that make that AI possible.
- 3:24So we'll begin our conversations with search, the problem of we have an AI,
- 3:28and we would like the AI to be able to search for solutions to some kind of problem,
- 3:32no matter what that problem might be.
- 3:33Whether it's trying to get driving directions from point A to point B,
- 3:37or trying to figure out how to play a game, given a tic-tac-toe game,
- 3:40for example, figuring out what move it ought to make.
- 3:43After that, we'll take a look at knowledge.
- 3:45Ideally, we want our AI to be able to know information,
- 3:48to be able to represent that information, and more importantly,
- 3:51to be able to draw inferences from that information,
- 3:53to be able to use the information it knows and draw additional conclusions.
- 3:57So we'll talk about how AI can be programmed in order to do just that.
- 4:02Then we'll explore the topic of uncertainty,
- 4:04talking about ideas of what happens if a computer isn't sure about a fact,
- 4:08but maybe is only sure with a certain probability.
- 4:11So we'll talk about some of the ideas behind probability,
- 4:13and how computers can begin to deal with uncertain events
- 4:16in order to be a little bit more intelligent in that sense as well.
- 4:20After that, we'll turn our attention to optimization,
- 4:23problems of when the computer is trying to optimize for some sort of goal,
- 4:27especially in a situation where there might be multiple ways
- 4:29that a computer might solve a problem, but we're looking for a better way,
- 4:33or potentially the best way, if that's at all possible.
- 4:36Then we'll take a look at machine learning, or learning more generally,
- 4:39and looking at how, when we have access to data,
- 4:41our computers can be programmed to be quite intelligent by learning from data
- 4:45and learning from experience, being able to perform a task better and better
- 4:48based on greater access to data.
- 4:50So your email, for example, where your email inbox somehow knows
- 4:54which of your emails are good emails and which of your emails are spam.
- 4:57These are all examples of computers being able to learn from past experiences
- 5:01and past data.
- 5:03We'll take a look, too, at how computers are
- 5:05able to draw inspiration from human intelligence,
- 5:08looking at the structure of the human brain,
- 5:10and how neural networks can be a computer analog to that sort of idea,
- 5:13and how, by taking advantage of a certain type of structure of a computer program,
- 5:17we can write neural networks that are able to perform tasks very, very
- 5:21effectively.
- 5:22And then finally, we'll turn our attention to language, not programming
- 5:25languages, but human languages that we speak every day.
- 5:28And taking a look at the challenges that come about as a computer tries
- 5:31to understand natural language, and how it is some of the natural language
- 5:35processing that occurs in modern artificial intelligence can actually
- 5:39work.
- 5:40But today, we'll begin our conversation with search, this problem
- 5:43of trying to figure out what to do when we have some sort of situation
- 5:47that the computer is in, some sort of environment that an agent is in,
- 5:50so to speak, and we would like for that agent
- 5:52to be able to somehow look for a solution to that problem.
- 5:56Now, these problems can come in any number of different types of formats.
- 5:59One example, for instance, might be something
- 6:01like this classic 15 puzzle with the sliding tiles that you might have seen.
- 6:04Where you're trying to slide the tiles in order
- 6:06to make sure that all the numbers line up in order.
- 6:09This is an example of what you might call a search problem.
- 6:12The 15 puzzle begins in an initially mixed up state,
- 6:15and we need some way of finding moves to make in order
- 6:18to return the puzzle to its solved state.
- 6:20But there are similar problems that you can frame in other ways.
- 6:23Trying to find your way through a maze, for example,
- 6:25is another example of a search problem.
- 6:27You begin in one place, you have some goal of where you're trying to get to,
- 6:31and you need to figure out the correct sequence of actions that will take you
- 6:34from that initial state to the goal.
- 6:36And while this is a little bit abstract, any time
- 6:38we talk about maze solving in this class,
- 6:40you can translate it to something a little more real world.
- 6:43Something like driving directions.
- 6:45If you ever wonder how Google Maps is able to figure out what is the best way
- 6:48for you to get from point A to point B, and what turns to make at what time,
- 6:52depending on traffic, for example, it's often some sort of search algorithm.
- 6:56You have an AI that is trying to get from an initial position
- 6:59to some sort of goal by taking some sequence of actions.
- 7:03So we'll start our conversations today by thinking
- 7:06about these types of search problems and what
- 7:08goes in to solving a search problem like this in order for an AI
- 7:11to be able to find a good solution.
- 7:14In order to do so, though, we're going to need
- 7:15to introduce a little bit of terminology, some of which I've already used.
- 7:19But the first term we'll need to think about is an agent.
- 7:22An agent is just some entity that perceives its environment.
- 7:25It somehow is able to perceive the things around it
- 7:27and act on that environment in some way.
- 7:30So in the case of the driving directions,
- 7:31your agent might be some representation of a car that
- 7:34is trying to figure out what actions to take in order
- 7:36to arrive at a destination.
- 7:38In the case of the 15 puzzle with the sliding tiles,
- 7:40the agent might be the AI or the person that
- 7:43is trying to solve that puzzle to try and figure out what tiles to move
- 7:46in order to get to that solution.
- 7:49Next, we introduce the idea of a state.
- 7:52A state is just some configuration of the agent in its environment.
- 7:56So in the 15 puzzle, for example, any state might be any one of these three,
- 8:00for example. A state is just some configuration of the tiles.
- 8:03And each of these states is different and is
- 8:05going to require a slightly different solution.
- 8:08A different sequence of actions will be needed in each one of these
- 8:11in order to get from this initial state to the goal, which
- 8:15is where we're trying to get.
- 8:16So the initial state, then, what is that?
- 8:18The initial state is just the state where the agent begins.
- 8:21It is one such state where we're going to start from.
- 8:24And this is going to be the starting point for our search algorithm,
- 8:27so to speak.
- 8:28We're going to begin with this initial state
- 8:29and then start to reason about it, to think about what actions might we
- 8:32apply to that initial state in order to figure out how to get from the beginning
- 8:37to the end, from the initial position to whatever our goal happens to be.
- 8:42And how do we make our way from that initial position to the goal?
- 8:44Well, ultimately, it's via taking actions.
- 8:47Actions are just choices that we can make in any given state.
- 8:50And in AI, we're always going to try to formalize these ideas a little bit
- 8:54more precisely, such that we could program them a little bit more
- 8:57mathematically, so to speak.
- 8:58So this will be a recurring theme.
- 9:00And we can more precisely define actions as a function.
- 9:04We're going to effectively define a function called actions that takes an
- 9:07input, s, where s is going to be some state that exists inside of our environment.
- 9:12And actions of s is going to take the state as input and return as output
- 9:17the set of all actions that can be executed in that state.
- 9:22And so it's possible that some actions are only valid in certain states
- 9:25and not in other states.
- 9:27And we'll see examples of that soon, too.
- 9:29So in the case of the 15 puzzle, for example,
- 9:31there are generally going to be four possible actions that we can do most of
- 9:35the time.
- 9:36We can slide a tile to the right, slide a tile to the left, slide a tile up,
- 9:39or slide a tile down, for example.
- 9:41And those are going to be the actions that are available to us.
- 9:45So somehow our AI, our program, needs some encoding
- 9:48of the state, which is often going to be in some numerical format,
- 9:51and some encoding of these actions.
- 9:53But it also needs some encoding of the relationship between these things.
- 9:56How do the states and actions relate to one another?
- 10:00And in order to do that, we'll introduce to our AI a transition model, which
- 10:04will be a description of what state we get after we perform some available
- 10:08action in some other state.
- 10:10And again, we can be a little bit more precise about this,
- 10:12define this transition model a little bit more formally, again, as a function.
- 10:17The function is going to be a function called result that this time takes two
- 10:20inputs.
- 10:21Input number one is s, some state.
- 10:24And input number two is a, some action.
- 10:27And the output of this function result is it
- 10:30is going to give us the state that we get after we perform action a in state s.
- 10:36So let's take a look at an example to see more precisely what this actually means.
- 10:39Here is an example of a state, of the 15 puzzle, for example.
- 10:43And here is an example of an action, sliding a tile to the right.
- 10:46What happens if we pass these as inputs to the result function?
- 10:50Again, the result function takes this board, this state, as its first input.
- 10:54And it takes an action as a second input.
- 10:57And of course, here, I'm describing things visually
- 10:59so that you can see visually what the state is and what the action is.
- 11:02In a computer, you might represent one of these actions
- 11:04as just some number that represents the action.
- 11:06Or if you're familiar with enums that allow
- 11:08you to enumerate multiple possibilities,
- 11:10it might be something like that.
- 11:11And this state might just be represented
- 11:13as an array or two-dimensional array of all of these numbers that exist.
- 11:17But here, we're going to show it visually just so you can see it.
- 11:20But when we take this state and this action,
- 11:23pass it into the result function, the output is a new state.
- 11:26The state we get after we take a tile and slide it to the right,
- 11:30and this is the state we get as a result.
- 11:32If we had a different action and a different state, for example,
- 11:35and pass that into the result function, we'd
- 11:37get a different answer altogether.
- 11:38So the result function needs to take care
- 11:41of figuring out how to take a state and take an action and get what results.
- 11:45And this is going to be our transition model that
- 11:48describes how it is that states and actions are related to each other.
- 11:52If we take this transition model and think about it more generally
- 11:55and across the entire problem, we can form what we might call a state space.
- 12:00The set of all of the states we can get from the initial state
- 12:03via any sequence of actions, by taking 0 or 1 or 2 or more actions in addition
- 12:08to that, so we could draw a diagram that looks something like this, where
- 12:12every state is represented here by a game board, and there are arrows
- 12:15that connect every state to every other state we can get to from that state.
- 12:20And the state space is much larger than what you see just here.
- 12:23This is just a sample of what the state space might actually look like.
- 12:27And in general, across many search problems,
- 12:29whether they're this particular 15 puzzle or driving directions or something else,
- 12:33the state space is going to look something like this.
- 12:36We have individual states and arrows that are connecting them.
- 12:40And oftentimes, just for simplicity, we'll
- 12:42simplify our representation of this entire thing as a graph, some sequence
- 12:47of nodes and edges that connect nodes.
- 12:50But you can think of this more abstract representation
- 12:52as the exact same idea.
- 12:54Each of these little circles or nodes is going
- 12:56to represent one of the states inside of our problem.
- 12:59And the arrows here represent the actions
- 13:01that we can take in any particular state, taking us
- 13:04from one particular state to another state, for example.
- 13:09All right.
- 13:10So now we have this idea of nodes that are representing these states,
- 13:14actions that can take us from one state to another,
- 13:16and a transition model that defines what happens after we
- 13:19take a particular action.
- 13:21So the next step we need to figure out is how
- 13:23we know when the AI is done solving the problem.
- 13:26The AI needs some way to know when it gets to the goal that it's found the goal.
- 13:30So the next thing we'll need to encode into our artificial intelligence
- 13:33is a goal test, some way to determine whether a given state is a goal state.
- 13:39In the case of something like driving directions, it might be pretty easy.
- 13:42If you're in a state that corresponds to whatever the user typed
- 13:45in as their intended destination, well, then you know you're in a goal state.
- 13:48In the 15 puzzle, it might be checking the numbers
- 13:51to make sure they're all in ascending order.
- 13:52But the AI needs some way to encode whether or not
- 13:55any state they happen to be in is a goal.
- 13:58And some problems might have one goal, like a maze
- 14:00where you have one initial position and one ending position,
- 14:03and that's the goal.
- 14:04In other more complex problems, you might imagine
- 14:06that there are multiple possible goals.
- 14:08That there are multiple ways to solve a problem,
- 14:10and we might not care which one the computer finds,
- 14:13as long as it does find a particular goal.
- 14:17However, sometimes the computer doesn't just care about finding a goal,
- 14:20but finding a goal well, or one with a low cost.
- 14:23And it's for that reason that the last piece of terminology
- 14:26that we'll use to define these search problems
- 14:28is something called a path cost.
- 14:30You might imagine that in the case of driving directions,
- 14:33it would be pretty annoying if I said I wanted directions from point A
- 14:36to point B, and the route that Google Maps gave me
- 14:38was a long route with lots of detours that were unnecessary that took longer
- 14:42than it should have for me to get to that destination.
- 14:45And it's for that reason that when we're formulating search problems,
- 14:48we'll often give every path some sort of numerical cost,
- 14:51some number telling us how expensive it is to take this particular option,
- 14:56and then tell our AI that instead of just finding
- 14:59a solution, some way of getting from the initial state to the goal,
- 15:02we'd really like to find one that minimizes this path cost.
- 15:06That is, less expensive, or takes less time,
- 15:09or minimizes some other numerical value.
- 15:12We can represent this graphically if we take a look at this graph again,
- 15:15and imagine that each of these arrows, each of these actions
- 15:18that we can take from one state to another state,
- 15:21has some sort of number associated with it.
- 15:23That number being the path cost of this particular action,
- 15:26where some of the costs for any particular action
- 15:29might be more expensive than the cost for some other action, for example.
- 15:33Although this will only happen in some sorts of problems.
- 15:35In other problems, we can simplify the diagram
- 15:38and just assume that the cost of any particular action is the same.
- 15:42And this is probably the case in something like the 15 puzzle,
- 15:45for example, where it doesn't really make a difference
- 15:47whether I'm moving right or moving left.
- 15:49The only thing that matters is the total number
- 15:52of steps that I have to take to get from point A to point B.
- 15:56And each of those steps is of equal cost.
- 15:58We can just assume it's of some constant cost like one.
- 16:03And so this now forms the basis for what we might consider to be a search problem.
- 16:07A search problem has some sort of initial state, some place where we begin,
- 16:11some sort of action that we can take or multiple actions
- 16:14that we can take in any given state.
- 16:16And it has a transition model.
- 16:17Some way of defining what happens when we go from one state
- 16:21and take one action, what state do we end up with as a result.
- 16:24In addition to that, we need some goal test
- 16:26to know whether or not we've reached a goal.
- 16:29And then we need a path cost function that
- 16:31tells us for any particular path, by following some sequence of actions,
- 16:35how expensive is that path.
- 16:37What does its cost in terms of money or time or some other resource
- 16:41that we are trying to minimize our usage of.
- 16:44And the goal ultimately is to find a solution.
- 16:46Where a solution in this case is just some sequence of actions
- 16:50that will take us from the initial state to the goal state.
- 16:52And ideally, we'd like to find not just any solution
- 16:55but the optimal solution, which is a solution that
- 16:58has the lowest path cost among all of the possible solutions.
- 17:02And in some cases, there might be multiple optimal solutions.
- 17:05But an optimal solution just means that there
- 17:07is no way that we could have done better in terms of finding that solution.
- 17:12So now we've defined the problem.
- 17:13And now we need to begin to figure out how it
- 17:15is that we're going to solve this kind of search problem.
- 17:18And in order to do so, you'll probably imagine
- 17:21that our computer is going to need to represent a whole bunch of data
- 17:24about this particular problem.
- 17:26We need to represent data about where we are in the problem.
- 17:28And we might need to be considering multiple different options at once.
- 17:32And oftentimes, when we're trying to package a whole bunch of data
- 17:35related to a state together, we'll do so using a data structure
- 17:38that we're going to call a node.
- 17:40A node is a data structure that is just going
- 17:42to keep track of a variety of different values.
- 17:44And specifically, in the case of a search problem,
- 17:47it's going to keep track of these four values in particular.
- 17:50Every node is going to keep track of a state, the state we're currently on.
- 17:54And every node is also going to keep track of a parent.
- 17:57A parent being the state before us or the node
- 18:00that we used in order to get to this current state.
- 18:03And this is going to be relevant because eventually, once we reach the goal node,
- 18:07once we get to the end, we want to know what sequence of actions
- 18:10we use in order to get to that goal.
- 18:12And the way we'll know that is by looking at these parents
- 18:16to keep track of what led us to the goal and what led us to that state
- 18:19and what led us to the state before that, so on and so forth,
- 18:22backtracking our way to the beginning so that we
- 18:25know the entire sequence of actions we needed in order
- 18:27to get from the beginning to the end.
- 18:30The node is also going to keep track of what action we took in order
- 18:33to get from the parent to the current state.
- 18:35And the node is also going to keep track of a path cost.
- 18:39In other words, it's going to keep track of the number
- 18:41that represents how long it took to get from the initial state
- 18:45to the state that we currently happen to be at.
- 18:47And we'll see why this is relevant as we
- 18:49start to talk about some of the optimizations
- 18:51that we can make in terms of these search problems more generally.
- 18:55So this is the data structure that we're going to use in order to solve
- 18:57the problem.
- 18:58And now let's talk about the approach.
- 19:00How might we actually begin to solve the problem?
- 19:03Well, as you might imagine, what we're going to do
- 19:05is we're going to start at one particular state,
- 19:08and we're just going to explore from there.
- 19:10The intuition is that from a given state,
- 19:12we have multiple options that we could take,
- 19:14and we're going to explore those options.
- 19:16And once we explore those options, we'll
- 19:18find that more options than that are going to make themselves available.
- 19:22And we're going to consider all of the available options
- 19:24to be stored inside of a single data structure that we'll call the frontier.
- 19:29The frontier is going to represent all of the things
- 19:31that we could explore next that we haven't yet explored or visited.
- 19:36So in our approach, we're going to begin the search algorithm
- 19:39by starting with a frontier that just contains one state.
- 19:42The frontier is going to contain the initial state,
- 19:45because at the beginning, that's the only state we know about.
- 19:47That is the only state that exists.
- 19:50And then our search algorithm is effectively going to follow a loop.
- 19:53We're going to repeat some process again and again and again.
- 19:57The first thing we're going to do is if the frontier is empty,
- 20:01then there's no solution.
- 20:02And we can report that there is no way to get to the goal.
- 20:05And that's certainly possible.
- 20:06There are certain types of problems that an AI might try to explore
- 20:09and realize that there is no way to solve that problem.
- 20:12And that's useful information for humans to know as well.
- 20:15So if ever the frontier is empty, that means there's nothing left to explore.
- 20:19And we haven't yet found a solution, so there is no solution.
- 20:22There's nothing left to explore.
- 20:24Otherwise, what we'll do is we'll remove a node from the frontier.
- 20:28So right now at the beginning, the frontier just contains one node
- 20:32representing the initial state.
- 20:33But over time, the frontier might grow.
- 20:35It might contain multiple states.
- 20:36And so here, we're just going to remove a single node from that frontier.
- 20:41If that node happens to be a goal, then we found a solution.
- 20:44So we remove a node from the frontier and ask ourselves, is this the goal?
- 20:48And we do that by applying the goal test that we talked about earlier,
- 20:51asking if we're at the destination.
- 20:53Or asking if all the numbers of the 15 puzzle happen to be in order.
- 20:56So if the node contains the goal, we found a solution.
- 20:59Great.
- 21:00We're done.
- 21:01And otherwise, what we'll need to do is we'll need to expand the node.
- 21:06And this is a term of art in artificial intelligence.
- 21:08To expand the node just means to look at all of the neighbors of that node.
- 21:12In other words, consider all of the possible actions
- 21:15that I could take from the state that this node is representing
- 21:18and what nodes could I get to from there.
- 21:21We're going to take all of those nodes, the next nodes
- 21:23that I can get to from this current one I'm looking at,
- 21:26and add those to the frontier.
- 21:28And then we'll repeat this process.
- 21:30So at a very high level, the idea is we start
- 21:32with a frontier that contains the initial state.
- 21:35And we're constantly removing a node from the frontier,
- 21:38looking at where we can get to next and adding those nodes to the frontier,
- 21:41repeating this process over and over until either we
- 21:44remove a node from the frontier and it contains a goal,
- 21:47meaning we've solved the problem, or we run into a situation
- 21:50where the frontier is empty, at which point we're left with no solution.
- 21:55So let's actually try and take the pseudocode,
- 21:57put it into practice by taking a look at an example of a sample search problem.
- 22:02So right here, I have a sample graph.
- 22:04A is connected to B via this action.
- 22:06B is connected to nodes C and D. C is connected to E. D is connected to F.
- 22:10And what I'd like to do is have my AI find a path from A to E.
- 22:16We want to get from this initial state to this goal state.
- 22:20So how are we going to do that?
- 22:22Well, we're going to start with a frontier that contains the initial state.
- 22:25This is going to represent our frontier.
- 22:27So our frontier initially will just contain
- 22:29A, that initial state where we're going to begin.
- 22:32And now we'll repeat this process.
- 22:34If the frontier is empty, no solution.
- 22:36That's not a problem, because the frontier is not empty.
- 22:38So we'll remove a node from the frontier as the one to consider next.
- 22:42There's only one node in the frontier.
- 22:44So we'll go ahead and remove it from the frontier.
- 22:46But now A, this initial node, this is the node we're currently considering.
- 22:51We follow the next step.
- 22:52We ask ourselves, is this node the goal?
- 22:55No, it's not.
- 22:55A is not the goal.
- 22:56E is the goal.
- 22:57So we don't return the solution.
- 22:59So instead, we go to this last step, expand the node,
- 23:02and add the resulting nodes to the frontier.
- 23:05What does that mean?
- 23:06Well, it means take this state A and consider where we could get to next.
- 23:10And after A, what we could get to next is only B.
- 23:14So that's what we get when we expand A. We find B.
- 23:16And we add B to the frontier.
- 23:18And now B is in the frontier.
- 23:20And we repeat the process again.
- 23:22We say, all right, the frontier is not empty.
- 23:24So let's remove B from the frontier.
- 23:26B is now the node that we're considering.
- 23:28We ask ourselves, is B the goal?
- 23:29No, it's not.
- 23:30So we go ahead and expand B and add its resulting nodes to the frontier.
- 23:35What happens when we expand B?
- 23:37In other words, what nodes can we get to from B?
- 23:40Well, we can get to C and D. So we'll go ahead and add C and D
- 23:43from the frontier.
- 23:44And now we have two nodes in the frontier, C and D.
- 23:47And we repeat the process again.
- 23:48We remove a node from the frontier.
- 23:50For now, I'll do so arbitrarily just by picking C.
- 23:52We'll see why later, how choosing which node you remove from the frontier
- 23:56is actually quite an important part of the algorithm.
- 23:58But for now, I'll arbitrarily remove C, say it's not the goal.
- 24:02So we'll add E, the next one, to the frontier.
- 24:05Then let's say I remove E from the frontier.
- 24:07And now I check I'm currently looking at state E. Is it a goal state?
- 24:11It is, because I'm trying to find a path from A to E. So I would return the goal.
- 24:15And that now would be the solution, that I'm now able to return the solution.
- 24:19And I have found a path from A to E.
- 24:23So this is the general idea, the general approach of this search algorithm,
- 24:26to follow these steps, constantly removing nodes from the frontier,
- 24:30until we're able to find a solution.
- 24:31So the next question you might reasonably ask is, what could go wrong here?
- 24:35What are the potential problems with an approach like this?
- 24:39And here's one example of a problem that could arise from this sort of approach.
- 24:42Imagine this same graph, same as before, with one change.
- 24:47The change being now, instead of just an arrow from A to B,
- 24:50we also have an arrow from B to A, meaning we can go in both directions.
- 24:54And this is true in something like the 15 puzzle, where when I slide a tile
- 24:57to the right, I could then slide a tile to the left
- 25:00to get back to the original position.
- 25:02I could go back and forth between A and B.
- 25:04And that's what these double arrows symbolize,
- 25:06the idea that from one state, I can get to another, and then I can get back.
- 25:10And that's true in many search problems.
- 25:12What's going to happen if I try to apply the same approach now?
- 25:16Well, I'll begin with A, same as before.
- 25:18And I'll remove A from the frontier.
- 25:20And then I'll consider where I can get to from A.
- 25:23And after A, the only place I can get to is B. So B goes into the frontier.
- 25:28Then I'll say, all right, let's take a look at B.
- 25:29That's the only thing left in the frontier.
- 25:31Where can I get to from B?
- 25:33Before, it was just C and D. But now, because of that reverse arrow,
- 25:37I can get to A or C or D. So all three, A, C, and D, all of those
- 25:43now go into the frontier.
- 25:44They are places I can get to from B. And now I remove one from the frontier.
- 25:48And maybe I'm unlucky, and maybe I pick A. And now I'm looking at A again.
- 25:53And I consider, where can I get to from A?
- 25:54And from A, well, I can get to B. And now we start to see the problem.
- 25:58But if I'm not careful, I go from A to B, and then back to A, and then to B again.
- 26:02And I could be going in this infinite loop, where I never make any progress,
- 26:05because I'm constantly just going back and forth between two states
- 26:09that I've already seen.
- 26:10So what is the solution to this?
- 26:12We need some way to deal with this problem.
- 26:14And the way that we can deal with this problem
- 26:16is by somehow keeping track of what we've already explored.
- 26:20And the logic is going to be, well, if we've already explored the state,
- 26:23there's no reason to go back to it.
- 26:25Once we've explored a state, don't go back to it.
- 26:27Don't bother adding it to the frontier.
- 26:29There's no need to.
- 26:31So here's going to be our revised approach, a better way
- 26:33to approach this sort of search problem.
- 26:35And it's going to look very similar, just with a couple of modifications.
- 26:39We'll start with a frontier that contains the initial state, same as before.
- 26:43But now we'll start with another data structure, which
- 26:46will just be a set of nodes that we've already explored.
- 26:49So what are the states we've explored?
- 26:51Initially, it's empty.
- 26:52We have an empty explored set.
- 26:55And now we repeat.
- 26:57If the frontier is empty, no solution, same as before.
- 27:00We remove a node from the frontier.
- 27:02We check to see if it's a goal state, return the solution.
- 27:04None of this is any different so far.
- 27:06But now what we're going to do is we're going to add the node
- 27:09to the explored state.
- 27:11So if it happens to be the case that we remove a node from the frontier
- 27:15and it's not the goal, we'll add it to the explored set
- 27:18so that we know we've already explored it.
- 27:19We don't need to go back to it again if it happens to come up later.
- 27:23And then the final step, we expand the node
- 27:26and we add the resulting nodes to the frontier.
- 27:28But before, we just always added the resulting nodes to the frontier.
- 27:31We're going to be a little clever about it this time.
- 27:34We're only going to add the nodes to the frontier
- 27:36if they aren't already in the frontier and if they aren't already
- 27:40in the explored set.
- 27:42So we'll check both the frontier and the explored set,
- 27:45make sure that the node isn't already in one of those two.
- 27:48And so long as it isn't, then we'll go ahead and add it to the frontier,
- 27:51but not otherwise.
- 27:53And so that revised approach is ultimately
- 27:55what's going to help make sure that we don't go back and forth between two
- 27:58nodes.
- 28:00Now, the one point that I've kind of glossed over here so far
- 28:02is this step here, removing a node from the frontier.
- 28:06Before, I just chose arbitrarily.
- 28:08Like, let's just remove a node and that's it.
- 28:10But it turns out it's actually quite important how
- 28:12we decide to structure our frontier, how we add and how we remove our nodes.
- 28:17The frontier is a data structure and we need
- 28:19to make a choice about in what order are we
- 28:21going to be removing elements.
- 28:23And one of the simplest data structures for adding and removing elements
- 28:27is something called a stack.
- 28:28And a stack is a data structure that is a last in, first out data type, which
- 28:33means the last thing that I add to the frontier
- 28:36is going to be the first thing that I remove from the frontier.
- 28:40So the most recent thing to go into the stack or the frontier in this case
- 28:44is going to be the node that I explore.
- 28:47So let's see what happens if I apply this stack-based approach to something
- 28:51like this problem, finding a path from A to E. What's going to happen?
- 28:56Well, again, we'll start with A and we'll say, all right,
- 28:58let's go ahead and look at A first.
- 29:00And then notice this time, we've added A to the explored set.
- 29:04A is something we've now explored.
- 29:06We have this data structure that's keeping track.
- 29:09We then say from A, we can get to B. And all right, from B, what can we do?
- 29:13Well, from B, we can explore B and get to both C and D.
- 29:17So we added C and then D. So now,
- 29:21when we explore a node, we're going to treat the frontier as a stack,
- 29:24last in, first out.
- 29:26D was the last one to come in.
- 29:27So we'll go ahead and explore that next and say, all right,
- 29:30where can we get to from D?
- 29:32Well, we can get to F. And so all right, we'll put F into the frontier.
- 29:36And now, because the frontier is a stack,
- 29:39F is the most recent thing that's gone in the stack.
- 29:42So F is what we'll explore next.
- 29:43We'll explore F and say, all right, where can we get to from F?
- 29:47Well, we can't get anywhere, so nothing gets added to the frontier.
- 29:50So now, what was the new most recent thing added to the frontier?
- 29:53Well, it's now C, the only thing left in the frontier.
- 29:55We'll explore that from which we can see, all right, from C, we can get to E.
- 29:59So E goes into the frontier.
- 30:01And then we say, all right, let's look at E. And E is now the solution.
- 30:04And now, we've solved the problem.
- 30:07So when we treat the frontier like a stack, a last in,
- 30:10first out data structure, that's the result we get.
- 30:13We go from A to B to D to F. And then we sort of backed up and went down to C
- 30:18and then E.
- 30:19And it's important to get a visual sense for how this algorithm is working.
- 30:23We went very deep in this search tree, so to speak,
- 30:25all the way until the bottom where we hit a dead end.
- 30:28And then we effectively backed up and explored this other route
- 30:32that we didn't try before.
- 30:33And it's this going very deep in the search tree idea,
- 30:36this way the algorithm ends up working when we use a stack
- 30:39that we call this version of the algorithm depth first search.
- 30:44Depth first search is the search algorithm
- 30:46where we always explore the deepest node in the frontier.
- 30:49We keep going deeper and deeper through our search tree.
- 30:52And then if we hit a dead end, we back up and we try something else instead.
- 30:57But depth first search is just one of the possible search options
- 31:00that we could use.
- 31:01It turns out that there's another algorithm called breadth first search,
- 31:05which behaves very similarly to depth first search with one difference.
- 31:08Instead of always exploring the deepest node in the search tree,
- 31:12the way the depth first search does, breadth first search
- 31:14is always going to explore the shallowest node in the frontier.
- 31:19So what does that mean?
- 31:20Well, it means that instead of using a stack which depth first search or DFS
- 31:24used, where the most recent item added to the frontier
- 31:27is the one we'll explore next, in breadth first search or BFS,
- 31:32we'll instead use a queue, where a queue is a first in first out data type,
- 31:37where the very first thing we add to the frontier
- 31:39is the first one we'll explore and they effectively form a line or a queue,
- 31:43where the earlier you arrive in the frontier, the earlier you get explored.
- 31:49So what would that mean for the same exact problem,
- 31:51finding a path from A to E?
- 31:53Well, we start with A, same as before, then we'll go ahead and have explored A
- 31:57and say, where can we get to from A?
- 31:59Well, from A, we can get to B, same as before.
- 32:01From B, same as before, we can get to C and D.
- 32:04So C and D get added to the frontier.
- 32:06This time, though, we added C to the frontier before D.
- 32:10So we'll explore C first.
- 32:12So C gets explored.
- 32:14And from C, where can we get to?
- 32:16Well, we can get to E. So E gets added to the frontier.
- 32:19But because D was explored before E, we'll look at D next.
- 32:24So we'll explore D and say, where can we get to from D?
- 32:26We can get to F. And only then will we say, all right, now we can get to E.
- 32:31And so what breadth first search or BFS did is we started here,
- 32:35we looked at both C and D, and then we looked at E.
- 32:39Effectively, we're looking at things one away from the initial state,
- 32:42then two away from the initial state, and only then,
- 32:45things that are three away from the initial state, unlike depth first search,
- 32:49which just went as deep as possible into the search tree
- 32:53until it hit a dead end and then ultimately had to back up.
- 32:56So these now are two different search algorithms
- 32:59that we could apply in order to try and solve a problem.
- 33:01And let's take a look at how these would actually work in practice
- 33:05with something like maze solving, for example.
- 33:07So here's an example of a maze.
- 33:09These empty cells represent places where our agent can move.
- 33:12These darkened gray cells represent walls that the agent can't pass through.
- 33:16And ultimately, our agent, our AI, is going to try to find a way
- 33:20to get from position A to position B via some sequence of actions,
- 33:25where those actions are left, right, up, and down.
- 33:28What will depth first search do in this case?
- 33:31Well, depth first search will just follow one path.
- 33:34If it reaches a fork in the road where it has multiple different options,
- 33:37depth first search is just, in this case, going to choose one.
- 33:40That doesn't a real preference.
- 33:41But it's going to keep following one until it hits a dead end.
- 33:45And when it hits a dead end, depth first search effectively
- 33:48goes back to the last decision point and tries the other path,
- 33:52fully exhausting this entire path.
- 33:54And when it realizes that, OK, the goal is not here,
- 33:56then it turns its attention to this path.
- 33:58It goes as deep as possible.
- 34:00When it hits a dead end, it backs up and then tries this other path,
- 34:04keeps going as deep as possible down one particular path.
- 34:07And when it realizes that that's a dead end, then it'll back up,
- 34:10and then ultimately find its way to the goal.
- 34:13And maybe you got lucky, and maybe you made a different choice earlier on.
- 34:16But ultimately, this is how depth first search is going to work.
- 34:19It's going to keep following until it hits a dead end.
- 34:22And when it hits a dead end, it backs up and looks for a different solution.
- 34:26And so one thing you might reasonably ask is,
- 34:28is this algorithm always going to work?
- 34:30Will it always actually find a way to get from the initial state?
- 34:33To the goal.
- 34:34And it turns out that as long as our maze is finite,
- 34:37as long as there are only finitely many spaces where we can travel,
- 34:40then, yes, depth first search is going to find a solution.
- 34:44Because eventually, it'll just explore everything.
- 34:46If the maze happens to be infinite and there's an infinite state space,
- 34:49which does exist in certain types of problems,
- 34:51then it's a slightly different story.
- 34:53But as long as our maze has finitely many squares,
- 34:56we're going to find a solution.
- 34:58The next question, though, that we want to ask is,
- 35:00is it going to be a good solution?
- 35:02Is it the optimal solution that we can find?
- 35:05And the answer there is not necessarily.
- 35:07And let's take a look at an example of that.
- 35:09In this maze, for example, we're again trying to find our way from A to B.
- 35:14And you notice here there are multiple possible solutions.
- 35:16We could go this way or we could go up in order to make our way from A to B.
- 35:21Now, if we're lucky, depth first search will choose this way and get to B.
- 35:25But there's no reason necessarily why depth first search
- 35:28would choose between going up or going to the right.
- 35:30It's sort of an arbitrary decision point because both
- 35:33are going to be added to the frontier.
- 35:35And ultimately, if we get unlucky, depth first search
- 35:38might choose to explore this path first because it's just a random choice
- 35:42at this point.
- 35:42It'll explore, explore, explore.
- 35:45And it'll eventually find the goal, this particular path,
- 35:48when in actuality there was a better path.
- 35:50There was a more optimal solution that used fewer steps,
- 35:54assuming we're measuring the cost of a solution based on the number of steps
- 35:58that we need to take.
- 35:59So depth first search, if we're unlucky,
- 36:01might end up not finding the best solution when a better solution is
- 36:05available.
- 36:07So that's DFS, depth first search.
- 36:09How does BFS, or breadth first search, compare?
- 36:12How would it work in this particular situation?
- 36:14Well, the algorithm is going to look very different visually
- 36:17in terms of how BFS explores.
- 36:20Because BFS looks at shallower nodes first, the idea is going to be,
- 36:24BFS will first look at all of the nodes that are one away from the initial state.
- 36:29Look here and look here, for example, just
- 36:31at the two nodes that are immediately next to this initial state.
- 36:36Then it'll explore nodes that are two away,
- 36:37looking at this state and that state, for example.
- 36:40Then it'll explore nodes that are three away, this state and that state.
- 36:43Whereas depth first search just picked one path and kept following it,
- 36:47breadth first search, on the other hand,
- 36:49is taking the option of exploring all of the possible paths
- 36:52as kind of at the same time bouncing back between them,
- 36:56looking deeper and deeper at each one, but making sure
- 36:58to explore the shallower ones or the ones that
- 37:01are closer to the initial state earlier.
- 37:04So we'll keep following this pattern, looking at things that are four away,
- 37:07looking at things that are five away, looking at things that are six away,
- 37:10until eventually we make our way to the goal.
- 37:14And in this case, it's true we had to explore some states that ultimately
- 37:17didn't lead us anywhere, but the path that we found to the goal
- 37:20was the optimal path.
- 37:22This is the shortest way that we could get to the goal.
- 37:25And so what might happen then in a larger maze?
- 37:28Well, let's take a look at something like this
- 37:30and how breadth first search is going to behave.
- 37:32Well, breadth first search, again, we'll just keep following the states
- 37:35until it receives a decision point.
- 37:37It could go either left or right.
- 37:39And while DFS just picked one and kept following that until it hit a dead end,
- 37:44BFS, on the other hand, will explore both.
- 37:47It'll say look at this node, then this node,
- 37:50and it'll look at this node, then that node.
- 37:52So on and so forth.
- 37:53And when it hits a decision point here, rather than pick one left or two
- 37:57right and explore that path, it'll again explore both,
- 38:01alternating between them, going deeper and deeper.
- 38:03We'll explore here, and then maybe here and here, and then keep going.
- 38:07Explore here and slowly make our way, you can visually
- 38:10see, further and further out.
- 38:12Once we get to this decision point, we'll explore both up and down
- 38:16until ultimately we make our way to the goal.
- 38:21And what you'll notice is, yes, breadth first search
- 38:24did find our way from A to B by following this particular path,
- 38:28but it needed to explore a lot of states in order to do so.
- 38:32And so we see some trade offs here between DFS and BFS,
- 38:35that in DFS, there may be some cases where there is some memory savings
- 38:39as compared to a breadth first approach, where breadth first search in this case
- 38:43had to explore a lot of states.
- 38:45But maybe that won't always be the case.
- 38:48So now let's actually turn our attention to some code
- 38:51and look at the code that we could actually
- 38:52write in order to implement something like depth first search or breadth
- 38:56first search in the context of solving a maze, for example.
- 39:01So I'll go ahead and go into my terminal.
- 39:03And what I have here inside of maze.py is an implementation
- 39:07of this same idea of maze solving.
- 39:09I've defined a class called node that in this case
- 39:12is keeping track of the state, the parent, in other words,
- 39:15the state before the state, and the action.
- 39:17In this case, we're not keeping track of the path cost
- 39:20because we can calculate the cost of the path at the end
- 39:22after we found our way from the initial state to the goal.
- 39:26In addition to this, I've defined a class called a stack frontier.
- 39:31And if unfamiliar with a class, a class is a way for me
- 39:34to define a way to generate objects in Python.
- 39:37It refers to an idea of object oriented programming, where the idea here
- 39:42is that I would like to create an object that is
- 39:44able to store all of my frontier data.
- 39:46And I would like to have functions, otherwise known
- 39:49as methods, on that object that I can use to manipulate the object.
- 39:53And so what's going on here, if unfamiliar with the syntax,
- 39:57is I have a function that initially creates a frontier that I'm
- 40:00going to represent using a list.
- 40:02And initially, my frontier is represented by the empty list.
- 40:05There's nothing in my frontier to begin with.
- 40:08I have an add function that adds something to the frontier
- 40:12as by appending it to the end of the list.
- 40:15I have a function that checks if the frontier contains
- 40:17a particular state.
- 40:19I have an empty function that checks if the frontier is empty.
- 40:22If the frontier is empty, that just means the length of the frontier is 0.
- 40:26And then I have a function for removing something from the frontier.
- 40:29I can't remove something from the frontier if the frontier is empty,
- 40:32so I check for that first.
- 40:33But otherwise, if the frontier isn't empty,
- 40:36recall that I'm implementing this frontier as a stack, a last in first
- 40:41out data structure, which means the last thing I add to the frontier,
- 40:45in other words, the last thing in the list, is the item
- 40:48that I should remove from this frontier.
- 40:51So what you'll see here is I have removed the last item of a list.
- 40:56And if you index into a Python list with negative 1,
- 40:59that gets you the last item in the list.
- 41:01Since 0 is the first item, negative 1 kind of wraps around
- 41:04and gets you to the last item in the list.
- 41:07So we give that the node.
- 41:09We call that node.
- 41:10We update the frontier here on line 28 to say,
- 41:12go ahead and remove that node that you just removed from the frontier.
- 41:16And then we return the node as a result.
- 41:18So this class here effectively implements the idea of a frontier.
- 41:23It gives me a way to add something to a frontier
- 41:25and a way to remove something from the frontier as a stack.
- 41:29I've also, just for good measure, implemented
- 41:31an alternative version of the same thing called a queue frontier, which
- 41:36in parentheses you'll see here, it inherits from a stack frontier,
- 41:39meaning it's going to do all the same things that the stack frontier did,
- 41:42except the way we remove a node from the frontier
- 41:45is going to be slightly different.
- 41:47Instead of removing from the end of the list the way we would in a stack,
- 41:50we're instead going to remove from the beginning of the list.
- 41:53Self.frontier 0 will get me the first node in the frontier, the first one
- 41:58that was added, and that is going to be the one
- 42:00that we return in the case of a queue.
- 42:03Then under here, I have a definition of a class called maze.
- 42:06This is going to handle the process of taking a sequence, a maze-like text
- 42:11file, and figuring out how to solve it.
- 42:13So it will take as input a text file that looks something like this,
- 42:16for example, where we see hash marks that are here representing walls,
- 42:20and I have the character A representing the starting position
- 42:23and the character B representing the ending position.
- 42:27And you can take a look at the code for parsing this text file right now.
- 42:30That's the less interesting part.
- 42:32The more interesting part is this solve function here,
- 42:35the solve function is going to figure out
- 42:37how to actually get from point A to point B.
- 42:41And here we see an implementation of the exact same idea
- 42:44we saw from a moment ago.
- 42:45We're going to keep track of how many states we've explored,
- 42:48just so we can report that data later.
- 42:50But I start with a node that represents just the start state.
- 42:55And I start with a frontier that, in this case, is a stack frontier.
- 43:00And given that I'm treating my frontier as a stack,
- 43:02you might imagine that the algorithm I'm using here is now depth-first search,
- 43:06because depth-first search, or DFS, uses a stack as its data structure.
- 43:11And initially, this frontier is just going to contain the start state.
- 43:16We initialize an explored set that initially is empty.
- 43:19There's nothing we've explored so far.
- 43:21And now here's our loop, that notion of repeating something again and again.
- 43:25First, we check if the frontier is empty by calling that empty function
- 43:29that we saw the implementation of a moment ago.
- 43:31And if the frontier is indeed empty, we'll
- 43:34go ahead and raise an exception, or a Python error, to say,
- 43:37sorry, there is no solution to this problem.
- 43:41Otherwise, we'll go ahead and remove a node from the frontier
- 43:44as by calling frontier.remove and update the number of states we've explored,
- 43:48because now we've explored one additional state.
- 43:51So we say self.numexplored plus equals 1, adding 1
- 43:55to the number of states we've explored.
- 43:57Once we remove a node from the frontier,
- 44:00recall that the next step is to see whether or not
- 44:02it's the goal, the goal test.
- 44:04And in the case of the maze, the goal is pretty easy.
- 44:06I check to see whether the state of the node is equal to the goal.
- 44:11Initially, when I set up the maze, I set up
- 44:13this value called goal, which is a property of the maze,
- 44:15so I can just check to see if the node is actually the goal.
- 44:19And if it is the goal, then what I want to do
- 44:22is backtrack my way towards figuring out what actions I took in order
- 44:26to get to this goal.
- 44:28And how do I do that?
- 44:29We'll recall that every node stores its parent, the node that came before it
- 44:33that we used to get to this node, and also the action used in order to get
- 44:37there.
- 44:37So I can create this loop where I'm constantly just looking
- 44:40at the parent of every node and keeping track for all of the parents
- 44:44what action I took to get from the parent to this current node.
- 44:47So this loop is going to keep repeating this process
- 44:50of looking through all of the parent nodes
- 44:52until we get back to the initial state, which
- 44:54has no parent, where node.parent is going to be equal to none.
- 44:59As I do so, I'm going to be building up the list of all of the actions
- 45:01that I'm following and the list of all the cells that are part of the solution.
- 45:05But I'll reverse them because when I build it up,
- 45:08going from the goal back to the initial state
- 45:10and building the sequence of actions from the goal to the initial state,
- 45:14but I want to reverse them in order to get the sequence of actions
- 45:16from the initial state to the goal.
- 45:19And that is ultimately going to be the solution.
- 45:23So all of that happens if the current state is equal to the goal.
- 45:27And otherwise, if it's not the goal, well,
- 45:29then I'll go ahead and add this state to the explored set to say,
- 45:32I've explored this state now.
- 45:34No need to go back to it if I come across it in the future.
- 45:37And then this logic here implements the idea of adding neighbors to the frontier.
- 45:42I'm saying, look at all of my neighbors, and I
- 45:44implemented a function called neighbors that you can take a look at.
- 45:47And for each of those neighbors, I'm going to check,
- 45:49is the state already in the frontier?
- 45:51Is the state already in the explored set?
- 45:54And if it's not in either of those, then I'll go ahead and add this new child
- 45:58node, this new node, to the frontier.
- 46:01So there's a fair amount of syntax here,
- 46:03but the key here is not to understand all the nuances of the syntax.
- 46:05So feel free to take a closer look at this file on your own
- 46:08to get a sense for how it is working.
- 46:10But the key is to see how this is an implementation
- 46:13of the same pseudocode, the same idea that we were describing a moment ago
- 46:16on the screen when we were looking at the steps
- 46:19that we might follow in order to solve this kind of search problem.
- 46:23So now let's actually see this in action.
- 46:25I'll go ahead and run maze.py on maze1.txt, for example.
- 46:31And what we'll see is here, we have a printout
- 46:34of what the maze initially looked like.
- 46:36And then here down below is after we've solved it.
- 46:39We had to explore 11 states in order to do it,
- 46:41and we found a path from A to B. And in this program,
- 46:45I just happened to generate a graphical representation of this as well.
- 46:48So I can open up maze.png, which is generated
- 46:50by this program, that shows you where in the darker color here are the walls,
- 46:54red is the initial state, green is the goal,
- 46:56and yellow is the path that was followed.
- 46:58We found a path from the initial state to the goal.
- 47:03But now let's take a look at a more sophisticated maze
- 47:06to see what might happen instead.
- 47:08Let's look now at maze2.txt.
- 47:10We're now here.
- 47:11We have a much larger maze.
- 47:13Again, we're trying to find our way from point A to point B.
- 47:16But now you'll imagine that depth-first search might not be so lucky.
- 47:19It might not get the goal on the first try.
- 47:22It might have to follow one path, then backtrack and explore something else
- 47:26a little bit later.
- 47:28So let's try this.
- 47:29We'll run python maze.py of maze2.txt, this time trying on this other maze.
- 47:34And now, depth-first search is able to find a solution.
- 47:38Here, as indicated by the stars, is a way to get from A to B.
- 47:42And we can represent this visually by opening up this maze.
- 47:45Here's what that maze looks like, and highlighted in yellow
- 47:48is the path that was found from the initial state to the goal.
- 47:52But how many states did we have to explore before we found that path?
- 47:57Well, recall that in my program, I was keeping
- 47:59track of the number of states that we've explored so far.
- 48:02And so I can go back to the terminal and see that, all right,
- 48:05in order to solve this problem, we had to explore 399 different states.
- 48:12And in fact, if I make one small modification of the program
- 48:14and tell the program at the end when we output this image,
- 48:17I added an argument called show explored.
- 48:21And if I set show explored equal to true and rerun this program,
- 48:26python maze.py, running it on maze2, and then I open the maze, what you'll see
- 48:30here is highlighted in red are all of the states
- 48:33that had to be explored to get from the initial state to the goal.
- 48:37Depth-first search, or DFS, didn't find its way to the goal right away.
- 48:41It made a choice to first explore this direction.
- 48:44And when it explored this direction, it had
- 48:46to follow every conceivable path all the way to the very end,
- 48:49even this long and winding one, in order to realize that, you know what?
- 48:52That's a dead end.
- 48:53And instead, the program needed to backtrack.
- 48:55After going this direction, it must have gone this direction.
- 48:58It got lucky here by just not choosing this path,
- 49:01but it got unlucky here, exploring this direction, exploring a bunch of states
- 49:05it didn't need to, and then likewise exploring
- 49:07all of this top part of the graph when it probably
- 49:10didn't need to do that either.
- 49:12So all in all, depth-first search here really not performing optimally,
- 49:16or probably exploring more states than it needs to.
- 49:19It finds an optimal solution, the best path to the goal,
- 49:22but the number of states needed to explore in order to do so,
- 49:25the number of steps I had to take, that was much higher.
- 49:29So let's compare.
- 49:30How would breadth-first search, or BFS, do on this exact same maze instead?
- 49:35And in order to do so, it's a very easy change.
- 49:37The algorithm for DFS and BFS is identical with the exception
- 49:42of what data structure we use to represent the frontier,
- 49:47that in DFS, I used a stack frontier, last in, first out,
- 49:51whereas in BFS, I'm going to use a queue frontier, first in, first out,
- 49:57where the first thing I add to the frontier is the first thing that I
- 50:00remove.
- 50:01So I'll go back to the terminal, rerun this program on the same maze,
- 50:06and now you'll see that the number of states
- 50:08we had to explore was only 77 as compared to almost 400
- 50:13when we used depth-first search.
- 50:15And we can see exactly why.
- 50:16We can see what happened if we open up maze.png now and take a look.
- 50:21Again, yellow highlight is the solution that breadth-first search found,
- 50:25which incidentally is the same solution that depth-first search found.
- 50:29They're both finding the best solution.
- 50:31But notice all the white unexplored cells.
- 50:33There was much fewer states that needed to be explored in order
- 50:37to make our way to the goal because breadth-first search operates
- 50:41a little more shallowly.
- 50:42It's exploring things that are close to the initial state
- 50:45without exploring things that are further away.
- 50:48So if the goal is not too far away, then breadth-first search
- 50:51can actually behave quite effectively on a maze that
- 50:53looks a little something like this.
- 50:56Now, in this case, both BFS and DFS ended up finding the same solution,
- 51:01but that won't always be the case.
- 51:03And in fact, let's take a look at one more example.
- 51:06For instance, maze3.txt.
- 51:09In maze3.txt, notice that here there are multiple ways
- 51:12that you could get from A to B. It's a relatively small maze,
- 51:16but let's look at what happens.
- 51:18If I use, and I'll go ahead and turn off show explored
- 51:21so we just see the solution.
- 51:24If I use BFS, breadth-first search, to solve maze3.txt,
- 51:30well, then we find a solution, and if I open up the maze,
- 51:33here is the solution that we found.
- 51:35It is the optimal one.
- 51:36With just four steps, we can get from the initial state
- 51:39to what the goal happens to be.
- 51:43But what happens if we tried to use depth-first search or DFS instead?
- 51:47Well, again, I'll go back up to my Q frontier, where Q frontier means
- 51:52that we're using breadth-first search, and I'll change it to a stack frontier,
- 51:57which means that now we'll be using depth-first search.
- 52:00I'll rerun pythonmaze.py, and now you'll see that we find the solution,
- 52:06but it is not the optimal solution.
- 52:09This instead is what our algorithm finds,
- 52:11and maybe depth-first search would have found the solution.
- 52:14It's possible, but it's not guaranteed that if we just
- 52:17happen to be unlucky, if we choose this state instead of that state,
- 52:21then depth-first search might find a longer route
- 52:24to get from the initial state to the goal.
- 52:27So we do see some trade-offs here, where depth-first search might not
- 52:30find the optimal solution.
- 52:32So at that point, it seems like breadth-first search is pretty good.
- 52:35Is that the best we can do, where it's going to find us the optimal solution,
- 52:38and we don't have to worry about situations
- 52:41where we might end up finding a longer path to the solution
- 52:44than what actually exists?
- 52:46Where the goal is far away from the initial state,
- 52:49and we might have to take lots of steps in order
- 52:51to get from the initial state to the goal, what ended up happening
- 52:55is that this algorithm, BFS, ended up exploring basically the entire graph,
- 52:59having to go through the entire maze in order
- 53:01to find its way from the initial state to the goal state.
- 53:05What we'd ultimately like is for our algorithm
- 53:08to be a little bit more intelligent.
- 53:10And now what would it mean for our algorithm to be a little bit more
- 53:13intelligent in this case?
- 53:16Well, let's look back to where breadth-first search might
- 53:18have been able to make a different decision
- 53:20and consider human intuition in this process as well.
- 53:23What might a human do when solving this maze
- 53:26that is different than what BFS ultimately chose to do?
- 53:30Well, the very first decision point that BFS made was right here,
- 53:35when it made five steps and ended up in a position
- 53:38where it had a fork in the row.
- 53:39It could either go left or it could go right.
- 53:41In these initial couple steps, there was no choice.
- 53:44There was only one action that could be taken from each of those states.
- 53:46And so the search algorithm did the only thing
- 53:49that any search algorithm could do, which is keep following that state
- 53:53after the next state.
- 53:54But this decision point is where things get a little bit interesting.
- 53:57Depth-first search, that very first search algorithm we looked at,
- 54:01chose to say, let's pick one path and exhaust that path.
- 54:04See if anything that way has the goal.
- 54:07And if not, then let's try the other way.
- 54:09Depth-first search took the alternative approach of saying,
- 54:12you know what, let's explore things that are shallow, close to us first.
- 54:16Look left and right, then back left and back right, so on and so forth,
- 54:20alternating between our options in the hopes of finding something nearby.
- 54:24But ultimately, what might a human do if confronted
- 54:27with a situation like this of go left or go right?
- 54:30Well, a human might visually see that, all right, I'm
- 54:33trying to get to state b, which is way up there,
- 54:36and going right just feels like it's closer to the goal.
- 54:39It feels like going right should be better than going left
- 54:42because I'm making progress towards getting to that goal.
- 54:45Now, of course, there are a couple of assumptions that I'm making here.
- 54:48I'm making the assumption that we can represent this grid
- 54:51as like a two-dimensional grid where I know the coordinates of everything.
- 54:55I know that a is in coordinate 0, 0, and b is in some other coordinate pair,
- 55:00and I know what coordinate I'm at now.
- 55:01So I can calculate that, yeah, going this way, that is closer to the goal.
- 55:05And that might be a reasonable assumption for some types of search problems,
- 55:08but maybe not in others.
- 55:10But for now, we'll go ahead and assume that,
- 55:12that I know what my current coordinate pair is,
- 55:15and I know the coordinate, x, y, of the goal that I'm trying to get to.
- 55:19And in this situation, I'd like an algorithm
- 55:22that is a little bit more intelligent, that somehow knows
- 55:25that I should be making progress towards the goal,
- 55:28and this is probably the way to do that because in a maze,
- 55:31moving in the coordinate direction of the goal
- 55:34is usually, though not always, a good thing.
- 55:37And so here we draw a distinction between two different types
- 55:40of search algorithms, uninformed search and informed search.
- 55:45Uninformed search algorithms are algorithms like DFS and BFS,
- 55:49the two algorithms that we just looked at, which
- 55:51are search strategies that don't use any problem-specific knowledge
- 55:55to be able to solve the problem.
- 55:57DFS and BFS didn't really care about the structure of the maze
- 56:01or anything about the way that a maze is in order to solve the problem.
- 56:05They just look at the actions available and choose from those actions,
- 56:08and it doesn't matter whether it's a maze or some other problem,
- 56:11the solution or the way that it tries to solve the problem
- 56:14is really fundamentally going to be the same.
- 56:17What we're going to take a look at now is an improvement
- 56:19upon uninformed search.
- 56:21We're going to take a look at informed search.
- 56:24Informed search are going to be search strategies
- 56:26that use knowledge specific to the problem
- 56:29to be able to better find a solution.
- 56:31And in the case of a maze, this problem-specific knowledge
- 56:35is something like if I'm in a square that is geographically closer to the goal,
- 56:40that is better than being in a square that is geographically further away.
- 56:45And this is something we can only know by thinking about this problem
- 56:49and reasoning about what knowledge might be helpful for our AI agent
- 56:54to know a little something about.
- 56:56There are a number of different types of informed search.
- 56:58Specifically, first, we're going to look at a particular type of search
- 57:01algorithm called greedy best-first search.
- 57:05Greedy best-first search, often abbreviated G-BFS,
- 57:08is a search algorithm that instead of expanding the deepest node like DFS
- 57:13or the shallowest node like BFS, this algorithm
- 57:16is always going to expand the node that it thinks is closest to the goal.
- 57:22Now, the search algorithm isn't going to know for sure
- 57:24whether it is the closest thing to the goal.
- 57:27Because if we knew what was closest to the goal all the time,
- 57:29then we would already have a solution.
- 57:31The knowledge of what is close to the goal,
- 57:33we could just follow those steps in order to get from the initial position
- 57:36to the solution.
- 57:37But if we don't know the solution, meaning
- 57:39we don't know exactly what's closest to the goal,
- 57:42instead we can use an estimate of what's closest to the goal,
- 57:46otherwise known as a heuristic, just some way of estimating whether or not
- 57:50we're close to the goal.
- 57:51And we'll do so using a heuristic function conventionally
- 57:54called h of n that takes a status input and returns
- 57:58our estimate of how close we are to the goal.
- 58:03So what might this heuristic function actually
- 58:05look like in the case of a maze solving algorithm?
- 58:08Where we're trying to solve a maze, what does the heuristic look like?
- 58:11Well, the heuristic needs to answer a question
- 58:14between these two cells, C and D, which one is better?
- 58:17Which one would I rather be in if I'm trying to find my way to the goal?
- 58:22Well, any human could probably look at this and tell you,
- 58:24you know what, D looks like it's better.
- 58:26Even if the maze is convoluted and you haven't thought about all the walls,
- 58:29D is probably better.
- 58:31And why is D better?
- 58:32Well, because if you ignore the wall, so let's just pretend
- 58:35the walls don't exist for a moment and relax the problem, so to speak,
- 58:40D, just in terms of coordinate pairs, is closer to this goal.
- 58:44It's fewer steps that I wouldn't take to get to the goal as compared to C,
- 58:49even if you ignore the walls.
- 58:50If you just know the xy-coordinate of C and the xy-coordinate of the goal,
- 58:55and likewise you know the xy-coordinate of D,
- 58:57you can calculate the D just geographically.
- 59:00Ignoring the walls looks like it's better.
- 59:03And so this is the heuristic function that we're going to use.
- 59:05And it's something called the Manhattan distance,
- 59:08one specific type of heuristic, where the heuristic is how many squares
- 59:12vertically and horizontally and then left to right,
- 59:15so not allowing myself to go diagonally, just either up or right
- 59:18or left or down.
- 59:19How many steps do I need to take to get from each of these cells to the goal?
- 59:24Well, as it turns out, D is much closer.
- 59:27There are fewer steps.
- 59:28It only needs to take six steps in order to get to that goal.
- 59:31Again, here, ignoring the walls.
- 59:33We've relaxed the problem a little bit.
- 59:35We're just concerned with if you do the math
- 59:38to subtract the x values from each other and the y values from each other,
- 59:41what is our estimate of how far we are away?
- 59:44We can estimate the D is closer to the goal than C is.
- 59:49And so now we have an approach.
- 59:51We have a way of picking which node to remove from the frontier.
- 59:54And at each stage in our algorithm, we're
- 59:56going to remove a node from the frontier.
- 59:57We're going to explore the node if it has the smallest
- 1:00:00value for this heuristic function, if it has the smallest
- 1:00:04Manhattan distance to the goal.
- 1:00:06And so what would this actually look like?
- 1:00:08Well, let me first label this graph, label this maze,
- 1:00:11with a number representing the value of this heuristic function,
- 1:00:14the value of the Manhattan distance from any of these cells.
- 1:00:18So from this cell, for example, we're one away from the goal.
- 1:00:21From this cell, we're two away from the goal, three away, four away.
- 1:00:24Here, we're five away because we have to go one to the right
- 1:00:27and then four up.
- 1:00:28From somewhere like here, the Manhattan distance is two.
- 1:00:32We're only two squares away from the goal geographically,
- 1:00:35even though in practice, we're going to have to take a longer path.
- 1:00:39But we don't know that yet.
- 1:00:40The heuristic is just some easy way to estimate
- 1:00:42how far we are away from the goal.
- 1:00:44And maybe our heuristic is overly optimistic.
- 1:00:47It thinks that, yeah, we're only two steps away.
- 1:00:49When in practice, when you consider the walls, it might be more steps.
- 1:00:53So the important thing here is that the heuristic isn't a guarantee of how
- 1:00:57many steps it's going to take.
- 1:00:59It is estimating.
- 1:01:01It's an attempt at trying to approximate.
- 1:01:03And it does seem generally the case that the squares that
- 1:01:06look closer to the goal have smaller values for the heuristic function
- 1:01:10than squares that are further away.
- 1:01:13So now, using greedy best-first search, what might this algorithm actually do?
- 1:01:18Well, again, for these first five steps, there's not much of a choice.
- 1:01:21We start at this initial state a, and we say, all right,
- 1:01:23we have to explore these five states.
- 1:01:26But now we have a decision point.
- 1:01:28Now we have a choice between going left and going right.
- 1:01:30And before, when DFS and BFS would just pick arbitrarily,
- 1:01:34because it just depends on the order you throw these two nodes into the frontier,
- 1:01:37and we didn't specify what order you put them into the frontier,
- 1:01:40only the order you take them out, here we can look at 13 and 11
- 1:01:45and say that, all right, this square is a distance of 11 away from the goal
- 1:01:50according to our heuristic, according to our estimate.
- 1:01:53And this one, we estimate to be 13 away from the goal.
- 1:01:57So between those two options, between these two choices,
- 1:02:00I'd rather have the 11.
- 1:02:02I'd rather be 11 steps away from the goal, so I'll go to the right.
- 1:02:06We're able to make an informed decision, because we know a little something
- 1:02:09more about this problem.
- 1:02:11So then we keep following, 10, 9, 8.
- 1:02:13Between the two 7s, we don't really have much of a way to know between those.
- 1:02:17So then we do just have to make an arbitrary choice.
- 1:02:20And you know what, maybe we choose wrong.
- 1:02:21But that's OK, because now we can still say, all right, let's try this 7.
- 1:02:26We say 7, 6, we have to make this choice,
- 1:02:29even though it increases the value of the heuristic function.
- 1:02:31But now we have another decision point, between 6 and 8, and between those two.
- 1:02:36And really, we're also considering this 13, but that's much higher.
- 1:02:39Between 6, 8, and 13, well, the 6 is the smallest value,
- 1:02:43so we'd rather take the 6.
- 1:02:45We're able to make an informed decision that going this way to the right
- 1:02:48is probably better than going down.
- 1:02:51So we turn this way, we go to 5.
- 1:02:53And now we find a decision point where we'll actually
- 1:02:55make a decision that we might not want to make,
- 1:02:57but there's unfortunately not too much of a way around this.
- 1:03:00We see 4 and 6.
- 1:03:014 looks closer to the goal, right?
- 1:03:03It's going up, and the goal is further up.
- 1:03:06So we end up taking that route, which ultimately leads us to a dead end.
- 1:03:09But that's OK, because we can still say, all right, now let's try the 6.
- 1:03:13And now follow this route that will ultimately lead us to the goal.
- 1:03:17And so this now is how greedy best-for-search
- 1:03:20might try to approach this problem by saying,
- 1:03:22whenever we have a decision between multiple nodes that we could explore,
- 1:03:26let's explore the node that has the smallest value of h of n,
- 1:03:30this heuristic function that is estimating how far I have to go.
- 1:03:35And it just so happens that in this case, we end up
- 1:03:37doing better in terms of the number of states we needed to explore
- 1:03:41than BFS needed to.
- 1:03:42BFS explored all of this section and all of that section,
- 1:03:46but we were able to eliminate that by taking advantage of this heuristic,
- 1:03:49this knowledge about how close we are to the goal or some estimate of that idea.
- 1:03:56So this seems much better.
- 1:03:57So wouldn't we always prefer an algorithm like this over an algorithm
- 1:04:01like breadth-first search?
- 1:04:03Well, maybe one thing to take into consideration
- 1:04:05is that we need to come up with a good heuristic, how good the heuristic is,
- 1:04:09is going to affect how good this algorithm is.
- 1:04:11And coming up with a good heuristic can oftentimes be challenging.
- 1:04:16But the other thing to consider is to ask the question,
- 1:04:18just as we did with the prior two algorithms, is this algorithm optimal?
- 1:04:22Will it always find the shortest path from the initial state to the goal?
- 1:04:28And to answer that question, let's take a look at this example for a moment.
- 1:04:32Take a look at this example.
- 1:04:33Again, we're trying to get from A to B. And again,
- 1:04:36I've labeled each of the cells with their Manhattan distance from the goal.
- 1:04:40The number of squares up and to the right,
- 1:04:42you would need to travel in order to get from that square to the goal.
- 1:04:46And let's think about, would greedy best-first search
- 1:04:49that always picks the smallest number end up finding the optimal solution?
- 1:04:55What is the shortest solution?
- 1:04:57And would this algorithm find it?
- 1:04:59And the important thing to realize is that right here is the decision point.
- 1:05:04We're estimated to be 12 away from the goal.
- 1:05:06And we have two choices.
- 1:05:08We can go to the left, which we estimate to be 13 away from the goal.
- 1:05:11Or we can go up, where we estimate it to be 11 away from the goal.
- 1:05:15And between those two, greedy best-first search
- 1:05:18is going to say the 11 looks better than the 13.
- 1:05:23And in doing so, greedy best-first search will end up
- 1:05:26finding this path to the goal.
- 1:05:28But it turns out this path is not optimal.
- 1:05:31There is a way to get to the goal using fewer steps.
- 1:05:33And it's actually this way, this way that ultimately involved fewer steps,
- 1:05:38even though it meant at this moment choosing the worst option between the two
- 1:05:43or what we estimated to be the worst option based on the heuristics.
- 1:05:47And so this is what we mean by this is a greedy algorithm.
- 1:05:50It's making the best decision locally.
- 1:05:52At this decision point, it looks like it's better to go here
- 1:05:55than it is to go to the 13.
- 1:05:57But in the big picture, it's not necessarily optimal.
- 1:06:00That it might find a solution when in actuality,
- 1:06:03there was a better solution available.
- 1:06:06So we would like some way to solve this problem.
- 1:06:09We like the idea of this heuristic, of being
- 1:06:12able to estimate the path, the distance between us and the goal.
- 1:06:16And that helps us to be able to make better decisions
- 1:06:18and to eliminate having to search through entire parts of this state space.
- 1:06:23But we would like to modify the algorithm so that we can achieve optimality,
- 1:06:27so that it can be optimal.
- 1:06:28And what is the way to do this?
- 1:06:30What is the intuition here?
- 1:06:31Well, let's take a look at this problem.
- 1:06:34In this initial problem, greedy best research
- 1:06:37found us this solution here, this long path.
- 1:06:40And the reason why it wasn't great is because, yes, the heuristic numbers
- 1:06:43went down pretty low.
- 1:06:44But later on, they started to build back up.
- 1:06:47They built back 8, 9, 10, 11, all the way up to 12 in this case.
- 1:06:52And so how might we go about trying to improve this algorithm?
- 1:06:55Well, one thing that we might realize is that if we go all the way
- 1:06:59through this algorithm, through this path, and we end up going to the 12,
- 1:07:03and we've had to take this many steps, who knows how many steps that is,
- 1:07:06just to get to this 12, we could have also, as an alternative,
- 1:07:11taken much fewer steps, just six steps, and ended up at this 13 here.
- 1:07:16And yes, 13 is more than 12, so it looks like it's not as good.
- 1:07:19But it required far fewer steps.
- 1:07:22It only took six steps to get to this 13 versus many more steps
- 1:07:25to get to this 12.
- 1:07:27And while greedy best research says, oh, well, 12 is better than 13,
- 1:07:30so pick the 12, we might more intelligently say,
- 1:07:33I'd rather be somewhere that heuristically looks
- 1:07:37like it takes slightly longer if I can get there much more quickly.
- 1:07:42And we're going to encode that idea, this general idea,
- 1:07:45into a more formal algorithm known as A star search.
- 1:07:49A star search is going to solve this problem
- 1:07:51by instead of just considering the heuristic,
- 1:07:54also considering how long it took us to get to any particular state.
- 1:07:58So the distinction is greedy best for search.
- 1:08:01If I am in a state right now, the only thing I care about
- 1:08:04is, what is the estimated distance, the heuristic value,
- 1:08:07between me and the goal?
- 1:08:09Whereas A star search will take into consideration
- 1:08:11two pieces of information.
- 1:08:13It'll take into consideration, how far do I estimate I am from the goal?
- 1:08:17But also, how far did I have to travel in order to get here?
- 1:08:21Because that is relevant, too.
- 1:08:23So we'll search algorithms by expanding the node
- 1:08:26with the lowest value of g of n plus h of n.
- 1:08:30h of n is that same heuristic that we were talking about a moment ago that's
- 1:08:33going to vary based on the problem.
- 1:08:35But g of n is going to be the cost to reach the node, how many steps
- 1:08:40I had to take, in this case, to get to my current position.
- 1:08:45So what does that search algorithm look like in practice?
- 1:08:48Well, let's take a look.
- 1:08:49Again, we've got the same maze.
- 1:08:51And again, I've labeled them with their Manhattan distance.
- 1:08:54This value is the h of n value, the heuristic
- 1:08:57estimate of how far each of these squares is away from the goal.
- 1:09:02But now, as we begin to explore states, we
- 1:09:04care not just about this heuristic value, but also about g of n,
- 1:09:08the number of steps I had to take in order to get there.
- 1:09:11And I care about summing those two numbers together.
- 1:09:14So what does that look like?
- 1:09:15On this very first step, I have taken one step.
- 1:09:19And now I am estimated to be 16 steps away from the goal.
- 1:09:22So the total value here is 17.
- 1:09:25Then I take one more step.
- 1:09:26I've now taken two steps.
- 1:09:28And I estimate myself to be 15 away from the goal, again, a total value of 17.
- 1:09:32Now I've taken three steps.
- 1:09:34And I'm estimated to be 14 away from the goal, so on and so forth.
- 1:09:37Four steps, an estimate of 13.
- 1:09:39Five steps, estimate of 12.
- 1:09:41And now here's a decision point.
- 1:09:44I could either be six steps away from the goal with a heuristic of 13
- 1:09:48for a total of 19, or I could be six steps away
- 1:09:52from the goal with a heuristic of 11 with an estimate of 17 for the total.
- 1:09:57So between 19 and 17, I'd rather take the 17, the 6 plus 11.
- 1:10:03So so far, no different than what we saw before.
- 1:10:05We're still taking this option because it appears to be better.
- 1:10:08And I keep taking this option because it appears to be better.
- 1:10:11But it's right about here that things get a little bit different.
- 1:10:15Now I could be 15 steps away from the goal with an estimated distance of 6.
- 1:10:21So 15 plus 6, total value of 21.
- 1:10:24Alternatively, I could be six steps away from the goal,
- 1:10:28because this is five steps away, so this is six steps away,
- 1:10:30with a total value of 13 as my estimate.
- 1:10:33So 6 plus 13, that's 19.
- 1:10:36So here, we would evaluate g of n plus h of n to be 19, 6 plus 13.
- 1:10:41Whereas here, we would be 15 plus 6, or 21.
- 1:10:46And so the intuition is 19 less than 21, pick here.
- 1:10:49But the idea is ultimately I'd rather be having taken fewer steps, get to a 13,
- 1:10:55than having taken 15 steps and be at a 6, because it
- 1:10:59means I've had to take more steps in order to get there.
- 1:11:01Maybe there's a better path this way.
- 1:11:04So instead, we'll explore this route.
- 1:11:07Now if we go one more, this is seven steps plus 14 is 21.
- 1:11:11So between those two, it's sort of a toss-up.
- 1:11:12We might end up exploring that one anyways.
- 1:11:15But after that, as these numbers start to get bigger in the heuristic values,
- 1:11:19and these heuristic values start to get smaller,
- 1:11:21you'll find that we'll actually keep exploring down this path.
- 1:11:25And you can do the math to see that at every decision point,
- 1:11:28A star search is going to make a choice based
- 1:11:31on the sum of how many steps it took me to get to my current position,
- 1:11:35and then how far I estimate I am from the goal.
- 1:11:39So while we did have to explore some of these states,
- 1:11:41the ultimate solution we found was, in fact, an optimal solution.
- 1:11:46It did find us the quickest possible way to get from the initial state
- 1:11:50to the goal.
- 1:11:51And it turns out that A star is an optimal search algorithm
- 1:11:55under certain conditions.
- 1:11:57So the conditions are H of n, my heuristic, needs to be admissible.
- 1:12:02What does it mean for a heuristic to be admissible?
- 1:12:04Well, a heuristic is admissible if it never overestimates the true cost.
- 1:12:08H of n always needs to either get it exactly right
- 1:12:12in terms of how far away I am, or it needs to underestimate.
- 1:12:16So we saw an example from before where the heuristic value was much smaller
- 1:12:20than the actual cost it would take.
- 1:12:22That's totally fine, but the heuristic value should never overestimate.
- 1:12:26It should never think that I'm further away from the goal than I actually am.
- 1:12:30And meanwhile, to make a stronger statement, H of n also needs to be
- 1:12:34consistent.
- 1:12:36And what does it mean for it to be consistent?
- 1:12:37Mathematically, it means that for every node, which we'll call n,
- 1:12:41and successor, the node after me, that I'll
- 1:12:43call n prime, where it takes a cost of C to make that step,
- 1:12:48the heuristic value of n needs to be less than or equal to the heuristic
- 1:12:52value of n prime plus the cost.
- 1:12:55So it's a lot of math, but in words what that ultimately means
- 1:12:58is that if I am here at this state right now,
- 1:13:01the heuristic value from me to the goal shouldn't
- 1:13:03be more than the heuristic value of my successor,
- 1:13:07the next place I could go to, plus however much
- 1:13:10it would cost me to just make that step from one step to the next step.
- 1:13:14And so this is just making sure that my heuristic is consistent between all
- 1:13:18of these steps that I might take.
- 1:13:20So as long as this is true, then A star search
- 1:13:22is going to find me an optimal solution.
- 1:13:25And this is where much of the challenge of solving these search problems
- 1:13:28can sometimes come in, that A star search is an algorithm that is known
- 1:13:32and you could write the code fairly easily,
- 1:13:34but it's choosing the heuristic.
- 1:13:35It can be the interesting challenge.
- 1:13:37The better the heuristic is, the better I'll
- 1:13:39be able to solve the problem in the fewer states that I'll have to explore.
- 1:13:43And I need to make sure that the heuristic satisfies
- 1:13:46these particular constraints.
- 1:13:48So all in all, these are some of the examples of search algorithms
- 1:13:52that might work, and certainly there are many more than just this.
- 1:13:55A star, for example, does have a tendency to use quite a bit of memory.
- 1:13:58So there are alternative approaches to A star
- 1:14:01that ultimately use less memory than this version of A star
- 1:14:04happens to use, and there are other search algorithms
- 1:14:07that are optimized for other cases as well.
- 1:14:11But now so far, we've only been looking at search algorithms
- 1:14:14where there is one agent.
- 1:14:17I am trying to find a solution to a problem.
- 1:14:19I am trying to navigate my way through a maze.
- 1:14:22I am trying to solve a 15 puzzle.
- 1:14:24I am trying to find driving directions from point A to point B.
- 1:14:28Sometimes in search situations, though, we'll
- 1:14:30enter an adversarial situation, where I am an agent trying
- 1:14:34to make intelligent decisions.
- 1:14:36And there's someone else who is fighting against me, so to speak,
- 1:14:39that has opposite objectives, someone where
- 1:14:41I am trying to succeed, someone else that wants me to fail.
- 1:14:45And this is most popular in something like a game, a game like Tic Tac Toe,
- 1:14:49where we've got this 3 by 3 grid, and x and o take turns,
- 1:14:53either writing an x or an o in any one of these squares.
- 1:14:56And the goal is to get three x's in a row if you're the x player,
- 1:14:59or three o's in a row if you're the o player.
- 1:15:02And computers have gotten quite good at playing games,
- 1:15:05Tic Tac Toe very easily, but even more complex games.
- 1:15:08And so you might imagine, what does an intelligent decision in a game
- 1:15:12look like?
- 1:15:13So maybe x makes an initial move in the middle, and o plays up here.
- 1:15:17What does an intelligent move for x now become?
- 1:15:20Where should you move if you were x?
- 1:15:22And it turns out there are a couple of possibilities.
- 1:15:24But if an AI is playing this game optimally,
- 1:15:27then the AI might play somewhere like the upper right,
- 1:15:30where in this situation, o has the opposite objective of x.
- 1:15:34x is trying to win the game to get three in a row diagonally here.
- 1:15:37And o is trying to stop that objective, opposite of the objective.
- 1:15:41And so o is going to place here to try to block.
- 1:15:44But now, x has a pretty clever move.
- 1:15:46x can make a move like this, where now x has two possible ways
- 1:15:51that x can win the game.
- 1:15:52x could win the game by getting three in a row across here.
- 1:15:55Or x could win the game by getting three in a row vertically this way.
- 1:15:58So it doesn't matter where o makes their next move.
- 1:16:00o could play here, for example, blocking the three in a row horizontally.
- 1:16:04But then x is going to win the game by getting a three in a row vertically.
- 1:16:09And so there's a fair amount of reasoning that's
- 1:16:11going on here in order for the computer to be able to solve a problem.
- 1:16:14And it's similar in spirit to the problems we've looked at so far.
- 1:16:17There are actions.
- 1:16:19There's some sort of state of the board and some transition
- 1:16:21from one action to the next.
- 1:16:23But it's different in the sense that this is now
- 1:16:25not just a classical search problem, but an adversarial search problem.
- 1:16:29That I am at the x player trying to find the best moves to make,
- 1:16:32but I know that there is some adversary that is trying to stop me.
- 1:16:36So we need some sort of algorithm to deal with these adversarial type of search
- 1:16:41situations.
- 1:16:42And the algorithm we're going to take a look at
- 1:16:44is an algorithm called Minimax, which works very well
- 1:16:47for these deterministic games where there are two players.
- 1:16:51It can work for other types of games as well.
- 1:16:52But we'll look right now at games where I make a move,
- 1:16:55then my opponent makes a move.
- 1:16:56And I am trying to win, and my opponent is trying to win also.
- 1:17:00Or in other words, my opponent is trying to get me to lose.
- 1:17:04And so what do we need in order to make this algorithm work?
- 1:17:07Well, any time we try and translate this human concept of playing a game,
- 1:17:10winning and losing to a computer, we want to translate it
- 1:17:14in terms that the computer can understand.
- 1:17:16And ultimately, the computer really just understands the numbers.
- 1:17:19And so we want some way of translating a game of x's and o's on a grid
- 1:17:23to something numerical, something the computer can understand.
- 1:17:26The computer doesn't normally understand notions of win or lose.
- 1:17:30But it does understand the concept of bigger and smaller.
- 1:17:34And so what we might do is we might take each of the possible ways
- 1:17:38that a tic-tac-toe game can unfold and assign a value or a utility
- 1:17:43to each one of those possible ways.
- 1:17:45And in a tic-tac-toe game, and in many types of games,
- 1:17:47there are three possible outcomes.
- 1:17:49The outcomes are o wins, x wins, or nobody wins.
- 1:17:54So player one wins, player two wins, or nobody wins.
- 1:17:58And for now, let's go ahead and assign each of these possible outcomes
- 1:18:02a different value.
- 1:18:04We'll say o winning, that'll have a value of negative 1.
- 1:18:07Nobody winning, that'll have a value of 0.
- 1:18:09And x winning, that will have a value of 1.
- 1:18:13So we've just assigned numbers to each of these three possible outcomes.
- 1:18:17And now we have two players, we have the x player and the o player.
- 1:18:22And we're going to go ahead and call the x player the max player.
- 1:18:26And we'll call the o player the min player.
- 1:18:29And the reason why is because in the min and max algorithm,
- 1:18:32the max player, which in this case is x, is aiming to maximize the score.
- 1:18:37These are the possible options for the score, negative 1, 0, and 1.
- 1:18:40x wants to maximize the score, meaning if at all possible,
- 1:18:44x would like this situation, where x wins the game,
- 1:18:48and we give it a score of 1.
- 1:18:49But if this isn't possible, if x needs to choose between these two options,
- 1:18:54negative 1, meaning o winning, or 0, meaning nobody winning,
- 1:18:58x would rather that nobody wins, score of 0,
- 1:19:01than a score of negative 1, o winning.
- 1:19:04So this notion of winning and losing and tying
- 1:19:07has been reduced mathematically to just this idea of try and maximize the score.
- 1:19:12The x player always wants the score to be bigger.
- 1:19:16And on the flip side, the min player, in this case o,
- 1:19:19is aiming to minimize the score.
- 1:19:20The o player wants the score to be as small as possible.
- 1:19:25So now we've taken this game of x's and o's and winning and losing
- 1:19:29and turned it into something mathematical,
- 1:19:30something where x is trying to maximize the score,
- 1:19:33o is trying to minimize the score.
- 1:19:35Let's now look at all of the parts of the game
- 1:19:37that we need in order to encode it in an AI
- 1:19:40so that an AI can play a game like tic-tac-toe.
- 1:19:44So the game is going to need a couple of things.
- 1:19:46We'll need some sort of initial state that will, in this case, call s0,
- 1:19:50which is how the game begins, like an empty tic-tac-toe board, for example.
- 1:19:54We'll also need a function called player, where the player function
- 1:20:00is going to take as input a state here represented by s.
- 1:20:04And the output of the player function is going to be which player's turn is it.
- 1:20:09We need to be able to give a tic-tac-toe board to the computer,
- 1:20:12run it through a function, and that function tells us whose turn it is.
- 1:20:16We'll need some notion of actions that we can take.
- 1:20:19We'll see examples of that in just a moment.
- 1:20:21We need some notion of a transition model, same as before.
- 1:20:24If I have a state and I take an action, I
- 1:20:26need to know what results as a consequence of it.
- 1:20:29I need some way of knowing when the game is over.
- 1:20:31So this is equivalent to kind of like a goal test,
- 1:20:34but I need some terminal test, some way to check
- 1:20:36to see if a state is a terminal state, where a terminal state means the game is
- 1:20:40over.
- 1:20:41In a classic game of tic-tac-toe, a terminal state
- 1:20:44means either someone has gotten three in a row
- 1:20:47or all of the squares of the tic-tac-toe board are filled.
- 1:20:50Either of those conditions make it a terminal state.
- 1:20:52In a game of chess, it might be something like when there is checkmate
- 1:20:55or if checkmate is no longer possible, that that becomes a terminal state.
- 1:21:00And then finally, we'll need a utility function, a function that takes a state
- 1:21:04and gives us a numerical value for that terminal state, some way of saying
- 1:21:08if x wins the game, that has a value of 1.
- 1:21:10If o is won the game, that has a value of negative 1.
- 1:21:13If nobody has won the game, that has a value of 0.
- 1:21:16So let's take a look at each of these in turn.
- 1:21:18The initial state, we can just represent in tic-tac-toe as the empty game board.
- 1:21:23This is where we begin.
- 1:21:24It's the place from which we begin this search.
- 1:21:27And again, I'll be representing these things visually,
- 1:21:29but you can imagine this really just being like an array
- 1:21:32or a two-dimensional array of all of these possible squares.
- 1:21:36Then we need the player function that, again, takes a state
- 1:21:39and tells us whose turn it is.
- 1:21:41Assuming x makes the first move, if I have an empty game board,
- 1:21:44then my player function is going to return x.
- 1:21:47And if I have a game board where x has made a move,
- 1:21:49then my player function is going to return o.
- 1:21:52The player function takes a tic-tac-toe game board
- 1:21:54and tells us whose turn it is.
- 1:21:58Next up, we'll consider the actions function.
- 1:22:01The actions function, much like it did in classical search,
- 1:22:04takes a state and gives us the set of all of the possible actions
- 1:22:08we can take in that state.
- 1:22:10So let's imagine it's o is turned to move in a game board that looks like this.
- 1:22:15What happens when we pass it into the actions function?
- 1:22:18So the actions function takes this state of the game as input,
- 1:22:22and the output is a set of possible actions.
- 1:22:25It's a set of I could move in the upper left
- 1:22:27or I could move in the bottom middle.
- 1:22:29So those are the two possible action choices
- 1:22:31that I have when I begin in this particular state.
- 1:22:36Now, just as before, when we had states and actions,
- 1:22:39we need some sort of transition model to tell us
- 1:22:41when we take this action in the state, what is the new state that we get.
- 1:22:45And here, we define that using the result function
- 1:22:48that takes a state as input as well as an action.
- 1:22:51And when we apply the result function to this state,
- 1:22:54saying that let's let o move in this upper left corner,
- 1:22:58the new state we get is this resulting state where o is in the upper left
- 1:23:01corner.
- 1:23:02And now, this seems obvious to someone who knows how to play tic-tac-toe.
- 1:23:04Of course, you play in the upper left corner.
- 1:23:06That's the board you get.
- 1:23:07But all of this information needs to be encoded into the AI.
- 1:23:11The AI doesn't know how to play tic-tac-toe until you
- 1:23:14tell the AI how the rules of tic-tac-toe work.
- 1:23:17And this function, defining this function here,
- 1:23:19allows us to tell the AI how this game actually works
- 1:23:23and how actions actually affect the outcome of the game.
- 1:23:27So the AI needs to know how the game works.
- 1:23:29The AI also needs to know when the game is over,
- 1:23:32as by defining a function called terminal that takes as input a state s,
- 1:23:36such that if we take a game that is not yet over,
- 1:23:39pass it into the terminal function, the output is false.
- 1:23:42The game is not over.
- 1:23:43But if we take a game that is over because x has gotten three in a row
- 1:23:47along that diagonal, pass that into the terminal function,
- 1:23:50then the output is going to be true because the game now is, in fact, over.
- 1:23:55And finally, we've told the AI how the game works
- 1:23:58in terms of what moves can be made and what happens when you make those moves.
- 1:24:01We've told the AI when the game is over.
- 1:24:03Now we need to tell the AI what the value of each of those states is.
- 1:24:07And we do that by defining this utility function that takes a state s
- 1:24:11and tells us the score or the utility of that state.
- 1:24:14So again, we said that if x wins the game, that utility is a value of 1,
- 1:24:18whereas if o wins the game, then the utility of that is negative 1.
- 1:24:23And the AI needs to know, for each of these terminal states
- 1:24:26where the game is over, what is the utility of that state?
- 1:24:30So if I give you a game board like this where the game is, in fact, over,
- 1:24:34and I ask the AI to tell me what the value of that state is, it could do so.
- 1:24:38The value of the state is 1.
- 1:24:42Where things get interesting, though, is if the game is not yet over.
- 1:24:46Let's imagine a game board like this, where in the middle of the game,
- 1:24:49it's o's turn to make a move.
- 1:24:52So how do we know it's o's turn to make a move?
- 1:24:54We can calculate that using the player function.
- 1:24:56We can say player of s, pass in the state, o is the answer.
- 1:25:00So we know it's o's turn to move.
- 1:25:02And now, what is the value of this board and what action should o take?
- 1:25:06Well, that's going to depend.
- 1:25:08We have to do some calculation here.
- 1:25:09And this is where the minimax algorithm really comes in.
- 1:25:13Recall that x is trying to maximize the score, which
- 1:25:16means that o is trying to minimize the score.
- 1:25:19So o would like to minimize the total value
- 1:25:22that we get at the end of the game.
- 1:25:25And because this game isn't over yet, we don't really
- 1:25:27know just yet what the value of this game board is.
- 1:25:30We have to do some calculation in order to figure that out.
- 1:25:34And so how do we do that kind of calculation?
- 1:25:36Well, in order to do so, we're going to consider,
- 1:25:39just as we might in a classical search situation,
- 1:25:41what actions could happen next and what states will that take us to.
- 1:25:46And it turns out that in this position, there are only two open squares,
- 1:25:50which means there are only two open places where o can make a move.
- 1:25:54o could either make a move in the upper left
- 1:25:57or o can make a move in the bottom middle.
- 1:26:00And minimax doesn't know right out of the box which of those moves
- 1:26:03is going to be better.
- 1:26:04So it's going to consider both.
- 1:26:06But now, we sort of run into the same situation.
- 1:26:08Now, I have two more game boards, neither of which is over.
- 1:26:11What happens next?
- 1:26:12And now, it's in this sense that minimax is
- 1:26:14what we'll call a recursive algorithm.
- 1:26:16It's going to now repeat the exact same process,
- 1:26:20although now considering it from the opposite perspective.
- 1:26:23It's as if I am now going to put myself, if I am the o player,
- 1:26:27I'm going to put myself in my opponent's shoes, my opponent as the x player,
- 1:26:31and consider what would my opponent do if they were in this position?
- 1:26:36What would my opponent do, the x player, if they were in that position?
- 1:26:40And what would then happen?
- 1:26:41Well, the other player, my opponent, the x player,
- 1:26:44is trying to maximize the score, whereas I
- 1:26:46am trying to minimize the score as the o player.
- 1:26:49So x is trying to find the maximum possible value that they can get.
- 1:26:53And so what's going to happen?
- 1:26:55Well, from this board position, x only has one choice.
- 1:26:58x is going to play here, and they're going to get three in a row.
- 1:27:01And we know that that board, x winning, that has a value of 1.
- 1:27:05If x wins the game, the value of that game board is 1.
- 1:27:09And so from this position, if this state can only ever
- 1:27:14lead to this state, it's the only possible option,
- 1:27:16and this state has a value of 1, then the maximum possible value
- 1:27:21that the x player can get from this game board is also 1.
- 1:27:24From here, the only place we can get is to a game with a value of 1,
- 1:27:27so this game board also has a value of 1.
- 1:27:31Now we consider this one over here.
- 1:27:33What's going to happen now?
- 1:27:34Well, x needs to make a move.
- 1:27:36The only move x can make is in the upper left, so x will go there.
- 1:27:39And in this game, no one wins the game.
- 1:27:41Nobody has three in a row.
- 1:27:42And so the value of that game board is 0.
- 1:27:45Nobody is 1.
- 1:27:47And so again, by the same logic, if from this board position
- 1:27:50the only place we can get to is a board where the value is 0,
- 1:27:53then this state must also have a value of 0.
- 1:27:57And now here comes the choice part, the idea of trying to minimize.
- 1:28:01I, as the o player, now know that if I make this choice moving in the upper
- 1:28:05left, that is going to result in a game with a value of 1,
- 1:28:09assuming everyone plays optimally.
- 1:28:11And if I instead play in the lower middle,
- 1:28:13choose this fork in the road, that is going
- 1:28:15to result in a game board with a value of 0.
- 1:28:17I have two options.
- 1:28:18I have a 1 and a 0 to choose from, and I need to pick.
- 1:28:22And as the min player, I would rather choose
- 1:28:25the option with the minimum value.
- 1:28:27So whenever a player has multiple choices,
- 1:28:29the min player will choose the option with the smallest value.
- 1:28:32The max player will choose the option with the largest value.
- 1:28:34Between the 1 and the 0, the 0 is smaller,
- 1:28:37meaning I'd rather tie the game than lose the game.
- 1:28:40And so this game board will say also has a value of 0,
- 1:28:44because if I am playing optimally, I will pick this fork in the road.
- 1:28:48I'll place my o here to block x's 3 in a row, x will move in the upper left,
- 1:28:53and the game will be over, and no one will have won the game.
- 1:28:56So this is now the logic of minimax, to consider all of the possible options
- 1:29:00that I can take, all of the actions that I can take,
- 1:29:03and then to put myself in my opponent's shoes.
- 1:29:05I decide what move I'm going to make now by considering
- 1:29:08what move my opponent will make on the next turn.
- 1:29:11And to do that, I consider what move I would make on the turn after that,
- 1:29:14so on and so forth, until I get all the way down
- 1:29:17to the end of the game, to one of these so-called terminal states.
- 1:29:21In fact, this very decision point, where I am trying to decide as the o player
- 1:29:25what to make a decision about, might have just
- 1:29:27been a part of the logic that the x player, my opponent, was using,
- 1:29:31the move before me.
- 1:29:32This might be part of some larger tree, where
- 1:29:35x is trying to make a move in this situation,
- 1:29:37and needs to pick between three different options in order
- 1:29:40to make a decision about what to happen.
- 1:29:42And the further and further away we are from the end of the game,
- 1:29:45the deeper this tree has to go.
- 1:29:47Because every level in this tree is going to correspond to one move,
- 1:29:51one move or action that I take, one move or action
- 1:29:55that my opponent takes, in order to decide what happens.
- 1:29:58And in fact, it turns out that if I am the x player in this position,
- 1:30:02and I recursively do the logic, and see I have a choice, three choices,
- 1:30:05in fact, one of which leads to a value of 0.
- 1:30:08If I play here, and if everyone plays optimally, the game will be a tie.
- 1:30:12If I play here, then o is going to win, and I'll lose playing optimally.
- 1:30:17Or here, where I, the x player, can win, well between a score of 0,
- 1:30:21and negative 1, and 1, I'd rather pick the board with a value of 1,
- 1:30:25because that's the maximum value I can get.
- 1:30:27And so this board would also have a maximum value of 1.
- 1:30:31And so this tree can get very, very deep, especially as the game
- 1:30:35starts to have more and more moves.
- 1:30:37And this logic works not just for tic-tac-toe,
- 1:30:39but any of these sorts of games, where I make a move,
- 1:30:41my opponent makes a move, and ultimately, we
- 1:30:44have these adversarial objectives.
- 1:30:46And we can simplify the diagram into a diagram that looks like this.
- 1:30:50This is a more abstract version of the minimax tree,
- 1:30:53where these are each states, but I'm no longer representing them
- 1:30:56as exactly like tic-tac-toe boards.
- 1:30:57This is just representing some generic game that might be tic-tac-toe,
- 1:31:01might be some other game altogether.
- 1:31:04Any of these green arrows that are pointing up,
- 1:31:06that represents a maximizing state.
- 1:31:08I would like the score to be as big as possible.
- 1:31:11And any of these red arrows pointing down,
- 1:31:13those are minimizing states, where the player is the min player,
- 1:31:16and they are trying to make the score as small as possible.
- 1:31:20So if you imagine in this situation, I am the maximizing player, this player
- 1:31:24here, and I have three choices.
- 1:31:26One choice gives me a score of 5, one choice gives me a score of 3,
- 1:31:30and one choice gives me a score of 9.
- 1:31:32Well, then between those three choices, my best option
- 1:31:36is to choose this 9 over here, the score that
- 1:31:38maximizes my options out of all the three options.
- 1:31:42And so I can give this state a value of 9, because among my three options,
- 1:31:46that is the best choice that I have available to me.
- 1:31:50So that's my decision now.
- 1:31:51You imagine it's like one move away from the end of the game.
- 1:31:55But then you could also ask a reasonable question,
- 1:31:57what might my opponent do two moves away from the end of the game?
- 1:32:01My opponent is the minimizing player.
- 1:32:03They are trying to make the score as small as possible.
- 1:32:05Imagine what would have happened if they had to pick which choice to make.
- 1:32:09One choice leads us to this state, where I, the maximizing player,
- 1:32:13am going to opt for 9, the biggest score that I can get.
- 1:32:16And 1 leads to this state, where I, the maximizing player,
- 1:32:21would choose 8, which is then the largest score that I can get.
- 1:32:25Now the minimizing player, forced to choose between a 9 or an 8,
- 1:32:28is going to choose the smallest possible score,
- 1:32:31which in this case is an 8.
- 1:32:33And that is then how this process would unfold,
- 1:32:35that the minimizing player in this case considers both of their options,
- 1:32:39and then all of the options that would happen as a result of that.
- 1:32:43So this now is a general picture of what the minimax algorithm looks like.
- 1:32:47Let's now try to formalize it using a little bit of pseudocode.
- 1:32:50So what exactly is happening in the minimax algorithm?
- 1:32:53Well, given a state s, we need to decide what to happen.
- 1:32:57The max player, if it's max's player's turn,
- 1:33:00then max is going to pick an action a in actions of s.
- 1:33:05Recall that actions is a function that takes a state
- 1:33:08and gives me back all of the possible actions that I can take.
- 1:33:11It tells me all of the moves that are possible.
- 1:33:15The max player is going to specifically pick an action a in this set of actions
- 1:33:19that gives me the highest value of min value of result of s and a.
- 1:33:26So what does that mean?
- 1:33:27Well, it means that I want to make the option that
- 1:33:30gives me the highest score of all of the actions a.
- 1:33:34But what score is that going to have?
- 1:33:35To calculate that, I need to know what my opponent, the min player,
- 1:33:38is going to do if they try to minimize the value of the state that results.
- 1:33:44So we say, what state results after I take this action?
- 1:33:48And what happens when the min player tries to minimize the value of that state?
- 1:33:53I consider that for all of my possible options.
- 1:33:56And after I've considered that for all of my possible options,
- 1:33:58I pick the action a that has the highest value.
- 1:34:02Likewise, the min player is going to do the same thing but backwards.
- 1:34:06They're also going to consider what are all of the possible actions they
- 1:34:09can take if it's their turn.
- 1:34:10And they're going to pick the action a that
- 1:34:12has the smallest possible value of all the options.
- 1:34:16And the way they know what the smallest possible value of all the options
- 1:34:19is is by considering what the max player is going to do by saying,
- 1:34:24what's the result of applying this action to the current state?
- 1:34:27And then what would the max player try to do?
- 1:34:29What value would the max player calculate for that particular state?
- 1:34:34So everyone makes their decision based on trying
- 1:34:36to estimate what the other person would do.
- 1:34:39And now we need to turn our attention to these two functions, max value
- 1:34:43and min value.
- 1:34:44How do you actually calculate the value of a state
- 1:34:47if you're trying to maximize its value?
- 1:34:50And how do you calculate the value of a state
- 1:34:52if you're trying to minimize the value?
- 1:34:53If you can do that, then we have an entire implementation
- 1:34:56of this min and max algorithm.
- 1:34:58So let's try it.
- 1:34:59Let's try and implement this max value function that takes a state
- 1:35:03and returns as output the value of that state
- 1:35:06if I'm trying to maximize the value of the state.
- 1:35:10Well, the first thing I can check for is to see if the game is over.
- 1:35:13Because if the game is over, in other words,
- 1:35:14if the state is a terminal state, then this is easy.
- 1:35:18I already have this utility function that tells me
- 1:35:21what the value of the board is.
- 1:35:22If the game is over, I just check, did x win, did o win, is it a tie?
- 1:35:26And this utility function just knows what the value of the state is.
- 1:35:30What's trickier is if the game isn't over.
- 1:35:32Because then I need to do this recursive reasoning about thinking,
- 1:35:35what is my opponent going to do on the next move?
- 1:35:39And I want to calculate the value of this state.
- 1:35:41And I want the value of the state to be as high as possible.
- 1:35:45And I'll keep track of that value in a variable called v.
- 1:35:48And if I want the value to be as high as possible,
- 1:35:50I need to give v an initial value.
- 1:35:53And initially, I'll just go ahead and set it to be as low as possible.
- 1:35:57Because I don't know what options are available to me yet.
- 1:36:00So initially, I'll set v equal to negative infinity, which
- 1:36:04seems a little bit strange.
- 1:36:06But the idea here is I want the value initially
- 1:36:08to be as low as possible.
- 1:36:09Because as I consider my actions, I'm always
- 1:36:12going to try and do better than v. And if I set v to negative infinity,
- 1:36:16I know I can always do better than that.
- 1:36:19So now I consider my actions.
- 1:36:21And this is going to be some kind of loop
- 1:36:22where for every action in actions of state,
- 1:36:26recall actions as a function that takes my state
- 1:36:29and gives me all the possible actions that I can use in that state.
- 1:36:32So for each one of those actions, I want to compare it to v and say,
- 1:36:37all right, v is going to be equal to the maximum of v and this expression.
- 1:36:44So what is this expression?
- 1:36:46Well, first it is get the result of taking the action in the state
- 1:36:50and then get the min value of that.
- 1:36:54In other words, let's say I want to find out from that state
- 1:36:58what is the best that the min player can do because they're
- 1:37:00going to try and minimize the score.
- 1:37:02So whatever the resulting score is of the min value of that state,
- 1:37:06compare it to my current best value and just pick the maximum of those two
- 1:37:10because I am trying to maximize the value.
- 1:37:12In short, what these three lines of code are doing
- 1:37:14are going through all of my possible actions and asking the question,
- 1:37:18how do I maximize the score given what my opponent is going to try to do?
- 1:37:24After this entire loop, I can just return v
- 1:37:26and that is now the value of that particular state.
- 1:37:30And for the min player, it's the exact opposite of this,
- 1:37:32the same logic just backwards.
- 1:37:35To calculate the minimum value of a state,
- 1:37:37first we check if it's a terminal state.
- 1:37:38If it is, we return its utility.
- 1:37:41Otherwise, we're going to now try to minimize the value of the state
- 1:37:45given all of my possible actions.
- 1:37:47So I need an initial value for v, the value of the state.
- 1:37:50And initially, I'll set it to infinity because I
- 1:37:53know I can always get something less than infinity.
- 1:37:56So by starting with v equals infinity, I make sure that the very first action
- 1:38:00I find, that will be less than this value of v.
- 1:38:03And then I do the same thing, loop over all of my possible actions.
- 1:38:07And for each of the results that we could get when the max player makes
- 1:38:10their decision, let's take the minimum of that and the current value of v.
- 1:38:15So after all is said and done, I get the smallest possible value of v
- 1:38:19that I then return back to the user.
- 1:38:22So that, in effect, is the pseudocode for Minimax.
- 1:38:25That is how we take a gain and figure out what the best move to make
- 1:38:28is by recursively using these max value and min value functions,
- 1:38:32where max value calls min value, min value calls max value back and forth,
- 1:38:36all the way until we reach a terminal state, at which point
- 1:38:39our algorithm can simply return the utility of that particular state.
- 1:38:45So what you might imagine is that this is going to start to be a long process,
- 1:38:48especially as games start to get more complex,
- 1:38:51as we start to add more moves and more possible options and games that
- 1:38:54might last quite a bit longer.
- 1:38:56So the next question to ask is, what sort of optimizations can we make here?
- 1:39:00How can we do better in order to use less space or take less time
- 1:39:05to be able to solve this kind of problem?
- 1:39:08And we'll take a look at a couple of possible optimizations.
- 1:39:10But for one, we'll take a look at this example.
- 1:39:13Again, returning to these up arrows and down arrows,
- 1:39:15let's imagine that I now am the max player, this green arrow.
- 1:39:20I am trying to make this score as high as possible.
- 1:39:23And this is an easy game where there are just two moves.
- 1:39:26I make a move, one of these three options.
- 1:39:29And then my opponent makes a move, one of these three options,
- 1:39:32based on what move I make.
- 1:39:33And as a result, we get some value.
- 1:39:36Let's look at the order in which I do these calculations
- 1:39:39and figure out if there are any optimizations I
- 1:39:41might be able to make to this calculation process.
- 1:39:44I'm going to have to look at these states one at a time.
- 1:39:47So let's say I start here on the left and say, all right,
- 1:39:49now I'm going to consider, what will the min player, my opponent,
- 1:39:52try to do here?
- 1:39:54Well, the min player is going to look at all three of their possible actions
- 1:39:57and look at their value, because these are terminal states.
- 1:40:00They're the end of the game.
- 1:40:01And so they'll see, all right, this node is a value of four, value of eight,
- 1:40:04value of five.
- 1:40:06And the min player is going to say, well, all right,
- 1:40:08between these three options, four, eight, and five, I'll take the smallest one.
- 1:40:13I'll take the four.
- 1:40:14So this state now has a value of four.
- 1:40:16Then I, as the max player, say, all right, if I take this action,
- 1:40:20it will have a value of four.
- 1:40:21That's the best that I can do, because min player
- 1:40:23is going to try and minimize my score.
- 1:40:25So now what if I take this option?
- 1:40:27We'll explore this next.
- 1:40:28And now explore what the min player would do if I choose this action.
- 1:40:32And the min player is going to say, all right, what are the three options?
- 1:40:35The min player has options between nine, three, and seven.
- 1:40:39And so three is the smallest among nine, three, and seven.
- 1:40:42So we'll go ahead and say this state has a value of three.
- 1:40:45So now I, as the max player, I have now explored two of my three options.
- 1:40:49I know that one of my options will guarantee me a score of four, at least.
- 1:40:53And one of my options will guarantee me a score of three.
- 1:40:57And now I consider my third option and say, all right, what happens here?
- 1:41:00Same exact logic.
- 1:41:01The min player is going to look at these three states, two, four, and six.
- 1:41:04I'll say the minimum possible option is two.
- 1:41:06So the min player wants the two.
- 1:41:08Now I, as the max player, have calculated all of the information
- 1:41:11by looking two layers deep, by looking at all of these nodes.
- 1:41:15And I can now say, between the four, the three, and the two, you know what?
- 1:41:18I'd rather take the four.
- 1:41:20Because if I choose this option, if my opponent plays optimally,
- 1:41:24they will try and get me to the four.
- 1:41:26But that's the best I can do.
- 1:41:27I can't guarantee a higher score.
- 1:41:29Because if I pick either of these two options, I might get a three
- 1:41:32or I might get a two.
- 1:41:33And it's true that down here is a nine.
- 1:41:36And that's the highest score out of any of the scores.
- 1:41:38So I might be tempted to say, you know what?
- 1:41:40Maybe I should take this option because I might get the nine.
- 1:41:43But if the min player is playing intelligently,
- 1:41:46if they're making the best moves at each possible option
- 1:41:48they have when they get to make a choice, I'll be left with a three.
- 1:41:52Whereas I could better, playing optimally,
- 1:41:54have guaranteed that I would get the four.
- 1:41:58So that is, in effect, the logic that I would use as a min and max player
- 1:42:01trying to maximize my score from that node there.
- 1:42:05But it turns out they took quite a bit of computation for me to figure that out.
- 1:42:08I had to reason through all of these nodes
- 1:42:10in order to draw this conclusion.
- 1:42:11And this is for a pretty simple game where I have three choices,
- 1:42:14my opponent has three choices, and then the game's over.
- 1:42:18So what I'd like to do is come up with some way to optimize this.
- 1:42:21Maybe I don't need to do all of this calculation
- 1:42:24to still reach the conclusion that, you know what, this action to the left,
- 1:42:28that's the best that I could do.
- 1:42:29Let's go ahead and try again and try to be a little more intelligent about how
- 1:42:33I go about doing this.
- 1:42:36So first, I start the exact same way.
- 1:42:38I don't know what to do initially, so I just
- 1:42:40have to consider one of the options and consider what the min player might do.
- 1:42:45Min has three options, four, eight, and five.
- 1:42:47And between those three options, min says four is the best they can do
- 1:42:51because they want to try to minimize the score.
- 1:42:54Now I, the max player, will consider my second option,
- 1:42:58making this move here, and considering what my opponent would do in response.
- 1:43:02What will the min player do?
- 1:43:04Well, the min player is going to, from that state, look at their options.
- 1:43:07And I would say, all right, nine is an option, three is an option.
- 1:43:12And if I am doing the math from this initial state,
- 1:43:14doing all this calculation, when I see a three,
- 1:43:17that should immediately be a red flag for me.
- 1:43:20Because when I see a three down here at this state,
- 1:43:23I know that the value of this state is going to be at most three.
- 1:43:28It's going to be three or something less than three,
- 1:43:30even though I haven't yet looked at this last action or even further actions
- 1:43:34if there were more actions that could be taken here.
- 1:43:37How do I know that?
- 1:43:37Well, I know that the min player is going to try to minimize my score.
- 1:43:42And if they see a three, the only way this
- 1:43:44could be something other than a three is if this remaining thing
- 1:43:47that I haven't yet looked at is less than three, which
- 1:43:50means there is no way for this value to be anything more than three
- 1:43:54because the min player can already guarantee a three
- 1:43:57and they are trying to minimize my score.
- 1:44:01So what does that tell me?
- 1:44:02Well, it tells me that if I choose this action,
- 1:44:04my score is going to be three or maybe even less than three if I'm unlucky.
- 1:44:09But I already know that this action will guarantee me a four.
- 1:44:13And so given that I know that this action guarantees me a score of four
- 1:44:17and this action means I can't do better than three,
- 1:44:20if I'm trying to maximize my options, there
- 1:44:22is no need for me to consider this triangle here.
- 1:44:25There is no value, no number that could go here
- 1:44:28that would change my mind between these two options.
- 1:44:30I'm always going to opt for this path that gets me a four as opposed
- 1:44:34to this path where the best I can do is a three if my opponent plays optimally.
- 1:44:39And this is going to be true for all the future states that I look at too.
- 1:44:43That if I look over here at what min player might do over here,
- 1:44:45if I see that this state is a two, I know that this state is at most a two
- 1:44:50because the only way this value could be something other than two
- 1:44:54is if one of these remaining states is less than a two
- 1:44:57and so the min player would opt for that instead.
- 1:45:00So even without looking at these remaining states,
- 1:45:03I as the maximizing player can know that choosing this path to the left
- 1:45:08is going to be better than choosing either of those two paths to the right
- 1:45:13because this one can't be better than three.
- 1:45:16This one can't be better than two.
- 1:45:17And so four in this case is the best that I can do.
- 1:45:21So in order to do this cut, and I can say now
- 1:45:23that this state has a value of four.
- 1:45:25So in order to do this type of calculation,
- 1:45:27I was doing a little bit more bookkeeping, keeping track of things,
- 1:45:31keeping track all the time of what is the best that I can do,
- 1:45:34what is the worst that I can do, and for each of these states
- 1:45:37saying, all right, well, if I already know that I can get a four,
- 1:45:41then if the best I can do at this state is a three,
- 1:45:44no reason for me to consider it, I can effectively prune this leaf
- 1:45:48and anything below it from the tree.
- 1:45:51And it's for that reason this approach, this optimization to minimax,
- 1:45:54is called alpha, beta pruning.
- 1:45:56Alpha and beta stand for these two values
- 1:45:58that you'll have to keep track of of the best you can do so far
- 1:46:01and the worst you can do so far.
- 1:46:02And pruning is the idea of if I have a big, long, deep search tree,
- 1:46:07I might be able to search it more efficiently
- 1:46:09if I don't need to search through everything,
- 1:46:11if I can remove some of the nodes to try and optimize the way that I
- 1:46:15look through this entire search space.
- 1:46:18So alpha, beta pruning can definitely save us a lot of time
- 1:46:21as we go about the search process by making our searches more efficient.
- 1:46:25But even then, it's still not great as games get more complex.
- 1:46:29Tic-tac-toe, fortunately, is a relatively simple game.
- 1:46:33And we might reasonably ask a question like,
- 1:46:35how many total possible tic-tac-toe games are there?
- 1:46:39You can think about it.
- 1:46:40You can try and estimate how many moves are there at any given point,
- 1:46:43how many moves long can the game last.
- 1:46:45It turns out there are about 255,000 possible tic-tac-toe games
- 1:46:52that can be played.
- 1:46:53But compare that to a more complex game, something
- 1:46:56like a game of chess, for example.
- 1:46:58Far more pieces, far more moves, games that last much longer.
- 1:47:01How many total possible chess games could there be?
- 1:47:05It turns out that after just four moves each, four moves by the white player,
- 1:47:08four moves by the black player, that there are
- 1:47:10288 billion possible chess games that can result from that situation,
- 1:47:15after just four moves each.
- 1:47:17And going even further, if you look at entire chess games
- 1:47:20and how many possible chess games there could be as a result there,
- 1:47:23there are more than 10 to the 29,000 possible chess games,
- 1:47:27far more chess games than could ever be considered.
- 1:47:30And this is a pretty big problem for the Minimax algorithm,
- 1:47:33because the Minimax algorithm starts with an initial state,
- 1:47:36considers all the possible actions, and all the possible actions
- 1:47:39after that, all the way until we get to the end of the game.
- 1:47:44And that's going to be a problem if the computer is going
- 1:47:46to need to look through this many states, which is far more than any computer
- 1:47:51could ever do in any reasonable amount of time.
- 1:47:54So what do we do in order to solve this problem?
- 1:47:57Instead of looking through all these states which
- 1:47:59is totally intractable for a computer, we need some better approach.
- 1:48:02And it turns out that better approach generally takes the form of something
- 1:48:05called depth-limited Minimax, where normally Minimax
- 1:48:09is depth-unlimited.
- 1:48:10We just keep going layer after layer, move after move,
- 1:48:13until we get to the end of the game.
- 1:48:15Depth-limited Minimax is instead going to say,
- 1:48:17you know what, after a certain number of moves, maybe I'll look 10 moves ahead,
- 1:48:21maybe I'll look 12 moves ahead, but after that point,
- 1:48:23I'm going to stop and not consider additional moves that
- 1:48:26might come after that, just because it would be computationally intractable
- 1:48:30to consider all of those possible options.
- 1:48:34But what do we do after we get 10 or 12 moves deep
- 1:48:36when we arrive at a situation where the game's not over?
- 1:48:40Minimax still needs a way to assign a score to that game board or game
- 1:48:43state to figure out what its current value is, which is easy to do
- 1:48:47if the game is over, but not so easy to do if the game is not yet over.
- 1:48:51So in order to do that, we need to add one additional feature
- 1:48:54to depth-limited Minimax called an evaluation function, which
- 1:48:57is just some function that is going to estimate the expected utility
- 1:49:01of a game from a given state.
- 1:49:04So in a game like chess, if you imagine that a game value of 1
- 1:49:07means white wins, negative 1 means black wins, 0 means it's a draw,
- 1:49:12then you might imagine that a score of 0.8
- 1:49:15means white is very likely to win, though certainly not guaranteed.
- 1:49:19And you would have an evaluation function
- 1:49:21that estimates how good the game state happens to be.
- 1:49:25And depending on how good that evaluation function is,
- 1:49:28that is ultimately what's going to constrain how good the AI is.
- 1:49:32The better the AI is at estimating how good or how bad
- 1:49:36any particular game state is, the better the AI
- 1:49:38is going to be able to play that game.
- 1:49:40If the evaluation function is worse and not as good as it estimating
- 1:49:44what the expected utility is, then it's going to be a whole lot harder.
- 1:49:47And you can imagine trying to come up with these evaluation functions.
- 1:49:51In chess, for example, you might write an evaluation function
- 1:49:54based on how many pieces you have as compared
- 1:49:56to how many pieces your opponent has, because each one has a value.
- 1:49:59And your evaluation function probably needs
- 1:50:02to be a little bit more complicated than that
- 1:50:04to consider other possible situations that might arise as well.
- 1:50:08And there are many other variants on Minimax that add additional features
- 1:50:11in order to help it perform better under these larger, more computationally
- 1:50:15untractable situations where we couldn't possibly
- 1:50:18explore all of the possible moves.
- 1:50:20So we need to figure out how to use evaluation functions and other techniques
- 1:50:25to be able to play these games ultimately better.
- 1:50:28But this now was a look at this kind of adversarial search, these search
- 1:50:31problems where we have situations where I am trying
- 1:50:35to play against some sort of opponent.
- 1:50:37And these search problems show up all over the place
- 1:50:40throughout artificial intelligence.
- 1:50:41We've been talking a lot today about more classical search problems,
- 1:50:44like trying to find directions from one location to another.
- 1:50:48But any time an AI is faced with trying to make a decision,
- 1:50:51like what do I do now in order to do something that is rational,
- 1:50:54or do something that is intelligent, or trying to play a game,
- 1:50:57like figuring out what move to make, these sort of algorithms
- 1:51:00can really come in handy.
- 1:51:01It turns out that for tic-tac-toe, the solution is pretty simple
- 1:51:04because it's a small game.
- 1:51:05XKCD has famously put together a web comic
- 1:51:08where he will tell you exactly what move to make as the optimal move to make
- 1:51:12no matter what your opponent happens to do.
- 1:51:14This type of thing is not quite as possible for a much larger game
- 1:51:17like Checkers or Chess, for example, where chess is totally computationally
- 1:51:21untractable for most computers to be able to explore all the possible states.
- 1:51:25So we really need our AI to be far more intelligent about how
- 1:51:29they go about trying to deal with these problems
- 1:51:31and how they go about taking this environment that they find themselves in
- 1:51:35and ultimately searching for one of these solutions.
- 1:51:38So this, then, was a look at search in artificial intelligence.
- 1:51:41Next time, we'll take a look at knowledge,
- 1:51:43thinking about how it is that our AIs are able to know information, reason
- 1:51:47about that information, and draw conclusions, all in our look at AI
- 1:51:51and the principles behind it.
- 1:51:52We'll see you next time.
- 1:51:55["AIMS INTRO MUSIC"]
- 1:52:13All right, welcome back, everyone, to an introduction
- 1:52:16to artificial intelligence with Python.
- 1:52:18Last time, we took a look at search problems, in particular,
- 1:52:20where we have AI agents that are trying to solve some sort of problem
- 1:52:24by taking actions in some sort of environment,
- 1:52:26whether that environment is trying to take actions by playing moves in a game
- 1:52:30or whether those actions are something like trying
- 1:52:32to figure out where to make turns in order to get driving directions
- 1:52:35from point A to point B. This time, we're
- 1:52:38going to turn our attention more generally to just this idea of knowledge,
- 1:52:42the idea that a lot of intelligence is based on knowledge,
- 1:52:44especially if we think about human intelligence.
- 1:52:47People know information.
- 1:52:48We know facts about the world.
- 1:52:50And using that information that we know, we're
- 1:52:52able to draw conclusions, reason about the information
- 1:52:55that we know in order to figure out how to do something
- 1:52:58or figure out some other piece of information
- 1:53:00that we conclude based on the information we already have available to us.
- 1:53:05What we'd like to focus on now is the ability
- 1:53:07to take this idea of knowledge and being able to reason based on knowledge
- 1:53:11and apply those ideas to artificial intelligence.
- 1:53:14In particular, we're going to be building what
- 1:53:16are known as knowledge-based agents, agents that
- 1:53:19are able to reason and act by representing knowledge internally.
- 1:53:23Somehow inside of our AI, they have some understanding
- 1:53:25of what it means to know something.
- 1:53:27And ideally, they have some algorithms or some techniques
- 1:53:30they can use based on that knowledge that they know in order to figure out
- 1:53:34the solution to a problem or figure out some additional piece of information
- 1:53:38that can be helpful in some sense.
- 1:53:40So what do we mean by reasoning based on knowledge
- 1:53:43to be able to draw conclusions?
- 1:53:44Well, let's look at a simple example drawn from the world of Harry Potter.
- 1:53:47We take one sentence that we know to be true.
- 1:53:50Imagine if it didn't rain, then Harry visited Hagrid today.
- 1:53:55So one fact that we might know about the world.
- 1:53:57And then we take another fact.
- 1:53:59Harry visited Hagrid or Dumbledore today, but not both.
- 1:54:02So it tells us something about the world, that Harry either visited
- 1:54:05Hagrid but not Dumbledore, or Harry visited Dumbledore but not Hagrid.
- 1:54:09And now we have a third piece of information about the world
- 1:54:12that Harry visited Dumbledore today.
- 1:54:14So we now have three pieces of information now, three facts.
- 1:54:17Inside of a knowledge base, so to speak, information that we know.
- 1:54:21And now we, as humans, can try and reason about this
- 1:54:23and figure out, based on this information, what additional information
- 1:54:27can we begin to conclude?
- 1:54:29And well, looking at these last two statements,
- 1:54:31Harry either visited Hagrid or Dumbledore but not both,
- 1:54:35and we know that Harry visited Dumbledore today, well,
- 1:54:38then it's pretty reasonable that we could draw the conclusion that,
- 1:54:40you know what, Harry must not have visited Hagrid today.
- 1:54:43Because based on a combination of these two statements,
- 1:54:46we can draw this inference, so to speak, a conclusion that Harry did not
- 1:54:50visit Hagrid today.
- 1:54:52But it turns out we can even do a little bit better than that,
- 1:54:54get some more information by taking a look at this first statement
- 1:54:57and reasoning about that.
- 1:54:59This first statement says, if it didn't rain,
- 1:55:01then Harry visited Hagrid today.
- 1:55:04So what does that mean?
- 1:55:05In all cases where it didn't rain, then we know that Harry visited Hagrid.
- 1:55:09But if we also know now that Harry did not visit Hagrid,
- 1:55:12then that tells us something about our initial premise
- 1:55:15that we were thinking about.
- 1:55:16In particular, it tells us that it did rain today, because we can reason,
- 1:55:21if it didn't rain, that Harry would have visited Hagrid.
- 1:55:24But we know for a fact that Harry did not visit Hagrid today.
- 1:55:28So it's this kind of reason, this sort of logical reasoning,
- 1:55:31where we use logic based on the information
- 1:55:33that we know in order to take information and reach conclusions that
- 1:55:38is going to be the focus of what we're going to be talking about today.
- 1:55:40How can we make our artificial intelligence
- 1:55:43logical so that they can perform the same kinds of deduction,
- 1:55:47the same kinds of reasoning that we've been doing so far?
- 1:55:50Of course, humans reason about logic generally
- 1:55:53in terms of human language.
- 1:55:54That I just now was speaking in English, talking in English about these
- 1:55:58sentences and trying to reason through how it
- 1:56:01is that they relate to one another.
- 1:56:02We're going to need to be a little bit more formal when
- 1:56:05we turn our attention to computers and being
- 1:56:07able to encode this notion of logic and truthhood and falsehood
- 1:56:11inside of a machine.
- 1:56:12So we're going to need to introduce a few more terms and a few symbols that
- 1:56:16will help us reason through this idea of logic
- 1:56:18inside of an artificial intelligence.
- 1:56:20And we'll begin with the idea of a sentence.
- 1:56:22Now, a sentence in a natural language like English
- 1:56:24is just something that I'm saying, like what I'm saying right now.
- 1:56:28In the context of AI, though, a sentence is just an assertion about the world
- 1:56:32in what we're going to call a knowledge representation language,
- 1:56:36some way of representing knowledge inside of our computers.
- 1:56:40And the way that we're going to spend most of today reasoning about knowledge
- 1:56:44is through a type of logic known as propositional logic.
- 1:56:47There are a number of different types of logic, some of which we'll touch on.
- 1:56:50But propositional logic is based on a logic of propositions,
- 1:56:54or just statements about the world.
- 1:56:56And so we begin in propositional logic with a notion of propositional symbols.
- 1:57:01We will have certain symbols that are oftentimes just letters,
- 1:57:04something like P or Q or R, where each of those symbols
- 1:57:07is going to represent some fact or sentence about the world.
- 1:57:11So P, for example, might represent the fact that it is raining.
- 1:57:15And so P is going to be a symbol that represents that idea.
- 1:57:19And Q, for example, might represent Harry visited Hagrid today.
- 1:57:22Each of these propositional symbols represents some sentence
- 1:57:26or some fact about the world.
- 1:57:29But in addition to just having individual facts about the world,
- 1:57:32we want some way to connect these propositional symbols together
- 1:57:36in order to reason more complexly about other facts that
- 1:57:39might exist inside of the world in which we're reasoning.
- 1:57:42So in order to do that, we'll need to introduce some additional symbols
- 1:57:45that are known as logical connectives.
- 1:57:47Now, there are a number of these logical connectives.
- 1:57:49But five of the most important, and the ones we're going to focus on today,
- 1:57:52are these five up here, each represented by a logical symbol.
- 1:57:56Not is represented by this symbol here, and is represented
- 1:58:00as sort of an upside down V, or is represented by a V shape.
- 1:58:04Implication, and we'll talk about what that means in just a moment,
- 1:58:07is represented by an arrow.
- 1:58:09And biconditional, again, we'll talk about what that means in a moment,
- 1:58:12is represented by these double arrows.
- 1:58:14But these five logical connectives are the main ones
- 1:58:17we're going to be focusing on in terms of thinking about how
- 1:58:20it is that a computer can reason about facts
- 1:58:22and draw conclusions based on the facts that it knows.
- 1:58:26But in order to get there, we need to take
- 1:58:28a look at each of these logical connectives
- 1:58:30and build up an understanding for what it is that they actually mean.
- 1:58:34So let's go ahead and begin with the not symbol, so this not symbol here.
- 1:58:38And what we're going to show for each of these logical connectives
- 1:58:41is what we're going to call a truth table, a table that
- 1:58:43demonstrates what this word not means when we attach it
- 1:58:47to a propositional symbol or any sentence inside of our logical language.
- 1:58:52And so the truth table for not is shown right here.
- 1:58:56If P, some propositional symbol, or some other sentence even, is false,
- 1:59:01then not P is true.
- 1:59:04And if P is true, then not P is false.
- 1:59:08So you can imagine that placing this not symbol
- 1:59:11in front of some sentence of propositional logic
- 1:59:14just says the opposite of that.
- 1:59:16So if, for example, P represented it is raining,
- 1:59:19then not P would represent the idea that it is not raining.
- 1:59:23And as you might expect, if P is false, meaning if the sentence,
- 1:59:27it is raining, is false, well then the sentence not P must be true.
- 1:59:32The sentence that it is not raining is therefore true.
- 1:59:36So not, you can imagine, just takes whatever is in P and it inverts it.
- 1:59:40It turns false into true and true into false,
- 1:59:43much analogously to what the English word not means,
- 1:59:46just taking whatever comes after it and inverting it to mean the opposite.
- 1:59:51Next up, and also very English-like, is this idea
- 1:59:53of and represented by this upside-down V shape or this point shape.
- 1:59:58And as opposed to just taking a single argument the way not does,
- 2:00:01we have P and we have not P. And is going to combine two different sentences
- 2:00:07in propositional logic together.
- 2:00:09So I might have one sentence P and another sentence Q,
- 2:00:12and I want to combine them together to say P and Q.
- 2:00:16And the general logic for what P and Q means
- 2:00:19is it means that both of its operands are true.
- 2:00:22P is true and also Q is true.
- 2:00:26And so here's what that truth table looks like.
- 2:00:29This time we have two variables, P and Q. And when we have two variables, each
- 2:00:33of which can be in two possible states, true or false,
- 2:00:36that leads to two squared or four possible combinations
- 2:00:41of truth and falsehood.
- 2:00:42So we have P is false and Q is false.
- 2:00:45We have P is false and Q is true.
- 2:00:47P is true and Q is false.
- 2:00:48And then P and Q both are true.
- 2:00:51And those are the only four possibilities for what P and Q could mean.
- 2:00:55And in each of those situations, this third column here, P and Q,
- 2:00:59is telling us a little bit about what it actually means for P and Q to be true.
- 2:01:03And we see that the only case where P and Q is true is in this fourth row
- 2:01:08here, where P happens to be true, Q also happens to be true.
- 2:01:12And in all other situations, P and Q is going to evaluate to false.
- 2:01:18So this, again, is much in line with what our intuition of and might mean.
- 2:01:21If I say P and Q, I probably mean that I expect both P and Q to be true.
- 2:01:29Next up, also potentially consistent with what we mean,
- 2:01:32is this word or, represented by this V shape, sort of an upside down and symbol.
- 2:01:37And or, as the name might suggest, is true if either of its arguments
- 2:01:41are true, as long as P is true or Q is true, then P or Q is going to be true.
- 2:01:47Which means the only time that P or Q is false
- 2:01:50is if both of its operands are false.
- 2:01:53If P is false and Q is false, then P or Q is going to be false.
- 2:01:58But in all other cases, at least one of the operands is true.
- 2:02:03Maybe they're both true, in which case P or Q is going to evaluate to true.
- 2:02:08Now, this is mostly consistent with the way
- 2:02:10that most people might use the word or, in the sense of speaking the word
- 2:02:14or in normal English, though there is sometimes when we might say
- 2:02:17or, where we mean P or Q, but not both, where we mean, sort of,
- 2:02:21it can only be one or the other.
- 2:02:23It's important to note that this symbol here, this or,
- 2:02:26means P or Q or both, that those are totally OK.
- 2:02:30As long as either or both of them are true,
- 2:02:33then the or is going to evaluate to be true, as well.
- 2:02:36It's only in the case where all of the operands
- 2:02:38are false that P or Q ultimately evaluates to false, as well.
- 2:02:43In logic, there's another symbol known as the exclusive or,
- 2:02:46which encodes this idea of exclusivity of one or the other, but not both.
- 2:02:51But we're not going to be focusing on that today.
- 2:02:53Whenever we talk about or, we're always talking about either or both,
- 2:02:56in this case, as represented by this truth table here.
- 2:03:01So that now is not an and an or.
- 2:03:04And next up is what we might call implication,
- 2:03:07as denoted by this arrow symbol.
- 2:03:09So we have P and Q. And this sentence here will generally
- 2:03:13read as P implies Q.
- 2:03:16And what P implies Q means is that if P is true, then Q is also true.
- 2:03:23So I might say something like, if it is raining, then I will be indoors.
- 2:03:27Meaning, it is raining implies I will be indoors,
- 2:03:31as the logical sentence that I'm saying there.
- 2:03:34And the truth table for this can sometimes be a little bit tricky.
- 2:03:37So obviously, if P is true and Q is true, then P implies Q. That's true.
- 2:03:44That definitely makes sense.
- 2:03:46And it should also stand to reason that when P is true and Q is false,
- 2:03:50then P implies Q is false.
- 2:03:52Because if I said to you, if it is raining, then I will be out indoors.
- 2:03:57And it is raining, but I'm not indoors?
- 2:04:01Well, then it would seem to be that my original statement was not true.
- 2:04:04P implies Q means that if P is true, then Q also needs to be true.
- 2:04:09And if it's not, well, then the statement is false.
- 2:04:13What's also worth noting, though, is what happens when P is false.
- 2:04:17When P is false, the implication makes no claim at all.
- 2:04:22If I say something like, if it is raining, then I will be indoors.
- 2:04:26And it turns out it's not raining.
- 2:04:28Then in that case, I am not making any statement
- 2:04:31as to whether or not I will be indoors or not.
- 2:04:33P implies Q just means that if P is true, Q must be true.
- 2:04:37But if P is not true, then we make no claim about whether or not Q
- 2:04:42is true at all.
- 2:04:43So in either case, if P is false, it doesn't matter what Q is.
- 2:04:46Whether it's false or true, we're not making any claim about Q whatsoever.
- 2:04:50We can still evaluate the implication to true.
- 2:04:53The only way that the implication is ever false
- 2:04:56is if our premise, P, is true, but the conclusion that we're drawing Q
- 2:05:01happens to be false.
- 2:05:03So in that case, we would say P does not imply Q in that case.
- 2:05:09Finally, the last connective that we'll discuss is this bi-conditional.
- 2:05:13You can think of a bi-conditional as a condition
- 2:05:15that goes in both directions.
- 2:05:17So originally, when I said something like, if it is raining,
- 2:05:20then I will be indoors.
- 2:05:22I didn't say what would happen if it wasn't raining.
- 2:05:24Maybe I'll be indoors, maybe I'll be outdoors.
- 2:05:27This bi-conditional, you can read as an if and only if.
- 2:05:31So I can say, I will be indoors if and only if it is raining,
- 2:05:36meaning if it is raining, then I will be indoors.
- 2:05:39And if I am indoors, it's reasonable to conclude that it is also raining.
- 2:05:43So this bi-conditional is only true when P and Q are the same.
- 2:05:48So if P is true and Q is true, then this bi-conditional is also true.
- 2:05:53P implies Q, but also the reverse is true.
- 2:05:56Q also implies P. So if P and Q both happen to be false,
- 2:06:01we would still say it's true.
- 2:06:02But in any of these other two situations,
- 2:06:04this P if and only if Q is going to ultimately evaluate to false.
- 2:06:08So a lot of trues and falses going on there,
- 2:06:11but these five basic logical connectives
- 2:06:13are going to form the core of the language of propositional logic,
- 2:06:16the language that we're going to use in order to describe ideas,
- 2:06:20and the language that we're going to use in order
- 2:06:21to reason about those ideas in order to draw conclusions.
- 2:06:26So let's now take a look at some of the additional terms
- 2:06:29that we'll need to know about in order to go about trying
- 2:06:31to form this language of propositional logic
- 2:06:33and writing AI that's actually able to understand this sort of logic.
- 2:06:37The next thing we're going to need is the notion of what
- 2:06:40is actually true about the world.
- 2:06:42We have a whole bunch of propositional symbols, P and Q and R and maybe others,
- 2:06:46but we need some way of knowing what actually is true in the world.
- 2:06:50Is P true or false?
- 2:06:51Is Q true or false?
- 2:06:52So on and so forth.
- 2:06:54And to do that, we'll introduce the notion of a model.
- 2:06:57A model just assigns a truth value, where a truth value is either true
- 2:07:02or false, to every propositional symbol.
- 2:07:05In other words, it's creating what we might call a possible world.
- 2:07:09So let me give an example.
- 2:07:10If, for example, I have two propositional symbols, P is it is raining
- 2:07:15and Q is it is a Tuesday, a model just takes each of these two symbols
- 2:07:21and assigns a truth value to them, either true or false.
- 2:07:24So here's a sample model.
- 2:07:26In this model, in other words, in this possible world,
- 2:07:29it is possible that P is true, meaning it is raining, and Q is false,
- 2:07:33meaning it is not a Tuesday.
- 2:07:36But there are other possible worlds or other models as well.
- 2:07:39There is some model where both of these variables are true,
- 2:07:41some model where both of these variables are false.
- 2:07:44In fact, if there are n variables that are propositional symbols like this
- 2:07:48that are either true or false, then the number of possible models
- 2:07:51is 2 to the n, because each of these possible models,
- 2:07:55possible variables within my model, could be set to either true or false
- 2:08:00if I don't know any information about it.
- 2:08:03So now that I have the symbols and the connectives
- 2:08:07that I'm going to need in order to construct these parts of knowledge,
- 2:08:11we need some way to represent that knowledge.
- 2:08:13And to do so, we're going to allow our AI access
- 2:08:15to what we'll call a knowledge base.
- 2:08:18And a knowledge base is really just a set of sentences
- 2:08:21that our AI knows to be true.
- 2:08:24Some set of sentences in propositional logic
- 2:08:27that are things that our AI knows about the world.
- 2:08:30And so we might tell our AI some information, information about a situation
- 2:08:35that it finds itself in, or a situation about a problem
- 2:08:38that it happens to be trying to solve.
- 2:08:39And we would give that information to the AI
- 2:08:41that the AI would store inside of its knowledge base.
- 2:08:44And what happens next is the AI would like
- 2:08:47to use that information in the knowledge base
- 2:08:49to be able to draw conclusions about the rest of the world.
- 2:08:53And what do those conclusions look like?
- 2:08:55Well, to understand those conclusions, we'll
- 2:08:56need to introduce one more idea, one more symbol.
- 2:08:59And that is the notion of entailment.
- 2:09:02So this sentence here, with this double turnstile in these Greek letters,
- 2:09:06this is the Greek letter alpha and the Greek letter beta.
- 2:09:08And we read this as alpha entails beta.
- 2:09:12And alpha and beta here are just sentences in propositional logic.
- 2:09:17And what this means is that alpha entails beta
- 2:09:20means that in every model, in other words,
- 2:09:23in every possible world in which sentence alpha is true,
- 2:09:28then sentence beta is also true.
- 2:09:31So if something entails something else, if alpha entails beta,
- 2:09:35it means that if I know alpha to be true, then beta must therefore also
- 2:09:40be true.
- 2:09:41So if my alpha is something like I know that it is a Tuesday in January,
- 2:09:47then a reasonable beta might be something like I know that it is January.
- 2:09:52Because in all worlds where it is a Tuesday in January,
- 2:09:55I know for sure that it must be January, just by definition.
- 2:09:59This first statement or sentence about the world
- 2:10:01entails the second statement.
- 2:10:03And we can reasonably use deduction based on that first sentence
- 2:10:07to figure out that the second sentence is, in fact, true as well.
- 2:10:12And ultimately, it's this idea of entailment
- 2:10:14that we're going to try and encode into our computer.
- 2:10:17We want our AI agent to be able to figure out
- 2:10:20what the possible entailments are.
- 2:10:22We want our AI to be able to take these three sentences, sentences like,
- 2:10:26if it didn't rain, Harry visited Hagrid.
- 2:10:28That Harry visited Hagrid or Dumbledore, but not both.
- 2:10:31And that Harry visited Dumbledore.
- 2:10:33And just using that information, we'd like our AI
- 2:10:36to be able to infer or figure out that using these three sentences inside
- 2:10:41of a knowledge base, we can draw some conclusions.
- 2:10:44In particular, we can draw the conclusions here that, one,
- 2:10:47Harry did not visit Hagrid today.
- 2:10:49And we can draw the entailment, too, that it did, in fact, rain today.
- 2:10:53And this process is known as inference.
- 2:10:56And that's what we're going to be focusing on today,
- 2:10:58this process of deriving new sentences from old ones,
- 2:11:01that I give you these three sentences, you put them
- 2:11:04in the knowledge base in, say, the AI.
- 2:11:06And the AI is able to use some sort of inference algorithm
- 2:11:09to figure out that these two sentences must also be true.
- 2:11:14And that is how we define inference.
- 2:11:16So let's take a look at an inference example
- 2:11:18to see how we might actually go about inferring things in a human sense
- 2:11:22before we take a more algorithmic approach
- 2:11:24to see how we could encode this idea of inference in AI.
- 2:11:27And we'll see there are a number of ways that we can actually achieve this.
- 2:11:30So again, we'll deal with a couple of propositional symbols.
- 2:11:33We'll deal with P, Q, and R. P is it is a Tuesday.
- 2:11:37Q is it is raining.
- 2:11:39And R is Harry will go for a run, three propositional symbols
- 2:11:42that we are just defining to mean this.
- 2:11:44We're not saying anything yet about whether they're true or false.
- 2:11:47We're just defining what they are.
- 2:11:50Now, we'll give ourselves or an AI access to a knowledge base,
- 2:11:53abbreviated to KB, the knowledge that we know about the world.
- 2:11:57We know this statement.
- 2:11:59All right.
- 2:11:59So let's try to parse it.
- 2:12:00The parentheses here are just used for precedent,
- 2:12:02so we can see what associates with what.
- 2:12:05But you would read this as P and not Q implies R.
- 2:12:11All right.
- 2:12:12So what does that mean?
- 2:12:13Let's put it piece by piece.
- 2:12:14P is it is a Tuesday.
- 2:12:16Q is it is raining, so not Q is it is not raining,
- 2:12:21and implies R is Harry will go for a run.
- 2:12:25So the way to read this entire sentence in human natural language
- 2:12:28at least is if it is a Tuesday and it is not raining,
- 2:12:33then Harry will go for a run.
- 2:12:35So if it is a Tuesday and it is not raining,
- 2:12:37then Harry will go for a run.
- 2:12:39And that is now inside of our knowledge base.
- 2:12:41And let's now imagine that our knowledge base has
- 2:12:43two other pieces of information as well.
- 2:12:45It has information that P is true, that it is a Tuesday.
- 2:12:49And we also have the information not Q, that it is not raining,
- 2:12:53that this sentence Q, it is raining, happens to be false.
- 2:12:57And those are the three sentences that we have access to.
- 2:12:59P and not Q implies R, P and not Q. Using that information,
- 2:13:05we should be able to draw some inferences.
- 2:13:08P and not Q is only true if both P and not Q are true.
- 2:13:14All right, we know that P is true and we know that not Q is true.
- 2:13:18So we know that this whole expression is true.
- 2:13:20And the definition of implication is if this whole thing on the left
- 2:13:24is true, then this thing on the right must also be true.
- 2:13:27So if we know that P and not Q is true, then R must be true as well.
- 2:13:31So the inference we should be able to draw from all of this
- 2:13:34is that R is true and we know that Harry will go for a run
- 2:13:38by taking this knowledge inside of our knowledge base
- 2:13:40and being able to reason based on that idea.
- 2:13:43And so this ultimately is the beginning of what
- 2:13:46we might consider to be some sort of inference algorithm,
- 2:13:49some process that we can use to try and figure out
- 2:13:52whether or not we can draw some conclusion.
- 2:13:55And ultimately, what these inference algorithms are going to answer
- 2:13:58is the central question about entailment.
- 2:14:00Given some query about the world, something
- 2:14:02we're wondering about the world, and we'll call that query alpha,
- 2:14:06the question we want to ask using these inference algorithms
- 2:14:09is does KB, our knowledge base, entail alpha?
- 2:14:14In other words, using only the information
- 2:14:16we know inside of our knowledge base, the knowledge that we have access to,
- 2:14:20can we conclude that this sentence alpha is true?
- 2:14:24And that's ultimately what we would like to do.
- 2:14:26So how can we do that?
- 2:14:28How can we go about writing an algorithm that
- 2:14:30can look at this knowledge base and figure out whether or not this query
- 2:14:33alpha is actually true?
- 2:14:35Well, it turns out there are a couple of different algorithms for doing so.
- 2:14:39And one of the simplest, perhaps, is known as model checking.
- 2:14:43Now, remember that a model is just some assignment
- 2:14:45of all of the propositional symbols inside of our language to a truth
- 2:14:49value, true or false.
- 2:14:51And you can think of a model as a possible world,
- 2:14:53that there are many possible worlds where different things might
- 2:14:55be true or false, and we can enumerate all of them.
- 2:14:59And the model checking algorithm does exactly that.
- 2:15:02So what does our model checking algorithm do?
- 2:15:04Well, if we wanted to determine if our knowledge base entails
- 2:15:08some query alpha, then we are going to enumerate all possible models.
- 2:15:13In other words, consider all possible values of true and false
- 2:15:16for our variables, all possible states in which our world can be in.
- 2:15:21And if in every model where our knowledge base is true,
- 2:15:25alpha is also true, then we know that the knowledge base entails alpha.
- 2:15:30So let's take a closer look at that sentence
- 2:15:32and try and figure out what it actually means.
- 2:15:34If we know that in every model, in other words, in every possible world,
- 2:15:38no matter what assignment of true and false to variables you give,
- 2:15:41if we know that whenever our knowledge is true, what
- 2:15:44we know to be true is true, that this query alpha is also true,
- 2:15:49well, then it stands to reason that as long as our knowledge base is true,
- 2:15:52then alpha must also be true.
- 2:15:56And so this is going to form the foundation of our model checking
- 2:15:58algorithm.
- 2:15:59We're going to enumerate all of the possible worlds
- 2:16:01and ask ourselves whenever the knowledge base is true, is alpha true?
- 2:16:05And if that's the case, then we know alpha to be true.
- 2:16:09And otherwise, there is no entailment.
- 2:16:11Our knowledge base does not entail alpha.
- 2:16:14All right.
- 2:16:15So this is a little bit abstract, but let's
- 2:16:17take a look at an example to try and put real propositional symbols
- 2:16:20to this idea.
- 2:16:22So again, we'll work with the same example.
- 2:16:24P is it is a Tuesday, Q is it is raining, R as Harry will go for a run.
- 2:16:29Our knowledge base contains these pieces of information.
- 2:16:32P and not Q implies R. We also know P.
- 2:16:35It is a Tuesday and not Q. It is not raining.
- 2:16:38And our query, our alpha in this case, the thing we want to ask is R.
- 2:16:43We want to know, is it guaranteed?
- 2:16:45Is it entailed that Harry will go for a run?
- 2:16:49So the first step is to enumerate all of the possible models.
- 2:16:52We have three propositional symbols here, P, Q, and R,
- 2:16:55which means we have 2 to the third power, or eight possible models.
- 2:16:59All false, false, false true, false true, false, false true, true, et cetera.
- 2:17:04Eight possible ways you could assign true and false to all of these models.
- 2:17:09And we might ask in each one of them, is the knowledge base true?
- 2:17:13Here are the set of things that we know.
- 2:17:15In which of these worlds could this knowledge base possibly apply to?
- 2:17:20In which world is this knowledge base true?
- 2:17:22Well, in the knowledge base, for example, we know P.
- 2:17:26We know it is a Tuesday, which means we know that these four first four rows
- 2:17:31where P is false, none of those are going to be true
- 2:17:35or are going to work for this particular knowledge base.
- 2:17:37Our knowledge base is not true in those worlds.
- 2:17:40Likewise, we also know not Q. We know that it is not raining.
- 2:17:46So any of these models where Q is true, like these two and these two here,
- 2:17:51those aren't going to work either because we know that Q is not true.
- 2:17:55And finally, we also know that P and not Q implies R,
- 2:18:00which means that when P is true or P is true here and Q is false,
- 2:18:04Q is false in these two, then R must be true.
- 2:18:08And if ever P is true, Q is false, but R is also false,
- 2:18:14well, that doesn't satisfy this implication here.
- 2:18:17That implication does not hold true under those situations.
- 2:18:21So we could say that for our knowledge base,
- 2:18:24we can conclude under which of these possible worlds
- 2:18:27is our knowledge base true and under which of the possible worlds
- 2:18:30is our knowledge base false.
- 2:18:31And it turns out there is only one possible world
- 2:18:35where our knowledge base is actually true.
- 2:18:37In some cases, there might be multiple possible worlds
- 2:18:39where the knowledge base is true.
- 2:18:40But in this case, it just so happens that there's only one, one possible world
- 2:18:44where we can definitively say something about our knowledge base.
- 2:18:48And in this case, we would look at the query.
- 2:18:50The query of R is R true, R is true, and so as a result,
- 2:18:56we can draw that conclusion.
- 2:18:58And so this is this idea of model check-in.
- 2:19:01Enumerate all the possible models and look in those possible models
- 2:19:04to see whether or not, if our knowledge base is true,
- 2:19:08is the query in question true as well.
- 2:19:11So let's now take a look at how we might actually go about writing this
- 2:19:14in a programming language like Python.
- 2:19:16Take a look at some actual code that would
- 2:19:18encode this notion of propositional symbols and logic
- 2:19:21and these connectives like and and or and not and implication and so forth
- 2:19:25and see what that code might actually look like.
- 2:19:28So I've written in advance a logic library that's
- 2:19:30more detailed than we need to worry about entirely today.
- 2:19:33But the important thing is that we have one class for every type
- 2:19:37of logical symbol or connective that we might have.
- 2:19:40So we just have one class for logical symbols, for example,
- 2:19:44where every symbol is going to represent and store
- 2:19:46some name for that particular symbol.
- 2:19:49And we also have a class for not that takes an operand.
- 2:19:52So we might say not one symbol to say something is not true
- 2:19:56or some other sentence is not true.
- 2:19:58We have one for and, one for or, so on and so forth.
- 2:20:02And I'll just demonstrate how this works.
- 2:20:03And you can take a look at the actual logic.py later on.
- 2:20:07But I'll go ahead and call this file harry.py.
- 2:20:11We're going to store information about this world of Harry Potter,
- 2:20:15for example.
- 2:20:16So I'll go ahead and import from my logic module.
- 2:20:19I'll import everything.
- 2:20:20And in this library, in order to create a symbol, you use capital S symbol.
- 2:20:25And I'll create a symbol for rain, to mean it is raining, for example.
- 2:20:30And I'll create a symbol for Hagrid, to mean Harry visited Hagrid,
- 2:20:35is what this symbol is going to mean.
- 2:20:36So this symbol means it is raining.
- 2:20:38This symbol means Harry visited Hagrid.
- 2:20:41And I'll add another symbol called Dumbledore for Harry visited Dumbledore.
- 2:20:49Now, I'd like to save these symbols so that I can use them later
- 2:20:52as I do some logical analysis.
- 2:20:54So I'll go ahead and save each one of them inside of a variable.
- 2:20:56So like rain, Hagrid, and Dumbledore, so you could call the variables anything.
- 2:21:02And now that I have these logical symbols,
- 2:21:04I can use logical connectives to combine them together.
- 2:21:07So for example, if I have a sentence like and rain and Hagrid,
- 2:21:14for example, which is not necessarily true, but just for demonstration,
- 2:21:18I can now try and print out sentence.formula, which
- 2:21:22is a function I wrote that takes a sentence in propositional logic
- 2:21:25and just prints it out so that we, the programmers,
- 2:21:27can now see this in order to get an understanding for how it actually works.
- 2:21:32So if I run python harry.py, what we'll see
- 2:21:36is this sentence in propositional logic, rain and Hagrid.
- 2:21:40This is the logical representation of what we have here in our Python program
- 2:21:44of saying and whose arguments are rain and Hagrid.
- 2:21:48So we're saying rain and Hagrid by encoding that idea.
- 2:21:51And this is quite common in Python object-oriented programming,
- 2:21:54where you have a number of different classes,
- 2:21:56and you pass arguments into them in order to create a new and object,
- 2:22:01for example, in order to represent this idea.
- 2:22:03But now what I'd like to do is somehow encode the knowledge
- 2:22:07that I have about the world in order to solve
- 2:22:09that problem from the beginning of class, where
- 2:22:11we talked about trying to figure out who Harry visited
- 2:22:14and trying to figure out if it's raining or if it's not raining.
- 2:22:17And so what knowledge do I have?
- 2:22:19I'll go ahead and create a new variable called knowledge.
- 2:22:22And what do I know?
- 2:22:23Well, I know the very first sentence that we talked about
- 2:22:25was the idea that if it is not raining, then Harry will visit Hagrid.
- 2:22:30So all right, how do I encode the idea that it is not raining?
- 2:22:33Well, I can use not and then the rain symbol.
- 2:22:36So here's me saying that it is not raining.
- 2:22:39And now the implication is that if it is not raining,
- 2:22:42then Harry visited Hagrid.
- 2:22:45So I'll wrap this inside of an implication to say,
- 2:22:48if it is not raining, this first argument to the implication
- 2:22:52will then Harry visited Hagrid.
- 2:22:56So I'm saying implication, the premise is that it's not raining.
- 2:23:00And if it is not raining, then Harry visited Hagrid.
- 2:23:04And I can print out knowledge.formula to see the logical formula
- 2:23:07equivalent of that same idea.
- 2:23:09So I run Python of harry.py.
- 2:23:11And this is the logical formula that we see
- 2:23:13as a result, which is a text-based version of what
- 2:23:16we were looking at before, that if it is not raining,
- 2:23:18then that implies that Harry visited Hagrid.
- 2:23:23But there was additional information that we had access to as well.
- 2:23:26In this case, we had access to the fact that Harry visited either Hagrid
- 2:23:31or Dumbledore.
- 2:23:32So how do I encode that?
- 2:23:34Well, this means that in my knowledge, I've really
- 2:23:36got multiple pieces of knowledge going on.
- 2:23:38I know one thing and another thing and another thing.
- 2:23:41So I'll go ahead and wrap all of my knowledge inside of an and.
- 2:23:44And I'll move things on to new lines just for good measure.
- 2:23:47But I know multiple things.
- 2:23:49So I'm saying knowledge is an and of multiple different sentences.
- 2:23:52I know multiple different sentences to be true.
- 2:23:55One such sentence that I know to be true is this implication,
- 2:23:59that if it is not raining, then Harry visited Hagrid.
- 2:24:02Another such sentence that I know to be true is or Hagrid Dumbledore.
- 2:24:08In other words, Hagrid or Dumbledore is true,
- 2:24:12because I know that Harry visited Hagrid or Dumbledore.
- 2:24:16But I know more than that, actually.
- 2:24:17That initial sentence from before said that Harry visited Hagrid or Dumbledore,
- 2:24:22but not both.
- 2:24:23So now I want a sentence that will encode the idea that Harry didn't
- 2:24:26visit both Hagrid and Dumbledore.
- 2:24:29Well, the notion of Harry visiting Hagrid and Dumbledore
- 2:24:33would be represented like this, and of Hagrid and Dumbledore.
- 2:24:38And if that is not true, if I want to say not that,
- 2:24:41then I'll just wrap this whole thing inside of a not.
- 2:24:46So now these three lines, line 8 says that if it is not raining,
- 2:24:50then Harry visited Hagrid.
- 2:24:51Line 9 says Harry visited Hagrid or Dumbledore.
- 2:24:55And line 10 says Harry didn't visit both Hagrid and Dumbledore,
- 2:25:01that it is not true that both the Hagrid symbol and the Dumbledore
- 2:25:04symbol are true.
- 2:25:05Only one of them can be true.
- 2:25:08And finally, the last piece of information that I knew
- 2:25:11was the fact that Harry visited Dumbledore.
- 2:25:15So these now are the pieces of knowledge that I know, one sentence
- 2:25:18and another sentence and another and another.
- 2:25:21And I can print out what I know just to see it a little bit more visually.
- 2:25:24And here now is a logical representation of the information
- 2:25:28that my computer is now internally representing
- 2:25:31using these various different Python objects.
- 2:25:33And again, take a look at logic.py if you want to take a look at how exactly
- 2:25:37it's implementing this, but no need to worry too much about all of the details
- 2:25:40there.
- 2:25:40We're here saying that if it is not raining, then Harry visited Hagrid.
- 2:25:44We're saying that Hagrid or Dumbledore is true.
- 2:25:47And we're saying it is not the case that Hagrid and Dumbledore is true,
- 2:25:52that they're not both true.
- 2:25:54And we also know that Dumbledore is true.
- 2:25:57So this long logical sentence represents our knowledge base.
- 2:26:01It is the thing that we know.
- 2:26:03And now what we'd like to do is we'd like to use model checking
- 2:26:06to ask a query, to ask a question like, based on this information,
- 2:26:10do I know whether or not it's raining?
- 2:26:12And we as humans were able to logic our way through it and figure out that,
- 2:26:15all right, based on these sentences, we can conclude this and that
- 2:26:18to figure out that, yes, it must have been raining.
- 2:26:20But now we'd like for the computer to do that as well.
- 2:26:23So let's take a look at the model checking algorithm
- 2:26:26that is going to follow that same pattern
- 2:26:27that we drew out in pseudocode a moment ago.
- 2:26:30So I've defined a function here in logic.py
- 2:26:32that you can take a look at called model check.
- 2:26:35Model check takes two arguments, the knowledge that I already know,
- 2:26:39and the query.
- 2:26:41And the idea is, in order to do model checking,
- 2:26:43I need to enumerate all of the possible models.
- 2:26:46And for each of the possible models, I need to ask myself,
- 2:26:49is the knowledge base true?
- 2:26:50And is the query true?
- 2:26:52So the first thing I need to do is somehow
- 2:26:54enumerate all of the possible models, meaning
- 2:26:57for all possible symbols that exist, I need
- 2:26:59to assign true and false to each one of them
- 2:27:02and see whether or not it's still true.
- 2:27:05And so here is the way we're going to do that.
- 2:27:07We're going to start.
- 2:27:08So I've defined another helper function internally
- 2:27:10that we'll get to in just a moment.
- 2:27:12But this function starts by getting all of the symbols in both the knowledge
- 2:27:17and the query, by figuring out what symbols am I dealing with.
- 2:27:20In this case, the symbols I'm dealing with are rain and Hagrid and Dumbledore,
- 2:27:24but there might be other symbols depending on the problem.
- 2:27:26And we'll take a look soon at some examples of situations
- 2:27:29where ultimately we're going to need some additional symbols in order
- 2:27:32to represent the problem.
- 2:27:34And then we're going to run this check all function, which
- 2:27:38is a helper function that's basically going to recursively call itself
- 2:27:41checking every possible configuration of propositional symbols.
- 2:27:46So we start out by looking at this check all function.
- 2:27:51And what do we do?
- 2:27:52So if not symbols means if we finish assigning all of the symbols.
- 2:27:57We've assigned every symbol a value.
- 2:27:58So far we haven't done that, but if we ever do, then we check.
- 2:28:03In this model, is the knowledge true?
- 2:28:05That's what this line is saying.
- 2:28:06If we evaluate the knowledge propositional logic formula
- 2:28:10using the model's assignment of truth values, is the knowledge true?
- 2:28:14If the knowledge is true, then we should return true only if the query is true.
- 2:28:19Because if the knowledge is true, we want the query
- 2:28:22to be true as well in order for there to be entailment.
- 2:28:25Otherwise, we don't know that there otherwise there won't be an entailment
- 2:28:29if there's ever a situation where what we know in our knowledge is true,
- 2:28:33but the query, the thing we're asking, happens to be false.
- 2:28:36So this line here is checking that same idea
- 2:28:38that in all worlds where the knowledge is true, the query must also be true.
- 2:28:44Otherwise, we can just return true because if the knowledge isn't true,
- 2:28:47then we don't care.
- 2:28:48This is equivalent to when we were enumerating
- 2:28:50this table from a moment ago.
- 2:28:52In all situations where the knowledge base wasn't true, all of these seven
- 2:28:56rows here, we didn't care whether or not our query was true or not.
- 2:29:00We only care to check whether the query is true
- 2:29:03when the knowledge base is actually true, which was just this green highlighted
- 2:29:06row right there.
- 2:29:08So that logic is encoded using that statement there.
- 2:29:12And otherwise, if we haven't assigned symbols yet,
- 2:29:15which we haven't seen anything yet, then the first thing we do
- 2:29:18is pop one of the symbols.
- 2:29:20I make a copy of the symbols first just to save an existing copy.
- 2:29:23But I pop one symbol off of the remaining symbols
- 2:29:26so that I just pick one symbol at random.
- 2:29:29And I create one copy of the model where that symbol is true.
- 2:29:33And I create a second copy of the model where that symbol is false.
- 2:29:38So I now have two copies of the model, one where the symbol is true
- 2:29:41and one where the symbol is false.
- 2:29:43And I need to make sure that this entailment holds in both of those models.
- 2:29:47So I recursively check all on the model where the statement is true
- 2:29:52and check all on the model where the statement is false.
- 2:29:57So again, you can take a look at that function
- 2:29:59to try to get a sense for how exactly this logic is working.
- 2:30:02But in effect, what it's doing is recursively
- 2:30:03calling this check all function again and again and again.
- 2:30:07And on every level of the recursion, we're
- 2:30:09saying let's pick a new symbol that we haven't yet assigned,
- 2:30:13assign it to true and assign it to false,
- 2:30:16and then check to make sure that the entailment holds in both cases.
- 2:30:19Because ultimately, I need to check every possible world.
- 2:30:22I need to take every combination of symbols
- 2:30:24and try every combination of true and false
- 2:30:27in order to figure out whether the entailment relation actually holds.
- 2:30:31So that function we've written for you.
- 2:30:34But in order to use that function inside of harry.py,
- 2:30:37what I'll write is something like this.
- 2:30:39I would like to model check based on the knowledge.
- 2:30:43And then I provide as a second argument what the query is,
- 2:30:46what the thing I want to ask is.
- 2:30:48And what I want to ask in this case is, is it raining?
- 2:30:51So model check again takes two arguments.
- 2:30:54The first argument is the information that I know, this knowledge,
- 2:30:57which in this case is this information that was given to me at the beginning.
- 2:31:01And the second argument, rain, is encoding the idea of the query.
- 2:31:06What am I asking?
- 2:31:07I would like to ask, based on this knowledge,
- 2:31:10do I know for sure that it is raining?
- 2:31:13And I can try and print out the result of that.
- 2:31:17And when I run this program, I see that the answer is true.
- 2:31:20That based on this information, I can conclusively
- 2:31:23say that it is raining, because using this model checking algorithm,
- 2:31:26we were able to check that in every world where this knowledge is true,
- 2:31:30it is raining.
- 2:31:31In other words, there is no world where this knowledge is true,
- 2:31:35and it is not raining.
- 2:31:36So you can conclude that it is, in fact, raining.
- 2:31:41And this sort of logic can be applied to a number
- 2:31:43of different types of problems, that if confronted with a problem where
- 2:31:47some sort of logical deduction can be used in order to try to solve it,
- 2:31:50you might try thinking about what propositional symbols you might
- 2:31:54need in order to represent that information,
- 2:31:56and what statements and propositional logic
- 2:31:58you might use in order to encode that information which you know.
- 2:32:03And this process of trying to take a problem
- 2:32:05and figure out what propositional symbols to use in order
- 2:32:08to encode that idea, or how to represent it logically,
- 2:32:11is known as knowledge engineering.
- 2:32:13That software engineers and AI engineers will take a problem
- 2:32:16and try and figure out how to distill it down
- 2:32:19into knowledge that is representable by a computer.
- 2:32:22And if we can take any general purpose problem, some problem
- 2:32:25that we find in the human world, and turn it
- 2:32:27into a problem that computers know how to solve
- 2:32:30as by using any number of different variables, well,
- 2:32:32then we can take a computer that is able to do something
- 2:32:35like model checking or some other inference algorithm
- 2:32:37and actually figure out how to solve that problem.
- 2:32:41So now we'll take a look at two or three examples of knowledge engineering
- 2:32:45and practice, of taking some problem and figuring out
- 2:32:47how we can apply logical symbols and use logical formulas
- 2:32:51to be able to encode that idea.
- 2:32:53And we'll start with a very popular board game in the US and the UK
- 2:32:57known as Clue.
- 2:32:58Now, in the game of Clue, there's a number of different factors
- 2:33:00that are going on.
- 2:33:01But the basic premise of the game, if you've never played it before,
- 2:33:04is that there are a number of different people.
- 2:33:06For now, we'll just use three, Colonel Mustard, Professor Plumb,
- 2:33:09and Miss Scarlet.
- 2:33:10There are a number of different rooms, like a ballroom, a kitchen,
- 2:33:12and a library.
- 2:33:13And there are a number of different weapons, a knife, a revolver, and a wrench.
- 2:33:17And three of these, one person, one room, and one weapon,
- 2:33:21is the solution to the mystery, the murderer and what room they were in
- 2:33:26and what weapon they happened to use.
- 2:33:28And what happens at the beginning of the game
- 2:33:30is that all these cards are randomly shuffled together.
- 2:33:32And three of them, one person, one room, and one weapon,
- 2:33:35are placed into a sealed envelope that we don't know.
- 2:33:37And we would like to figure out, using some sort of logical process,
- 2:33:41what's inside the envelope, which person, which room, and which weapon.
- 2:33:45And we do so by looking at some, but not all, of these cards here,
- 2:33:50by looking at these cards to try and figure out what might be going on.
- 2:33:54And so this is a very popular game.
- 2:33:56But let's now try and formalize it and see
- 2:33:58if we could train a computer to be able to play this game by reasoning
- 2:34:01through it logically.
- 2:34:04So in order to do this, we'll begin by thinking about what
- 2:34:06propositional symbols we're ultimately going to need.
- 2:34:09Remember, again, that propositional symbols are just some symbol,
- 2:34:12some variable, that can be either true or false in the world.
- 2:34:17And so in this case, the propositional symbols
- 2:34:20are really just going to correspond to each of the possible things that
- 2:34:25could be inside the envelope.
- 2:34:26Mustard is a propositional symbol that, in this case,
- 2:34:29will just be true if Colonel Mustard is inside the envelope,
- 2:34:32if he is the murderer, and false otherwise.
- 2:34:35And likewise for Plum, for Professor Plum, and Scarlet, for Miss Scarlet.
- 2:34:38And likewise for each of the rooms and for each of the weapons.
- 2:34:41We have one propositional symbol for each of these ideas.
- 2:34:46Then using those propositional symbols, we
- 2:34:48can begin to create logical sentences, create knowledge
- 2:34:52that we know about the world.
- 2:34:54So for example, we know that someone is the murderer,
- 2:34:57that one of the three people is, in fact, the murderer.
- 2:35:00And how would we encode that?
- 2:35:01Well, we don't know for sure who the murderer is.
- 2:35:04But we know it is one person or the second person or the third person.
- 2:35:09So I could say something like this.
- 2:35:10Mustard or Plum or Scarlet.
- 2:35:13And this piece of knowledge encodes that one of these three people
- 2:35:17is the murderer.
- 2:35:17We don't know which, but one of these three things must be true.
- 2:35:22What other information do we know?
- 2:35:24Well, we know that, for example, one of the rooms
- 2:35:26must have been the room in the envelope.
- 2:35:28The crime was committed either in the ballroom or the kitchen or the library.
- 2:35:33Again, right now, we don't know which.
- 2:35:34But this is knowledge we know at the outset,
- 2:35:36knowledge that one of these three must be inside the envelope.
- 2:35:40And likewise, we can say the same thing about the weapon,
- 2:35:42that it was either the knife or the revolver or the wrench,
- 2:35:45that one of those weapons must have been the weapon of choice
- 2:35:48and therefore the weapon in the envelope.
- 2:35:51And then as the game progresses, the gameplay
- 2:35:53works by people get various different cards.
- 2:35:55And using those cards, you can deduce information.
- 2:35:59That if someone gives you a card, for example,
- 2:36:01I have the Professor Plum card in my hand,
- 2:36:04then I know the Professor Plum card can't be inside the envelope.
- 2:36:07I know that Professor Plum is not the criminal,
- 2:36:11so I know a piece of information like not Plum, for example.
- 2:36:15I know that Professor Plum has to be false.
- 2:36:18This propositional symbol is not true.
- 2:36:21And sometimes I might not know for sure that a particular card is not
- 2:36:24in the middle, but sometimes someone will make a guess
- 2:36:27and I'll know that one of three possibilities is not true.
- 2:36:30Someone will guess Colonel Mustard in the library with the revolver
- 2:36:33or something to that effect.
- 2:36:35And in that case, a card might be revealed that I don't see.
- 2:36:38But if it is a card and it is either Colonel Mustard or the revolver
- 2:36:43or the library, then I know that at least one of them
- 2:36:46can't be in the middle.
- 2:36:47So I know something like it is either not Mustard
- 2:36:51or it is not the library or it is not the revolver.
- 2:36:55Now maybe multiple of these are not true,
- 2:36:57but I know that at least one of Mustard, Library, and Revolver
- 2:37:01must, in fact, be false.
- 2:37:03And so this now is a propositional logic representation
- 2:37:07of this game of Clue, a way of encoding the knowledge that we
- 2:37:10know inside this game using propositional logic
- 2:37:13that a computer algorithm, something like model checking
- 2:37:15that we saw a moment ago, can actually look at and understand.
- 2:37:19So let's now take a look at some code to see
- 2:37:21how this algorithm might actually work in practice.
- 2:37:26All right, so I'm now going to open up a file called Clue.py, which
- 2:37:30I've started already.
- 2:37:31And what we'll see here is I've defined a couple of things.
- 2:37:33To find some symbols initially, notice I
- 2:37:35have a symbol for Colonel Mustard, a symbol for Professor Plum,
- 2:37:38a symbol for Miss Scarlett, all of which
- 2:37:40I've put inside of this list of characters.
- 2:37:42I have a symbol for Ballroom and Kitchen and Library
- 2:37:45inside of a list of rooms.
- 2:37:46And then I have symbols for Knife and Revolver and Wrench.
- 2:37:49These are my weapons.
- 2:37:50And so all of these characters and rooms and weapons altogether,
- 2:37:53those are my symbols.
- 2:37:55And now I also have this check knowledge function.
- 2:37:59And what the check knowledge function does is it takes my knowledge
- 2:38:02and it's going to try and draw conclusions about what I know.
- 2:38:07So for example, we'll loop over all of the possible symbols
- 2:38:10and we'll check, do I know that that symbol is true?
- 2:38:13And a symbol is going to be something like Professor Plum
- 2:38:15or the Knife or the Library.
- 2:38:17And if I know that it is true, in other words,
- 2:38:19I know that it must be the card in the envelope,
- 2:38:22then I'm going to print out using a function called
- 2:38:24cprint, which prints things in color.
- 2:38:26I'm going to print out the word yes, and I'm
- 2:38:28going to print that in green, just to make it very clear to us.
- 2:38:32If we're not sure that the symbol is true,
- 2:38:35maybe I can check to see if I'm sure that the symbol is not true.
- 2:38:38Like if I know for sure that it is not Professor Plum, for example.
- 2:38:42And I do that by running model check again,
- 2:38:44this time checking if my knowledge is not the symbol,
- 2:38:48if I know for sure that the symbol is not true.
- 2:38:52And if I don't know for sure that the symbol is not true,
- 2:38:55because I say if not model check, meaning I'm not sure that the symbol is
- 2:38:59false, well, then I'll go ahead and print out maybe next to the symbol.
- 2:39:03Because maybe the symbol is true, maybe it's not, I don't actually know.
- 2:39:07So what knowledge do I actually have?
- 2:39:10Well, let's try and represent my knowledge now.
- 2:39:12So my knowledge is, I know a couple of things, so I'll put them in an and.
- 2:39:16And I know that one of the three people must be the criminal.
- 2:39:20So I know or mustard, plum, scarlet.
- 2:39:23This is my way of encoding that it is either Colonel Mustard or Professor
- 2:39:26Plum or Miss Scarlet.
- 2:39:28I know that it must have happened in one of the rooms.
- 2:39:31So I know or ballroom, kitchen, library, for example.
- 2:39:36And I know that one of the weapons must have been used as well.
- 2:39:38So I know or knife, revolver, wrench.
- 2:39:43So that might be my initial knowledge, that I
- 2:39:45know that it must have been one of the people,
- 2:39:47I know it must have been in one of the rooms,
- 2:39:48and I know that it must have been one of the weapons.
- 2:39:51And I can see what that knowledge looks like as a formula
- 2:39:54by printing out knowledge.formula.
- 2:39:56So I'll run python clue.py.
- 2:39:58And here now is the information that I know in logical format.
- 2:40:02I know that it is Colonel Mustard or Professor Plum or Miss Scarlet.
- 2:40:05And I know that it is the ballroom, the kitchen, or the library.
- 2:40:08And I know that it is the knife, the revolver, or the wrench.
- 2:40:11But I don't know much more than that.
- 2:40:13I can't really draw any firm conclusions.
- 2:40:16And in fact, we can see that if I try and do,
- 2:40:19let me go ahead and run my knowledge check function on my knowledge.
- 2:40:24Knowledge check is this function that I, or check knowledge rather,
- 2:40:27is this function that I just wrote that looks over all of the symbols
- 2:40:31and tries to see what conclusions I can actually
- 2:40:33draw about any of the symbols.
- 2:40:36So I'll go ahead and run clue.py and see what it is that I know.
- 2:40:41And it seems that I don't really know anything for sure.
- 2:40:43I have all three people are maybes, all three of the rooms are maybes,
- 2:40:47all three of the weapons are maybes.
- 2:40:48I don't really know anything for certain just yet.
- 2:40:52But now let me try and add some additional information
- 2:40:54and see if additional information, additional knowledge,
- 2:40:57can help us to logically reason our way through this process.
- 2:41:00And we are just going to provide the information.
- 2:41:02Our AI is going to take care of doing the inference
- 2:41:05and figuring out what conclusions it's able to draw.
- 2:41:09So I start with some cards.
- 2:41:11And those cards tell me something.
- 2:41:12So if I have the kernel mustard card, for example,
- 2:41:15I know that the mustard symbol must be false.
- 2:41:19In other words, mustard is not the one in the envelope,
- 2:41:22is not the criminal.
- 2:41:23So I can say, knowledge supports something called,
- 2:41:26every and in this library supports dot add,
- 2:41:30which is a way of adding knowledge or adding
- 2:41:32an additional logical sentence to an and clause.
- 2:41:35So I can say, knowledge dot add, not mustard.
- 2:41:40I happen to know, because I have the mustard card,
- 2:41:42that kernel mustard is not the suspect.
- 2:41:44And maybe I have a couple of other cards too.
- 2:41:46Maybe I also have a card for the kitchen.
- 2:41:49So I know it's not the kitchen.
- 2:41:50And maybe I have another card that says that it is not the revolver.
- 2:41:54So I have three cards, kernel mustard, the kitchen, and the revolver.
- 2:41:57And I encode that into my AI this way by saying, it's not kernel mustard,
- 2:42:01it's not the kitchen, and it's not the revolver.
- 2:42:04And I know those to be true.
- 2:42:06So now, when I rerun clue.py, we'll see that I've
- 2:42:09been able to eliminate some possibilities.
- 2:42:12Before, I wasn't sure if it was the knife or the revolver or the wrench.
- 2:42:15If a knife was maybe, a revolver was maybe, wrench is maybe.
- 2:42:18Now I'm down to just the knife and the wrench.
- 2:42:21Between those two, I don't know which one it is.
- 2:42:23They're both maybes.
- 2:42:24But I've been able to eliminate the revolver, which
- 2:42:27is one that I know to be false, because I have the revolver card.
- 2:42:31And so additional information might be acquired
- 2:42:34over the course of this game.
- 2:42:36And we would represent that just by adding knowledge to our knowledge set
- 2:42:41or knowledge base that we've been building here.
- 2:42:43So if, for example, we additionally got the information
- 2:42:46that someone made a guess, someone guessed like Miss Scarlet
- 2:42:49in the library with the wrench.
- 2:42:51And we know that a card was revealed, which
- 2:42:53means that one of those three cards, either Miss Scarlet
- 2:42:56or the library or the wrench, one of those at minimum
- 2:42:59must not be inside of the envelope.
- 2:43:02So I could add some knowledge, say knowledge.add.
- 2:43:05And I'm going to add an or clause, because I don't know for sure which one
- 2:43:09it's not, but I know one of them is not in the envelope.
- 2:43:12So it's either not Scarlet, or it's not the library,
- 2:43:15and or supports multiple arguments.
- 2:43:17I can say it's also or not the wrench.
- 2:43:20So at least one of those needs a Scarlet library and wrench.
- 2:43:23At least one of those needs to be false.
- 2:43:25I don't know which, though.
- 2:43:26Maybe it's multiple.
- 2:43:27Maybe it's just one, but at least one I know needs to hold.
- 2:43:32And so now if I rerun clue.py, I don't actually
- 2:43:35have any additional information just yet.
- 2:43:37Nothing I can say conclusively.
- 2:43:38I still know that maybe it's Professor Plum, maybe it's Miss Scarlet.
- 2:43:41I haven't eliminated any options.
- 2:43:44But let's imagine that I get some more information,
- 2:43:46that someone shows me the Professor Plum card, for example.
- 2:43:50So I say, all right, let's go back here, knowledge.add, not Plum.
- 2:43:57So I have the Professor Plum card.
- 2:43:58I know the Professor Plum is not in the middle.
- 2:44:00I rerun clue.py.
- 2:44:02And right now, I'm able to draw some conclusions.
- 2:44:04Now I've been able to eliminate Professor Plum,
- 2:44:07and the only person it could left remaining be is Miss Scarlet.
- 2:44:10So I know, yes, Miss Scarlet, this variable must be true.
- 2:44:14And I've been able to infer that based on the information I already had.
- 2:44:17Now between the ballroom and the library and the knife and the wrench,
- 2:44:20for those two, I'm still not sure.
- 2:44:22So let's add one more piece of information.
- 2:44:25Let's say that I know that it's not the ballroom.
- 2:44:28Someone has shown me the ballroom card, so I know it's not the ballroom.
- 2:44:30Which means at this point, I should be able to conclude that it's the library.
- 2:44:33Let's see.
- 2:44:35I'll say knowledge.add, not the ballroom.
- 2:44:40And we'll go ahead and run that.
- 2:44:43And it turns out that after all of this, not only can I conclude that I
- 2:44:46know that it's the library, but I also know that the weapon was the knife.
- 2:44:49And that might have been an inference that was a little bit trickier, something
- 2:44:52I wouldn't have realized immediately, but the AI,
- 2:44:55via this model checking algorithm, is able to draw that conclusion,
- 2:44:58that we know for sure that it must be Miss Scarlet in the library with the knife.
- 2:45:02And how did we know that?
- 2:45:03Well, we know it from this or clause up here,
- 2:45:07that we know that it's either not Scarlet, or it's not the library,
- 2:45:11or it's not the wrench.
- 2:45:13And given that we know that it is Miss Scarlet,
- 2:45:16and we know that it is the library, then the only remaining option for the weapon
- 2:45:20is that it is not the wrench, which means that it must be the knife.
- 2:45:24So we as humans now can go back and reason through that,
- 2:45:26even though it might not have been immediately clear.
- 2:45:28And that's one of the advantages of using an AI or some sort of algorithm
- 2:45:32in order to do this, is that the computer can exhaust all of these possibilities
- 2:45:36and try and figure out what the solution actually should be.
- 2:45:40And so for that reason, it's often helpful to be
- 2:45:43able to represent knowledge in this way.
- 2:45:45Knowledge engineering, some situation where
- 2:45:47we can use a computer to be able to represent knowledge
- 2:45:50and draw conclusions based on that knowledge.
- 2:45:52And any time we can translate something into propositional logic symbols
- 2:45:56like this, this type of approach can be useful.
- 2:45:59So you might be familiar with logic puzzles,
- 2:46:01where you have to puzzle your way through trying to figure something out.
- 2:46:04This is what a classic logic puzzle might look like.
- 2:46:06Something like Gilderoy, Minerva, Pomona, and Horace each
- 2:46:09belong to a different one of the four houses, Gryffindor, Hufflepuff, Ravenclaw,
- 2:46:14and Slytherin.
- 2:46:15And then we have some information.
- 2:46:16The Gilderoy belongs to Gryffindor or Ravenclaw, Pomona
- 2:46:20does not belong in Slytherin, and Minerva does belong to Gryffindor.
- 2:46:24So we have a couple pieces of information.
- 2:46:26And using that information, we need to be
- 2:46:28able to draw some conclusions about which person should
- 2:46:31be assigned to which house.
- 2:46:33And again, we can use the exact same idea to try and implement this notion.
- 2:46:37So we need some propositional symbols.
- 2:46:39And in this case, the propositional symbols
- 2:46:41are going to get a little more complex, although we'll
- 2:46:43see ways to make this a little bit cleaner later on.
- 2:46:46But we'll need 16 propositional symbols, one for each person and house.
- 2:46:51So we need to say, remember, every propositional symbol
- 2:46:54is either true or false.
- 2:46:56So Gilderoy Gryffindor is either true or false.
- 2:46:59Either he's in Gryffindor or he is not.
- 2:47:01Likewise, Gilderoy Hufflepuff also true or false.
- 2:47:03Either it is true or it's false.
- 2:47:05And that's true for every combination of person and house
- 2:47:09that we could come up with.
- 2:47:10We have some sort of propositional symbol for each one of those.
- 2:47:14Using this type of knowledge, we can then
- 2:47:17begin to think about what types of logical sentences
- 2:47:20we can say about the puzzle.
- 2:47:22That if we know what will before even think about the information we were
- 2:47:25given, we can think about the premise of the problem,
- 2:47:28that every person is assigned to a different house.
- 2:47:31So what does that tell us?
- 2:47:32Well, it tells us sentences like this.
- 2:47:34It tells us like Pomona Slytherin implies not Pomona Hufflepuff.
- 2:47:39Something like if Pomona is in Slytherin,
- 2:47:42then we know that Pomona is not in Hufflepuff.
- 2:47:44And we know this for all four people and for all combinations of houses,
- 2:47:48that no matter what person you pick, if they're in one house,
- 2:47:51then they're not in some other house.
- 2:47:53So I'll probably have a whole bunch of knowledge statements
- 2:47:56that are of this form, that if we know Pomona is in Slytherin,
- 2:47:59then we know Pomona is not in Hufflepuff.
- 2:48:01We were also given the information that each person
- 2:48:04is in a different house.
- 2:48:05So I also have pieces of knowledge that look something like this.
- 2:48:08Minerva Ravenclaw implies not Gilderoy Ravenclaw.
- 2:48:13If they're all in different houses, then if Minerva is in Ravenclaw,
- 2:48:16then we know the Gilderoy is not in Ravenclaw as well.
- 2:48:20And I have a whole bunch of similar sentences
- 2:48:22like this that are expressing that idea for other people and other houses
- 2:48:26as well.
- 2:48:27And so in addition to sentences of these form,
- 2:48:29I also have the knowledge that was given to me.
- 2:48:32Information like Gilderoy was in Gryffindor or in Ravenclaw
- 2:48:35that would be represented like this, Gilderoy Gryffindor or Gilderoy
- 2:48:39Ravenclaw.
- 2:48:40And then using these sorts of sentences,
- 2:48:42I can begin to draw some conclusions about the world.
- 2:48:46So let's see an example of this.
- 2:48:48We'll go ahead and actually try and implement this logic puzzle
- 2:48:50to see if we can figure out what the answer is.
- 2:48:53I'll go ahead and open up puzzle.py, where I've already
- 2:48:56started to implement this sort of idea.
- 2:48:58I've defined a list of people and a list of houses.
- 2:49:01And I've so far created one symbol for every person and for every house.
- 2:49:06That's what this double four loop is doing, looping over all people,
- 2:49:09looping over all houses, creating a new symbol for each of them.
- 2:49:13And then I've added some information.
- 2:49:16I know that every person belongs to a house,
- 2:49:19so I've added the information for every person that person Gryffindor
- 2:49:24or person Hufflepuff or person Ravenclaw or person Slytherin,
- 2:49:28that one of those four things must be true.
- 2:49:30Every person belongs to a house.
- 2:49:33What other information do I know?
- 2:49:34I also know that only one house per person,
- 2:49:37so no person belongs to multiple houses.
- 2:49:41So how does this work?
- 2:49:42Well, this is going to be true for all people.
- 2:49:44So I'll loop over every person.
- 2:49:47And then I need to loop over all different pairs of houses.
- 2:49:51The idea is I want to encode the idea that if Minerva is in Gryffindor,
- 2:49:54then Minerva can't be in Ravenclaw.
- 2:49:57So I'll loop over all houses, each one.
- 2:49:59And I'll loop over all houses again, h2.
- 2:50:02And as long as they're different, h1 not equal to h2,
- 2:50:06then I'll add to my knowledge base this piece of information.
- 2:50:09That implication, in other words, an if then, if the person is in h1,
- 2:50:14then I know that they are not in house h2.
- 2:50:18So these lines here are encoding the notion that for every person,
- 2:50:22if they belong to house one, then they are not in house two.
- 2:50:25And the other piece of logic we need to encode
- 2:50:27is the idea that every house can only have one person.
- 2:50:30In other words, if Pomona is in Hufflepuff,
- 2:50:33then nobody else is allowed to be in Hufflepuff either.
- 2:50:35And that's the same logic, but sort of backwards.
- 2:50:37I loop over all of the houses and loop over all different pairs of people.
- 2:50:42So I loop over people once, loop over people again,
- 2:50:45and only do this when the people are different, p1 not equal to p2.
- 2:50:50And I add the knowledge that if, as given by the implication,
- 2:50:54if person one belongs to the house, then it
- 2:50:58is not the case that person two belongs to the same house.
- 2:51:03So here I'm just encoding the knowledge that
- 2:51:05represents the problem's constraints.
- 2:51:07I know that everyone's in a different house.
- 2:51:09I know that any person can only belong to one house.
- 2:51:12And I can now take my knowledge and try and print out the information
- 2:51:17that I happen to know.
- 2:51:18So I'll go ahead and print out knowledge.formula,
- 2:51:22just to see this in action, and I'll go ahead and skip this for now.
- 2:51:24But we'll come back to this in a second.
- 2:51:26Let's print out the knowledge that I know by running Python puzzle.py.
- 2:51:31It's a lot of information, a lot that I have to scroll through,
- 2:51:34because there are 16 different variables all going on.
- 2:51:36But the basic idea, if we scroll up to the very top,
- 2:51:39is I see my initial information.
- 2:51:41Gilderoy is either in Gryffindor, or Gilderoy is in Hufflepuff,
- 2:51:44or Gilderoy is in Ravenclaw, or Gilderoy is in Slytherin,
- 2:51:48and then way more information as well.
- 2:51:50So this is quite messy, more than we really want to be looking at.
- 2:51:54And soon, too, we'll see ways of representing
- 2:51:55this a little bit more nicely using logic.
- 2:51:58But for now, we can just say these are the variables
- 2:52:00that we're dealing with.
- 2:52:01And now we'd like to add some information.
- 2:52:05So the information we're going to add is Gilderoy is in Gryffindor,
- 2:52:09or he is in Ravenclaw.
- 2:52:10So that knowledge was given to us.
- 2:52:12So I'll go ahead and say knowledge.add.
- 2:52:15And I know that either or Gilderoy Gryffindor or Gilderoy Ravenclaw.
- 2:52:26One of those two things must be true.
- 2:52:29I also know that Pomona was not in Slytherin,
- 2:52:32so I can say knowledge.add not this symbol, not the Pomona-Slytherin
- 2:52:37symbol.
- 2:52:38And then I can add the knowledge that Minerva is in Gryffindor
- 2:52:42by adding the symbol Minerva Gryffindor.
- 2:52:46So those are the pieces of knowledge that I know.
- 2:52:49And this loop here at the bottom just loops over all of my symbols,
- 2:52:52checks to see if the knowledge entails that symbol
- 2:52:56by calling this model check function again.
- 2:52:58And if it does, if we know the symbol is true, we print out the symbol.
- 2:53:03So now I can run Python, puzzle.py, and Python
- 2:53:07is going to solve this puzzle for me.
- 2:53:08We're able to conclude that Gilderoy belongs to Ravenclaw,
- 2:53:11Pomona belongs to Hufflepuff, Minerva to Gryffindor, and Horace to Slytherin
- 2:53:15just by encoding this knowledge inside the computer,
- 2:53:18although it was quite tedious to do in this case.
- 2:53:20And as a result, we were able to get the conclusion from that as well.
- 2:53:24And you can imagine this being applied to many sorts
- 2:53:27of different deductive situations.
- 2:53:29So not only these situations where we're trying
- 2:53:31to deal with Harry Potter characters in this puzzle,
- 2:53:33but if you've ever played games like Mastermind, where
- 2:53:35you're trying to figure out which order different colors go in
- 2:53:39and trying to make predictions about it, I
- 2:53:40could tell you, for example, let's play a simplified version of Mastermind
- 2:53:44where there are four colors, red, blue, green, and yellow,
- 2:53:47and they're in some order, but I'm not telling you what order.
- 2:53:51You just have to make a guess, and I'll tell you
- 2:53:53of red, blue, green, and yellow how many of the four
- 2:53:55you got in the right position.
- 2:53:57So a simplified version of this game, you
- 2:53:59might make a guess like red, blue, green, yellow,
- 2:54:01and I would tell you something like two of those four
- 2:54:05are in the correct position, but the other two are not.
- 2:54:08And then you could reasonably make a guess and say, all right,
- 2:54:10look at this, blue, red, green, yellow.
- 2:54:13Try switching two of them around, and this time maybe I tell you,
- 2:54:16you know what, none of those are in the correct position.
- 2:54:19And the question then is, all right, what is the correct order
- 2:54:23of these four colors?
- 2:54:24And we as humans could begin to reason this through.
- 2:54:26All right, well, if none of these were correct,
- 2:54:28but two of these were correct, well, it must have been
- 2:54:31because I switched the red and the blue, which means red and blue here
- 2:54:34must be correct, which means green and yellow are probably not correct.
- 2:54:37You can begin to do this sort of deductive reasoning.
- 2:54:40And we can also equivalently try and take this
- 2:54:42and encode it inside of our computer as well.
- 2:54:45And it's going to be very similar to the logic puzzle
- 2:54:48that we just did a moment ago.
- 2:54:49So I won't spend too much time on this code because it is fairly similar.
- 2:54:52But again, we have a whole bunch of colors
- 2:54:54and four different positions in which those colors can be.
- 2:54:58And then we have some additional knowledge.
- 2:55:00And I encode all of that knowledge.
- 2:55:02And you can take a look at this code on your own time.
- 2:55:04But I just want to demonstrate that when we run this code,
- 2:55:07run python mastermind.py and run and see what we get,
- 2:55:12we ultimately are able to compute red 0 in the 0 position,
- 2:55:16blue in the 1 position, yellow in the 2 position,
- 2:55:19and green in the 3 position as the ordering of those symbols.
- 2:55:24Now, ultimately, what you might have noticed
- 2:55:25is this process was taking quite a long time.
- 2:55:28And in fact, model checking is not a particularly efficient algorithm, right?
- 2:55:32What I need to do in order to model check
- 2:55:34is take all of my possible different variables
- 2:55:36and enumerate all of the possibilities that they could be in.
- 2:55:39If I have n variables, I have 2 to the n possible worlds
- 2:55:44that I need to be looking through in order
- 2:55:45to perform this model checking algorithm.
- 2:55:48And this is probably not tractable, especially
- 2:55:50as we start to get to much larger and larger sets of data
- 2:55:53where you have many, many more variables that are at play.
- 2:55:56Right here, we only have a relatively small number of variables.
- 2:55:59So this sort of approach can actually work.
- 2:56:01But as the number of variables increases, model checking
- 2:56:04becomes less and less good of a way of trying
- 2:56:07to solve these sorts of problems.
- 2:56:09So while it might have been OK for something like Mastermind
- 2:56:12to conclude that this is indeed the correct sequence where all four
- 2:56:15are in the correct position, what we'd like to do
- 2:56:17is come up with some better ways to be able to make inferences rather than
- 2:56:21just enumerate all of the possibilities.
- 2:56:24And to do so, what we'll transition to next
- 2:56:26is the idea of inference rules, some sort of rules
- 2:56:29that we can apply to take knowledge that already exists
- 2:56:33and translate it into new forms of knowledge.
- 2:56:36And the general way we'll structure an inference rule
- 2:56:38is by having a horizontal line here.
- 2:56:40Anything above the line is going to represent a premise, something
- 2:56:44that we know to be true.
- 2:56:45And then anything below the line will be the conclusion
- 2:56:48that we can arrive at after we apply the logic from the inference rule
- 2:56:53that we're going to demonstrate.
- 2:56:54So we'll do some of these inference rules
- 2:56:56by demonstrating them in English first, but then translating them
- 2:56:59into the world of propositional logic so you
- 2:57:01can see what those inference rules actually look like.
- 2:57:04So for example, let's imagine that I have access
- 2:57:07to two pieces of information.
- 2:57:08I know, for example, that if it is raining,
- 2:57:11then Harry is inside, for example.
- 2:57:14And let's say I also know it is raining.
- 2:57:16Then most of us could reasonably then look at this information
- 2:57:19and conclude that, all right, Harry must be inside.
- 2:57:23This inference rule is known as modus ponens,
- 2:57:27and it's phrased more formally in logic as this.
- 2:57:29If we know that alpha implies beta, in other words, if alpha, then beta,
- 2:57:35and we also know that alpha is true, then we
- 2:57:38should be able to conclude that beta is also true.
- 2:57:41We can apply this inference rule to take these two pieces of information
- 2:57:45and generate this new piece of information.
- 2:57:47Notice that this is a totally different approach from the model checking
- 2:57:51approach, where the approach was look at all of the possible worlds
- 2:57:54and see what's true in each of these worlds.
- 2:57:56Here, we're not dealing with any specific world.
- 2:57:59We're just dealing with the knowledge that we know
- 2:58:01and what conclusions we can arrive at based on that knowledge.
- 2:58:04That I know that A implies B, and I know A, and the conclusion is B.
- 2:58:10And this should seem like a relatively obvious rule.
- 2:58:12But of course, if alpha, then beta, and we know alpha,
- 2:58:16then we should be able to conclude that beta is also true.
- 2:58:19And that's going to be true for many, but maybe even
- 2:58:21all of the inference rules that we'll take a look at.
- 2:58:23You should be able to look at them and say,
- 2:58:25yeah, of course that's going to be true.
- 2:58:27But it's putting these all together, figuring out the right combination
- 2:58:30of inference rules that can be applied that ultimately
- 2:58:32is going to allow us to generate interesting knowledge inside of our AI.
- 2:58:38So that's modus ponensis application of implication,
- 2:58:41that if we know alpha and we know that alpha implies beta,
- 2:58:44then we can conclude beta.
- 2:58:47Let's take a look at another example.
- 2:58:48Fairly straightforward, something like Harry is friends with Ron and Hermione.
- 2:58:52Based on that information, we can reasonably
- 2:58:54conclude Harry is friends with Hermione.
- 2:58:56That must also be true.
- 2:58:58And this inference rule is known as and elimination.
- 2:59:01And what and elimination says is that if we have a situation where alpha
- 2:59:06and beta are both true, I have information alpha and beta,
- 2:59:11well then, just alpha is true.
- 2:59:14Or likewise, just beta is true.
- 2:59:16That if I know that both parts are true, then one of those parts
- 2:59:19must also be true.
- 2:59:21Again, something obvious from the point of view of human intuition,
- 2:59:24but a computer needs to be told this kind of information.
- 2:59:27To be able to apply the inference rule, we
- 2:59:28need to tell the computer that this is an inference rule that you can apply,
- 2:59:32so the computer has access to it and is able to use it
- 2:59:35in order to translate information from one form to another.
- 2:59:39In addition to that, let's take a look at another example of an inference
- 2:59:42rule, something like it is not true that Harry did not pass the test.
- 2:59:48Bit of a tricky sentence to parse.
- 2:59:50I'll read it again.
- 2:59:50It is not true, or it is false, that Harry did not pass the test.
- 2:59:54Well, if it is false that Harry did not pass the test,
- 2:59:58then the only reasonable conclusion is that Harry did pass the test.
- 3:00:02And so this, instead of being and elimination,
- 3:00:05is what we call double negation elimination.
- 3:00:07That if we have two negatives inside of our premise,
- 3:00:10then we can just remove them altogether.
- 3:00:12They cancel each other out.
- 3:00:13One turns true to false, and the other one turns false back into true.
- 3:00:17Phrased a little bit more formally, we say
- 3:00:19that if the premise is not alpha, then the conclusion
- 3:00:23we can draw is just alpha.
- 3:00:25We can say that alpha is true.
- 3:00:28We'll take a look at a couple more of these.
- 3:00:30If I have it is raining, then Harry is inside.
- 3:00:33How do I reframe this?
- 3:00:35Well, this one is a little bit trickier.
- 3:00:37But if I know if it is raining, then Harry is inside,
- 3:00:41then I conclude one of two things must be true.
- 3:00:43Either it is not raining, or Harry is inside.
- 3:00:48Now, this one's trickier.
- 3:00:49So let's think about it a little bit.
- 3:00:50This first premise here, if it is raining, then Harry is inside,
- 3:00:54is saying that if I know that it is raining, then Harry must be inside.
- 3:00:59So what is the other possible case?
- 3:01:01Well, if Harry is not inside, then I know that it must not be raining.
- 3:01:06So one of those two situations must be true.
- 3:01:09Either it's not raining, or it is raining, in which case Harry is inside.
- 3:01:14So the conclusion I can draw is either it is not raining,
- 3:01:18or it is raining, so therefore, Harry is inside.
- 3:01:22And so this is a way to translate if-then statements into or statements.
- 3:01:28And this is known as implication elimination.
- 3:01:31And this is similar to what we actually did in the beginning
- 3:01:33when we were first looking at those very first sentences
- 3:01:35about Harry and Hagrid and Dumbledore.
- 3:01:37And phrased a little bit more formally, this
- 3:01:39says that if I have the implication, alpha implies beta,
- 3:01:43that I can draw the conclusion that either not alpha or beta,
- 3:01:49because there are only two possibilities.
- 3:01:50Either alpha is true or alpha is not true.
- 3:01:54So one of those possibilities is alpha is not true.
- 3:01:57But if alpha is true, well, then we can draw the conclusion
- 3:02:00that beta must be true.
- 3:02:01So either alpha is not true or alpha is true, in which case beta is also true.
- 3:02:07So this is one way to turn an implication into just a statement about or.
- 3:02:12In addition to eliminating implications,
- 3:02:14we can also eliminate biconditionals as well.
- 3:02:17So let's take an English example, something like,
- 3:02:19it is raining if and only if Harry is inside.
- 3:02:23And this if and only if really sounds like that biconditional,
- 3:02:26that double arrow sign that we saw in propositional logic not too long ago.
- 3:02:31And what does this actually mean if we were to translate this?
- 3:02:33Well, this means that if it is raining, then Harry is inside.
- 3:02:37And if Harry is inside, then it is raining,
- 3:02:40that this implication goes both ways.
- 3:02:43And this is what we would call biconditional elimination,
- 3:02:45that I can take a biconditional, a if and only if b,
- 3:02:50and translate that into something like this, a implies b, and b implies a.
- 3:02:56So many of these inference rules are taking logic that uses certain symbols
- 3:03:00and turning them into different symbols, taking an implication
- 3:03:03and turning it into an or, or taking a biconditional
- 3:03:06and turning it into implication.
- 3:03:08And another example of it would be something like this.
- 3:03:11It is not true that both Harry and Ron passed the test.
- 3:03:16Well, all right, how do we translate that?
- 3:03:17What does that mean?
- 3:03:18Well, if it is not true that both of them passed the test, well,
- 3:03:22then the reasonable conclusion we might draw
- 3:03:25is that at least one of them didn't pass the test.
- 3:03:28So the conclusion is either Harry did not pass the test
- 3:03:31or Ron did not pass the test, or both.
- 3:03:33This is not an exclusive or.
- 3:03:35But if it is true that it is not true that both Harry and Ron passed the test,
- 3:03:40well, then either Harry didn't pass the test or Ron didn't pass the test.
- 3:03:45And this type of law is one of De Morgan's laws.
- 3:03:48Quite famous in logic where the idea is that we can turn an and into an or.
- 3:03:52We can say we can take this and that both Harry and Ron passed the test
- 3:03:56and turn it into an or by moving the nots around.
- 3:03:59So if it is not true that Harry and Ron passed the test,
- 3:04:03well, then either Harry did not pass the test
- 3:04:05or Ron did not pass the test either.
- 3:04:08And the way we frame that more formally using logic is to say this.
- 3:04:12If it is not true that alpha and beta, well, then either not alpha or not beta.
- 3:04:20The way I like to think about this is that if you
- 3:04:22have a negation in front of an and expression,
- 3:04:25you move the negation inwards, so to speak,
- 3:04:27moving the negation into each of these individual sentences
- 3:04:31and then flip the and into an or.
- 3:04:34So the negation moves inwards and the and flips into an or.
- 3:04:37So I go from not a and b to not a or not b.
- 3:04:43And there's actually a reverse of De Morgan's law
- 3:04:45that goes in the other direction for something like this.
- 3:04:48If I say it is not true that Harry or Ron passed the test,
- 3:04:52meaning neither of them passed the test, well, then the conclusion I can draw
- 3:04:56is that Harry did not pass the test and Ron did not pass the test.
- 3:05:01So in this case, instead of turning an and into an or,
- 3:05:04we're turning an or into an and.
- 3:05:06But the idea is the same.
- 3:05:07And this, again, is another example of De Morgan's laws.
- 3:05:10And the way that works is that if I have not a or b this time,
- 3:05:15the same logic is going to apply.
- 3:05:17I'm going to move the negation inwards.
- 3:05:19And I'm going to flip this time, flip the or into an and.
- 3:05:22So if not a or b, meaning it is not true that a or b or alpha or beta,
- 3:05:28then I can say not alpha and not beta, moving the negation inwards
- 3:05:34in order to make that conclusion.
- 3:05:36So those are De Morgan's laws and a couple other inference rules
- 3:05:38that are worth just taking a look at.
- 3:05:40One is the distributive law that works this way.
- 3:05:43So if I have alpha and beta or gamma, well, then much in the same way
- 3:05:49that you can use in math, use distributive laws to distribute
- 3:05:52operands like addition and multiplication,
- 3:05:55I can do a similar thing here, where I can say if alpha and beta or gamma,
- 3:06:01then I can say something like alpha and beta or alpha and gamma,
- 3:06:06that I've been able to distribute this and sign throughout this expression.
- 3:06:11So this is an example of the distributive property
- 3:06:13or the distributive law as applied to logic in much the same way
- 3:06:16that you would distribute a multiplication over the addition
- 3:06:19of something, for example.
- 3:06:22This works the other way too.
- 3:06:23So if, for example, I have alpha or beta and gamma,
- 3:06:27I can distribute the or throughout the expression.
- 3:06:30I can say alpha or beta and alpha or gamma.
- 3:06:34So the distributive law works in that way too.
- 3:06:36And it's helpful if I want to take an or and move it into the expression.
- 3:06:40And we'll see an example soon of why it is that we might actually
- 3:06:43care to do something like that.
- 3:06:46All right, so now we've seen a lot of different inference rules.
- 3:06:49And the question now is, how can we use those inference rules to actually try
- 3:06:53and draw some conclusions, to actually try and prove something about entailment,
- 3:06:57proving that given some initial knowledge base,
- 3:06:59we would like to find some way to prove that a query is true?
- 3:07:04Well, one way to think about it is actually
- 3:07:06to think back to what we talked about last time
- 3:07:08when we talked about search problems.
- 3:07:10Recall again that search problems have some sort of initial state.
- 3:07:13They have actions that you can take from one state to another
- 3:07:16as defined by a transition model that tells you
- 3:07:18how to get from one state to another.
- 3:07:20We talked about testing to see if you were at a goal.
- 3:07:22And then some path cost function to see how many steps
- 3:07:26did you have to take or how costly was the solution that you found.
- 3:07:31Now that we have these inference rules that
- 3:07:33take some set of sentences in propositional logic
- 3:07:36and get us some new set of sentences in propositional logic,
- 3:07:40we can actually treat those sentences or those sets of sentences
- 3:07:44as states inside of a search problem.
- 3:07:47So if we want to prove that some query is true,
- 3:07:49prove that some logical theorem is true,
- 3:07:52we can treat theorem proving as a form of a search problem.
- 3:07:55I can say that we begin in some initial state, where
- 3:07:59that initial state is the knowledge base that I begin with,
- 3:08:02the set of all of the sentences that I know to be true.
- 3:08:05What actions are available to me?
- 3:08:07Well, the actions are any of the inference rules
- 3:08:09that I can apply at any given time.
- 3:08:12The transition model just tells me after I apply the inference rule,
- 3:08:16here is the new set of all of the knowledge
- 3:08:18that I have, which will be the old set of knowledge,
- 3:08:20plus some additional inference that I've been able to draw,
- 3:08:23much as in the same way we saw what we got when we applied those inference
- 3:08:26rules and got some sort of conclusion.
- 3:08:28That conclusion gets added to our knowledge base,
- 3:08:31and our transition model will encode that.
- 3:08:34What is the goal test?
- 3:08:35Well, our goal test is checking to see if we
- 3:08:38have proved the statement we're trying to prove,
- 3:08:40if the thing we're trying to prove is inside of our knowledge base.
- 3:08:44And the path cost function, the thing we're trying to minimize,
- 3:08:47is maybe the number of inference rules that we needed to use,
- 3:08:50the number of steps, so to speak, inside of our proof.
- 3:08:54And so here we've been able to apply the same types of ideas
- 3:08:57that we saw last time with search problems
- 3:08:59to something like trying to prove something about knowledge
- 3:09:02by taking our knowledge and framing it in terms
- 3:09:05that we can understand as a search problem with an initial state,
- 3:09:08with actions, with a transition model.
- 3:09:10So this shows a couple of things, one being how versatile search problems
- 3:09:14are, that they can be the same types of algorithms
- 3:09:16that we use to solve a maze or figure out
- 3:09:19how to get from point A to point B inside of driving directions,
- 3:09:22for example, can also be used as a theorem proving
- 3:09:25method of taking some sort of starting knowledge base
- 3:09:28and trying to prove something about that knowledge.
- 3:09:31So this, yet again, is a second way, in addition to model checking,
- 3:09:35to try and prove that certain statements are true.
- 3:09:38But it turns out there's yet another way that we can try and apply inference.
- 3:09:42And we'll talk about this now, which is not the only way, but certainly one
- 3:09:45of the most common, which is known as resolution.
- 3:09:48And resolution is based on another inference rule
- 3:09:51that we'll take a look at now, quite a powerful inference rule that
- 3:09:54will let us prove anything that can be proven about a knowledge base.
- 3:09:58And it's based on this basic idea.
- 3:10:01Let's say I know that either Ron is in the Great Hall
- 3:10:05or Hermione is in the library.
- 3:10:08And let's say I also know that Ron is not in the Great Hall.
- 3:10:12Based on those two pieces of information, what can I conclude?
- 3:10:16Well, I could pretty reasonably conclude that Hermione
- 3:10:18must be in the library.
- 3:10:20How do I know that?
- 3:10:21Well, it's because these two statements, these two
- 3:10:24what we'll call complementary literals, literals that complement each other,
- 3:10:28they're opposites of each other, seem to conflict with each other.
- 3:10:32This sentence tells us that either Ron is in the Great Hall
- 3:10:35or Hermione is in the library.
- 3:10:37So if we know that Ron is not in the Great Hall,
- 3:10:40that conflicts with this one, which means Hermione must be in the library.
- 3:10:45And this we can frame as a more general rule
- 3:10:48known as the unit resolution rule, a rule that says that if we have p or q
- 3:10:54and we also know not p, well then from that we can reasonably conclude q.
- 3:11:00That if p or q are true and we know that p is not true,
- 3:11:03the only possibility is for q to then be true.
- 3:11:07And this, it turns out, is quite a powerful inference rule
- 3:11:10in terms of what it can do, in part because we can quickly
- 3:11:13start to generalize this rule.
- 3:11:14This q right here doesn't need to just be a single propositional symbol.
- 3:11:19It could be multiple, all chained together in a single clause,
- 3:11:22as we'll call it.
- 3:11:23So if I had something like p or q1 or q2 or q3, so on and so forth, up until qn,
- 3:11:29so I had n different other variables, and I have not p,
- 3:11:34well then what happens when these two complement each other
- 3:11:37is that these two clauses resolve, so to speak,
- 3:11:40to produce a new clause that is just q1 or q2 all the way up to qn.
- 3:11:46And in an or, the order of the arguments in the or doesn't actually matter.
- 3:11:49The p doesn't need to be the first thing.
- 3:11:50It could have been in the middle.
- 3:11:52But the idea here is that if I have p in one clause and not
- 3:11:56p in the other clause, well then I know that one of these remaining things
- 3:11:59must be true.
- 3:12:00I've resolved them in order to produce a new clause.
- 3:12:04But it turns out we can generalize this idea even further, in fact,
- 3:12:08and display even more power that we can have with this resolution rule.
- 3:12:12So let's take another example.
- 3:12:14Let's say, for instance, that I know the same piece of information
- 3:12:17that either Ron is in the Great Hall or Hermione is in the library.
- 3:12:21And the second piece of information I know
- 3:12:23is that Ron is not in the Great Hall or Harry is sleeping.
- 3:12:29So it's not just a single piece of information.
- 3:12:31I have two different clauses.
- 3:12:33And we'll define clauses more precisely in just a moment.
- 3:12:37What do I know here?
- 3:12:38Well again, for any propositional symbol like Ron is in the Great Hall,
- 3:12:42there are only two possibilities.
- 3:12:44Either Ron is in the Great Hall, in which case, based on resolution,
- 3:12:48we know that Harry must be sleeping, or Ron is not in the Great Hall,
- 3:12:53in which case we know based on the same rule
- 3:12:56that Hermione must be in the library.
- 3:12:59Based on those two things in combination,
- 3:13:01I can say based on these two premises that I
- 3:13:03can conclude that either Hermione is in the library or Harry is sleeping.
- 3:13:10So again, because these two conflict with each other,
- 3:13:13I know that one of these two must be true.
- 3:13:15And you can take a closer look and try and reason through that logic.
- 3:13:18Make sure you convince yourself that you believe this conclusion.
- 3:13:22Stated more generally, we can name this resolution rule
- 3:13:25by saying that if we know p or q is true,
- 3:13:28and we also know that not p or r is true,
- 3:13:33we resolve these two clauses together to get a new clause, q or r,
- 3:13:37that either q or r must be true.
- 3:13:41And again, much as in the last case, q and r
- 3:13:43don't need to just be single propositional symbols.
- 3:13:46It could be multiple symbols.
- 3:13:48So if I had a rule that had p or q1 or q2 or q3, so on and so forth,
- 3:13:52up until qn, where n is just some number.
- 3:13:55And likewise, I had not p or r1 or r2, so on and so forth, up until rm,
- 3:14:02where m, again, is just some other number.
- 3:14:05I can resolve these two clauses together to get one of these must be true,
- 3:14:09q1 or q2 up until qn or r1 or r2 up until rm.
- 3:14:14And this is just a generalization of that same rule we saw before.
- 3:14:19Each of these things here are what we're going to call a clause,
- 3:14:23where a clause is formally defined as a disjunction of literals,
- 3:14:27where a disjunction means it's a bunch of things that are connected with or.
- 3:14:31Disjunction means things connected with or.
- 3:14:34Conjunction, meanwhile, is things connected with and.
- 3:14:37And a literal is either a propositional symbol
- 3:14:40or the opposite of a propositional symbol.
- 3:14:42So it's something like p or q or not p or not q.
- 3:14:46Those are all propositional symbols or not of the propositional symbols.
- 3:14:50And we call those literals.
- 3:14:52And so a clause is just something like this, p or q or r, for example.
- 3:14:57Meanwhile, what this gives us an ability to do
- 3:15:00is it gives us an ability to turn logic, any logical sentence,
- 3:15:04into something called conjunctive normal form.
- 3:15:07A conjunctive normal form sentence is a logical sentence
- 3:15:11that is a conjunction of clauses.
- 3:15:14Recall, again, conjunction means things are connected to one another using and.
- 3:15:18And so a conjunction of clauses means it is an and of individual clauses,
- 3:15:23each of which has ors in it.
- 3:15:25So something like this, a or b or c, and d or not e, and f or g.
- 3:15:32Everything in parentheses is one clause.
- 3:15:35All of the clauses are connected to each other using an and.
- 3:15:38And everything in the clause is separated using an or.
- 3:15:43And this is just a standard form that we can translate a logical sentence
- 3:15:46into that just makes it easy to work with and easy to manipulate.
- 3:15:50And it turns out that we can take any sentence in logic
- 3:15:53and turn it into conjunctive normal form just
- 3:15:56by applying some inference rules and transformations to it.
- 3:15:59So we'll take a look at how we can actually do that.
- 3:16:03So what is the process for taking a logical formula
- 3:16:06and converting it into conjunctive normal form, otherwise known as c and f?
- 3:16:10Well, the process looks a little something like this.
- 3:16:12We need to take all of the symbols that are not
- 3:16:14part of conjunctive normal form.
- 3:16:16The bi-conditionals and the implications and so forth,
- 3:16:18and turn them into something that is more closely like conjunctive normal
- 3:16:23form.
- 3:16:24So the first step will be to eliminate bi-conditionals,
- 3:16:26those if and only if double arrows.
- 3:16:29And we know how to eliminate bi-conditionals
- 3:16:31because we saw there was an inference rule to do just that.
- 3:16:34Any time I have an expression like alpha if and only if beta,
- 3:16:38I can turn that into alpha implies beta and beta implies alpha
- 3:16:43based on that inference rule we saw before.
- 3:16:46Likewise, in addition to eliminating bi-conditionals,
- 3:16:48I can eliminate implications as well, the if then arrows.
- 3:16:52And I can do that using the same inference rule we saw before too,
- 3:16:56taking alpha implies beta and turning that into not alpha or beta
- 3:17:01because that is logically equivalent to this first thing here.
- 3:17:06Then we can move knots inwards because we don't
- 3:17:08want knots on the outsides of our expressions.
- 3:17:10Conjunctive normal form requires that it's just claws and claws
- 3:17:14and claws and claws.
- 3:17:15Any knots need to be immediately next to propositional symbols.
- 3:17:19But we can move those knots around using De Morgan's laws
- 3:17:22by taking something like not A and B and turn it into not A or not B,
- 3:17:29for example, using De Morgan's laws to manipulate that.
- 3:17:31And after that, all we'll be left with are ands and ors.
- 3:17:34And those are easy to deal with.
- 3:17:35We can use the distributive law to distribute the ors
- 3:17:39so that the ors end up on the inside of the expression, so to speak,
- 3:17:42and the ands end up on the outside.
- 3:17:45So this is the general pattern for how we'll take a formula
- 3:17:47and convert it into conjunctive normal form.
- 3:17:50And let's now take a look at an example of how we would do this
- 3:17:53and explore then why it is that we would want to do something like this.
- 3:17:57Here's how we can do it.
- 3:17:58Let's take this formula, for example.
- 3:18:00P or Q implies R. And I'd like to convert this into conjunctive normal form,
- 3:18:06where it's all ands of clauses, and every clause is a disjunctive clause.
- 3:18:10It's ors together.
- 3:18:12So what's the first thing I need to do?
- 3:18:14Well, this is an implication.
- 3:18:15So let me go ahead and remove that implication.
- 3:18:18Using the implication inference rule, I can turn P or Q into P or Q implies R
- 3:18:25into not P or Q or R. So that's the first step.
- 3:18:29I've gotten rid of the implication.
- 3:18:32And next, I can get rid of the not on the outside of this expression, too.
- 3:18:36I can move the nots inwards so they're closer to the literals themselves
- 3:18:41by using De Morgan's laws.
- 3:18:43And De Morgan's law says that not P or Q is equivalent to not P and not Q.
- 3:18:50Again, here, just applying the inference rules
- 3:18:52that we've already seen in order to translate these statements.
- 3:18:57And now, I have two things that are separated by an or,
- 3:19:00where this thing on the inside is an and.
- 3:19:03What I'd really like to move the ors so the ors are on the inside,
- 3:19:06because conjunctive normal form means I need clause and clause
- 3:19:10and clause and clause.
- 3:19:11And so to do that, I can use the distributive law.
- 3:19:14If I have not P and not Q or R, I can distribute the or R to both of these
- 3:19:21to get not P or R and not Q or R using the distributive law.
- 3:19:26And this now here at the bottom is in conjunctive normal form.
- 3:19:30It is a conjunction and and of disjunctions of clauses
- 3:19:35that just are separated by ors.
- 3:19:38So this process can be used by any formula to take a logical sentence
- 3:19:42and turn it into this conjunctive normal form, where
- 3:19:44I have clause and clause and clause and clause and clause and so on.
- 3:19:49So why is this helpful?
- 3:19:50Why do we even care about taking all these sentences
- 3:19:52and converting them into this form?
- 3:19:54It's because once they're in this form where we have these clauses,
- 3:19:58these clauses are the inputs to the resolution inference rule
- 3:20:02that we saw a moment ago, that if I have two clauses where there's
- 3:20:05something that conflicts or something complementary
- 3:20:08between those two clauses, I can resolve them
- 3:20:10to get a new clause, to draw a new conclusion.
- 3:20:13And we call this process inference by resolution,
- 3:20:16using the resolution rule to draw some sort of inference.
- 3:20:19And it's based on the same idea, that if I have P or Q, this clause,
- 3:20:23and I have not P or R, that I can resolve these two clauses together
- 3:20:28to get Q or R as the resulting clause, a new piece of information
- 3:20:32that I didn't have before.
- 3:20:35Now, a couple of key points that are worth noting about this
- 3:20:37before we talk about the actual algorithm.
- 3:20:39One thing is that, let's imagine we have P or Q or S,
- 3:20:43and I also have not P or R or S. The resolution rule
- 3:20:48says that because this P conflicts with this not P,
- 3:20:51we would resolve to put everything else together to get Q or S or R or S.
- 3:20:57But it turns out that this double S is redundant, or S here and or S there.
- 3:21:01It doesn't change the meaning of the sentence.
- 3:21:03So in resolution, when we do this resolution process,
- 3:21:06we'll usually also do a process known as factoring,
- 3:21:08where we take any duplicate variables that show up
- 3:21:11and just eliminate them.
- 3:21:12So Q or S or R or S just becomes Q or R or S. The S only needs to appear once,
- 3:21:18no need to include it multiple times.
- 3:21:22Now, one final question worth considering
- 3:21:24is what happens if I try to resolve P and not P together?
- 3:21:28If I know that P is true and I know that not P is true,
- 3:21:32well, resolution says I can merge these clauses together
- 3:21:35and look at everything else.
- 3:21:37Well, in this case, there is nothing else,
- 3:21:39so I'm left with what we might call the empty clause.
- 3:21:42I'm left with nothing.
- 3:21:43And the empty clause is always false.
- 3:21:46The empty clause is equivalent to just being false.
- 3:21:49And that's pretty reasonable because it's impossible for both P and not P
- 3:21:55to both hold at the same time.
- 3:21:57P is either true or it's not true, which
- 3:21:59means that if P is true, then this must be false.
- 3:22:02And if this is true, then this must be false.
- 3:22:05There is no way for both of these to hold at the same time.
- 3:22:07So if ever I try and resolve these two, it's a contradiction,
- 3:22:11and I'll end up getting this empty clause where the empty clause I
- 3:22:14can call equivalent to false.
- 3:22:17And this idea that if I resolve these two contradictory terms,
- 3:22:21I get the empty clause, this is the basis for our inference
- 3:22:25by resolution algorithm.
- 3:22:26Here's how we're going to perform inference by resolution
- 3:22:29at a very high level.
- 3:22:31We want to prove that our knowledge base entails some query alpha,
- 3:22:35that based on the knowledge we have, we can prove conclusively
- 3:22:39that alpha is going to be true.
- 3:22:41How are we going to do that?
- 3:22:43Well, in order to do that, we're going to try
- 3:22:45to prove that if we know the knowledge and not alpha,
- 3:22:49that that would be a contradiction.
- 3:22:51And this is a common technique in computer science
- 3:22:53more generally, this idea of proving something by contradiction.
- 3:22:57If I want to prove that something is true,
- 3:23:00I can do so by first assuming that it is false
- 3:23:04and showing that it would be contradictory,
- 3:23:06showing that it leads to some contradiction.
- 3:23:08And if the thing I'm trying to prove, if when I assume it's false,
- 3:23:11leads to a contradiction, then it must be true.
- 3:23:14And that's the logical approach or the idea behind a proof by contradiction.
- 3:23:18And that's what we're going to do here.
- 3:23:20We want to prove that this query alpha is true.
- 3:23:23So we're going to assume that it's not true.
- 3:23:26We're going to assume not alpha.
- 3:23:28And we're going to try and prove that it's a contradiction.
- 3:23:30If we do get a contradiction, well, then we
- 3:23:32know that our knowledge entails the query alpha.
- 3:23:36If we don't get a contradiction, there is no entailment.
- 3:23:39This is this idea of a proof by contradiction
- 3:23:41of assuming the opposite of what you're trying to prove.
- 3:23:44And if you can demonstrate that that's a contradiction,
- 3:23:46then what you're proving must be true.
- 3:23:49But more formally, how do we actually do this?
- 3:23:51How do we check that knowledge base and not alpha
- 3:23:56is going to lead to a contradiction?
- 3:23:58Well, here is where resolution comes into play.
- 3:24:01To determine if our knowledge base entails some query alpha,
- 3:24:05we're going to convert knowledge base and not alpha
- 3:24:08to conjunctive normal form, that form where
- 3:24:10we have a whole bunch of clauses that are all anded together.
- 3:24:14And when we have these individual clauses,
- 3:24:16now we can keep checking to see if we can use resolution
- 3:24:21to produce a new clause.
- 3:24:23We can take any pair of clauses and check,
- 3:24:26is there some literal that is the opposite of each other
- 3:24:29or complementary to each other in both of them?
- 3:24:32For example, I have a p in one clause and a not p in another clause.
- 3:24:35Or an r in one clause and a not r in another clause.
- 3:24:39If ever I have that situation where once I
- 3:24:41convert to conjunctive normal form and I have a whole bunch of clauses,
- 3:24:44I see two clauses that I can resolve to produce a new clause, then I'll do so.
- 3:24:49This process occurs in a loop.
- 3:24:50I'm going to keep checking to see if I can use resolution
- 3:24:53to produce a new clause and keep using those new clauses
- 3:24:56to try to generate more new clauses after that.
- 3:25:00Now, it just so may happen that eventually we
- 3:25:03may produce the empty clause, the clause we were talking about before.
- 3:25:06If I resolve p and not p together, that produces the empty clause
- 3:25:11and the empty clause we know to be false.
- 3:25:14Because we know that there's no way for both p and not p
- 3:25:18to both simultaneously be true.
- 3:25:21So if ever we produce the empty clause, then we have a contradiction.
- 3:25:25And if we have a contradiction, that's exactly what we were trying
- 3:25:27to do in a fruit by contradiction.
- 3:25:29If we have a contradiction, then we know that our knowledge base
- 3:25:32must entail this query alpha.
- 3:25:34And we know that alpha must be true.
- 3:25:37And it turns out, and we won't go into the proof here,
- 3:25:39but you can show that otherwise, if you don't produce the empty clause,
- 3:25:43then there is no entailment.
- 3:25:45If we run into a situation where there are no more new clauses to add,
- 3:25:48we've done all the resolution that we can do,
- 3:25:50and yet we still haven't produced the empty clause,
- 3:25:53then there is no entailment in this case.
- 3:25:56And this now is the resolution algorithm.
- 3:25:58And it's very abstract looking, especially this idea of like,
- 3:26:01what does it even mean to have the empty clause?
- 3:26:03So let's take a look at an example, actually
- 3:26:05try and prove some entailment by using this inference by resolution process.
- 3:26:11So here's our question.
- 3:26:12We have this knowledge base.
- 3:26:14Here is the knowledge that we know, A or B, and not B or C, and not C.
- 3:26:21And we want to know if all of this entails A.
- 3:26:25So this is our knowledge base here, this whole log thing.
- 3:26:28And our query alpha is just this propositional symbol, A.
- 3:26:33So what do we do?
- 3:26:34Well, first, we want to prove by contradiction.
- 3:26:36So we want to first assume that A is false,
- 3:26:39and see if that leads to some sort of contradiction.
- 3:26:42So here is what we're going to start with, A or B, and not B or C, and not C.
- 3:26:46This is our knowledge base.
- 3:26:48And we're going to assume not A. We're going
- 3:26:51to assume that the thing we're trying to prove is, in fact, false.
- 3:26:56And so this is now in conjunctive normal form,
- 3:26:59and I have four different clauses.
- 3:27:01I have A or B. I have not B or C. I have not C, and I have not A.
- 3:27:08And now, I can begin to just pick two clauses that I can resolve,
- 3:27:12and apply the resolution rule to them.
- 3:27:15And so looking at these four clauses, I see, all right, these two clauses
- 3:27:20are ones I can resolve.
- 3:27:21I can resolve them because there are complementary literals
- 3:27:25that show up in them.
- 3:27:26There's a C here, and a not C here.
- 3:27:28So just looking at these two clauses, if I know that not B or C is true,
- 3:27:34and I know that C is not true, well, then I
- 3:27:36can resolve these two clauses to say, all right, not B, that must be true.
- 3:27:41I can generate this new clause as a new piece of information
- 3:27:45that I now know to be true.
- 3:27:47And all right, now I can repeat this process, do the process again.
- 3:27:50Can I use resolution again to get some new conclusion?
- 3:27:54Well, it turns out I can.
- 3:27:55I can use that new clause I just generated, along with this one here.
- 3:27:58There are complementary literals.
- 3:28:00This B is complementary to, or conflicts with, this not B over here.
- 3:28:06And so if I know that A or B is true, and I know that B is not true,
- 3:28:12well, then the only remaining possibility is that A must be true.
- 3:28:15So now we have A. That is a new clause that I've been able to generate.
- 3:28:19And now, I can do this one more time.
- 3:28:21I'm looking for two clauses that can be resolved,
- 3:28:23and you might programmatically do this by just looping
- 3:28:25over all possible pairs of clauses and checking
- 3:28:28for complementary literals in each.
- 3:28:30And here, I can say, all right, I found two clauses, not A and A,
- 3:28:34that conflict with each other.
- 3:28:36And when I resolve these two together, well,
- 3:28:38this is the same as when we were resolving P and not P from before.
- 3:28:42When I resolve these two clauses together, I get rid of the As,
- 3:28:45and I'm left with the empty clause.
- 3:28:48And the empty clause we know to be false, which means we have a contradiction,
- 3:28:51which means we can safely say that this whole knowledge base does entail A.
- 3:28:56That if this sentence is true, that we know that A for sure is also true.
- 3:29:02So this now, using inference by resolution,
- 3:29:04is an entirely different way to take some statement
- 3:29:07and try and prove that it is, in fact, true.
- 3:29:10Instead of enumerating all of the possible worlds
- 3:29:12that we might be in in order to try to figure out in which cases
- 3:29:15is the knowledge base true and in which cases are query true,
- 3:29:18instead we use this resolution algorithm to say,
- 3:29:22let's keep trying to figure out what conclusions we can draw
- 3:29:25and see if we reach a contradiction.
- 3:29:27And if we reach a contradiction, then that
- 3:29:28tells us something about whether our knowledge actually
- 3:29:31entails the query or not.
- 3:29:33And it turns out there are many different algorithms that
- 3:29:35can be used for inference.
- 3:29:37What we've just looked at here are just a couple of them.
- 3:29:39And in fact, all of this is just based on one particular type of logic.
- 3:29:44It's based on propositional logic, where we have these individual symbols
- 3:29:47and we connect them using and and or and not and implies and by conditionals.
- 3:29:52But propositional logic is not the only kind of logic that exists.
- 3:29:56And in fact, we see that there are limitations
- 3:29:58that exist in propositional logic, especially
- 3:30:01as we saw in examples like with the mastermind example
- 3:30:06or with the example with the logic puzzle where
- 3:30:08we had different Hogwarts house people that belong to different houses
- 3:30:12and we were trying to figure out who belonged to which houses.
- 3:30:15There were a lot of different propositional symbols that we needed
- 3:30:18in order to represent some fairly basic ideas.
- 3:30:21So now is the final topic that we'll take a look at just before we end class
- 3:30:24today is one final type of logic different from propositional logic
- 3:30:28known as first order logic, which is a little bit more powerful than
- 3:30:32propositional logic and is going to make it easier for us
- 3:30:34to express certain types of ideas.
- 3:30:37In propositional logic, if we think back to that puzzle
- 3:30:39with the people in the Hogwarts houses, we had a whole bunch of symbols.
- 3:30:43And every symbol could only be true or false.
- 3:30:46We had a symbol for Minerva Gryffindor, which was either true of Minerva
- 3:30:49within Gryffindor and false otherwise, and likewise
- 3:30:51for Minerva Hufflepuff and Minerva Ravenclaw and Minerva Slytherin
- 3:30:55and so forth.
- 3:30:56But this was starting to get quite redundant.
- 3:30:58We wanted some way to be able to express that there
- 3:31:01is a relationship between these propositional symbols,
- 3:31:03that Minerva shows up in all of them.
- 3:31:05And also, I would have liked to have not have had so many different symbols
- 3:31:09to represent what really was a fairly straightforward problem.
- 3:31:13So first order logic will give us a different way
- 3:31:15of trying to deal with this idea by giving us two different types of symbols.
- 3:31:19We're going to have constant symbols that are going to represent objects
- 3:31:23like people or houses.
- 3:31:24And then predicate symbols, which you can think of as relations or functions
- 3:31:29that take an input and evaluate them to true or false, for example,
- 3:31:33that tell us whether or not some property of some constant
- 3:31:37or some pair of constants or multiple constants actually holds.
- 3:31:41So we'll see an example of that in just a moment.
- 3:31:43For now, in this same problem, our constant symbols
- 3:31:46might be objects, things like people or houses.
- 3:31:49So Minerva, Pomona, Horace, Gilderoy, those are all constant symbols,
- 3:31:53as are my four houses, Gryffindor, Hufflepuff, Ravenclaw, and Slytherin.
- 3:31:58Predicates, meanwhile, these predicate symbols
- 3:32:00are going to be properties that might hold true or false
- 3:32:03of these individual constants.
- 3:32:06So person might hold true of Minerva, but it
- 3:32:09would be false for Gryffindor because Gryffindor is not a person.
- 3:32:12And house is going to hold true for Ravenclaw,
- 3:32:15but it's not going to hold true for Horace, for example,
- 3:32:17because Horace is a person.
- 3:32:19And belongs to, meanwhile, is going to be some relation that
- 3:32:23is going to relate people to their houses.
- 3:32:26And it's going to only tell me when someone belongs to a house or does not.
- 3:32:30So let's take a look at some examples of what a sentence in first order logic
- 3:32:35might actually look like.
- 3:32:36A sentence might look like something like this.
- 3:32:38Person Minerva, with Minerva in parentheses, and person being a predicate
- 3:32:42symbol, Minerva being a constant symbol.
- 3:32:45This sentence in first order logic effectively
- 3:32:48means Minerva is a person, or the person property applies to the Minerva object.
- 3:32:54So if I want to say something like Minerva is a person,
- 3:32:56here is how I express that idea using first order logic.
- 3:33:00Meanwhile, I can say something like, house Gryffindor,
- 3:33:03to likewise express the idea that Gryffindor is a house.
- 3:33:07I can do that this way.
- 3:33:08And all of the same logical connectives that we
- 3:33:10saw in propositional logic, those are going to work here too.
- 3:33:13And or implication by conditional not.
- 3:33:16In fact, I can use not to say something like, not house Minerva.
- 3:33:20And this sentence in first order logic means something like,
- 3:33:24Minerva is not a house.
- 3:33:26It is not true that the house property applies to Minerva.
- 3:33:31Meanwhile, in addition to some of these predicate symbols
- 3:33:34that just take a single argument, some of our predicate symbols
- 3:33:36are going to express binary relations, relations
- 3:33:39between two of its arguments.
- 3:33:42So I could say something like, belongs to, and then two inputs, Minerva
- 3:33:46and Gryffindor, to express the idea that Minerva belongs to Gryffindor.
- 3:33:51And so now here's the key difference, or one of the key differences,
- 3:33:54between this and propositional logic.
- 3:33:56In propositional logic, I needed one symbol for Minerva Gryffindor,
- 3:34:00and one symbol for Minerva Hufflepuff, and one
- 3:34:02symbol for all the other people's Gryffindor and Hufflepuff variables.
- 3:34:06In this case, I just need one symbol for each of my people,
- 3:34:10and one symbol for each of my houses.
- 3:34:13And then I can express as a predicate something like, belongs to,
- 3:34:16and say, belongs to Minerva Gryffindor, to express the idea that Minerva
- 3:34:21belongs to Gryffindor House.
- 3:34:23So already we can see that first order logic is quite expressive in being
- 3:34:27able to express these sorts of sentences using the existing constant symbols
- 3:34:32and predicates that already exist, while minimizing the number of new symbols
- 3:34:36that I need to create.
- 3:34:37I can just use eight symbols for people for houses,
- 3:34:40instead of 16 symbols for every possible combination of each.
- 3:34:46But first order logic gives us a couple of additional features
- 3:34:49that we can use to express even more complex ideas.
- 3:34:52And these more additional features are generally known as quantifiers.
- 3:34:56And there are two main quantifiers in first order logic,
- 3:34:58the first of which is universal quantification.
- 3:35:01Universal quantification lets me express an idea
- 3:35:04like something is going to be true for all values of a variable.
- 3:35:09Like for all values of x, some statement is going to hold true.
- 3:35:13So what might a sentence in universal quantification look like?
- 3:35:16Well, we're going to use this upside down a to mean for all.
- 3:35:21So upside down ax means for all values of x, where x is any object,
- 3:35:26this is going to hold true.
- 3:35:28Belongs to x Gryffindor implies not belongs to x Hufflepuff.
- 3:35:36So let's try and parse this out.
- 3:35:38This means that for all values of x, if this holds true,
- 3:35:42if x belongs to Gryffindor, then this does not hold true.
- 3:35:46x does not belong to Hufflepuff.
- 3:35:50So translated into English, this sentence
- 3:35:52is saying something like for all objects x, if x belongs to Gryffindor,
- 3:35:57then x does not belong to Hufflepuff, for example.
- 3:36:00Or a phrase even more simply, anyone in Gryffindor
- 3:36:03is not in Hufflepuff, simplified way of saying the same thing.
- 3:36:07So this universal quantification lets us express
- 3:36:10an idea like something is going to hold true for all values
- 3:36:14of a particular variable.
- 3:36:16In addition to universal quantification though,
- 3:36:18we also have existential quantification.
- 3:36:21Whereas universal quantification said that something
- 3:36:24is going to be true for all values of a variable,
- 3:36:27existential quantification says that some expression is going
- 3:36:30to be true for some value of a variable, at least one value of the variable.
- 3:36:36So let's take a look at a sample sentence using existential quantification.
- 3:36:40One such sentence looks like this.
- 3:36:42There exists an x.
- 3:36:43This backwards e stands for exists.
- 3:36:46And here we're saying there exists an x such that house x and belongs
- 3:36:51to Minerva x.
- 3:36:53In other words, there exists some object x where x is a house
- 3:36:57and Minerva belongs to x.
- 3:37:00Or phrased a little more succinctly in English,
- 3:37:02I'm here just saying Minerva belongs to a house.
- 3:37:05There's some object that is a house and Minerva belongs to a house.
- 3:37:10And combining this universal and existential quantification,
- 3:37:13we can create far more sophisticated logical statements
- 3:37:16than we were able to just using propositional logic.
- 3:37:19I could combine these to say something like this.
- 3:37:21For all x, person x implies there exists
- 3:37:26a y such that house y and belongs to xy.
- 3:37:30All right.
- 3:37:31So a lot of stuff going on there, a lot of symbols.
- 3:37:33Let's try and parse it out and just understand what it's saying.
- 3:37:36Here we're saying that for all values of x, if x is a person,
- 3:37:41then this is true.
- 3:37:43So in other words, I'm saying for all people,
- 3:37:45and we call that person x, this statement is going to be true.
- 3:37:48What statement is true of all people?
- 3:37:50Well, there exists a y that is a house, so there exists some house,
- 3:37:55and x belongs to y.
- 3:37:58In other words, I'm saying that for all people out there,
- 3:38:01there exists some house such that x, the person, belongs to y, the house.
- 3:38:07This is phrased more succinctly.
- 3:38:08I'm saying that every person belongs to a house, that for all x,
- 3:38:12if x is a person, then there exists a house that x belongs to.
- 3:38:17And so we can now express a lot more powerful ideas using this idea now
- 3:38:20of first order logic.
- 3:38:21And it turns out there are many other kinds of logic out there.
- 3:38:24There's second order logic and other higher order logic,
- 3:38:27each of which allows us to express more and more complex ideas.
- 3:38:30But all of it, in this case, is really in pursuit
- 3:38:33of the same goal, which is the representation of knowledge.
- 3:38:36We want our AI agents to be able to know information,
- 3:38:39to represent that information, whether that's
- 3:38:41using propositional logic or first order logic or some other logic,
- 3:38:45and then be able to reason based on that, to be able to draw conclusions,
- 3:38:49make inferences, figure out whether there's
- 3:38:50some sort of entailment relationship, as by using some sort of inference
- 3:38:54algorithm, something like inference by resolution or model checking
- 3:38:58or any number of these other algorithms that we can use in order
- 3:39:01to take information that we know and translate it to additional conclusions.
- 3:39:06So all of this has helped us to create AI that
- 3:39:08is able to represent information about what it knows and what it doesn't know.
- 3:39:13Next time, though, we'll take a look at how we can make our AI even more
- 3:39:16powerful by not just encoding information that we know for sure to be true
- 3:39:20and not to be true, but also to take a look at uncertainty,
- 3:39:23to look at what happens if AI thinks that something might be probable
- 3:39:27or maybe not very probable or somewhere in between those two extremes,
- 3:39:31all in the pursuit of trying to build our intelligent systems
- 3:39:34to be even more intelligent.
- 3:39:36We'll see you next time.
- 3:39:39Thank you.
- 3:39:57All right, welcome back, everyone, to an introduction
- 3:39:59to artificial intelligence with Python.
- 3:40:02And last time, we took a look at how it is that AI inside of our computers
- 3:40:05can represent knowledge.
- 3:40:07We represented that knowledge in the form of logical sentences
- 3:40:10in a variety of different logical languages.
- 3:40:12And the idea was we wanted our AI to be able to represent knowledge
- 3:40:15or information and somehow use those pieces of information
- 3:40:19to be able to derive new pieces of information by inference,
- 3:40:22to be able to take some information and deduce
- 3:40:24some additional conclusions based on the information
- 3:40:27that it already knew for sure.
- 3:40:29But in reality, when we think about computers and we think about AI,
- 3:40:32very rarely are our machines going to be able to know things for sure.
- 3:40:35Oftentimes, there's going to be some amount of uncertainty
- 3:40:38in the information that our AIs or our computers are dealing with,
- 3:40:41where it might believe something with some probability,
- 3:40:44as we'll soon discuss what probability is all about and what it means,
- 3:40:46but not entirely for certain.
- 3:40:48And we want to use the information that it has some knowledge about,
- 3:40:51even if it doesn't have perfect knowledge,
- 3:40:53to still be able to make inferences, still be able to draw conclusions.
- 3:40:57So you might imagine, for example, in the context of a robot that
- 3:41:00has some sensors and is exploring some environment,
- 3:41:02it might not know exactly where it is or exactly what's around it,
- 3:41:06but it does have access to some data that can allow it
- 3:41:08to draw inferences with some probability.
- 3:41:10There's some likelihood that one thing is true or another.
- 3:41:13Or you can imagine in context where there is a little bit more
- 3:41:15randomness and uncertainty, something like predicting the weather,
- 3:41:18where you might not be able to know for sure what tomorrow's weather is
- 3:41:21with 100% certainty, but you can probably infer with some probability
- 3:41:26what tomorrow's weather is going to be based on maybe today's weather
- 3:41:29and yesterday's weather and other data that you might have access
- 3:41:32to as well.
- 3:41:33And so oftentimes, we can distill this in terms of just possible events
- 3:41:36that might happen and what the likelihood of those events are.
- 3:41:39This comes a lot in games, for example, where there is an element of chance
- 3:41:43inside of those games.
- 3:41:44So you imagine rolling a dice.
- 3:41:45You're not sure exactly what the die roll is going to be,
- 3:41:48but you know it's going to be one of these possibilities from 1 to 6,
- 3:41:52for example.
- 3:41:53And so here now, we introduce the idea of probability theory.
- 3:41:56And what we'll take a look at today is beginning
- 3:41:58by looking at the mathematical foundations of probability theory,
- 3:42:01getting an understanding for some of the key concepts within probability,
- 3:42:05and then diving into how we can use probability and the ideas
- 3:42:08that we look at mathematically to represent some ideas in terms of models
- 3:42:12that we can put into our computers in order to program an AI that
- 3:42:15is able to use information about probability to draw inferences,
- 3:42:19to make some judgments about the world with some probability
- 3:42:22or likelihood of being true.
- 3:42:25So probability ultimately boils down to this idea
- 3:42:27that there are possible worlds that we're here representing
- 3:42:30using this little Greek letter omega.
- 3:42:32And the idea of a possible world is that when I roll a die,
- 3:42:36there are six possible worlds that could result from it.
- 3:42:38I could roll a 1, or a 2, or a 3, or a 4, or a 5, or a 6.
- 3:42:42And each of those are a possible world.
- 3:42:45And each of those possible worlds has some probability of being true,
- 3:42:49the probability that I do roll a 1, or a 2, or a 3, or something else.
- 3:42:53And we represent that probability like this, using the capital letter P.
- 3:42:57And then in parentheses, what it is that we want the probability of.
- 3:43:00So this right here would be the probability of some possible world
- 3:43:04as represented by the little letter omega.
- 3:43:07Now, there are a couple of basic axioms of probability
- 3:43:09that become relevant as we consider how we deal with probability
- 3:43:13and how we think about it.
- 3:43:14First and foremost, every probability value
- 3:43:16must range between 0 and 1 inclusive.
- 3:43:20So the smallest value any probability can have is the number 0,
- 3:43:23which is an impossible event.
- 3:43:25Something like I roll a die, and the die is a 7 is the roll that I get.
- 3:43:28If the die only has numbers 1 through 6, the event that I roll a 7
- 3:43:33is impossible, so it would have probability 0.
- 3:43:36And on the other end of the spectrum, probability
- 3:43:38can range all the way up to the positive number 1,
- 3:43:40meaning an event is certain to happen, that I roll a die
- 3:43:43and the number is less than 10, for example.
- 3:43:46That is an event that is guaranteed to happen if the only sides on my die
- 3:43:49are 1 through 6, for instance.
- 3:43:51And then they can range through any real number in between these two values.
- 3:43:55Where, generally speaking, a higher value for the probability
- 3:43:58means an event is more likely to take place,
- 3:44:00and a lower value for the probability means the event is less
- 3:44:03likely to take place.
- 3:44:05And the other key rule for probability looks a little bit like this.
- 3:44:08This sigma notation, if you haven't seen it before,
- 3:44:11refers to summation, the idea that we're going
- 3:44:13to be adding up a whole sequence of values.
- 3:44:16And this sigma notation is going to come up a couple of times today,
- 3:44:19because as we deal with probability, oftentimes we're
- 3:44:21adding up a whole bunch of individual values or individual probabilities
- 3:44:25to get some other value.
- 3:44:26So we'll see this come up a couple of times.
- 3:44:28But what this notation means is that if I sum up
- 3:44:31all of the possible worlds omega that are in big omega, which
- 3:44:35represents the set of all the possible worlds,
- 3:44:38meaning I take for all of the worlds in the set of possible worlds
- 3:44:42and add up all of their probabilities, what I ultimately get is the number 1.
- 3:44:47So if I take all the possible worlds, add up
- 3:44:48what each of their probabilities is, I should get the number 1 at the end,
- 3:44:52meaning all probabilities just need to sum to 1.
- 3:44:55So for example, if I take dice, for example,
- 3:44:57and if you imagine I have a fair die with numbers 1 through 6
- 3:45:00and I roll the die, each one of these rolls
- 3:45:02has an equal probability of taking place.
- 3:45:04And the probability is 1 over 6, for example.
- 3:45:07So each of these probabilities is between 0 and 1, 0 meaning impossible
- 3:45:12and 1 meaning for certain.
- 3:45:13And if you add up all of these probabilities
- 3:45:15for all of the possible worlds, you get the number 1.
- 3:45:18And we can represent any one of those probabilities like this.
- 3:45:22The probability that we roll the number 2, for example,
- 3:45:25is just 1 over 6.
- 3:45:27Every six times we roll the die, we'd expect that one time, for instance,
- 3:45:31the die might come up as a 2.
- 3:45:33Its probability is not certain, but it's a little more than nothing,
- 3:45:36for instance.
- 3:45:38And so this is all fairly straightforward for just a single die.
- 3:45:40But things get more interesting as our models of the world
- 3:45:43get a little bit more complex.
- 3:45:44Let's imagine now that we're not just dealing with a single die,
- 3:45:47but we have two dice, for example.
- 3:45:49I have a red die here and a blue die there,
- 3:45:51and I care not just about what the individual roll is,
- 3:45:54but I care about the sum of the two rolls.
- 3:45:56In this case, the sum of the two rolls is the number 3.
- 3:46:00How do I begin to now reason about what does the probability look like
- 3:46:04if instead of having one die, I now have two dice?
- 3:46:07Well, what we might imagine is that we could first consider
- 3:46:09what are all of the possible worlds.
- 3:46:12And in this case, all of the possible worlds
- 3:46:14are just every combination of the red and blue die that I could come up with.
- 3:46:18For the red die, it could be a 1 or a 2 or a 3 or a 4 or a 5 or a 6.
- 3:46:22And for each of those possibilities, the blue die, likewise,
- 3:46:25could also be either 1 or 2 or 3 or 4 or 5 or 6.
- 3:46:30And it just so happens that in this particular case,
- 3:46:33each of these possible combinations is equally likely.
- 3:46:36Equally likely are all of these various different possible worlds.
- 3:46:39That's not always going to be the case.
- 3:46:41If you imagine more complex models that we could try to build and things
- 3:46:44that we could try to represent in the real world,
- 3:46:46it's probably not going to be the case that every single possible world is
- 3:46:49always equally likely.
- 3:46:50But in the case of fair dice, where in any given die roll,
- 3:46:53any one number has just as good a chance of coming up as any other number,
- 3:46:57we can consider all of these possible worlds to be equally likely.
- 3:47:01But even though all of the possible worlds are equally likely,
- 3:47:04that doesn't necessarily mean that their sums are equally likely.
- 3:47:07So if we consider what the sum is of all of these two, so 1 plus 1,
- 3:47:10that's a 2.
- 3:47:112 plus 1 is a 3.
- 3:47:12And consider for each of these possible pairs of numbers
- 3:47:15what their sum ultimately is, we can notice that there are some patterns
- 3:47:18here, where it's not entirely the case that every number comes up
- 3:47:22equally likely.
- 3:47:23If you consider 7, for example, what's the probability that when I roll two
- 3:47:26dice, their sum is 7?
- 3:47:28There are several ways this can happen.
- 3:47:30There are six possible worlds where the sum is 7.
- 3:47:33It could be a 1 and a 6, or a 2 and a 5, or a 3 and a 4, a 4 and a 3,
- 3:47:37and so forth.
- 3:47:39But if you instead consider what's the probability that I roll two dice,
- 3:47:42and the sum of those two die rolls is 12, for example,
- 3:47:45we're looking at this diagram, there's only one possible world in which that
- 3:47:49can happen.
- 3:47:50And that's the possible world where both the red die and the blue die
- 3:47:54both come up as sixes to give us a sum total of 12.
- 3:47:58So based on just taking a look at this diagram,
- 3:48:00we see that some of these probabilities are likely different.
- 3:48:03The probability that the sum is a 7 must be greater than the probability
- 3:48:07that the sum is a 12.
- 3:48:08And we can represent that even more formally by saying, OK, the probability
- 3:48:11that we sum to 12 is 1 out of 36.
- 3:48:15Out of the 36 equally likely possible worlds,
- 3:48:186 squared because we have six options for the red die and six
- 3:48:22options for the blue die, out of those 36 options,
- 3:48:24only one of them sums to 12.
- 3:48:27Whereas on the other hand, the probability
- 3:48:29that if we take two dice rolls and they sum up to the number 7, well,
- 3:48:33out of those 36 possible worlds, there were six worlds where the sum was 7.
- 3:48:37And so we get 6 over 36, which we can simplify as a fraction to just 1
- 3:48:42over 6.
- 3:48:43So here now, we're able to represent these different ideas
- 3:48:46of probability, representing some events that might be more likely
- 3:48:49and then other events that are less likely as well.
- 3:48:52And these sorts of judgments, where we're figuring out just in the abstract
- 3:48:55what is the probability that this thing takes place,
- 3:48:58are generally known as unconditional probabilities.
- 3:49:01Some degree of belief we have in some proposition,
- 3:49:04some fact about the world, in the absence of any other evidence.
- 3:49:07Without knowing any additional information, if I roll a die,
- 3:49:10what's the chance it comes up as a 2?
- 3:49:12Or if I roll two dice, what's the chance that the sum of those two die
- 3:49:15rolls is a 7?
- 3:49:17But usually when we're thinking about probability, especially when
- 3:49:20we're thinking about training in AI to intelligently
- 3:49:22be able to know something about the world
- 3:49:24and make predictions based on that information,
- 3:49:26it's not unconditional probability that our AI is dealing with,
- 3:49:30but rather conditional probability, probability
- 3:49:32where rather than having no original knowledge,
- 3:49:35we have some initial knowledge about the world
- 3:49:37and how the world actually works.
- 3:49:39So conditional probability is the degree of belief in a proposition
- 3:49:43given some evidence that has already been revealed to us.
- 3:49:47So what does this look like?
- 3:49:49Well, it looks like this in terms of notation.
- 3:49:51We're going to represent conditional probability as probability of A
- 3:49:56and then this vertical bar and then B. And the way to read this
- 3:49:59is the thing on the left-hand side of the vertical bar
- 3:50:02is what we want the probability of.
- 3:50:05Here now, I want the probability that A is true,
- 3:50:08that it is the real world, that it is the event that actually does take place.
- 3:50:12And then on the right side of the vertical bar is our evidence,
- 3:50:14the information that we already know for certain about the world.
- 3:50:18For example, that B is true.
- 3:50:21So the way to read this entire expression
- 3:50:23is what is the probability of A given B, the probability that A is true,
- 3:50:28given that we already know that B is true.
- 3:50:31And this type of judgment, conditional probability,
- 3:50:34the probability of one thing given some other fact,
- 3:50:37comes up quite a lot when we think about the types of calculations
- 3:50:40we might want our AI to be able to do.
- 3:50:42For example, we might care about the probability of rain today
- 3:50:45given that we know that it rained yesterday.
- 3:50:47We could think about the probability of rain today just in the abstract.
- 3:50:51What is the chance that today it rains?
- 3:50:52But usually, we have some additional evidence.
- 3:50:54I know for certain that it rained yesterday.
- 3:50:57And so I would like to calculate the probability that it rains today
- 3:51:00given that I know that it rained yesterday.
- 3:51:03Or you might imagine that I want to know the probability that my optimal
- 3:51:06route to my destination changes given the current traffic condition.
- 3:51:09So whether or not traffic conditions change,
- 3:51:12that might change the probability that this route is actually the optimal route.
- 3:51:16Or you might imagine in a medical context,
- 3:51:18I want to know the probability that a patient has a particular disease given
- 3:51:22some results of some tests that have been performed on that patient.
- 3:51:25And I have some evidence, the results of that test,
- 3:51:28and I would like to know the probability that a patient has
- 3:51:31a particular disease.
- 3:51:33So this notion of conditional probability comes up everywhere.
- 3:51:35So we begin to think about what we would like to reason about,
- 3:51:38but being able to reason a little more intelligently
- 3:51:40by taking into account evidence that we already have.
- 3:51:43We're more able to get an accurate result for what is the likelihood
- 3:51:46that someone has this disease if we know this evidence, the results of the test,
- 3:51:50as opposed to if we were just calculating the unconditional probability of saying,
- 3:51:55what is the probability they have the disease without any evidence
- 3:51:58to try and back up our result one way or the other.
- 3:52:03So now that we've got this idea of what conditional probability is,
- 3:52:06the next question we have to ask is, all right,
- 3:52:08how do we calculate conditional probability?
- 3:52:10How do we figure out mathematically, if I have an expression like this,
- 3:52:13how do I get a number from that?
- 3:52:15What does conditional probability actually mean?
- 3:52:17Well, the formula for conditional probability
- 3:52:19looks a little something like this.
- 3:52:21The probability of a given b, the probability that a is true,
- 3:52:25given that we know that b is true, is equal to this fraction,
- 3:52:29the probability that a and b are true, divided by just the probability
- 3:52:34that b is true.
- 3:52:35And the way to intuitively try to think about this
- 3:52:37is that if I want to know the probability that a is true, given
- 3:52:40that b is true, well, I want to consider all the ways they could both be true out
- 3:52:46of the only worlds that I care about are the worlds where b is already true.
- 3:52:50I can sort of ignore all the cases where b isn't true,
- 3:52:52because those aren't relevant to my ultimate computation.
- 3:52:55They're not relevant to what it is that I want to get information about.
- 3:52:59So let's take a look at an example.
- 3:53:01Let's go back to that example of rolling two dice and the idea
- 3:53:04that those two dice might sum up to the number 12.
- 3:53:06We discussed earlier that the unconditional probability
- 3:53:09that if I roll two dice and they sum to 12 is 1 out of 36,
- 3:53:13because out of the 36 possible worlds that I might care about,
- 3:53:16in only one of them is the sum of those two dice 12.
- 3:53:19It's only when red is 6 and blue is also 6.
- 3:53:22But let's say now that I have some additional information.
- 3:53:25I now want to know what is the probability that the two dice sum to 12,
- 3:53:29given that I know that the red die was a 6.
- 3:53:33So I already have some evidence.
- 3:53:35I already know the red die is a 6.
- 3:53:36I don't know what the blue die is.
- 3:53:38That information isn't given to me in this expression.
- 3:53:41But given the fact that I know that the red die rolled a 6,
- 3:53:44what is the probability that we sum to 12?
- 3:53:47And so we can begin to do the math using that expression from before.
- 3:53:50Here, again, are all of the possibilities,
- 3:53:52all of the possible combinations of red die being 1 through 6
- 3:53:55and blue die being 1 through 6.
- 3:53:58And I might consider first, all right, what
- 3:54:00is the probability of my evidence, my B variable, where I want to know,
- 3:54:04what is the probability that the red die is a 6?
- 3:54:07Well, the probability that the red die is a 6 is just 1 out of 6.
- 3:54:11So these 1 out of 6 options are really the only worlds
- 3:54:14that I care about here now.
- 3:54:16All the rest of them are irrelevant to my calculation,
- 3:54:19because I already have this evidence that the red die was a 6,
- 3:54:22so I don't need to care about all of the other possibilities that could result.
- 3:54:26So now, in addition to the fact that the red die rolled as a 6
- 3:54:29and the probability of that, the other piece of information
- 3:54:32I need to know in order to calculate this conditional probability
- 3:54:35is the probability that both of my variables, A and B, are true.
- 3:54:39The probability that both the red die is a 6, and they all sum to 12.
- 3:54:44So what is the probability that both of these things happen?
- 3:54:47Well, it only happens in one possible case in 1 out of these 36 cases,
- 3:54:51and it's the case where both the red and the blue die are equal to 6.
- 3:54:55This is a piece of information that we already knew.
- 3:54:57And so this probability is equal to 1 over 36.
- 3:55:01And so to get the conditional probability that the sum is 12,
- 3:55:05given that I know that the red dice is equal to 6,
- 3:55:08well, I just divide these two values together,
- 3:55:10and 1 over 36 divided by 1 over 6 gives us this probability of 1 over 6.
- 3:55:16Given that I know that the red die rolled a value of 6,
- 3:55:19the probability that the sum of the two dice is 12 is also 1 over 6.
- 3:55:25And that probably makes intuitive sense to you, too,
- 3:55:27because if the red die is a 6, the only way for me to get to a 12
- 3:55:30is if the blue die also rolls a 6, and we
- 3:55:33know that the probability of the blue die rolling a 6 is 1 over 6.
- 3:55:37So in this case, the conditional probability seems fairly straightforward.
- 3:55:40But this idea of calculating a conditional probability
- 3:55:44by looking at the probability that both of these events take place
- 3:55:47is an idea that's going to come up again and again.
- 3:55:49This is the definition now of conditional probability.
- 3:55:52And we're going to use that definition as we
- 3:55:54think about probability more generally to be
- 3:55:56able to draw conclusions about the world.
- 3:55:59This, again, is that formula.
- 3:56:00The probability of A given B is equal to the probability
- 3:56:04that A and B take place divided by the probability of B.
- 3:56:08And you'll see this formula sometimes written in a couple of different ways.
- 3:56:11You could imagine algebraically multiplying both sides of this equation
- 3:56:15by probability of B to get rid of the fraction,
- 3:56:18and you'll get an expression like this.
- 3:56:20The probability of A and B, which is this expression over here,
- 3:56:24is just the probability of B times the probability of A given B.
- 3:56:28Or you could represent this equivalently since A and B in this expression
- 3:56:31are interchangeable.
- 3:56:32A and B is the same thing as B and A. You could imagine also
- 3:56:36representing the probability of A and B as the probability of A
- 3:56:41times the probability of B given A, just switching all of the A's and B's.
- 3:56:45These three are all equivalent ways of trying
- 3:56:47to represent what joint probability means.
- 3:56:49And so you'll sometimes see all of these equations,
- 3:56:52and they might be useful to you as you begin to reason about probability
- 3:56:55and to think about what values might be taking place in the real world.
- 3:57:00Now, sometimes when we deal with probability,
- 3:57:02we don't just care about a Boolean event like did this happen
- 3:57:05or did this not happen.
- 3:57:06Sometimes we might want the ability to represent variable values
- 3:57:10in a probability space where some variable might take
- 3:57:13on multiple different possible values.
- 3:57:16And in probability, we call a variable in probability theory
- 3:57:19a random variable.
- 3:57:21A random variable in probability is just some variable in probability theory
- 3:57:25that has some domain of values that it can take on.
- 3:57:28So what do I mean by this?
- 3:57:29Well, what I mean is I might have a random variable that is just
- 3:57:32called roll, for example, that has six possible values.
- 3:57:36Roll is my variable, and the possible values, the domain of values
- 3:57:39that it can take on are 1, 2, 3, 4, 5, and 6.
- 3:57:43And I might like to know the probability of each.
- 3:57:45In this case, they happen to all be the same.
- 3:57:47But in other random variables, that might not be the case.
- 3:57:50For example, I might have a random variable
- 3:57:52to represent the weather, for example, where the domain of values
- 3:57:55it could take on are things like sun or cloudy or rainy or windy or snowy.
- 3:57:59And each of those might have a different probability.
- 3:58:02And I care about knowing what is the probability that the weather equals
- 3:58:05sun or that the weather equals clouds, for instance.
- 3:58:08And I might like to do some mathematical calculations
- 3:58:11based on that information.
- 3:58:12Other random variables might be something like traffic.
- 3:58:15What are the odds that there is no traffic or light traffic or heavy traffic?
- 3:58:18Traffic, in this case, is my random variable.
- 3:58:21And the values that that random variable can take on are here.
- 3:58:24It's either none or light or heavy.
- 3:58:26And I, the person doing these calculations,
- 3:58:28I, the person encoding these random variables into my computer,
- 3:58:32need to make the decision as to what these possible values actually are.
- 3:58:36You might imagine, for example, for a flight.
- 3:58:38If I care about whether or not I make it or do a flight on time,
- 3:58:41my flight has a couple of possible values that it could take on.
- 3:58:43My flight could be on time.
- 3:58:45My flight could be delayed.
- 3:58:46My flight could be canceled.
- 3:58:47So flight, in this case, is my random variable.
- 3:58:51And these are the values that it can take on.
- 3:58:54And often, I want to know something about the probability
- 3:58:57that my random variable takes on each of those possible values.
- 3:59:00And this is what we then call a probability distribution.
- 3:59:04A probability distribution takes a random variable
- 3:59:07and gives me the probability for each of the possible values in its domain.
- 3:59:12So in the case of this flight, for example, my probability distribution
- 3:59:15might look something like this.
- 3:59:16My probability distribution says the probability
- 3:59:19that the random variable flight is equal to the value on time is 0.6.
- 3:59:25Or otherwise, put into more English human-friendly terms,
- 3:59:28the likelihood that my flight is on time is 60%, for example.
- 3:59:32And in this case, the probability that my flight is delayed is 30%.
- 3:59:35The probability that my flight is canceled is 10% or 0.1.
- 3:59:39And if you sum up all of these possible values,
- 3:59:42the sum is going to be 1, right?
- 3:59:43If you take all of the possible worlds, here
- 3:59:46are my three possible worlds for the value of the random variable flight,
- 3:59:49add them all up together, the result needs
- 3:59:52to be the number 1 per that axiom of probability theory
- 3:59:55that we've discussed before.
- 3:59:57So this now is one way of representing this probability
- 4:00:00distribution for the random variable flight.
- 4:00:03Sometimes you'll see it represented a little bit more concisely
- 4:00:06that this is pretty verbose for really just trying
- 4:00:08to express three possible values.
- 4:00:10And so often, you'll instead see the same notation
- 4:00:13representing using a vector.
- 4:00:15And all a vector is is a sequence of values.
- 4:00:17As opposed to just a single value, I might have multiple values.
- 4:00:21And so I could extend instead, represent this idea this way.
- 4:00:25Bold p, so a larger p, generally meaning the probability distribution
- 4:00:29of this variable flight is equal to this vector represented in angle brackets.
- 4:00:35The probability distribution is 0.6, 0.3, and 0.1.
- 4:00:39And I would just have to know that this probability distribution is
- 4:00:42in order of on time or delayed and canceled
- 4:00:46to know how to interpret this vector.
- 4:00:48To mean the first value in the vector is the probability
- 4:00:51that my flight is on time.
- 4:00:52The second value in the vector is the probability that my flight is delayed.
- 4:00:56And the third value in the vector is the probability
- 4:00:58that my flight is canceled.
- 4:01:00And so this is just an alternate way of representing this idea,
- 4:01:03a little more verbosely.
- 4:01:05But oftentimes, you'll see us just talk about a probability distribution
- 4:01:08over a random variable.
- 4:01:10And whenever we talk about that, what we're really doing
- 4:01:12is trying to figure out the probabilities of each of the possible values
- 4:01:16that that random variable can take on.
- 4:01:17But this notation is just a little bit more succinct,
- 4:01:20even though it can sometimes be a little confusing,
- 4:01:22depending on the context in which you see it.
- 4:01:24So we'll start to look at examples where we use this sort of notation
- 4:01:27to describe probability and to describe events that might take place.
- 4:01:33A couple of other important ideas to know with regards to probability theory.
- 4:01:37One is this idea of independence.
- 4:01:39And independence refers to the idea that the knowledge of one event
- 4:01:43doesn't influence the probability of another event.
- 4:01:46So for example, in the context of my two dice rolls,
- 4:01:48where I had the red die and the blue die, the probability
- 4:01:51that I roll the red die and the blue die,
- 4:01:54those two events, red die and blue die, are independent.
- 4:01:57Knowing the result of the red die doesn't change
- 4:02:00the probabilities for the blue die.
- 4:02:01It doesn't give me any additional information
- 4:02:03about what the value of the blue die is ultimately going to be.
- 4:02:06But that's not always going to be the case.
- 4:02:08You might imagine that in the case of weather, something
- 4:02:11like clouds and rain, those are probably not independent.
- 4:02:15But if it is cloudy, that might increase the probability that later
- 4:02:18in the day it's going to rain.
- 4:02:20So some information informs some other event or some other random variable.
- 4:02:24So independence refers to the idea that one event doesn't influence the other.
- 4:02:29And if they're not independent, then there might be some relationship.
- 4:02:34So mathematically, formally, what does independence actually mean?
- 4:02:37Well, recall this formula from before, that the probability of A and B
- 4:02:42is the probability of A times the probability of B given A.
- 4:02:46And the more intuitive way to think about this
- 4:02:48is that to know how likely it is that A and B happen,
- 4:02:51well, let's first figure out the likelihood that A happens.
- 4:02:54And then given that we know that A happens,
- 4:02:56let's figure out the likelihood that B happens
- 4:02:58and multiply those two things together.
- 4:03:01But if A and B were independent, meaning knowing A
- 4:03:05doesn't change anything about the likelihood that B is true,
- 4:03:09well, then the probability of B given A, meaning the probability that B is true,
- 4:03:14given that I know A is true, well, that I know A is true
- 4:03:17shouldn't really make a difference if these two things are independent,
- 4:03:20that A shouldn't influence B at all.
- 4:03:22So the probability of B given A is really just the probability of B.
- 4:03:27If it is true that A and B are independent.
- 4:03:30And so this right here is one example of a definition
- 4:03:33for what it means for A and B to be independent.
- 4:03:36The probability of A and B is just the probability
- 4:03:39of A times the probability of B. Anytime you find two events A and B
- 4:03:44where this relationship holds, then you can say that A and B are independent.
- 4:03:49So an example of that might be the dice that we were taking a look at before.
- 4:03:53Here, if I wanted the probability of red being a 6 and blue being a 6,
- 4:03:58well, that's just the probability that red is a 6 multiplied
- 4:04:01by the probability that blue is a 6.
- 4:04:03It's both equal to 1 over 36.
- 4:04:05So I can say that these two events are independent.
- 4:04:10What wouldn't be independent, for example, would be an example.
- 4:04:13So this, for example, has a probability of 1 over 36,
- 4:04:16as we talked about before.
- 4:04:17But what wouldn't be independent would be a case like this,
- 4:04:20the probability that the red die rolls a 6 and the red die rolls a 4.
- 4:04:26If you just naively took, OK, red die 6, red die 4,
- 4:04:29well, if I'm only rolling the die once, you
- 4:04:31might imagine the naive approach is to say, well, each of these
- 4:04:34has a probability of 1 over 6.
- 4:04:35So multiply them together, and the probability is 1 over 36.
- 4:04:39But of course, if you're only rolling the red die once,
- 4:04:41there's no way you could get two different values for the red die.
- 4:04:45It couldn't both be a 6 and a 4.
- 4:04:48So the probability should be 0.
- 4:04:50But if you were to multiply probability of red 6 times
- 4:04:53probability of red 4, well, that would equal 1 over 36.
- 4:04:57But of course, that's not true.
- 4:04:58Because we know that there is no way, probability 0,
- 4:05:01that when we roll the red die once, we get both a 6 and a 4,
- 4:05:06because only one of those possibilities can actually be the result.
- 4:05:10And so we can say that the event that red roll is 6
- 4:05:14and the event that red roll is 4, those two events are not independent.
- 4:05:18If I know that the red roll is a 6, I know that the red roll cannot possibly
- 4:05:23be a 4, so these things are not independent.
- 4:05:25And instead, if I wanted to calculate the probability,
- 4:05:28I would need to use this conditional probability
- 4:05:31as the regular definition of the probability of two events taking place.
- 4:05:36And the probability of this now, well, the probability
- 4:05:38of the red roll being a 6, that's 1 over 6.
- 4:05:41But what's the probability that the roll is a 4 given that the roll is a 6?
- 4:05:45Well, this is just 0, because there's no way for the red roll to be a 4,
- 4:05:50given that we already know the red roll is a 6.
- 4:05:53And so the value, if we do add all that multiplication, is we get the number 0.
- 4:05:59So this idea of conditional probability is going to come up again and again,
- 4:06:02especially as we begin to reason about multiple different random variables
- 4:06:06that might be interacting with each other in some way.
- 4:06:08And this gets us to one of the most important rules
- 4:06:10in probability theory, which is known as Bayes rule.
- 4:06:14And it turns out that just using the information we've already
- 4:06:17learned about probability and just applying a little bit of algebra,
- 4:06:20we can actually derive Bayes rule for ourselves.
- 4:06:23But it's a very important rule when it comes to inference
- 4:06:26and thinking about probability in the context of what
- 4:06:28it is that a computer can do or what a mathematician could
- 4:06:31do by having access to information about probability.
- 4:06:34So let's go back to these equations to be able to derive Bayes rule ourselves.
- 4:06:39We know the probability of A and B, the likelihood that A and B take place,
- 4:06:43is the likelihood of B, and then the likelihood of A,
- 4:06:47given that we know that B is already true.
- 4:06:49And likewise, the probability of A given A and B
- 4:06:52is the probability of A times the probability of B,
- 4:06:56given that we know that A is already true.
- 4:06:58This is sort of a symmetric relationship where
- 4:07:00it doesn't matter the order of A and B and B and A mean the same thing.
- 4:07:04And so in these equations, we can just swap out A and B
- 4:07:07to be able to represent the exact same idea.
- 4:07:09So we know that these two equations are already true.
- 4:07:12We've seen that already.
- 4:07:13And now let's just do a little bit of algebraic manipulation of this stuff.
- 4:07:17Both of these expressions on the right-hand side
- 4:07:19are equal to the probability of A and B. So what I can do
- 4:07:24is take these two expressions on the right-hand side
- 4:07:26and just set them equal to each other.
- 4:07:28If they're both equal to the probability of A and B,
- 4:07:32then they both must be equal to each other.
- 4:07:34So probability of A times probability of B given A
- 4:07:38is equal to the probability of B times the probability of A given B.
- 4:07:44And now all we're going to do is do a little bit of division.
- 4:07:47I'm going to divide both sides by P of A. And now I get what is Bayes' rule.
- 4:07:53The probability of B given A is equal to the probability of B
- 4:07:59times the probability of A given B divided by the probability of A.
- 4:08:03And sometimes in Bayes' rule, you'll see the order
- 4:08:05of these two arguments switched.
- 4:08:06So instead of B times A given B, it'll be A given B times B.
- 4:08:10That ultimately doesn't matter because in multiplication,
- 4:08:12you can switch the order of the two things you're multiplying,
- 4:08:15and it doesn't change the result. But this here right now
- 4:08:18is the most common formulation of Bayes' rule.
- 4:08:21The probability of B given A is equal to the probability of A given
- 4:08:26B times the probability of B divided by the probability of A.
- 4:08:31And this rule, it turns out, is really important
- 4:08:33when it comes to trying to infer things about the world,
- 4:08:36because it means you can express one conditional probability,
- 4:08:39the conditional probability of B given A, using knowledge
- 4:08:44about the probability of A given B, using the reverse
- 4:08:47of that conditional probability.
- 4:08:49So let's first do a little bit of an example with this,
- 4:08:51just to see how we might use it, and then explore
- 4:08:54what this means a little bit more generally.
- 4:08:56So we're going to construct a situation where I have some information.
- 4:08:59There are two events that I care about, the idea
- 4:09:02that it's cloudy in the morning and the idea
- 4:09:05that it is rainy in the afternoon.
- 4:09:07Those are two different possible events that could take place,
- 4:09:10cloudy in the morning, or the AM, rainy in the PM.
- 4:09:13And what I care about is, given clouds in the morning,
- 4:09:17what is the probability of rain in the afternoon?
- 4:09:19A reasonable question I might ask, in the morning,
- 4:09:22I look outside, or an AI's camera looks outside
- 4:09:24and sees that there are clouds in the morning.
- 4:09:27And we want to conclude, we want to figure out what is the probability
- 4:09:30that in the afternoon, there is going to be rain.
- 4:09:34Of course, in the abstract, we don't have access
- 4:09:36to this kind of information, but we can use data
- 4:09:38to begin to try and figure this out.
- 4:09:40So let's imagine now that I have access to some pieces of information.
- 4:09:44I have access to the idea that 80% of rainy afternoons
- 4:09:48start out with a cloudy morning.
- 4:09:50And you might imagine that I could have gathered this data just
- 4:09:52by looking at data over a sequence of time,
- 4:09:54that I know that 80% of the time when it's raining in the afternoon,
- 4:09:58it was cloudy that morning.
- 4:10:01I also know that 40% of days have cloudy mornings.
- 4:10:04And I also know that 10% of days have rainy afternoons.
- 4:10:08And now using this information, I would like to figure out,
- 4:10:12given clouds in the morning, what is the probability
- 4:10:15that it rains in the afternoon?
- 4:10:16I want to know the probability of afternoon rain given morning clouds.
- 4:10:21And I can do that, in particular, using this fact, the probability of,
- 4:10:26so if I know that 80% of rainy afternoons start with cloudy mornings,
- 4:10:29then I know the probability of cloudy mornings given rainy afternoons.
- 4:10:34So using sort of the reverse conditional probability,
- 4:10:36I can figure that out.
- 4:10:38Expressed in terms of Bayes rule, this is what that would look like.
- 4:10:41Probability of rain given clouds is the probability of clouds given rain
- 4:10:46times the probability of rain divided by the probability of clouds.
- 4:10:50Here I'm just substituting in for the values of a and b
- 4:10:53from that equation of Bayes rule from before.
- 4:10:55And then I can just do the math.
- 4:10:56I have this information.
- 4:10:57I know that 80% of the time, if it was raining,
- 4:11:00then there were clouds in the morning.
- 4:11:01So 0.8 here.
- 4:11:03Probability of rain is 0.1, because 10% of days were rainy,
- 4:11:06and 40% of days were cloudy.
- 4:11:08I do the math, and I can figure out the answer is 0.2.
- 4:11:11So the probability that it rains in the afternoon,
- 4:11:14given that it was cloudy in the morning, is 0.2 in this case.
- 4:11:19And this now is an application of Bayes rule,
- 4:11:22the idea that using one conditional probability,
- 4:11:24we can get the reverse conditional probability.
- 4:11:27And this is often useful when one of the conditional probabilities
- 4:11:31might be easier for us to know about or easier for us to have data about.
- 4:11:34And using that information, we can calculate
- 4:11:37the other conditional probability.
- 4:11:39So what does this look like?
- 4:11:40Well, it means that knowing the probability of cloudy mornings
- 4:11:43given rainy afternoons, we can calculate the probability
- 4:11:47of rainy afternoons given cloudy mornings.
- 4:11:50Or, for example, more generally, if we know the probability
- 4:11:54of some visible effect, some effect that we can see and observe,
- 4:11:58given some unknown cause that we're not sure about,
- 4:12:02well, then we can calculate the probability of that unknown cause
- 4:12:05given the visible effect.
- 4:12:08So what might that look like?
- 4:12:10Well, in the context of medicine, for example,
- 4:12:12I might know the probability of some medical test result given a disease.
- 4:12:17Like, I know that if someone has a disease,
- 4:12:19then x% of the time the medical test result will show up as this,
- 4:12:23for instance.
- 4:12:24And using that information, then I can calculate, all right,
- 4:12:26what is the probability that given I know the medical test result, what
- 4:12:31is the likelihood that someone has the disease?
- 4:12:33This is the piece of information that is usually easier to know,
- 4:12:36easier to immediately have access to data for.
- 4:12:38And this is the information that I actually want to calculate.
- 4:12:42Or I might want to know, for example, if I
- 4:12:44know that some probability of counterfeit bills
- 4:12:48have blurry text around the edges, because counterfeit printers aren't
- 4:12:51nearly as good at printing text precisely.
- 4:12:53So I have some information about, given that something
- 4:12:56is a counterfeit bill, like x% of counterfeit bills
- 4:12:59have blurry text, for example.
- 4:13:01And using that information, then I can calculate some piece of information
- 4:13:04that I might want to know, like, given that I know there's blurry text
- 4:13:08on a bill, what is the probability that that bill is counterfeit?
- 4:13:12So given one conditional probability, I can
- 4:13:14calculate the other conditional probability as well.
- 4:13:19And so now we've taken a look at a couple of different types of probability.
- 4:13:22And we've looked at unconditional probability,
- 4:13:24where I just look at what is the probability of this event occurring,
- 4:13:27given no additional evidence that I might have access to.
- 4:13:31And we've also looked at conditional probability,
- 4:13:33where I have some sort of evidence, and I
- 4:13:35would like to, using that evidence, be able to calculate some other
- 4:13:38probability as well.
- 4:13:40And the other kind of probability that will be important for us to think about
- 4:13:43is joint probability.
- 4:13:45And this is when we're considering the likelihood
- 4:13:47of multiple different events simultaneously.
- 4:13:50And so what do we mean by this?
- 4:13:52For example, I might have probability distributions
- 4:13:55that look a little something like this.
- 4:13:56Like, oh, I want to know the probability distribution of clouds
- 4:13:59in the morning.
- 4:14:00And that distribution looks like this.
- 4:14:0240% of the time, C, which is my random variable here,
- 4:14:06is equal to it's cloudy.
- 4:14:07And 60% of the time, it's not cloudy.
- 4:14:10So here is just a simple probability distribution
- 4:14:13that is effectively telling me that 40% of the time, it's cloudy.
- 4:14:17I might also have a probability distribution for rain in the afternoon,
- 4:14:20where 10% of the time, or with probability 0.1,
- 4:14:24it is raining in the afternoon.
- 4:14:25And with probability 0.9, it is not raining in the afternoon.
- 4:14:30And using just these two pieces of information,
- 4:14:34I don't actually have a whole lot of information
- 4:14:36about how these two variables relate to each other.
- 4:14:39But I could if I had access to their joint probability,
- 4:14:42meaning for every combination of these two things,
- 4:14:45meaning morning cloudy and afternoon rain, morning cloudy and afternoon not
- 4:14:49rain, morning not cloudy and afternoon rain,
- 4:14:52and morning not cloudy and afternoon not raining,
- 4:14:54if I had access to values for each of those four,
- 4:14:57I'd have more information.
- 4:14:58So information that'd be organized in a table like this,
- 4:15:02and this, rather than just a probability distribution,
- 4:15:05is a joint probability distribution.
- 4:15:07It tells me the probability distribution of each
- 4:15:10of the possible combinations of values that these random variables can take on.
- 4:15:15So if I want to know what is the probability that on any given day
- 4:15:19it is both cloudy and rainy, well, I would say, all right,
- 4:15:22we're looking at cases where it is cloudy and cases where it is raining.
- 4:15:26And the intersection of those two, that row in that column, is 0.08.
- 4:15:30So that is the probability that it is both cloudy and rainy using
- 4:15:35that information.
- 4:15:36And using this conditional probability table,
- 4:15:39using this joint probability table, I can
- 4:15:41begin to draw other pieces of information about things like conditional
- 4:15:46probability.
- 4:15:47So I might ask a question like, what is the probability distribution of clouds
- 4:15:51given that I know that it is raining?
- 4:15:53Meaning I know for sure that it's raining.
- 4:15:56Tell me the probability distribution over whether it's cloudy or not,
- 4:15:59given that I know already that it is, in fact, raining.
- 4:16:02And here I'm using C to stand for that random variable.
- 4:16:05I'm looking for a distribution, meaning the answer to this
- 4:16:07is not going to be a single value.
- 4:16:09It's going to be two values, a vector of two values,
- 4:16:12where the first value is probability of clouds,
- 4:16:14the second value is probability that it is not cloudy,
- 4:16:17but the sum of those two values is going to be 1.
- 4:16:19Because when you add up the probabilities of all of the possible worlds,
- 4:16:23the result that you get must be the number 1.
- 4:16:26And well, what do we know about how to calculate a conditional probability?
- 4:16:30Well, we know that the probability of A given B
- 4:16:33is the probability of A and B divided by the probability of B.
- 4:16:38So what does this mean?
- 4:16:40Well, it means that I can calculate the probability of clouds
- 4:16:43given that it's raining as the probability of clouds and raining
- 4:16:49divided by the probability of rain.
- 4:16:50And this comma here for the probability distribution
- 4:16:53of clouds and rain, this comma sort of stands in for the word and.
- 4:16:57You'll sort of see in the logical operator and and the comma
- 4:16:59used interchangeably.
- 4:17:01This means the probability distribution over the clouds
- 4:17:04and knowing the fact that it is raining divided
- 4:17:06by the probability of rain.
- 4:17:09And the interesting thing to note here and what
- 4:17:11we'll often do in order to simplify our mathematics
- 4:17:13is that dividing by the probability of rain,
- 4:17:16the probability of rain here is just some numerical constant.
- 4:17:19It is some number.
- 4:17:20Dividing by probability of rain is just dividing by some constant,
- 4:17:24or in other words, multiplying by the inverse of that constant.
- 4:17:27And it turns out that oftentimes we can just not
- 4:17:30worry about what the exact value of this is
- 4:17:32and just know that it is, in fact, a constant value.
- 4:17:36And we'll see why in a moment.
- 4:17:37So instead of expressing this as this joint probability divided
- 4:17:41by the probability of rain, sometimes we'll
- 4:17:43just represent it as alpha times the numerator here,
- 4:17:47the probability distribution of C, this variable,
- 4:17:50and that we know that it is raining, for instance.
- 4:17:53So all we've done here is said this value of 1 over the probability of rain,
- 4:17:57that's really just a constant we're going to divide by or equivalently
- 4:18:00multiply by the inverse of at the end.
- 4:18:02We'll just call it alpha for now and deal with it a little bit later.
- 4:18:06But the key idea here now, and this is an idea that's going to come up again,
- 4:18:09is that the conditional distribution of C given rain
- 4:18:14is proportional to, meaning just some factor multiplied
- 4:18:17by the joint probability of C and rain being true.
- 4:18:22And so how do we figure this out?
- 4:18:23Well, this is going to be the probability that it
- 4:18:25is cloudy given that it's raining, which is 0.08,
- 4:18:28and the probability that it's not cloudy given
- 4:18:30that it's raining, which is 0.02.
- 4:18:32And so we get alpha times here now is that probability distribution.
- 4:18:370.08 is clouds and rain.
- 4:18:400.02 is not cloudy and rain.
- 4:18:43But of course, 0.08 and 0.02 don't sum up to the number 1.
- 4:18:47And we know that in a probability distribution,
- 4:18:50if you consider all of the possible values,
- 4:18:52they must sum up to a probability of 1.
- 4:18:55And so we know that we just need to figure out
- 4:18:57some constant to normalize, so to speak, these values, something
- 4:19:01we can multiply or divide by to get it so that all these probabilities sum up
- 4:19:05to 1, and it turns out that if we multiply both numbers by 10,
- 4:19:08then we can get that result of 0.8 and 0.2.
- 4:19:11The proportions are still equivalent, but now 0.8 plus 0.2,
- 4:19:15those sum up to the number 1.
- 4:19:18So take a look at this and see if you can understand step by step
- 4:19:21how it is we're getting from one point to another.
- 4:19:23The key idea here is that by using the joint probabilities,
- 4:19:27these probabilities that it is both cloudy and rainy
- 4:19:31and that it is not cloudy and rainy, I can take that information
- 4:19:35and figure out the conditional probability given that it's raining.
- 4:19:39What is the chance that it's cloudy versus not cloudy?
- 4:19:41Just by multiplying by some normalization constant, so to speak.
- 4:19:46And this is what a computer can begin to use
- 4:19:48to be able to interact with these various different types of probabilities.
- 4:19:52And it turns out there are a number of other probability rules
- 4:19:55that are going to be useful to us as we begin
- 4:19:57to explore how we can actually use this information to encode
- 4:20:01into our computers some more complex analysis that we might want to do
- 4:20:05about probability and distributions and random variables
- 4:20:08that we might be interacting with.
- 4:20:10So here are a couple of those important probability rules.
- 4:20:12One of the simplest rules is just this negation rule.
- 4:20:15What is the probability of not event A?
- 4:20:19So A is an event that has some probability,
- 4:20:21and I would like to know what is the probability that A does not occur.
- 4:20:25And it turns out it's just 1 minus P of A, which makes sense.
- 4:20:29Because if those are the two possible cases, either A happens or A
- 4:20:33doesn't happen, then when you add up those two cases, you must get 1,
- 4:20:37which means that P of not A must just be 1 minus P of A.
- 4:20:42Because P of A and P of not A must sum up to the number 1.
- 4:20:46They must include all of the possible cases.
- 4:20:49We've seen an expression for calculating the probability of A and B.
- 4:20:53We might also reasonably want to calculate the probability of A or B.
- 4:20:57What is the probability that one thing happens or another thing happens?
- 4:21:01So for example, I might want to calculate what is the probability
- 4:21:04that if I roll two dice, a red die and a blue die, what is the likelihood
- 4:21:07that A is a 6 or B is a 6, like one or the other?
- 4:21:11And what you might imagine you could do, and the wrong way to approach it,
- 4:21:14would be just to say, all right, well, A comes up as a 6 with the red die
- 4:21:19comes up as a 6 with probability 1 over 6.
- 4:21:21The same for the blue die, it's also 1 over 6.
- 4:21:23Add them together, and you get 2 over 6, otherwise known as 1 third.
- 4:21:27But this suffers from a problem of over counting,
- 4:21:30that we've double counted the case, where both A and B, both the red die
- 4:21:34and the blue die, both come up as a 6-roll.
- 4:21:37And I've counted that instance twice.
- 4:21:39So to resolve this, the actual expression for calculating the probability of A
- 4:21:43or B uses what we call the inclusion-exclusion formula.
- 4:21:47So I take the probability of A, add it to the probability of B.
- 4:21:51That's all same as before.
- 4:21:52But then I need to exclude the cases that I've double counted.
- 4:21:56So I subtract from that the probability of A and B.
- 4:22:01And that gets me the result for A or B. I consider all the cases where A is true
- 4:22:05and all the cases where B is true.
- 4:22:07And if you imagine this is like a Venn diagram of cases where A is true,
- 4:22:09cases where B is true, I just need to subtract out the middle
- 4:22:12to get rid of the cases that I have overcounted by double counting them
- 4:22:16inside of both of these individual expressions.
- 4:22:21One other rule that's going to be quite helpful
- 4:22:23is a rule called marginalization.
- 4:22:25So marginalization is answering the question
- 4:22:27of how do I figure out the probability of A using some other variable
- 4:22:31that I might have access to, like B?
- 4:22:33Even if I don't know additional information about it,
- 4:22:35I know that B, some event, can have two possible states, either B
- 4:22:40happens or B doesn't happen, assuming it's a Boolean, true or false.
- 4:22:44And well, what that means is that for me to be
- 4:22:47able to calculate the probability of A, there are only two cases.
- 4:22:50Either A happens and B happens, or A happens and B doesn't happen.
- 4:22:55And those are two disjoint, meaning they can't both happen together.
- 4:22:58Either B happens or B doesn't happen.
- 4:23:01They're disjoint or separate cases.
- 4:23:03And so I can figure out the probability of A
- 4:23:05just by adding up those two cases.
- 4:23:07The probability that A is true is the probability that A and B is true,
- 4:23:13plus the probability that A is true and B isn't true.
- 4:23:16So by marginalizing, I've looked at the two possible cases
- 4:23:19that might take place, either B happens or B doesn't happen.
- 4:23:23And in either of those cases, I look at what's
- 4:23:25the probability that A happens.
- 4:23:27And if I add those together, well, then I get the probability
- 4:23:30that A happens as a whole.
- 4:23:32So take a look at that rule.
- 4:23:33It doesn't matter what B is or how it's related to A.
- 4:23:36So long as I know these joint distributions,
- 4:23:39I can figure out the overall probability of A.
- 4:23:42And this can be a useful way if I have a joint distribution,
- 4:23:44like the joint distribution of A and B, to just figure out
- 4:23:48some unconditional probability, like the probability of A.
- 4:23:51And we'll see examples of this soon as well.
- 4:23:54Now, sometimes these might not just be random,
- 4:23:55might not just be variables that are events that are like they happened
- 4:23:58or they didn't happen, like B is here.
- 4:24:00They might be some broader probability distribution
- 4:24:03where there are multiple possible values.
- 4:24:05And so here, in order to use this marginalization rule,
- 4:24:08I need to sum up not just over B and not B,
- 4:24:11but for all of the possible values that the other random variable could take
- 4:24:15on.
- 4:24:16And so here, we'll see a version of this rule for random variables.
- 4:24:19And it's going to include that summation notation
- 4:24:21to indicate that I'm summing up, adding up a whole bunch of individual values.
- 4:24:25So here's the rule.
- 4:24:26Looks a lot more complicated, but it's actually
- 4:24:28the equivalent exactly the same rule.
- 4:24:30What I'm saying here is that if I have two random variables, one called x
- 4:24:35and one called y, well, the probability that x is equal to some value x sub i,
- 4:24:41this is just some value that this variable takes on.
- 4:24:43How do I figure it out?
- 4:24:45Well, I'm going to sum up over j, where j is going
- 4:24:48to range over all of the possible values that y can take on.
- 4:24:53Well, let's look at the probability that x equals xi and y equals yj.
- 4:24:58So the exact same rule, the only difference here
- 4:25:00is now I'm summing up over all of the possible values
- 4:25:03that y can take on, saying let's add up all of those possible cases
- 4:25:06and look at this joint distribution, this joint probability,
- 4:25:10that x takes on the value I care about, given all of the possible values for y.
- 4:25:15And if I add all those up, then I can get
- 4:25:18this unconditional probability of what x is equal to,
- 4:25:22whether or not x is equal to some value x sub i.
- 4:25:26So let's take a look at this rule, because it
- 4:25:27does look a little bit complicated.
- 4:25:29Let's try and put a concrete example to it.
- 4:25:31Here again is that same joint distribution from before.
- 4:25:34I have cloud, not cloudy, rainy, not rainy.
- 4:25:38And maybe I want to access some variable.
- 4:25:40I want to know what is the probability that it is cloudy.
- 4:25:44Well, marginalization says that if I have this joint distribution
- 4:25:48and I want to know what is the probability that it is cloudy,
- 4:25:51well, I need to consider the other variable, the variable that's not here,
- 4:25:55the idea that it's rainy.
- 4:25:56And I consider the two cases, either it's raining or it's not raining.
- 4:26:00And I just sum up the values for each of those possibilities.
- 4:26:04In other words, the probability that it is cloudy
- 4:26:07is equal to the sum of the probability that it's cloudy and it's rainy
- 4:26:12and the probability that it's cloudy and it is not raining.
- 4:26:17And so these now are values that I have access to.
- 4:26:20These are values that are just inside of this joint probability table.
- 4:26:24What is the probability that it is both cloudy and rainy?
- 4:26:27Well, it's just the intersection of these two here, which is 0.08.
- 4:26:31And the probability that it's cloudy and not raining is, all right,
- 4:26:34here's cloudy, here's not raining.
- 4:26:36It's 0.32.
- 4:26:37So it's 0.08 plus 0.32, which just gives us equal to 0.4.
- 4:26:42That is the unconditional probability that it is, in fact, cloudy.
- 4:26:46And so marginalization gives us a way to go from these joint distributions
- 4:26:50to just some individual probability that I might care about.
- 4:26:53And you'll see a little bit later why it is that we care about that
- 4:26:56and why that's actually useful to us as we begin
- 4:26:59doing some of these calculations.
- 4:27:01Last rule we'll take a look at before transitioning
- 4:27:04to something a little bit different is this rule of conditioning,
- 4:27:06very similar to the marginalization rule.
- 4:27:09But it says that, again, if I have two events, a and b,
- 4:27:12but instead of having access to their joint probabilities,
- 4:27:15I have access to their conditional probabilities,
- 4:27:17how they relate to each other.
- 4:27:19Well, again, if I want to know the probability that a happens,
- 4:27:22and I know that there's some other variable b, either b happens or b
- 4:27:26doesn't happen, and so I can say that the probability of a
- 4:27:30is the probability of a given b times the probability of b, meaning b happened.
- 4:27:35And given that I know b happened, what's the likelihood that a happened?
- 4:27:39And then I consider the other case, that b didn't happen.
- 4:27:42So here's the probability that b didn't happen.
- 4:27:44And here's the probability that a happens,
- 4:27:47given that I know that b didn't happen.
- 4:27:49And this is really the equivalent rule just
- 4:27:51using conditional probability instead of joint probability,
- 4:27:55where I'm saying let's look at both of these two cases and condition on b.
- 4:27:59Look at the case where b happens, and look at the case where b doesn't happen,
- 4:28:03and look at what probabilities I get as a result.
- 4:28:06And just as in the case of marginalization,
- 4:28:08where there was an equivalent rule for random variables
- 4:28:10that could take on multiple possible values in a domain of possible values,
- 4:28:14here, too, conditioning has the same equivalent rule.
- 4:28:17Again, there's a summation to mean I'm summing over
- 4:28:19all of the possible values that some random variable y could take on.
- 4:28:23But if I want to know what is the probability that x takes on this value,
- 4:28:27then I'm going to sum up over all the values j that y could take on,
- 4:28:31and say, all right, what's the chance that y takes on that value yj?
- 4:28:35And multiply it by the conditional probability
- 4:28:38that x takes on this value, given that y took on that value yj.
- 4:28:42So equivalent rule just using conditional probabilities
- 4:28:46instead of joint probabilities.
- 4:28:47And using the equation we know about joint probabilities,
- 4:28:50we can translate between these two.
- 4:28:53So all right, we've seen a whole lot of mathematics,
- 4:28:55and we've just laid the foundation for mathematics.
- 4:28:57And no need to worry if you haven't seen probability in too much detail
- 4:29:00up until this point.
- 4:29:02These are the foundations of the ideas that are going to come up
- 4:29:05as we begin to explore how we can now take these ideas from probability
- 4:29:09and begin to apply them to represent something inside of our computer,
- 4:29:12something inside of the AI agent we're trying to design that
- 4:29:16is able to represent information and probabilities
- 4:29:18and the likelihoods between various different events.
- 4:29:22So there are a number of different probabilistic models
- 4:29:24that we can generate, but the first of the models
- 4:29:26we're going to talk about are what are known as Bayesian networks.
- 4:29:30And a Bayesian network is just going to be some network of random variables,
- 4:29:34connected random variables that are going to represent
- 4:29:37the dependence between these random variables.
- 4:29:39The odds are most random variables in this world
- 4:29:43are not independent from each other, but there's
- 4:29:45some relationship between things that are happening that we care about.
- 4:29:48If it is rainy today, that might increase the likelihood
- 4:29:51that my flight or my train gets delayed, for example.
- 4:29:54There are some dependence between these random variables,
- 4:29:57and a Bayesian network is going to be able to capture those dependencies.
- 4:30:01So what is a Bayesian network?
- 4:30:03What is its actual structure, and how does it work?
- 4:30:06Well, a Bayesian network is going to be a directed graph.
- 4:30:08And again, we've seen directed graphs before.
- 4:30:10They are individual nodes with arrows or edges
- 4:30:13that connect one node to another node pointing in a particular direction.
- 4:30:18And so this directed graph is going to have nodes
- 4:30:20as well, where each node in this directed graph
- 4:30:23is going to represent a random variable, something like the weather,
- 4:30:27or something like whether my train was on time or delayed.
- 4:30:30And we're going to have an arrow from a node x to a node y
- 4:30:34to mean that x is a parent of y.
- 4:30:37So that'll be our notation.
- 4:30:38If there's an arrow from x to y, x is going to be considered a parent of y.
- 4:30:42And the reason that's important is because each of these nodes
- 4:30:46is going to have a probability distribution that we're
- 4:30:48going to store along with it, which is the distribution of x
- 4:30:52given some evidence, given the parents of x.
- 4:30:56So the way to more intuitively think about this
- 4:30:58is the parents seem to be thought of as sort of causes for some effect
- 4:31:01that we're going to observe.
- 4:31:04And so let's take a look at an actual example of a Bayesian network
- 4:31:07and think about the types of logic that might be involved
- 4:31:09in reasoning about that network.
- 4:31:11Let's imagine for a moment that I have an appointment out of town,
- 4:31:15and I need to take a train in order to get to that appointment.
- 4:31:18So what are the things I might care about?
- 4:31:19Well, I care about getting to my appointment on time.
- 4:31:22Whether I make it to my appointment and I'm able to attend it
- 4:31:24or I miss the appointment.
- 4:31:26And you might imagine that that's influenced by the train,
- 4:31:29that the train is either on time or it's delayed, for example.
- 4:31:33But that train itself is also influenced.
- 4:31:36Whether the train is on time or not depends maybe on the rain.
- 4:31:39Is there no rain?
- 4:31:40Is it light rain?
- 4:31:41Is there heavy rain?
- 4:31:42And it might also be influenced by other variables too.
- 4:31:44It might be influenced as well by whether or not
- 4:31:47there's maintenance on the train track, for example.
- 4:31:49If there is maintenance on the train track,
- 4:31:51that probably increases the likelihood that my train is delayed.
- 4:31:55And so we can represent all of these ideas
- 4:31:57using a Bayesian network that looks a little something like this.
- 4:32:01Here I have four nodes representing four random variables
- 4:32:05that I would like to keep track of.
- 4:32:06I have one random variable called rain that
- 4:32:08can take on three possible values in its domain, either none or light
- 4:32:12or heavy, for no rain, light rain, or heavy rain.
- 4:32:16I have a variable called maintenance for whether or not
- 4:32:18there is maintenance on the train track, which
- 4:32:20it has two possible values, just either yes or no.
- 4:32:22Either there is maintenance or there's no maintenance happening on the track.
- 4:32:26Then I have a random variable for the train indicating whether or not
- 4:32:28the train was on time or not.
- 4:32:30That random variable has two possible values in its domain.
- 4:32:33The train is either on time or the train is delayed.
- 4:32:37And then finally, I have a random variable
- 4:32:39for whether I make it to my appointment.
- 4:32:41For my appointment down here, I have a random variable
- 4:32:43called appointment that itself has two possible values, attend and miss.
- 4:32:49And so here are the possible values.
- 4:32:50Here are my four nodes, each of which represents a random variable, each
- 4:32:54of which has a domain of possible values that it can take on.
- 4:32:58And the arrows, the edges pointing from one node to another,
- 4:33:01encode some notion of dependence inside of this graph,
- 4:33:05that whether I make it to my appointment or not
- 4:33:08is dependent upon whether the train is on time or delayed.
- 4:33:12And whether the train is on time or delayed
- 4:33:14is dependent on two things given by the two arrows pointing at this node.
- 4:33:18It is dependent on whether or not there was maintenance on the train track.
- 4:33:22And it is also dependent upon whether or not it was raining
- 4:33:25or whether it is raining.
- 4:33:27And just to make things a little complicated,
- 4:33:29let's say as well that whether or not there is maintenance on the track,
- 4:33:32this too might be influenced by the rain.
- 4:33:34That if there's heavier rain, well, maybe it's
- 4:33:37less likely that it's going to be maintenance on the train track that day
- 4:33:40because they're more likely to want to do maintenance on the track on days
- 4:33:43when it's not raining, for example.
- 4:33:45And so these nodes might have different relationships between them.
- 4:33:47But the idea is that we can come up with a probability distribution
- 4:33:51for any of these nodes based only upon its parents.
- 4:33:56And so let's look node by node at what this probability distribution might
- 4:33:59actually look like.
- 4:34:00And we'll go ahead and begin with this root node, this rain node here,
- 4:34:03which is at the top, and has no arrows pointing into it, which
- 4:34:07means its probability distribution is not
- 4:34:10going to be a conditional distribution.
- 4:34:11It's not based on anything.
- 4:34:13I just have some probability distribution over the possible values
- 4:34:17for the rain random variable.
- 4:34:20And that distribution might look a little something like this.
- 4:34:23None, light and heavy, each have a possible value.
- 4:34:25Here I'm saying the likelihood of no rain is 0.7, of light rain is 0.2,
- 4:34:31of heavy rain is 0.1, for example.
- 4:34:33So here is a probability distribution for this root node in this Bayesian
- 4:34:38network.
- 4:34:39And let's now consider the next node in the network, maintenance.
- 4:34:42Track maintenance is yes or no.
- 4:34:44And the general idea of what this distribution is going to encode,
- 4:34:47at least in this story, is the idea that the heavier the rain is,
- 4:34:52the less likely it is that there's going to be maintenance on the track.
- 4:34:55Because the people that are doing maintenance on the track probably
- 4:34:57want to wait until a day when it's not as rainy in order
- 4:35:00to do the track maintenance, for example.
- 4:35:02And so what might that probability distribution look like?
- 4:35:05Well, this now is going to be a conditional probability distribution,
- 4:35:08that here are the three possible values for the rain random variable, which
- 4:35:12I'm here just going to abbreviate to R, either no rain, light rain,
- 4:35:15or heavy rain.
- 4:35:17And for each of those possible values, either there
- 4:35:19is yes track maintenance or no track maintenance.
- 4:35:22And those have probabilities associated with them.
- 4:35:25That I see here that if it is not raining,
- 4:35:30then there is a probability of 0.4 that there's track maintenance
- 4:35:33and a probability of 0.6 that there isn't.
- 4:35:36But if there's heavy rain, then here the chance
- 4:35:38that there is track maintenance is 0.1 and the chance
- 4:35:41that there is not track maintenance is 0.9.
- 4:35:44Each of these rows is going to sum up to 1.
- 4:35:47Because each of these represent different values
- 4:35:49of whether or not it's raining, the three possible values
- 4:35:52that that random variable can take on.
- 4:35:54And each is associated with its own probability distribution
- 4:35:57that is ultimately all going to add up to the number 1.
- 4:36:02So that there is our distribution for this random variable called maintenance,
- 4:36:05about whether or not there is maintenance on the train track.
- 4:36:09And now let's consider the next variable.
- 4:36:11Here we have a node inside of our Bayesian network called train
- 4:36:15that has two possible values, on time and delayed.
- 4:36:18And this node is going to be dependent upon the two nodes that
- 4:36:21are pointing towards it, that whether or not
- 4:36:23the train is on time or delayed depends on whether or not
- 4:36:27there is track maintenance.
- 4:36:28And it depends on whether or not there is rain,
- 4:36:30that heavier rain probably means more likely that my train is delayed.
- 4:36:35And if there is track maintenance, that also probably
- 4:36:38means it's more likely that my train is delayed as well.
- 4:36:41And so you could construct a larger probability distribution,
- 4:36:45a conditional probability distribution, that instead
- 4:36:47of conditioning on just one variable, as was the case here,
- 4:36:51is now conditioning on two variables, conditioning
- 4:36:54both on rain represented by r and on maintenance represented by yes.
- 4:36:58Again, each of these rows has two values that sum up to the number 1,
- 4:37:02one for whether the train is on time, one for whether the train is delayed.
- 4:37:06And here I can say something like, all right,
- 4:37:08if I know there was light rain and track maintenance, well, OK,
- 4:37:12that would be r is light and m is yes.
- 4:37:16Well, then there is a probability of 0.6 that my train is on time,
- 4:37:19and a probability of 0.4 the train is delayed.
- 4:37:23And you can imagine gathering this data just
- 4:37:25by looking at real world data, looking at data about, all right,
- 4:37:28if I knew that it was light rain and there was track maintenance,
- 4:37:31how often was a train delayed or not delayed?
- 4:37:33And you could begin to construct this thing.
- 4:37:35The interesting thing is intelligently, being
- 4:37:37able to try to figure out how might you go about ordering these things,
- 4:37:40what things might influence other nodes inside of this Bayesian network.
- 4:37:46And the last thing I care about is whether or not I make it to my appointment.
- 4:37:50So did I attend or miss the appointment?
- 4:37:52And ultimately, whether I attend or miss the appointment,
- 4:37:55it is influenced by track maintenance, because it's indirectly this idea that,
- 4:37:59all right, if there is track maintenance,
- 4:38:01well, then my train might more likely be delayed.
- 4:38:02And if my train is more likely to be delayed,
- 4:38:04then I'm more likely to miss my appointment.
- 4:38:06But what we encode in this Bayesian network
- 4:38:09are just what we might consider to be more direct relationships.
- 4:38:12So the train has a direct influence on the appointment.
- 4:38:15And given that I know whether the train is on time or delayed,
- 4:38:18knowing whether there's track maintenance isn't
- 4:38:20going to give me any additional information that I didn't already have.
- 4:38:24That if I know train, these other nodes that are up above
- 4:38:27isn't really going to influence the result.
- 4:38:30And so here we might represent it using another conditional probability
- 4:38:34distribution that looks a little something like this.
- 4:38:36The train can take on two possible values.
- 4:38:39Either my train is on time or my train is delayed.
- 4:38:42And for each of those two possible values,
- 4:38:44I have a distribution for what are the odds that I'm
- 4:38:46able to attend the meeting and what are the odds that I missed the meeting.
- 4:38:49And obviously, if my train is on time, I'm
- 4:38:51much more likely to be able to attend the meeting
- 4:38:53than if my train is delayed, in which case I'm more likely to miss that
- 4:38:57meeting.
- 4:38:59So all of these nodes put all together here represent this Bayesian network,
- 4:39:03this network of random variables whose values I ultimately care about,
- 4:39:07and that have some sort of relationship between them,
- 4:39:09some sort of dependence where these arrows from one node to another
- 4:39:13indicate some dependence, that I can calculate
- 4:39:15the probability of some node given the parents that happen to exist there.
- 4:39:21So now that we've been able to describe the structure of this Bayesian
- 4:39:24network and the relationships between each of these nodes
- 4:39:27by associating each of the nodes in the network with a probability
- 4:39:30distribution, whether that's an unconditional probability distribution
- 4:39:34in the case of this root node here, like rain,
- 4:39:36and a conditional probability distribution in the case
- 4:39:39of all of the other nodes whose probabilities are
- 4:39:42dependent upon the values of their parents,
- 4:39:44we can begin to do some computation and calculation using
- 4:39:47the information inside of that table.
- 4:39:50So let's imagine, for example, that I just
- 4:39:51wanted to compute something simple like the probability of light rain.
- 4:39:55How would I get the probability of light rain?
- 4:39:57Well, light rain, rain here is a root node.
- 4:40:01And so if I wanted to calculate that probability,
- 4:40:03I could just look at the probability distribution for rain
- 4:40:06and extract from it the probability of light rains, just a single value
- 4:40:10that I already have access to.
- 4:40:12But we could also imagine wanting to compute more complex joint
- 4:40:16probabilities, like the probability that there is light rain and also
- 4:40:21no track maintenance.
- 4:40:22This is a joint probability of two values, light rain and no track
- 4:40:27maintenance.
- 4:40:27And the way I might do that is first by starting by saying, all right,
- 4:40:30well, let me get the probability of light rain.
- 4:40:33But now I also want the probability of no track maintenance.
- 4:40:36But of course, this node is dependent upon the value of rain.
- 4:40:41So what I really want is the probability of no track maintenance,
- 4:40:44given that I know that there was light rain.
- 4:40:47And so the expression for calculating this idea that the probability of light
- 4:40:51rain and no track maintenance is really just the probability of light rain
- 4:40:56and the probability that there is no track maintenance,
- 4:40:58given that I know that there already is light rain.
- 4:41:01So I take the unconditional probability of light rain,
- 4:41:05multiply it by the conditional probability of no track maintenance,
- 4:41:09given that I know there is light rain.
- 4:41:12And you can continue to do this again and again for every variable
- 4:41:15that you want to add into this joint probability
- 4:41:18that I might want to calculate.
- 4:41:19If I wanted to know the probability of light rain and no track maintenance
- 4:41:23and a delayed train, well, that's going to be the probability of light rain,
- 4:41:27multiplied by the probability of no track maintenance, given light rain,
- 4:41:31multiplied by the probability of a delayed train, given light rain
- 4:41:36and no track maintenance.
- 4:41:37Because whether the train is on time or delayed
- 4:41:39is dependent upon both of these other two variables.
- 4:41:42And so I have two pieces of evidence that go
- 4:41:45into the calculation of that conditional probability.
- 4:41:48And each of these three values is just a value
- 4:41:51that I can look up by looking at one of these individual probability
- 4:41:55distributions that is encoded into my Bayesian network.
- 4:41:59And if I wanted a joint probability over all four of the variables,
- 4:42:03something like the probability of light rain and no track maintenance
- 4:42:06and a delayed train and I miss my appointment,
- 4:42:09well, that's going to be multiplying four different values, one
- 4:42:12from each of these individual nodes.
- 4:42:14It's going to be the probability of light rain,
- 4:42:16then of no track maintenance given light rain, then of a delayed train,
- 4:42:20given light rain and no track maintenance.
- 4:42:22And then finally, for this node here, for whether I
- 4:42:25make it to my appointment or not, it's not
- 4:42:26dependent upon these two variables, given
- 4:42:29that I know whether or not the train is on time.
- 4:42:31I only need to care about the conditional probability
- 4:42:34that I miss my train, or that I miss my appointment,
- 4:42:37given that the train happens to be delayed.
- 4:42:39And so that's represented here by four probabilities, each of which
- 4:42:43is located inside of one of these probability distributions
- 4:42:47for each of the nodes, all multiplied together.
- 4:42:50And so I can take a variable like that and figure out
- 4:42:52what the joint probability is by multiplying
- 4:42:55a whole bunch of these individual probabilities from the Bayesian network.
- 4:42:59But of course, just as with last time, where what I really wanted to do
- 4:43:02was to be able to get new pieces of information,
- 4:43:05here, too, this is what we're going to want to do with our Bayesian network.
- 4:43:08In the context of knowledge, we talked about the problem of inference.
- 4:43:11Given things that I know to be true, can I draw conclusions,
- 4:43:14make deductions about other facts about the world that I also know to be true?
- 4:43:19And what we're going to do now is apply the same sort of idea to probability.
- 4:43:23Using information about which I have some knowledge,
- 4:43:26whether some evidence or some probabilities,
- 4:43:28can I figure out not other variables for certain,
- 4:43:32but can I figure out the probabilities of other variables
- 4:43:35taking on particular values?
- 4:43:36And so here, we introduce the problem of inference in a probabilistic setting,
- 4:43:41in a case where variables might not necessarily be true for sure,
- 4:43:44but they might be random variables that take on different values
- 4:43:48with some probability.
- 4:43:50So how do we formally define what exactly this inference problem actually
- 4:43:53is?
- 4:43:54Well, the inference problem has a couple of parts to it.
- 4:43:57We have some query, some variable x that we
- 4:43:59want to compute the distribution for.
- 4:44:01Maybe I want the probability that I miss my train,
- 4:44:04or I want the probability that there is track maintenance,
- 4:44:08something that I want information about.
- 4:44:11And then I have some evidence variables.
- 4:44:13Maybe it's just one piece of evidence.
- 4:44:14Maybe it's multiple pieces of evidence.
- 4:44:16But I've observed certain variables for some sort of event.
- 4:44:20So for example, I might have observed that it is raining.
- 4:44:23This is evidence that I have.
- 4:44:24I know that there is light rain, or I know that there is heavy rain.
- 4:44:27And that is evidence I have.
- 4:44:28And using that evidence, I want to know what is the probability
- 4:44:32that my train is delayed, for example.
- 4:44:34And that is a query that I might want to ask based on this evidence.
- 4:44:38So I have a query, some variable.
- 4:44:39Evidence, which are some other variables that I
- 4:44:41have observed inside of my Bayesian network.
- 4:44:44And of course, that does leave some hidden variables.
- 4:44:46Why?
- 4:44:47These are variables that are not evidence variables and not query variables.
- 4:44:52So you might imagine in the case where I know whether or not it's raining,
- 4:44:55and I want to know whether my train is going to be delayed or not,
- 4:44:59the hidden variable, the thing I don't have access to,
- 4:45:02is something like, is there maintenance on the track?
- 4:45:04Or am I going to make or not make my appointment, for example?
- 4:45:07These are variables that I don't have access to.
- 4:45:09They're hidden because they're not things I observed,
- 4:45:12and they're also not the query, the thing that I'm asking.
- 4:45:14And so ultimately, what we want to calculate
- 4:45:17is I want to know the probability distribution of x given
- 4:45:21e, the event that I observed.
- 4:45:22So given that I observed some event, I observed that it is raining,
- 4:45:25I would like to know what is the distribution over the possible values
- 4:45:29of the train random variable.
- 4:45:31Is it on time?
- 4:45:32Is it delayed?
- 4:45:33What's the likelihood it's going to be there?
- 4:45:35And it turns out we can do this calculation just
- 4:45:37using a lot of the probability rules that we've already seen in action.
- 4:45:42And ultimately, we're going to take a look at the math
- 4:45:44at a little bit of a high level, at an abstract level.
- 4:45:46But ultimately, we can allow computers and programming libraries
- 4:45:49that already exist to begin to do some of this math for us.
- 4:45:52But it's good to get a general sense for what's actually happening
- 4:45:55when this inference process takes place.
- 4:45:57Let's imagine, for example, that I want to compute the probability
- 4:46:00distribution of the appointment random variable given some evidence,
- 4:46:05given that I know that there was light rain
- 4:46:07and no track maintenance.
- 4:46:08So there's my evidence, these two variables that I observe the values of.
- 4:46:12I observe the value of rain.
- 4:46:14I know there's light rain.
- 4:46:15And I know that there is no track maintenance going on today.
- 4:46:18And what I care about knowing, my query, is this random variable appointment.
- 4:46:22I want to know the distribution of this random variable appointment,
- 4:46:25like what is the chance that I'm able to attend my appointment?
- 4:46:28What is the chance that I miss my appointment given this evidence?
- 4:46:32And the hidden variable, the information that I don't have access to,
- 4:46:35is this variable train.
- 4:46:36This is information that is not part of the evidence
- 4:46:38that I see, not something that I observe.
- 4:46:41But it is also not the query that I'm asking for.
- 4:46:44And so what might this inference procedure look like?
- 4:46:47Well, if you recall back from when we were defining conditional probability
- 4:46:50and doing math with conditional probabilities,
- 4:46:52we know that a conditional probability is proportional to the joint
- 4:46:57probability.
- 4:46:58And we remembered this by recalling that the probability of A given
- 4:47:01B is just some constant factor alpha multiplied by the probability of A
- 4:47:06and B. That constant factor alpha turns out
- 4:47:08to be like dividing over the probability of B.
- 4:47:10But the important thing is that it's just some constant multiplied
- 4:47:14by the joint distribution, the probability
- 4:47:17that all of these individual things happen.
- 4:47:19So in this case, I can take the probability of the appointment random
- 4:47:23variable given light rain and no track maintenance
- 4:47:27and say that is just going to be proportional, some constant alpha,
- 4:47:30multiplied by the joint probability, the probability
- 4:47:33of a particular value for the appointment random variable
- 4:47:36and light rain and no track maintenance.
- 4:47:40Well, all right, how do I calculate this, probability of appointment
- 4:47:43and light rain and no track maintenance, when what I really care about
- 4:47:46is knowing I need all four of these values
- 4:47:48to be able to calculate a joint distribution across everything
- 4:47:52because in a particular appointment depends upon the value of train?
- 4:47:56Well, in order to do that, here I can begin to use that marginalization
- 4:47:59trick, that there are only two ways I can get
- 4:48:02any configuration of an appointment, light rain, and no track maintenance.
- 4:48:05Either this particular setting of variables
- 4:48:07happens and the train is on time, or this particular setting of variables
- 4:48:12happens and the train is delayed.
- 4:48:13Those are two possible cases that I would want to consider.
- 4:48:17And if I add those two cases up, well, then I
- 4:48:19get the result just by adding up all of the possibilities
- 4:48:23for the hidden variable or variables that there are multiple.
- 4:48:26But since there's only one hidden variable here, train, all I need to do
- 4:48:30is iterate over all the possible values for that hidden variable train
- 4:48:34and add up their probabilities.
- 4:48:36So this probability expression here becomes probability distribution
- 4:48:40over appointment, light, no rain, and train is on time,
- 4:48:44and the probability distribution over the appointment, light rain,
- 4:48:47no track maintenance, and that the train is delayed, for example.
- 4:48:51So I take both of the possible values for train, go ahead and add them up.
- 4:48:55These are just joint probabilities that we saw earlier,
- 4:48:57how to calculate just by going parent, parent, parent, parent,
- 4:48:59and calculating those probabilities and multiplying them together.
- 4:49:03And then you'll need to normalize them at the end,
- 4:49:05speaking at a high level, to make sure that everything adds up to the number 1.
- 4:49:09So the formula for how you do this in a process known as inference by enumeration
- 4:49:13looks a little bit complicated, but ultimately it looks like this.
- 4:49:16And let's now try to distill what it is that all of these symbols actually mean.
- 4:49:20Let's start here.
- 4:49:21What I care about knowing is the probability of x, my query variable,
- 4:49:25given some sort of evidence.
- 4:49:28What do I know about conditional probabilities?
- 4:49:30Well, a conditional probability is proportional to the joint probability.
- 4:49:34So it is some alpha, some normalizing constant,
- 4:49:37multiplied by this joint probability of x and evidence.
- 4:49:41And how do I calculate that?
- 4:49:42Well, to do that, I'm going to marginalize
- 4:49:45over all of the hidden variables, all the variables
- 4:49:47that I don't directly observe the values for.
- 4:49:50I'm basically going to iterate over all of the possibilities
- 4:49:53that it could happen and just sum them all up.
- 4:49:55And so I can translate this into a sum over all y,
- 4:49:58which ranges over all the possible hidden variables and the values
- 4:50:02that they could take on, and adds up all of those possible individual
- 4:50:06probabilities.
- 4:50:07And that is going to allow me to do this process of inference by enumeration.
- 4:50:11Now, ultimately, it's pretty annoying if we as humans
- 4:50:14have to do all this math for ourselves.
- 4:50:16But turns out this is where computers and AI can be particularly helpful,
- 4:50:19that we can program a computer to understand a Bayesian network,
- 4:50:22to be able to understand these inference procedures,
- 4:50:25and to be able to do these calculations.
- 4:50:27And using the information you've seen here,
- 4:50:29you could implement a Bayesian network from scratch yourself.
- 4:50:31But turns out there are a lot of libraries, especially written in Python,
- 4:50:34that allow us to make it easier to do this sort of probabilistic inference,
- 4:50:38to be able to take a Bayesian network and do these sorts of calculations,
- 4:50:41so that you don't need to know and understand all of the underlying math,
- 4:50:44though it's helpful to have a general sense for how it works.
- 4:50:46But you just need to be able to describe the structure of the network
- 4:50:49and make queries in order to be able to produce the result.
- 4:50:53And so let's take a look at an example of that right now.
- 4:50:56It turns out that there are a lot of possible libraries
- 4:50:59that exist in Python for doing this sort of inference.
- 4:51:01It doesn't matter too much which specific library you use.
- 4:51:04They all behave in fairly similar ways.
- 4:51:05But the library I'm going to use here is one known as pomegranate.
- 4:51:08And here inside of model.py, I have defined a Bayesian network,
- 4:51:13just using the structure and the syntax that the pomegranate library expects.
- 4:51:17And what I'm effectively doing is just, in Python,
- 4:51:20creating nodes to represent each of the nodes of the Bayesian network
- 4:51:24that you saw me describe a moment ago.
- 4:51:26So here on line four, after I've imported pomegranate,
- 4:51:29I'm defining a variable called rain that is going
- 4:51:31to represent a node inside of my Bayesian network.
- 4:51:35It's going to be a node that follows this distribution, where
- 4:51:39there are three possible values, none for no rain, light for light rain,
- 4:51:42heavy for heavy rain.
- 4:51:43And these are the probabilities of each of those taking place.
- 4:51:460.7 is the likelihood of no rain, 0.2 for light rain, 0.1 for heavy rain.
- 4:51:53Then after that, we go to the next variable,
- 4:51:55the variable for track maintenance, for example,
- 4:51:57which is dependent upon that rain variable.
- 4:52:00And this, instead of being an unconditional distribution,
- 4:52:03is a conditional distribution, as indicated
- 4:52:05by a conditional probability table here.
- 4:52:07And the idea is that I'm following this is conditional
- 4:52:11on the distribution of rain.
- 4:52:13So if there is no rain, then the chance that there is, yes, track maintenance
- 4:52:17is 0.4.
- 4:52:17If there's no rain, the chance that there is no track maintenance is 0.6.
- 4:52:21Likewise, for light rain, I have a distribution.
- 4:52:23For heavy rain, I have a distribution as well.
- 4:52:25But I'm effectively encoding the same information
- 4:52:27you saw represented graphically a moment ago.
- 4:52:29But I'm telling this Python program that the maintenance node
- 4:52:33obeys this particular conditional probability distribution.
- 4:52:37And we do the same thing for the other random variables as well.
- 4:52:40Train was a node inside my distribution that
- 4:52:44was a conditional probability table with two parents.
- 4:52:47It was dependent not only on rain, but also on track maintenance.
- 4:52:51And so here I'm saying something like, given
- 4:52:53that there is no rain and, yes, track maintenance,
- 4:52:55the probability that my train is on time is 0.8.
- 4:52:59And the probability that it's delayed is 0.2.
- 4:53:01And likewise, I can do the same thing for all
- 4:53:03of the other possible values of the parents of the train node
- 4:53:07inside of my Bayesian network by saying, for all of those possible values,
- 4:53:12here is the distribution that the train node should follow.
- 4:53:16Then I do the same thing for an appointment
- 4:53:18based on the distribution of the variable train.
- 4:53:21Then at the end, what I do is actually construct this network
- 4:53:24by describing what the states of the network are
- 4:53:27and by adding edges between the dependent nodes.
- 4:53:30So I create a new Bayesian network, add states to it, one for rain,
- 4:53:33one for maintenance, one for the train, one for the appointment.
- 4:53:36And then I add edges connecting the related pieces.
- 4:53:40Rain has an arrow to maintenance because rain influences track maintenance.
- 4:53:44Rain also influences the train.
- 4:53:46Maintenance also influences the train.
- 4:53:48And train influences whether I make it to my appointment
- 4:53:50and bake just finalizes the model and does some additional computation.
- 4:53:54So the specific syntax of this is not really the important part.
- 4:53:57Pomegranate just happens to be one of several different libraries
- 4:54:00that can all be used for similar purposes.
- 4:54:02And you could describe and define a library for yourself
- 4:54:05that implemented similar things.
- 4:54:07But the key idea here is that someone can design a library
- 4:54:11for a general Bayesian network that has nodes that are based upon its parents.
- 4:54:15And then all a programmer needs to do using one of those libraries
- 4:54:18is to define what those nodes and what those probability distributions are.
- 4:54:23And we can begin to do some interesting logic based on it.
- 4:54:26So let's try doing that conditional or joint probability calculation
- 4:54:30that we saw us do by hand before by going into likelihood.py, where
- 4:54:36here I'm importing the model that I just defined a moment ago.
- 4:54:40And here I'd just like to calculate model.probability, which
- 4:54:42calculates the probability for a given observation.
- 4:54:46And I'd like to calculate the probability of no rain, no track maintenance,
- 4:54:51my train is on time, and I'm able to attend the meeting.
- 4:54:54So sort of the optimal scenario that there is no rain and no maintenance
- 4:54:58on the track, my train is on time, and I'm able to attend the meeting.
- 4:55:01What is the probability that all of that actually happens?
- 4:55:04And I can calculate that using the library and just print out its probability.
- 4:55:08And so I'll go ahead and run python of likelihood.py.
- 4:55:12And I see that, OK, the probability is about 0.34.
- 4:55:16So about a third of the time, everything goes right for me in this case.
- 4:55:20No rain, no track maintenance, train is on time,
- 4:55:22and I'm able to attend the meeting.
- 4:55:24But I could experiment with this, try and calculate other probabilities as well.
- 4:55:28What's the probability that everything goes right up until the train,
- 4:55:31but I still miss my meeting?
- 4:55:33So no rain, no track maintenance, train is on time,
- 4:55:37but I miss the appointment.
- 4:55:39Let's calculate that probability.
- 4:55:41And all right, that has a probability of about 0.04.
- 4:55:44So about 4% of the time, the train will be on time,
- 4:55:47there won't be any rain, no track maintenance,
- 4:55:49and yet I'll still miss the meeting.
- 4:55:52And so this is really just an implementation
- 4:55:54of the calculation of the joint probabilities that we did before.
- 4:55:57What this library is likely doing is first figuring out
- 4:56:00the probability of no rain, then figuring out
- 4:56:03the probability of no track maintenance given no rain,
- 4:56:06then the probability that my train is on time given both of these values,
- 4:56:10and then the probability that I miss my appointment given that I
- 4:56:13know that the train was on time.
- 4:56:15So this, again, is the calculation of that joint probability.
- 4:56:18And turns out we can also begin to have our computer solve inference problems
- 4:56:22as well, to begin to infer, based on information, evidence that we see,
- 4:56:26what is the likelihood of other variables also being true.
- 4:56:30So let's go into inference.py, for example.
- 4:56:33We're here, I'm again importing that exact same model from before,
- 4:56:36importing all the nodes and all the edges
- 4:56:38and the probability distribution that is encoded there as well.
- 4:56:42And now there's a function for doing some sort of prediction.
- 4:56:45And here, into this model, I pass in the evidence that I observe.
- 4:56:50So here, I've encoded into this Python program the evidence
- 4:56:54that I have observed.
- 4:56:55I have observed the fact that the train is delayed.
- 4:56:58And that is the value for one of the four random variables
- 4:57:01inside of this Bayesian network.
- 4:57:03And using that information, I would like to be able to draw inspiration
- 4:57:07and figure out inferences about the values
- 4:57:09of the other random variables that are inside of my Bayesian network.
- 4:57:13I would like to make predictions about everything else.
- 4:57:15So all of the actual computational logic is happening in just these three lines,
- 4:57:19where I'm making this call to this prediction.
- 4:57:21Down below, I'm just iterating over all of the states and all the predictions
- 4:57:25and just printing them out so that we can visually see what the results are.
- 4:57:29But let's find out, given the train is delayed,
- 4:57:31what can I predict about the values of the other random variables?
- 4:57:35Let's go ahead and run python inference.py.
- 4:57:38I run that, and all right, here is the result that I get.
- 4:57:41Given the fact that I know that the train is delayed,
- 4:57:44this is evidence that I have observed.
- 4:57:46Well, given that there is a 45% chance or a 46% chance
- 4:57:50that there was no rain, a 31% chance there was light rain,
- 4:57:52a 23% chance there was heavy rain, I can see a probability distribution
- 4:57:56of a track maintenance and a probability distribution
- 4:57:58over whether I'm able to attend or miss my appointment.
- 4:58:01Now, we know that whether I attend or miss the appointment,
- 4:58:04that is only dependent upon the train being delayed or not delayed.
- 4:58:07It shouldn't depend on anything else.
- 4:58:10So let's imagine, for example, that I knew that there was heavy rain.
- 4:58:14That shouldn't affect the distribution for making the appointment.
- 4:58:18And indeed, if I go up here and add some evidence,
- 4:58:21say that I know that the value of rain is heavy.
- 4:58:23That is evidence that I now have access to.
- 4:58:25I now have two pieces of evidence.
- 4:58:27I know that the rain is heavy, and I know that my train is delayed.
- 4:58:31I can calculate the probability by running this inference procedure again
- 4:58:35and seeing the result. I know that the rain is heavy.
- 4:58:37I know my train is delayed.
- 4:58:39The probability distribution for track maintenance changed.
- 4:58:42Given that I know that there's heavy rain,
- 4:58:44now it's more likely that there is no track maintenance, 88%,
- 4:58:48as opposed to 64% from here before.
- 4:58:51And now, what is the probability that I make the appointment?
- 4:58:55Well, that's the same as before.
- 4:58:57It's still going to be attend the appointment with probability 0.6,
- 4:59:00missed the appointment with probability 0.4,
- 4:59:03because it was only dependent upon whether or not
- 4:59:05my train was on time or delayed.
- 4:59:07And so this here is implementing that idea of that inference algorithm
- 4:59:11to be able to figure out, based on the evidence that I have,
- 4:59:14what can we infer about the values of the other variables that exist as well.
- 4:59:18So inference by enumeration is one way of doing this inference procedure,
- 4:59:22just looping over all of the values the hidden variables could take on
- 4:59:26and figuring out what the probability is.
- 4:59:29Now, it turns out this is not particularly efficient.
- 4:59:31And there are definitely optimizations you can make by avoiding repeated work.
- 4:59:35If you're calculating the same sort of probability multiple times,
- 4:59:38there are ways of optimizing the program to avoid
- 4:59:40having to recalculate the same probabilities again and again.
- 4:59:44But even then, as the number of variables get large,
- 4:59:47as the number of possible values of variables could take on, get large,
- 4:59:50we're going to start to have to do a lot of computation,
- 4:59:52a lot of calculation, to be able to do this inference.
- 4:59:55And at that point, it might start to get unreasonable,
- 4:59:58in terms of the amount of time that it would take
- 5:00:00to be able to do this sort of exact inference.
- 5:00:04And it's for that reason that oftentimes, when
- 5:00:06it comes towards probability and things we're not entirely sure about,
- 5:00:09we don't always care about doing exact inference
- 5:00:11and knowing exactly what the probability is.
- 5:00:14But if we can approximate the inference procedure,
- 5:00:17do some sort of approximate inference, that that can be pretty good as well.
- 5:00:21That if I don't know the exact probability,
- 5:00:23but I have a general sense for the probability
- 5:00:25that I can get increasingly accurate with more time,
- 5:00:28that that's probably pretty good, especially
- 5:00:30if I can get that to happen even faster.
- 5:00:33So how could I do approximate inference inside of a Bayesian network?
- 5:00:37Well, one method is through a procedure known as sampling.
- 5:00:40In the process of sampling, I'm going to take
- 5:00:42a sample of all of the variables inside of this Bayesian network here.
- 5:00:46And how am I going to sample?
- 5:00:47Well, I'm going to sample one of the values from each of these nodes
- 5:00:51according to their probability distribution.
- 5:00:54So how might I take a sample of all these nodes?
- 5:00:56Well, I'll start at the root.
- 5:00:57I'll start with rain.
- 5:00:58Here's the distribution for rain.
- 5:00:59And I'll go ahead and, using a random number generator or something like it,
- 5:01:03randomly pick one of these three values.
- 5:01:05I'll pick none with probability 0.7, light with probability 0.2,
- 5:01:09and heavy with probability 0.1.
- 5:01:11So I'll randomly just pick one of them according to that distribution.
- 5:01:14And maybe in this case, I pick none, for example.
- 5:01:17Then I do the same thing for the other variable.
- 5:01:19Maintenance also has a probability distribution.
- 5:01:22And I'm going to sample.
- 5:01:23Now, there are three probability distributions here.
- 5:01:26But I'm only going to sample from this first row here,
- 5:01:29because I've observed already in my sample that the value of rain is none.
- 5:01:33So given that rain is none, I'm going to sample from this distribution to say,
- 5:01:37all right, what should the value of maintenance be?
- 5:01:40And in this case, maintenance is going to be, let's just say yes,
- 5:01:42which happens 40% of the time in the event that there is no rain, for example.
- 5:01:47And we'll sample all of the rest of the nodes in this way as well,
- 5:01:50that I want to sample from the train distribution.
- 5:01:52And I'll sample from this first row here, where there is no rain,
- 5:01:56but there is track maintenance.
- 5:01:58And I'll sample 80% of the time.
- 5:02:00I'll say the train is on time.
- 5:02:0120% of the time, I'll say the train is delayed.
- 5:02:04And finally, we'll do the same thing for whether I make it to my appointment
- 5:02:07or not.
- 5:02:07Did I attend or miss the appointment?
- 5:02:09We'll sample based on this distribution and maybe say
- 5:02:11that in this case, I attend the appointment, which
- 5:02:13happens 90% of the time when the train is actually on time.
- 5:02:18So by going through these nodes, I can very quickly just do some sampling
- 5:02:22and get a sample of the possible values that could come up
- 5:02:26from going through this entire Bayesian network
- 5:02:28according to those probability distributions.
- 5:02:31And where this becomes powerful is if I do this not once,
- 5:02:34but I do this thousands or tens of thousands of times
- 5:02:36and generate a whole bunch of samples all using this distribution.
- 5:02:39I get different samples.
- 5:02:41Maybe some of them are the same.
- 5:02:42But I get a value for each of the possible variables that could come up.
- 5:02:47And so then if I'm ever faced with a question,
- 5:02:49a question like, what is the probability that the train is on time,
- 5:02:53you could do an exact inference procedure.
- 5:02:55This is no different than the inference problem we had before
- 5:02:58where I could just marginalize, look at all the possible other values
- 5:03:01of the variables, and do the computation of inference by enumeration
- 5:03:05to find out this probability exactly.
- 5:03:07But I could also, if I don't care about the exact probability,
- 5:03:10just sample it, approximate it to get close.
- 5:03:12And this is a powerful tool in AI where we don't need to be right 100%
- 5:03:16of the time or we don't need to be exactly right.
- 5:03:18If we just need to be right with some probability,
- 5:03:20we can often do so more effectively, more efficiently.
- 5:03:23And so if here now are all of those possible samples,
- 5:03:26I'll highlight the ones where the train is on time.
- 5:03:30I'm ignoring the ones where the train is delayed.
- 5:03:32And in this case, there's like six out of eight of the samples
- 5:03:35have the train is arriving on time.
- 5:03:37And so maybe in this case, I can say that in six out of eight cases,
- 5:03:40that's the likelihood that the train is on time.
- 5:03:43And with eight samples, that might not be a great prediction.
- 5:03:45But if I had thousands upon thousands of samples,
- 5:03:48then this could be a much better inference procedure
- 5:03:51to be able to do these sorts of calculations.
- 5:03:53So this is a direct sampling method to just do a bunch of samples
- 5:03:56and then figure out what the probability of some event is.
- 5:04:00Now, this from before was an unconditional probability.
- 5:04:03What is the probability that the train is on time?
- 5:04:07And I did that by looking at all the samples and figuring out, right,
- 5:04:09here are the ones where the train is on time.
- 5:04:12But sometimes what I want to calculate is not an unconditional probability,
- 5:04:16but rather a conditional probability, something
- 5:04:18like what is the probability that there is light rain,
- 5:04:21given that the train is on time, something to that effect.
- 5:04:24And to do that kind of calculation, well, what I might do
- 5:04:28is here are all the samples that I have.
- 5:04:31And I want to calculate a probability distribution,
- 5:04:33given that I know that the train is on time.
- 5:04:36So to be able to do that, I can kind of look
- 5:04:38at the two cases where the train was delayed and ignore or reject them,
- 5:04:43sort of exclude them from the possible samples that I'm considering.
- 5:04:47And now I want to look at these remaining cases where the train is on time.
- 5:04:50Here are the cases where there is light rain.
- 5:04:53And I say, OK, these are two out of the six possible cases.
- 5:04:56That can give me an approximation for the probability of light rain,
- 5:05:00given the fact that I know the train was on time.
- 5:05:03And I did that in almost exactly the same way,
- 5:05:05just by adding an additional step, by saying that, all right,
- 5:05:08when I take each sample, let me reject all of the samples that
- 5:05:12don't match my evidence and only consider
- 5:05:14the samples that do match what it is that I have in my evidence
- 5:05:19that I want to make some sort of calculation about.
- 5:05:21And it turns out, using the libraries that we've had for Bayesian networks,
- 5:05:25we can begin to implement this same sort of idea,
- 5:05:28like implement rejection sampling, which is what this method is called,
- 5:05:31to be able to figure out some probability, not via direct inference,
- 5:05:35but instead by sampling.
- 5:05:37So what I have here is a program called sample.py.
- 5:05:39Imports the exact same model.
- 5:05:41And what I define first is a program to generate a sample.
- 5:05:45And the way I generate a sample is just by looping over all of the states.
- 5:05:48The states need to be in some sort of order
- 5:05:50to make sure I'm looping in the correct order.
- 5:05:52But effectively, if it is a conditional distribution,
- 5:05:55I'm going to sample based on the parents.
- 5:05:58And otherwise, I'm just going to directly sample
- 5:06:00the variable, like rain, which has no parents.
- 5:06:02It's just an unconditional distribution and keep
- 5:06:05track of all those parent samples and return the final sample.
- 5:06:08The exact syntax of this, again, not particularly important.
- 5:06:11It just happens to be part of the implementation details
- 5:06:13of this particular library.
- 5:06:15The interesting logic is down below.
- 5:06:17Now that I have the ability to generate a sample,
- 5:06:20if I want to know the distribution of the appointment random variable,
- 5:06:24given that the train is delayed, well, then I
- 5:06:26can begin to do calculations like this.
- 5:06:28Let me take 10,000 samples and assemble all my results
- 5:06:32in this list called data.
- 5:06:33I'll go ahead and loop n times, in this case, 10,000 times.
- 5:06:36I'll generate a sample.
- 5:06:38And I want to know the distribution of appointment,
- 5:06:41given that the train is delayed.
- 5:06:43So according to rejection sampling, I'm only
- 5:06:45going to consider samples where the train is delayed.
- 5:06:47If the train is not delayed, I'm not going to consider those values at all.
- 5:06:51So I'm going to say, all right, if I take the sample,
- 5:06:53look at the value of the train random variable, if the train is delayed,
- 5:06:57well, let me go ahead and add to my data
- 5:06:59that I'm collecting the value of the appointment random variable
- 5:07:02that it took on in this particular sample.
- 5:07:05So I'm only considering the samples where the train is delayed.
- 5:07:08And for each of those samples, considering what the value of appointment
- 5:07:11is, and then at the end, I'm using a Python class called
- 5:07:14counter, which quickly counts up all the values inside of a data set.
- 5:07:18So I can take this list of data and figure out
- 5:07:20how many times was my appointment made and how many times was my appointment
- 5:07:25missed.
- 5:07:27And so this here, with just a couple lines of code,
- 5:07:29is an implementation of rejection sampling.
- 5:07:32And I can run it by going ahead and running Python sample.py.
- 5:07:37And when I do that, here is the result I get.
- 5:07:39This is the result of the counter.
- 5:07:411,251 times, I was able to attend the meeting.
- 5:07:45And 856 times, I was able to miss the meeting.
- 5:07:48And you can imagine, by doing more and more samples,
- 5:07:51I'll be able to get a better and better, more accurate result.
- 5:07:54And this is a randomized process.
- 5:07:55It's going to be an approximation of the probability.
- 5:07:58If I run it a different time, you'll notice the numbers are similar, 12,
- 5:08:0172, and 905.
- 5:08:03But they're not identical because there's some randomization, some likelihood
- 5:08:07that things might be higher or lower.
- 5:08:09And so this is why we generally want to try and use more samples so that we
- 5:08:12can have a greater amount of confidence in our result,
- 5:08:15be more sure about the result that we're getting of whether or not
- 5:08:18it accurately reflects or represents the actual underlying probabilities that
- 5:08:23are inherent inside of this distribution.
- 5:08:26And so this, then, was an instance of rejection sampling.
- 5:08:29And it turns out there are a number of other sampling methods
- 5:08:32that you could use to begin to try to sample.
- 5:08:34One problem that rejection sampling has is
- 5:08:37that if the evidence you're looking for is a fairly unlikely event,
- 5:08:41well, you're going to be rejecting a lot of samples.
- 5:08:44Like if I'm looking for the probability of x given some evidence e,
- 5:08:48if e is very unlikely to occur, like occurs maybe one every 1,000 times,
- 5:08:52then I'm only going to be considering 1 out of every 1,000 samples that I do,
- 5:08:56which is a pretty inefficient method for trying to do this sort of calculation.
- 5:08:59I'm throwing away a lot of samples.
- 5:09:01And it takes computational effort to be able to generate those samples.
- 5:09:05So I'd like to not have to do something like that.
- 5:09:07So there are other sampling methods that can try and address this.
- 5:09:09One such sampling method is called likelihood weighting.
- 5:09:13In likelihood weighting, we follow a slightly different procedure.
- 5:09:16And the goal is to avoid needing to throw out samples
- 5:09:20that didn't match the evidence.
- 5:09:22And so what we'll do is we'll start by fixing the values for the evidence
- 5:09:26variables.
- 5:09:26Rather than sample everything, we're going
- 5:09:29to fix the values of the evidence variables and not sample those.
- 5:09:33Then we're going to sample all the other non-evidence variables
- 5:09:36in the same way, just using the Bayesian network looking
- 5:09:38at the probability distributions, sampling all the non-evidence variables.
- 5:09:43But then what we need to do is weight each sample by its likelihood.
- 5:09:48If our evidence is really unlikely, we want
- 5:09:50to make sure that we've taken into account how likely was the evidence
- 5:09:53to actually show up in the sample.
- 5:09:55If I have a sample where the evidence was much more
- 5:09:58likely to show up than another sample, then I
- 5:10:00want to weight the more likely one higher.
- 5:10:02So we're going to weight each sample by its likelihood, where likelihood is just
- 5:10:06defined as the probability of all the evidence.
- 5:10:09Given all the evidence we have, what is the probability
- 5:10:11that it would happen in that particular sample?
- 5:10:14So before, all of our samples were weighted equally.
- 5:10:16They all had a weight of 1 when we were calculating
- 5:10:19the overall average.
- 5:10:20In this case, we're going to weight each sample,
- 5:10:22multiply each sample by its likelihood in order
- 5:10:25to get the more accurate distribution.
- 5:10:28So what would this look like?
- 5:10:30Well, if I ask the same question, what is the probability of light rain,
- 5:10:33given that the train is on time, when I do the sampling procedure
- 5:10:36and start by trying to sample, I'm going to start by fixing the evidence
- 5:10:40variable.
- 5:10:41I'm already going to have in my sample the train is on time.
- 5:10:44That way, I don't have to throw out anything.
- 5:10:46I'm only sampling things where I know the value of the variables that
- 5:10:50are my evidence are what I expect them to be.
- 5:10:53So I'll go ahead and sample from rain.
- 5:10:55And maybe this time, I sample light rain instead of no rain.
- 5:10:58Then I'll sample from track maintenance and say,
- 5:11:00maybe, yes, there's track maintenance.
- 5:11:01Then for train, well, I've already fixed it in place.
- 5:11:04Train was an evidence variable.
- 5:11:06So I'm not going to bother sampling again.
- 5:11:09I'll just go ahead and move on.
- 5:11:10I'll move on to appointment and go ahead and sample from appointment as well.
- 5:11:14So now I've generated a sample.
- 5:11:16I've generated a sample by fixing this evidence variable
- 5:11:19and sampling the other three.
- 5:11:22And the last step is now weighting the sample.
- 5:11:24How much weight should it have?
- 5:11:25And the weight is based on how probable is it
- 5:11:28that the train was actually on time, this evidence actually happened,
- 5:11:32given the values of these other variables, light rain and the fact
- 5:11:35that, yes, there was track maintenance.
- 5:11:37Well, to do that, I can just go back to the train variable
- 5:11:39and say, all right, if there was light rain and track maintenance,
- 5:11:43the likelihood of my evidence, the likelihood that my train was on time,
- 5:11:46is 0.6.
- 5:11:48And so this particular sample would have a weight of 0.6.
- 5:11:52And I could repeat the sampling procedure again and again.
- 5:11:55Each time every sample would be given a weight
- 5:11:57according to the probability of the evidence that I see associated with it.
- 5:12:02And there are other sampling methods that exist as well,
- 5:12:04but all of them are designed to try and get it the same idea,
- 5:12:07to approximate the inference procedure of figuring out the value of a variable.
- 5:12:13So we've now dealt with probability as it
- 5:12:15pertains to particular variables that have these discrete values.
- 5:12:18But what we haven't really considered is how values might change over time.
- 5:12:22That we've considered something like a variable for rain,
- 5:12:25where rain can take on values of none or light rain or heavy rain.
- 5:12:28But in practice, usually when we consider values for variables like rain,
- 5:12:32we like to consider it for over time, how do the values of these variables
- 5:12:37change?
- 5:12:37What do we do with when we're dealing with uncertainty
- 5:12:40over a period of time, which can come up in the context of weather,
- 5:12:43for example, if I have sunny days and I have rainy days.
- 5:12:46And I'd like to know not just what is the probability that it's raining now,
- 5:12:51but what is the probability that it rains tomorrow,
- 5:12:53or the day after that, or the day after that.
- 5:12:55And so to do this, we're going to introduce
- 5:12:57a slightly different kind of model.
- 5:12:58But here, we're going to have a random variable, not just one for the weather,
- 5:13:02but for every possible time step.
- 5:13:05And you can define time step however you like.
- 5:13:07A simple way is just to use days as your time step.
- 5:13:10And so we can define a variable called x sub t, which
- 5:13:13is going to be the weather at time t.
- 5:13:16So x sub 0 might be the weather on day 0.
- 5:13:19x sub 1 might be the weather on day 1, so on and so forth.
- 5:13:22x sub 2 is the weather on day 2.
- 5:13:24But as you can imagine, if we start to do this
- 5:13:26over longer and longer periods of time, there's
- 5:13:28an incredible amount of data that might go into this.
- 5:13:30If you're keeping track of data about the weather for a year,
- 5:13:33now suddenly you might be trying to predict the weather tomorrow,
- 5:13:36given 365 days of previous pieces of evidence.
- 5:13:40And that's a lot of evidence to have to deal with and manipulate and calculate.
- 5:13:43Probably nobody knows what the exact conditional probability distribution
- 5:13:47is for all of those combinations of variables.
- 5:13:49And so when we're trying to do this inference inside of a computer,
- 5:13:52when we're trying to reasonably do this sort of analysis,
- 5:13:56it's helpful to make some simplifying assumptions,
- 5:13:58some assumptions about the problem that we can just assume are true,
- 5:14:01to make our lives a little bit easier.
- 5:14:03Even if they're not totally accurate assumptions,
- 5:14:05if they're close to accurate or approximate, they're usually pretty good.
- 5:14:09And the assumption we're going to make is called the Markov assumption, which
- 5:14:13is the assumption that the current state depends only
- 5:14:16on a finite fixed number of previous states.
- 5:14:19So the current day's weather depends not on all the previous day's weather
- 5:14:23for the rest of all of history, but the current day's weather
- 5:14:26I can predict just based on yesterday's weather,
- 5:14:29or just based on the last two days weather, or the last three days weather.
- 5:14:32But oftentimes, we're going to deal with just the one previous state
- 5:14:36that helps to predict this current state.
- 5:14:39And by putting a whole bunch of these random variables together,
- 5:14:42using this Markov assumption, we can create what's called a Markov chain,
- 5:14:46where a Markov chain is just some sequence of random variables
- 5:14:49where each of the variables distribution follows that Markov assumption.
- 5:14:53And so we'll do an example of this where the Markov assumption is,
- 5:14:56I can predict the weather.
- 5:14:57Is it sunny or rainy?
- 5:14:58And we'll just consider those two possibilities for now,
- 5:15:01even though there are other types of weather.
- 5:15:02But I can predict each day's weather just on the prior day's weather,
- 5:15:06using today's weather, I can come up with a probability distribution
- 5:15:10for tomorrow's weather.
- 5:15:11And here's what this weather might look like.
- 5:15:13It's formatted in terms of a matrix, as you might describe it,
- 5:15:16as rows and columns of values, where on the left-hand side,
- 5:15:21I have today's weather, represented by the variable x sub t.
- 5:15:25And over here in the columns, I have tomorrow's weather,
- 5:15:28represented by the variable x sub t plus 1, t plus 1 day's weather instead.
- 5:15:34And what this matrix is saying is, if today is sunny,
- 5:15:38well, then it's more likely than not that tomorrow is also sunny.
- 5:15:42Oftentimes, the weather stays consistent for multiple days in a row.
- 5:15:45And for example, let's say that if today is sunny,
- 5:15:47our model says that tomorrow, with probability 0.8, it will also be sunny.
- 5:15:52And with probability 0.2, it will be raining.
- 5:15:55And likewise, if today is raining, then it's more likely than not
- 5:15:59that tomorrow is also raining.
- 5:16:01With probability 0.7, it'll be raining. With probability 0.3, it will be sunny.
- 5:16:06So this matrix, this description of how it is we transition from one state
- 5:16:10to the next state is what we're going to call the transition model.
- 5:16:14And using the transition model, you can begin
- 5:16:16to construct this Markov chain by just predicting,
- 5:16:20given today's weather, what's the likelihood of tomorrow's weather
- 5:16:23happening.
- 5:16:23And you can imagine doing a similar sampling procedure,
- 5:16:27where you take this information, you sample what tomorrow's weather is
- 5:16:30going to be.
- 5:16:31Using that, you sample the next day's weather.
- 5:16:33And the result of that is you can form this Markov chain of like x0,
- 5:16:38time and time, day zero is sunny, the next day is sunny,
- 5:16:40maybe the next day it changes to raining, then raining, then raining.
- 5:16:43And the pattern that this Markov chain follows,
- 5:16:46given the distribution that we had access to, this transition model here,
- 5:16:50is that when it's sunny, it tends to stay sunny for a little while.
- 5:16:53The next couple of days tend to be sunny too.
- 5:16:55And when it's raining, it tends to be raining as well.
- 5:16:59And so you get a Markov chain that looks like this,
- 5:17:01and you can do analysis on this.
- 5:17:02You can say, given that today is raining, what is the probability
- 5:17:06that tomorrow is raining?
- 5:17:07Or you can begin to ask probability questions
- 5:17:09like, what is the probability of this sequence of five values, sun, sun,
- 5:17:13rain, rain, rain, and answer those sorts of questions too.
- 5:17:17And it turns out there are, again, many Python libraries
- 5:17:19for interacting with models like this of probabilities
- 5:17:23that have distributions and random variables that
- 5:17:25are based on previous variables according to this Markov assumption.
- 5:17:29And pomegranate2 has ways of dealing with these sorts of variables.
- 5:17:32So I'll go ahead and go into the chain directory,
- 5:17:39where I have some information about Markov chains.
- 5:17:42And here, I've defined a file called model.py,
- 5:17:45where I've defined in a very similar syntax.
- 5:17:47And again, the exact syntax doesn't matter so much as the idea
- 5:17:50that I'm encoding this information into a Python program
- 5:17:54so that the program has access to these distributions.
- 5:17:56I've here defined some starting distribution.
- 5:17:59So every Markov model begins at some point in time,
- 5:18:02and I need to give it some starting distribution.
- 5:18:04And so we'll just say, you know at the start, you can pick 50-50 between sunny
- 5:18:08and rainy.
- 5:18:09We'll say it's sunny 50% of the time, rainy 50% of the time.
- 5:18:13And then down below, I've here defined the transition model,
- 5:18:16how it is that I transition from one day to the next.
- 5:18:19And here, I've encoded that exact same matrix from before,
- 5:18:22that if it was sunny today, then with probability 0.8,
- 5:18:24it will be sunny tomorrow.
- 5:18:26And it'll be rainy tomorrow with probability 0.2.
- 5:18:29And I likewise have another distribution for if it was raining today instead.
- 5:18:34And so that alone defines the Markov model.
- 5:18:36You can begin to answer questions using that model.
- 5:18:39But one thing I'll just do is sample from the Markov chain.
- 5:18:42It turns out there is a method built into this Markov chain library
- 5:18:45that allows me to sample 50 states from the chain,
- 5:18:48basically just simulating like 50 instances of weather.
- 5:18:52And so let me go ahead and run this.
- 5:18:54Python model.py.
- 5:18:57And when I run it, what I get is that it's
- 5:18:59going to sample from this Markov chain 50 states, 50 days worth of weather
- 5:19:04that it's just going to randomly sample.
- 5:19:06And you can imagine sampling many times to be able to get more data,
- 5:19:09to be able to do more analysis.
- 5:19:10But here, for example, it's sunny two days in a row,
- 5:19:13rainy a whole bunch of days in a row before it changes back to sun.
- 5:19:17And so you get this model that follows the distribution
- 5:19:20that we originally described, that follows the distribution of sunny days
- 5:19:23tend to lead to more sunny days.
- 5:19:25Rainy days tend to lead to more rainy days.
- 5:19:29And that then is a Markov model.
- 5:19:31And Markov models rely on us knowing the values
- 5:19:34of these individual states.
- 5:19:35I know that today is sunny or that today is raining.
- 5:19:38And using that information, I can draw some sort of inference
- 5:19:41about what tomorrow is going to be like.
- 5:19:44But in practice, this often isn't the case.
- 5:19:46It often isn't the case that I know for certain what
- 5:19:49the exact state of the world is.
- 5:19:51Oftentimes, the state of the world is exactly unknown.
- 5:19:54But I'm able to somehow sense some information about that state,
- 5:19:58that a robot or an AI doesn't have exact knowledge
- 5:20:01about the world around it.
- 5:20:02But it has some sort of sensor, whether that sensor is a camera
- 5:20:05or sensors that detect distance or just a microphone that is sensing audio,
- 5:20:09for example.
- 5:20:09It is sensing data.
- 5:20:11And using that data, that data is somehow related
- 5:20:14to the state of the world, even if it doesn't actually know,
- 5:20:17our AI doesn't know, what the underlying true state of the world
- 5:20:20actually is.
- 5:20:22And for that, we need to get into the world of sensor models,
- 5:20:25the way of describing how it is that we translate
- 5:20:28what the hidden state, the underlying true state of the world,
- 5:20:31is with what the observation, what it is that the AI knows or the AI has
- 5:20:36access to, actually is.
- 5:20:38And so for example, a hidden state might be a robot's position.
- 5:20:42If a robot is exploring new uncharted territory,
- 5:20:45the robot likely doesn't know exactly where it is.
- 5:20:48But it does have an observation.
- 5:20:49It has robot sensor data, where it can sense how far away
- 5:20:52are possible obstacles around it.
- 5:20:54And using that information, using the observed information that it has,
- 5:20:58it can infer something about the hidden state.
- 5:21:01Because what the true hidden state is influences those observations.
- 5:21:05Whatever the robot's true position is affects or has some effect
- 5:21:10upon what the sensor data of the robot is able to collect is,
- 5:21:13even if the robot doesn't actually know for certain what its true position is.
- 5:21:18Likewise, if you think about a voice recognition or a speech recognition
- 5:21:21program that listens to you and is able to respond to you, something
- 5:21:25like Alexa or what Apple and Google are doing with their voice recognition
- 5:21:29as well, that you might imagine that the hidden state, the underlying state,
- 5:21:33is what words are actually spoken.
- 5:21:35The true nature of the world contains you saying
- 5:21:38a particular sequence of words, but your phone or your smart home device
- 5:21:42doesn't know for sure exactly what words you said.
- 5:21:45The only observation that the AI has access to is some audio waveforms.
- 5:21:50And those audio waveforms are, of course, dependent upon this hidden state.
- 5:21:54And you can infer, based on those audio waveforms,
- 5:21:57what the words spoken likely were.
- 5:22:00But you might not know with 100% certainty what that hidden state actually
- 5:22:04is.
- 5:22:05And it might be a task to try and predict, given this observation,
- 5:22:08given these audio waveforms, can you figure out what the actual words spoken
- 5:22:12are.
- 5:22:13And likewise, you might imagine on a website, true user engagement.
- 5:22:16Might be information you don't directly have access to.
- 5:22:19But you can observe data, like website or app analytics,
- 5:22:22about how often was this button clicked or how often are people interacting
- 5:22:25with a page in a particular way.
- 5:22:26And you can use that to infer things about your users as well.
- 5:22:30So this type of problem comes up all the time
- 5:22:33when we're dealing with AI and trying to infer things about the world.
- 5:22:36That often AI doesn't really know the hidden true state of the world.
- 5:22:40All the AI has access to is some observation
- 5:22:43that is related to the hidden true state.
- 5:22:45But it's not direct.
- 5:22:47There might be some noise there.
- 5:22:48The audio waveform might have some additional noise
- 5:22:50that might be difficult to parse.
- 5:22:52The sensor data might not be exactly correct.
- 5:22:54There's some noise that might not allow you to conclude with certainty what
- 5:22:57the hidden state is, but can allow you to infer what it might be.
- 5:23:01And so the simple example we'll take a look at here
- 5:23:04is imagining the hidden state as the weather, whether it's sunny or rainy
- 5:23:07or not.
- 5:23:07And imagine you are programming an AI inside of a building that maybe has
- 5:23:11access to just a camera to inside the building.
- 5:23:14And all you have access to is an observation
- 5:23:17as to whether or not employees are bringing
- 5:23:19an umbrella into the building or not.
- 5:23:21You can detect whether it's an umbrella or not.
- 5:23:24And so you might have an observation as to whether or not
- 5:23:26an umbrella is brought into the building or not.
- 5:23:28And using that information, you want to predict whether it's sunny or rainy,
- 5:23:32even if you don't know what the underlying weather is.
- 5:23:35So the underlying weather might be sunny or rainy.
- 5:23:37And if it's raining, obviously people are more likely to bring an umbrella.
- 5:23:41And so whether or not people bring an umbrella, your observation,
- 5:23:44tells you something about the hidden state.
- 5:23:46And of course, this is a bit of a contrived example,
- 5:23:48but the idea here is to think about this more broadly in terms of more
- 5:23:51generally, any time you observe something,
- 5:23:54it having to do with some underlying hidden state.
- 5:23:57And so to try and model this type of idea where
- 5:23:59we have these hidden states and observations,
- 5:24:02rather than just use a Markov model, which has state, state, state, state,
- 5:24:05each of which is connected by that transition matrix that we described
- 5:24:08before, we're going to use what we call a hidden Markov model.
- 5:24:12Very similar to a Markov model, but this is going
- 5:24:14to allow us to model a system that has hidden states
- 5:24:17that we don't directly observe, along with some observed event
- 5:24:21that we do actually see.
- 5:24:23And so in addition to that transition model that we still
- 5:24:25need of saying, given the underlying state of the world,
- 5:24:28if it's sunny or rainy, what's the probability of tomorrow's weather?
- 5:24:32We also need another model that, given some state,
- 5:24:35is going to give us an observation of green, yes, someone brings
- 5:24:38an umbrella into the office, or red, no, nobody brings umbrellas into the office.
- 5:24:43And so the observation might be that if it's sunny,
- 5:24:46then odds are nobody is going to bring an umbrella to the office.
- 5:24:49But maybe some people are just being cautious,
- 5:24:51and they do bring an umbrella to the office anyways.
- 5:24:54And if it's raining, then with much higher probability,
- 5:24:57then people are going to bring umbrellas into the office.
- 5:24:59But maybe if the rain was unexpected, people didn't bring an umbrella.
- 5:25:02And so it might have some other probability as well.
- 5:25:05And so using the observations, you can begin
- 5:25:07to predict with reasonable likelihood what the underlying state is,
- 5:25:11even if you don't actually get to observe the underlying state,
- 5:25:15if you don't get to see what the hidden state is actually equal to.
- 5:25:18This here we'll often call the sensor model.
- 5:25:21It's also often called the emission probabilities,
- 5:25:23because the state, the underlying state, emits some sort of emission
- 5:25:27that you then observe.
- 5:25:29And so that can be another way of describing that same idea.
- 5:25:32And the sensor Markov assumption that we're going to use
- 5:25:35is this assumption that the evidence variable, the thing we observe,
- 5:25:38the emission that gets produced, depends only on the corresponding state,
- 5:25:43meaning it can predict whether or not people will bring umbrellas or not
- 5:25:46entirely dependent just on whether it is sunny or rainy today.
- 5:25:50Of course, again, this assumption might not hold in practice,
- 5:25:53that in practice, it might depend whether or not
- 5:25:55people bring umbrellas, might depend not just on today's weather,
- 5:25:58but also on yesterday's weather and the day before.
- 5:26:00But for simplification purposes, it can be helpful to apply this sort
- 5:26:04of assumption just to allow us to be able to reason
- 5:26:07about these probabilities a little more easily.
- 5:26:09And if we're able to approximate it, we can still often get a very good answer.
- 5:26:14And so what these hidden Markov models end up looking like
- 5:26:16is a little something like this, where now, rather than just have
- 5:26:20one chain of states, like sun, sun, rain, rain, rain,
- 5:26:23we instead have this upper level, which is the underlying state of the world.
- 5:26:29Is it sunny or is it rainy?
- 5:26:30And those are connected by that transition matrix we described before.
- 5:26:34But each of these states produces an emission,
- 5:26:37produces an observation that I see, that on this day, it was sunny
- 5:26:41and people didn't bring umbrellas.
- 5:26:43And on this day, it was sunny, but people did bring umbrellas.
- 5:26:46And on this day, it was raining and people did bring umbrellas,
- 5:26:48and so on and so forth.
- 5:26:49And so each of these underlying states represented
- 5:26:52by x sub t for x sub 1, 0, 1, 2, so on and so forth,
- 5:26:56produces some sort of observation or emission,
- 5:26:59which is what the e stands for, e sub 0, e sub 1, e sub 2, so on and so forth.
- 5:27:04And so this, too, is a way of trying to represent this idea.
- 5:27:07And what you want to think about is that these underlying states are
- 5:27:10the true nature of the world, the robot's position as it moves over time,
- 5:27:14and that produces some sort of sensor data that might be observed,
- 5:27:17or what people are actually saying and using the emission data of what
- 5:27:21audio waveforms do you detect in order to process that data
- 5:27:24and try and figure it out.
- 5:27:26And there are a number of possible tasks that you might want to do
- 5:27:29given this kind of information.
- 5:27:30And one of the simplest is trying to infer something
- 5:27:33about the future or the past or about these sort of hidden states that
- 5:27:37might exist.
- 5:27:38And so the tasks that you'll often see, and we're not
- 5:27:40going to go into the mathematics of these tasks,
- 5:27:42but they're all based on the same idea of conditional probabilities
- 5:27:45and using the probability distributions we
- 5:27:48have to draw these sorts of conclusions.
- 5:27:51One task is called filtering, which is given observations from the start
- 5:27:55until now, calculate the distribution for the current state,
- 5:27:59meaning given information about from the beginning of time until now,
- 5:28:03on which days do people bring an umbrella or not bring an umbrella,
- 5:28:06can I calculate the probability of the current state that today,
- 5:28:10is it sunny or is it raining?
- 5:28:12Another task that might be possible is prediction,
- 5:28:14which is looking towards the future.
- 5:28:16Given observations about people bringing umbrellas
- 5:28:18from the beginning of when we started counting time until now,
- 5:28:22can I figure out the distribution that tomorrow is it sunny or is it
- 5:28:25raining?
- 5:28:26And you can also go backwards as well by a smoothing,
- 5:28:29where I can say given observations from start until now,
- 5:28:32calculate the distributions for some past state.
- 5:28:35Like I know that today people brought umbrellas and tomorrow people
- 5:28:38brought umbrellas.
- 5:28:39And so given two days worth of data of people bringing umbrellas,
- 5:28:42what's the probability that yesterday it was raining?
- 5:28:45And that I know that people brought umbrellas today,
- 5:28:47that might inform that decision as well.
- 5:28:50It might influence those probabilities.
- 5:28:52And there's also a most likely explanation task,
- 5:28:56in addition to other tasks that might exist as well, which
- 5:28:58is combining some of these given observations from the start up
- 5:29:01until now, figuring out the most likely sequence of states.
- 5:29:04And this is what we're going to take a look at now, this idea that if I
- 5:29:07have all these observations, umbrella, no umbrella, umbrella, no umbrella,
- 5:29:11can I calculate the most likely states of sun, rain, sun, rain, and whatnot
- 5:29:15that actually represented the true weather that
- 5:29:18would produce these observations?
- 5:29:20And this is quite common when you're trying to do something like voice
- 5:29:23recognition, for example, that you have these emissions of the audio waveforms,
- 5:29:27and you would like to calculate based on all of the observations
- 5:29:30that you have, what is the most likely sequence of actual words, or syllables,
- 5:29:34or sounds that the user actually made when they were speaking
- 5:29:38to this particular device, or other tasks that might come up in that context
- 5:29:41as well.
- 5:29:43And so we can try this out by going ahead and going into the HMM directory,
- 5:29:47HMM for Hidden Markov Model.
- 5:29:50And here, what I've done is I've defined a model where this model first defines
- 5:29:57my possible state, sun, and rain, along with their emission probabilities,
- 5:30:02the observation model, or the emission model, where here, given
- 5:30:06that I know that it's sunny, the probability
- 5:30:09that I see people bring an umbrella is 0.2,
- 5:30:11the probability of no umbrella is 0.8.
- 5:30:14And likewise, if it's raining, then people
- 5:30:16are more likely to bring an umbrella.
- 5:30:18Umbrella has probability 0.9, no umbrella has probability 0.1.
- 5:30:21So the actual underlying hidden states, those states are sun and rain,
- 5:30:26but the things that I observe, the observations that I can see,
- 5:30:29are either umbrella or no umbrella as the things that I observe as a result.
- 5:30:35So this then, I also need to add to it a transition matrix, same as before,
- 5:30:39saying that if today is sunny, then tomorrow is more likely to be sunny.
- 5:30:43And if today is rainy, then tomorrow is more likely to be raining.
- 5:30:47As of before, I give it some starting probabilities,
- 5:30:49saying at first, 50-50 chance for whether it's sunny or rainy.
- 5:30:53And then I can create the model based on that information.
- 5:30:56Again, the exact syntax of this is not so important,
- 5:30:59so much as it is the data that I am now encoding into a program,
- 5:31:02such that now I can begin to do some inference.
- 5:31:06So I can give my program, for example, a list of observations,
- 5:31:10umbrella, umbrella, no umbrella, umbrella, umbrella, so on and so forth,
- 5:31:13no umbrella, no umbrella.
- 5:31:14And I would like to calculate, I would like to figure out the most likely
- 5:31:18explanation for these observations.
- 5:31:20What is likely is whether rain, rain, is this rain,
- 5:31:23or is it more likely that this was actually sunny,
- 5:31:25and then it switched back to it being rainy?
- 5:31:28And that's an interesting question.
- 5:31:29We might not be sure, because it might just
- 5:31:31be that it just so happened on this rainy day,
- 5:31:34people decided not to bring an umbrella.
- 5:31:36Or it could be that it switched from rainy to sunny back to rainy,
- 5:31:40which doesn't seem too likely, but it certainly could happen.
- 5:31:43And using the data we give to the hidden Markov model,
- 5:31:46our model can begin to predict these answers, can begin to figure it out.
- 5:31:49So we're going to go ahead and just predict these observations.
- 5:31:53And then for each of those predictions, go ahead and print out
- 5:31:56what the prediction is.
- 5:31:56And this library just so happens to have a function called
- 5:31:59predict that does this prediction process for me.
- 5:32:03So I'll run python sequence.py.
- 5:32:06And the result I get is this.
- 5:32:07This is the prediction based on the observations
- 5:32:10of what all of those states are likely to be.
- 5:32:12And it's likely to be rain and rain.
- 5:32:14In this case, it thinks that what most likely happened
- 5:32:16is that it was sunny for a day and then went back to being rainy.
- 5:32:19But in different situations, if it was rainy for longer maybe,
- 5:32:22or if the probabilities were slightly different,
- 5:32:24you might imagine that it's more likely that it was rainy all the way through.
- 5:32:27And it just so happened on one rainy day, people decided not to bring umbrellas.
- 5:32:32And so here, too, Python libraries can begin
- 5:32:35to allow for the sort of inference procedure.
- 5:32:38And by taking what we know and by putting it
- 5:32:40in terms of these tasks that already exist,
- 5:32:43these general tasks that work with hidden Markov models,
- 5:32:45then any time we can take an idea and formulate it as a hidden Markov model,
- 5:32:50formulate it as something that has hidden states
- 5:32:52and observed emissions that result from those states,
- 5:32:55then we can take advantage of these algorithms
- 5:32:57that are known to exist for trying to do this sort of inference.
- 5:33:01So now we've seen a couple of ways that AI can begin to deal with uncertainty.
- 5:33:05We've taken a look at probability and how we can use probability
- 5:33:08to describe numerically things that are likely or more likely or less
- 5:33:11likely to happen than other events or other variables.
- 5:33:14And using that information, we can begin to construct
- 5:33:17these standard types of models, things like Bayesian networks and Markov
- 5:33:20chains and hidden Markov models that all allow us to be able to describe
- 5:33:25how particular events relate to other events
- 5:33:27or how the values of particular variables relate to other variables,
- 5:33:30not for certain, but with some sort of probability distribution.
- 5:33:34And by formulating things in terms of these models that already exist,
- 5:33:37we can take advantage of Python libraries that
- 5:33:39implement these sort of models already and allow us just
- 5:33:42to be able to use them to produce some sort of resulting effect.
- 5:33:46So all of this then allows our AI to begin
- 5:33:48to deal with these sort of uncertain problems
- 5:33:50so that our AI doesn't need to know things for certain
- 5:33:53but can infer based on information it doesn't know.
- 5:33:56Next time, we'll take a look at additional types of problems
- 5:33:59that we can solve by taking advantage of AI-related algorithms,
- 5:34:02even beyond the world of the types of problems we've already explored.
- 5:34:05We'll see you next time.
- 5:34:08OK.
- 5:34:27Welcome back, everyone, to an introduction to artificial intelligence
- 5:34:30with Python.
- 5:34:31And now, so far, we've taken a look at a couple
- 5:34:32of different types of problems.
- 5:34:34We've seen classical search problems where
- 5:34:36we're trying to get from an initial state to a goal
- 5:34:38by figuring out some optimal path.
- 5:34:40We've taken a look at adversarial search where
- 5:34:42we have a game-playing agent that is trying to make the best move.
- 5:34:45We've seen knowledge-based problems where we're trying to use logic
- 5:34:48and inference to be able to figure out and draw
- 5:34:50some additional conclusions.
- 5:34:51And we've seen some probabilistic models as well where we might not
- 5:34:54have certain information about the world,
- 5:34:56but we want to use the knowledge about probabilities that we do have
- 5:34:59to be able to draw some conclusions.
- 5:35:01Today, we're going to turn our attention to another category of problems
- 5:35:04generally known as optimization problems, where optimization is really
- 5:35:08all about choosing the best option from a set of possible options.
- 5:35:12And we've already seen optimization in some contexts,
- 5:35:14like game-playing, where we're trying to create an AI that
- 5:35:17chooses the best move out of a set of possible moves.
- 5:35:19But what we'll take a look at today is a category of types of problems
- 5:35:23and algorithms to solve them that can be used
- 5:35:25in order to deal with a broader range of potential optimization problems.
- 5:35:29And the first of the algorithms that we'll take a look at
- 5:35:32is known as a local search.
- 5:35:34And local search differs from search algorithms
- 5:35:36we've seen before in the sense that the search algorithms we've
- 5:35:38looked at so far, which are things like breadth-first search or A-star search,
- 5:35:42for example, generally maintain a whole bunch of different paths
- 5:35:45that we're simultaneously exploring, and we're
- 5:35:47looking at a bunch of different paths at once trying
- 5:35:50to find our way to the solution.
- 5:35:51On the other hand, in local search, this is going
- 5:35:53to be a search algorithm that's really just going to maintain a single node,
- 5:35:57looking at a single state.
- 5:35:59And we'll generally run this algorithm by maintaining that single node
- 5:36:02and then moving ourselves to one of the neighboring nodes
- 5:36:05throughout this search process.
- 5:36:07And this is generally useful in context not like these problems, which
- 5:36:10we've seen before, like a maze-solving situation where
- 5:36:13we're trying to find our way from the initial state to the goal
- 5:36:16by following some path.
- 5:36:17But local search is most applicable when we really
- 5:36:20don't care about the path at all, and all we care about
- 5:36:23is what the solution is.
- 5:36:24And in the case of solving a maze, the solution was always obvious.
- 5:36:27You could point to the solution.
- 5:36:28You know exactly what the goal is, and the real question
- 5:36:31is, what is the path to get there?
- 5:36:33But local search is going to come up in cases
- 5:36:35where figuring out exactly what the solution is,
- 5:36:37exactly what the goal looks like, is actually the heart of the challenge.
- 5:36:41And to give an example of one of these kinds of problems,
- 5:36:44we'll consider a scenario where we have two types of buildings,
- 5:36:46for example.
- 5:36:47We have houses and hospitals.
- 5:36:49And our goal might be in a world that's formatted as this grid,
- 5:36:52where we have a whole bunch of houses, a house here, house here,
- 5:36:55two houses over there, maybe we want to try and find a way
- 5:36:58to place two hospitals on this map.
- 5:37:01So maybe a hospital here and a hospital there.
- 5:37:04And the problem now is we want to place two hospitals on the map,
- 5:37:07but we want to do so with some sort of objective.
- 5:37:09And our objective in this case is to try and minimize
- 5:37:12the distance of any of the houses from a hospital.
- 5:37:16So you might imagine, all right, what's the distance
- 5:37:18from each of the houses to their nearest hospital?
- 5:37:20There are a number of ways we could calculate that distance.
- 5:37:23But one way is using a heuristic we've looked at before,
- 5:37:25which is the Manhattan distance, this idea of how many rows
- 5:37:28and columns would you have to move inside of this grid layout in order
- 5:37:32to get to a hospital, for example.
- 5:37:34And it turns out, if you take each of these four houses
- 5:37:36and figure out, all right, how close are they to their nearest hospital,
- 5:37:39you get something like this, where this house is three away from a hospital,
- 5:37:42this house is six away, and these two houses are each four away.
- 5:37:46And if you add all those numbers up together,
- 5:37:48you get a total cost of 17, for example.
- 5:37:51So for this particular configuration of hospitals, a hospital here
- 5:37:55and a hospital there, that state, we might say,
- 5:37:58has a cost of 17.
- 5:37:59And the goal of this problem now that we would
- 5:38:01like to apply a search algorithm to figure out
- 5:38:04is, can you solve this problem to find a way to minimize that cost?
- 5:38:08Minimize the total amount if you sum up all of the distances
- 5:38:11from all the houses to the nearest hospital.
- 5:38:14How can we minimize that final value?
- 5:38:16And if we think about this problem a little bit more abstractly,
- 5:38:19abstracting away from this specific problem
- 5:38:21and thinking more generally about problems like it,
- 5:38:23you can often formulate these problems by thinking about them
- 5:38:26as a state-space landscape, as we'll soon call it.
- 5:38:29Here in this diagram of a state-space landscape,
- 5:38:32each of these vertical bars represents a particular state
- 5:38:35that our world could be in.
- 5:38:37So for example, each of these vertical bars
- 5:38:39represents a particular configuration of two hospitals.
- 5:38:43And the height of this vertical bar is generally
- 5:38:45going to represent some function of that state, some value of that state.
- 5:38:50So maybe in this case, the height of the vertical bar
- 5:38:52represents what is the cost of this particular configuration
- 5:38:56of hospitals in terms of what is the sum total of all the distances
- 5:38:59from all of the houses to their nearest hospital.
- 5:39:03And generally speaking, when we have a state-space landscape,
- 5:39:06we want to do one of two things.
- 5:39:08We might be trying to maximize the value of this function,
- 5:39:12trying to find a global maximum, so to speak, of this state-space landscape,
- 5:39:16a single state whose value is higher than all of the other states
- 5:39:20that we could possibly choose from.
- 5:39:22And generally in this case, when we're trying to find a global maximum,
- 5:39:25we'll call the function that we're trying to optimize
- 5:39:27some objective function, some function that
- 5:39:30measures for any given state how good is that state,
- 5:39:34such that we can take any state, pass it into the objective function,
- 5:39:37and get a value for how good that state is.
- 5:39:39And ultimately, what our goal is is to find one of these states
- 5:39:42that has the highest possible value for that objective function.
- 5:39:46An equivalent but reversed problem is the problem
- 5:39:49of finding a global minimum, some state that has a value
- 5:39:52after you pass it into this function that is lower than all of the other
- 5:39:55possible values that we might choose from.
- 5:39:57And generally speaking, when we're trying to find a global minimum,
- 5:40:00we call the function that we're calculating a cost function.
- 5:40:03Generally, each state has some sort of cost,
- 5:40:05whether that cost is a monetary cost, or a time cost,
- 5:40:08or in the case of the houses and hospitals,
- 5:40:10we've been looking at just now, a distance cost in terms
- 5:40:13of how far away each of the houses is from a hospital.
- 5:40:17And we're trying to minimize the cost, find
- 5:40:19the state that has the lowest possible value of that cost.
- 5:40:23So these are the general types of ideas we
- 5:40:25might be trying to go for within a state space landscape,
- 5:40:28trying to find a global maximum, or trying to find a global minimum.
- 5:40:32And how exactly do we do that?
- 5:40:33We'll recall that in local search, we generally
- 5:40:36operate this algorithm by maintaining just a single state,
- 5:40:39just some current state represented inside of some node,
- 5:40:41maybe inside of a data structure, where we're
- 5:40:43keeping track of where we are currently.
- 5:40:46And then ultimately, what we're going to do is from that state,
- 5:40:49move to one of its neighbor states.
- 5:40:51So in this case, represented in this one-dimensional space
- 5:40:54by just the state immediately to the left or to the right of it.
- 5:40:57But for any different problem, you might define
- 5:40:58what it means for there to be a neighbor of a particular state.
- 5:41:02In the case of a hospital, for example, that we were just looking at,
- 5:41:05a neighbor might be moving one hospital one space to the left
- 5:41:08or to the right or up or down.
- 5:41:10Some state that is close to our current state, but slightly different,
- 5:41:14and as a result, might have a slightly different value
- 5:41:17in terms of its objective function or in terms of its cost function.
- 5:41:21So this is going to be our general strategy in local search,
- 5:41:24to be able to take a state, maintaining some current node,
- 5:41:27and move where we're looking at in the state space landscape
- 5:41:29in order to try to find a global maximum or a global minimum somehow.
- 5:41:33And perhaps the simplest of algorithms that we
- 5:41:35could use to implement this idea of local search
- 5:41:38is an algorithm known as hill climbing.
- 5:41:41And the basic idea of hill climbing is, let's
- 5:41:43say I'm trying to maximize the value of my state.
- 5:41:46I'm trying to figure out where the global maximum is.
- 5:41:49I'm going to start at a state.
- 5:41:50And generally, what hill climbing is going to do
- 5:41:53is it's going to consider the neighbors of that state,
- 5:41:55that from this state, all right, I could go left or I could go right,
- 5:41:58and this neighbor happens to be higher and this neighbor happens to be lower.
- 5:42:01And in hill climbing, if I'm trying to maximize the value,
- 5:42:04I'll generally pick the highest one I can between the state
- 5:42:07to the left and right of me.
- 5:42:08This one is higher.
- 5:42:10So I'll go ahead and move myself to consider that state instead.
- 5:42:13And then I'll repeat this process, continually looking at all of my neighbors
- 5:42:17and picking the highest neighbor, doing the same thing,
- 5:42:19looking at my neighbors, picking the highest of my neighbors,
- 5:42:21until I get to a point like right here, where I consider both of my neighbors
- 5:42:25and both of my neighbors have a lower value than I do.
- 5:42:29This current state has a value that is higher than any of its neighbors.
- 5:42:32And at that point, the algorithm terminates.
- 5:42:34And I can say, all right, here I have now found the solution.
- 5:42:38And the same thing works in exactly the opposite way
- 5:42:40for trying to find a global minimum.
- 5:42:42But the algorithm is fundamentally the same.
- 5:42:44If I'm trying to find a global minimum and say my current state starts here,
- 5:42:47I'll continually look at my neighbors, pick the lowest value
- 5:42:50that I possibly can, until I eventually, hopefully,
- 5:42:53find that global minimum, a point at which when
- 5:42:55I look at both of my neighbors, they each have a higher value.
- 5:42:58And I'm trying to minimize the total score or cost or value
- 5:43:02that I get as a result of calculating some sort of cost function.
- 5:43:06So we can formulate this graphical idea in terms of pseudocode.
- 5:43:09And the pseudocode for hill climbing might look like this.
- 5:43:12We define some function called hill climb that
- 5:43:15takes as input the problem that we're trying to solve.
- 5:43:17And generally, we're going to start in some sort of initial state.
- 5:43:21So I'll start with a variable called current
- 5:43:23that is keeping track of my initial state, like an initial configuration
- 5:43:26of hospitals.
- 5:43:27And maybe some problems lend themselves to an initial state,
- 5:43:30some place where you begin.
- 5:43:31In other cases, maybe not, in which case we might just randomly
- 5:43:34generate some initial state, just by choosing two locations for hospitals
- 5:43:38at random, for example, and figuring out from there
- 5:43:41how we might be able to improve.
- 5:43:42But that initial state, we're going to store inside of current.
- 5:43:46And now, here comes our loop, some repetitive process
- 5:43:48we're going to do again and again until the algorithm terminates.
- 5:43:52And what we're going to do is first say, let's
- 5:43:55figure out all of the neighbors of the current state.
- 5:43:57From my state, what are all of the neighboring
- 5:43:59states for some definition of what it means to be a neighbor?
- 5:44:02And I'll go ahead and choose the highest value of all of those neighbors
- 5:44:06and save it inside of this variable called neighbor.
- 5:44:09So keep track of the highest-valued neighbor.
- 5:44:11This is in the case where I'm trying to maximize the value.
- 5:44:14In the case where I'm trying to minimize the value,
- 5:44:15you might imagine here, you'll pick the neighbor
- 5:44:17with the lowest possible value.
- 5:44:18But these ideas are really fundamentally interchangeable.
- 5:44:21And it's possible, in some cases, there might be multiple neighbors
- 5:44:24that each have an equally high value or an equally low value
- 5:44:28in the minimizing case.
- 5:44:29And in that case, we can just choose randomly from among them.
- 5:44:31Choose one of them and save it inside of this variable neighbor.
- 5:44:35And then the key question to ask is, is this neighbor better
- 5:44:39than my current state?
- 5:44:41And if the neighbor, the best neighbor that I was able to find,
- 5:44:44is not better than my current state, well, then the algorithm is over.
- 5:44:48And I'll just go ahead and return the current state.
- 5:44:50If none of my neighbors are better, then I may as well stay where I am,
- 5:44:53is the general logic of the hill climbing algorithm.
- 5:44:56But otherwise, if the neighbor is better, then I may as well
- 5:44:59move to that neighbor.
- 5:45:00So you might imagine setting current equal to neighbor, where the general idea
- 5:45:04is if I'm at a current state and I see a neighbor that is better than me,
- 5:45:07then I'll go ahead and move there.
- 5:45:08And then I'll repeat the process, continually moving to a better neighbor
- 5:45:11until I reach a point at which none of my neighbors are better than I am.
- 5:45:15And at that point, we'd say the algorithm can just terminate there.
- 5:45:19So let's take a look at a real example of this
- 5:45:21with these houses and hospitals.
- 5:45:23So we've seen now that if we put the hospitals in these two locations,
- 5:45:26that has a total cost of 17.
- 5:45:28And now we need to define, if we're going to implement this hill climbing
- 5:45:31algorithm, what it means to take this particular configuration
- 5:45:34of hospitals, this particular state, and get a neighbor of that state.
- 5:45:39And a simple definition of neighbor might be just,
- 5:45:42let's pick one of the hospitals and move it by one square, the left or right
- 5:45:46or up or down, for example.
- 5:45:48And that would mean we have six possible neighbors
- 5:45:50from this particular configuration.
- 5:45:52We could take this hospital and move it to any of these three possible squares,
- 5:45:56or we take this hospital and move it to any of those three possible squares.
- 5:46:00And each of those would generate a neighbor.
- 5:46:02And what I might do is say, all right, here's
- 5:46:04the locations and the distances between each of the houses
- 5:46:07and their nearest hospital.
- 5:46:09Let me consider all of the neighbors and see if any of them
- 5:46:12can do better than a cost of 17.
- 5:46:14And it turns out there are a couple of ways that we could do that.
- 5:46:17And it doesn't matter if we randomly choose
- 5:46:19among all the ways that are the best.
- 5:46:20But one such possible way is by taking a look at this hospital here
- 5:46:24and considering the directions in which it might move.
- 5:46:27If we hold this hospital constant, if we take this hospital
- 5:46:30and move it one square up, for example, that doesn't really help us.
- 5:46:33It gets closer to the house up here, but it gets further away
- 5:46:36from the house down here.
- 5:46:37And it doesn't really change anything for the two houses
- 5:46:40along the left-hand side.
- 5:46:41But if we take this hospital on the right and move it one square down,
- 5:46:45it's the opposite problem.
- 5:46:46It gets further away from the house up above,
- 5:46:49and it gets closer to the house down below.
- 5:46:51The real idea, the goal should be to be able to take this hospital
- 5:46:54and move it one square to the left.
- 5:46:56By moving it one square to the left, we move it closer
- 5:46:59to both of these houses on the right without changing anything
- 5:47:02about the houses on the left.
- 5:47:03For them, this hospital is still the closer one, so they aren't affected.
- 5:47:06So we're able to improve the situation by picking a neighbor that
- 5:47:10results in a decrease in our total cost.
- 5:47:13And so we might do that.
- 5:47:14Move ourselves from this current state to a neighbor
- 5:47:16by just taking that hospital and moving it.
- 5:47:19And at this point, there's not a whole lot
- 5:47:21that can be done with this hospital.
- 5:47:22But there's still other optimizations we can make, other neighbors
- 5:47:25we can move to that are going to have a better value.
- 5:47:27If we consider this hospital, for example,
- 5:47:29we might imagine that right now it's a bit far up,
- 5:47:32that both of these houses are a little bit lower.
- 5:47:34So we might be able to do better by taking this hospital
- 5:47:37and moving it one square down, moving it down so that now instead
- 5:47:40of a cost of 15, we're down to a cost of 13
- 5:47:43for this particular configuration.
- 5:47:45And we can do even better by taking the hospital
- 5:47:47and moving it one square to the left.
- 5:47:49Now instead of a cost of 13, we have a cost of 11,
- 5:47:52because this house is one away from the hospital.
- 5:47:54This one is four away.
- 5:47:56This one is three away.
- 5:47:57And this one is also three away.
- 5:47:59So we've been able to do much better than that initial cost
- 5:48:02that we had using the initial configuration.
- 5:48:04Just by taking every state and asking ourselves the question,
- 5:48:07can we do better by just making small incremental changes,
- 5:48:11moving to a neighbor, moving to a neighbor,
- 5:48:12and moving to a neighbor after that?
- 5:48:15And now at this point, we can potentially see that at this point,
- 5:48:18the algorithm is going to terminate.
- 5:48:20There's actually no neighbor we can move to
- 5:48:22that is going to improve the situation, get us a cost that is less than 11.
- 5:48:27Because if we take this hospital and move it upper to the right,
- 5:48:29well, that's going to make it further away.
- 5:48:31If we take it and move it down, that doesn't really change the situation.
- 5:48:34It gets further away from this house but closer to that house.
- 5:48:37And likewise, the same story was true for this hospital.
- 5:48:40Any neighbor we move it to, up, left, down, or right,
- 5:48:42is either going to make it further away from the houses and increase the cost,
- 5:48:46or it's going to have no effect on the cost whatsoever.
- 5:48:51And so the question we might now ask is, is this the best we could do?
- 5:48:54Is this the best placement of the hospitals we could possibly have?
- 5:48:57And it turns out the answer is no, because there's a better way
- 5:49:00that we could place these hospitals.
- 5:49:02And in particular, there are a number of ways you could do this.
- 5:49:05But one of the ways is by taking this hospital here
- 5:49:07and moving it to this square, for example, moving it diagonally
- 5:49:10by one square, which was not part of our definition of neighbor.
- 5:49:13We could only move left, right, up, or down.
- 5:49:15But this is, in fact, better.
- 5:49:17It has a total cost of 9.
- 5:49:18It is now closer to both of these houses.
- 5:49:21And as a result, the total cost is less.
- 5:49:24But we weren't able to find it, because in order to get there,
- 5:49:27we had to go through a state that actually wasn't any better than the current
- 5:49:31state that we had been on previously.
- 5:49:33And so this appears to be a limitation, or a concern you might have
- 5:49:36as you go about trying to implement a hill climbing algorithm,
- 5:49:39is that it might not always give you the optimal solution.
- 5:49:43If we're trying to maximize the value of any particular state,
- 5:49:46we're trying to find the global maximum, a concern
- 5:49:49might be that we could get stuck at one of the local maxima,
- 5:49:53highlighted here in blue, where a local maxima is any state whose value is
- 5:49:57higher than any of its neighbors.
- 5:49:59If we ever find ourselves at one of these two states
- 5:50:02when we're trying to maximize the value of the state,
- 5:50:04we're not going to make any changes.
- 5:50:05We're not going to move left or right.
- 5:50:07We're not going to move left here, because those states are worse.
- 5:50:10But yet, we haven't found the global optimum.
- 5:50:13We haven't done as best as we could do.
- 5:50:15And likewise, in the case of the hospitals, what we're ultimately
- 5:50:18trying to do is find a global minimum, find a value that
- 5:50:20is lower than all of the others.
- 5:50:22But we have the potential to get stuck at one of the local minima,
- 5:50:26any of these states whose value is lower than all of its neighbors,
- 5:50:30but still not as low as the local minima.
- 5:50:33And so the takeaway here is that it's not always
- 5:50:36going to be the case that when we run this naive hill climbing algorithm,
- 5:50:40that we're always going to find the optimal solution.
- 5:50:42There are things that could go wrong.
- 5:50:43If we started here, for example, and tried to maximize our value as much
- 5:50:47as possible, we might move to the highest possible neighbor,
- 5:50:50move to the highest possible neighbor, move to the highest possible neighbor,
- 5:50:54and stop, and never realize that there's actually a better state way over there
- 5:50:57that we could have gone to instead.
- 5:51:00And other problems you might imagine just by taking a look at this state
- 5:51:03space landscape are these various different types of plateaus,
- 5:51:06something like this flat local maximum here,
- 5:51:09where all six of these states each have the exact same value.
- 5:51:12And so in the case of the algorithm we showed before,
- 5:51:15none of the neighbors are better, so we might just
- 5:51:17get stuck at this flat local maximum.
- 5:51:19And even if you allowed yourself to move to one of the neighbors,
- 5:51:22it wouldn't be clear which neighbor you would ultimately move to,
- 5:51:25and you could get stuck here as well.
- 5:51:27And there's another one over here.
- 5:51:28This one is called a shoulder.
- 5:51:30It's not really a local maximum, because there's still
- 5:51:32places where we can go higher, not a local minimum, because we can go lower.
- 5:51:35So we can still make progress, but it's still this flat area,
- 5:51:38where if you have a local search algorithm,
- 5:51:40there's potential to get lost here, unable to make some upward or downward
- 5:51:44progress, depending on whether we're trying to maximize or minimize it,
- 5:51:48and therefore another potential for us to be
- 5:51:50able to find a solution that might not actually be the optimal solution.
- 5:51:54And so because of this potential, the potential that hill climbing
- 5:51:57has to not always find us the optimal result,
- 5:52:00it turns out there are a number of different varieties and variations
- 5:52:03on the hill climbing algorithm that help to solve the problem better
- 5:52:07depending on the context, and depending on the specific type of problem,
- 5:52:10some of these variants might be more applicable than others.
- 5:52:13What we've taken a look at so far is a version of hill climbing
- 5:52:16generally called steepest ascent hill climbing,
- 5:52:19where the idea of steepest ascent hill climbing
- 5:52:21is we are going to choose the highest valued neighbor,
- 5:52:24in the case where we're trying to maximize or the lowest valued neighbor
- 5:52:27in cases where we're trying to minimize.
- 5:52:28But generally speaking, if I have five neighbors
- 5:52:31and they're all better than my current state,
- 5:52:33I will pick the best one of those five.
- 5:52:36Now, sometimes that might work pretty well.
- 5:52:37It's sort of a greedy approach of trying to take the best operation
- 5:52:40at any particular time step, but it might not always work.
- 5:52:43There might be cases where actually I want
- 5:52:45to choose an option that is slightly better than me,
- 5:52:47but maybe not the best one because that later on might
- 5:52:50lead to a better outcome ultimately.
- 5:52:52So there are other variants that we might consider
- 5:52:54of this basic hill climbing algorithm.
- 5:52:56One is known as stochastic hill climbing.
- 5:52:58And in this case, we choose randomly from all of our higher value neighbors.
- 5:53:02So if I'm at my current state and there are five neighbors that
- 5:53:04are all better than I am, rather than choosing the best one,
- 5:53:07as steep as the set would do, stochastic will just choose
- 5:53:10randomly from one of them, thinking that if it's better, then it's better.
- 5:53:13And maybe there's a potential to make forward progress,
- 5:53:16even if it is not locally the best option I could possibly choose.
- 5:53:20First choice hill climbing ends up just choosing the very first highest
- 5:53:24valued neighbor that it follows, behaving on a similar idea,
- 5:53:27rather than consider all of the neighbors.
- 5:53:28As soon as we find a neighbor that is better than our current state,
- 5:53:31we'll go ahead and move there.
- 5:53:33There may be some efficiency improvements there
- 5:53:35and maybe has the potential to find a solution
- 5:53:37that the other strategies weren't able to find.
- 5:53:39And with all of these variants, we still suffer from the same potential risk,
- 5:53:43this risk that we might end up at a local minimum or a local maximum.
- 5:53:48And we can reduce that risk by repeating the process multiple times.
- 5:53:52So one variant of hill climbing is random restart hill climbing,
- 5:53:55where the general idea is we'll conduct hill climbing multiple times.
- 5:53:59If we apply steepest descent hill climbing, for example,
- 5:54:02we'll start at some random state, try and figure out
- 5:54:04how to solve the problem and figure out what
- 5:54:06is the local maximum or local minimum we get to.
- 5:54:09And then we'll just randomly restart and try again,
- 5:54:11choose a new starting configuration, try and figure out
- 5:54:14what the local maximum or minimum is, and do this some number of times.
- 5:54:17And then after we've done it some number of times,
- 5:54:19we can pick the best one out of all of the ones that we've taken a look at.
- 5:54:23So there's another option we have access to as well.
- 5:54:26And then, although I said that generally local search will usually
- 5:54:29just keep track of a single node and then move to one of its neighbors,
- 5:54:33there are variants of hill climbing that are known as local beam searches,
- 5:54:36where rather than keep track of just one current best state,
- 5:54:39we're keeping track of k highest valued neighbors, such that rather than
- 5:54:43starting at one random initial configuration,
- 5:54:46I might start with 3 or 4 or 5, randomly generate all the neighbors,
- 5:54:50and then pick the 3 or 4 or 5 best of all of the neighbors that I find,
- 5:54:54and continually repeat this process, with the idea
- 5:54:57being that now I have more options that I'm considering,
- 5:55:00more ways that I could potentially navigate myself
- 5:55:02to the optimal solution that might exist for a particular problem.
- 5:55:07So let's now take a look at some actual code that
- 5:55:09can implement some of these kinds of ideas, something
- 5:55:11like steepest ascent hill climbing, for example,
- 5:55:14for trying to solve this hospital problem.
- 5:55:17So I'm going to go ahead and go into my hospitals directory, where
- 5:55:20I've actually set up the basic framework for solving this type of problem.
- 5:55:24I'll go ahead and go into hospitals.py, and we'll
- 5:55:26take a look at the code we've created here.
- 5:55:28I've defined a class that is going to represent the state space.
- 5:55:32So the space has a height, and a width, and also some number of hospitals.
- 5:55:36So you can configure how big is your map, how many hospitals should go here.
- 5:55:41We have a function for adding a new house to the state space,
- 5:55:44and then some functions that are going to get
- 5:55:45me all of the available spaces for if I want to randomly place hospitals
- 5:55:49in particular locations.
- 5:55:50And here now is the hill climbing algorithm.
- 5:55:54So what are we going to do in the hill climbing algorithm?
- 5:55:56Well, we're going to start by randomly initializing
- 5:55:59where the hospitals are going to go.
- 5:56:01We don't know where the hospitals should actually be,
- 5:56:03so let's just randomly place them.
- 5:56:05So here I'm running a loop for each of the hospitals that I have.
- 5:56:08I'm going to go ahead and add a new hospital at some random location.
- 5:56:13So I basically get all of the available spaces,
- 5:56:15and I randomly choose one of them as where
- 5:56:17I would like to add this particular hospital.
- 5:56:20I have some logging output and generating some images,
- 5:56:23which we'll take a look at a little bit later.
- 5:56:25But here is the key idea.
- 5:56:27So I'm going to just keep repeating this algorithm.
- 5:56:30I could specify a maximum of how many times I want it to run,
- 5:56:33or I could just run it up until it hits a local maximum or local minimum.
- 5:56:37And now we'll basically consider all of the hospitals
- 5:56:40that could potentially move.
- 5:56:41So consider each of the two hospitals or more hospitals
- 5:56:43if they're more than that.
- 5:56:45And consider all of the places where that hospital could move to,
- 5:56:49some neighbor of that hospital that we can move the neighbor to.
- 5:56:53And then see, is this going to be better than where we were currently?
- 5:56:58So if it is going to be better, then we'll
- 5:57:00go ahead and update our best neighbor and keep
- 5:57:02track of this new best neighbor that we found.
- 5:57:05And then afterwards, we can ask ourselves the question,
- 5:57:08if best neighbor cost is greater than or equal
- 5:57:10to the cost of the current set of hospitals,
- 5:57:13meaning if the cost of our best neighbor is greater than the current cost,
- 5:57:18meaning our best neighbor is worse than our current state,
- 5:57:21well, then we shouldn't make any changes at all.
- 5:57:23And we should just go ahead and return the current set of hospitals.
- 5:57:27But otherwise, we can update our hospitals
- 5:57:29in order to change them to one of the best neighbors.
- 5:57:32And if there are multiple that are all equivalent,
- 5:57:34I'm here using random.choice to say go ahead and choose one randomly.
- 5:57:38So this is really just a Python implementation of that same idea
- 5:57:41that we were just talking about, this idea of taking a current state,
- 5:57:44some current set of hospitals, generating all of the neighbors,
- 5:57:48looking at all of the ways we could take one hospital
- 5:57:50and move it one square to the left or right or up or down,
- 5:57:53and then figuring out, based on all of that information, which
- 5:57:56is the best neighbor or the set of all the best neighbors,
- 5:57:59and then choosing from one of those.
- 5:58:02And each time, we go ahead and generate an image in order to do that.
- 5:58:05And so now what we're doing is if we look down at the bottom,
- 5:58:08I'm going to randomly generate a space with height 10 and width 20.
- 5:58:12And I'll say go ahead and put three hospitals somewhere in the space.
- 5:58:16I'll randomly generate 15 houses that I just go ahead
- 5:58:18and add in random locations.
- 5:58:20And now I'm going to run this hill climbing algorithm in order
- 5:58:23to try and figure out where we should place those hospitals.
- 5:58:27So we'll go ahead and run this program by running Python hospitals.
- 5:58:31And we see that we started.
- 5:58:32Our initial state had a cost of 72, but we
- 5:58:35were able to continually find neighbors that were able to decrease that cost,
- 5:58:38decrease to 69, 66, 63, so on and so forth, all the way down to 53,
- 5:58:43as the best neighbor we were able to ultimately find.
- 5:58:46And we can take a look at what that looked like
- 5:58:48by just opening up these files.
- 5:58:50So here, for example, was the initial configuration.
- 5:58:53We randomly selected a location for each of these 15 different houses
- 5:58:57and then randomly selected locations for one, two, three hospitals
- 5:59:01that were just located somewhere inside of the state space.
- 5:59:04And if you add up all the distances from each of the houses
- 5:59:07to their nearest hospital, you get a total cost of about 72.
- 5:59:11And so now the question is, what neighbors can we move to
- 5:59:14that improve the situation?
- 5:59:16And it looks like the first one the algorithm found
- 5:59:18was by taking this house that was over there on the right
- 5:59:21and just moving it to the left.
- 5:59:23And that probably makes sense because if you
- 5:59:25look at the houses in that general area, really these five houses look like
- 5:59:29they're probably the ones that are going to be closest to this hospital over here.
- 5:59:33Moving it to the left decreases the total distance, at least
- 5:59:36to most of these houses, though it does increase that distance for one of them.
- 5:59:40And so we're able to make these improvements to the situation
- 5:59:43by continually finding ways that we can move these hospitals around
- 5:59:47until we eventually settle at this particular state that
- 5:59:50has a cost of 53, where we figured out a position for each of the hospitals.
- 5:59:54And now none of the neighbors that we could move to
- 5:59:57are actually going to improve the situation.
- 5:59:59We can take this hospital and this hospital and that hospital
- 6:00:02and look at each of the neighbors.
- 6:00:03And none of those are going to be better than this particular configuration.
- 6:00:07And again, that's not to say that this is the best we could do.
- 6:00:10There might be some other configuration of hospitals
- 6:00:12that is a global minimum.
- 6:00:14And this might just be a local minimum that is the best of all of its neighbors,
- 6:00:18but maybe not the best in the entire possible state space.
- 6:00:21And you could search through the entire state space
- 6:00:24by considering all of the possible configurations for hospitals.
- 6:00:27But ultimately, that's going to be very time intensive,
- 6:00:29especially as our state space gets bigger and there
- 6:00:31might be more and more possible states.
- 6:00:33It's going to take quite a long time to look through all of them.
- 6:00:36And so being able to use these sort of local search algorithms
- 6:00:39can often be quite good for trying to find the best solution we can do.
- 6:00:42And especially if we don't care about doing the best possible
- 6:00:45and we just care about doing pretty good and finding
- 6:00:47a pretty good placement of those hospitals,
- 6:00:50then these methods can be particularly powerful.
- 6:00:53But of course, we can try and mitigate some of this concern
- 6:00:56by instead of using hill climbing to use random restart,
- 6:00:59this idea of rather than just hill climb one time,
- 6:01:02we can hill climb multiple times and say,
- 6:01:04try hill climbing a whole bunch of times on the exact same map
- 6:01:07and figure out what is the best one that we've been able to find.
- 6:01:10And so I've here implemented a function for random restart
- 6:01:14that restarts some maximum number of times.
- 6:01:17And what we're going to do is repeat that number of times this process of just
- 6:01:22go ahead and run the hill climbing algorithm,
- 6:01:24figure out what the cost is of getting from all the houses to the hospitals,
- 6:01:28and then figure out is this better than we've done so far.
- 6:01:31So I can try this exact same idea where instead of running hill climbing,
- 6:01:35I'll go ahead and run random restart.
- 6:01:37And I'll randomly restart maybe 20 times, for example.
- 6:01:41And we'll go ahead and now I'll remove all the images
- 6:01:44and then rerun the program.
- 6:01:46And now we started by finding a original state.
- 6:01:49When we initially ran hill climbing, the best cost
- 6:01:51we were able to find was 56.
- 6:01:53Each of these iterations is a different iteration of the hill climbing
- 6:01:56algorithm.
- 6:01:57We're running hill climbing not one time, but 20 times here,
- 6:02:00each time going until we find a local minimum in this case.
- 6:02:04And we look and see each time did we do better
- 6:02:06than we did the best time we've done so far.
- 6:02:09So we went from 56 to 46.
- 6:02:11This one was greater, so we ignored it.
- 6:02:12This one was 41, which was less, so we went ahead and kept that one.
- 6:02:16And for all of the remaining 16 times that we
- 6:02:18tried to implement hill climbing and we tried to run the hill climbing
- 6:02:21algorithm, we couldn't do any better than that 41.
- 6:02:25Again, maybe there is a way to do better that we just didn't find,
- 6:02:28but it looks like that way ended up being a pretty good solution
- 6:02:31to the problem.
- 6:02:32That was attempt number three, starting from counting at zero.
- 6:02:36So we can take a look at that, open up number three.
- 6:02:39And this was the state that happened to have a cost of 41,
- 6:02:42that after running the hill climbing algorithm
- 6:02:45on some particular random initial configuration of hospitals,
- 6:02:48this is what we found was the local minimum in terms
- 6:02:51of trying to minimize the cost.
- 6:02:53And it looks like we did pretty well.
- 6:02:54This hospital is pretty close to this region.
- 6:02:56This one is pretty close to these houses here.
- 6:02:58This hospital looks about as good as we can do
- 6:03:01for trying to capture those houses over on that side.
- 6:03:03And so these sorts of algorithms can be quite useful
- 6:03:06for trying to solve these problems.
- 6:03:09But the real problem with many of these different types of hill climbing,
- 6:03:12steepest of sense, stochastic, first choice, and so forth,
- 6:03:15is that they never make a move that makes our situation worse.
- 6:03:18They're always going to take ourselves in our current state,
- 6:03:21look at the neighbors, and consider can we do better than our current state
- 6:03:24and move to one of those neighbors.
- 6:03:26Which of those neighbors we choose might vary among these various different
- 6:03:29types of algorithms, but we never go from a current position
- 6:03:32to a position that is worse than our current position.
- 6:03:35And ultimately, that's what we're going to need to do
- 6:03:37if we want to be able to find a global maximum or a global minimum.
- 6:03:40Because sometimes if we get stuck, we want
- 6:03:42to find some way of dislodging ourselves from our local maximum
- 6:03:46or local minimum in order to find the global maximum or the global minimum
- 6:03:50or increase the probability that we do find it.
- 6:03:52And so the most popular technique for trying
- 6:03:54to approach the problem from that angle is a technique known
- 6:03:57as simulated annealing, simulated because it's modeling
- 6:04:00after a real physical process of annealing, where you can think about this
- 6:04:03in terms of physics, a physical situation where
- 6:04:06you have some system of particles.
- 6:04:08And you might imagine that when you heat up
- 6:04:10a particular physical system, there's a lot of energy there.
- 6:04:12Things are moving around quite randomly.
- 6:04:14But over time, as the system cools down, it eventually
- 6:04:17settles into some final position.
- 6:04:20And that's going to be the general idea of simulated annealing.
- 6:04:23We're going to simulate that process of some high temperature system where
- 6:04:27things are moving around randomly quite frequently,
- 6:04:29but over time decreasing that temperature until we eventually
- 6:04:32settle at our ultimate solution.
- 6:04:35And the idea is going to be if we have some state space landscape that
- 6:04:38looks like this and we begin at its initial state here,
- 6:04:42if we're looking for a global maximum and we're
- 6:04:44trying to maximize the value of the state,
- 6:04:46our traditional hill climbing algorithms would just take the state
- 6:04:50and look at the two neighbor ones and always
- 6:04:52pick the one that is going to increase the value of the state.
- 6:04:55But if we want some chance of being able to find the global maximum,
- 6:04:58we can't always make good moves.
- 6:05:01We have to sometimes make bad moves and allow ourselves
- 6:05:04to make a move in a direction that actually seems for now
- 6:05:08to make our situation worse such that later we
- 6:05:11can find our way up to that global maximum in terms
- 6:05:14of trying to solve that problem.
- 6:05:16Of course, once we get up to this global maximum,
- 6:05:18once we've done a whole lot of the searching,
- 6:05:20then we probably don't want to be moving to states
- 6:05:22that are worse than our current state.
- 6:05:24And so this is where this metaphor for annealing
- 6:05:26starts to come in, where we want to start making more random moves
- 6:05:30and over time start to make fewer of those random moves based
- 6:05:33on a particular temperature schedule.
- 6:05:36So the basic outline looks something like this.
- 6:05:38Early on in simulated annealing, we have a higher temperature state.
- 6:05:42And what we mean by a higher temperature state
- 6:05:44is that we are more likely to accept neighbors that
- 6:05:47are worse than our current state.
- 6:05:49We might look at our neighbors.
- 6:05:50And if one of our neighbors is worse than the current state,
- 6:05:53especially if it's not all that much worse,
- 6:05:54if it's pretty close but just slightly worse,
- 6:05:57then we might be more likely to accept that and go ahead
- 6:05:59and move to that neighbor anyways.
- 6:06:02But later on as we run simulated annealing,
- 6:06:04we're going to decrease that temperature.
- 6:06:06And at a lower temperature, we're going to be less likely to accept neighbors
- 6:06:10that are worse than our current state.
- 6:06:12Now to formalize this and put a little bit of pseudocode to it,
- 6:06:15here is what that algorithm might look like.
- 6:06:17We have a function called simulated annealing
- 6:06:19that takes as input the problem we're trying to solve
- 6:06:21and also potentially some maximum number of times
- 6:06:24we might want to run the simulated annealing process, how many different
- 6:06:27neighbors we're going to try and look for.
- 6:06:29And that value is going to vary based on the problem you're trying to solve.
- 6:06:33We'll, again, start with some current state
- 6:06:34that will be equal to the initial state of the problem.
- 6:06:37But now we need to repeat this process over and over
- 6:06:40for max number of times.
- 6:06:42Repeat some process some number of times where we're first
- 6:06:45going to calculate a temperature.
- 6:06:48And this temperature function takes the current time t
- 6:06:51starting at 1 going all the way up to max
- 6:06:53and then gives us some temperature that we can use in our computation,
- 6:06:57where the idea is that this temperature is going to be higher early on
- 6:07:01and it's going to be lower later on.
- 6:07:02So there are a number of ways this temperature function could often work.
- 6:07:05One of the simplest ways is just to say it
- 6:07:07is like the proportion of time that we still have remaining.
- 6:07:10Out of max units of time, how much time do we have remaining?
- 6:07:14You start off with a lot of that time remaining.
- 6:07:16And as time goes on, the temperature is going to decrease
- 6:07:18because you have less and less of that remaining time still available to you.
- 6:07:22So we calculate a temperature for the current time.
- 6:07:25And then we pick a random neighbor of the current state.
- 6:07:28No longer are we going to be picking the best neighbor that we possibly can
- 6:07:31or just one of the better neighbors that we can.
- 6:07:33We're going to pick a random neighbor.
- 6:07:34It might be better.
- 6:07:35It might be worse.
- 6:07:36But we're going to calculate that.
- 6:07:37We're going to calculate delta E, E for energy in this case,
- 6:07:40which is just how much better is the neighbor than the current state.
- 6:07:45So if delta E is positive, that means the neighbor
- 6:07:47is better than our current state.
- 6:07:49If delta E is negative, that means the neighbor
- 6:07:51is worse than our current state.
- 6:07:53And so we can then have a condition that looks like this.
- 6:07:56If delta E is greater than 0, that means the neighbor state
- 6:07:59is better than our current state.
- 6:08:01And if ever that situation arises, we'll just go ahead and update current
- 6:08:05to be that neighbor.
- 6:08:06Same as before, move where we are currently to be the neighbor
- 6:08:09because the neighbor is better than our current state.
- 6:08:11We'll go ahead and accept that.
- 6:08:13But now the difference is that whereas before, we never,
- 6:08:16ever wanted to take a move that made our situation worse,
- 6:08:19now we sometimes want to make a move that is actually
- 6:08:22going to make our situation worse because sometimes we're
- 6:08:24going to need to dislodge ourselves from a local minimum or local maximum
- 6:08:27to increase the probability that we're able to find the global minimum
- 6:08:31or the global maximum a little bit later.
- 6:08:34And so how do we do that?
- 6:08:35How do we decide to sometimes accept some state that might actually be worse?
- 6:08:39Well, we're going to accept a worse state with some probability.
- 6:08:43And that probability needs to be based on a couple of factors.
- 6:08:46It needs to be based in part on the temperature,
- 6:08:49where if the temperature is higher, we're more likely to move to a worse
- 6:08:52neighbor.
- 6:08:52And if the temperature is lower, we're less likely to move to a worse neighbor.
- 6:08:56But it also, to some degree, should be based on delta E.
- 6:09:00If the neighbor is much worse than the current state,
- 6:09:03we probably want to be less likely to choose that
- 6:09:05than if the neighbor is just a little bit worse than the current state.
- 6:09:09So again, there are a couple of ways you could calculate this.
- 6:09:12But it turns out one of the most popular is just
- 6:09:14to calculate E to the power of delta E over T, where E is just a constant.
- 6:09:19Delta E over T are based on delta E and T here.
- 6:09:22We calculate that value.
- 6:09:24And that'll be some value between 0 and 1.
- 6:09:26And that is the probability with which we should just say, all right,
- 6:09:29let's go ahead and move to that neighbor.
- 6:09:31And it turns out that if you do the math for this value,
- 6:09:33when delta E is such that the neighbor is not
- 6:09:36that much worse than the current state, that's
- 6:09:38going to be more likely that we're going to go ahead and move to that state.
- 6:09:41And likewise, when the temperature is lower,
- 6:09:43we're going to be less likely to move to that neighboring state as well.
- 6:09:47So now this is the big picture for simulated annealing,
- 6:09:49this process of taking the problem and going ahead and generating
- 6:09:53random neighbors will always move to a neighbor
- 6:09:55if it's better than our current state.
- 6:09:56But even if the neighbor is worse than our current state,
- 6:09:59we'll sometimes move there depending on how much worse it is
- 6:10:03and also based on the temperature.
- 6:10:04And as a result, the hope, the goal of this whole process
- 6:10:07is that as we begin to try and find our way to the global maximum
- 6:10:11or the global minimum, we can dislodge ourselves
- 6:10:14if we ever get stuck at a local maximum or local minimum
- 6:10:17in order to eventually make our way to exploring
- 6:10:19the part of the state space that is going to be the best.
- 6:10:22And then as the temperature decreases, eventually we settle there
- 6:10:25without moving around too much from what we've
- 6:10:27found to be the globally best thing that we can do thus far.
- 6:10:31So at the very end, we just return whatever the current state happens to be.
- 6:10:35And that is the conclusion of this algorithm.
- 6:10:37We've been able to figure out what the solution is.
- 6:10:40And these types of algorithms have a lot of different applications.
- 6:10:44Any time you can take a problem and formulate it
- 6:10:46as something where you can explore a particular configuration
- 6:10:49and then ask, are any of the neighbors better
- 6:10:51than this current configuration and have some way of measuring that,
- 6:10:54then there is an applicable case for these hill climbing, simulated annealing
- 6:10:58types of algorithms.
- 6:10:59So sometimes it can be for facility location type problems,
- 6:11:02like for when you're trying to plan a city and figure out
- 6:11:05where the hospitals should be.
- 6:11:06But there are definitely other applications as well.
- 6:11:08And one of the most famous problems in computer science
- 6:11:11is the traveling salesman problem.
- 6:11:13Traveling salesman problem generally is formulated like this.
- 6:11:16I have a whole bunch of cities here indicated by these dots.
- 6:11:19And what I'd like to do is find some route that
- 6:11:22takes me through all of the cities and ends up back where I started.
- 6:11:25So some route that starts here, goes through all these cities,
- 6:11:29and ends up back where I originally started.
- 6:11:32And what I might like to do is minimize the total distance
- 6:11:35that I have to travel or the total cost of taking this entire path.
- 6:11:40And you can imagine this is a problem that's very applicable in situations
- 6:11:43like when delivery companies are trying to deliver things
- 6:11:46to a whole bunch of different houses, they
- 6:11:48want to figure out, how do I get from the warehouse
- 6:11:51to all these various different houses and get back again,
- 6:11:53all using as minimal time and distance and energy as possible.
- 6:11:57So you might want to try to solve these sorts of problems.
- 6:12:00But it turns out that solving this particular kind of problem
- 6:12:03is very computationally difficult.
- 6:12:05It is a very computationally expensive task to be able to figure it out.
- 6:12:09This falls under the category of what are known as NP-complete problems,
- 6:12:12problems that there is no known efficient way to try and solve
- 6:12:16these sorts of problems.
- 6:12:17And so what we ultimately have to do is come up with some approximation,
- 6:12:21some ways of trying to find a good solution, even if we're not
- 6:12:25going to find the globally best solution that we possibly can,
- 6:12:27at least not in a feasible or tractable amount of time.
- 6:12:30And so what we could do is take the traveling salesman problem
- 6:12:34and try to formulate it using local search and ask a question like, all right,
- 6:12:38I can pick some state, some configuration, some route between all
- 6:12:41of these nodes.
- 6:12:42And I can measure the cost of that state, figure out what the distance is.
- 6:12:46And I might now want to try to minimize that cost as much as possible.
- 6:12:49And then the only question now is, what does it
- 6:12:51mean to have a neighbor of this state?
- 6:12:54What does it mean to take this particular route
- 6:12:55and have some neighboring route that is close to it but slightly different
- 6:12:59and such that it might have a different total distance?
- 6:13:01And there are a number of different definitions
- 6:13:03for what a neighbor of a traveling salesman configuration might look like.
- 6:13:07But one way is just to say, a neighbor is
- 6:13:09what happens if we pick two of these edges between nodes
- 6:13:13and switch them effectively.
- 6:13:16So for example, I might pick these two edges here,
- 6:13:19these two that just happened across this node goes here, this node goes there,
- 6:13:23and go ahead and switch them.
- 6:13:24And what that process will generally look like
- 6:13:26is removing both of these edges from the graph, taking this node,
- 6:13:31and connecting it to the node it wasn't connected to.
- 6:13:33So connecting it up here instead.
- 6:13:35We'll need to take these arrows that were originally
- 6:13:37going this way and reverse them, so move them going the other way,
- 6:13:40and then just fill in that last remaining blank,
- 6:13:42add an arrow that goes in that direction instead.
- 6:13:45So by taking two edges and just switching them,
- 6:13:48I have been able to consider one possible neighbor
- 6:13:51of this particular configuration.
- 6:13:53And it looks like this neighbor is actually better.
- 6:13:55It looks like this probably travels a shorter distance in order
- 6:13:57to get through all the cities through this route
- 6:14:00than the current state did.
- 6:14:02And so you could imagine implementing this idea inside of a hill climbing
- 6:14:05or simulated annealing algorithm, where we repeat this process
- 6:14:08to try and take a state of this traveling salesman problem,
- 6:14:11look at all the neighbors, and then move to the neighbors if they're better,
- 6:14:14or maybe even move to the neighbors if they're worse,
- 6:14:16until we eventually settle upon some best solution
- 6:14:20that we've been able to find.
- 6:14:21And it turns out that these types of approximation algorithms,
- 6:14:24even if they don't always find the very best solution,
- 6:14:26can often do pretty well at trying to find solutions that are helpful too.
- 6:14:32So that then was a look at local search, a particular category of algorithms
- 6:14:36that can be used for solving a particular type of problem,
- 6:14:38where we don't really care about the path to the solution.
- 6:14:41I didn't care about the steps I took to decide
- 6:14:43where the hospitals should go.
- 6:14:44I just cared about the solution itself.
- 6:14:46I just care about where the hospitals should be,
- 6:14:49or what the route through the traveling salesman journey really ought to be.
- 6:14:53Another type of algorithm that might come up
- 6:14:55are known as these categories of linear programming types of problems.
- 6:14:59And linear programming often comes up in the context
- 6:15:01where we're trying to optimize for some mathematical function.
- 6:15:04But oftentimes, linear programming will come up
- 6:15:07when we might have real numbered values.
- 6:15:10So it's not just discrete fixed values that we might have,
- 6:15:13but any decimal values that we might want to be able to calculate.
- 6:15:16And so linear programming is a family of types of problems
- 6:15:19where we might have a situation that looks like this, where
- 6:15:22the goal of linear programming is to minimize a cost function.
- 6:15:26And you can invert the numbers and say try and maximize it,
- 6:15:29but often we'll frame it as trying to minimize a cost function that
- 6:15:32has some number of variables, x1, x2, x3, all the way up to xn,
- 6:15:36just some number of variables that are involved,
- 6:15:38things that I want to know the values to.
- 6:15:41And this cost function might have coefficients
- 6:15:43in front of those variables.
- 6:15:45And this is what we would call a linear equation,
- 6:15:47where we just have all of these variables that might be multiplied
- 6:15:50by a coefficient and then add it together.
- 6:15:52We're not going to square anything or cube anything,
- 6:15:53because that'll give us different types of equations.
- 6:15:56With linear programming, we're just dealing with linear equations
- 6:15:59in addition to linear constraints, where a constraint is going
- 6:16:03to look something like if we sum up this particular equation that
- 6:16:07is just some linear combination of all of these variables,
- 6:16:10it is less than or equal to some bound b.
- 6:16:13And we might have a whole number of these various different constraints
- 6:16:16that we might place onto our linear programming exercise.
- 6:16:21And likewise, just as we can have constraints that are saying this linear
- 6:16:24equation is less than or equal to some bound b,
- 6:16:27it might also be equal to something.
- 6:16:28That if you want some sum of some combination of variables
- 6:16:31to be equal to a value, you can specify that.
- 6:16:33And we can also maybe specify that each variable has lower and upper bounds,
- 6:16:37that it needs to be a positive number, for example,
- 6:16:39or it needs to be a number that is less than 50, for example.
- 6:16:42And there are a number of other choices that we
- 6:16:44can make there for defining what the bounds of a variable are.
- 6:16:47But it turns out that if you can take a problem
- 6:16:50and formulate it in these terms, formulate the problem as your goal
- 6:16:54is to minimize a cost function, and you're
- 6:16:56minimizing that cost function subject to particular constraints,
- 6:17:00subjects to equations that are of the form like this of some sequence
- 6:17:03of variables is less than a bound or is equal to some particular value,
- 6:17:07then there are a number of algorithms that already
- 6:17:10exist for solving these sorts of problems.
- 6:17:13So let's go ahead and take a look at an example.
- 6:17:16Here's an example of a problem that might come up
- 6:17:18in the world of linear programming.
- 6:17:19Often, this is going to come up when we're
- 6:17:21trying to optimize for something.
- 6:17:23And we want to be able to do some calculations,
- 6:17:25and we have constraints on what we're trying to optimize.
- 6:17:27And so it might be something like this.
- 6:17:29In the context of a factory, we have two machines, x1 and x2.
- 6:17:34x1 costs $50 an hour to run.
- 6:17:36x2 costs $80 an hour to run.
- 6:17:38And our goal, what we're trying to do, our objective,
- 6:17:41is to minimize the total cost.
- 6:17:45So that's what we'd like to do.
- 6:17:46But we need to do so subject to certain constraints.
- 6:17:49So there might be a labor constraint that x1
- 6:17:51requires five units of labor per hour, x2 requires two units of labor per hour,
- 6:17:56and we have a total of 20 units of labor that we have to spend.
- 6:18:00So this is a constraint.
- 6:18:01We have no more than 20 units of labor that we can spend,
- 6:18:04and we have to spend it across x1 and x2, each of which
- 6:18:08requires a different amount of labor.
- 6:18:10And we might also have a constraint like this
- 6:18:13that tells us x1 is going to produce 10 units of output per hour,
- 6:18:16x2 is going to produce 12 units of output per hour,
- 6:18:19and the company needs 90 units of output.
- 6:18:22So we have some goal, something we need to achieve.
- 6:18:24We need to achieve 90 units of output, but there are some constraints
- 6:18:28that x1 can only produce 10 units of output per hour,
- 6:18:31x2 produces 12 units of output per hour.
- 6:18:34These types of problems come up quite frequently,
- 6:18:36and you can start to notice patterns in these types of problems,
- 6:18:39problems where I am trying to optimize for some goal, minimizing cost,
- 6:18:43maximizing output, maximizing profits, or something like that.
- 6:18:46And there are constraints that are placed on that process.
- 6:18:50And so now we just need to formulate this problem
- 6:18:52in terms of linear equations.
- 6:18:55So let's start with this first point.
- 6:18:56Two machines, x1 and x2, x costs $50 an hour, x2 costs $80 an hour.
- 6:19:01Here we can come up with an objective function that might look like this.
- 6:19:05This is our cost function, rather.
- 6:19:0750 times x1 plus 80 times x2, where x1 is going
- 6:19:11to be a variable representing how many hours do we run machine x1 for,
- 6:19:15x2 is going to be a variable representing how many hours
- 6:19:18are we running machine x2 for.
- 6:19:20And what we're trying to minimize is this cost function, which
- 6:19:23is just how much it costs to run each of these machines per hour summed up.
- 6:19:27This is an example of a linear equation, just some combination
- 6:19:31of these variables plus coefficients that are placed in front of them.
- 6:19:34And I would like to minimize that total value.
- 6:19:37But I need to do so subject to these constraints.
- 6:19:40x1 requires 50 units of labor per hour, x2 requires 2,
- 6:19:44and we have a total of 20 units of labor to spend.
- 6:19:46And so that gives us a constraint of this form.
- 6:19:505 times x1 plus 2 times x2 is less than or equal to 20.
- 6:19:5420 is the total number of units of labor we have to spend.
- 6:19:57And that's spent across x1 and x2, each of which
- 6:20:00requires a different number of units of labor per hour, for example.
- 6:20:05And finally, we have this constraint here.
- 6:20:07x1 produces 10 units of output per hour, x2 produces 12,
- 6:20:10and we need 90 units of output.
- 6:20:13And so this might look something like this.
- 6:20:15That 10x1 plus 12x2, this is amount of output per hour,
- 6:20:20it needs to be at least 90.
- 6:20:21We can do better or great, but it needs to be at least 90.
- 6:20:25And if you recall from my formulation before,
- 6:20:27I said that generally speaking in linear programming,
- 6:20:29we deal with equals constraints or less than or equal to constraints.
- 6:20:33So we have a greater than or equal to sign here.
- 6:20:35That's not a problem.
- 6:20:36Whenever we have a greater than or equal to sign,
- 6:20:38we can just multiply the equation by negative 1,
- 6:20:40and that'll flip it around to a less than or equals negative 90,
- 6:20:44for example, instead of a greater than or equal to 90.
- 6:20:47And that's going to be an equivalent expression
- 6:20:49that we can use to represent this problem.
- 6:20:51So now that we have this cost function and these constraints
- 6:20:55that it's subject to, it turns out there are a number of algorithms
- 6:20:58that can be used in order to solve these types of problems.
- 6:21:02And these problems go a little bit more into geometry and linear algebra
- 6:21:05than we're really going to get into.
- 6:21:06But the most popular of these types of algorithms
- 6:21:09are simplex, which was one of the first algorithms discovered
- 6:21:12for trying to solve linear programs.
- 6:21:14And later on, a class of interior point algorithms
- 6:21:17can be used to solve this type of problem as well.
- 6:21:20The key is not to understand exactly how these algorithms work,
- 6:21:23but to realize that these algorithms exist for efficiently finding solutions
- 6:21:27any time we have a problem of this particular form.
- 6:21:30And so we can take a look, for example, at the production directory here,
- 6:21:39where here I have a file called production.py, where here I'm
- 6:21:43using scipy, which was the library for a lot of science-related functions
- 6:21:47within Python.
- 6:21:49And I can go ahead and just run this optimization function
- 6:21:52in order to run a linear program.
- 6:21:54.linprog here is going to try and solve this linear program for me,
- 6:21:58where I provide to this expression, to this function call,
- 6:22:01all of the data about my linear program.
- 6:22:03So it needs to be in a particular format, which
- 6:22:05might be a little confusing at first.
- 6:22:07But this first argument to scipy.optimize.linprogramming
- 6:22:11is the cost function, which is in this case just an array or a list that
- 6:22:15has 50 and 80, because my original cost function was 50 times x1 plus 80
- 6:22:20times x2.
- 6:22:21So I just tell Python, 50 and 80, those are the coefficients
- 6:22:25that I am now trying to optimize for.
- 6:22:27And then I provide all of the constraints.
- 6:22:30So the constraints, and I wrote them up above in comments,
- 6:22:33is the constraint 1 is 5x1 plus 2x2 is less than or equal to 20.
- 6:22:39And constraint 2 is negative 10x1 plus negative 12x2
- 6:22:44is less than or equal to negative 90.
- 6:22:47And so scipy expects these constraints to be in a particular format.
- 6:22:51It first expects me to provide all of the coefficients
- 6:22:54for the upper bound equations, ub just for upper bound,
- 6:22:58where the coefficients of the first equation
- 6:23:00are 5 and 2, because we have 5x1 and 2x2.
- 6:23:03And the coefficients for the second equation
- 6:23:06are negative 10 and negative 12, because I have negative 10x1 plus negative 12x2.
- 6:23:12And then here, we provide it as a separate argument,
- 6:23:14just to keep things separate, what the actual bound is.
- 6:23:17What is the upper bound for each of these constraints?
- 6:23:20Well, for the first constraint, the upper bound is 20.
- 6:23:22That was constraint number 1.
- 6:23:24And then for constraint number 2, the upper bound is 90.
- 6:23:28So a bit of a cryptic way of representing it.
- 6:23:30It's not quite as simple as just writing the mathematical equations.
- 6:23:33What really is being expected here are all of the coefficients
- 6:23:36and all of the numbers that are in these equations
- 6:23:39by first providing the coefficients for the cost function,
- 6:23:42then providing all the coefficients for the inequality constraints,
- 6:23:45and then providing all of the upper bounds for those inequality constraints.
- 6:23:50And once all of that information is there,
- 6:23:52then we can run any of these interior point algorithms or the simplex algorithm.
- 6:23:57Even if you don't understand how it works,
- 6:23:59you can just run the function and figure out what the result should be.
- 6:24:02And here, I said if the result is a success,
- 6:24:04we were able to solve this problem.
- 6:24:06Go ahead and print out what the value of x1 and x2 should be.
- 6:24:10Otherwise, go ahead and print out no solution.
- 6:24:13And so if I run this program by running python production.py,
- 6:24:19it takes a second to calculate.
- 6:24:21But then we see here is what the optimal solution should be.
- 6:24:24x1 should run for 1.5 hours.
- 6:24:26x2 should run for 6.25 hours.
- 6:24:30And we were able to do this by just formulating the problem
- 6:24:33as a linear equation that we were trying to optimize,
- 6:24:36some cost that we were trying to minimize,
- 6:24:38and then some constraints that were placed on that.
- 6:24:40And many, many problems fall into this category of problems
- 6:24:43that you can solve if you can just figure out how to use equations
- 6:24:47and use these constraints to represent that general idea.
- 6:24:51And that's a theme that's going to come up a couple of times today,
- 6:24:53where we want to be able to take some problem
- 6:24:55and reduce it down to some problem we know
- 6:24:57how to solve in order to begin to find a solution
- 6:25:01and to use existing methods that we can use in order
- 6:25:04to find a solution more effectively or more efficiently.
- 6:25:08And it turns out that these types of problems, where we have constraints,
- 6:25:11show up in other ways too.
- 6:25:13And there's an entire class of problems that's more generally just known
- 6:25:16as constraint satisfaction problems.
- 6:25:18And we're going to now take a look at how you might formulate a constraint
- 6:25:21satisfaction problem and how you might go about solving a constraint
- 6:25:24satisfaction problem.
- 6:25:26But the basic idea of a constraint satisfaction problem
- 6:25:28is we have some number of variables that need to take on some values.
- 6:25:32And we need to figure out what values each of those variables should take on.
- 6:25:35But those variables are subject to particular constraints
- 6:25:39that are going to limit what values those variables can actually take on.
- 6:25:43So let's take a look at a real world example, for example.
- 6:25:46Let's look at exam scheduling, that I have
- 6:25:48four students here, students 1, 2, 3, and 4.
- 6:25:51Each of them is taking some number of different classes.
- 6:25:53Classes here are going to be represented by letters.
- 6:25:56So student 1 is enrolled in courses A, B, and C. Student 2
- 6:26:00is enrolled in courses B, D, and E, so on and so forth.
- 6:26:04And now, say university, for example, is trying
- 6:26:07to schedule exams for all of these courses.
- 6:26:10But there are only three exam slots on Monday, Tuesday, and Wednesday.
- 6:26:13And we have to schedule an exam for each of these courses.
- 6:26:17But the constraint now, the constraint we
- 6:26:19have to deal with with the scheduling, is
- 6:26:21that we don't want anyone to have to take two exams on the same day.
- 6:26:25We would like to try and minimize that or eliminate it if at all possible.
- 6:26:29So how do we begin to represent this idea?
- 6:26:31How do we structure this in a way that a computer with an AI algorithm
- 6:26:35can begin to try and solve the problem?
- 6:26:37Well, let's in particular just look at these classes that we might take
- 6:26:41and represent each of the courses as some node inside of a graph.
- 6:26:45And what we'll do is we'll create an edge between two nodes in this graph
- 6:26:49if there is a constraint between those two nodes.
- 6:26:54So what does this mean?
- 6:26:55Well, we can start with student 1, who's enrolled in courses A, B, and C.
- 6:26:59What that means is that A and B can't have an exam at the same time.
- 6:27:03A and C can't have an exam at the same time.
- 6:27:06And B and C also can't have an exam at the same time.
- 6:27:09And I can represent that in this graph by just drawing edges.
- 6:27:12One edge between A and B, one between B and C,
- 6:27:15and then one between C and A. And that encodes now the idea
- 6:27:18that between those nodes, there is a constraint.
- 6:27:21And in particular, the constraint happens to be
- 6:27:23that these two can't be equal to each other,
- 6:27:25though there are other types of constraints that are possible,
- 6:27:28depending on the type of problem that you're trying to solve.
- 6:27:31And then we can do the same thing for each of the other students.
- 6:27:34So for student 2, who's enrolled in courses B, D, and E,
- 6:27:36well, that means B, D, and E, those all need
- 6:27:39to have edges that connect each other as well.
- 6:27:41Student 3 is enrolled in courses C, E, and F. So we'll go ahead
- 6:27:44and take C, E, and F and connect those by drawing edges between them too.
- 6:27:48And then finally, student 4 is enrolled in courses E, F, and G.
- 6:27:52And we can represent that by drawing edges between E, F, and G,
- 6:27:55although E and F already had an edge between them.
- 6:27:57We don't need another one, because this constraint
- 6:27:59is just encoding the idea that course E and course F cannot have
- 6:28:03an exam on the same day.
- 6:28:05So this then is what we might call the constraint graph.
- 6:28:09There's some graphical representation of all of my variables,
- 6:28:13so to speak, and the constraints between those possible variables.
- 6:28:16Where in this particular case, each of the constraints
- 6:28:19represents an inequality constraint, that an edge between B and D
- 6:28:23means whatever value the variable B takes on cannot be the value
- 6:28:27that the variable D takes on as well.
- 6:28:30So what then actually is a constraint satisfaction problem?
- 6:28:33Well, a constraint satisfaction problem is just some set of variables, x1
- 6:28:38all the way through xn, some set of domains for each of those variables.
- 6:28:42So every variable needs to take on some values.
- 6:28:45Maybe every variable has the same domain,
- 6:28:47but maybe each variable has a slightly different domain.
- 6:28:49And then there's a set of constraints, and we'll just call a set C,
- 6:28:52that is some constraints that are placed upon these variables,
- 6:28:55like x1 is not equal to x2.
- 6:28:58But there could be other forms too, like maybe x1 equals x2 plus 1
- 6:29:02if these variables are taking on numerical values in their domain,
- 6:29:05for example.
- 6:29:06The types of constraints are going to vary based on the types of problems.
- 6:29:10And constraint satisfaction shows up all over the place as well,
- 6:29:14in any situation where we have variables that
- 6:29:16are subject to particular constraints.
- 6:29:19So one popular game is Sudoku, for example, this 9 by 9 grid
- 6:29:23where you need to fill in numbers in each of these cells,
- 6:29:25but you want to make sure there's never a duplicate number in any row,
- 6:29:29or in any column, or in any grid of 3 by 3 cells, for example.
- 6:29:34So what might this look like as a constraint satisfaction problem?
- 6:29:37Well, my variables are all of the empty squares in the puzzle.
- 6:29:41So represented here is just like an x comma y coordinate, for example,
- 6:29:45as all of the squares where I need to plug in a value,
- 6:29:48where I don't know what value it should take on.
- 6:29:50The domain is just going to be all of the numbers from 1 through 9,
- 6:29:54any value that I could fill in to one of these cells.
- 6:29:57So that is going to be the domain for each of these variables.
- 6:30:00And then the constraints are going to be of the form,
- 6:30:02like this cell can't be equal to this cell, can't be equal to this cell,
- 6:30:05can't be, and all of these need to be different, for example,
- 6:30:08and same for all of the rows, and the columns, and the 3 by 3 squares as well.
- 6:30:12So those constraints are going to enforce what values are actually allowed.
- 6:30:17And we can formulate the same idea in the case of this exam scheduling
- 6:30:21problem, where the variables we have are the different courses, a up through g.
- 6:30:25The domain for each of these variables is going to be Monday, Tuesday,
- 6:30:29and Wednesday.
- 6:30:30Those are the possible values each of the variables can take on,
- 6:30:33that in this case just represent when is the exam for that class.
- 6:30:38And then the constraints are of this form, a is not equal to b,
- 6:30:41a is not equal to c, meaning a and b can't have an exam on the same day,
- 6:30:45a and c can't have an exam on the same day.
- 6:30:48Or more formally, these two variables cannot take on the same value
- 6:30:53within their domain.
- 6:30:56So that then is this formulation of a constraint satisfaction problem
- 6:31:00that we can begin to use to try and solve this problem.
- 6:31:03And constraints can come in a number of different forms.
- 6:31:05There are hard constraints, which are constraints
- 6:31:07that must be satisfied for a correct solution.
- 6:31:10So something like in the Sudoku puzzle, you cannot have this cell
- 6:31:14and this cell that are in the same row take on the same value.
- 6:31:17That is a hard constraint.
- 6:31:18But problems can also have soft constraints,
- 6:31:21where these are constraints that express some notion of preference,
- 6:31:24that maybe a and b can't have an exam on the same day,
- 6:31:27but maybe someone has a preference that a's exam is earlier than b's exam.
- 6:31:32It doesn't need to be the case with some expression
- 6:31:34that some solution is better than another solution.
- 6:31:37And in that case, you might formulate the problem
- 6:31:39as trying to optimize for maximizing people's preferences.
- 6:31:43You want people's preferences to be satisfied as much as possible.
- 6:31:46In this case, though, we'll mostly just deal with hard constraints,
- 6:31:49constraints that must be met in order to have a correct solution to the problem.
- 6:31:54So we want to figure out some assignment of these variables
- 6:31:57to their particular values that is ultimately
- 6:32:00going to give us a solution to the problem
- 6:32:02by allowing us to assign some day to each of the classes
- 6:32:05such that we don't have any conflicts between classes.
- 6:32:09So it turns out that we can classify the constraints
- 6:32:11in a constraint satisfaction problem into a number of different categories.
- 6:32:16The first of those categories are perhaps the simplest
- 6:32:18of the types of constraints, which are known as unary constraints,
- 6:32:21where unary constraint is a constraint that just involves a single variable.
- 6:32:26For example, a unary constraint might be something like,
- 6:32:28a does not equal Monday, meaning Course A cannot have its exam on Monday.
- 6:32:33If for some reason the instructor for the course
- 6:32:35isn't available on Monday, you might have a constraint in your problem
- 6:32:38that looks like this, something that just has a single variable a in it,
- 6:32:41and maybe says a is not equal to Monday, or a is equal to something,
- 6:32:44or in the case of numbers greater than or less than something,
- 6:32:47a constraint that just has one variable, we consider to be a unary constraint.
- 6:32:51And this is in contrast to something like a binary constraint, which
- 6:32:55is a constraint that involves two variables, for example.
- 6:32:58So this would be a constraint like the ones we were looking at before.
- 6:33:01Something like a does not equal b is an example of a binary constraint,
- 6:33:06because it is a constraint that has two variables involved in it, a and b.
- 6:33:10And we represented that using some arc or some edge that
- 6:33:14connects variable a to variable b.
- 6:33:17And using this knowledge of, OK, what is a unary constraint?
- 6:33:20What is a binary constraint?
- 6:33:21There are different types of things we can
- 6:33:23say about a particular constraint satisfaction problem.
- 6:33:27And one thing we can say is we can try and make the problem node consistent.
- 6:33:31So what does node consistency mean?
- 6:33:33Node consistency means that we have all of the values
- 6:33:36in a variable's domain satisfying that variable's unary constraints.
- 6:33:41So for each of the variables inside of our constraint satisfaction problem,
- 6:33:45if all of the values satisfy the unary constraints
- 6:33:48for that particular variable, we can say that the entire problem is node
- 6:33:53consistent, or we can even say that a particular variable is
- 6:33:56node consistent if we just want to make one node consistent within itself.
- 6:34:00So what does that actually look like?
- 6:34:02Let's look at now a simplified example, where
- 6:34:04instead of having a whole bunch of different classes,
- 6:34:06we just have two classes, a and b, each of which
- 6:34:09has an exam on either Monday or Tuesday or Wednesday.
- 6:34:12So this is the domain for the variable a,
- 6:34:14and this is the domain for the variable b.
- 6:34:17And now let's imagine we have these constraints, a not equal to Monday,
- 6:34:21b not equal to Tuesday, b not equal to Monday, a not equal to b.
- 6:34:24So those are the constraints that we have on this particular problem.
- 6:34:28And what we can now try to do is enforce node consistency.
- 6:34:32And node consistency just means we make sure
- 6:34:35that all of the values for any variable's domain satisfy its unary constraints.
- 6:34:41And so we could start by trying to make node a node consistent.
- 6:34:45Is it consistent?
- 6:34:46Does every value inside of a's domain satisfy its unary constraints?
- 6:34:51Well, initially, we'll see that Monday does not satisfy a's unary constraints,
- 6:34:55because we have a constraint, a unary constraint here,
- 6:34:58that a is not equal to Monday.
- 6:35:00But Monday is still in a's domain.
- 6:35:03And so this is something that is not node consistent,
- 6:35:06because we have Monday in the domain.
- 6:35:07But this is not a valid value for this particular node.
- 6:35:11And so how do we make this node consistent?
- 6:35:13Well, to make the node consistent, what we'll do
- 6:35:15is we'll just go ahead and remove Monday from a's domain.
- 6:35:18Now a can only be on Tuesday or Wednesday,
- 6:35:21because we had this constraint that said a is not equal to Monday.
- 6:35:25And at this point now, a is node consistent.
- 6:35:28For each of the values that a can take on, Tuesday and Wednesday,
- 6:35:31there is no constraint that is a unary constraint that conflicts with that idea.
- 6:35:36There is no constraint that says that a can't be Tuesday.
- 6:35:39There is no unary constraint that says that a cannot be on Wednesday.
- 6:35:43And so now we can turn our attention to b.
- 6:35:44b also has a domain, Monday, Tuesday, and Wednesday.
- 6:35:47And we can begin to see whether those variables satisfy
- 6:35:51the unary constraints as well.
- 6:35:53Well, here is a unary constraint, b is not equal to Tuesday.
- 6:35:56And that does not appear to be satisfied by this domain of Monday, Tuesday,
- 6:35:59and Wednesday, because Tuesday, this possible value
- 6:36:03that the variable b could take on is not consistent with this unary constraint,
- 6:36:07that b is not equal to Tuesday.
- 6:36:09So to solve that problem, we'll go ahead and remove Tuesday from b's domain.
- 6:36:13Now b's domain only contains Monday and Wednesday.
- 6:36:16But as it turns out, there's yet another unary constraint
- 6:36:18that we placed on the variable b, which is here.
- 6:36:21b is not equal to Monday.
- 6:36:23And that means that this value, Monday, inside of b's domain,
- 6:36:27is not consistent with b's unary constraints,
- 6:36:30because we have a constraint that says the b cannot be Monday.
- 6:36:33And so we can remove Monday from b's domain.
- 6:36:35And now we've made it through all of the unary constraints.
- 6:36:38We've not yet considered this constraint, which is a binary constraint.
- 6:36:41But we've considered all of the unary constraints,
- 6:36:44all of the constraints that involve just a single variable.
- 6:36:47And we've made sure that every node is consistent with those unary constraints.
- 6:36:51So we can say that now we have enforced node consistency,
- 6:36:55that for each of these possible nodes, we can pick any of these values
- 6:36:59in the domain.
- 6:37:00And there won't be a unary constraint that is violated as a result of it.
- 6:37:05So node consistency is fairly easy to enforce.
- 6:37:07We just take each node, make sure the values in the domain
- 6:37:10satisfy the unary constraints.
- 6:37:12Where things get a little bit more interesting
- 6:37:14is when we consider different types of consistency,
- 6:37:17something like arc consistency, for example.
- 6:37:20And arc consistency refers to when all of the values in a variable's domain
- 6:37:25satisfy the variable's binary constraints.
- 6:37:28So when we're looking at trying to make a arc consistent,
- 6:37:31we're no longer just considering the unary constraints that involve a.
- 6:37:35We're trying to consider all of the binary constraints
- 6:37:38that involve a as well.
- 6:37:39So any edge that connects a to another variable
- 6:37:43inside of that constraint graph that we were taking a look at before.
- 6:37:47Put a little bit more formally, arc consistency.
- 6:37:50And arc really is just another word for an edge
- 6:37:52that connects two of these nodes inside of our constraint graph.
- 6:37:55We can define arc consistency a little more precisely like this.
- 6:37:59In order to make some variable x arc consistent with respect
- 6:38:03to some other variable y, we need to remove any element from x's domain
- 6:38:09to make sure that every choice for x, every choice in x's domain,
- 6:38:14has a possible choice for y.
- 6:38:17So put another way, if I have a variable x
- 6:38:19and I want to make x an arc consistent, then
- 6:38:21I'm going to look at all of the possible values that x can take on
- 6:38:25and make sure that for all of those possible values,
- 6:38:28there is still some choice that I can make for y,
- 6:38:31if there's some arc between x and y, to make sure
- 6:38:34that y has a possible option that I can choose as well.
- 6:38:39So let's look at an example of that going back to this example from before.
- 6:38:42We enforced node consistency already by saying
- 6:38:45that a can only be on Tuesday or Wednesday
- 6:38:47because we knew that a could not be on Monday.
- 6:38:49And we also said that b's only domain only
- 6:38:51consists of Wednesday because we know that b does not equal Tuesday
- 6:38:55and also b does not equal Monday.
- 6:38:58So now let's begin to consider arc consistency.
- 6:39:01Let's try and make a arc consistent with b.
- 6:39:05And what that means is to make a arc consistent with respect to b
- 6:39:08means that for any choice we make in a's domain,
- 6:39:11there is some choice we can make in b's domain that is going to be consistent.
- 6:39:16And we can try that.
- 6:39:17For a, we can choose Tuesday as a possible value for a.
- 6:39:20If I choose Tuesday for a, is there a value
- 6:39:23for b that satisfies the binary constraint?
- 6:39:26Well, yes, b Wednesday would satisfy this constraint
- 6:39:29that a does not equal b because Tuesday does not equal Wednesday.
- 6:39:33However, if we chose Wednesday for a, well, then
- 6:39:37there is no choice in b's domain that satisfies this binary constraint.
- 6:39:42There is no way I can choose something for b that satisfies a does not equal b
- 6:39:47because I know b must be Wednesday.
- 6:39:49And so if ever I run into a situation like this
- 6:39:52where I see that here is a possible value for a such
- 6:39:55that there is no choice of value for b that satisfies the binary constraint,
- 6:39:59well, then this is not arc consistent.
- 6:40:02And to make it arc consistent, I would need to take Wednesday
- 6:40:05and remove it from a's domain.
- 6:40:07Because Wednesday was not going to be a possible choice I can make for a
- 6:40:11because it wasn't consistent with this binary constraint for b.
- 6:40:14There was no way I could choose Wednesday for a
- 6:40:17and still have an available solution by choosing something for b as well.
- 6:40:22So here now, I've been able to enforce arc consistency.
- 6:40:25And in doing so, I've actually solved this entire problem,
- 6:40:28that given these constraints where a and b can have exams on either Monday
- 6:40:32or Tuesday or Wednesday, the only solution, as it would appear,
- 6:40:35is that a's exam must be on Tuesday and b's exam must be on Wednesday.
- 6:40:40And that is the only option available to me.
- 6:40:43So if we want to apply our consistency to a larger graph,
- 6:40:46not just looking at one particular pair of our consistency,
- 6:40:49there are ways we can do that too.
- 6:40:51And we can begin to formalize what the pseudocode would look like
- 6:40:53for trying to write an algorithm that enforces arc consistency.
- 6:40:57And we'll start by defining a function called revise.
- 6:41:01Revise is going to take as input a CSP, otherwise
- 6:41:03known as a constraint satisfaction problem,
- 6:41:06and also two variables, x and y.
- 6:41:08And what revise is going to do is it is going
- 6:41:11to make x arc consistent with respect to y,
- 6:41:15meaning remove anything from x's domain that
- 6:41:18doesn't allow for a possible option for y.
- 6:41:21How does this work?
- 6:41:22Well, we'll go ahead and first keep track of whether or not
- 6:41:25we've made a revision.
- 6:41:26Revise is ultimately going to return true or false.
- 6:41:29It'll return true in the event that we did make a revision to x's domain.
- 6:41:33It'll return false if we didn't make any change to x's domain.
- 6:41:37And we'll see in a moment why that's going to be helpful.
- 6:41:39But we start by saying revised equals false.
- 6:41:41We haven't made any changes.
- 6:41:43Then we'll say, all right, let's go ahead and loop over all
- 6:41:46of the possible values in x's domain.
- 6:41:49So loop over x's domain for each little x in x's domain.
- 6:41:53I want to make sure that for each of those choices,
- 6:41:55I have some available choice in y that satisfies the binary constraints that
- 6:42:00are defined inside of my CSP, inside of my constraint
- 6:42:03satisfaction problem.
- 6:42:05So if ever it's the case that there is no value y in y's domain that
- 6:42:11satisfies the constraint for x and y, well, if that's the case,
- 6:42:15that means that this value x shouldn't be in x's domain.
- 6:42:19So we'll go ahead and delete x from x's domain.
- 6:42:22And I'll set revised equal to true because I did change x's domain.
- 6:42:26I changed x's domain by removing little x.
- 6:42:29And I removed little x because it wasn't art consistent.
- 6:42:33There was no way I could choose a value for y
- 6:42:35that would satisfy this xy constraint.
- 6:42:38So in this case, we'll go ahead and set revised equal true.
- 6:42:41And we'll do this again and again for every value in x's domain.
- 6:42:44Sometimes it might be fine.
- 6:42:46In other cases, it might not allow for a possible choice for y,
- 6:42:49in which case we need to remove this value from x's domain.
- 6:42:53And at the end, we just return revised to indicate whether or not
- 6:42:56we actually made a change.
- 6:42:59So this function, then, this revised function
- 6:43:01is effectively an implementation of what you saw me do graphically a moment ago.
- 6:43:04And it makes one variable, x, arc consistent with another variable,
- 6:43:09in this case, y.
- 6:43:10But generally speaking, when we want to enforce our consistency,
- 6:43:14we'll often want to enforce our consistency not just for a single arc,
- 6:43:17but for the entire constraint satisfaction problem.
- 6:43:20And it turns out there's an algorithm to do that as well.
- 6:43:22And that algorithm is known as AC3.
- 6:43:25AC3 takes a constraint satisfaction problem.
- 6:43:27And it enforces our consistency across the entire problem.
- 6:43:32How does it do that?
- 6:43:33Well, it's going to basically maintain a queue or basically just a line
- 6:43:36of all of the arcs that it needs to make consistent.
- 6:43:39And over time, we might remove things from that queue
- 6:43:42as we begin dealing with our consistency.
- 6:43:44And we might need to add things to that queue as well
- 6:43:47if there are more things we need to make arc consistent.
- 6:43:50So we'll go ahead and start with a queue that
- 6:43:52contains all of the arcs in the constraint satisfaction problem,
- 6:43:56all of the edges that connect two nodes that
- 6:43:58have some sort of binary constraint between them.
- 6:44:02And now, as long as the queue is non-empty, there is work to be done.
- 6:44:06The queue is all of the things that we need to make arc consistent.
- 6:44:10So as long as the queue is non-empty, there's still things we have to do.
- 6:44:13What do we have to do?
- 6:44:15Well, we'll start by de-queuing from the queue,
- 6:44:17remove something from the queue.
- 6:44:19And strictly speaking, it doesn't need to be a queue,
- 6:44:21but a queue is a traditional way of doing this.
- 6:44:23We'll de-queue from the queue, and that'll give us an arc, x and y,
- 6:44:27these two variables where I would like to make x arc consistent with y.
- 6:44:32So how do we make x arc consistent with y?
- 6:44:35Well, we can go ahead and just use that revise function
- 6:44:38that we talked about a moment ago.
- 6:44:39We called the revise function, passing as input the constraint satisfaction
- 6:44:43problem, and also these variables x and y,
- 6:44:46because I want to make x arc consistent with y.
- 6:44:49In other words, remove any values from x's domain
- 6:44:52that don't leave an available option for y.
- 6:44:55And recall, what does revised return?
- 6:44:57Well, it returns true if we actually made a change,
- 6:45:00if we removed something from x's domain, because there
- 6:45:04wasn't an available option for y, for example.
- 6:45:06And it returns false if we didn't make any change to x's domain at all.
- 6:45:10And it turns out if revised returns false, if we didn't make any changes,
- 6:45:14well, then there's not a whole lot more work
- 6:45:15to be done here for this arc.
- 6:45:17We can just move ahead to the next arc that's in the queue.
- 6:45:20But if we did make a change, if we did reduce x's domain
- 6:45:24by removing values from x's domain, well, then what we might realize
- 6:45:28is that this creates potential problems later on,
- 6:45:31that it might mean that some arc that was arc consistent with x,
- 6:45:35that node might no longer be arc consistent with x,
- 6:45:38because while there used to be an option that we could choose for x,
- 6:45:41now there might not be, because now we might have removed something
- 6:45:44from x that was necessary for some other arc to be arc consistent.
- 6:45:49And so if ever we did revise x's domain,
- 6:45:52we're going to need to add some things to the queue, some additional arcs
- 6:45:55that we might want to check.
- 6:45:57How do we do that?
- 6:45:58Well, first thing we want to check is to make sure that x's domain is not 0.
- 6:46:02If x's domain is 0, that means there are no available options for x at all.
- 6:46:07And that means that there's no way you can solve the constraint satisfaction
- 6:46:10problem.
- 6:46:10If we've removed everything from x's domain,
- 6:46:13we'll go ahead and just return false here to indicate there's
- 6:46:15no way to solve the problem, because there's nothing left in x's domain.
- 6:46:19But otherwise, if there are things left in x's domain,
- 6:46:23but fewer things than before, well, then what we'll do
- 6:46:26is we'll loop over each variable z that is in all of x's neighbors,
- 6:46:31except for y, y we already handled.
- 6:46:33But we'll consider all of x's other's neighbors and ask ourselves,
- 6:46:37all right, will that arc from each of those z's to x,
- 6:46:41that arc might no longer be arc consistent,
- 6:46:43because while for each z, there might have been a possible option
- 6:46:46we could choose for x to correspond with each of z's possible values,
- 6:46:50now there might not be, because we removed some elements from x's domain.
- 6:46:54And so what we'll do here is we'll go ahead and enqueue,
- 6:46:57adding something to the queue, this arc zx for all of those neighbors z.
- 6:47:02So we need to add back some arcs to the queue
- 6:47:05in order to continue to enforce arc consistency.
- 6:47:08At the very end, if we make it through all this process,
- 6:47:11then we can return true.
- 6:47:13But this now is AC3, this algorithm for enforcing arc consistency
- 6:47:18on a constraint satisfaction problem.
- 6:47:20And the big idea is really just keep track of all of the arcs
- 6:47:23that we might need to make arc consistent,
- 6:47:25make it arc consistent by calling the revise function.
- 6:47:28And if we did revise it, then there are some new arcs
- 6:47:31that might need to be added to the queue in order
- 6:47:33to make sure that everything is still arc consistent, even
- 6:47:36after we've removed some of the elements from a particular variable's
- 6:47:40domain.
- 6:47:42So what then would happen if we tried to enforce arc consistency
- 6:47:46on a graph like this, on a graph where each of these variables
- 6:47:48has a domain of Monday, Tuesday, and Wednesday?
- 6:47:51Well, it turns out that by enforcing arc consistency on this graph,
- 6:47:55well, it can solve some types of problems.
- 6:47:57Nothing actually changes here.
- 6:47:59For any particular arc, just considering two variables,
- 6:48:03there's always a way for me to just, for any of the choices
- 6:48:05I make for one of them, make a choice for the other one,
- 6:48:08because there are three options, and I just need the two
- 6:48:11to be different from each other.
- 6:48:12So this is actually quite easy to just take an arc
- 6:48:15and just declare that it is arc consistent,
- 6:48:17because if I pick Monday for D, then I just
- 6:48:19pick something that isn't Monday for B. In arc consistency,
- 6:48:23we only consider consistency between a binary constraint between two nodes,
- 6:48:28and we're not really considering all of the rest of the nodes yet.
- 6:48:32So just using AC3, the enforcement of arc consistency,
- 6:48:36that can sometimes have the effect of reducing domains
- 6:48:39to make it easier to find solutions, but it will not always actually
- 6:48:42solve the problem.
- 6:48:44We might still need to somehow search to try and find a solution.
- 6:48:48And we can use classical traditional search algorithms to try to do so.
- 6:48:52You'll recall that a search problem generally consists of these parts.
- 6:48:55We have some initial state, some actions, a transition model
- 6:48:59that takes me from one state to another state,
- 6:49:01a goal test to tell me have I satisfied my objective correctly,
- 6:49:05and then some path cost function, because in the case of like maze solving,
- 6:49:09I was trying to get to my goal as quickly as possible.
- 6:49:12So you could formulate a CSP, or a constraint satisfaction problem,
- 6:49:16as one of these types of search problems.
- 6:49:18The initial state will just be an empty assignment,
- 6:49:22where an assignment is just a way for me to assign any particular variable
- 6:49:26to any particular value.
- 6:49:27So if an empty assignment is no variables that are assigned to any values
- 6:49:30yet, then the action I can take is adding some new variable equals value
- 6:49:37pair to that assignment, saying for this assignment,
- 6:49:40let me add a new value for this variable.
- 6:49:43And the transition model just defines what happens when you take that action.
- 6:49:46You get a new assignment that has that variable equal to that value inside
- 6:49:50of it.
- 6:49:51The goal test is just checking to make sure all the variables have been assigned
- 6:49:54and making sure all the constraints have been satisfied.
- 6:49:57And the path cost function is sort of irrelevant.
- 6:50:00I don't really care about what the path really is.
- 6:50:02I just care about finding some assignment that actually satisfies
- 6:50:06all of the constraints.
- 6:50:07So really, all the paths have the same cost.
- 6:50:09I don't really care about the path to the goal.
- 6:50:12I just care about the solution itself, much as we've talked about now before.
- 6:50:17The problem here, though, is that if we just implement this naive search
- 6:50:20algorithm just by implementing like breadth-first search or depth-first
- 6:50:23search, this is going to be very, very inefficient.
- 6:50:25And there are ways we can take advantage of efficiencies
- 6:50:28in the structure of a constraint satisfaction problem itself.
- 6:50:31And one of the key ideas is that we can really just order these variables.
- 6:50:37And it doesn't matter what order we assign variables in.
- 6:50:39The assignment a equals 2 and then b equals 8
- 6:50:43is identical to the assignment of b equals 8 and then a equals 2.
- 6:50:47Switching the order doesn't really change anything
- 6:50:50about the fundamental nature of that assignment.
- 6:50:53And so there are some ways that we can try and revise
- 6:50:56this idea of a search algorithm to apply it specifically
- 6:50:59for a problem like a constraint satisfaction problem.
- 6:51:02And it turns out the search algorithm we'll generally
- 6:51:04use when talking about constraint satisfaction problems
- 6:51:06is something known as backtracking search.
- 6:51:09And the big idea of backtracking search is we'll
- 6:51:11go ahead and make assignments from variables to values.
- 6:51:14And if ever we get stuck, we arrive at a place
- 6:51:17where there is no way we can make any forward progress while still
- 6:51:20preserving the constraints that we need to enforce,
- 6:51:23we'll go ahead and backtrack and try something else instead.
- 6:51:27So the very basic sketch of what backtracking search looks like
- 6:51:30is it looks like this.
- 6:51:32Function called backtrack that takes as input an assignment
- 6:51:35and a constraint satisfaction problem.
- 6:51:37So initially, we don't have any assigned variables.
- 6:51:40So when we begin backtracking search, this assignment
- 6:51:42is just going to be the empty assignment with no variables inside of it.
- 6:51:46But we'll see later this is going to be a recursive function.
- 6:51:49So backtrack takes as input the assignment and the problem.
- 6:51:53If the assignment is complete, meaning all of the variables have been assigned,
- 6:51:57we just return that assignment.
- 6:51:59That, of course, won't be true initially,
- 6:52:00because we start with an empty assignment.
- 6:52:02But over time, we might add things to that assignment.
- 6:52:05So if ever the assignment actually is complete, then we're done.
- 6:52:08Then just go ahead and return that assignment.
- 6:52:10But otherwise, there is some work to be done.
- 6:52:13So what we'll need to do is select an unassigned variable
- 6:52:17for this particular problem.
- 6:52:18So we need to take the problem, look at the variables that have already
- 6:52:21been assigned, and pick a variable that has not yet been assigned.
- 6:52:26And I'll go ahead and take that variable.
- 6:52:28And then I need to consider all of the values in that variable's domain.
- 6:52:32So we'll go ahead and call this domain values function.
- 6:52:34We'll talk a little more about that later, that takes a variable
- 6:52:37and just gives me back an ordered list of all of the values in its domain.
- 6:52:42So I've taken a random unselected variable.
- 6:52:44I'm going to loop over all of the possible values.
- 6:52:47And the idea is, let me just try all of these values
- 6:52:50as possible values for the variable.
- 6:52:53So if the value is consistent with the assignment so far,
- 6:52:56it doesn't violate any of the constraints,
- 6:52:59well then let's go ahead and add variable equals value to the assignment
- 6:53:02because it's so far consistent.
- 6:53:04And now let's recursively call backtrack to try and make
- 6:53:08the rest of the assignments also consistent.
- 6:53:10So I'll go ahead and call backtrack on this new assignment
- 6:53:13that I've added the variable equals value to.
- 6:53:17And now I recursively call backtrack and see what the result is.
- 6:53:20And if the result isn't a failure, well then let me just return that result.
- 6:53:27And otherwise, what else could happen?
- 6:53:30Well, if it turns out the result was a failure, well then
- 6:53:32that means this value was probably a bad choice
- 6:53:35for this particular variable because when I assigned
- 6:53:37this variable equal to that value, eventually down the road
- 6:53:41I ran into a situation where I violated constraints.
- 6:53:43There was nothing more I could do.
- 6:53:45So now I'll remove variable equals value from the assignment,
- 6:53:48effectively backtracking to say, all right, that value didn't work.
- 6:53:52Let's try another value instead.
- 6:53:55And then at the very end, if we were never
- 6:53:57able to return a complete assignment, we'll just go ahead and return failure
- 6:54:00because that means that none of the values worked for this particular
- 6:54:04variable.
- 6:54:05This now is the idea for backtracking search,
- 6:54:07to take each of the variables, try values for them,
- 6:54:10and recursively try backtracking search, see if we can make progress.
- 6:54:14And if ever we run into a dead end, we run
- 6:54:16into a situation where there is no possible value we can choose
- 6:54:19that satisfies the constraints, we return failure.
- 6:54:22And that propagates up, and eventually we
- 6:54:24make a different choice by going back and trying something else instead.
- 6:54:29So let's put this algorithm into practice.
- 6:54:31Let's actually try and use backtracking search to solve this problem now,
- 6:54:35where I need to figure out how to assign each of these courses
- 6:54:37to an exam slot on Monday or Tuesday or Wednesday in such a way
- 6:54:41that it satisfies these constraints, that each of these edges
- 6:54:44mean those two classes cannot have an exam on the same day.
- 6:54:47So I can start by just starting at a node.
- 6:54:50It doesn't really matter which I start with,
- 6:54:51but in this case, I'll just start with A.
- 6:54:54And I'll ask the question, all right, let me loop over the values in the domain.
- 6:54:57And maybe in this case, I'll just start with Monday and say, all right,
- 6:55:00let's go ahead and assign A to Monday.
- 6:55:02We'll just go and order Monday, Tuesday, Wednesday.
- 6:55:04And now let's consider node B. So I've made an assignment to A,
- 6:55:08so I recursively call backtrack with this new part of the assignment.
- 6:55:11And now I'm looking to pick another unassigned variable like B.
- 6:55:14And I'll say, all right, maybe I'll start with Monday,
- 6:55:16because that's the very first value in B's domain.
- 6:55:18And I ask, all right, does Monday violate any constraints?
- 6:55:22And it turns out, yes, it does.
- 6:55:23It violates this constraint here between A and B,
- 6:55:26because A and B are now both on Monday, and that doesn't work,
- 6:55:29because B can't be on the same day as A. So that doesn't work.
- 6:55:33So we might instead try Tuesday, try the next value in B's domain.
- 6:55:37And is that consistent with the assignment so far?
- 6:55:39Well, yeah, B, Tuesday, A, Monday, that is consistent so far,
- 6:55:43because they're not on the same day.
- 6:55:44So that's good.
- 6:55:45Now we can recursively call backtrack.
- 6:55:47Try again.
- 6:55:48Pick another unassigned variable, something like D, and say, all right,
- 6:55:51let's go through its possible values.
- 6:55:53Is Monday consistent with this assignment?
- 6:55:55Well, yes, it is.
- 6:55:56B and D are on different days, Monday versus Tuesday.
- 6:55:59And A and B are also on different days, Monday versus Tuesday.
- 6:56:02So that's fine so far, too.
- 6:56:04We'll go ahead and try again.
- 6:56:05Maybe we'll go to this variable here, E. Say, can we make that consistent?
- 6:56:09Let's go through the possible values.
- 6:56:10We've recursively called backtrack.
- 6:56:12We might start with Monday and say, all right, that's not consistent,
- 6:56:15because D and E now have exams on the same day.
- 6:56:19So we might try Tuesday instead, going to the next one.
- 6:56:21Ask, is that consistent?
- 6:56:23Well, no, it's not, because B and E, those have exams on the same day.
- 6:56:27And so we try, all right, is Wednesday consistent?
- 6:56:29And in turn, it's like, all right, yes, it is.
- 6:56:31Wednesday is consistent, because D and E now
- 6:56:33have exams on different days.
- 6:56:34B and E now have exams on different days.
- 6:56:37All seems to be well so far.
- 6:56:38I recursively call backtrack, select another unassigned variable,
- 6:56:43we'll say maybe choose C this time, and say, all right,
- 6:56:45let's try the values that C could take on.
- 6:56:48Let's start with Monday.
- 6:56:49And it turns out that's not consistent, because now A and C both
- 6:56:53have exams on the same day.
- 6:56:55So I try Tuesday and say, that's not consistent either,
- 6:56:57because B and C now have exams on the same day.
- 6:57:00And then I say, all right, let's go ahead and try Wednesday.
- 6:57:04But that's not consistent either, because C and E each have
- 6:57:08exams on the same day too.
- 6:57:09So now we've gone through all the possible values for C, Monday, Tuesday,
- 6:57:13and Wednesday.
- 6:57:14And none of them are consistent.
- 6:57:15There is no way we can have a consistent assignment.
- 6:57:18Backtrack, in this case, will return a failure.
- 6:57:21And so then we'd say, all right, we have to backtrack back to here.
- 6:57:24Well, now for E, we've tried all of Monday, Tuesday, and Wednesday.
- 6:57:28And none of those work, because Wednesday, which seemed to work,
- 6:57:31turned out to be a failure.
- 6:57:33So that means there's no possible way we can assign E.
- 6:57:36So that's a failure too.
- 6:57:37We have to go back up to D, which means that Monday assignment to D,
- 6:57:41that must be wrong.
- 6:57:41We must try something else.
- 6:57:43So we can try, all right, what if instead of Monday, we try Tuesday?
- 6:57:47Tuesday, it turns out, is not consistent,
- 6:57:49because B and D now have an exam on the same day.
- 6:57:51But Wednesday, as it turns out, works.
- 6:57:55And now we can begin to mix and forward progress again.
- 6:57:57We go back to E and say, all right, which of these values works?
- 6:58:00Monday turns out to work by not violating any constraints.
- 6:58:03Then we go up to C now.
- 6:58:05Monday doesn't work, because it violates a constraint.
- 6:58:08Violates two, actually.
- 6:58:09Tuesday doesn't work, because it violates a constraint as well.
- 6:58:12But Wednesday does work.
- 6:58:13Then we can go to the next variable, F, and say, all right, does Monday work?
- 6:58:16We'll know.
- 6:58:17It violates a constraint.
- 6:58:18But Tuesday does work.
- 6:58:19And then finally, we can look at the last variable, G,
- 6:58:21recursively calling backtrack one more time.
- 6:58:24Monday is inconsistent.
- 6:58:25That violates a constraint.
- 6:58:27Tuesday also violates a constraint.
- 6:58:29But Wednesday, that doesn't violate a constraint.
- 6:58:33And so now at this point, we recursively call backtrack one last time.
- 6:58:36We now have a satisfactory assignment of all of the variables.
- 6:58:40And at this point, we can say that we are now done.
- 6:58:42We have now been able to successfully assign a variable or a value
- 6:58:47to each one of these variables in such a way
- 6:58:49that we're not violating any constraints.
- 6:58:51We're going to go ahead and have classes A and E have their exams on Monday.
- 6:58:55Classes B and F can have their exams on Tuesday.
- 6:58:58And classes C, D, and G can have their exams on Wednesday.
- 6:59:02And there's no violated constraints that might come up there.
- 6:59:06So that then was a graphical look at how this might work.
- 6:59:08Let's now take a look at some code we could use to actually try
- 6:59:11and solve this problem as well.
- 6:59:14So here I'll go ahead and go into the scheduling directory.
- 6:59:20We're here now.
- 6:59:21We'll start by looking at schedule0.py.
- 6:59:24We're here.
- 6:59:25I define a list of variables, A, B, C, D, E, F, G.
- 6:59:28Those are all different classes.
- 6:59:31Then underneath that, I define my list of constraints.
- 6:59:34So constraint A and B. That is a constraint
- 6:59:36because they can't be on the same day.
- 6:59:38Likewise, A and C, B and C, so on and so forth,
- 6:59:40enforcing those exact same constraints.
- 6:59:43And here then is what the backtracking function might look like.
- 6:59:47First, if the assignment is complete, if I've
- 6:59:50made an assignment of every variable to a value,
- 6:59:54go ahead and just return that assignment.
- 6:59:56Then we'll select an unassigned variable from that assignment.
- 7:00:00Then for each of the possible values in the domain, Monday, Tuesday,
- 7:00:03Wednesday, let's go ahead and create a new assignment that
- 7:00:06assigns the variable to that value.
- 7:00:09I'll call this consistent function, which I'll show you in a moment,
- 7:00:11that just checks to make sure this new assignment is consistent.
- 7:00:14But if it is consistent, we'll go ahead and call backtrack
- 7:00:17to go ahead and continue trying to run backtracking search.
- 7:00:20And as long as the result is not none, meaning it wasn't a failure,
- 7:00:24we can go ahead and return that result.
- 7:00:26But if we make it through all the values and nothing works, then it is a failure.
- 7:00:31There's no solution.
- 7:00:32We go ahead and return none here.
- 7:00:35What do these functions do?
- 7:00:36Select unassigned variable is just going to choose a variable not yet assigned.
- 7:00:40So it's going to loop over all the variables.
- 7:00:42And if it's not already assigned, we'll go ahead and just return that variable.
- 7:00:46And what does the consistent function do?
- 7:00:48Well, the consistent function goes through all the constraints.
- 7:00:51And if we have a situation where we've assigned both of those values
- 7:00:56to variables, but they are the same, well,
- 7:00:59then that is a violation of the constraint, in which case we'll return false.
- 7:01:03But if nothing is inconsistent, then the assignment is consistent
- 7:01:06and will return true.
- 7:01:08And then all the program does is it calls backtrack
- 7:01:12on an empty assignment, an empty dictionary that has no variable assigned
- 7:01:15and no values yet, save that inside a solution,
- 7:01:18and then print out that solution.
- 7:01:21So by running this now, I can run Python schedule0.py.
- 7:01:27And what I get as a result of that is an assignment
- 7:01:29of all these variables to values.
- 7:01:31And it turns out we assign a to Monday as we would expect, b to Tuesday,
- 7:01:35c to Wednesday, exactly the same type of thing
- 7:01:37we were talking about before, an assignment of each of these variables
- 7:01:40to values that doesn't violate any constraints.
- 7:01:43And I had to do a fair amount of work in order
- 7:01:45to implement this idea myself.
- 7:01:47I had to write the backtrack function that went ahead
- 7:01:49and went through this process of recursively trying
- 7:01:51to do this backtracking search.
- 7:01:53But it turns out the constraint satisfaction problems are so popular
- 7:01:56that there exist many libraries that already implement this type of idea.
- 7:02:00Again, as with before, the specific library
- 7:02:03is not as important as the fact that libraries do exist.
- 7:02:06This is just one example of a Python constraint library,
- 7:02:09where now, rather than having to do all the work from scratch
- 7:02:13inside of schedule1.py, I'm just taking advantage
- 7:02:15of a library that implements a lot of these ideas already.
- 7:02:19So here, I create a new problem, add variables to it
- 7:02:22with particular domains.
- 7:02:24I add a whole bunch of these individual constraints,
- 7:02:27where I call addConstraint and pass in a function describing
- 7:02:30what the constraint is.
- 7:02:32And the constraint basically says the function that takes two variables, x
- 7:02:35and y, and makes sure that x is not equal to y,
- 7:02:38enforcing the idea that these two classes cannot have exams on the same day.
- 7:02:43And then, for any constraint satisfaction problem,
- 7:02:46I can call getSolutions to get all the solutions to that problem.
- 7:02:50And then, for each of those solutions, print out
- 7:02:53what that solution happens to be.
- 7:02:55And if I run python schedule1.py, and now see,
- 7:02:59there are actually a number of different solutions
- 7:03:01that can be used to solve the problem.
- 7:03:03There are, in fact, six different solutions, assignments of variables
- 7:03:06to values that will give me a satisfactory answer to this constraint
- 7:03:10satisfaction problem.
- 7:03:13So this then was an implementation of a very basic backtracking search method,
- 7:03:17where really we just went through each of the variables,
- 7:03:19picked one that wasn't assigned, tried the possible values
- 7:03:22the variable could take on.
- 7:03:23And then, if it worked, if it didn't violate any constraints,
- 7:03:27then we kept trying other variables.
- 7:03:28And if ever we hit a dead end, we had to backtrack.
- 7:03:31But ultimately, we might be able to be a little bit more
- 7:03:34intelligent about how we do this in order
- 7:03:36to improve the efficiency of how we solve these sorts of problems.
- 7:03:39And one thing we might imagine trying to do
- 7:03:41is going back to this idea of inference, using the knowledge we
- 7:03:44know to be able to draw conclusions in order
- 7:03:47to make the rest of the problem solving process a little bit easier.
- 7:03:51And let's now go back to where we got stuck in this problem the first time.
- 7:03:55When we were solving this constraint satisfaction problem, we dealt with B.
- 7:03:59And then we went on to D. And we went ahead and just assigned D to Monday,
- 7:04:03because that seemed to work with the assignment so far.
- 7:04:05It didn't violate any constraints.
- 7:04:07But it turned out that later on that choice turned out to be a bad one,
- 7:04:11that that choice wasn't consistent with the rest of the values
- 7:04:15that we could take on here.
- 7:04:16And the question is, is there anything we
- 7:04:18could do to avoid getting into a situation like this,
- 7:04:21avoid trying to go down a path that's ultimately not going to lead anywhere
- 7:04:25by taking advantage of knowledge that we have initially?
- 7:04:28And it turns out we do have that kind of knowledge.
- 7:04:30We can look at just the structure of this graph so far.
- 7:04:33And we can say that right now C's domain, for example,
- 7:04:37contains values Monday, Tuesday, and Wednesday.
- 7:04:41And based on those values, we can say that this graph is not arc consistent.
- 7:04:46Recall that arc consistency is all about making sure
- 7:04:49that for every possible value for a particular node,
- 7:04:52that there is some other value that we are able to choose.
- 7:04:55And as we can see here, Monday and Tuesday
- 7:04:58are not going to be possible values that we can choose for C.
- 7:05:01They're not going to be consistent with a node like B, for example,
- 7:05:06because B is equal to Tuesday, which means that C cannot be Tuesday.
- 7:05:09And because A is equal to Monday, C also cannot be Monday.
- 7:05:13So using that information, by making C arc consistent with A and B,
- 7:05:18we could remove Monday and Tuesday from C's domain
- 7:05:21and just leave C with Wednesday, for example.
- 7:05:25And if we continued to try and enforce arc consistency,
- 7:05:28we'd see there are some other conclusions we can draw as well.
- 7:05:31We see that B's only option is Tuesday and C's only option is Wednesday.
- 7:05:35And so if we want to make E arc consistent,
- 7:05:38well, E can't be Tuesday, because that wouldn't be arc consistent with B.
- 7:05:42And E can't be Wednesday, because that wouldn't be arc consistent with C.
- 7:05:45So we can go ahead and say E and just set that equal to Monday, for example.
- 7:05:49And then we can begin to do this process again and again,
- 7:05:51that in order to make D arc consistent with B and E,
- 7:05:54then D would have to be Wednesday.
- 7:05:56That's the only possible option.
- 7:05:57And likewise, we can make the same judgments for F and G as well.
- 7:06:01And it turns out that without having to do any additional search,
- 7:06:04just by enforcing arc consistency, we were
- 7:06:07able to actually figure out what the assignment of all the variables
- 7:06:10should be without needing to backtrack at all.
- 7:06:14And the way we did that is by interleaving this search process
- 7:06:18and the inference step, by this step of trying to enforce arc consistency.
- 7:06:22And the algorithm to do this is often called just the maintaining arc
- 7:06:26consistency algorithm, which just enforces arc consistency every time
- 7:06:30we make a new assignment of a value to an existing variable.
- 7:06:34So sometimes we can enforce our consistency using that AC3 algorithm
- 7:06:38at the very beginning of the problem before we even begin searching
- 7:06:41in order to limit the domain of the variables
- 7:06:43in order to make it easier to search.
- 7:06:45But we can also take advantage of the interleaving
- 7:06:48of enforcing our consistency with search such that every time in the search
- 7:06:52process we make a new assignment, we go ahead and enforce arc consistency
- 7:06:56as well to make sure that we're just eliminating
- 7:06:59possible values from domains whenever possible.
- 7:07:02And how do we do this?
- 7:07:03Well, this is really equivalent to just every time
- 7:07:06we make a new assignment to a variable x.
- 7:07:09We'll go ahead and call our AC3 algorithm,
- 7:07:12this algorithm that enforces arc consistency on a constraint satisfaction
- 7:07:15problem.
- 7:07:16And we go ahead and call that, starting it
- 7:07:18with a Q, not of all of the arcs, which we did originally,
- 7:07:22but just of all of the arcs that we want to make arc consistent with x,
- 7:07:26this thing that we have just made an assignment to.
- 7:07:28So all arcs yx, where y is a neighbor of x, something
- 7:07:33that shares a constraint with x, for example.
- 7:07:36And by maintaining arc consistency in the backtracking search process,
- 7:07:40we can ultimately make our search process a little bit more efficient.
- 7:07:44And so this is the revised version of this backtrack function.
- 7:07:47Same as before, the changes here are highlighted in yellow.
- 7:07:50Every time we add a new variable equals value to our assignment,
- 7:07:54we'll go ahead and run this inference procedure, which
- 7:07:56might do a number of different things.
- 7:07:57But one thing it could do is call the maintaining arc consistency
- 7:08:00algorithm to make sure we're able to enforce arc consistency on the problem.
- 7:08:05And we might be able to draw new inferences as a result of that process.
- 7:08:09Get new guarantees of this variable needs to be equal to that value,
- 7:08:13for example.
- 7:08:14That might happen one time.
- 7:08:15It might happen many times.
- 7:08:16And so long as those inferences are not a failure,
- 7:08:19as long as they don't lead to a situation where there is no possible way
- 7:08:22to make forward progress, well, then we can go ahead and add those inferences,
- 7:08:26those new knowledge, that new pieces of knowledge
- 7:08:28I know about what variables should be assigned to what values,
- 7:08:31I can add those to the assignment in order to more quickly make forward
- 7:08:35progress by taking advantage of information that I can just deduce,
- 7:08:38information I know based on the rest of the structure
- 7:08:41of the constraint satisfaction problem.
- 7:08:44And the only other change I'll need to make now
- 7:08:46is if it turns out this value doesn't work, well, then down here,
- 7:08:49I'll go ahead and need to remove not only variable equals value,
- 7:08:52but also any of those inferences that I made,
- 7:08:54remove that from the assignment as well.
- 7:08:57So here, then, we're often able to solve the problem by backtracking less
- 7:09:01than we might originally have needed to, just
- 7:09:03by taking advantage of the fact that every time we
- 7:09:05make a new assignment of one variable to one value,
- 7:09:08that might reduce the domains of other variables as well.
- 7:09:12And we can use that information to begin to more quickly draw conclusions
- 7:09:15in order to try and solve the problem more efficiently as well.
- 7:09:19And it turns out there are other heuristics
- 7:09:21we can use to try and improve the efficiency of our search process
- 7:09:25as well.
- 7:09:25And it really boils down to a couple of these functions
- 7:09:28that I've talked about, but we haven't really
- 7:09:30talked about how they're working.
- 7:09:32And one of them is this function here, select unassigned variable,
- 7:09:37where we're selecting some variable in the constraint satisfaction problem
- 7:09:40that has not yet been assigned.
- 7:09:42So far, I've sort of just been selecting variables randomly,
- 7:09:45just like picking one variable and one unassigned variable in order
- 7:09:48to decide, all right, this is the variable
- 7:09:50that we're going to assign next, and then going from there.
- 7:09:53But it turns out that by being a little bit intelligent,
- 7:09:55by following certain heuristics, we might be
- 7:09:57able to make the search process much more efficient just
- 7:10:00by choosing very carefully which variable we should explore next.
- 7:10:05So some of those heuristics include the minimum remaining values,
- 7:10:09or MRV heuristic, which generally says that if I
- 7:10:12have a choice between which variable I should select,
- 7:10:14I should select the variable with the smallest domain,
- 7:10:18the variable that has the fewest number of remaining values left.
- 7:10:21With the idea being, if there are only two remaining values left,
- 7:10:24well, I may as well prune one of them very quickly in order
- 7:10:27to get to the other, because one of those two has got to be the solution,
- 7:10:30if a solution does exist.
- 7:10:33Sometimes minimum remaining values might not give a conclusive result
- 7:10:37if all the nodes have the same number of remaining values, for example.
- 7:10:40And in that case, another heuristic that can be helpful to look at
- 7:10:43is the degree heuristic.
- 7:10:45The degree of a node is the number of nodes that are attached to that node,
- 7:10:49the number of nodes that are constrained by that particular node.
- 7:10:52And if you imagine which variable should I choose,
- 7:10:54should I choose a variable that has a high degree that
- 7:10:57is connected to a lot of different things,
- 7:10:59or a variable with a low degree that is not
- 7:11:01connected to a lot of different things, well,
- 7:11:03it can often make sense to choose the variable that
- 7:11:06has the highest degree that is connected to the most other nodes
- 7:11:09as the thing you would search first.
- 7:11:11Why is that the case?
- 7:11:12Well, it's because by choosing a variable with a high degree,
- 7:11:16that is immediately going to constrain the rest of the variables more,
- 7:11:20and it's more likely to be able to eliminate large sections of the state
- 7:11:23space that you don't need to search through at all.
- 7:11:26So what could this actually look like?
- 7:11:29Let's go back to this search problem here.
- 7:11:31In this particular case, I've made an assignment here.
- 7:11:34I've made an assignment here.
- 7:11:35And the question is, what should I look at next?
- 7:11:38And according to the minimum remaining values heuristic,
- 7:11:41what I should choose is the variable that has the fewest
- 7:11:44remaining possible values.
- 7:11:46And in this case, that's this node here, node
- 7:11:48C, that only has one variable left in this domain, which in this case
- 7:11:51is Wednesday, which is a very reasonable choice of a next assignment
- 7:11:55to make, because I know it's the only option, for example.
- 7:11:58I know that the only possible option for C is Wednesday,
- 7:12:01so I may as well make that assignment and then potentially explore
- 7:12:04the rest of the space after that.
- 7:12:07But meanwhile, at the very start of the problem,
- 7:12:09when I didn't have any knowledge of what nodes should have what values yet,
- 7:12:12I still had to pick what node should be the first one that I try and assign
- 7:12:16a value to.
- 7:12:17And I arbitrarily just chose the one at the top, node A originally.
- 7:12:20But we can be more intelligent about that.
- 7:12:23We can look at this particular graph.
- 7:12:25All of them have domains of the same size, domain of size 3.
- 7:12:28So minimum remaining values doesn't really help us there.
- 7:12:31But we might notice that node E has the highest degree.
- 7:12:34It is connected to the most things.
- 7:12:37And so perhaps it makes sense to begin our search,
- 7:12:39rather than starting at node A at the very top,
- 7:12:41start with the node with the highest degree.
- 7:12:43Start by searching from node E, because from there,
- 7:12:46that's going to much more easily allow us to enforce
- 7:12:49the constraints that are nearby, eliminating
- 7:12:51large portions of the search space that I might not need to search through.
- 7:12:55And in fact, by starting with E, we can immediately then assign other variables.
- 7:12:59And following that, we can actually assign the rest of the variables
- 7:13:02without needing to do any backtracking at all,
- 7:13:04even if I'm not using this inference procedure.
- 7:13:06Just by starting with a node that has a high degree,
- 7:13:09that is going to very quickly restrict the possible values
- 7:13:12that other nodes can take on.
- 7:13:14So that then is how we can go about selecting
- 7:13:17an unassigned variable in a particular order.
- 7:13:19Rather than randomly picking a variable, if we're
- 7:13:22a little bit intelligent about how we choose it,
- 7:13:24we can make our search process much, much more efficient
- 7:13:26by making sure we don't have to search through portions of the search space
- 7:13:30that ultimately aren't going to matter.
- 7:13:32The other variable we haven't really talked about,
- 7:13:34the other function here, is this domain values function.
- 7:13:37This domain values function that takes a variable
- 7:13:40and gives me back a sequence of all of the values
- 7:13:43inside of that variable's domain.
- 7:13:45The naive way to approach it is what we did before,
- 7:13:47which is just go in order, go Monday, then Tuesday, then Wednesday.
- 7:13:51But the problem is that going in that order
- 7:13:53might not be the most efficient order to search in,
- 7:13:55that sometimes it might be more efficient to choose values
- 7:13:59that are likely to be solutions first and then go to other values.
- 7:14:04Now, how do you assess whether a value is
- 7:14:06likelier to lead to a solution or less likely to lead to a solution?
- 7:14:10Well, one thing you can take a look at is how many constraints get added,
- 7:14:15how many things get removed from domains as you
- 7:14:17make this new assignment of a variable to this particular value.
- 7:14:21And the heuristic we can use here is the least constraining value heuristic,
- 7:14:26which is the idea that we should return variables in order
- 7:14:28based on the number of choices that are ruled out for neighboring values.
- 7:14:32And I want to start with the least constraining value, the value that
- 7:14:36rules out the fewest possible options.
- 7:14:40And the idea there is that if all I care about doing
- 7:14:43is finding a solution, if I start with a value that
- 7:14:47rules out a lot of other choices, I'm ruling out a lot of possibilities
- 7:14:51that maybe is going to make it less likely that this particular choice
- 7:14:55leads to a solution.
- 7:14:56Whereas on the other hand, if I have a variable
- 7:14:58and I start by choosing a value that doesn't rule out very much,
- 7:15:02well, then I still have a lot of space where there might be a solution
- 7:15:05that I could ultimately find.
- 7:15:06And this might seem a little bit counterintuitive and a little bit at odds
- 7:15:09with what we were talking about before, where I said,
- 7:15:12when you're picking a variable, you should
- 7:15:14pick the variable that is going to have the fewest possible values remaining.
- 7:15:18But here, I want to pick the value for the variable
- 7:15:20that is the least constraining.
- 7:15:22But the general idea is that when I am picking a variable,
- 7:15:25I would like to prune large portions of the search space
- 7:15:27by just choosing a variable that is going to allow me to quickly eliminate
- 7:15:30possible options.
- 7:15:32Whereas here, within a particular variable,
- 7:15:34as I'm considering values that that variable could take on,
- 7:15:37I would like to just find a solution.
- 7:15:40And so what I want to do is ultimately choose
- 7:15:42a value that still leaves open the possibility of me finding a solution
- 7:15:46to be as likely as possible.
- 7:15:48By not ruling out many options, I leave open the possibility
- 7:15:51that I can still find a solution without needing
- 7:15:54to go back later and backtrack.
- 7:15:56So an example of that might be in this particular situation here,
- 7:15:59if I'm trying to choose a variable for a value for node C here,
- 7:16:03that C is equal to either Tuesday or Wednesday.
- 7:16:06We know it can't be Monday because it conflicts with this domain here,
- 7:16:09where we already know that A is Monday, so C must be Tuesday or Wednesday.
- 7:16:13And the question is, should I try Tuesday first,
- 7:16:16or should I try Wednesday first?
- 7:16:18And if I try Tuesday, what gets ruled out?
- 7:16:21Well, one option gets ruled out here, a second option gets ruled out here,
- 7:16:25and a third option gets ruled out here.
- 7:16:27So choosing Tuesday would rule out three possible options.
- 7:16:30And what about choosing Wednesday?
- 7:16:32Well, choosing Wednesday would rule out one option here,
- 7:16:35and it would rule out one option there.
- 7:16:37And so I have two choices.
- 7:16:38I can choose Tuesday that rules out three options,
- 7:16:41or Wednesday that rules out two options.
- 7:16:43And according to the least constraining value heuristic,
- 7:16:46what I should probably do is go ahead and choose Wednesday,
- 7:16:49the one that rules out the fewest number of possible options,
- 7:16:52leaving open as many chances as possible for me
- 7:16:55to eventually find the solution inside of the state space.
- 7:16:58And ultimately, if you continue this process,
- 7:17:00we will find the solution, an assignment of variables, two values,
- 7:17:05that allows us to give each of these exams, each of these classes,
- 7:17:09an exam date that doesn't conflict with anyone
- 7:17:12that happens to be enrolled in two classes at the same time.
- 7:17:16So the big takeaway now with all of this is
- 7:17:18that there are a number of different ways we can formulate a problem.
- 7:17:21The ways we've looked at today are we can formulate a problem
- 7:17:24as a local search problem, a problem where we're looking at a current node
- 7:17:27and moving to a neighbor based on whether that neighbor is better
- 7:17:30or worse than the current node that we are looking at.
- 7:17:33We looked at formulating problems as linear programs,
- 7:17:35where just by putting things in terms of equations and constraints,
- 7:17:38we're able to solve problems a little bit more efficiently.
- 7:17:41And we saw formulating a problem as a constraint satisfaction problem,
- 7:17:45creating this graph of all of the constraints
- 7:17:48that connect two variables that have some constraint between them,
- 7:17:51and using that information to be able to figure out
- 7:17:54what the solution should be.
- 7:17:56And so the takeaway of all of this now is
- 7:17:58that if we have some problem in artificial intelligence
- 7:18:00that we would like to use AI to be able to solve them,
- 7:18:03whether that's trying to figure out where hospitals should be
- 7:18:05or trying to solve the traveling salesman problem,
- 7:18:07trying to optimize productions and costs and whatnot,
- 7:18:10or trying to figure out how to satisfy certain constraints,
- 7:18:13whether that's in a Sudoku puzzle, or whether that's
- 7:18:15in trying to figure out how to schedule exams for a university,
- 7:18:18or any number of a wide variety of types of problems,
- 7:18:21if we can formulate that problem as one of these sorts of problems,
- 7:18:24then we can use these known algorithms, these algorithms
- 7:18:27for enforcing art consistency and backtracking search,
- 7:18:30these hill climbing and simulated annealing algorithms,
- 7:18:33these simplex algorithms and interior point algorithms that
- 7:18:36can be used to solve linear programs, that we
- 7:18:38can use those techniques to begin to solve a whole wide variety of problems
- 7:18:42all in this world of optimization inside of artificial intelligence.
- 7:18:46This was an introduction to artificial intelligence with Python for today.
- 7:18:49We will see you next time.
- 7:18:52["
- 7:19:11All right.
- 7:19:11Welcome back, everyone, to an introduction
- 7:19:13to artificial intelligence with Python.
- 7:19:15Now, so far in this class, we've used AI to solve
- 7:19:17a number of different problems, giving AI instructions
- 7:19:20for how to search for a solution, or how to satisfy certain constraints in order
- 7:19:24to find its way from some input point to some output point
- 7:19:27in order to solve some sort of problem.
- 7:19:29Today, we're going to turn to the world of learning,
- 7:19:31in particular the idea of machine learning, which generally refers
- 7:19:34to the idea where we are not going to give the computer explicit instructions
- 7:19:38for how to perform a task, but rather we are going to give the computer access
- 7:19:42to information in the form of data, or patterns that it can learn from,
- 7:19:45and let the computer try and figure out what those patterns are,
- 7:19:48try and understand that data to be able to perform a task on its own.
- 7:19:52Now, machine learning comes in a number of different forms,
- 7:19:54and it's a very wide field.
- 7:19:56So today, we'll explore some of the foundational algorithms and ideas
- 7:20:00that are behind a lot of the different areas within machine learning.
- 7:20:03And one of the most popular is the idea of supervised machine learning,
- 7:20:07or just supervised learning.
- 7:20:08And supervised learning is a particular type of task.
- 7:20:11It refers to the task where we give the computer access
- 7:20:14to a data set, where that data set consists of input-output pairs.
- 7:20:19And what we would like the computer to do
- 7:20:21is we would like our AI to be able to figure out
- 7:20:23some function that maps inputs to outputs.
- 7:20:27So we have a whole bunch of data that generally consists
- 7:20:29of some kind of input, some evidence, some information
- 7:20:32that the computer will have access to.
- 7:20:33And we would like the computer, based on that input information,
- 7:20:36to predict what some output is going to be.
- 7:20:40And we'll give it some data so that the computer can train its model on
- 7:20:43and begin to understand how it is that this information works
- 7:20:46and how it is that the inputs and outputs relate to each other.
- 7:20:49But ultimately, we hope that our computer
- 7:20:51will be able to figure out some function that, given those inputs,
- 7:20:54is able to get those outputs.
- 7:20:56There are a couple of different tasks within supervised learning.
- 7:20:59The one we'll focus on and start with is known as classification.
- 7:21:02And classification is the problem where, if I give you a whole bunch of inputs,
- 7:21:07you need to figure out some way to map those inputs into discrete categories,
- 7:21:11where you can decide what those categories are,
- 7:21:13and it's the job of the computer to predict what those categories are
- 7:21:16going to be.
- 7:21:17So that might be, for example, I give you information
- 7:21:19about a bank note, like a US dollar, and I'm asking you to predict for me,
- 7:21:23does it belong to the category of authentic bank notes,
- 7:21:26or does it belong to the category of counterfeit bank notes?
- 7:21:29You need to categorize the input, and we want
- 7:21:31to train the computer to figure out some function
- 7:21:33to be able to do that calculation.
- 7:21:36Another example might be the case of weather,
- 7:21:38someone we've talked about a little bit so far in this class,
- 7:21:40where we would like to predict on a given day,
- 7:21:43is it going to rain on that day?
- 7:21:44Is it going to be cloudy on that day?
- 7:21:46And before we've seen how we could do this, if we really give the computer
- 7:21:49all the exact probabilities for if these are the conditions,
- 7:21:53what's the probability of rain?
- 7:21:54Oftentimes, we don't have access to that information, though.
- 7:21:57But what we do have access to is a whole bunch of data.
- 7:22:00So if we wanted to be able to predict something like,
- 7:22:02is it going to rain or is it not going to rain,
- 7:22:04we would give the computer historical information about days
- 7:22:07when it was raining and days when it was not raining
- 7:22:10and ask the computer to look for patterns in that data.
- 7:22:14So what might that data look like?
- 7:22:15Well, we could structure that data in a table like this.
- 7:22:18This might be what our table looks like, where for any particular day,
- 7:22:21going back, we have information about that day's humidity,
- 7:22:24that day's air pressure, and then importantly, we have a label,
- 7:22:28something where the human has said that on this particular day,
- 7:22:31it was raining or it was not raining.
- 7:22:33So you could fill in this table with a whole bunch of data.
- 7:22:35And what makes this what we would call a supervised learning exercise
- 7:22:39is that a human has gone in and labeled each of these data points,
- 7:22:42said that on this day, when these were the values for the humidity and pressure,
- 7:22:45that day was a rainy day and this day was a not rainy day.
- 7:22:49And what we would like the computer to be able to do then
- 7:22:51is to be able to figure out, given these inputs, given the humidity
- 7:22:55and the pressure, can the computer predict what label
- 7:22:58should be associated with that day?
- 7:22:59Does that day look more like it's going to be a day that rains
- 7:23:02or does it look more like a day when it's not going to rain?
- 7:23:06Put a little bit more mathematically, you can think of this as a function
- 7:23:10that takes two inputs, the inputs being the data points
- 7:23:13that our computer will have access to, things like humidity and pressure.
- 7:23:16So we could write a function f that takes
- 7:23:18as input both humidity and pressure.
- 7:23:20And then the output is going to be what category
- 7:23:24we would ascribe to these particular input points, what label
- 7:23:27we would associate with that input.
- 7:23:29So we've seen a couple of example data points here,
- 7:23:31where given this value for humidity and this value for pressure,
- 7:23:34we predict, is it going to rain or is it not going to rain?
- 7:23:37And that's information that we just gathered from the world.
- 7:23:40We measured on various different days what the humidity and pressure were.
- 7:23:44We observed whether or not we saw rain or no rain on that particular day.
- 7:23:48And this function f is what we would like to approximate.
- 7:23:51Now, the computer and we humans don't really
- 7:23:53know exactly how this function f works.
- 7:23:55It's probably quite a complex function.
- 7:23:57So what we're going to do instead is attempt to estimate it.
- 7:24:01We would like to come up with a hypothesis function.
- 7:24:03h, which is going to try to approximate what f does.
- 7:24:08We want to come up with some function h that will also take the same inputs
- 7:24:12and will also produce an output, rain or no rain.
- 7:24:15And ideally, we'd like these two functions to agree as much as possible.
- 7:24:20So the goal then of the supervised learning classification tasks
- 7:24:23is going to be to figure out, what does that function h look like?
- 7:24:26How can we begin to estimate, given all of this information, all of this data,
- 7:24:30what category or what label should be assigned to a particular data point?
- 7:24:35So where could you begin doing this?
- 7:24:37Well, a reasonable thing to do, especially in this situation,
- 7:24:39I have two numerical values, is I could try
- 7:24:42to plot this on a graph that has two axes, an x-axis and a y-axis.
- 7:24:47And in this case, we're just going to be using two numerical values as input.
- 7:24:50But these same types of ideas scale as you add more and more inputs as well.
- 7:24:54We'll be plotting things in two dimensions.
- 7:24:56But as we soon see, you could add more inputs
- 7:24:58and just imagine things in multiple dimensions.
- 7:25:00And while we humans have trouble conceptualizing anything really
- 7:25:04beyond three dimensions, at least visually,
- 7:25:06a computer has no problem with trying to imagine things
- 7:25:08in many, many more dimensions, that for a computer,
- 7:25:11each dimension is just some separate number that it is keeping track of.
- 7:25:14So it wouldn't be unreasonable for a computer to think in 10 dimensions
- 7:25:17or 100 dimensions to be able to try to solve a problem.
- 7:25:20But for now, we've got two inputs.
- 7:25:22So we'll graph things along two axes, an x-axis, which will here
- 7:25:25represent humidity, and a y-axis, which here represents pressure.
- 7:25:29And what we might do is say, let's take all of the days
- 7:25:32that were raining and just try to plot them on this graph
- 7:25:35and see where they fall on this graph.
- 7:25:37And here might be all of the rainy days, where each rainy day is
- 7:25:40one of these blue dots here that corresponds
- 7:25:42to a particular value for humidity and a particular value for pressure.
- 7:25:46And then I might do the same thing with the days that were not rainy.
- 7:25:49So take all the not rainy days, figure out
- 7:25:51what their values were for each of these two inputs,
- 7:25:53and go ahead and plot them on this graph as well.
- 7:25:56And I've here plotted them in red.
- 7:25:58So blue here stands for a rainy day.
- 7:26:00Red here stands for a not rainy day.
- 7:26:02And this then is the input that my computer
- 7:26:04has access to all of this input.
- 7:26:07And what I would like the computer to be able to do
- 7:26:09is to train a model such that if I'm ever presented with a new input that
- 7:26:13doesn't have a label associated with it, something like this white dot here,
- 7:26:18I would like to predict, given those values for each of the two inputs,
- 7:26:21should we classify it as a blue dot, a rainy day,
- 7:26:24or should we classify it as a red dot, a not rainy day?
- 7:26:28And if you're just looking at this picture graphically, trying to say,
- 7:26:30all right, this white dot, does it look like it belongs to the blue category,
- 7:26:34or does it look like it belongs to the red category,
- 7:26:36I think most people would agree that it probably belongs to the blue category.
- 7:26:40And why is that?
- 7:26:41Well, it looks like it's close to other blue dots.
- 7:26:45And that's not a very formal notion, but it's a notion
- 7:26:47that we'll formalize in just a moment.
- 7:26:49That because it seems to be close to this blue dot here,
- 7:26:52nothing else is closer to it, then we might
- 7:26:54say that it should be categorized as blue.
- 7:26:56It should fall into that category of, I think
- 7:26:58that day is going to be a rainy day based on that input.
- 7:27:01Might not be totally accurate, but it's a pretty good guess.
- 7:27:04And this type of algorithm is actually a very popular and common machine
- 7:27:08learning algorithm known as nearest neighbor classification.
- 7:27:11It's an algorithm for solving these classification-type problems.
- 7:27:14And in nearest neighbor classification, it's going to perform this algorithm.
- 7:27:18What it will do is, given an input, it will
- 7:27:20choose the class of the nearest data point to that input.
- 7:27:24By class, we just here mean category, like rain or no rain,
- 7:27:27counterfeit or not counterfeit.
- 7:27:29And we choose the category or the class based on the nearest data point.
- 7:27:34So given all that data, we just looked at,
- 7:27:36is the nearest data point a blue point or is it a red point?
- 7:27:39And depending on the answer to that question,
- 7:27:42we were able to make some sort of judgment.
- 7:27:44We were able to say something like, we think it's going to be blue
- 7:27:47or we think it's going to be red.
- 7:27:49So likewise, we could apply this to other data points
- 7:27:51that we encounter as well.
- 7:27:52If suddenly this data point comes about, well, its nearest data is red.
- 7:27:56So we would go ahead and classify this as a red point, not raining.
- 7:28:00Things get a little bit trickier, though, when you look at a point
- 7:28:03like this white point over here and you ask the same sort of question.
- 7:28:07Should it belong to the category of blue points, the rainy days?
- 7:28:10Or should it belong to the category of red points, the not rainy days?
- 7:28:14Now, nearest neighbor classification would say the way you solve this problem
- 7:28:18is look at which point is nearest to that point.
- 7:28:21You look at this nearest point and say it's red.
- 7:28:23It's a not rainy day.
- 7:28:24And therefore, according to nearest neighbor classification,
- 7:28:27I would say that this unlabeled point, well, that should also be red.
- 7:28:30It should also be classified as a not rainy day.
- 7:28:33But your intuition might think that that's a reasonable judgment to make,
- 7:28:37that it's the closest thing is a not rainy day.
- 7:28:39So may as well guess that it's a not rainy day.
- 7:28:41But it's probably also reasonable to look at the bigger picture of things
- 7:28:44to say, yes, it is true that the nearest point to it was a red point.
- 7:28:49But it's surrounded by a whole bunch of other blue points.
- 7:28:52So looking at the bigger picture, there's potentially
- 7:28:55an argument to be made that this point should actually be blue.
- 7:28:59And with only this data, we actually don't know for sure.
- 7:29:01We are given some input, something we're trying to predict.
- 7:29:04And we don't necessarily know what the output is going to be.
- 7:29:07So in this case, which one is correct is difficult to say.
- 7:29:10But oftentimes, considering more than just a single neighbor,
- 7:29:13considering multiple neighbors can sometimes give us a better result.
- 7:29:18And so there's a variant on the nearest neighbor classification algorithm
- 7:29:21that is known as the K nearest neighbor classification algorithm,
- 7:29:25where K is some parameter, some number that we choose,
- 7:29:28for how many neighbors are we going to look at.
- 7:29:30So one nearest neighbor classification is what we saw before.
- 7:29:34Just pick the one nearest neighbor and use that category.
- 7:29:37But with K nearest neighbor classification,
- 7:29:39where K might be 3, or 5, or 7, to say look at the 3, or 5, or 7 closest
- 7:29:44neighbors, closest data points to that point, works a little bit differently.
- 7:29:48This algorithm, we'll give it an input.
- 7:29:50Choose the most common class out of the K nearest data points to that input.
- 7:29:55So if we look at the five nearest points, and three of them say it's raining,
- 7:29:59and two of them say it's not raining, we'll
- 7:30:01go with the three instead of the two, because each one effectively
- 7:30:05gets one vote towards what they believe the category ought to be.
- 7:30:09And ultimately, you choose the category that has the most votes
- 7:30:12as a consequence of that.
- 7:30:14So K nearest neighbor classification, fairly straightforward one
- 7:30:17to understand intuitively.
- 7:30:18You just look at the neighbors and figure out what the answer might be.
- 7:30:21And it turns out this can work very, very well
- 7:30:24for solving a whole variety of different types of classification problems.
- 7:30:28But not every model is going to work under every situation.
- 7:30:31And so one of the things we'll take a look at today, especially
- 7:30:33in the context of supervised machine learning,
- 7:30:35is that there are a number of different approaches to machine learning,
- 7:30:38a number of different algorithms that we can apply,
- 7:30:40all solving the same type of problem, all solving some kind of classification
- 7:30:44problem where we want to take inputs and organize it
- 7:30:47into different categories.
- 7:30:49And no one algorithm is necessarily always
- 7:30:51going to be better than some other algorithm.
- 7:30:53They each have their trade-offs.
- 7:30:54And maybe depending on the data, one type of algorithm
- 7:30:57is going to be better suited to trying to model
- 7:30:59that information than some other algorithm.
- 7:31:01And so this is what a lot of machine learning research ends up being about,
- 7:31:04that when you're trying to apply machine learning techniques,
- 7:31:06you're often looking not just at one particular algorithm,
- 7:31:09but trying multiple different algorithms,
- 7:31:11trying to see what is going to give you the best results for trying
- 7:31:14to predict some function that maps inputs to outputs.
- 7:31:18So what then are the drawbacks of K nearest neighbor classification?
- 7:31:22Well, there are a couple.
- 7:31:23One might be that in a naive approach, at least, it could be fairly slow
- 7:31:27to have to go through and measure the distance between a point
- 7:31:30and every single one of these points that exist here.
- 7:31:32Now, there are ways of trying to get around that.
- 7:31:33There are data structures that can help to make it more quickly
- 7:31:36to be able to find these neighbors.
- 7:31:38There are also techniques you can use to try and prune some of this data,
- 7:31:41remove some of the data points so that you're only
- 7:31:43left with the relevant data points just to make it a little bit easier.
- 7:31:47But ultimately, what we might like to do is come up
- 7:31:49with another way of trying to do this classification.
- 7:31:53And one way of trying to do the classification
- 7:31:55was looking at what are the neighboring points.
- 7:31:57But another way might be to try to look at all of the data
- 7:32:01and see if we can come up with some decision boundary, some boundary that
- 7:32:05will separate the rainy days from the not rainy days.
- 7:32:08And in the case of two dimensions, we can do that by drawing a line,
- 7:32:11for example.
- 7:32:12So what we might want to try to do is just find some line,
- 7:32:15find some separator that divides the rainy days, the blue points over here,
- 7:32:20from the not rainy days, the red points over there.
- 7:32:22We're now trying a different approach in contrast
- 7:32:25with the nearest neighbor approach, which just
- 7:32:27looked at local data around the input data point that we cared about.
- 7:32:31Now what we're doing is trying to use a technique known as linear regression
- 7:32:35to find some sort of line that will separate the two halves from each other.
- 7:32:39Now sometimes it'll actually be possible to come up
- 7:32:42with some line that perfectly separates all the rainy days
- 7:32:45from the not rainy days.
- 7:32:46Realistically, though, this is probably cleaner
- 7:32:49than many data sets will actually be.
- 7:32:50Oftentimes, data is messier.
- 7:32:52There are outliers.
- 7:32:53There's random noise that happens inside of a particular system.
- 7:32:56And what we'd like to do is still be able to figure out
- 7:32:59what a line might look like.
- 7:33:00So in practice, the data will not always be linearly separable.
- 7:33:04Or linearly separable refers to some data set
- 7:33:07where I could draw a line just to separate the two halves of it perfectly.
- 7:33:11Instead, you might have a situation like this,
- 7:33:13where there are some rainy points that are on this side of the line
- 7:33:16and some not rainy points that are on that side of the line.
- 7:33:19And there may not be a line that perfectly separates
- 7:33:23what path of the inputs from the other half,
- 7:33:25that perfectly separates all the rainy days from the not rainy days.
- 7:33:29But we can still say that this line does a pretty good job.
- 7:33:33And we'll try to formalize a little bit later
- 7:33:34what we mean when we say something like this line does a pretty good job
- 7:33:38of trying to make that prediction.
- 7:33:40But for now, let's just say we're looking for a line that
- 7:33:42does as good of a job as we can at trying to separate one category of things
- 7:33:47from another category of things.
- 7:33:49So let's now try to formalize this a little bit more mathematically.
- 7:33:53We want to come up with some sort of function, some way we can define this
- 7:33:56line.
- 7:33:57And our inputs are things like humidity and pressure in this case.
- 7:34:01So our inputs we might call x1 is going to represent humidity,
- 7:34:05and x2 is going to represent pressure.
- 7:34:08These are inputs that we are going to provide to our machine learning
- 7:34:11algorithm.
- 7:34:12And given those inputs, we would like for our model
- 7:34:14to be able to predict some sort of output.
- 7:34:17And we are going to predict that using our hypothesis function, which
- 7:34:20we called h.
- 7:34:21Our hypothesis function is going to take as input x1 and x2, humidity
- 7:34:26and pressure in this case.
- 7:34:27And you can imagine if we didn't just have two inputs,
- 7:34:29we had three or four or five inputs or more,
- 7:34:31we could have this hypothesis function take all of those as input.
- 7:34:35And we'll see examples of that a little bit later as well.
- 7:34:38And now the question is, what does this hypothesis function do?
- 7:34:42Well, it really just needs to measure, is this data point
- 7:34:46on one side of the boundary, or is it on the other side of the boundary?
- 7:34:51And how do we formalize that boundary?
- 7:34:53Well, the boundary is generally going to be
- 7:34:55a linear combination of these input variables,
- 7:34:59at least in this particular case.
- 7:35:01So what we're trying to do when we say linear combination
- 7:35:03is take each of these inputs and multiply them
- 7:35:06by some number that we're going to have to figure out.
- 7:35:08We'll generally call that number a weight for how important
- 7:35:11should these variables be in trying to determine the answer.
- 7:35:14So we'll weight each of these variables with some weight,
- 7:35:17and we might add a constant to it just to try and make
- 7:35:19the function a little bit different.
- 7:35:21And the result, we just need to compare.
- 7:35:23Is it greater than 0, or is it less than 0 to say,
- 7:35:26does it belong on one side of the line or the other side of the line?
- 7:35:30So what that mathematical expression might look like is this.
- 7:35:33We would take each of my variables, x1 and x2, multiply them by some weight.
- 7:35:38I don't yet know what that weight is, but it's
- 7:35:40going to be some number, weight 1 and weight 2.
- 7:35:43And maybe we just want to add some other weight 0 to it,
- 7:35:46because the function might require us to shift the entire value up or down
- 7:35:50by a certain amount.
- 7:35:51And then we just compare.
- 7:35:52If we do all this math, is it greater than or equal to 0?
- 7:35:55If so, we might categorize that data point as a rainy day.
- 7:35:58And otherwise, we might say, no rain.
- 7:36:02So the key here, then, is that this expression
- 7:36:05is how we are going to calculate whether it's a rainy day or not.
- 7:36:08We're going to do a bunch of math where we take each of the variables,
- 7:36:11multiply them by a weight, maybe add an extra weight to it,
- 7:36:14see if the result is greater than or equal to 0.
- 7:36:17And using that result of that expression,
- 7:36:19we're able to determine whether it's raining or not raining.
- 7:36:22This expression here is in this case going to refer to just some line.
- 7:36:26If you were to plot that graphically, it would just be some line.
- 7:36:29And what the line actually looks like depends upon these weights.
- 7:36:33x1 and x2 are the inputs, but these weights
- 7:36:35are really what determine the shape of that line, the slope of that line,
- 7:36:39and what that line actually looks like.
- 7:36:42So we then would like to figure out what these weights should be.
- 7:36:45We can choose whatever weights we want, but we
- 7:36:47want to choose weights in such a way that if you pass in a rainy day's
- 7:36:51humidity and pressure, then you end up with a result that
- 7:36:53is greater than or equal to 0.
- 7:36:55And we would like it such that if we passed into our hypothesis
- 7:36:57function a not rainy day's inputs, then the output that we get
- 7:37:01should be not raining.
- 7:37:03So before we get there, let's try and formalize this a little bit more
- 7:37:06mathematically just to get a sense for how it is that you'll often see this
- 7:37:10if you ever go further into supervised machine learning
- 7:37:12and explore this idea.
- 7:37:14One thing is that generally for these categories,
- 7:37:16we'll sometimes just use the names of the categories like rain and not rain.
- 7:37:20Often mathematically, if we're trying to do comparisons between these things,
- 7:37:23it's easier just to deal in the world of numbers.
- 7:37:25So we could just say 1 and 0, 1 for raining, 0 for not raining.
- 7:37:30So we do all this math.
- 7:37:31And if the result is greater than or equal to 0,
- 7:37:34we'll go ahead and say our hypothesis function outputs 1, meaning raining.
- 7:37:37And otherwise, it outputs 0, meaning not raining.
- 7:37:41And oftentimes, this type of expression will instead
- 7:37:45express using vector mathematics.
- 7:37:47And all a vector is, if you're not familiar with the term,
- 7:37:50is it refers to a sequence of numerical values.
- 7:37:53You could represent that in Python using a list of numerical values
- 7:37:56or a tuple with numerical values.
- 7:37:59And here, we have a couple of sequences of numerical values.
- 7:38:02One of our vectors, one of our sequences of numerical values,
- 7:38:06are all of these individual weights, w0, w1, and w2.
- 7:38:11So we could construct what we'll call a weight vector,
- 7:38:14and we'll see why this is useful in a moment,
- 7:38:16called w, generally represented using a boldface w, that
- 7:38:19is just a sequence of these three weights, weight 0, weight 1,
- 7:38:23and weight 2.
- 7:38:24And to be able to calculate, based on those weights,
- 7:38:26whether we think a day is raining or not raining,
- 7:38:30we're going to multiply each of those weights by one of our input variables.
- 7:38:35That w2, this weight, is going to be multiplied by input variable x2.
- 7:38:39w1 is going to be multiplied by input variable x1.
- 7:38:42And w0, well, it's not being multiplied by anything.
- 7:38:46But to make sure the vectors are the same length,
- 7:38:48and we'll see why that's useful in just a second,
- 7:38:50we'll just go ahead and say w0 is being multiplied by 1.
- 7:38:54Because you can multiply by something by 1,
- 7:38:55and you end up getting the exact same number.
- 7:38:58So in addition to the weight vector w, we'll
- 7:39:00also have an input vector that we'll call x that has three values, 1,
- 7:39:05again, because we're just multiplying w0 by 1 eventually, and then x1 and x2.
- 7:39:11So here, then, we've represented two distinct vectors, a vector of weights
- 7:39:14that we need to somehow learn.
- 7:39:16The goal of our machine learning algorithm
- 7:39:18is to learn what this weight vector is supposed to be.
- 7:39:21We could choose any arbitrary set of numbers,
- 7:39:23and it would produce a function that tries to predict rain or not rain,
- 7:39:26but it probably wouldn't be very good.
- 7:39:28What we want to do is come up with a good choice of these weights
- 7:39:32so that we're able to do the accurate predictions.
- 7:39:34And then this input vector represents a particular input
- 7:39:38to the function, a data point for which we would like to estimate,
- 7:39:41is that day a rainy day, or is that day a not rainy day?
- 7:39:45And so that's going to vary just depending
- 7:39:47on what input is provided to our function, what
- 7:39:49it is that we are trying to estimate.
- 7:39:51And then to do the calculation, we want to calculate this expression here,
- 7:39:55and it turns out that expression is what we would call the dot product
- 7:39:59of these two vectors.
- 7:40:00The dot product of two vectors just means taking each of the terms
- 7:40:04in the vectors and multiplying them together, w0 multiply it by 1,
- 7:40:08w1 multiply it by x1, w2 multiply it by x2,
- 7:40:11and that's why these vectors need to be the same length.
- 7:40:14And then we just add all of the results together.
- 7:40:17So the dot product of w and x, our weight vector and our input vector,
- 7:40:22that's just going to be w0 times 1, or just w0,
- 7:40:26plus w1 times x1, multiplying these two terms together,
- 7:40:30plus w2 times x2, multiplying those terms together.
- 7:40:35So we have our weight vector, which we need to figure out.
- 7:40:38We need our machine learning algorithm to figure out
- 7:40:39what the weights should be.
- 7:40:41We have the input vector representing the data point
- 7:40:44that we're trying to predict a category for, predict a label for.
- 7:40:47And we're able to do that calculation by taking this dot product, which
- 7:40:51you'll often see represented in vector form.
- 7:40:53But if you haven't seen vectors before, you
- 7:40:54can think of it as identical to just this mathematical expression,
- 7:40:57just doing the multiplication, adding the results together,
- 7:41:01and then seeing whether the result is greater than or equal to 0 or not.
- 7:41:04This expression here is identical to the expression
- 7:41:07that we're calculating to see whether or not
- 7:41:09that answer is greater than or equal to 0 in this case.
- 7:41:14And so for that reason, you'll often see the hypothesis function
- 7:41:17written as something like this, a simpler representation where
- 7:41:20the hypothesis takes as input some input vector x, some humidity
- 7:41:25and pressure for some day.
- 7:41:26And we want to predict an output like rain or no rain or 1 or 0
- 7:41:30if we choose to represent things numerically.
- 7:41:33And the way we do that is by taking the dot product of the weights
- 7:41:37and our input.
- 7:41:38If it's greater than or equal to 0, we'll go ahead and say the output is 1.
- 7:41:42Otherwise, the output is going to be 0.
- 7:41:44And this hypothesis, we say, is parameterized by the weights.
- 7:41:49Depending on what weights we choose, we'll
- 7:41:51end up getting a different hypothesis.
- 7:41:53If we choose the weights randomly, we're probably
- 7:41:55not going to get a very good hypothesis function.
- 7:41:57We'll get a 1 or a 0.
- 7:41:58But it's probably not accurately going to reflect
- 7:42:01whether we think a day is going to be rainy or not rainy.
- 7:42:04But if we choose the weights right, we can often
- 7:42:06do a pretty good job of trying to estimate whether we think
- 7:42:09the output of the function should be a 1 or a 0.
- 7:42:13And so the question, then, is how to figure out
- 7:42:16what these weights should be, how to be able to tune those parameters.
- 7:42:19And there are a number of ways you can do that.
- 7:42:21One of the most common is known as the perceptron learning rule.
- 7:42:25And we'll see more of this later.
- 7:42:27But the idea of the perceptron learning rule,
- 7:42:29and we're not going to get too deep into the mathematics,
- 7:42:30we'll mostly just introduce it more conceptually,
- 7:42:33is to say that given some data point that we would like to learn from,
- 7:42:37some data point that has an input x and an output y, where
- 7:42:41y is like 1 for rain or 0 for not rain, then we're going to update the weights.
- 7:42:46And we'll look at the formula in just a moment.
- 7:42:48But the big picture idea is that we can start with random weights,
- 7:42:51but then learn from the data.
- 7:42:53Take the data points one at a time.
- 7:42:55And for each one of the data points, figure out, all right,
- 7:42:58what parameters do we need to change inside of the weights
- 7:43:02in order to better match that input point.
- 7:43:05And so that is the value of having access to a lot of data
- 7:43:07in the supervised machine learning algorithm,
- 7:43:09is that you take each of the data points and maybe look at them multiple times
- 7:43:13and constantly try and figure out whether you
- 7:43:15need to shift your weights in order to better create some weight vector that
- 7:43:19is able to correctly or more accurately try to estimate what the output should
- 7:43:24be, whether we think it's going to be raining
- 7:43:25or whether we think it's not going to be raining.
- 7:43:28So what does that weight update look like?
- 7:43:30Without going into too much of the mathematics,
- 7:43:32we're going to update each of the weights to be the result of the original
- 7:43:35weight plus some additional expression.
- 7:43:39And to understand this expression, y, well,
- 7:43:41y is what the actual output is.
- 7:43:44And hypothesis of x, the input, that's going to be what we thought the input
- 7:43:50was.
- 7:43:51And so I can replace this by saying what the actual value was minus what
- 7:43:55our estimate was.
- 7:43:56And based on the difference between the actual value and what our estimate was,
- 7:44:01we might want to change our hypothesis, change the way
- 7:44:04that we do that estimation.
- 7:44:06If the actual value and the estimate were the same thing,
- 7:44:08meaning we were correctly able to predict what category
- 7:44:11this data point belonged to, well, then actual value minus estimate,
- 7:44:14that's just going to be 0, which means this whole term on the right-hand side
- 7:44:18goes to be 0, and the weight doesn't change.
- 7:44:20Weight i, where i is like weight 1 or weight 2 or weight 0,
- 7:44:24weight i just stays at weight i.
- 7:44:26And none of the weights change if we were able to correctly predict
- 7:44:29what category the input belonged to.
- 7:44:32But if our hypothesis didn't correctly predict what category the input
- 7:44:36belonged to, well, then maybe then we need to make some changes, adjust
- 7:44:40the weights so that we're better able to predict this kind of data
- 7:44:43point in the future.
- 7:44:45And what is the way we might do that?
- 7:44:47Well, if the actual value was bigger than the estimate, then,
- 7:44:51and for now we'll go ahead and assume that these x's are positive values,
- 7:44:54then if the actual value was bigger than the estimate,
- 7:44:57well, that means we need to increase the weight in order
- 7:45:00to make it such that the output is bigger,
- 7:45:02and therefore we're more likely to get to the right actual value.
- 7:45:06And so if the actual value is bigger than the estimate,
- 7:45:08then actual value minus estimate, that'll be a positive number.
- 7:45:11And so you imagine we're just adding some positive number to the weight
- 7:45:14just to increase it ever so slightly.
- 7:45:16And likewise, the inverse case is true, that if the actual value
- 7:45:19was less than the estimate, the actual value was 0,
- 7:45:23but we estimated 1, meaning it actually was not raining,
- 7:45:26but we predicted it was going to be raining.
- 7:45:28Well, then we want to decrease the value of the weight,
- 7:45:31because then in that case, we want to try and lower
- 7:45:33the total value of computing that dot product in order
- 7:45:36to make it less likely that we would predict that it would actually
- 7:45:39be raining.
- 7:45:40So no need to get too deep into the mathematics of that,
- 7:45:43but the general idea is that every time we encounter some data point,
- 7:45:46we can adjust these weights accordingly to try and make
- 7:45:49the weights better line up with the actual data that we have access to.
- 7:45:53And you can repeat this process with data point after data point
- 7:45:56until eventually, hopefully, your algorithm
- 7:45:58converges to some set of weights that do a pretty good job of trying
- 7:46:02to figure out whether a day is going to be rainy or not raining.
- 7:46:05And just as a final point about this particular equation,
- 7:46:08this value alpha here is generally what we'll call the learning rate.
- 7:46:12It's just some parameter, some number we choose
- 7:46:15for how quickly we're actually going to be updating these weight values.
- 7:46:18So that if alpha is bigger, then we're going
- 7:46:20to update these weight values by a lot.
- 7:46:22And if alpha is smaller, then we'll update the weight values by less.
- 7:46:25And you can choose a value of alpha.
- 7:46:26Depending on the problem, different values
- 7:46:29might suit the situation better or worse than others.
- 7:46:32So after all of that, after we've done this training process of take
- 7:46:36all this data and using this learning rule,
- 7:46:38look at all the pieces of data and use each piece of data as an indication
- 7:46:43to us of do the weights stay the same, do we increase the weights,
- 7:46:45do we decrease the weights, and if so, by how much?
- 7:46:48What you end up with is effectively a threshold function.
- 7:46:52And we can look at what the threshold function looks like like this.
- 7:46:56On the x-axis here, we have the output of that function,
- 7:46:58taking the weights, taking the dot product of it with the input.
- 7:47:03And on the y-axis, we have what the output is going to be,
- 7:47:050, which in this case represented not raining,
- 7:47:08and 1, which in this case represented raining.
- 7:47:11And the way that our hypothesis function works is it calculates this value.
- 7:47:16And if it's greater than 0 or greater than some threshold value,
- 7:47:20then we declare that it's a rainy day.
- 7:47:22And otherwise, we declare that it's a not rainy day.
- 7:47:25And this then graphically is what that function looks like,
- 7:47:28that initially when the value of this dot product is small, it's not raining,
- 7:47:32it's not raining, it's not raining.
- 7:47:33But as soon as it crosses that threshold,
- 7:47:36we suddenly say, OK, now it's raining, now it's raining, now it's raining.
- 7:47:39And the way to interpret this kind of representation
- 7:47:42is that anything on this side of the line, that
- 7:47:44would be the category of data points where we say, yes, it's raining.
- 7:47:47Anything that falls on this side of the line
- 7:47:49are the data points where we would say, it's not raining.
- 7:47:52And again, we want to choose some value for the weights
- 7:47:54that results in a function that does a pretty good job of trying
- 7:47:57to do this estimation.
- 7:48:00But one tricky thing with this type of hard threshold
- 7:48:04is that it only leaves two possible outcomes.
- 7:48:07We plug in some data as input.
- 7:48:09And the output we get is raining or not raining.
- 7:48:13And there's no room for anywhere in between.
- 7:48:15And maybe that's what you want.
- 7:48:17Maybe all you want is given some data point,
- 7:48:19you would like to be able to classify it into one or two or more
- 7:48:22of these various different categories.
- 7:48:24But it might also be the case that you care about knowing
- 7:48:28how strong that prediction is, for example.
- 7:48:31So if we go back to this instance here, where we have rainy days
- 7:48:34on this side of the line, not rainy days on that side of the line,
- 7:48:38you might imagine that let's look now at these two white data points.
- 7:48:41This data point here that we would like to predict a label or a category for.
- 7:48:46And this data point over here that we would also
- 7:48:48like to predict a label or a category for.
- 7:48:51It seems likely that you could pretty confidently
- 7:48:53say that this data point, that should be a rainy day.
- 7:48:56Seems close to the other rainy days if we're
- 7:48:58going by the nearest neighbor strategy.
- 7:49:00It's on this side of the line if we're going by the strategy of just saying,
- 7:49:04which side of the line does it fall on by figuring out
- 7:49:07what those weights should be.
- 7:49:08And if we're using the line strategy of just which side of the line
- 7:49:11does it fall on, which side of this decision boundary,
- 7:49:14well, we'd also say that this point here is also a rainy day
- 7:49:18because it falls on the side of the line that corresponds to rainy days.
- 7:49:23But it's likely that even in this case, we
- 7:49:25would know that we don't feel nearly as confident about this data
- 7:49:29point on the left as compared to this data point on the right.
- 7:49:33That for this one on the right, we can feel very confident
- 7:49:35that yes, it's a rainy day.
- 7:49:37This one, it's pretty close to the line if we're judging just by distance.
- 7:49:41And so you might be less sure.
- 7:49:44But our threshold function doesn't allow for a notion of less sure
- 7:49:48or more sure about something.
- 7:49:50It's what we would call a hard threshold.
- 7:49:51It's once you've crossed this line, then immediately we say,
- 7:49:55yes, this is going to be a rainy day.
- 7:49:57Anywhere before it, we're going to say it's not a rainy day.
- 7:50:00And that may not be helpful in a number of cases.
- 7:50:03One, this is not a particularly easy function to deal with.
- 7:50:06As you get deeper into the world of machine learning
- 7:50:08and are trying to do things like taking derivatives of these curves
- 7:50:11with this type of function makes things challenging.
- 7:50:14But the other challenge is that we don't really
- 7:50:16have any notion of gradation between things.
- 7:50:17We don't have a notion of yes, this is a very strong belief
- 7:50:21that it's going to be raining as opposed to it's probably more likely than not
- 7:50:25that it's going to be raining, but maybe not totally sure about that either.
- 7:50:30So what we can do by taking advantage of a technique known
- 7:50:32as logistic regression is instead of using this hard threshold
- 7:50:36type of function, we can use instead a logistic function, something
- 7:50:39we might call a soft threshold.
- 7:50:41And that's going to transform this into looking something
- 7:50:45a little more like this, something that more nicely curves.
- 7:50:48And as a result, the possible output values are no longer just 0 and 1,
- 7:50:520 for not raining, 1 for raining.
- 7:50:55But you can actually get any real numbered value between 0 and 1.
- 7:50:59But if you're way over on this side, then you get a value of 0.
- 7:51:03OK, it's not going to be raining, and we're pretty sure about that.
- 7:51:05And if you're over on this side, you get a value of 1.
- 7:51:07And yes, we're very sure that it's going to be raining.
- 7:51:10But in between, you could get some real numbered value,
- 7:51:13where a value like 0.7 might mean we think it's going to rain.
- 7:51:17It's more probable that it's going to rain than not based on the data.
- 7:51:20But we're not as confident as some of the other data points might be.
- 7:51:25So one of the advantages of the soft threshold
- 7:51:27is that it allows us to have an output that could be some real number that
- 7:51:30potentially reflects some sort of probability, the likelihood that we
- 7:51:34think that this particular data point belongs to that particular category.
- 7:51:39And there are some other nice mathematical properties of that as well.
- 7:51:43So that then is two different approaches to trying
- 7:51:46to solve this type of classification problem.
- 7:51:48One is this nearest neighbor type of approach,
- 7:51:51where you just take a data point and look at the data points that are nearby
- 7:51:54to try and estimate what category we think it belongs to.
- 7:51:58And the other approach is the approach of saying, all right,
- 7:52:01let's just try and use linear regression,
- 7:52:03figure out what these weights should be, adjust the weights in order
- 7:52:06to figure out what line or what decision boundary is going
- 7:52:09to best separate these two categories.
- 7:52:12It turns out that another popular approach, a very popular approach
- 7:52:15if you just have a data set and you want to start
- 7:52:17trying to do some learning on it, is what we call the support vector machine.
- 7:52:20And we're not going to go too much into the mathematics of the support vector
- 7:52:23machine, but we'll at least explore it graphically to see what it is
- 7:52:26that it looks like.
- 7:52:27And the idea or the motivation behind the support vector machine
- 7:52:31is the idea that there are actually a lot of different lines
- 7:52:34that we could draw, a lot of different decision boundaries
- 7:52:37that we could draw to separate two groups.
- 7:52:39So for example, I had the red data points over here
- 7:52:41and the blue data points over here.
- 7:52:43One possible line I could draw is a line like this,
- 7:52:47that this line here would separate the red points from the blue points.
- 7:52:50And it does so perfectly.
- 7:52:51All the red points are on one side of the line.
- 7:52:54All the blue points are on the other side of the line.
- 7:52:56But this should probably make you a little bit nervous.
- 7:52:59If you come up with a model and the model comes up
- 7:53:02with a line that looks like this.
- 7:53:03And the reason why is that you worry about how well
- 7:53:06it's going to generalize to other data points that are not necessarily
- 7:53:10in the data set that we have access to.
- 7:53:12For example, if there was a point that fell like right here,
- 7:53:15for example, on the right side of the line, well, then based on that,
- 7:53:19we might want to guess that it is, in fact, a red point,
- 7:53:23but it falls on the side of the line where instead we
- 7:53:25would estimate that it's a blue point instead.
- 7:53:29And so based on that, this line is probably not a great choice
- 7:53:32just because it is so close to these various data points.
- 7:53:36We might instead prefer like a diagonal line
- 7:53:38that just goes diagonally through the data set like we've seen before.
- 7:53:41But there too, there's a lot of diagonal lines that we could draw as well.
- 7:53:44For example, I could draw this diagonal line here, which also successfully
- 7:53:48separates all the red points from all of the blue points.
- 7:53:51From the perspective of something like just trying
- 7:53:54to figure out some setting of weights that allows
- 7:53:56us to predict the correct output, this line
- 7:53:58will predict the correct output for this particular set of data
- 7:54:02every single time because the red points are on one side,
- 7:54:04the blue points are on the other.
- 7:54:06But yet again, you should probably be a little nervous
- 7:54:08because this line is so close to these red points,
- 7:54:11even though we're able to correctly predict on the input data,
- 7:54:15if there was a point that fell somewhere in this general area,
- 7:54:18our algorithm, this model, would say that, yeah, we think it's a blue point,
- 7:54:22when in actuality, it might belong to the red category instead
- 7:54:26just because it looks like it's close to the other red points.
- 7:54:29What we really want to be able to say, given this data, how can you generalize
- 7:54:33this as best as possible, is to come up with a line like this that
- 7:54:37seems like the intuitive line to draw.
- 7:54:39And the reason why it's intuitive is because it
- 7:54:41seems to be as far apart as possible from the red data and the blue data.
- 7:54:47So that if we generalize a little bit and assume
- 7:54:49that maybe we have some points that are different from the input
- 7:54:51but still slightly further away, we can still
- 7:54:54say that something on this side probably red, something on that side
- 7:54:58probably blue, and we can make those judgments that way.
- 7:55:01And that is what support vector machines are designed to do.
- 7:55:04They're designed to try and find what we call the maximum margin separator,
- 7:55:08where the maximum margin separator is just
- 7:55:10some boundary that maximizes the distance between the groups of points
- 7:55:14rather than come up with some boundary that's
- 7:55:16very close to one set or the other, where in the case
- 7:55:19before, we wouldn't have cared.
- 7:55:20As long as we're categorizing the input well, that seems all we need to do.
- 7:55:24The support vector machine will try and find this maximum margin separator,
- 7:55:28some way of trying to maximize that particular distance.
- 7:55:31And it does so by finding what we call the support vectors, which
- 7:55:35are the vectors that are closest to the line,
- 7:55:37and trying to maximize the distance between the line
- 7:55:40and those particular points.
- 7:55:42And it works that way in two dimensions.
- 7:55:44It also works in higher dimensions, where we're not
- 7:55:46looking for some line that separates the two data points,
- 7:55:49but instead looking for what we generally call a hyperplane,
- 7:55:52some decision boundary, effectively, that separates one set of data
- 7:55:57from the other set of data.
- 7:55:59And this ability of support vector machines
- 7:56:00to work in higher dimensions actually has a number of other applications
- 7:56:04as well.
- 7:56:04But one is that it helpfully deals with cases
- 7:56:07where data may not be linearly separable.
- 7:56:10So we talked about linear separability before,
- 7:56:12this idea that you can take data and just draw a line or some linear
- 7:56:16combination of the inputs that allows us to perfectly separate
- 7:56:20the two sets from each other.
- 7:56:21There are some data sets that are not linearly separable.
- 7:56:24And some were even two.
- 7:56:26You would not be able to find a good line at all
- 7:56:29that would try to do that kind of separation.
- 7:56:32Something like this, for example.
- 7:56:34Or if you imagine here are the red points and the blue points
- 7:56:37around it.
- 7:56:38If you try to find a line that divides the red points from the blue points,
- 7:56:43it's actually going to be difficult, if not impossible,
- 7:56:45to do that any line you choose, well, if you draw a line here,
- 7:56:49then you ignore all of these blue points that should actually
- 7:56:52be blue and not red.
- 7:56:53Anywhere else you draw a line, there's going to be a lot of error,
- 7:56:56a lot of mistakes, a lot of what we'll soon
- 7:56:58call loss to that line that you draw, a lot of points
- 7:57:02that you're going to categorize incorrectly.
- 7:57:04What we really want is to be able to find a better decision boundary that
- 7:57:08may not be just a straight line through this two dimensional space.
- 7:57:12And what support vector machines can do is
- 7:57:14they can begin to operate in higher dimensions
- 7:57:16and be able to find some other decision boundary,
- 7:57:19like the circle in this case, that actually
- 7:57:21is able to separate one of these sets of data
- 7:57:24from the other set of data a lot better.
- 7:57:26So oftentimes in data sets where the data is not linearly separable,
- 7:57:30support vector machines by working in higher dimensions
- 7:57:33can actually figure out a way to solve that kind of problem effectively.
- 7:57:37So that then, three different approaches to trying
- 7:57:39to solve these sorts of problems.
- 7:57:41We've seen support vector machines.
- 7:57:42We've seen trying to use linear regression and the perceptron learning
- 7:57:46rule to be able to figure out how to categorize inputs and outputs.
- 7:57:49We've seen the nearest neighbor approach.
- 7:57:51No one necessarily better than any other again.
- 7:57:54It's going to depend on the data set, the information you have access to.
- 7:57:57It's going to depend on what the function looks like that you're ultimately
- 7:58:00trying to predict.
- 7:58:01And this is where a lot of research and experimentation
- 7:58:04can be involved in trying to figure out how it
- 7:58:06is to best perform that kind of estimation.
- 7:58:09But classification is only one of the tasks
- 7:58:12that you might encounter in supervised machine learning.
- 7:58:14Because in classification, what we're trying to predict
- 7:58:17is some discrete category.
- 7:58:19We're trying to predict red or blue, rain or not rain,
- 7:58:22authentic or counterfeit.
- 7:58:24But sometimes what we want to predict is a real numbered value.
- 7:58:28And for that, we have a related problem, not classification,
- 7:58:31but instead known as regression.
- 7:58:33And regression is the supervised learning problem
- 7:58:35where we try and learn a function mapping inputs to outputs same as before.
- 7:58:39But instead of the outputs being discrete categories, things
- 7:58:43like rain or not rain, in a regression problem,
- 7:58:46the output values are generally continuous values, some real number
- 7:58:50that we would like to predict.
- 7:58:51This happens all the time as well.
- 7:58:53You might imagine that a company might take this approach
- 7:58:55if it's trying to figure out, for instance, what
- 7:58:58the effect of its advertising is.
- 7:58:59How do advertising dollars spent translate
- 7:59:02into sales for the company's product, for example?
- 7:59:05And so they might like to try to predict some function that
- 7:59:08takes as input the amount of money spent on advertising.
- 7:59:11And here, we're just going to use one input.
- 7:59:13But again, you could scale this up to many more inputs as well
- 7:59:15if you have a lot of different kinds of data you have access to.
- 7:59:18And the goal is to learn a function that given this amount of spending
- 7:59:21on advertising, we're going to get this amount in sales.
- 7:59:23And you might judge, based on having access to a whole bunch of data,
- 7:59:27like for every past month, here is how much we spent on advertising,
- 7:59:30and here is what sales were.
- 7:59:32And we would like to predict some sort of hypothesis function
- 7:59:36that, again, given the amount spent on advertising,
- 7:59:39we can predict, in this case, some real number, some number estimate
- 7:59:43of how much sales we expect that company to do in this month
- 7:59:47or in this quarter or whatever unit of time
- 7:59:49we're choosing to measure things in.
- 7:59:51And so again, the approach to solving this type of problem,
- 7:59:54we could try using a linear regression type approach where we take this data
- 7:59:58and we just plot it.
- 7:59:59On the x-axis, we have advertising dollars spent.
- 8:00:02On the y-axis, we have sales.
- 8:00:04And we might just want to try and draw a line that
- 8:00:07does a pretty good job of trying to estimate
- 8:00:09this relationship between advertising and sales.
- 8:00:12And in this case, unlike before, we're not
- 8:00:14trying to separate the data points into discrete categories.
- 8:00:17But instead, in this case, we're just trying
- 8:00:19to find a line that approximates this relationship between advertising
- 8:00:24and sales so that if we want to figure out what the estimated sales are
- 8:00:27for a particular advertising budget, you just look it up in this line,
- 8:00:31figure out for this amount of advertising,
- 8:00:33we would have this amount of sales and just try
- 8:00:35and make the estimate that way.
- 8:00:37And so you can try and come up with a line, again,
- 8:00:39figuring out how to modify the weights using various different techniques
- 8:00:42to try and make it so that this line fits as well as possible.
- 8:00:47So with all of these approaches, then, to trying to solve machine learning
- 8:00:51style problems, the question becomes, how do we evaluate these approaches?
- 8:00:54How do we evaluate the various different hypotheses
- 8:00:58that we could come up with?
- 8:00:59Because each of these algorithms will give us some sort of hypothesis,
- 8:01:02some function that maps inputs to outputs,
- 8:01:05and we want to know, how well does that function work?
- 8:01:09And you can think of evaluating these hypotheses
- 8:01:11and trying to get a better hypothesis as kind of like an optimization problem.
- 8:01:16In an optimization problem, as you recall from before,
- 8:01:19we were either trying to maximize some objective function
- 8:01:23by trying to find a global maximum, or we
- 8:01:26were trying to minimize some cost function by trying to find some global
- 8:01:30minimum.
- 8:01:31And in the case of evaluating these hypotheses, one thing we might say
- 8:01:34is that this cost function, the thing we're trying to minimize,
- 8:01:38we might be trying to minimize what we would call a loss function.
- 8:01:42And what a loss function is, is it is a function
- 8:01:44that is going to estimate for us how poorly our function performs.
- 8:01:49More formally, it's like a loss of utility
- 8:01:51by whenever we predict something that is wrong, that is a loss of utility.
- 8:01:55That's going to add to the output of our loss function.
- 8:01:59And you could come up with any loss function
- 8:02:01that you want, just some mathematical way of estimating,
- 8:02:03given each of these data points, given what the actual output is,
- 8:02:06and given what our projected output is, our estimate,
- 8:02:10you could calculate some sort of numerical loss for it.
- 8:02:12But there are a couple of popular loss functions
- 8:02:14that are worth discussing, just so that you've seen them before.
- 8:02:18When it comes to discrete categories, things like rain or not rain,
- 8:02:21counterfeit or not counterfeit, one approaches the 0, 1 loss function.
- 8:02:26And the way that works is for each of the data points,
- 8:02:29our loss function takes as input what the actual output is,
- 8:02:32like whether it was actually raining or not raining,
- 8:02:35and takes our prediction into account.
- 8:02:37Did we predict, given this data point, that it was raining or not raining?
- 8:02:41And if the actual value equals the prediction, well, then the 0, 1 loss
- 8:02:45function will just say the loss is 0.
- 8:02:47There was no loss of utility, because we were able to predict correctly.
- 8:02:51And otherwise, if the actual value was not the same thing
- 8:02:54as what we predicted, well, then in that case, our loss is 1.
- 8:02:58We lost something, lost some utility, because what we predicted
- 8:03:01was the output of the function, was not what it actually was.
- 8:03:05And the goal, then, in a situation like this
- 8:03:07would be to come up with some hypothesis that minimizes
- 8:03:11the total empirical loss, the total amount that we've lost,
- 8:03:14if you add up for all these data points what the actual output is
- 8:03:17and what your hypothesis would have predicted.
- 8:03:21So in this case, for example, if we go back to classifying days as raining
- 8:03:24or not raining, and we came up with this decision boundary,
- 8:03:27how would we evaluate this decision boundary?
- 8:03:29How much better is it than drawing the line here or drawing the line there?
- 8:03:33Well, we could take each of the input data points,
- 8:03:35and each input data point has a label, whether it was raining
- 8:03:38or whether it was not raining.
- 8:03:40And we could compare it to the prediction,
- 8:03:41whether we predicted it would be raining or not raining,
- 8:03:44and assign it a numerical value as a result.
- 8:03:47So for example, these points over here, they were all rainy days,
- 8:03:51and we predicted they would be raining, because they
- 8:03:53fall on the bottom side of the line.
- 8:03:55So they have a loss of 0, nothing lost from those situations.
- 8:03:58And likewise, same is true for some of these points over here,
- 8:04:01where it was not raining and we predicted it would not be raining either.
- 8:04:05Where we do have loss are points like this point here and that point there,
- 8:04:09where we predicted that it would not be raining,
- 8:04:13but in actuality, it's a blue point.
- 8:04:14It was raining.
- 8:04:15Or likewise here, we predicted that it would be raining,
- 8:04:18but in actuality, it's a red point.
- 8:04:20It was not raining.
- 8:04:21And so as a result, we miscategorized these data points
- 8:04:25that we were trying to train on.
- 8:04:27And as a result, there is some loss here.
- 8:04:29One loss here, there, here, and there, for a total loss of 4,
- 8:04:33for example, in this case.
- 8:04:34And that might be how we would estimate or how we would say
- 8:04:37that this line is better than a line that goes somewhere else
- 8:04:41or a line that's further down, because this line might minimize the loss.
- 8:04:45So there is no way to do better than just these four points of loss
- 8:04:50if you're just drawing a straight line through our space.
- 8:04:54So the 0, 1 loss function checks.
- 8:04:56Did we get it right?
- 8:04:57Did we get it wrong?
- 8:04:57If we got it right, the loss is 0, nothing lost.
- 8:05:00If we got it wrong, then our loss function for that data point says 1.
- 8:05:04And we add up all of those losses across all of our data points
- 8:05:07to get some sort of empirical loss, how much we
- 8:05:10have lost across all of these original data points
- 8:05:13that our algorithm had access to.
- 8:05:16There are other forms of loss as well that work especially well when
- 8:05:19we deal with more real valued cases, cases
- 8:05:21like the mapping between advertising budget and amount
- 8:05:24that we do in sales, for example.
- 8:05:26Because in that case, you care not just that you get the number exactly right,
- 8:05:30but you care how close you were to the actual value.
- 8:05:33If the actual value is you did like $2,800 in sales
- 8:05:37and you predicted that you would do $2,900 in sales,
- 8:05:40maybe that's pretty good.
- 8:05:42That's much better than if you had predicted you'd do $1,000 in sales,
- 8:05:45for example.
- 8:05:46And so we would like our loss function to be
- 8:05:48able to take that into account as well, take into account not just
- 8:05:53whether the actual value and the expected value are exactly the same,
- 8:05:57but also take into account how far apart they were.
- 8:06:01And so for that one approach is what we call L1 loss.
- 8:06:05L1 loss doesn't just look at whether actual and predicted
- 8:06:08are equal to each other, but we take the absolute value
- 8:06:11of the actual value minus the predicted value.
- 8:06:15In other words, we just ask how far apart were the actual and predicted
- 8:06:19values, and we sum that up across all of the data points
- 8:06:23to be able to get what our answer ultimately is.
- 8:06:26So what might this actually look like for our data set?
- 8:06:29Well, if we go back to this representation
- 8:06:31where we had advertising along the x-axis, sales along the y-axis,
- 8:06:35our line was our prediction, our estimate for any given
- 8:06:38amount of advertising, what we predicted sales was going to be.
- 8:06:42And our L1 loss is just how far apart vertically along the sales axis
- 8:06:48our prediction was from each of the data points.
- 8:06:51So we could figure out exactly how far apart
- 8:06:53our prediction was from each of the data points
- 8:06:55and figure out as a result of that what our loss is overall
- 8:06:59for this particular hypothesis just by adding up
- 8:07:02all of these various different individual losses for each of these data
- 8:07:05points.
- 8:07:06And our goal then is to try and minimize that loss,
- 8:07:08to try and come up with some line that minimizes what the utility loss is
- 8:07:13by judging how far away our estimate amount of sales
- 8:07:16is from the actual amount of sales.
- 8:07:18And turns out there are other loss functions as well.
- 8:07:21One that's quite popular is the L2 loss.
- 8:07:23The L2 loss, instead of just using the absolute value,
- 8:07:26like how far away the actual value is from the predicted value,
- 8:07:30it uses the square of actual minus predicted.
- 8:07:33So how far apart are the actual and predicted value?
- 8:07:36And it squares that value, effectively penalizing much more harshly anything
- 8:07:41that is a worse prediction.
- 8:07:43So you imagine if you have two data points
- 8:07:45that you predict as being one value away from their actual value,
- 8:07:50as opposed to one data point that you predict as being two away
- 8:07:53from its actual value, the L2 loss function
- 8:07:56will more harshly penalize that one that is two away,
- 8:08:00because it's going to square, however, much the differences
- 8:08:03between the actual value and the predicted value.
- 8:08:05And depending on the situation, you might
- 8:08:07want to choose a loss function depending on what you care about minimizing.
- 8:08:10If you really care about minimizing the error on more outlier cases,
- 8:08:14then you might want to consider something like this.
- 8:08:15But if you've got a lot of outliers, and you don't necessarily
- 8:08:18care about modeling them, then maybe an L1 loss function is preferable.
- 8:08:21But there are trade-offs here that you need to decide,
- 8:08:23based on a particular set of data.
- 8:08:26But what you do run the risk of with any of these loss functions,
- 8:08:29with anything that we're trying to do, is a problem known as overfitting.
- 8:08:33And overfitting is a big problem that you can encounter in machine learning,
- 8:08:36which happens anytime a model fits too closely with a data set,
- 8:08:41and as a result, fails to generalize.
- 8:08:44We would like our model to be able to accurately predict
- 8:08:48data and inputs and output pairs for the data that we have access to.
- 8:08:52But the reason we wanted to do so is because we
- 8:08:55want our model to generalize well to data that we haven't seen before.
- 8:08:59I would like to take data from the past year
- 8:09:01of whether it was raining or not raining,
- 8:09:03and use that data to generalize it towards the future.
- 8:09:06Say, in the future, is it going to be raining or not raining?
- 8:09:09Or if I have a whole bunch of data on what counterfeit and not counterfeit
- 8:09:12US dollar bills look like in the past when people have encountered them,
- 8:09:16I'd like to train a computer to be able to, in the future,
- 8:09:19generalize to other dollar bills that I might see as well.
- 8:09:24And the problem with overfitting is that if you try and tie yourself
- 8:09:28too closely to the data set that you're training your model on,
- 8:09:32you can end up not generalizing very well.
- 8:09:35So what does this look like?
- 8:09:36Well, we might imagine the rainy day and not rainy day
- 8:09:38example again from here, where the blue points indicate rainy days
- 8:09:41and the red points indicate not rainy days.
- 8:09:43And we decided that we felt pretty comfortable with drawing a line
- 8:09:47like this as the decision boundary between rainy days and not rainy days.
- 8:09:52So we can pretty comfortably say that points on this side
- 8:09:55more likely to be rainy days, points on that side more
- 8:09:57likely to be not rainy days.
- 8:09:59But the loss, the empirical loss, isn't zero in this particular case
- 8:10:04because we didn't categorize everything perfectly.
- 8:10:07There was this one outlier, this one day that it wasn't raining,
- 8:10:10but yet our model still predicts that it is raining.
- 8:10:13But that doesn't necessarily mean our model is bad.
- 8:10:15It just means the model isn't 100% accurate.
- 8:10:18If you really wanted to try and find a hypothesis that
- 8:10:21resulted in minimizing the loss, you could come up
- 8:10:25with a different decision boundary.
- 8:10:26It wouldn't be a line, but it would look something like this.
- 8:10:30This decision boundary does separate all of the red points
- 8:10:34from all of the blue points because the red points fall
- 8:10:37on this side of this decision boundary, the blue points
- 8:10:40fall on the other side of the decision boundary.
- 8:10:42But this, we would probably argue, is not as good of a prediction.
- 8:10:47Even though it seems to be more accurate based
- 8:10:50on all of the available training data that we
- 8:10:53have for training this machine learning model,
- 8:10:55we might say that it's probably not going to generalize well.
- 8:10:58That if there were other data points like here and there,
- 8:11:00we might still want to consider those to be rainy days
- 8:11:03because we think this was probably just an outlier.
- 8:11:06So if the only thing you care about is minimizing the loss on the data
- 8:11:10you have available to you, you run the risk of overfitting.
- 8:11:13And this can happen in the classification case.
- 8:11:15It can also happen in the regression case,
- 8:11:18that here we predicted what we thought was a pretty good line relating
- 8:11:21advertising to sales, trying to predict what sales were going
- 8:11:24to be for a given amount of advertising.
- 8:11:26But I could come up with a line that does a better job of predicting
- 8:11:29the training data, and it would be something that looks like this,
- 8:11:32just connecting all of the various different data points.
- 8:11:35And now there is no loss at all.
- 8:11:37Now I've perfectly predicted, given any advertising, what sales are.
- 8:11:41And for all the data available to me, it's going to be accurate.
- 8:11:45But it's probably not going to generalize very well.
- 8:11:47I have overfit my model on the training data that is available to me.
- 8:11:52And so in general, we want to avoid overfitting.
- 8:11:54We'd like strategies to make sure that we haven't overfit our model
- 8:11:58to a particular data set.
- 8:12:00And there are a number of ways that you could try to do this.
- 8:12:02One way is by examining what it is that we're optimizing for.
- 8:12:05In an optimization problem, all we do is we say, there is some cost,
- 8:12:10and I want to minimize that cost.
- 8:12:12And so far, we've defined that cost function, the cost of a hypothesis,
- 8:12:17just as being equal to the empirical loss of that hypothesis,
- 8:12:21like how far away are the actual data points, the outputs,
- 8:12:25away from what I predicted them to be based on that particular hypothesis.
- 8:12:29And if all you're trying to do is minimize cost, meaning minimizing
- 8:12:32the loss in this case, then the result is going to be that you might overfit,
- 8:12:36that to minimize cost, you're going to try and find a way to perfectly match
- 8:12:41all the input data.
- 8:12:42And that might happen as a result of overfitting
- 8:12:46on that particular input data.
- 8:12:48So in order to address this, you could add something to the cost function.
- 8:12:52What counts as cost will not just loss, but also
- 8:12:56some measure of the complexity of the hypothesis.
- 8:12:59The word the complexity of the hypothesis is something
- 8:13:02that you would need to define for how complicated does our line look.
- 8:13:06This is sort of an Occam's razor-style approach
- 8:13:08where we want to give preference to a simpler decision boundary,
- 8:13:12like a straight line, for example, some simpler curve, as opposed
- 8:13:15to something far more complex that might represent the training data better
- 8:13:19but might not generalize as well.
- 8:13:21We'll generally say that a simpler solution is probably the better solution
- 8:13:26and probably the one that is more likely to generalize well to other inputs.
- 8:13:31So we measure what the loss is, but we also measure the complexity.
- 8:13:34And now that all gets taken into account when we consider the overall cost,
- 8:13:38that yes, something might have less loss if it better predicts the training
- 8:13:42data, but if it's much more complex, it still
- 8:13:45might not be the best option that we have.
- 8:13:48And we need to come up with some balance between loss and complexity.
- 8:13:51And for that reason, you'll often see this represented
- 8:13:54as multiplying the complexity by some parameter that we have to choose,
- 8:13:58parameter lambda in this case, where we're saying if lambda is a greater
- 8:14:02value, then we really want to penalize more complex hypotheses.
- 8:14:06Whereas if lambda is smaller, we're going to penalize more complex hypotheses
- 8:14:10a little bit, and it's up to the machine learning programmer
- 8:14:14to decide where they want to set that value of lambda
- 8:14:17for how much do I want to penalize a more complex hypothesis that
- 8:14:21might fit the data a little better.
- 8:14:23And again, there's no one right answer to a lot of these things,
- 8:14:25but depending on the data set, depending on the data you have available to you
- 8:14:29and the problem you're trying to solve, your choice of these parameters
- 8:14:32may vary, and you may need to experiment a little bit
- 8:14:34to figure out what the right choice of that is ultimately going to be.
- 8:14:38This process, then, of considering not only loss,
- 8:14:41but also some measure of the complexity is known as regularization.
- 8:14:45Regularization is the process of penalizing a hypothesis that
- 8:14:49is more complex in order to favor a simpler hypothesis that is more
- 8:14:54likely to generalize well, more likely to be
- 8:14:56able to apply to other situations that are dealing with other input points
- 8:15:01unlike the ones that we've necessarily seen before.
- 8:15:04So oftentimes, you'll see us add some regularizing term
- 8:15:08to what we're trying to minimize in order to avoid this problem of overfitting.
- 8:15:14Now, another way of making sure we don't overfit
- 8:15:17is to run some experiments and to see whether or not
- 8:15:20we are able to generalize our model that we've created to other data sets
- 8:15:25as well.
- 8:15:26And it's for that reason that oftentimes when you're
- 8:15:28doing a machine learning experiment, when you've got some data
- 8:15:30and you want to try and come up with some function that predicts,
- 8:15:33given some input, what the output is going to be,
- 8:15:36you don't necessarily want to do your training on all of the data
- 8:15:39you have available to you that you could employ
- 8:15:42a method known as holdout cross-validation,
- 8:15:45where in holdout cross-validation, we split up our data.
- 8:15:48We split up our data into a training set and a testing set.
- 8:15:53The training set is the set of data that we're
- 8:15:55going to use to train our machine learning model.
- 8:15:57And the testing set is the set of data that we're
- 8:16:00going to use in order to test to see how well our machine learning
- 8:16:04model actually performed.
- 8:16:06So the learning happens on the training set.
- 8:16:08We figure out what the parameters should be.
- 8:16:10We figure out what the right model is.
- 8:16:12And then we see, all right, now that we've trained the model,
- 8:16:15we'll see how well it does at predicting things
- 8:16:17inside of the testing set, some set of data that we haven't seen before.
- 8:16:22And the hope then is that we're going to be
- 8:16:24able to predict the testing set pretty well
- 8:16:26if we're able to generalize based on the training
- 8:16:29data that's available to us.
- 8:16:31If we've overfit the training data, though,
- 8:16:32and we're not able to generalize, well, then when we look at the testing set,
- 8:16:36it's likely going to be the case that we're not
- 8:16:38going to predict things in the testing set nearly as effectively.
- 8:16:42So this is one method of cross-validation,
- 8:16:44validating to make sure that the work we have done
- 8:16:46is actually going to generalize to other data sets as well.
- 8:16:49And there are other statistical techniques we can use as well.
- 8:16:52One of the downsides of this just hold out cross-validation
- 8:16:55is if you say I just split it 50-50, I train using 50% of the data
- 8:17:00and test using the other 50%, or you could choose other percentages as well,
- 8:17:04is that there is a fair amount of data that I am now not using to train,
- 8:17:08that I might be able to get a better model as a result, for example.
- 8:17:12So one approach is known as k-fold cross-validation.
- 8:17:16In k-fold cross-validation, rather than just divide things into two sets
- 8:17:20and run one experiment, we divide things into k different sets.
- 8:17:24So maybe I divide things up into 10 different sets
- 8:17:27and then run 10 different experiments.
- 8:17:30So if I split up my data into 10 different sets of data,
- 8:17:33then what I'll do is each time for each of my 10 experiments,
- 8:17:37I will hold out one of those sets of data, where I'll say,
- 8:17:40let me train my model on these nine sets,
- 8:17:43and then test to see how well it predicts on set number 10.
- 8:17:47And then pick another set of nine sets to train on,
- 8:17:50and then test it on the other one that I held out,
- 8:17:52where each time I train the model on everything
- 8:17:55minus the one set that I'm holding out, and then
- 8:17:57test to see how well our model performs on the test that I did hold out.
- 8:18:02And what you end up getting is 10 different results,
- 8:18:0410 different answers for how accurately our model worked.
- 8:18:07And oftentimes, you could just take the average of those 10
- 8:18:09to get an approximation for how well we think our model performs overall.
- 8:18:14But the key idea is separating the training data from the testing data,
- 8:18:18because you want to test your model on data
- 8:18:20that is different from what you trained the model on.
- 8:18:23Because the training, you want to avoid overfitting.
- 8:18:25You want to be able to generalize.
- 8:18:26And the way you test whether you're able to generalize
- 8:18:29is by looking at some data that you haven't seen before
- 8:18:32and seeing how well we're actually able to perform.
- 8:18:36And so if we want to actually implement any of these techniques
- 8:18:38inside of a programming language like Python, number of ways we could do that.
- 8:18:42We could write this from scratch on our own,
- 8:18:45but there are libraries out there that allow
- 8:18:46us to take advantage of existing implementations of these algorithms,
- 8:18:50that we can use the same types of algorithms
- 8:18:53in a lot of different situations.
- 8:18:54And so there's a library, very popular one, known as Scikit-learn,
- 8:18:58which allows us in Python to be able to very quickly get
- 8:19:01set up with a lot of these different machine learning models.
- 8:19:03This library has already written an algorithm
- 8:19:06for nearest neighbor classification, for doing perceptron learning,
- 8:19:09for doing a bunch of other types of inference and supervised learning
- 8:19:12that we haven't yet talked about.
- 8:19:14But using it, we can begin to try actually testing how these methods work
- 8:19:19and how accurately they perform.
- 8:19:22So let's go ahead and take a look at one approach
- 8:19:24to trying to solve this type of problem.
- 8:19:26All right, so I'm first going to pull up banknotes.csv, which
- 8:19:30is a whole bunch of data provided by UC Irvine, which
- 8:19:33is information about various different banknotes
- 8:19:36that people took pictures of various different banknotes
- 8:19:38and measured various different properties of those banknotes.
- 8:19:41And in particular, some human categorized each of those banknotes
- 8:19:45as either a counterfeit banknote or as not counterfeit.
- 8:19:48And so what you're looking at here is each row represents one banknote.
- 8:19:52This is formatted as a CSV spreadsheet, where just comma separated values
- 8:19:55separating each of these various different fields.
- 8:19:58We have four different input values for each of these data points,
- 8:20:03just information, some measurement that was made on the banknote.
- 8:20:06And what those measurements exactly are aren't as important as the fact
- 8:20:09that we do have access to this data.
- 8:20:11But more importantly, we have access for each of these data points
- 8:20:14to a label, where 0 indicates something like this was not a counterfeit bill,
- 8:20:19meaning it was an authentic bill.
- 8:20:20And a data point labeled 1 means that it is a counterfeit bill,
- 8:20:25at least according to the human researcher who labeled this particular data.
- 8:20:29So we have a whole bunch of data representing
- 8:20:31a whole bunch of different data points, each of which
- 8:20:33has these various different measurements that
- 8:20:35were made on that particular bill, and each of which
- 8:20:38has an output value, 0 or 1, 0 meaning it was a genuine bill, 1 meaning
- 8:20:44it was a counterfeit bill.
- 8:20:46And what we would like to do is use supervised learning
- 8:20:48to begin to predict or model some sort of function that
- 8:20:51can take these four values as input and predict what the output would be.
- 8:20:55We want our learning algorithm to find some sort of pattern
- 8:20:58that is able to predict based on these measurements, something
- 8:21:01that you could measure just by taking a photo of a bill,
- 8:21:03predict whether that bill is authentic or whether that bill is counterfeit.
- 8:21:09And so how can we do that?
- 8:21:10Well, I'm first going to open up banknote0.py
- 8:21:13and see how it is that we do this.
- 8:21:15I'm first importing a lot of things from Scikit-learn,
- 8:21:18but importantly, I'm going to set my model equal to the perceptron model,
- 8:21:23which is one of those models that we talked about before.
- 8:21:25We're just going to try and figure out some setting of weights
- 8:21:28that is able to divide our data into two different groups.
- 8:21:31Then I'm going to go ahead and read data in for my file from banknotes.csv.
- 8:21:36And basically, for every row, I'm going to separate that row
- 8:21:39into the first four values of that row, which is the evidence for that row.
- 8:21:44And then the label, where if the final column in that row is a 0,
- 8:21:49the label is authentic.
- 8:21:51And otherwise, it's going to be counterfeit.
- 8:21:53So I'm effectively reading data in from the CSV file,
- 8:21:56dividing into a whole bunch of rows where each row has some evidence,
- 8:22:00those four input values that are going to be inputs to my hypothesis function.
- 8:22:04And then the label, the output, whether it is authentic or counterfeit,
- 8:22:07that is the thing that I am then trying to predict.
- 8:22:10So the next step is that I would like to split up my data set
- 8:22:12into a training set and a testing set, some set of data
- 8:22:15that I would like to train my machine learning model on,
- 8:22:18and some set of data that I would like to use to test that model,
- 8:22:21see how well it performed.
- 8:22:22So what I'll do is I'll go ahead and figure out length of the data,
- 8:22:25how many data points do I have.
- 8:22:27I'll go ahead and take half of them, save that number as a number called holdout.
- 8:22:30That is how many items I'm going to hold out for my data set
- 8:22:33to save for the testing phase.
- 8:22:35I'll randomly shuffle the data so it's in some random order.
- 8:22:38And then I'll say my testing set will be all of the data up to the holdout.
- 8:22:43So I'll take holdout many data items, and that will be my testing set.
- 8:22:47My training data will be everything else, the information
- 8:22:51that I'm going to train my model on.
- 8:22:53And then I'll say I need to divide my training data into two different sets.
- 8:22:58I need to divide it into my x values, where x here represents the inputs.
- 8:23:03So the x values, the x values that I'm going to train on,
- 8:23:06are basically for every row in my training set,
- 8:23:09I'm going to get the evidence for that row, those four values,
- 8:23:12where it's basically a vector of four numbers, where
- 8:23:14that is going to be all of the input.
- 8:23:16And then I need the y values.
- 8:23:18What are the outputs that I want to learn from,
- 8:23:20the labels that belong to each of these various different input points?
- 8:23:23Well, that's going to be the same thing for each row in the training data.
- 8:23:26But this time, I take that row and get what its label is,
- 8:23:29whether it is authentic or counterfeit.
- 8:23:31So I end up with one list of all of these vectors of my input data,
- 8:23:36and one list, which follows the same order,
- 8:23:38but is all of the labels that correspond with each of those vectors.
- 8:23:42And then to train my model, which in this case is just this perceptron model,
- 8:23:46I just call model.fit, pass in the training data,
- 8:23:49and what the labels for those training data are.
- 8:23:52And scikit-learn will take care of fitting the model,
- 8:23:54will do the entire algorithm for me.
- 8:23:57And then when it's done, I can then test to see how well that model performed.
- 8:24:01So I can say, let me get all of these input vectors
- 8:24:04for what I want to test on.
- 8:24:05So for each row in my testing data set, go ahead and get the evidence.
- 8:24:09And the y values, those are what the actual values were
- 8:24:13for each of the rows in the testing data set, what the actual label is.
- 8:24:17But then I'm going to generate some predictions.
- 8:24:19I'm going to use this model and try and predict,
- 8:24:22based on the testing vectors, I want to predict what the output is.
- 8:24:26And my goal then is to now compare y testing with predictions.
- 8:24:31I want to see how well my predictions, based on the model,
- 8:24:34actually reflect what the y values were, what the output is,
- 8:24:38that were actually labeled.
- 8:24:39Because I now have this label data, I can assess how well the algorithm worked.
- 8:24:44And so now I can just compute how well we did.
- 8:24:47I'm going to, this zip function basically just lets
- 8:24:49me look through two different lists, one by one at the same time.
- 8:24:53So for each actual value and for each predicted value,
- 8:24:57if the actual is the same thing as what I predicted,
- 8:24:59I'll go ahead and increment the counter by one.
- 8:25:01Otherwise, I'll increment my incorrect counter by one.
- 8:25:04And so at the end, I can print out, here are the results,
- 8:25:06here's how many I got right, here's how many I got wrong,
- 8:25:09and here was my overall accuracy, for example.
- 8:25:12So I can go ahead and run this.
- 8:25:14I can run python banknote0.py.
- 8:25:17And it's going to train on half the data set
- 8:25:20and then test on half the data set.
- 8:25:21And here are the results for my perceptron model.
- 8:25:24In this case, it correctly was able to classify 679 bills as correctly
- 8:25:29either authentic or counterfeit and incorrectly classified seven of them
- 8:25:33for an overall accuracy of close to 99% accurate.
- 8:25:37So on this particular data set, using this perceptron model,
- 8:25:40we were able to predict very well what the output was going to be.
- 8:25:44And we can try different models, too, that scikit-learn
- 8:25:46makes it very easy just to swap out one model for another model.
- 8:25:50So instead of the perceptron model, I can use the support vector machine
- 8:25:55using the SVC, otherwise known as a support vector classifier,
- 8:25:59using a support vector machine to classify things
- 8:26:01into two different groups.
- 8:26:03And now see, all right, how well does this perform?
- 8:26:07And all right, this time, we were able to correctly predict 682
- 8:26:10and incorrectly predicted four for accuracy of 99.4%.
- 8:26:15And we could even try the k-neighbors classifier as the model instead.
- 8:26:20And this takes a parameter, n neighbors, for how many neighbors
- 8:26:24do you want to look at?
- 8:26:25Let's just look at one neighbor, the one nearest neighbor,
- 8:26:27and use that to predict.
- 8:26:29Go ahead and run this as well.
- 8:26:31And it looks like, based on the k-neighbors classifier,
- 8:26:33looking at just one neighbor, we were able to correctly classify
- 8:26:36685 data points, incorrectly classified one.
- 8:26:40Maybe let's try three neighbors instead, instead of just using one neighbor.
- 8:26:43Do more of a k-nearest neighbors approach,
- 8:26:45where I look at the three nearest neighbors and see how that performs.
- 8:26:48And that one, in this case, seems to have gotten 100% of all of the predictions
- 8:26:54correctly described as either authentic banknotes
- 8:26:58or as counterfeit banknotes.
- 8:27:00And we could run these experiments multiple times,
- 8:27:02because I'm randomly reorganizing the data every time.
- 8:27:05We're technically training these on slightly different data sets.
- 8:27:07And so you might want to run multiple experiments to really see
- 8:27:10how well they're actually going to perform.
- 8:27:12But in short, they all perform very well.
- 8:27:14And while some of them perform slightly better than others here,
- 8:27:16that might not always be the case for every data set.
- 8:27:19But you can begin to test now by very quickly putting together
- 8:27:22these machine learning models using Scikit-learn
- 8:27:24to be able to train on some training set and then
- 8:27:27test on some testing set as well.
- 8:27:29And this splitting up into training groups and testing groups and testing
- 8:27:33happens so often that Scikit-learn has functions built in for trying to do it.
- 8:27:37I did it all by hand just now.
- 8:27:39But if we take a look at banknotes one, we
- 8:27:41take advantage of some other features that exist in Scikit-learn,
- 8:27:45where we can really simplify a lot of our logic,
- 8:27:48that there is a function built into Scikit-learn called train test split,
- 8:27:52which will automatically split data into a training group and a testing group.
- 8:27:56I just have to say what proportion should be in the testing group, something
- 8:27:59like 0.5, half the data inside the testing group.
- 8:28:02Then I can fit the model on the training data,
- 8:28:05make the predictions on the testing data, and then just count up.
- 8:28:08And Scikit-learn has some nice methods for just counting up
- 8:28:11how many times our testing data match the predictions,
- 8:28:15how many times our testing data didn't match the predictions.
- 8:28:18So very quickly, you can write programs with not all that many lines of code.
- 8:28:21It's maybe like 40 lines of code to get through all of these predictions.
- 8:28:25And then as a result, see how well we're able to do.
- 8:28:28So these types of libraries can allow us, without really knowing
- 8:28:31the implementation details of these algorithms,
- 8:28:33to be able to use the algorithms in a very practical way
- 8:28:36to be able to solve these types of problems.
- 8:28:40So that then was supervised learning, this task
- 8:28:42of given a whole set of data, some input output pairs,
- 8:28:45we would like to learn some function that maps those inputs to those outputs.
- 8:28:50But turns out there are other forms of learning as well.
- 8:28:52And another popular type of machine learning, especially nowadays,
- 8:28:55is known as reinforcement learning.
- 8:28:58And the idea of reinforcement learning is rather than just
- 8:29:00being given a whole data set at the beginning of input output pairs,
- 8:29:04reinforcement learning is all about learning from experience.
- 8:29:07In reinforcement learning, our agent, whether it's
- 8:29:10like a physical robot that's trying to make actions in the world
- 8:29:13or just some virtual agent that is a program running somewhere,
- 8:29:16our agent is going to be given a set of rewards or punishments
- 8:29:20in the form of numerical values.
- 8:29:22But you can think of them as reward or punishment.
- 8:29:24And based on that, it learns what actions to take in the future,
- 8:29:28that our agent, our AI, will be put in some sort of environment.
- 8:29:32It will make some actions.
- 8:29:33And based on the actions that it makes, it learns something.
- 8:29:36It either gets a reward when it does something well,
- 8:29:38it gets a punishment when it does something poorly,
- 8:29:40and it learns what to do or what not to do in the future
- 8:29:44based on those individual experiences.
- 8:29:47And so what this will often look like is it will often
- 8:29:50start with some agent, some AI, which might, again, be a physical robot,
- 8:29:54if you're imagining a physical robot moving around,
- 8:29:56but it can also just be a program.
- 8:29:58And our agent is situated in their environment,
- 8:30:01where the environment is where they're going to make their actions,
- 8:30:04and it's what's going to give them rewards or punishments
- 8:30:06for various actions that they're in.
- 8:30:09So for example, the environment is going to start off
- 8:30:12by putting our agent inside of a state.
- 8:30:14Our agent has some state that, in a game,
- 8:30:17might be the state of the game that the agent is playing.
- 8:30:19In a world that the agent is exploring might
- 8:30:21be some position inside of a grid representing the world
- 8:30:24that they're exploring.
- 8:30:25But the agent is in some sort of state.
- 8:30:28And in that state, the agent needs to choose to take an action.
- 8:30:32The agent likely has multiple actions they can choose from,
- 8:30:34but they pick an action.
- 8:30:36So they take an action in a particular state.
- 8:30:39And as a result of that, the agent will generally
- 8:30:42get two things in response as we model them.
- 8:30:44The agent gets a new state that they find themselves in.
- 8:30:47After being in this state, taking one action,
- 8:30:50they end up in some other state.
- 8:30:52And they're also given some sort of numerical reward,
- 8:30:55positive meaning reward, meaning it was a good thing,
- 8:30:58negative generally meaning they did something bad,
- 8:31:00they received some sort of punishment.
- 8:31:03And that is all the information the agent has.
- 8:31:06It's told what state it's in.
- 8:31:08It makes some sort of action.
- 8:31:10And based on that, it ends up in another state.
- 8:31:12And it ends up getting some particular reward.
- 8:31:14And it needs to learn, based on that information, what actions
- 8:31:17to begin to take in the future.
- 8:31:19And so you could imagine generalizing this to a lot
- 8:31:21of different situations.
- 8:31:22This is oftentimes how you train if you've ever seen those robots that
- 8:31:26are now able to walk around the way humans do.
- 8:31:29It would be quite difficult to program the robot in exactly the right way
- 8:31:32to get it to walk the way humans do.
- 8:31:34You could instead train it through reinforcement learning,
- 8:31:36give it some sort of numerical reward every time it does something good,
- 8:31:40like take steps forward, and punish it every time it does something
- 8:31:43bad, like fall over, and then let the AI just
- 8:31:46learn based on that sequence of rewards, based
- 8:31:48on trying to take various different actions.
- 8:31:51You can begin to have the agent learn what to do in the future
- 8:31:54and what not to do.
- 8:31:56So in order to begin to formalize this, the first thing we need to do
- 8:31:59is formalize this notion of what we mean about states and actions and rewards,
- 8:32:03like what does this world look like?
- 8:32:05And oftentimes, we'll formulate this world
- 8:32:07as what's known as a Markov decision process, similar in spirit
- 8:32:11to Markov chains, which you might recall from before.
- 8:32:14But a Markov decision process is a model that we
- 8:32:16can use for decision making, for an agent trying
- 8:32:19to make decisions in its environment.
- 8:32:21And it's a model that allows us to represent the various different states
- 8:32:25that an agent can be in, the various different actions that they can take,
- 8:32:28and also what the reward is for taking one action as opposed to another action.
- 8:32:35So what then does it actually look like?
- 8:32:37Well, if you recall a Markov chain from before,
- 8:32:40a Markov chain looked a little something like this,
- 8:32:43where we had a whole bunch of these individual states,
- 8:32:45and each state immediately transitioned to another state
- 8:32:48based on some probability distribution.
- 8:32:50We saw this in the context of the weather before, where if it was sunny,
- 8:32:54we said with some probability, it'll be sunny the next day.
- 8:32:56With some other probability, it'll be rainy, for example.
- 8:32:59But we could also imagine generalizing this.
- 8:33:02It's not just sun and rain anymore.
- 8:33:04We just have these states, where one state leads to another state
- 8:33:07according to some probability distribution.
- 8:33:09But in this original model, there was no agent
- 8:33:12that had any control over this process.
- 8:33:14It was just entirely probability based, where with some probability,
- 8:33:17we moved to this next state.
- 8:33:18But maybe it's going to be some other state with some other probability.
- 8:33:22What we'll now have is the ability for the agent in this state
- 8:33:26to choose from a set of actions, where maybe instead of just one path
- 8:33:29forward, they have three different choices of actions that each lead up
- 8:33:33down different paths.
- 8:33:34And even this is a bit of an oversimplification,
- 8:33:36because in each of these states, you might imagine more branching points
- 8:33:39where there are more decisions that can be taken as well.
- 8:33:42So we've extended the Markov chain to say that from a state,
- 8:33:46you now have available action choices.
- 8:33:48And each of those actions might be associated
- 8:33:50with its own probability distribution of going to various different states.
- 8:33:55Then in addition, we'll add another extension,
- 8:33:58where any time you move from a state, taking an action,
- 8:34:01going into this other state, we can associate a reward with that outcome,
- 8:34:07saying either r is positive, meaning some positive reward,
- 8:34:10or r is negative, meaning there was some sort of punishment.
- 8:34:13And this then is what we'll consider to be a Markov decision process.
- 8:34:16That a Markov decision process has some initial set
- 8:34:18of states, of states in the world that we can be in.
- 8:34:21We have some set of actions that, given a state,
- 8:34:24I can say, what are the actions that are available to me in that state,
- 8:34:28an action that I can choose from?
- 8:34:30Then we have some transition model.
- 8:34:32The transition model before just said that, given my current state,
- 8:34:36what is the probability that I end up in that next state or this other state?
- 8:34:39The transition model now has effectively two things we're conditioning on.
- 8:34:44We're saying, given that I'm in this state and that I take this action,
- 8:34:48what's the probability that I end up in this next state?
- 8:34:52Now maybe we live in a very deterministic world in this Markov decision process.
- 8:34:56We're given a state and given an action.
- 8:34:58We know for sure what next state we'll end up in.
- 8:35:00But maybe there's some randomness in the world
- 8:35:02that when you take in a state and you take an action,
- 8:35:04you might not always end up in the exact same state.
- 8:35:07There might be some probabilities involved there as well.
- 8:35:09The Markov decision process can handle both of those possible cases.
- 8:35:14And then finally, we have a reward function, generally called r,
- 8:35:18that in this case says, what is the reward for being in this state,
- 8:35:21taking this action, and then getting to s prime this next state?
- 8:35:26So I'm in this original state.
- 8:35:27I take this action.
- 8:35:28I get to this next state.
- 8:35:29What is the reward for doing that process?
- 8:35:32And you can add up these rewards every time you take an action
- 8:35:35to get the total amount of rewards that an agent might
- 8:35:38get from interacting in a particular environment
- 8:35:41modeled using this Markov decision process.
- 8:35:44So what might this actually look like in practice?
- 8:35:46Well, let's just create a little simulated world here
- 8:35:49where I have this agent that is just trying to navigate its way.
- 8:35:52This agent is this yellow dot here, like a robot in the world,
- 8:35:55trying to navigate its way through this grid.
- 8:35:57And ultimately, it's trying to find its way to the goal.
- 8:36:00And if it gets to the green goal, then it's going to get some sort of reward.
- 8:36:04But then we might also have some red squares that are places
- 8:36:08where you get some sort of punishment, some bad place where we don't want
- 8:36:11the agent to go.
- 8:36:12And if it ends up in the red square, then our agent
- 8:36:14is going to get some sort of punishment as a result of that.
- 8:36:18But the agent originally doesn't know all of these details.
- 8:36:21It doesn't know that these states are associated with punishments.
- 8:36:24But maybe it does know that this state is associated with a reward.
- 8:36:27Maybe it doesn't.
- 8:36:28But it just needs to sort of interact with the environment
- 8:36:30to try and figure out what to do and what not to do.
- 8:36:33So the first thing the agent might do is,
- 8:36:35given no additional information, if it doesn't know what the punishments are,
- 8:36:39it doesn't know where the rewards are, it just might try and take an action.
- 8:36:43And it takes an action and ends up realizing
- 8:36:45that it got some sort of punishment.
- 8:36:47And so what does it learn from that experience?
- 8:36:49Well, it might learn that when you're in this state in the future,
- 8:36:53don't take the action move to the right, that that is a bad action to take.
- 8:36:57That in the future, if you ever find yourself back in the state,
- 8:36:59don't take this action of going to the right
- 8:37:02when you're in this particular state, because that leads to punishment.
- 8:37:05That might be the intuition at least.
- 8:37:06And so you could try doing other actions.
- 8:37:08You move up, all right, that didn't lead to any immediate rewards.
- 8:37:11Maybe try something else.
- 8:37:12Then maybe try something else.
- 8:37:14And all right, now you found that you got another punishment.
- 8:37:17And so you learn something from that experience.
- 8:37:18So the next time you do this whole process,
- 8:37:20you know that if you ever end up in this square,
- 8:37:22you shouldn't take the down action, because being in this state
- 8:37:26and taking that action ultimately leads to some sort of punishment,
- 8:37:30a negative reward, in other words.
- 8:37:33And this process repeats.
- 8:37:34You might imagine just letting our agent explore the world,
- 8:37:37learning over time what states tend to correspond with poor actions,
- 8:37:41learning over time what states correspond with poor actions,
- 8:37:43until eventually, if it tries enough things randomly,
- 8:37:47it might find that eventually when you get to this state,
- 8:37:50if you take the up action in this state, it
- 8:37:53might find that you actually get a reward from that.
- 8:37:56And what it can learn from that is that if you're in this state,
- 8:37:59you should take the up action, because that leads to a reward.
- 8:38:02And over time, you can also learn that if you're in this state,
- 8:38:05you should take the left action, because that leads to this state that also
- 8:38:08lets you eventually get to the reward.
- 8:38:10So you begin to learn over time not only which actions
- 8:38:14are good in particular states, but also which actions are bad,
- 8:38:18such that once you know some sequence of good actions that
- 8:38:20leads you to some sort of reward, our agent can just follow those
- 8:38:24instructions, follow the experience that it has learned.
- 8:38:27We didn't tell the agent what the goal was.
- 8:38:30We didn't tell the agent where the punishments were.
- 8:38:32But the agent can begin to learn from this experience
- 8:38:35and learn to begin to perform these sorts of tasks better in the future.
- 8:38:40And so let's now try to formalize this idea, formalize the idea
- 8:38:43that we would like to be able to learn in this state taking this action,
- 8:38:47is that a good thing or a bad thing?
- 8:38:49There are lots of different models for reinforcement learning.
- 8:38:51We're just going to look at one of them today.
- 8:38:53And the one that we're going to look at is a method known as Q-learning.
- 8:38:57And what Q-learning is all about is about learning
- 8:38:59a function, a function Q, that takes inputs S and A, where S is a state
- 8:39:05and A is an action that you take in that state.
- 8:39:07And what this Q function is going to do is it is going to estimate the value.
- 8:39:12How much reward will I get from taking this action in this state?
- 8:39:18Originally, we don't know what this Q function should be.
- 8:39:21But over time, based on experience, based on trying things out
- 8:39:24and seeing what the result is, I would like to try and learn
- 8:39:28what Q of SA is for any particular state and any particular action
- 8:39:32that I might take in that state.
- 8:39:34So what is the approach?
- 8:39:35Well, the approach originally is we'll start with Q SA equal to 0 for all
- 8:39:40states S and for all actions A. That initially,
- 8:39:43before I've ever started anything, before I've had any experiences,
- 8:39:47I don't know the value of taking any action in any given state.
- 8:39:50So I'm going to assume that the value is just 0 all across the board.
- 8:39:55But then as I interact with the world, as I experience rewards or punishments,
- 8:39:59or maybe I go to a cell where I don't get either reward or a punishment,
- 8:40:03I want to somehow update my estimate of Q SA.
- 8:40:07I want to continually update my estimate of Q SA
- 8:40:10based on the experiences and rewards and punishments that I've received,
- 8:40:13such that in the future, my knowledge of what actions are good
- 8:40:17and what states will be better.
- 8:40:19So when we take an action and receive some sort of reward,
- 8:40:22I want to estimate the new value of Q SA.
- 8:40:25And I estimate that based on a couple of different things.
- 8:40:28I estimate it based on the reward that I'm getting from taking this action
- 8:40:32and getting into the next state.
- 8:40:33But assuming the situation isn't over, assuming there are still
- 8:40:37future actions that I might take as well,
- 8:40:40I also need to take into account the expected future rewards.
- 8:40:44That if you imagine an agent interacting with the environment,
- 8:40:47then sometimes you'll take an action and get a reward,
- 8:40:49but then you can keep taking more actions and get more rewards,
- 8:40:52that these both are relevant, both the current reward
- 8:40:55I'm getting from this current step and also my future reward.
- 8:40:58And it might be the case that I'll want to take a step that
- 8:41:01doesn't immediately lead to a reward, because later on down the line,
- 8:41:05I know it will lead to more rewards as well.
- 8:41:07So there's a balancing act between current rewards
- 8:41:10that the agent experiences and future rewards
- 8:41:13that the agent experiences as well.
- 8:41:16And then we need to update QSA.
- 8:41:19So we estimate the value of QSA based on the current reward
- 8:41:22and the expected future rewards.
- 8:41:24And then we need to update this Q function
- 8:41:26to take into account this new estimate.
- 8:41:29Now, we already, as we go through this process,
- 8:41:31we'll already have an estimate for what we think the value is.
- 8:41:35Now we have a new estimate, and then somehow we
- 8:41:37need to combine these two estimates together,
- 8:41:39and we'll look at more formal ways that we can actually begin to do that.
- 8:41:43So to actually show you what this formula looks like,
- 8:41:45here is the approach we'll take with Q learning.
- 8:41:47We're going to, again, start with Q of S and A being equal to 0 for all states.
- 8:41:52And then every time we take an action A in state S and observer reward R,
- 8:41:59we're going to update our value, our estimate, for Q of SA.
- 8:42:04And the idea is that we're going to figure out
- 8:42:06what the new value estimate is minus what our existing value estimate is.
- 8:42:12And so we have some preconceived notion for what the value is
- 8:42:15for taking this action in this state.
- 8:42:17Maybe our expectation is we currently think the value is 10.
- 8:42:21But then we're going to estimate what we now think it's going to be.
- 8:42:24Maybe the new value estimate is something like 20.
- 8:42:27So there's a delta of 10 that our new value estimate
- 8:42:30is 10 points higher than what our current value estimate happens to be.
- 8:42:35And so we have a couple of options here.
- 8:42:37We need to decide how much we want to adjust
- 8:42:40our current expectation of what the value is
- 8:42:42of taking this action in this particular state.
- 8:42:45And what that difference is, how much we add or subtract
- 8:42:49from our existing notion of how much do we expect the value to be,
- 8:42:52is dependent on this parameter alpha, also called a learning rate.
- 8:42:56And alpha represents, in effect, how much we value new information
- 8:43:01compared to how much we value old information.
- 8:43:04An alpha value of 1 means we really value new information.
- 8:43:08But if we have a new estimate, then it doesn't
- 8:43:10matter what our old estimate is.
- 8:43:12We're only going to consider our new estimate
- 8:43:14because we always just want to take into consideration our new information.
- 8:43:18So the way that works is that if you imagine alpha being 1,
- 8:43:21well, then we're taking the old value of QSA
- 8:43:25and then adding 1 times the new value minus the old value.
- 8:43:29And that just leaves us with the new value.
- 8:43:31So when alpha is 1, all we take into consideration
- 8:43:34is what our new estimate happens to be.
- 8:43:37But over time, as we go through a lot of experiences,
- 8:43:40we already have some existing information.
- 8:43:42We might have tried taking this action nine times already.
- 8:43:46And now we just tried it a 10th time.
- 8:43:48And we don't only want to consider this 10th experience.
- 8:43:51I also want to consider the fact that my prior nine experiences, those
- 8:43:54were meaningful, too.
- 8:43:55And that's data I don't necessarily want to lose.
- 8:43:58And so this alpha controls that decision,
- 8:44:01controls how important is the new information.
- 8:44:030 would mean ignore all the new information.
- 8:44:06Just keep this Q value the same.
- 8:44:091 means replace the old information entirely with the new information.
- 8:44:13And somewhere in between, keep some sort of balance between these two values.
- 8:44:17We can put this equation a little bit more formally as well.
- 8:44:21The old value estimate is our old estimate
- 8:44:23for what the value is of taking this action in a particular state.
- 8:44:27That's just Q of SNA.
- 8:44:30So we have it once here, and we're going to add something to it.
- 8:44:33We're going to add alpha times the new value estimate
- 8:44:35minus the old value estimate.
- 8:44:37But the old value estimate, we just look up by calling this Q function.
- 8:44:42And what then is the new value estimate?
- 8:44:44Based on this experience we have just taken,
- 8:44:46what is our new estimate for the value of taking
- 8:44:48this action in this particular state?
- 8:44:51Well, it's going to be composed of two parts.
- 8:44:54It's going to be composed of what reward did I just
- 8:44:56get from taking this action in this state.
- 8:45:00And then it's going to be, what can I expect my future rewards
- 8:45:03to be from this point forward?
- 8:45:05So it's going to be R, some reward I'm getting right now,
- 8:45:10plus whatever I estimate I'm going to get in the future.
- 8:45:14And how do I estimate what I'm going to get in the future?
- 8:45:16Well, it's a bit of another call to this Q function.
- 8:45:19It's going to be take the maximum across all possible actions
- 8:45:23I could take next and say, all right, of all of these possible actions
- 8:45:27I could take, which one is going to have the highest reward?
- 8:45:31And so this then looks a little bit complicated.
- 8:45:33This is going to be our notion for how we're
- 8:45:35going to perform this kind of update.
- 8:45:37I have some estimate, some old estimate, for what the value is
- 8:45:41of taking this action in this state.
- 8:45:44And I'm going to update it based on new information
- 8:45:46that I experience some reward.
- 8:45:48I predict what my future reward is going to be.
- 8:45:51And using that I update what I estimate the reward will
- 8:45:54be for taking this action in this particular state.
- 8:45:57And there are other additions you might make to this algorithm as well.
- 8:46:00Sometimes it might not be the case that future rewards
- 8:46:03you want to wait equally to current rewards.
- 8:46:05Maybe you want an agent that values reward now over reward later.
- 8:46:10And so sometimes you can even add another term in here, some other parameter,
- 8:46:13where you discount future rewards and say future rewards are not
- 8:46:17as valuable as rewards immediately.
- 8:46:19That getting reward in the current time step
- 8:46:21is better than waiting a year and getting rewards later.
- 8:46:24But that's something up to the programmer
- 8:46:26to decide what that parameter ought to be.
- 8:46:29But the big picture idea of this entire formula
- 8:46:32is to say that every time we experience some new reward,
- 8:46:35we take that into account.
- 8:46:36We update our estimate of how good is this action.
- 8:46:40And then in the future, we can make decisions based on that algorithm.
- 8:46:44Once we have some good estimate for every state and for every action,
- 8:46:48what the value is of taking that action, then we
- 8:46:50can do something like implement a greedy decision making policy.
- 8:46:54That if I am in a state and I want to know what action
- 8:46:57should I take in that state, well, then I
- 8:47:00consider for all of my possible actions, what is the value of QSA?
- 8:47:05What is my estimated value of taking that action in that state?
- 8:47:08And I will just pick the action that has the highest value
- 8:47:12after I evaluate that expression.
- 8:47:15So I pick the action that has the highest value.
- 8:47:17And based on that, that tells me what action I should take.
- 8:47:19At any given state that I'm in, I can just greedily say across all my actions,
- 8:47:24this action gives me the highest expected value.
- 8:47:27And so I'll go ahead and choose that action as the action that I take as well.
- 8:47:33But there is a downside to this kind of approach.
- 8:47:36And then downside comes up in a situation like this,
- 8:47:38where we know that there is some solution that gets me to the reward.
- 8:47:44And our agent has been able to figure that out.
- 8:47:46But it might not necessarily be the best way or the fastest way.
- 8:47:49If the agent is allowed to explore a little bit more,
- 8:47:52it might find that it can get the reward faster
- 8:47:55by taking some other route instead, by going through this particular path
- 8:47:59that is a faster way to get to that ultimate goal.
- 8:48:04And maybe we would like for the agent to be able to figure that out as well.
- 8:48:07But if the agent always takes the actions that it knows to be best,
- 8:48:11well, when it gets to this particular square,
- 8:48:13it doesn't know that this is a good action because it's never really tried it.
- 8:48:17But it knows that going down eventually leads its way to this reward.
- 8:48:21So it might learn in the future that it should just always take this route
- 8:48:25and it's never going to explore and go along that route instead.
- 8:48:29So in reinforcement learning, there is this tension
- 8:48:32between exploration and exploitation.
- 8:48:35And exploitation generally refers to using knowledge that the AI already has.
- 8:48:40The AI already knows that this is a move that leads to reward.
- 8:48:43So we'll go ahead and use that move.
- 8:48:45And exploration is all about exploring other actions
- 8:48:49that we may not have explored as thoroughly before
- 8:48:51because maybe one of these actions, even if I don't know anything about it,
- 8:48:54might lead to better rewards faster or to more rewards in the future.
- 8:49:00And so an agent that only ever exploits information and never explores
- 8:49:04might be able to get reward, but it might not maximize its rewards
- 8:49:07because it doesn't know what other possibilities are out there,
- 8:49:10possibilities that we only know about by taking advantage of exploration.
- 8:49:15And so how can we try and address this?
- 8:49:17Well, one possible solution is known as the Epsilon greedy algorithm,
- 8:49:21where we set Epsilon equal to how often we want to just make a random move,
- 8:49:26where occasionally we will just make a random move in order to say,
- 8:49:29let's try to explore and see what happens.
- 8:49:33And then the logic of the algorithm will be with probability 1 minus Epsilon,
- 8:49:38choose the estimated best move.
- 8:49:40In a greedy case, we'd always choose the best move.
- 8:49:43But in Epsilon greedy, we're most of the time
- 8:49:46going to choose the best move or sometimes going to choose the best move.
- 8:49:50But sometimes with probability Epsilon, we're
- 8:49:53going to choose a random move instead.
- 8:49:56So every time we're faced with the ability to take an action,
- 8:49:58sometimes we're going to choose the best move.
- 8:50:00Sometimes we're just going to choose a random move.
- 8:50:03So this type of algorithm can be quite powerful in a reinforcement learning
- 8:50:07context by not always just choosing the best possible move right now,
- 8:50:11but sometimes, especially early on, allowing yourself
- 8:50:14to make random moves that allow you to explore various different possible
- 8:50:18states and actions more, and maybe over time,
- 8:50:20you might decrease your value of Epsilon.
- 8:50:23More and more often, choosing the best move
- 8:50:25after you're more confident that you've explored
- 8:50:27what all of the possibilities actually are.
- 8:50:30So we can put this into practice.
- 8:50:32And one very common application of reinforcement learning
- 8:50:34is in game playing, that if you want to teach an agent how to play a game,
- 8:50:38you just let the agent play the game a whole bunch.
- 8:50:41And then the reward signal happens at the end of the game.
- 8:50:44When the game is over, if our AI won the game,
- 8:50:47it gets a reward of like 1, for example.
- 8:50:49And if it lost the game, it gets a reward of negative 1.
- 8:50:53And from that, it begins to learn what actions are good
- 8:50:56and what actions are bad.
- 8:50:57You don't have to tell the AI what's good and what's bad,
- 8:50:59but the AI figures it out based on that reward.
- 8:51:01Winning the game is some signal, losing the game is some signal,
- 8:51:04and based on all of that, it begins to figure out
- 8:51:07what decisions it should actually make.
- 8:51:09So one very simple game, which you may have played before, is a game called
- 8:51:13Nim.
- 8:51:13And in the game of Nim, you've got a whole bunch of objects
- 8:51:16in a whole bunch of different piles, where here I've
- 8:51:18represented each pile as an individual row.
- 8:51:20So you've got one object in the first pile,
- 8:51:22three in the second pile, five in the third pile, seven in the fourth pile.
- 8:51:26And the game of Nim is a two player game
- 8:51:28where players take turns removing objects from piles.
- 8:51:31And the rule is that on any given turn, you
- 8:51:34were allowed to remove as many objects as you want from any one of these piles,
- 8:51:39any one of these rows.
- 8:51:40You have to remove at least one object, but you
- 8:51:42remove as many as you want from exactly one of the piles.
- 8:51:46And whoever takes the last object loses.
- 8:51:50So player one might remove four from this pile here.
- 8:51:54Player two might remove four from this pile here.
- 8:51:57So now we've got four piles left, one, three, one, and three.
- 8:52:00Player one might remove the entirety of the second pile.
- 8:52:03Player two, if they're being strategic, might remove two from the third pile.
- 8:52:09Now we've got three piles left, each with one object left.
- 8:52:13Player one might remove one from one pile.
- 8:52:15Player two removes one from the other pile.
- 8:52:17And now player one is left with choosing this one object from the last pile,
- 8:52:22at which point player one loses the game.
- 8:52:24So fairly simple game.
- 8:52:25Piles of objects, any turn you choose how many objects
- 8:52:28to remove from a pile, whoever removes the last object loses.
- 8:52:33And this is the type of game you could encode into an AI fairly easily,
- 8:52:36because the states are really just four numbers.
- 8:52:39Every state is just how many objects in each of the four piles.
- 8:52:43And the actions are things like, how many
- 8:52:45am I going to remove from each one of these individual piles?
- 8:52:49And the reward happens at the end, that if you
- 8:52:51were the player that had to remove the last object,
- 8:52:53then you get some sort of punishment.
- 8:52:55But if you were not, and the other player
- 8:52:57had to remove the last object, well, then you get some sort of reward.
- 8:53:01So we could actually try and show a demonstration of this,
- 8:53:04that I've implemented an AI to play the game of Nim.
- 8:53:08All right, so here, what we're going to do is create an AI
- 8:53:11as a result of training the AI on some number of games,
- 8:53:15that the AI is going to play against itself, where the idea is the AI will
- 8:53:18play games against itself, learn from each of those experiences,
- 8:53:22and learn what to do in the future.
- 8:53:23And then I, the human, will play against the AI.
- 8:53:26So initially, we'll say train zero times,
- 8:53:28meaning we're not going to let the AI play any practice games against itself
- 8:53:32in order to learn from its experiences.
- 8:53:34We're just going to see how well it plays.
- 8:53:36And it looks like there are four piles.
- 8:53:38I can choose how many I remove from any one of the piles.
- 8:53:41So maybe from pile three, I will remove five objects, for example.
- 8:53:46So now, AI chose to take one item from pile zero.
- 8:53:50So I'm left with these piles now, for example.
- 8:53:53And so here, I could choose maybe to say, I
- 8:53:55would like to remove from pile two, I'll remove all five of them,
- 8:54:00for example.
- 8:54:01And so AI chose to take two away from pile one.
- 8:54:04Now I'm left with one pile that has one object, one pile that has two objects.
- 8:54:08So from pile three, I will remove two objects.
- 8:54:11And now I've left the AI with no choice but to take that last one.
- 8:54:15And so the game is over, and I was able to win.
- 8:54:17But I did so because the AI was really just playing randomly.
- 8:54:20It didn't have any prior experience that it was using in order
- 8:54:23to make these sorts of judgments.
- 8:54:24Now let me let the AI train itself on 10,000 games.
- 8:54:29I'm going to let the AI play 10,000 games of nim against itself.
- 8:54:32Every time it wins or loses, it's going to learn from that experience
- 8:54:36and learn in the future what to do and what not to do.
- 8:54:39So here then, I'll go ahead and run this again.
- 8:54:42And now you see the AI running through a whole bunch of training games,
- 8:54:4510,000 training games against itself.
- 8:54:47And now it's going to let me make these sorts of decisions.
- 8:54:50So now I'm going to play against the AI.
- 8:54:52Maybe I'll remove one from pile three.
- 8:54:55And the AI took everything from pile three, so I'm left with three piles.
- 8:54:59I'll go ahead and from pile two maybe remove three items.
- 8:55:04And the AI removes one item from pile zero.
- 8:55:07I'm left with two piles, each of which has two items in it.
- 8:55:10I'll remove one from pile one, I guess.
- 8:55:14And the AI took two from pile two, leaving me with no choice
- 8:55:17but to take one away from pile one.
- 8:55:20So it seems like after playing 10,000 games of nim against itself,
- 8:55:24the AI has learned something about what states and what actions tend to be good
- 8:55:28and has begun to learn some sort of pattern for how
- 8:55:31to predict what actions are going to be good
- 8:55:33and what actions are going to be bad in any given state.
- 8:55:37So reinforcement learning can be a very powerful technique
- 8:55:39for achieving these sorts of game-playing agents, agents
- 8:55:42that are able to play a game well just by learning from experience,
- 8:55:45whether that's playing against other people
- 8:55:47or by playing against itself and learning from those experiences as well.
- 8:55:51Now, nim is a bit of an easy game to use reinforcement learning for
- 8:55:55because there are so few states.
- 8:55:57There are only states that are as many as how many different objects
- 8:55:59are in each of these various different piles.
- 8:56:02You might imagine that it's going to be harder if you think of a game like chess
- 8:56:06or games where there are many, many more states and many, many more actions
- 8:56:09that you can imagine taking, where it's not
- 8:56:11going to be as easy to learn for every state and for every action
- 8:56:15what the value is going to be.
- 8:56:17So oftentimes in that case, we can't necessarily
- 8:56:20learn exactly what the value is for every state and for every action,
- 8:56:23but we can approximate it.
- 8:56:25So much as we saw with minimax, so we could use a depth-limiting approach
- 8:56:28to stop calculating at a certain point in time,
- 8:56:31we can do a similar type of approximation known
- 8:56:34as function approximation in a reinforcement learning context
- 8:56:37where instead of learning a value of q for every state and every action,
- 8:56:42we just have some function that estimates what the value is
- 8:56:46for taking this action in this particular state that
- 8:56:49might be based on various different features of the state
- 8:56:53that the agent happens to be in, where you might have
- 8:56:55to choose what those features actually are.
- 8:56:58But you can begin to learn some patterns that generalize beyond one
- 8:57:02specific state and one specific action that you can begin to learn
- 8:57:05if certain features tend to be good things or bad things.
- 8:57:08Reinforcement learning can allow you, using a very similar mechanism,
- 8:57:11to generalize beyond one particular state and say,
- 8:57:14if this other state looks kind of like this state,
- 8:57:17then maybe the similar types of actions that worked in one state
- 8:57:20will also work in another state as well.
- 8:57:23And so this type of approach can be quite helpful
- 8:57:25as you begin to deal with reinforcement learning that
- 8:57:27exist in larger and larger state spaces where it's just not feasible
- 8:57:31to explore all of the possible states that could actually exist.
- 8:57:36So there, then, are two of the main categories of reinforcement learning.
- 8:57:39Supervised learning, where you have labeled input and output pairs,
- 8:57:42and reinforcement learning, where an agent learns from rewards or punishments
- 8:57:46that it receives.
- 8:57:47The third major category of machine learning
- 8:57:49that we'll just touch on briefly is known as unsupervised learning.
- 8:57:53And unsupervised learning happens when we have data
- 8:57:56without any additional feedback, without labels,
- 8:57:59that in the supervised learning case, all of our data had labels.
- 8:58:02We labeled the data point with whether that was a rainy day or not rainy day.
- 8:58:06And using those labels, we were able to infer what the pattern was.
- 8:58:09Or we labeled data as a counterfeit banknote or not a counterfeit.
- 8:58:13And using those labels, we were able to draw inferences and patterns
- 8:58:16to figure out what does a banknote look like versus not.
- 8:58:20In unsupervised learning, we don't have any access to any of those labels.
- 8:58:25But we still would like to learn some of those patterns.
- 8:58:28And one of the tasks that you might want to perform in unsupervised learning
- 8:58:31is something like clustering, where clustering is just
- 8:58:34the task of, given some set of objects, organize it
- 8:58:37into distinct clusters, groups of objects that are similar to one another.
- 8:58:42And there's lots of applications for clustering.
- 8:58:44It comes up in genetic research, where you might have
- 8:58:47a whole bunch of different genes and you want to cluster them into similar genes
- 8:58:50if you're trying to analyze them across a population or across species.
- 8:58:54It comes up in an image if you want to take all the pixels of an image,
- 8:58:57cluster them into different parts of the image.
- 8:58:59Comes a lot up in market research if you want to divide your consumers
- 8:59:03into different groups so you know which groups to target with certain types
- 8:59:06of product advertisements, for example, and a number of other contexts
- 8:59:10as well in which clustering can be very applicable.
- 8:59:13One technique for clustering is an algorithm known as k-means clustering.
- 8:59:17And what k-means clustering is going to do
- 8:59:20is it is going to divide all of our data points into k different clusters.
- 8:59:24And it's going to do so by repeating this process of assigning points
- 8:59:28to clusters and then moving around those clusters at centers.
- 8:59:32We're going to define a cluster by its center, the middle of the cluster,
- 8:59:36and then assign points to that cluster based on which
- 8:59:39center is closest to that point.
- 8:59:42And I'll show you an example of that now.
- 8:59:44Here, for example, I have a whole bunch of unlabeled data,
- 8:59:47just various data points that are in some sort of graphical space.
- 8:59:51And I would like to group them into various different clusters.
- 8:59:55But I don't know how to do that originally.
- 8:59:57And let's say I want to assign like three clusters to this group.
- 9:00:00And you have to choose how many clusters you want in k-means clustering
- 9:00:03that you could try multiple and see how well those values perform.
- 9:00:06But I'll start just by randomly picking some places
- 9:00:09to put the centers of those clusters.
- 9:00:12Maybe I have a blue cluster, a red cluster, and a green cluster.
- 9:00:15And I'm going to start with the centers of those clusters
- 9:00:18just being in these three locations here.
- 9:00:20And what k-means clustering tells us to do
- 9:00:23is once I have the centers of the clusters,
- 9:00:25assign every point to a cluster based on which cluster center it is closest to.
- 9:00:32So we end up with something like this, where all of these points
- 9:00:35are closer to the blue cluster center than any other cluster center.
- 9:00:40All of these points here are closer to the green cluster
- 9:00:43center than any other cluster center.
- 9:00:45And then these two points plus these points over here,
- 9:00:48those are all closest to the red cluster center instead.
- 9:00:53So here then is one possible assignment of all these points
- 9:00:57to three different clusters.
- 9:00:58But it's not great that it seems like in this red cluster,
- 9:01:01these points are kind of far apart.
- 9:01:02In this green cluster, these points are kind of far apart.
- 9:01:05It might not be my ideal choice of how I would cluster
- 9:01:08these various different data points.
- 9:01:10But k-means clustering is an iterative process
- 9:01:13that after I do this, there is a next step, which
- 9:01:16is that after I've assigned all of the points to the cluster center
- 9:01:19that it is nearest to, we are going to re-center the clusters,
- 9:01:24meaning take the cluster centers, these diamond shapes here,
- 9:01:27and move them to the middle, or the average,
- 9:01:30effectively, of all of the points that are in that cluster.
- 9:01:33So we'll take this blue point, this blue center,
- 9:01:36and go ahead and move it to the middle or to the center of all
- 9:01:39of the points that were assigned to the blue cluster,
- 9:01:41moving it slightly to the right in this case.
- 9:01:43And we'll do the same thing for red.
- 9:01:45We'll move the cluster center to the middle of all of these points,
- 9:01:49weighted by how many points there are.
- 9:01:51There are more points over here, so the red center ends up
- 9:01:55moving a little bit further that way.
- 9:01:56And likewise, for the green center, there are many more points
- 9:01:59on this side of the green center.
- 9:02:01So the green center ends up being pulled a little bit further
- 9:02:04in this direction.
- 9:02:06So we re-center all of the clusters, and then we repeat the process.
- 9:02:10We go ahead and now reassign all of the points to the cluster center
- 9:02:14that they are now closest to.
- 9:02:16And now that we've moved around the cluster centers,
- 9:02:18these cluster assignments might change.
- 9:02:20That this point originally was closer to the red cluster center,
- 9:02:23but now it's actually closer to the blue cluster center.
- 9:02:26Same goes for this point as well.
- 9:02:28And these three points that were originally closer to the green cluster
- 9:02:31center are now closer to the red cluster center instead.
- 9:02:36So we can reassign what colors or which clusters each of these data points
- 9:02:41belongs to, and then repeat the process again,
- 9:02:43moving each of these cluster means and the middles of the clusterism
- 9:02:47to the mean, the average, of all of the other points that happen to be there,
- 9:02:52and repeat the process again.
- 9:02:54Go ahead and assign each of the points to the cluster
- 9:02:57that they are closest to.
- 9:02:58So once we reach a point where we've assigned all the points to clusters
- 9:03:01to the cluster that they are nearest to, and nothing changed,
- 9:03:05we've reached a sort of equilibrium in this situation,
- 9:03:07where no points are changing their allegiance.
- 9:03:09And as a result, we can declare this algorithm is now over.
- 9:03:12And we now have some assignment of each of these points
- 9:03:15into three different clusters.
- 9:03:17And it looks like we did a pretty good job of trying
- 9:03:19to identify which points are more similar to one another
- 9:03:22than they are to points in other groups.
- 9:03:24So we have the green cluster down here, this blue cluster here,
- 9:03:27and then this red cluster over there as well.
- 9:03:30And we did so without any access to some labels
- 9:03:33to tell us what these various different clusters were.
- 9:03:35We just used an algorithm in an unsupervised sense
- 9:03:38without any of those labels to figure out which points
- 9:03:41belonged to which categories.
- 9:03:43And again, lots of applications for this type of clustering technique.
- 9:03:47And there are many more algorithms in each of these various different fields
- 9:03:50within machine learning, supervised and reinforcement and unsupervised.
- 9:03:54But those are many of the big picture foundational ideas
- 9:03:57that underlie a lot of these techniques, where these are the problems
- 9:04:00that we're trying to solve.
- 9:04:01And we try and solve those problems using
- 9:04:03a number of different methods of trying to take data and learn
- 9:04:06patterns in that data, whether that's trying
- 9:04:08to find neighboring data points that are similar
- 9:04:10or trying to minimize some sort of loss function
- 9:04:13or any number of other techniques that allow us to begin to try
- 9:04:17to solve these sorts of problems.
- 9:04:19That then was a look at some of the principles
- 9:04:21that are at the foundation of modern machine learning,
- 9:04:23this ability to take data and learn from that data
- 9:04:26so that the computer can perform a task even
- 9:04:28if they haven't explicitly been given instructions
- 9:04:31in order to do so.
- 9:04:32Next time, we'll continue this conversation about machine learning,
- 9:04:35looking at other techniques we can use for solving these sorts of problems.
- 9:04:38We'll see you then.
- 9:04:41All right, welcome back, everyone, to an introduction
- 9:05:01to artificial intelligence with Python.
- 9:05:03Now, last time, we took a look at machine learning,
- 9:05:05a set of techniques that computers can use in order to take a set of data
- 9:05:09and learn some patterns inside of that data,
- 9:05:11learn how to perform a task even if we the programmers didn't
- 9:05:14give the computer explicit instructions for how to perform that task.
- 9:05:18Today, we transition to one of the most popular techniques and tools
- 9:05:21within machine learning, that of neural networks.
- 9:05:24And neural networks were inspired as early as the 1940s
- 9:05:27by researchers who were thinking about how it is that humans learn,
- 9:05:30studying neuroscience in the human brain and trying
- 9:05:33to see whether or not we could apply those same ideas to computers
- 9:05:36as well and model computer learning off of human learning.
- 9:05:39So how is the brain structured?
- 9:05:41Well, very simply put, the brain consists of a whole bunch of neurons.
- 9:05:45And those neurons are connected to one another
- 9:05:47and communicate with one another in some way.
- 9:05:49In particular, if you think about the structure of a biological neural
- 9:05:52network, something like this, there are a couple of key properties
- 9:05:55that scientists observed.
- 9:05:57One was that these neurons are connected to each other
- 9:05:59and receive electrical signals from one another,
- 9:06:01that one neuron can propagate electrical signals to another neuron.
- 9:06:06And another point is that neurons process those input signals
- 9:06:09and then can be activated, that a neuron becomes activated at a certain point
- 9:06:12and then can propagate further signals onto neurons in the future.
- 9:06:16And so the question then became, could we
- 9:06:18take this biological idea of how it is that humans learn with brains
- 9:06:22and with neurons and apply that to a machine as well,
- 9:06:25in effect designing an artificial neural network, or an ANN,
- 9:06:29which will be a mathematical model for learning
- 9:06:31that is inspired by these biological neural networks?
- 9:06:34And what artificial neural networks will allow us to do
- 9:06:37is they will first be able to model some sort of mathematical function.
- 9:06:40Every time you look at a neural network, which
- 9:06:42we'll see more of later today, each one of them
- 9:06:44is really just some mathematical function that
- 9:06:46is mapping certain inputs to particular outputs based
- 9:06:50on the structure of the network, that depending on where we place
- 9:06:53particular units inside of this neural network,
- 9:06:55that's going to determine how it is that the network is going to function.
- 9:06:59And in particular, artificial neural networks
- 9:07:01are going to lend themselves to a way that we can learn what the network's
- 9:07:05parameters should be.
- 9:07:07We'll see more on that in just a moment.
- 9:07:08But in effect, we want a model such that it
- 9:07:11is easy for us to be able to write some code that
- 9:07:13allows for the network to be able to figure out
- 9:07:16how to model the right mathematical function given
- 9:07:18a particular set of input data.
- 9:07:20So in order to create our artificial neural network,
- 9:07:23instead of using biological neurons, we're just
- 9:07:25going to use what we're going to call units, units inside of a neural
- 9:07:28network, which we can represent kind of like a node in a graph, which
- 9:07:31will here be represented just by a blue circle like this.
- 9:07:34And these artificial units, these artificial neurons,
- 9:07:37can be connected to one another.
- 9:07:39So here, for instance, we have two units that
- 9:07:41are connected by this edge inside of this graph, effectively.
- 9:07:46And so what we're going to do now is think
- 9:07:48of this idea as some sort of mapping from inputs to outputs.
- 9:07:51So we have one unit that is connected to another unit
- 9:07:54that we might think of this side of the input and that side of the output.
- 9:07:58And what we're trying to do then is to figure out
- 9:08:00how to solve a problem, how to model some sort of mathematical function.
- 9:08:04And this might take the form of something
- 9:08:05we saw last time, which was something like we have certain inputs,
- 9:08:08like variables x1 and x2.
- 9:08:10And given those inputs, we want to perform some sort of task,
- 9:08:13a task like predicting whether or not it's going to rain.
- 9:08:16And ideally, we'd like some way, given these inputs, x1 and x2,
- 9:08:20which stand for some sort of variables to do with the weather,
- 9:08:23we would like to be able to predict, in this case, a Boolean classification.
- 9:08:27Is it going to rain, or is it not going to rain?
- 9:08:30And we did this last time by way of a mathematical function.
- 9:08:33We defined some function, h, for our hypothesis function,
- 9:08:36that took as input x1 and x2, the two inputs that we cared about processing,
- 9:08:41in order to determine whether we thought it was going to rain
- 9:08:44or whether we thought it was not going to rain.
- 9:08:46The question then becomes, what does this hypothesis function
- 9:08:48do in order to make that determination?
- 9:08:51And we decided last time to use a linear combination of these input variables
- 9:08:56to determine what the output should be.
- 9:08:58So our hypothesis function was equal to something like this.
- 9:09:02Weight 0 plus weight 1 times x1 plus weight 2 times x2.
- 9:09:07So what's going on here is that x1 and x2, those are input variables,
- 9:09:11the inputs to this hypothesis function.
- 9:09:15And each of those input variables is being multiplied
- 9:09:17by some weight, which is just some number.
- 9:09:20So x1 is being multiplied by weight 1, x2 is being multiplied by weight 2.
- 9:09:25And we have this additional weight, weight 0,
- 9:09:27that doesn't get multiplied by an input variable at all,
- 9:09:30that just serves to either move the function up
- 9:09:32or move the function's value down.
- 9:09:33You can think of this as either a weight that's just
- 9:09:36multiplied by some dummy value, like the number 1.
- 9:09:38It's multiplied by 1, and so it's not multiplied by anything.
- 9:09:41Or sometimes, you'll see in the literature,
- 9:09:43people call this variable weight 0 a bias,
- 9:09:46so that you can think of these variables as slightly different.
- 9:09:48We have weights that are multiplied by the input,
- 9:09:50and we separately add some bias to the result as well.
- 9:09:54You'll hear both of those terminologies used
- 9:09:56when people talk about neural networks and machine learning.
- 9:09:59So in effect, what we've done here is that in order
- 9:10:02to define a hypothesis function, we just need to decide and figure out
- 9:10:06what these weights should be to determine
- 9:10:08what values to multiply by our inputs to get some sort of result.
- 9:10:12Of course, at the end of this, what we need to do
- 9:10:14is make some sort of classification, like rainy or not rainy.
- 9:10:18And to do that, we use some sort of function
- 9:10:20that defines some sort of threshold.
- 9:10:22And so we saw, for instance, the step function,
- 9:10:25which is defined as 1 if the result of multiplying the weights by the inputs
- 9:10:30is at least 0, otherwise it's 0.
- 9:10:32And you can think of this line down the middle
- 9:10:34as kind of like a dotted line.
- 9:10:35Effectively, it stays at 0 all the way up to one point,
- 9:10:38and then the function steps or jumps up to 1.
- 9:10:41So it's 0 before it reaches some threshold,
- 9:10:43and then it's 1 after it reaches a particular threshold.
- 9:10:46And so this was one way we could define what
- 9:10:49will come to call an activation function, a function that
- 9:10:51determines when it is that this output becomes active, changes to 1
- 9:10:56instead of being a 0.
- 9:10:58But we also saw that if we didn't just want a purely binary classification,
- 9:11:02we didn't want purely 1 or 0, but we wanted
- 9:11:04to allow for some in-between real numbered values,
- 9:11:07we could use a different function.
- 9:11:09And there are a number of choices, but the one that we looked at
- 9:11:11was the logistic sigmoid function that has sort of an s-shaped curve,
- 9:11:15where we could represent this as a probability that
- 9:11:18may be somewhere in between the probability of rain
- 9:11:20or something like 0.5.
- 9:11:22Maybe a little bit later, the probability of rain is 0.8.
- 9:11:25And so rather than just have a binary classification of 0 or 1,
- 9:11:29we could allow for numbers that are in between as well.
- 9:11:32And it turns out there are many other different types of activation
- 9:11:35functions, where an activation function just
- 9:11:37takes the output of multiplying the weights together and adding that bias,
- 9:11:41and then figuring out what the actual output should be.
- 9:11:43Another popular one is the rectified linear unit, otherwise known as ReLU.
- 9:11:48And the way that works is that it just takes its input
- 9:11:50and takes the maximum of that input and 0.
- 9:11:52So if it's positive, it remains unchanged.
- 9:11:55But if it's 0, if it's negative, it goes ahead and levels out at 0.
- 9:11:59And there are other activation functions that we could choose as well.
- 9:12:02But in short, each of these activation functions,
- 9:12:04you can just think of as a function that gets applied
- 9:12:07to the result of all of this computation.
- 9:12:10We take some function g and apply it to the result of all of that calculation.
- 9:12:15And this then is what we saw last time, the way
- 9:12:17of defining some hypothesis function that takes in inputs,
- 9:12:20calculate some linear combination of those inputs,
- 9:12:23and then passes it through some sort of activation function to get our output.
- 9:12:28And this actually turns out to be the model for the simplest of neural
- 9:12:32networks, that we're going to instead represent this mathematical idea
- 9:12:36graphically by using a structure like this.
- 9:12:39Here then is a neural network that has two inputs.
- 9:12:42We can think of this as x1 and this as x2.
- 9:12:44And then one output, which you can think of as classifying whether or not
- 9:12:48we think it's going to rain or not rain, for example,
- 9:12:50in this particular instance.
- 9:12:52And so how exactly does this model work?
- 9:12:54Well, each of these two inputs represents one of our input variables,
- 9:12:57x1 and x2.
- 9:12:59And notice that these inputs are connected to this output via these edges,
- 9:13:05which are going to be defined by their weights.
- 9:13:06So these edges each have a weight associated with them, weight 1 and weight
- 9:13:102.
- 9:13:12And then this output unit, what it's going to do
- 9:13:14is it is going to calculate an output based on those inputs
- 9:13:17and based on those weights.
- 9:13:19This output unit is going to multiply all the inputs by their weights,
- 9:13:23add in this bias term, which you can think of as an extra w0 term
- 9:13:26that gets added into it, and then we pass it through an activation function.
- 9:13:31So this then is just a graphical way of representing the same idea
- 9:13:34we saw last time just mathematically.
- 9:13:36And we're going to call this a very simple neural network.
- 9:13:40And we'd like for this neural network to be
- 9:13:42able to learn how to calculate some function,
- 9:13:44that we want some function for the neural network to learn.
- 9:13:46And the neural network is going to learn what should the values of w0,
- 9:13:50w1, and w2 be?
- 9:13:52What should the activation function be in order
- 9:13:54to get the result that we would expect?
- 9:13:57So we can actually take a look at an example of this.
- 9:13:59What then is a very simple function that we might calculate?
- 9:14:02Well, if we recall back from when we were looking at propositional logic,
- 9:14:06one of the simplest functions we looked at
- 9:14:07was something like the or function that takes two inputs, x and y,
- 9:14:12and outputs 1, otherwise known as true, if either one of the inputs
- 9:14:16or both of them are 1, and outputs of 0 if both of the inputs are 0 or false.
- 9:14:22So this then is the or function.
- 9:14:23And this was the truth table for the or function,
- 9:14:25that as long as either of the inputs are 1, the output of the function is 1,
- 9:14:29and the only case where the output is 0 is where both of the inputs are 0.
- 9:14:34So the question is, how could we take this and train a neural network
- 9:14:38to be able to learn this particular function?
- 9:14:40What would those weights look like?
- 9:14:42Well, we could do something like this.
- 9:14:44Here's our neural network.
- 9:14:45And I'll propose that in order to calculate the or function,
- 9:14:48we're going to use a value of 1 for each of the weights.
- 9:14:52And we'll use a bias of negative 1.
- 9:14:55And then we'll just use this step function as our activation function.
- 9:14:59How then does this work?
- 9:15:00Well, if I wanted to calculate something like 0 or 0,
- 9:15:04which we know to be 0 because false or false is false, then what are we going
- 9:15:08to do?
- 9:15:08Well, our output unit is going to calculate this input multiplied
- 9:15:12by the weight, 0 times 1, that's 0.
- 9:15:14Same thing here, 0 times 1, that's 0.
- 9:15:17And we'll add to that the bias minus 1.
- 9:15:21So that'll give us a result of negative 1.
- 9:15:23If we plot that on our activation function, negative 1 is here.
- 9:15:26It's before the threshold, which means either 0 or 1.
- 9:15:30It's only 1 after the threshold.
- 9:15:32Since negative 1 is before the threshold,
- 9:15:34the output that this unit provides is going to be 0.
- 9:15:38And that's what we would expect it to be, that 0 or 0 should be 0.
- 9:15:43What if instead we had had 1 or 0, where this is the number 1?
- 9:15:47Well, in this case, in order to calculate what the output is going to be,
- 9:15:50we again have to do this weighted sum, 1 times 1, that's 1.
- 9:15:550 times 1, that's 0.
- 9:15:57Sum of that so far is 1.
- 9:15:59Add negative 1 to that.
- 9:16:00Well, then the output is 0.
- 9:16:02And if we plot 0 on the step function, 0 ends up being here.
- 9:16:05It's just at the threshold.
- 9:16:07And so the output here is going to be 1, because the output of 1 or 0,
- 9:16:11that's 1.
- 9:16:12So that's what we would expect as well.
- 9:16:13And just for one more example, if I had 1 or 1, what would the result be?
- 9:16:17Well, 1 times 1 is 1.
- 9:16:191 times 1 is 1.
- 9:16:20The sum of those is 2.
- 9:16:22I add the bias term to that.
- 9:16:23I get the number 1.
- 9:16:241 plotted on this graph is way over there.
- 9:16:27That's well beyond the threshold.
- 9:16:28And so this output is going to be 1 as well.
- 9:16:31The output is always 0 or 1, depending on whether or not
- 9:16:34we're past the threshold.
- 9:16:35And this neural network then models the OR function, a very simple function,
- 9:16:39definitely.
- 9:16:40But it still is able to model it correctly.
- 9:16:42If I give it the inputs, it will tell me what x1 or x2 happens to be.
- 9:16:48And you could imagine trying to do this for other functions as well.
- 9:16:50A function like the AND function, for instance, that takes two inputs
- 9:16:55and calculates whether both x and y are true.
- 9:16:59So if x is 1 and y is 1, then the output of x and y is 1.
- 9:17:04But in all the other cases, the output is 0.
- 9:17:07How could we model that inside of a neural network as well?
- 9:17:10Well, it turns out we could do it in the same way,
- 9:17:13except instead of negative 1 as the bias,
- 9:17:16we can use negative 2 as the bias instead.
- 9:17:20What does that end up looking like?
- 9:17:21Well, if I had 1 and 1, that should be 1, because 1 true and true
- 9:17:25is equal to true.
- 9:17:27Well, I take 1 times 1, that's 1.
- 9:17:291 times 1 is 1.
- 9:17:30I get a total sum of 2 so far.
- 9:17:32Now I add the bias of negative 2, and I get the value 0.
- 9:17:35And 0, when I plot it on the activation function,
- 9:17:38is just past that threshold, and so the output is going to be 1.
- 9:17:42But if I had any other input, for example, like 1 and 0,
- 9:17:46well, the weighted sum of these is 1 plus 0 is going to be 1.
- 9:17:51Minus 2 is going to give us negative 1, and negative 1
- 9:17:53is not past that threshold, and so the output is going to be 0.
- 9:17:58So those then are some very simple functions
- 9:18:01that we can model using a neural network that has two inputs and one output,
- 9:18:05where our goal is to be able to figure out what those weights should be
- 9:18:08in order to determine what the output should be.
- 9:18:11And you could imagine generalizing this to calculate more complex functions
- 9:18:14as well, that maybe, given the humidity and the pressure,
- 9:18:17we want to calculate what's the probability that it's going to rain,
- 9:18:20for example.
- 9:18:20Or we might want to do a regression-style problem.
- 9:18:22We're given some amount of advertising, and given what month it is maybe,
- 9:18:26we want to predict what our expected sales are
- 9:18:28going to be for that particular month.
- 9:18:30So you could imagine these inputs and outputs being different as well.
- 9:18:34And it turns out that in some problems, we're not just
- 9:18:36going to have two inputs, and the nice thing about these neural networks
- 9:18:39is that we can compose multiple units together,
- 9:18:42make our networks more complex just by adding more units
- 9:18:46into this particular neural network.
- 9:18:48So the network we've been looking at has two inputs and one output.
- 9:18:52But we could just as easily say, let's go ahead and have three inputs in there,
- 9:18:56or have even more inputs, where we could arbitrarily
- 9:18:58decide however many inputs there are to our problem, all going
- 9:19:02to be calculating some sort of output that we care about figuring out
- 9:19:06the value of.
- 9:19:07How then does the math work for figuring out that output?
- 9:19:10Well, it's going to work in a very similar way.
- 9:19:12In the case of two inputs, we had two weights indicated by these edges,
- 9:19:16and we multiplied the weights by the numbers, adding this bias term.
- 9:19:20And we'll do the same thing in the other cases as well.
- 9:19:22If I have three inputs, you'll imagine multiplying
- 9:19:25each of these three inputs by each of these weights.
- 9:19:27If I had five inputs instead, we're going to do the same thing.
- 9:19:31Here I'm saying sum up from 1 to 5, xi multiplied by weight i.
- 9:19:35So take each of the five input variables, multiply them
- 9:19:38by their corresponding weight, and then add the bias to that.
- 9:19:41So this would be a case where there are five inputs into this neural network,
- 9:19:45for example.
- 9:19:46But there could be more, arbitrarily many nodes
- 9:19:48that we want inside of this neural network, where each time we're just
- 9:19:51going to sum up all of those input variables multiplied by their weight
- 9:19:54and then add the bias term at the very end.
- 9:19:57And so this allows us to be able to represent problems
- 9:20:00that have even more inputs just by growing the size of our neural network.
- 9:20:05Now, the next question we might ask is a question about how it
- 9:20:08is that we train these neural networks.
- 9:20:10In the case of the or function and the and function,
- 9:20:13they were simple enough functions that I could just tell you,
- 9:20:16like here, what the weights should be.
- 9:20:17And you could probably reason through it yourself
- 9:20:19what the weights should be in order to calculate the output that you want.
- 9:20:23But in general, with functions like predicting sales
- 9:20:26or predicting whether or not it's going to rain,
- 9:20:27these are much trickier functions to be able to figure out.
- 9:20:30We would like the computer to have some mechanism
- 9:20:33of calculating what it is that the weights should be,
- 9:20:36how it is to set the weights so that our neural network is
- 9:20:39able to accurately model the function that we
- 9:20:41care about trying to estimate.
- 9:20:43And it turns out that the strategy for doing this,
- 9:20:45inspired by the domain of calculus, is a technique called gradient descent.
- 9:20:49And what gradient descent is, it is an algorithm
- 9:20:52for minimizing loss when you're training a neural network.
- 9:20:55And recall that loss refers to how bad our hypothesis
- 9:20:59function happens to be, that we can define certain loss functions.
- 9:21:03And we saw some examples of loss functions last time that just give us
- 9:21:06a number for any particular hypothesis, saying,
- 9:21:09how poorly does it model the data?
- 9:21:11How many examples does it get wrong?
- 9:21:13How are they worse or less bad as compared to other hypothesis functions
- 9:21:17that we might define?
- 9:21:19And this loss function is just a mathematical function.
- 9:21:22And when you have a mathematical function,
- 9:21:24in calculus what you could do is calculate
- 9:21:26something known as the gradient, which you can think of as like a slope.
- 9:21:29It's the direction the loss function is moving at any particular point.
- 9:21:32And what it's going to tell us is, in which direction
- 9:21:36should we be moving these weights in order to minimize the amount of loss?
- 9:21:41And so generally speaking, we won't get into the calculus of it.
- 9:21:43But the high level idea for gradient descent
- 9:21:46is going to look something like this.
- 9:21:47If we want to train a neural network, we'll go ahead and start just
- 9:21:51by choosing the weights randomly.
- 9:21:52Just pick random weights for all of the weights in the neural network.
- 9:21:56And then we'll use the input data that we have access
- 9:21:58to in order to train the network, in order
- 9:22:00to figure out what the weights should actually be.
- 9:22:02So we'll repeat this process again and again.
- 9:22:05The first step is we're going to calculate the gradient based
- 9:22:08on all of the data points.
- 9:22:09So we'll look at all the data and figure out
- 9:22:11what the gradient is at the place where we currently
- 9:22:13are for the current setting of the weights, which
- 9:22:15means in which direction should we move the weights in order
- 9:22:19to minimize the total amount of loss, in order to make our solution better.
- 9:22:24And once we've calculated that gradient, which direction
- 9:22:26we should move in the loss function, well,
- 9:22:29then we can just update those weights according to the gradient.
- 9:22:32Take a small step in the direction of those weights
- 9:22:35in order to try to make our solution a little bit better.
- 9:22:37And the size of the step that we take, that's going to vary.
- 9:22:40And you can choose that when you're training a particular neural network.
- 9:22:43But in short, the idea is going to be take all the data points,
- 9:22:46figure out based on those data points in what direction
- 9:22:48the weights should move, and then move the weights one small step
- 9:22:52in that direction.
- 9:22:53And if you repeat that process over and over again,
- 9:22:55adjusting the weights a little bit at a time based on all the data points,
- 9:22:58eventually you should end up with a pretty good solution
- 9:23:02to trying to solve this sort of problem.
- 9:23:04At least that's what we would hope to happen.
- 9:23:06Now, if you look at this algorithm, a good question
- 9:23:08to ask anytime you're analyzing an algorithm
- 9:23:10is what is going to be the expensive part of doing the calculation?
- 9:23:14What's going to take a lot of work to try to figure out?
- 9:23:17What is going to be expensive to calculate?
- 9:23:19And in particular, in the case of gradient descent,
- 9:23:22the really expensive part is this all data points part right here,
- 9:23:26having to take all of the data points and using all of those data points
- 9:23:30figure out what the gradient is at this particular setting of all
- 9:23:34of the weights.
- 9:23:34Because odds are in a big machine learning problem
- 9:23:37where you're trying to solve a big problem with a lot of data,
- 9:23:39you have a lot of data points in order to calculate.
- 9:23:41And figuring out the gradient based on all of those data points
- 9:23:44is going to be expensive.
- 9:23:46And you'll have to do it many times.
- 9:23:47You'll likely repeat this process again and again and again,
- 9:23:50going through all the data points, taking one small step over and over
- 9:23:54as you try and figure out what the optimal setting of those weights
- 9:23:57happens to be.
- 9:23:59It turns out that we would ideally like to be
- 9:24:02able to train our neural networks faster,
- 9:24:04to be able to more quickly converge to some sort of solution that
- 9:24:07is going to be a good solution to the problem.
- 9:24:10So in that case, there are alternatives to just standard gradient descent,
- 9:24:13which looks at all of the data points at once.
- 9:24:15We can employ a method like stochastic gradient descent,
- 9:24:18which will randomly just choose one data point at a time
- 9:24:22to calculate the gradient based on, instead of calculating it
- 9:24:25based on all of the data points.
- 9:24:27So the idea there is that we have some setting of the weights.
- 9:24:30We pick a data point.
- 9:24:31And based on that one data point, we figure out in which direction
- 9:24:34should we move all of the weights and move the weights in that small
- 9:24:37direction, then take another data point and do that again
- 9:24:39and repeat this process again and again,
- 9:24:41maybe looking at each of the data points multiple times,
- 9:24:44but each time only using one data point to calculate the gradient,
- 9:24:48to calculate which direction we should move in.
- 9:24:51Now, just using one data point instead of all of the data points
- 9:24:55probably gives us a less accurate estimate of what the gradient actually
- 9:24:58is.
- 9:24:59But on the plus side, it's going to be much faster
- 9:25:01to be able to calculate, that we can much more quickly calculate
- 9:25:04what the gradient is based on one data point,
- 9:25:07instead of calculating based on all of the data points
- 9:25:09and having to do all of that computational work again and again.
- 9:25:13So there are trade-offs here between looking at all of the data points
- 9:25:16and just looking at one data point.
- 9:25:18And it turns out that a middle ground that is also quite popular
- 9:25:21is a technique called mini-batch gradient descent, where the idea there
- 9:25:24is instead of looking at all of the data versus just a single point,
- 9:25:28we instead divide our data set up into small batches, groups of data points,
- 9:25:32where you can decide how big a particular batch is.
- 9:25:34But in short, you're just going to look at a small number of points
- 9:25:37at any given time, hopefully getting a more accurate estimate of the gradient,
- 9:25:41but also not requiring all of the computational effort needed
- 9:25:44to look at every single one of these data points.
- 9:25:48So gradient descent, then, is this technique
- 9:25:50that we can use in order to train these neural networks,
- 9:25:53in order to figure out what the setting of all of these weights
- 9:25:56should be if we want some way to try and get
- 9:25:59an accurate notion of how it is that this function should work,
- 9:26:02some way of modeling how to transform the inputs into particular outputs.
- 9:26:08Now, so far, the networks that we've taken a look at
- 9:26:11have all been structured similar to this.
- 9:26:13We have some number of inputs, maybe two or three or five or more.
- 9:26:17And then we have one output that is just predicting like rain or no rain
- 9:26:21or just predicting one particular value.
- 9:26:23But often in machine learning problems, we
- 9:26:25don't just care about one output.
- 9:26:27We might care about an output that has multiple different values
- 9:26:31associated with it.
- 9:26:32So in the same way that we could take a neural network
- 9:26:35and add units to the input layer, we can likewise add inputs or add outputs
- 9:26:40to the output layer as well.
- 9:26:41Instead of just one output, you could imagine we have two outputs,
- 9:26:44or we could have four outputs, for example,
- 9:26:47where in each case, as we add more inputs or add more outputs,
- 9:26:50if we want to keep this network fully connected between these two layers,
- 9:26:54we just need to add more weights, that now each of these input nodes
- 9:26:58has four weights associated with each of the four outputs.
- 9:27:02And that's true for each of these various different input nodes.
- 9:27:06So as we add nodes, we add more weights in order
- 9:27:09to make sure that each of the inputs can somehow
- 9:27:11be connected to each of the outputs so that each output
- 9:27:14value can be calculated based on what the value of the input happens to be.
- 9:27:19So what might a case be where we want multiple different output values?
- 9:27:23Well, you might consider that in the case of weather predicting,
- 9:27:26for example, we might not just care whether it's raining or not raining.
- 9:27:30There might be multiple different categories of weather
- 9:27:33that we would like to categorize the weather into.
- 9:27:35With just a single output variable, we can do a binary classification,
- 9:27:39like rain or no rain, for instance, 1 or 0.
- 9:27:42But it doesn't allow us to do much more than that.
- 9:27:45With multiple output variables, I might be
- 9:27:47able to use each one to predict something a little different.
- 9:27:50Maybe I want to categorize the weather into one of four different categories,
- 9:27:54something like is it going to be raining or sunny or cloudy or snowy.
- 9:27:58And I now have four output variables that
- 9:27:59can be used to represent maybe the probability that it is
- 9:28:03rainy as opposed to sunny as opposed to cloudy or as opposed to snowy.
- 9:28:08How then would this neural network work?
- 9:28:10Well, we have some input variables that represent some data
- 9:28:13that we have collected about the weather.
- 9:28:15Each of those inputs gets multiplied by each of these various different weights.
- 9:28:18We have more multiplications to do, but these
- 9:28:20are fairly quick mathematical operations to perform.
- 9:28:24And then what we get is after passing them
- 9:28:25through some sort of activation function in the outputs,
- 9:28:28we end up getting some sort of number, where that number, you might imagine,
- 9:28:32you could interpret as a probability, like a probability that it is one
- 9:28:36category as opposed to another category.
- 9:28:38So here we're saying that based on the inputs,
- 9:28:40we think there is a 10% chance that it's raining, a 60% chance that it's sunny,
- 9:28:45a 20% chance of cloudy, a 10% chance that it's snowy.
- 9:28:48And given that output, if these represent a probability distribution,
- 9:28:52well, then you could just pick whichever one has the highest value,
- 9:28:55in this case, sunny, and say that, well, most likely, we
- 9:28:58think that this categorization of inputs means that the output should be snowy
- 9:29:04or should be sunny.
- 9:29:05And that is what we would expect the weather to be in this particular instance.
- 9:29:09And so this allows us to do these sort of multi-class classifications,
- 9:29:13where instead of just having a binary classification, 1 or 0,
- 9:29:17we can have as many different categories as we want.
- 9:29:20And we can have our neural network output these probabilities
- 9:29:23over which categories are more likely than other categories.
- 9:29:27And using that data, we're able to draw some sort of inference
- 9:29:30on what it is that we should do.
- 9:29:33So this was sort of the idea of supervised machine learning.
- 9:29:35I can give this neural network a whole bunch of data,
- 9:29:38a whole bunch of input data corresponding to some label, some output data,
- 9:29:42like we know that it was raining on this day,
- 9:29:45we know that it was sunny on that day.
- 9:29:46And using all of that data, the algorithm
- 9:29:49can use gradient descent to figure out what all of the weights
- 9:29:52should be in order to create some sort of model that hopefully allows us
- 9:29:55a way to predict what we think the weather is going to be.
- 9:29:59But neural networks have a lot of other applications as well.
- 9:30:02You could imagine applying the same sort of idea to a reinforcement learning
- 9:30:06sort of example as well, where you remember that in reinforcement
- 9:30:09learning, what we wanted to do is train some sort of agent
- 9:30:13to learn what action to take, depending on what state
- 9:30:16they currently happen to be in.
- 9:30:17So depending on the current state of the world,
- 9:30:19we wanted the agent to pick from one of the available actions
- 9:30:23that is available to them.
- 9:30:24And you might model that by having each of these input variables
- 9:30:28represent some information about the state, some data about what state
- 9:30:33our agent is currently in.
- 9:30:34And then the output, for example, could be each
- 9:30:37of the various different actions that our agent could take,
- 9:30:40action 1, 2, 3, and 4.
- 9:30:42And you might imagine that this network would work in the same way,
- 9:30:45but based on these particular inputs, we go ahead and calculate values
- 9:30:48for each of these outputs.
- 9:30:50And those outputs could model which action is better than other actions.
- 9:30:53And we could just choose, based on looking at those outputs,
- 9:30:56which action we should take.
- 9:30:59And so these neural networks are very broadly applicable,
- 9:31:01that all they're really doing is modeling some mathematical function.
- 9:31:05So anything that we can frame as a mathematical function,
- 9:31:07something like classifying inputs into various different categories
- 9:31:11or figuring out based on some input state what action we should take,
- 9:31:15these are all mathematical functions that we could attempt to model
- 9:31:18by taking advantage of this neural network structure,
- 9:31:21and in particular, taking advantage of this technique, gradient descent,
- 9:31:25that we can use in order to figure out what the weights should
- 9:31:27be in order to do this sort of calculation.
- 9:31:31Now, how is it that you would go about training a neural network that
- 9:31:33has multiple outputs instead of just one?
- 9:31:36Well, with just a single output, we could see what the output for that value
- 9:31:40should be, and then you update all of the weights that corresponded to it.
- 9:31:44And when we have multiple outputs, at least in this particular case,
- 9:31:47we can really think of this as four separate neural networks,
- 9:31:51that really we just have one network here that has these three inputs
- 9:31:55corresponding with these three weights corresponding to this one output value.
- 9:32:00And the same thing is true for this output value.
- 9:32:02This output value effectively defines yet another neural network
- 9:32:06that has these same three inputs, but a different set of weights
- 9:32:09that correspond to this output.
- 9:32:11And likewise, this output has its own set of weights as well,
- 9:32:14and same thing for the fourth output too.
- 9:32:17And so if you wanted to train a neural network that had four outputs instead
- 9:32:20of just one, in this case where the inputs are directly
- 9:32:23connected to the outputs, you could really
- 9:32:25think of this as just training four independent neural networks.
- 9:32:28We know what the outputs for each of these four
- 9:32:31should be based on our input data, and using that data,
- 9:32:34we can begin to figure out what all of these individual weights should be.
- 9:32:37And maybe there's an additional step at the end
- 9:32:39to make sure that we turn these values into a probability distribution such
- 9:32:43that we can interpret which one is better than another
- 9:32:46or more likely than another as a category or something like that.
- 9:32:50So this then seems like it does a pretty good job of taking inputs
- 9:32:53and trying to predict what outputs should be.
- 9:32:55And we'll see some real examples of this in just a moment as well.
- 9:32:58But it's important then to think about what the limitations
- 9:33:01of this sort of approach is, of just taking some linear combination
- 9:33:05of inputs and passing it into some sort of activation function.
- 9:33:09And it turns out that when we do this in the case of binary classification,
- 9:33:12trying to predict does it belong to one category or another,
- 9:33:16we can only predict things that are linearly separable.
- 9:33:20Because we're taking a linear combination of inputs
- 9:33:22and using that to define some decision boundary or threshold,
- 9:33:26then what we get is a situation where if we have this set of data,
- 9:33:29we can predict a line that separates linearly the red points from the blue
- 9:33:35points, but a single unit that is making a binary classification, otherwise
- 9:33:39known as a perceptron, can't deal with a situation like this, where we've
- 9:33:44seen this type of situation before, where there is no straight line that
- 9:33:48just goes straight through the data that will divide the red points away
- 9:33:51from the blue points.
- 9:33:52It's a more complex decision boundary.
- 9:33:55The decision boundary somehow needs to capture the things inside of this
- 9:33:58circle.
- 9:33:59And there isn't really a line that will allow us to deal with that.
- 9:34:03So this is the limitation of the perceptron,
- 9:34:05these units that just make these binary decisions based on their inputs,
- 9:34:08that a single perceptron is only capable of learning
- 9:34:12a linearly separable decision boundary.
- 9:34:15All it can do is define a line.
- 9:34:17And sure, it can give us probabilities based
- 9:34:19on how close to that decision boundary we are,
- 9:34:21but it can only really decide based on a linear decision boundary.
- 9:34:26And so this doesn't seem like it's going to generalize well
- 9:34:29to situations where real world data is involved,
- 9:34:32because real world data often isn't linearly separable.
- 9:34:34It often isn't the case that we can just draw a line through the data
- 9:34:38and be able to divide it up into multiple groups.
- 9:34:41So what then is the solution to this?
- 9:34:43Well, what was proposed was the idea of a multilayer neural network,
- 9:34:47that so far all of the neural networks we've seen
- 9:34:49have had a set of inputs and a set of outputs,
- 9:34:52and the inputs are connected to those outputs.
- 9:34:55But in a multilayer neural network, this is going
- 9:34:57to be an artificial neural network that has an input layer still.
- 9:35:00It has an output layer, but also has one or more hidden layers in between.
- 9:35:06Other layers of artificial neurons or units
- 9:35:09that are going to calculate their own values as well.
- 9:35:12So instead of a neural network that looks like this with three inputs
- 9:35:15and one output, you might imagine in the middle
- 9:35:17here injecting a hidden layer, something like this.
- 9:35:21This is a hidden layer that has four nodes.
- 9:35:23You could choose how many nodes or units end up going into the hidden layer.
- 9:35:26You can have multiple hidden layers as well.
- 9:35:29And so now each of these inputs isn't directly connected to the output.
- 9:35:33Each of the inputs is connected to this hidden layer.
- 9:35:36And then all of the nodes in the hidden layer, those
- 9:35:38are connected to the one output.
- 9:35:41And so this is just another step that we can
- 9:35:43take towards calculating more complex functions.
- 9:35:46Each of these hidden units will calculate its output value,
- 9:35:49otherwise known as its activation, based on a linear combination
- 9:35:53of all the inputs.
- 9:35:55And once we have values for all of these nodes,
- 9:35:57as opposed to this just being the output, we do the same thing again.
- 9:36:00Calculate the output for this node based on multiplying
- 9:36:04each of the values for these units by their weights as well.
- 9:36:07So in effect, the way this works is that we start with inputs.
- 9:36:10They get multiplied by weights in order to calculate values for the hidden nodes.
- 9:36:14Those get multiplied by weights in order to figure out
- 9:36:16what the ultimate output is going to be.
- 9:36:19And the advantage of layering things like this
- 9:36:22is it gives us an ability to model more complex functions,
- 9:36:25that instead of just having a single decision boundary, a single line
- 9:36:29dividing the red points from the blue points, each of these hidden nodes
- 9:36:33can learn a different decision boundary.
- 9:36:35And we can combine those decision boundaries
- 9:36:37to figure out what the ultimate output is going to be.
- 9:36:41And as we begin to imagine more complex situations,
- 9:36:43you could imagine each of these nodes learning some useful property
- 9:36:47or learning some useful feature of all of the inputs
- 9:36:50and us somehow learning how to combine those features together
- 9:36:53in order to get the output that we actually want.
- 9:36:56Now, the natural question when we begin to look at this now
- 9:36:59is to ask the question of, how do we train a neural network that
- 9:37:02has hidden layers inside of it?
- 9:37:04And this turns out to initially be a bit of a tricky question,
- 9:37:07because the input data that we are given is we
- 9:37:10are given values for all of the inputs, and we're
- 9:37:13given what the value of the output should be, what the category is,
- 9:37:16for example.
- 9:37:18But the input data doesn't tell us what the values for all of these nodes
- 9:37:22should be.
- 9:37:22So we don't know how far off each of these nodes actually
- 9:37:26is because we're only given data for the inputs and the outputs.
- 9:37:29The reason this is called the hidden layer
- 9:37:31is because the data that is made available to us
- 9:37:34doesn't tell us what the values for all of these intermediate nodes
- 9:37:38should actually be.
- 9:37:39And so the strategy people came up with was
- 9:37:42to say that if you know what the error or the losses on the output node,
- 9:37:48well, then based on what these weights are,
- 9:37:50if one of these weights is higher than another,
- 9:37:52you can calculate an estimate for how much
- 9:37:55the error from this node was due to this part of the hidden node,
- 9:38:00or this part of the hidden layer, or this part of the hidden layer,
- 9:38:03based on the values of these weights, in effect saying
- 9:38:05that based on the error from the output, I can back propagate the error
- 9:38:10and figure out an estimate for what the error is for each of these nodes
- 9:38:14in the hidden layer as well.
- 9:38:15And there's some more calculus here that we won't get into the details of,
- 9:38:18but the idea of this algorithm is known as back propagation.
- 9:38:21It's an algorithm for training a neural network
- 9:38:24with multiple different hidden layers.
- 9:38:26And the idea for this, the pseudocode for it,
- 9:38:28will again be if we want to run gradient descent with back propagation.
- 9:38:31We'll start with a random choice of weights, as we did before.
- 9:38:35And now we'll go ahead and repeat the training process again and again.
- 9:38:38But what we're going to do each time is now
- 9:38:41we're going to calculate the error for the output layer first.
- 9:38:43We know the output and what it should be,
- 9:38:45and we know what we calculated so we can figure out what the error there is.
- 9:38:49But then we're going to repeat for every layer,
- 9:38:52starting with the output layer, moving back into the hidden layer,
- 9:38:55then the hidden layer before that if there are multiple hidden layers,
- 9:38:58going back all the way to the very first hidden layer,
- 9:39:00assuming there are multiple, we're going to propagate the error back one layer.
- 9:39:05Whatever the error was from the output, figure out
- 9:39:07what the error should be a layer before that
- 9:39:09based on what the values of those weights are.
- 9:39:11And then we can update those weights.
- 9:39:14So graphically, the way you might think about this
- 9:39:17is that we first start with the output.
- 9:39:18We know what the output should be.
- 9:39:20We know what output we calculated.
- 9:39:22And based on that, we can figure out, all right,
- 9:39:23how do we need to update those weights?
- 9:39:25Backpropagating the error to these nodes.
- 9:39:28And using that, we can figure out how we should update these weights.
- 9:39:31And you might imagine if there are multiple layers,
- 9:39:33we could repeat this process again and again
- 9:39:35to begin to figure out how all of these weights should be updated.
- 9:39:39And this backpropagation algorithm is really
- 9:39:41the key algorithm that makes neural networks possible.
- 9:39:44It makes it possible to take these multi-level structures
- 9:39:47and be able to train those structures depending
- 9:39:50on what the values of these weights are in order
- 9:39:52to figure out how it is that we should go about updating those weights in
- 9:39:56order to create some function that is able to minimize
- 9:39:59the total amount of loss, to figure out some good setting of the weights
- 9:40:02that will take the inputs and translate it into the output that we expect.
- 9:40:07And this works, as we said, not just for a single hidden layer.
- 9:40:10But you can imagine multiple hidden layers, where each hidden layer we just
- 9:40:13define however many nodes we want, where each of the nodes in one layer,
- 9:40:17we can connect to the nodes in the next layer,
- 9:40:19defining more and more complex networks that
- 9:40:22are able to model more and more complex types of functions.
- 9:40:26And so this type of network is what we might call a deep neural network,
- 9:40:30part of a larger family of deep learning algorithms,
- 9:40:33if you've ever heard that term.
- 9:40:34And all deep learning is about is it's using multiple layers
- 9:40:38to be able to predict and be able to model higher level
- 9:40:41features inside of the input, to be able to figure out
- 9:40:44what the output should be.
- 9:40:45And so a deep neural network is just a neural network
- 9:40:47that has multiple of these hidden layers,
- 9:40:49where we start at the input, calculate values for this layer,
- 9:40:52then this layer, then this layer, and then ultimately get an output.
- 9:40:55And this allows us to be able to model more and more sophisticated types
- 9:40:59of functions, that each of these layers can calculate something
- 9:41:02a little bit different, and we can combine that information
- 9:41:05to figure out what the output should be.
- 9:41:08Of course, as with any situation of machine learning,
- 9:41:11as we begin to make our models more and more complex,
- 9:41:13to model more and more complex functions, the risk we run
- 9:41:17is something like overfitting.
- 9:41:18And we talked about overfitting last time in the context of overfitting
- 9:41:22based on when we were training our models to be
- 9:41:25able to learn some sort of decision boundary,
- 9:41:27where overfitting happens when we fit too closely to the training data.
- 9:41:31And as a result, we don't generalize well to other situations as well.
- 9:41:36And one of the risks we run with a far more complex neural network that
- 9:41:40has many, many different nodes is that we might overfit based on the input
- 9:41:44data.
- 9:41:44We might grow over reliant on certain nodes
- 9:41:46to calculate things just purely based on the input data that
- 9:41:49doesn't allow us to generalize very well to the output.
- 9:41:53And there are a number of strategies for dealing with overfitting.
- 9:41:56But one of the most popular in the context of neural networks
- 9:41:59is a technique known as dropout.
- 9:42:01And what dropout does is it, when we're training the neural network,
- 9:42:04what we'll do in dropout is temporarily remove units,
- 9:42:08temporarily remove these artificial neurons from our network chosen at
- 9:42:11random.
- 9:42:12And the goal here is to prevent over-reliance on certain units.
- 9:42:16What generally happens in overfitting is that we
- 9:42:18begin to over-rely on certain units inside the neural network
- 9:42:21to be able to tell us how to interpret the input data.
- 9:42:24What dropout will do is randomly remove some of these units
- 9:42:28in order to reduce the chance that we over-rely on certain units
- 9:42:31to make our neural network more robust, to be able to handle the situations
- 9:42:35even when we just drop out particular neurons entirely.
- 9:42:39So the way that might work is we have a network like this.
- 9:42:42And as we're training it, when we go about trying
- 9:42:44to update the weights the first time, we'll just randomly pick
- 9:42:47some percentage of the nodes to drop out of the network.
- 9:42:49It's as if those nodes aren't there at all.
- 9:42:51It's as if the weights associated with those nodes aren't there at all.
- 9:42:54And we'll train it this way.
- 9:42:56Then the next time we update the weights, we'll pick a different set
- 9:42:58and just go ahead and train that way.
- 9:42:59And then again, randomly choose and train with other nodes
- 9:43:02that have been dropped out as well.
- 9:43:04And the goal of that is that after the training process,
- 9:43:07if you train by dropping out random nodes inside of this neural network,
- 9:43:10you hopefully end up with a network that's a little bit more robust,
- 9:43:13that doesn't rely too heavily on any one particular node,
- 9:43:16but more generally learns how to approximate a function in general.
- 9:43:21So that then is a look at some of these techniques
- 9:43:24that we can use in order to implement a neural network,
- 9:43:27to get at the idea of taking this input, passing it
- 9:43:30through these various different layers in order to produce some sort of output.
- 9:43:34And what we'd like to do now is take those ideas and put them into code.
- 9:43:37And to do that, there are a number of different machine learning libraries,
- 9:43:40neural network libraries that we can use that allow us to get access
- 9:43:44to someone's implementation of back propagation and all of these hidden
- 9:43:47layers.
- 9:43:48And one of the most popular, developed by Google, is known as TensorFlow,
- 9:43:52a library that we can use for quickly creating neural networks and modeling
- 9:43:55them and running them on some sample data to see what the output is going
- 9:43:59to be.
- 9:44:00And before we actually start writing code,
- 9:44:01we'll go ahead and take a look at TensorFlow's playground, which
- 9:44:04will be an opportunity for us just to play around with this idea of neural
- 9:44:08networks in different layers, just to get a sense for what
- 9:44:10it is that we can do by taking advantage of neural networks.
- 9:44:15So let's go ahead and go into TensorFlow's playground, which
- 9:44:18you can go to by visiting that URL from before.
- 9:44:20And what we're going to do now is we're going to try and learn the decision
- 9:44:24boundary for this particular output.
- 9:44:27I want to learn to separate the orange points from the blue points.
- 9:44:30And I'd like to learn some sort of setting of weights inside of a neural
- 9:44:34network that will be able to separate those from each other.
- 9:44:37The features we have access to, our input data,
- 9:44:40are the x value and the y value, so the two values along each of the two axes.
- 9:44:44And what I'll do now is I can set particular parameters,
- 9:44:47like what activation function I would like to use.
- 9:44:50And I'll just go ahead and press play and see what happens.
- 9:44:53And what happens here is that you'll see that just
- 9:44:56by using these two input features, the x value and the y value,
- 9:45:00with no hidden layers, just take the input, x and y values,
- 9:45:04and figure out what the decision boundary is.
- 9:45:06Our neural network learns pretty quickly that in order
- 9:45:08to divide these two points, we should just use this line.
- 9:45:11This line acts as a decision boundary that
- 9:45:13separates this group of points from that group of points,
- 9:45:16and it does it very well.
- 9:45:17You can see up here what the loss is.
- 9:45:19The training loss is 0, meaning we were able to perfectly model separating
- 9:45:24these two points from each other inside of our training data.
- 9:45:27So this was a fairly simple case of trying
- 9:45:30to apply a neural network because the data is very clean.
- 9:45:33It's very nicely linearly separable.
- 9:45:35We could just draw a line that separates all of those points from each other.
- 9:45:39Let's now consider a more complex case.
- 9:45:42So I'll go ahead and pause the simulation,
- 9:45:44and we'll go ahead and look at this data set here.
- 9:45:47This data set is a little bit more complex now.
- 9:45:50In this data set, we still have blue and orange points
- 9:45:52that we'd like to separate from each other.
- 9:45:54But there's no single line that we can draw
- 9:45:56that is going to be able to figure out how to separate the blue from the orange,
- 9:45:59because the blue is located in these two quadrants,
- 9:46:02and the orange is located here and here.
- 9:46:04It's a more complex function to be able to learn.
- 9:46:07So let's see what happens.
- 9:46:09If we just try and predict based on those inputs, the x and y coordinates,
- 9:46:13what the output should be, I'll press Play.
- 9:46:16And what you'll notice is that we're not really
- 9:46:18able to draw much of a conclusion, that we're not
- 9:46:21able to very cleanly see how we should divide the orange points from the blue
- 9:46:25points, and you don't see a very clean separation there.
- 9:46:30So it seems like we don't have enough sophistication inside of our network
- 9:46:34to be able to model something that is that complex.
- 9:46:37We need a better model for this neural network.
- 9:46:39And I'll do that by adding a hidden layer.
- 9:46:42So now I have a hidden layer that has two neurons inside of it.
- 9:46:45So I have two inputs that then go to two neurons
- 9:46:49inside of a hidden layer that then go to our output.
- 9:46:52And now I'll press Play.
- 9:46:54And what you'll notice here is that we're able to do slightly better.
- 9:46:57We're able to now say, all right, these points are definitely blue.
- 9:47:00These points are definitely orange.
- 9:47:02We're still struggling a little bit with these points up here, though.
- 9:47:05And what we can do is we can see for each of these hidden neurons,
- 9:47:08what is it exactly that these hidden neurons are doing?
- 9:47:11Each hidden neuron is learning its own decision boundary.
- 9:47:15And we can see what that boundary is.
- 9:47:16This first neuron is learning, all right,
- 9:47:19this line that seems to separate some of the blue points
- 9:47:22from the rest of the points.
- 9:47:24This other hidden neuron is learning another line
- 9:47:27that seems to be separating the orange points in the lower right
- 9:47:29from the rest of the points.
- 9:47:31So that's why we're able to figure out these two areas in the bottom region.
- 9:47:36But we're still not able to perfectly classify all of the points.
- 9:47:40So let's go ahead and add another neuron.
- 9:47:42Now we've got three neurons inside of our hidden layer
- 9:47:46and see what we're able to learn now.
- 9:47:48All right, well, now we seem to be doing a better job.
- 9:47:50By learning three different decision boundaries, which
- 9:47:53each of the three neurons inside of our hidden layer,
- 9:47:55we're able to much better figure out how to separate these blue points
- 9:47:59from the orange points.
- 9:48:00And we can see what each of these hidden neurons is learning.
- 9:48:03Each one is learning a slightly different decision boundary.
- 9:48:06And then we're combining those decision boundaries together
- 9:48:09to figure out what the overall output should be.
- 9:48:11And then we can try it one more time by adding a fourth neuron there
- 9:48:15and try learning that.
- 9:48:17And it seems like now we can do even better at trying
- 9:48:19to separate the blue points from the orange points.
- 9:48:21But we were only able to do this by adding a hidden layer,
- 9:48:24by adding some layer that is learning some other boundaries
- 9:48:27and combining those boundaries to determine the output.
- 9:48:30And the strength, the size and thickness of these lines
- 9:48:33indicate how high these weights are, how important each of these inputs
- 9:48:37is for making this sort of calculation.
- 9:48:40And we can do maybe one more simulation.
- 9:48:42Let's go ahead and try this on a data set that looks like this.
- 9:48:46Go ahead and get rid of the hidden layer.
- 9:48:47Here now we're trying to separate the blue points from the orange points
- 9:48:51where all the blue points are located, again,
- 9:48:53inside of a circle effectively.
- 9:48:54So we're not going to be able to learn a line.
- 9:48:57Notice I press Play.
- 9:48:58And we're really not able to draw any sort of classification at all
- 9:49:01because there is no line that cleanly separates the blue points
- 9:49:04from the orange points.
- 9:49:06So let's try to solve this by introducing a hidden layer.
- 9:49:10I'll go ahead and press Play.
- 9:49:12And all right, with two neurons in a hidden layer,
- 9:49:14we're able to do a little better because we effectively
- 9:49:17learned two different decision boundaries.
- 9:49:18We learned this line here.
- 9:49:20And we learned this line on the right-hand side.
- 9:49:23And right now we're just saying, all right, well, if it's in between,
- 9:49:25we'll call it blue.
- 9:49:25And if it's outside, we'll call it orange.
- 9:49:27So not great, but certainly better than before,
- 9:49:30that we're learning one decision boundary and another.
- 9:49:33And based on those, we can figure out what the output should be.
- 9:49:36But let's now go ahead and add a third neuron and see what happens now.
- 9:49:42I go ahead and train it.
- 9:49:43And now, using three different decision boundaries
- 9:49:46that are learned by each of these hidden neurons,
- 9:49:48we're able to much more accurately model this distinction
- 9:49:51between blue points and orange points.
- 9:49:53We're able to figure out maybe with these three decision boundaries,
- 9:49:56combining them together, you can imagine figuring out
- 9:49:58what the output should be and how to make that sort of classification.
- 9:50:02And so the goal here is just to get a sense for having more neurons
- 9:50:05in these hidden layers allows us to learn more structure in the data,
- 9:50:09allows us to figure out what the relevant and important decision
- 9:50:12boundaries are.
- 9:50:13And then using this backpropagation algorithm,
- 9:50:15we're able to figure out what the values of these weights should be
- 9:50:18in order to train this network to be able to classify one category of points
- 9:50:23away from another category of points instead.
- 9:50:26And this is ultimately what we're going to be trying
- 9:50:28to do whenever we're training a neural network.
- 9:50:32So let's go ahead and actually see an example of this.
- 9:50:34You'll recall from last time that we had this banknotes file
- 9:50:38that included information about counterfeit banknotes as opposed
- 9:50:41to authentic banknotes, where I had four different values for each banknote
- 9:50:45and then a categorization of whether that banknote is considered
- 9:50:48to be authentic or a counterfeit note.
- 9:50:51And what I wanted to do was, based on that input information,
- 9:50:55figure out some function that could calculate
- 9:50:57based on the input information what category it belonged to.
- 9:51:00And what I've written here in banknotes.py
- 9:51:02is a neural network that will learn just that, a network that
- 9:51:05learns based on all of the input whether or not
- 9:51:08we should categorize a banknote as authentic or as counterfeit.
- 9:51:13The first step is the same as what we saw from last time.
- 9:51:15I'm really just reading the data in and getting it
- 9:51:17into an appropriate format.
- 9:51:19And so this is where more of the writing Python code on your own
- 9:51:22comes in, in terms of manipulating this data,
- 9:51:25massaging the data into a format that will be understood
- 9:51:28by a machine learning library like scikit-learn or like TensorFlow.
- 9:51:32And so here I separate it into a training and a testing set.
- 9:51:35And now what I'm doing down below is I'm creating a neural network.
- 9:51:40Here I'm using TF, which stands for TensorFlow.
- 9:51:42Up above, I said import TensorFlow as TF, TF just an abbreviation that we'll
- 9:51:47often use so we don't need to write out TensorFlow
- 9:51:49every time we want to use anything inside of the library.
- 9:51:52I'm using TF.keras.
- 9:51:55Keras is an API, a set of functions that we
- 9:51:57can use in order to manipulate neural networks inside of TensorFlow.
- 9:52:02And it turns out there are other machine learning libraries
- 9:52:04that also use the Keras API.
- 9:52:06But here I'm saying, all right, go ahead and give me
- 9:52:08a model that is a sequential model, a sequential neural network,
- 9:52:12meaning one layer after another.
- 9:52:14And now I'm going to add to that model what layers
- 9:52:17I want inside of my neural network.
- 9:52:20So here I'm saying model.add.
- 9:52:22Go ahead and add a dense layer.
- 9:52:24And when we say a dense layer, we mean a layer that is just each
- 9:52:28of the nodes inside of the layer is going to be connected
- 9:52:30to each of the nodes from the previous layer.
- 9:52:32So we have a densely connected layer.
- 9:52:35This layer is going to have eight units inside of it.
- 9:52:38So it's going to be a hidden layer inside of a neural network
- 9:52:40with eight different units, eight artificial neurons, each of which
- 9:52:43might learn something different.
- 9:52:45And I just sort of chose eight arbitrarily.
- 9:52:47You could choose a different number of hidden nodes inside of the layer.
- 9:52:50And as we saw before, depending on the number of units
- 9:52:53there are inside of your hidden layer, more units
- 9:52:56means you can learn more complex functions.
- 9:52:58So maybe you can more accurately model the training data.
- 9:53:01But it comes at the cost.
- 9:53:02More units means more weights that you need to figure out how to update.
- 9:53:05So it might be more expensive to do that calculation.
- 9:53:08And you also run the risk of overfitting on the data.
- 9:53:10If you have too many units and you learn to just
- 9:53:13overfit on the training data, that's not good either.
- 9:53:15So there is a balance.
- 9:53:16And there's often a testing process where you'll train on some data
- 9:53:20and maybe validate how well you're doing on a separate set of data,
- 9:53:23often called a validation set, to see, all right, which setting of parameters.
- 9:53:26How many layers should I have?
- 9:53:28How many units should be in each layer?
- 9:53:29Which one of those performs the best on the validation set?
- 9:53:32So you can do some testing to figure out what these hyper parameters, so called,
- 9:53:36should be equal to.
- 9:53:38Next, I specify what the input shape is.
- 9:53:41Meaning, all right, what does my input look like?
- 9:53:43My input has four values.
- 9:53:44And so the input shape is just four, because we have four inputs.
- 9:53:48And then I specify what the activation function is.
- 9:53:51And the activation function, again, we can choose.
- 9:53:53There are a number of different activation functions.
- 9:53:55Here I'm using relu, which you might recall from earlier.
- 9:53:59And then I'll add an output layer.
- 9:54:01So I have my hidden layer.
- 9:54:02Now I'm adding one more layer that will just have one unit,
- 9:54:05because all I want to do is predict something
- 9:54:07like counterfeit build or authentic build.
- 9:54:10So I just need a single unit.
- 9:54:12And the activation function I'm going to use here
- 9:54:14is that sigmoid activation function, which, again,
- 9:54:16was that S-shaped curve that just gave us a probability of what
- 9:54:20is the probability that this is a counterfeit build,
- 9:54:24as opposed to an authentic build.
- 9:54:26So that, then, is the structure of my neural network,
- 9:54:29a sequential neural network that has one hidden layer with eight units inside
- 9:54:32of it, and then one output layer that just has a single unit inside of it.
- 9:54:37And I can choose how many units there are.
- 9:54:38I can choose the activation function.
- 9:54:40Then I'm going to compile this model.
- 9:54:44TensorFlow gives you a choice of how you would like to optimize the weights.
- 9:54:48There are various different algorithms for doing that.
- 9:54:50What type of loss function you want to use.
- 9:54:52Again, many different options for doing that.
- 9:54:54And then how I want to evaluate my model, well, I care about accuracy.
- 9:54:57I care about how many of my points am I able to classify correctly
- 9:55:01versus not correctly as counterfeit or not counterfeit.
- 9:55:04And I would like it to report to me how accurate my model is performing.
- 9:55:09Then, now that I've defined that model, I
- 9:55:12call model.fit to say go ahead and train the model.
- 9:55:15Train it on all the training data plus all of the training labels.
- 9:55:19So labels for each of those pieces of training data.
- 9:55:22And I'm saying run it for 20 epics, meaning go ahead and go
- 9:55:25through each of these training points 20 times, effectively.
- 9:55:28Go through the data 20 times and keep trying to update the weights.
- 9:55:31If I did it for more, I could train for even longer
- 9:55:33and maybe get a more accurate result.
- 9:55:36But then after I fit it on all the data, I'll go ahead and just test it.
- 9:55:39I'll evaluate my model using model.evaluate built into TensorFlow
- 9:55:43that is just going to tell me how well do I perform on the testing data.
- 9:55:47So ultimately, this is just going to give me some numbers that tell me
- 9:55:50how well we did in this particular case.
- 9:55:54So now what I'm going to do is go into banknotes and go ahead and run
- 9:55:57banknotes.py.
- 9:55:59And what's going to happen now is it's going to read in all of that training
- 9:56:02data.
- 9:56:02It's going to generate a neural network with all my inputs,
- 9:56:05my eight hidden units inside my layer, and then an output unit.
- 9:56:10And now what it's doing is it's training.
- 9:56:11It's training 20 times.
- 9:56:13And each time you can see how my accuracy is increasing on my training data.
- 9:56:17It starts off the very first time not very accurate,
- 9:56:20though better than random, something like 79% of the time.
- 9:56:23It's able to accurately classify one bill from another.
- 9:56:26But as I keep training, notice this accuracy value
- 9:56:29improves and improves and improves until after I've trained through all
- 9:56:33the data points 20 times, it looks like my accuracy is above 99% on the training
- 9:56:39data.
- 9:56:40And here's where I tested it on a whole bunch of testing data.
- 9:56:43And it looks like in this case, I was also like 99.8% accurate.
- 9:56:48So just using that, I was able to generate a neural network that
- 9:56:51can detect counterfeit bills from authentic bills based on this input
- 9:56:54data 99.8% of the time, at least based on this particular testing data.
- 9:56:59And I might want to test it with more data as well,
- 9:57:01just to be confident about that.
- 9:57:03But this is really the value of using a machine learning library like TensorFlow.
- 9:57:06And there are others available for Python and other languages as well.
- 9:57:10But all I have to do is define the structure of the network
- 9:57:13and define the data that I'm going to pass into the network.
- 9:57:16And then TensorFlow runs the backpropagation algorithm
- 9:57:19for learning what all of those weights should be,
- 9:57:22for figuring out how to train this neural network to be
- 9:57:24able to accurately, as accurately as possible,
- 9:57:27figure out what the output values should be there as well.
- 9:57:31And so this then was a look at what it is that neural networks can do just
- 9:57:36using these sequences of layer after layer after layer.
- 9:57:39And you can begin to imagine applying these to much more general problems.
- 9:57:43And one big problem in computing and artificial intelligence
- 9:57:45more generally is the problem of computer vision.
- 9:57:49Computer vision is all about computational methods
- 9:57:51for analyzing and understanding images.
- 9:57:54You might have pictures that you want the computer to figure out
- 9:57:57how to deal with, how to process those images
- 9:57:59and figure out how to produce some sort of useful result out of this.
- 9:58:02You've seen this in the context of social media websites
- 9:58:05that are able to look at a photo that contains a whole bunch of faces.
- 9:58:08And it's able to figure out what's a picture of whom
- 9:58:10and label those and tag them with appropriate people.
- 9:58:13This is becoming increasingly relevant as we
- 9:58:15begin to discuss self-driving cars, that these cars now have cameras.
- 9:58:19And we would like for the computer to have some sort of algorithm
- 9:58:22that looks at the image and figures out what color is the light, what cars
- 9:58:26are around us and in what direction, for example.
- 9:58:29And so computer vision is all about taking an image and figuring out
- 9:58:33what sort of computation, what sort of calculation
- 9:58:35we can do with that image.
- 9:58:36It's also relevant in the context of something like handwriting recognition.
- 9:58:40This, what you're looking at, is an example of the MNIST data set.
- 9:58:43It's a big data set just of handwritten digits
- 9:58:46that we could use to ideally try and figure out
- 9:58:48how to predict, given someone's handwriting, given a photo of a digit
- 9:58:52that they have drawn, can you predict whether it's a 0, 1, 2, 3, 4, 5, 6, 7, 8,
- 9:58:57or 9, for example.
- 9:58:58So this sort of handwriting recognition is yet another task
- 9:59:01that we might want to use computer vision tasks and tools
- 9:59:04to be able to apply it towards.
- 9:59:05This might be a task that we might care about.
- 9:59:08So how, then, can we use neural networks to be
- 9:59:11able to solve a problem like this?
- 9:59:13Well, neural networks rely upon some sort of input
- 9:59:15where that input is just numerical data.
- 9:59:17We have a whole bunch of units where each one of them
- 9:59:19just represents some sort of number.
- 9:59:22And so in the context of something like handwriting recognition
- 9:59:24or in the context of just an image, you might imagine that an image is really
- 9:59:29just a grid of pixels, grid of dots where each dot has some sort of color.
- 9:59:34And in the context of something like handwriting recognition,
- 9:59:36you might imagine that if you just fill in each of these dots in a particular
- 9:59:39way, you can generate a 2 or an 8, for example,
- 9:59:42based on which dots happen to be shaded in and which dots are not.
- 9:59:46And we can represent each of these pixel values just using numbers.
- 9:59:50So for a particular pixel, for example, 0 might represent entirely black.
- 9:59:55Depending on how you're representing color,
- 9:59:57it's often common to represent color values on a 0 to 255 range
- 10:00:02so that you can represent a color using 8 bits for a particular value,
- 10:00:06like how much white is in the image.
- 10:00:08So 0 might represent all black.
- 10:00:10255 might represent entirely white as a pixel.
- 10:00:14And somewhere in between might represent some shade of gray, for example.
- 10:00:18But you might imagine not just having a single slider that
- 10:00:20determines how much white is in the image,
- 10:00:22but if you had a color image, you might imagine
- 10:00:24three different numerical values, a red, green, and blue value,
- 10:00:28where the red value controls how much red is in the image.
- 10:00:30We have one value for controlling how much green is in the pixel
- 10:00:33and one value for how much blue is in the pixel as well.
- 10:00:36And depending on how it is that you set these values of red, green, and blue,
- 10:00:40you can get a different color.
- 10:00:42And so any pixel can really be represented, in this case,
- 10:00:45by three numerical values, a red value, a green value, and a blue value.
- 10:00:50And if you take a whole bunch of these pixels, assemble them together
- 10:00:54inside of a grid of pixels, then you really
- 10:00:56just have a whole bunch of numerical values
- 10:00:59that you can use in order to perform some sort of prediction task.
- 10:01:03And so what you might imagine doing is using the same techniques
- 10:01:05we talked about before, just design a neural network
- 10:01:08with a lot of inputs, that for each of the pixels,
- 10:01:12we might have one or three different inputs
- 10:01:13in the case of a color image, a different input that
- 10:01:16is just connected to a deep neural network, for example.
- 10:01:20And this deep neural network might take all of the pixels
- 10:01:22inside of the image of what digit a person drew.
- 10:01:27And the output might be like 10 neurons that
- 10:01:29classify it as a 0, or a 1, or a 2, or a 3,
- 10:01:32or just tells us in some way what that digit happens to be.
- 10:01:36Now, there are a couple of drawbacks to this approach.
- 10:01:39The first drawback to the approach is just the size of this input array,
- 10:01:42that we have a whole bunch of inputs.
- 10:01:44If we have a big image that has a lot of different channels,
- 10:01:47we're looking at a lot of inputs, and therefore a lot of weights
- 10:01:50that we have to calculate.
- 10:01:51And a second problem is the fact that by flattening everything
- 10:01:55into just this structure of all the pixels,
- 10:01:58we've lost access to a lot of the information
- 10:02:00about the structure of the image that's relevant,
- 10:02:03that really, when a person looks at an image,
- 10:02:05they're looking at particular features of the image.
- 10:02:08They're looking at curves.
- 10:02:09They're looking at shapes.
- 10:02:09They're looking at what things can you identify
- 10:02:11in different regions of the image, and maybe put those things together
- 10:02:14in order to get a better picture of what the overall image is about.
- 10:02:18And by just turning it into pixel values for each of the pixels,
- 10:02:22sure, you might be able to learn that structure,
- 10:02:24but it might be challenging in order to do so.
- 10:02:26It might be helpful to take advantage of the fact
- 10:02:28that you can use properties of the image itself, the fact
- 10:02:31that it's structured in a particular way, to be
- 10:02:33able to improve the way that we learn based on that image too.
- 10:02:37So in order to figure out how we can train our neural networks to better
- 10:02:40be able to deal with images, we'll introduce a couple of ideas,
- 10:02:43a couple of algorithms that we can apply that
- 10:02:45allow us to take the image and extract some useful information out
- 10:02:50of that image.
- 10:02:50And the first idea we'll introduce is the notion of image convolution.
- 10:02:54And what image convolution is all about is it's about filtering an image,
- 10:02:58sort of extracting useful or relevant features out of the image.
- 10:03:01And the way we do that is by applying a particular filter that
- 10:03:05basically adds the value for every pixel with the values
- 10:03:09for all of the neighboring pixels to it, according
- 10:03:11to some sort of kernel matrix, which we'll see in a moment,
- 10:03:14is going to allow us to weight these pixels in various different ways.
- 10:03:17And the goal of image convolution, then, is
- 10:03:19to extract some sort of interesting or useful features out of an image,
- 10:03:22to be able to take a pixel and, based on its neighboring pixels,
- 10:03:26maybe predict some sort of valuable information.
- 10:03:29Something like taking a pixel and looking at its neighboring pixels,
- 10:03:32you might be able to predict whether or not
- 10:03:33there's some sort of curve inside the image,
- 10:03:35or whether it's forming the outline of a particular line or a shape,
- 10:03:38for example.
- 10:03:39And that might be useful if you're trying to use
- 10:03:42all of these various different features to combine them
- 10:03:44to say something meaningful about an image as a whole.
- 10:03:48So how, then, does image convolution work?
- 10:03:50Well, we start with a kernel matrix.
- 10:03:52And the kernel matrix looks something like this.
- 10:03:54And the idea of this is that, given a pixel that will be the middle pixel,
- 10:03:58we're going to multiply each of the neighboring pixels
- 10:04:00by these values in order to get some sort of result
- 10:04:04by summing up all the numbers together.
- 10:04:06So if I take this kernel, which you can think of as a filter
- 10:04:09that I'm going to apply to the image, and let's say that I take this image.
- 10:04:13This is a 4 by 4 image.
- 10:04:14We'll think of it as just a black and white image,
- 10:04:16where each one is just a single pixel value.
- 10:04:19So somewhere between 0 and 255, for example.
- 10:04:22So we have a whole bunch of individual pixel values like this.
- 10:04:25And what I'd like to do is apply this kernel, this filter, so to speak,
- 10:04:30to this image.
- 10:04:32And the way I'll do that is, all right, the kernel is 3 by 3.
- 10:04:35You can imagine a 5 by 5 kernel or a larger kernel, too.
- 10:04:38And I'll take it and just first apply it to the first 3
- 10:04:41by 3 section of the image.
- 10:04:43And what I'll do is I'll take each of these pixel values,
- 10:04:46multiply it by its corresponding value in the filter matrix,
- 10:04:50and add all of the results together.
- 10:04:53So here, for example, I'll say 10 times 0, plus 20 times negative 1,
- 10:04:59plus 30 times 0, so on and so forth, doing all of this calculation.
- 10:05:03And at the end, if I take all these values,
- 10:05:05multiply them by their corresponding value in the kernel,
- 10:05:08add the results together, for this particular set of 9 pixels,
- 10:05:11I get the value of 10, for example.
- 10:05:14And then what I'll do is I'll slide this 3 by 3 grid, effectively, over.
- 10:05:19I'll slide the kernel by 1 to look at the next 3 by 3 section.
- 10:05:24Here, I'm just sliding it over by 1 pixel.
- 10:05:26But you might imagine a different stride length,
- 10:05:28or maybe I jump by multiple pixels at a time if you really wanted to.
- 10:05:31You have different options here.
- 10:05:32But here, I'm just sliding over, looking at the next 3 by 3 section.
- 10:05:35And I'll do the same math, 20 times 0, plus 30 times negative 1,
- 10:05:40plus 40 times 0, plus 20 times negative 1, so on and so forth, plus 30 times 5.
- 10:05:45And what I end up getting is the number 20.
- 10:05:47Then you can imagine shifting over to this one, doing the same thing,
- 10:05:50calculating the number 40, for example, and then doing the same thing here,
- 10:05:54and calculating a value there as well.
- 10:05:56And so what we have now is what we'll call a feature map.
- 10:06:00We have taken this kernel, applied it to each
- 10:06:03of these various different regions, and what we get
- 10:06:06is some representation of a filtered version of that image.
- 10:06:11And so to give a more concrete example of why
- 10:06:13it is that this kind of thing could be useful,
- 10:06:14let's take this kernel matrix, for example, which is quite a famous one,
- 10:06:18that has an 8 in the middle, and then all of the neighboring pixels
- 10:06:22get a negative 1.
- 10:06:23And let's imagine we wanted to apply that to a 3
- 10:06:26by 3 part of an image that looks like this, where all the values are the same.
- 10:06:31They're all 20, for instance.
- 10:06:33Well, in this case, if you do 20 times 8, and then subtract 20, subtract 20,
- 10:06:38subtract 20 for each of the eight neighbors, well, the result of that
- 10:06:40is you just get that expression, which comes out to be 0.
- 10:06:44You multiplied 20 by 8, but then you subtracted
- 10:06:4720 eight times, according to that particular kernel.
- 10:06:50The result of all that is just 0.
- 10:06:52So the takeaway here is that when a lot of the pixels are the same value,
- 10:06:56we end up getting a value close to 0.
- 10:06:59If, though, we had something like this, 20 is along this first row,
- 10:07:02then 50 is in the second row, and 50 is in the third row, well,
- 10:07:05then when you do this, because it's the same kind of math, 20 times negative 1,
- 10:07:0820 times negative 1, so on and so forth, then I get a higher value,
- 10:07:12a value like 90 in this particular case.
- 10:07:15And so the more general idea here is that by applying this kernel, negative 1s,
- 10:07:218 in the middle, and then negative 1s, what I get
- 10:07:23is when this middle value is very different from the neighboring values,
- 10:07:29like 50 is greater than these 20s, then you'll
- 10:07:31end up with a value higher than 0.
- 10:07:34If this number is higher than its neighbors,
- 10:07:36you end up getting a bigger output.
- 10:07:38But if this value is the same as all of its neighbors,
- 10:07:41then you get a lower output, something like 0.
- 10:07:43And it turns out that this sort of filter can therefore
- 10:07:46be used in something like detecting edges in an image.
- 10:07:49Or I want to detect the boundaries between various different objects
- 10:07:53inside of an image.
- 10:07:54I might use a filter like this, which is able to tell
- 10:07:57whether the value of this pixel is different
- 10:08:00from the values of the neighboring pixel,
- 10:08:02if it's greater than the values of the pixels that happen to surround it.
- 10:08:06And so we can use this in terms of image filtering.
- 10:08:09And so I'll show you an example of that.
- 10:08:11I have here in filter.py a file that uses Python's image library,
- 10:08:17or PIL, to do some image filtering.
- 10:08:21I go ahead and open an image.
- 10:08:23And then all I'm going to do is apply a kernel to that image.
- 10:08:26It's going to be a 3 by 3 kernel, same kind of kernel we saw before.
- 10:08:30And here is the kernel.
- 10:08:31This is just a list representation of the same matrix
- 10:08:34that I showed you a moment ago.
- 10:08:36It's negative 1, negative 1, negative 1.
- 10:08:38The second row is negative 1, 8, negative 1.
- 10:08:40And the third row is all negative 1s.
- 10:08:43And then at the end, I'm going to go ahead and show the filtered image.
- 10:08:47So if, for example, I go into convolution directory
- 10:08:53and I open up an image, like bridge.png, this
- 10:08:56is what an input image might look like, just an image of a bridge over a river.
- 10:09:02Now I'm going to go ahead and run this filter program on the bridge.
- 10:09:07And what I get is this image here.
- 10:09:10Just by taking the original image and applying that filter
- 10:09:13to each 3 by 3 grid, I've extracted all of the boundaries,
- 10:09:17all of the edges inside the image that separate one part of the image
- 10:09:20from another.
- 10:09:21So here I've got a representation of boundaries
- 10:09:24between particular parts of the image.
- 10:09:26And you might imagine that if a machine learning algorithm is
- 10:09:28trying to learn what an image is of, a filter like this could be pretty useful.
- 10:09:33Maybe the machine learning algorithm doesn't
- 10:09:35care about all of the details of the image.
- 10:09:38It just cares about certain useful features.
- 10:09:40It cares about particular shapes that are
- 10:09:42able to help it determine that based on the image,
- 10:09:45this is going to be a bridge, for example.
- 10:09:47And so this type of idea of image convolution
- 10:09:50can allow us to apply filters to images that allow us to extract useful results
- 10:09:55out of those images, taking an image and extracting its edges, for example.
- 10:09:59And you might imagine many other filters that
- 10:10:01could be applied to an image that are able to extract particular values as
- 10:10:05well.
- 10:10:05And a filter might have separate kernels for the red values, the green values,
- 10:10:08and the blue values that are all summed together at the end,
- 10:10:11such that you could have particular filters looking for,
- 10:10:14is there red in this part of the image?
- 10:10:15Are there green in other parts of the image?
- 10:10:17You can begin to assemble these relevant and useful filters
- 10:10:20that are able to do these calculations as well.
- 10:10:24So that then was the idea of image convolution,
- 10:10:26applying some sort of filter to an image to be
- 10:10:29able to extract some useful features out of that image.
- 10:10:32But all the while, these images are still pretty big.
- 10:10:35There's a lot of pixels involved in the image.
- 10:10:38And realistically speaking, if you've got a really big image,
- 10:10:40that poses a couple of problems.
- 10:10:42One, it means a lot of input going into the neural network.
- 10:10:45But two, it also means that we really have
- 10:10:48to care about what's in each particular pixel.
- 10:10:50Whereas realistically, we often, if you're looking at an image,
- 10:10:54you don't care whether something is in one particular pixel versus the pixel
- 10:10:58immediately to the right of it.
- 10:10:59They're pretty close together.
- 10:11:01You really just care about whether there's a particular feature
- 10:11:03in some region of the image.
- 10:11:05And maybe you don't care about exactly which pixel it happens to be in.
- 10:11:09And so there's a technique we can use known as pooling.
- 10:11:11And what pooling is, is it means reducing the size of an input
- 10:11:15by sampling from regions inside of the input.
- 10:11:18So we're going to take a big image and turn it into a smaller image
- 10:11:22by using pooling.
- 10:11:23And in particular, one of the most popular types of pooling
- 10:11:25is called max pooling.
- 10:11:27And what max pooling does is it pools just
- 10:11:29by choosing the maximum value in a particular region.
- 10:11:33So for example, let's imagine I had this 4 by 4 image.
- 10:11:36But I wanted to reduce its dimensions.
- 10:11:38I wanted to make it a smaller image so that I have fewer inputs to work with.
- 10:11:42Well, what I could do is I could apply a 2 by 2 max pool,
- 10:11:47where the idea would be that I'm going to first look at this 2 by 2 region
- 10:11:50and say, what is the maximum value in that region?
- 10:11:53Well, it's the number 50.
- 10:11:54So we'll go ahead and just use the number 50.
- 10:11:57And then we'll look at this 2 by 2 region.
- 10:11:58What is the maximum value here?
- 10:12:00It's 110, so that's going to be my value.
- 10:12:02Likewise here, the maximum value looks like 20.
- 10:12:04Go ahead and put that there.
- 10:12:05Then for this last region, the maximum value was 40.
- 10:12:09So we'll go ahead and use that.
- 10:12:10And what I have now is a smaller representation
- 10:12:14of this same original image that I obtained just
- 10:12:17by picking the maximum value from each of these regions.
- 10:12:21So again, the advantages here are now I only
- 10:12:25have to deal with a 2 by 2 input instead of a 4 by 4.
- 10:12:27And you can imagine shrinking the size of an image even more.
- 10:12:31But in addition to that, I'm now able to make my analysis
- 10:12:36independent of whether a particular value was in this pixel or this pixel.
- 10:12:40I don't care if the 50 was here or here.
- 10:12:42As long as it was generally in this region,
- 10:12:45I'll still get access to that value.
- 10:12:47So it makes our algorithms a little bit more robust as well.
- 10:12:51So that then is pooling, taking the size of the image,
- 10:12:54reducing it a little bit by just sampling from particular regions
- 10:12:58inside of the image.
- 10:12:59And now we can put all of these ideas together, pooling, image convolution,
- 10:13:03and neural networks all together into another type of neural network
- 10:13:06called a convolutional neural network, or a CNN, which
- 10:13:10is a neural network that uses this convolution step usually
- 10:13:14in the context of analyzing an image, for example.
- 10:13:18And so the way that a convolutional neural network works
- 10:13:20is that we start with some sort of input image, some grid of pixels.
- 10:13:24But rather than immediately put that into the neural network layers
- 10:13:27that we've seen before, we'll start by applying a convolution step,
- 10:13:31where the convolution step involves applying
- 10:13:33some number of different image filters to our original image
- 10:13:36in order to get what we call a feature map, the result of applying
- 10:13:40some filter to an image.
- 10:13:41And we could do this once, but in general, we'll do this multiple times,
- 10:13:45getting a whole bunch of different feature maps, each of which
- 10:13:48might extract some different relevant feature out of the image,
- 10:13:51some different important characteristic of the image
- 10:13:53that we might care about using in order to calculate
- 10:13:56what the result should be.
- 10:13:58And in the same way that when we train neural networks,
- 10:14:01we can train neural networks to learn the weights between particular units
- 10:14:04inside of the neural networks, we can also train neural networks
- 10:14:07to learn what those filters should be, what
- 10:14:09the values of the filters should be in order
- 10:14:11to get the most useful, most relevant information out of the original image
- 10:14:15just by figuring out what setting of those filter values,
- 10:14:18the values inside of that kernel, results in minimizing the loss function,
- 10:14:23minimizing how poorly our hypothesis actually
- 10:14:26performs in figuring out the classification of a particular image,
- 10:14:30for example.
- 10:14:32So we first apply this convolution step, get a whole bunch
- 10:14:34of these various different feature maps.
- 10:14:36But these feature maps are quite large.
- 10:14:38There's a lot of pixel values that happen to be here.
- 10:14:41And so a logical next step to take is a pooling step,
- 10:14:44where we reduce the size of these images by using max pooling,
- 10:14:48for example, extracting the maximum value from any particular region.
- 10:14:51There are other pooling methods that exist as well,
- 10:14:53depending on the situation.
- 10:14:54You could use something like average pooling,
- 10:14:57where instead of taking the maximum value from a region,
- 10:14:59you take the average value from a region, which has its uses as well.
- 10:15:03But in effect, what pooling will do is it will take these feature maps
- 10:15:07and reduce their dimensions so that we end up
- 10:15:09with smaller grids with fewer pixels.
- 10:15:12And this then is going to be easier for us to deal with.
- 10:15:14It's going to mean fewer inputs that we have to worry about.
- 10:15:16And it's also going to mean we're more resilient,
- 10:15:19more robust against potential movements of particular values,
- 10:15:22just by one pixel, when ultimately we really
- 10:15:24don't care about those one-pixel differences that
- 10:15:27might arise in the original image.
- 10:15:30And now, after we've done this pooling step,
- 10:15:32now we have a whole bunch of values that we can then flatten out and just put
- 10:15:36into a more traditional neural network.
- 10:15:38So we go ahead and flatten it, and then we
- 10:15:40end up with a traditional neural network that
- 10:15:42has one input for each of these values in each of these resulting feature
- 10:15:46maps after we do the convolution and after we do the pooling step.
- 10:15:51And so this then is the general structure of a convolutional network.
- 10:15:54We begin with the image, apply convolution, apply pooling,
- 10:15:58flatten the results, and then put that into a more traditional neural
- 10:16:01network that might itself have hidden layers.
- 10:16:03You can have deep convolutional networks that
- 10:16:05have hidden layers in between this flattened layer and the eventual output
- 10:16:09to be able to calculate various different features of those values.
- 10:16:13But this then can help us to be able to use convolution and pooling
- 10:16:17to use our knowledge about the structure of an image
- 10:16:19to be able to get better results, to be able to train our networks faster
- 10:16:23in order to better capture particular parts of the image.
- 10:16:27And there's no reason necessarily why you can only use these steps once.
- 10:16:30In fact, in practice, you'll often use convolution and pooling
- 10:16:33multiple times in multiple different steps.
- 10:16:36See, what you might imagine doing is starting with an image,
- 10:16:39first applying convolution to get a whole bunch of maps,
- 10:16:42then applying pooling, then applying convolution again,
- 10:16:45because these maps are still pretty big.
- 10:16:48You can apply convolution to try and extract relevant features out
- 10:16:51of this result. Then take those results, apply pooling
- 10:16:55in order to reduce their dimensions, and then take that
- 10:16:57and feed it into a neural network that maybe has fewer inputs.
- 10:17:01So here I have two different convolution and pooling steps.
- 10:17:04I do convolution and pooling once, and then I do convolution and pooling
- 10:17:08a second time, each time extracting useful features
- 10:17:11from the layer before it, each time using pooling
- 10:17:14to reduce the dimensions of what you're ultimately looking at.
- 10:17:17And the goal now of this sort of model is that in each of these steps,
- 10:17:21you can begin to learn different types of features of the original image.
- 10:17:25That maybe in the first step, you learn very low level features.
- 10:17:28Just learn and look for features like edges and curves and shapes,
- 10:17:31because based on pixels and their neighboring values, you can figure out,
- 10:17:36all right, what are the edges?
- 10:17:37What are the curves?
- 10:17:38What are the various different shapes that might be present there?
- 10:17:41But then once you have a mapping that just represents
- 10:17:43where the edges and curves and shapes happen to be,
- 10:17:46you can imagine applying the same sort of process again
- 10:17:49to begin to look for higher level features, look for objects,
- 10:17:51maybe look for people's eyes and facial recognition, for example.
- 10:17:55Maybe look for more complex shapes like the curves on a particular number
- 10:17:59if you're trying to recognize a digit in a handwriting recognition sort
- 10:18:02of scenario.
- 10:18:03And then after all of that, now that you have these results that
- 10:18:06represent these higher level features, you
- 10:18:08can pass them into a neural network, which is really just a deep neural
- 10:18:12network that looks like this, where you might imagine
- 10:18:14making a binary classification or classifying into multiple categories
- 10:18:18or performing various different tasks on this sort of model.
- 10:18:23So convolutional neural networks can be quite powerful and quite popular
- 10:18:26when it comes towards trying to analyze images.
- 10:18:28We don't strictly need them.
- 10:18:29We could have just used a vanilla neural network
- 10:18:32that just operates with layer after layer, as we've seen before.
- 10:18:35But these convolutional neural networks can be quite helpful,
- 10:18:38in particular, because of the way they model
- 10:18:40the way a human might look at an image, that instead of a human looking
- 10:18:43at every single pixel simultaneously and trying to convolve all of them
- 10:18:46by multiplying them together, you might imagine
- 10:18:48that what convolution is really doing is looking
- 10:18:50at various different regions of the image
- 10:18:53and extracting relevant information and features out
- 10:18:56of those parts of the image, the same way
- 10:18:57that a human might have visual receptors that
- 10:18:59are looking at particular parts of what they see
- 10:19:02and using those combining them to figure out
- 10:19:04what meaning they can draw from all of those various different inputs.
- 10:19:09And so you might imagine applying this to a situation
- 10:19:11like handwriting recognition.
- 10:19:13So we'll go ahead and see an example of that now,
- 10:19:16where I'll go ahead and open up handwriting.py.
- 10:19:19Again, what we do here is we first import TensorFlow.
- 10:19:23And then TensorFlow, it turns out, has a few data sets
- 10:19:26that are built into the library that you can just immediately access.
- 10:19:30And one of the most famous data sets in machine learning
- 10:19:33is the MNIST data set, which is just a data
- 10:19:35set of a whole bunch of samples of people's handwritten digits.
- 10:19:38I showed you a slide of that a little while ago.
- 10:19:41And what we can do is just immediately access
- 10:19:43that data set which is built into the library
- 10:19:45so that if I want to do something like train
- 10:19:47on a whole bunch of handwritten digits, I can just use the data set
- 10:19:50that is provided to me.
- 10:19:52Of course, if I had my own data set of handwritten images,
- 10:19:55I can apply the same idea.
- 10:19:56I'd first just need to take those images and turn them
- 10:19:59into an array of pixels, because that's the way that these
- 10:20:02are going to be formatted.
- 10:20:03They're going to be formatted as, effectively,
- 10:20:05an array of individual pixels.
- 10:20:08Now there's a bit of reshaping I need to do,
- 10:20:10just turning the data into a format that I
- 10:20:12can put into my convolutional neural network.
- 10:20:14So this is doing things like taking all the values
- 10:20:17and dividing them by 255.
- 10:20:19If you remember, these color values tend to range from 0 to 255.
- 10:20:22So I can divide them by 255 just to put them
- 10:20:25into 0 to 1 range, which might be a little bit easier to train on.
- 10:20:29And then doing various other modifications to the data
- 10:20:32just to get it into a nice usable format.
- 10:20:34But here's the interesting and important part.
- 10:20:37Here is where I create the convolutional neural network, the CNN,
- 10:20:41where here I'm saying, go ahead and use a sequential model.
- 10:20:44And before I could use model.add to say add a layer, add a layer, add a layer,
- 10:20:47another way I could define it is just by passing as input
- 10:20:50to this sequential neural network a list of all of the layers that I want.
- 10:20:55And so here, the very first layer in my model is a convolution layer,
- 10:21:00where I'm first going to apply convolution to my image.
- 10:21:03I'm going to use 13 different filters.
- 10:21:05So my model is going to learn 32, rather, 32 different filters
- 10:21:09that I would like to learn on the input image, where each filter is going
- 10:21:13to be a 3 by 3 kernel.
- 10:21:15So we saw those 3 by 3 kernels before, where
- 10:21:17we could multiply each value in a 3 by 3 grid by a value,
- 10:21:20multiply it, and add all the results together.
- 10:21:22So here, I'm going to learn 32 different of these 3 by 3 filters.
- 10:21:27I can, again, specify my activation function.
- 10:21:29And I specify what my input shape is.
- 10:21:32My input shape in the banknotes case was just 4.
- 10:21:34I had 4 inputs.
- 10:21:36My input shape here is going to be 28, 28, 1,
- 10:21:40because for each of these handwritten digits,
- 10:21:42it turns out that the MNIST data set organizes their data.
- 10:21:46Each image is a 28 by 28 pixel grid.
- 10:21:49So we're going to have a 28 by 28 pixel grid.
- 10:21:51And each one of those images only has one channel value.
- 10:21:54These handwritten digits are just black and white.
- 10:21:56So there's just a single color value representing
- 10:21:59how much black or how much white.
- 10:22:00You might imagine that in a color image, if you
- 10:22:02were doing this sort of thing, you might have three different channels,
- 10:22:05a red, a green, and a blue channel, for example.
- 10:22:07But in the case of just handwriting recognition,
- 10:22:09recognizing a digit, we're just going to use a single value for,
- 10:22:12like, shaded in or not shaded in.
- 10:22:14And it might range, but it's just a single color value.
- 10:22:18And that, then, is the very first layer of our neural network,
- 10:22:22a convolutional layer that will take the input
- 10:22:24and learn a whole bunch of different filters
- 10:22:26that we can apply to the input to extract meaningful features.
- 10:22:30Next step is going to be a max pooling layer, also built right
- 10:22:34into TensorFlow, where this is going to be a layer that
- 10:22:37is going to use a pool size of 2 by 2, meaning
- 10:22:40we're going to look at 2 by 2 regions inside of the image
- 10:22:43and just extract the maximum value.
- 10:22:45Again, we've seen why this can be helpful.
- 10:22:47It'll help to reduce the size of our input.
- 10:22:49And once we've done that, we'll go ahead and flatten all of the units
- 10:22:53just into a single layer that we can then
- 10:22:55pass into the rest of the neural network.
- 10:22:57And now, here's the rest of the neural network.
- 10:23:00Here, I'm saying, let's add a hidden layer to my neural network
- 10:23:02with 128 units, so a whole bunch of hidden units
- 10:23:06inside of the hidden layer.
- 10:23:07And just to prevent overfitting, I can add a dropout to that.
- 10:23:11Say, you know what, when you're training, randomly dropout half
- 10:23:14of the nodes from this hidden layer just to make sure
- 10:23:16we don't become over-reliant on any particular node,
- 10:23:19we begin to really generalize and stop ourselves from overfitting.
- 10:23:22So TensorFlow allows us, just by adding a single line,
- 10:23:25to add dropout into our model as well, such that when it's training,
- 10:23:28it will perform this dropout step in order
- 10:23:31to help make sure that we don't overfit on this particular data.
- 10:23:36And then finally, I add an output layer.
- 10:23:38The output layer is going to have 10 units, one for each category
- 10:23:42that I would like to classify digits into, so 0 through 9,
- 10:23:4510 different categories.
- 10:23:47And the activation function I'm going to use here
- 10:23:49is called the softmax activation function.
- 10:23:52And in short, what the softmax activation function is going to do
- 10:23:55is it's going to take the output and turn it
- 10:23:57into a probability distribution.
- 10:23:59So ultimately, it's going to tell me, what
- 10:24:01did we estimate the probability is that this
- 10:24:03is a 2 versus a 3 versus a 4.
- 10:24:06And so it will turn it into that probability distribution for me.
- 10:24:10Next up, I'll go ahead and compile my model
- 10:24:12and fit it on all of my training data.
- 10:24:15And then I can evaluate how well the neural network performs.
- 10:24:19And then I've added to my Python program,
- 10:24:21if I've provided a command line argument like the name of a file,
- 10:24:24I'm going to go ahead and save the model to a file.
- 10:24:27And so this can be quite useful too.
- 10:24:29Once you've done the training step, which could take some time in terms
- 10:24:31of taking all the time, going through the data,
- 10:24:34running back propagation with gradient descent to be able to say, all right,
- 10:24:38how should we adjust the weight to this particular model?
- 10:24:40You end up calculating values for these weights,
- 10:24:42calculating values for these filters.
- 10:24:44You'd like to remember that information so you can use it later.
- 10:24:47And so TensorFlow allows us to just save a model to a file,
- 10:24:51such that later, if we want to use the model we've learned,
- 10:24:53use the weights that we've learned to make some sort of new prediction,
- 10:24:57we can just use the model that already exists.
- 10:25:00So what we're doing here is after we've done all the calculation,
- 10:25:03we go ahead and save the model to a file, such
- 10:25:07that we can use it a little bit later.
- 10:25:09So for example, if I go into digits, I'm going to run handwriting.py.
- 10:25:17I won't save it this time.
- 10:25:18We'll just run it and go ahead and see what happens.
- 10:25:20What will happen is we need to go through the model in order
- 10:25:22to train on all of these samples of handwritten digits.
- 10:25:26The MNIST data set gives us thousands and thousands
- 10:25:28of sample handwritten digits in the same format
- 10:25:31that we can use in order to train.
- 10:25:33And so now what you're seeing is this training process.
- 10:25:35And unlike the banknotes case, where there was much fewer data points,
- 10:25:39the data was very, very simple, here this data is more complex
- 10:25:42and this training process takes time.
- 10:25:44And so this is another one of those cases where when training neural networks,
- 10:25:48this is why computational power is so important that oftentimes you
- 10:25:52see people wanting to use sophisticated GPUs in order
- 10:25:55to more efficiently be able to do this sort of neural network training.
- 10:25:59It also speaks to the reason why more data can be helpful.
- 10:26:02The more sample data points you have, the better
- 10:26:04you can begin to do this training.
- 10:26:06So here we're going through 60,000 different samples of handwritten digits.
- 10:26:10And I said we're going to go through them 10 times.
- 10:26:13We're going to go through the data set 10 times, training each time,
- 10:26:16hopefully improving upon our weights with every time
- 10:26:18we run through this data set.
- 10:26:20And we can see over here on the right what the accuracy is each time
- 10:26:23we go ahead and run this model, that the first time it
- 10:26:26looks like we got an accuracy of about 92% of the digits
- 10:26:29correct based on this training set.
- 10:26:31We increased that to 96% or 97%.
- 10:26:34And every time we run this, we're going to see hopefully the accuracy
- 10:26:38improve as we continue to try and use that gradient descent,
- 10:26:41that process of trying to run the algorithm,
- 10:26:43to minimize the loss that we get in order to more accurately
- 10:26:46predict what the output should be.
- 10:26:49And what this process is doing is it's learning not only the weights,
- 10:26:52but it's learning the features to use, the kernel matrix
- 10:26:55to use when performing that convolution step.
- 10:26:57Because this is a convolutional neural network,
- 10:26:59where I'm first performing those convolutions
- 10:27:02and then doing the more traditional neural network structure,
- 10:27:05this is going to learn all of those individual steps as well.
- 10:27:09And so here we see the TensorFlow provides me with some very nice output,
- 10:27:12telling me about how many seconds are left with each of these training
- 10:27:15runs that allows me to see just how well we're doing.
- 10:27:18So we'll go ahead and see how this network performs.
- 10:27:21It looks like we've gone through the data set seven times.
- 10:27:23We're going through it an eighth time now.
- 10:27:26And at this point, the accuracy is pretty high.
- 10:27:28We saw we went from 92% up to 97%.
- 10:27:32Now it looks like 98%.
- 10:27:33And at this point, it seems like things are starting to level out.
- 10:27:36It's probably a limit to how accurate we can ultimately be
- 10:27:39without running the risk of overfitting.
- 10:27:41Of course, with enough nodes, you would just
- 10:27:42memorize the input and overfit upon them.
- 10:27:44But we'd like to avoid doing that.
- 10:27:46And Dropout will help us with this.
- 10:27:48But now we see we're almost done finishing our training step.
- 10:27:53We're at 55,000.
- 10:27:55All right, we finished training.
- 10:27:56And now it's going to go ahead and test for us on 10,000 samples.
- 10:28:00And it looks like on the testing set, we were at 98.8% accurate.
- 10:28:04So we ended up doing pretty well, it seems,
- 10:28:06on this testing set to see how accurately can we
- 10:28:10predict these handwritten digits.
- 10:28:13And so what we could do then is actually test it out.
- 10:28:15I've written a program called Recognition.py using PyGame.
- 10:28:19If you pass it a model that's been trained,
- 10:28:21and I pre-trained an example model using this input data, what we can do
- 10:28:26is see whether or not we've been able to train
- 10:28:27this convolutional neural network to be able to predict handwriting,
- 10:28:31for example.
- 10:28:32So I can try, just like drawing a handwritten digit.
- 10:28:35I'll go ahead and draw the number 2, for example.
- 10:28:39So there's my number 2.
- 10:28:40Again, this is messy.
- 10:28:41If you tried to imagine, how would you write a program with just ifs
- 10:28:44and thens to be able to do this sort of calculation,
- 10:28:46it would be tricky to do so.
- 10:28:48But here I'll press Classify, and all right,
- 10:28:50it seems I was able to correctly classify that what I drew was the number 2.
- 10:28:53I'll go ahead and reset it, try it again.
- 10:28:55We'll draw an 8, for example.
- 10:28:57So here is an 8.
- 10:29:00Press Classify.
- 10:29:01And all right, it predicts that the digit that I drew was an 8.
- 10:29:05And the key here is this really begins to show the power of what
- 10:29:08the neural network is doing, somehow looking
- 10:29:09at various different features of these different pixels,
- 10:29:12figuring out what the relevant features are,
- 10:29:14and figuring out how to combine them to get a classification.
- 10:29:17And this would be a difficult task to provide explicit instructions
- 10:29:21to the computer on how to do, to use a whole bunch of ifs ands
- 10:29:24to process all these pixel values to figure out
- 10:29:27what the handwritten digit is.
- 10:29:28Everyone's going to draw their 8s a little bit differently.
- 10:29:31If I drew the 8 again, it would look a little bit different.
- 10:29:33And yet, ideally, we want to train a network to be robust enough
- 10:29:37so that it begins to learn these patterns on its own.
- 10:29:40All I said was, here is the structure of the network,
- 10:29:43and here is the data on which to train the network.
- 10:29:45And the network learning algorithm just tries
- 10:29:47to figure out what is the optimal set of weights, what
- 10:29:50is the optimal set of filters to use them in order
- 10:29:52to be able to accurately classify a digit into one category or another.
- 10:29:57Just going to show the power of these sorts of convolutional neural
- 10:30:00networks.
- 10:30:02And so that then was a look at how we can use convolutional neural networks
- 10:30:06to begin to solve problems with regards to computer vision,
- 10:30:10the ability to take an image and begin to analyze it.
- 10:30:13So this is the type of analysis you might imagine
- 10:30:15that's happening in self-driving cars that
- 10:30:18are able to figure out what filters to apply to an image
- 10:30:21to understand what it is that the computer is looking at,
- 10:30:24or the same type of idea that might be applied
- 10:30:26to facial recognition and social media to be
- 10:30:28able to determine how to recognize faces in an image as well.
- 10:30:31You can imagine a neural network that instead of classifying
- 10:30:34into one of 10 different digits could instead classify like,
- 10:30:38is this person A or is this person B, trying
- 10:30:40to tell those people apart just based on convolution.
- 10:30:45And so now what we'll take a look at is yet another type of neural network
- 10:30:48that can be quite popular for certain types of tasks.
- 10:30:50But to do so, we'll try to generalize and think about our neural network
- 10:30:54a little bit more abstractly.
- 10:30:55That here we have a sample deep neural network
- 10:30:58where we have this input layer, a whole bunch of different hidden layers
- 10:31:01that are performing certain types of calculations,
- 10:31:04and then an output layer here that just generates some sort of output
- 10:31:07that we care about calculating.
- 10:31:09But we could imagine representing this a little more simply like this.
- 10:31:14Here is just a more abstract representation of our neural network.
- 10:31:17We have some input that might be like a vector
- 10:31:20of a whole bunch of different values as our input.
- 10:31:22That gets passed into a network that performs some sort of calculation
- 10:31:25or computation, and that network produces some sort of output.
- 10:31:29That output might be a single value.
- 10:31:31It might be a whole bunch of different values.
- 10:31:33But this is the general structure of the neural network that we've seen.
- 10:31:36There is some sort of input that gets fed into the network.
- 10:31:39And using that input, the network calculates what the output should be.
- 10:31:43And this sort of model for a neural network
- 10:31:46is what we might call a feed-forward neural network.
- 10:31:49Feed-forward neural networks have connections only in one direction.
- 10:31:52They move from one layer to the next layer to the layer after that,
- 10:31:56such that the inputs pass through various different hidden layers
- 10:31:59and then ultimately produce some sort of output.
- 10:32:02So feed-forward neural networks were very helpful
- 10:32:05for solving these types of classification problems that we saw before.
- 10:32:08We have a whole bunch of input.
- 10:32:10We want to learn what setting of weights will allow us
- 10:32:12to calculate the output effectively.
- 10:32:14But there are some limitations on feed-forward neural networks
- 10:32:16that we'll see in a moment.
- 10:32:17In particular, the input needs to be of a fixed shape,
- 10:32:20like a fixed number of neurons are in the input layer.
- 10:32:23And there's a fixed shape for the output,
- 10:32:24like a fixed number of neurons in the output layer.
- 10:32:28And that has some limitations of its own.
- 10:32:30And a possible solution to this, and we'll
- 10:32:33see examples of the types of problems we can solve for this in just a second,
- 10:32:36is instead of just a feed-forward neural network,
- 10:32:38where there are only connections in one direction from left to right
- 10:32:41effectively across the network, we could also imagine a recurrent neural
- 10:32:46network, where a recurrent neural network generates
- 10:32:48output that gets fed back into itself as input for future runs of that network.
- 10:32:54So whereas in a traditional neural network,
- 10:32:57we have inputs that get fed into the network, that get fed into the output.
- 10:33:00And the only thing that determines the output
- 10:33:02is based on the original input and based on the calculation
- 10:33:05we do inside of the network itself.
- 10:33:08This goes in contrast with a recurrent neural network,
- 10:33:11where in a recurrent neural network, you can imagine output from the network
- 10:33:14feeding back to itself into the network again as input
- 10:33:18for the next time you do the calculations inside of the network.
- 10:33:22What this allows is it allows the network to maintain some sort of state,
- 10:33:27to store some sort of information that can be used on future runs of the network.
- 10:33:33Previously, the network just defined some weights,
- 10:33:35and we passed inputs through the network, and it generated outputs.
- 10:33:38But the network wasn't saving any information based on those inputs
- 10:33:42to be able to remember for future iterations or for future runs.
- 10:33:45What a recurrent neural network will let us do
- 10:33:47is let the network store information that gets passed back in as input
- 10:33:51to the network again the next time we try and perform some sort of action.
- 10:33:55And this is particularly helpful when dealing with sequences of data.
- 10:34:00So we'll see a real world example of this right now, actually.
- 10:34:02Microsoft has developed an AI known as the caption bot.
- 10:34:07And what the caption bot does is it says,
- 10:34:09I can understand the content of any photograph,
- 10:34:11and I'll try to describe it as well as any human.
- 10:34:13I'll analyze your photo, but I won't store it or share it.
- 10:34:16And so what Microsoft's caption bot seems to be claiming to do
- 10:34:19is it can take an image and figure out what's in the image
- 10:34:22and just give us a caption to describe it.
- 10:34:25So let's try it out.
- 10:34:26Here, for example, is an image of Harvard Square.
- 10:34:29It's some people walking in front of one of the buildings at Harvard Square.
- 10:34:32I'll go ahead and take the URL for that image,
- 10:34:34and I'll paste it into caption bot and just press Go.
- 10:34:39So caption bot is analyzing the image, and then it
- 10:34:41says, I think it's a group of people walking
- 10:34:44in front of a building, which seems amazing.
- 10:34:46The AI is able to look at this image and figure out what's in the image.
- 10:34:50And the important thing to recognize here
- 10:34:52is that this is no longer just a classification task.
- 10:34:55We saw being able to classify images with a convolutional neural network
- 10:34:58where the job was take the image and then figure out,
- 10:35:01is it a 0 or a 1 or a 2, or is it this person's face or that person's face?
- 10:35:05What seems to be happening here is the input is an image,
- 10:35:09and we know how to get networks to take input of images,
- 10:35:12but the output is text.
- 10:35:14It's a sentence.
- 10:35:15It's a phrase, like a group of people walking in front of a building.
- 10:35:19And this would seem to pose a challenge for our more traditional feed-forward
- 10:35:23neural networks, for the reason being that in traditional neural networks,
- 10:35:28we just have a fixed-size input and a fixed-size output.
- 10:35:31There are a certain number of neurons in the input to our neural network
- 10:35:35and a certain number of outputs for our neural network,
- 10:35:37and then some calculation that goes on in between.
- 10:35:39But the size of the inputs and the number of values in the input
- 10:35:42and the number of values in the output, those
- 10:35:44are always going to be fixed based on the structure of the neural network.
- 10:35:49And that makes it difficult to imagine how a neural network could take an image
- 10:35:52like this and say it's a group of people walking in front of the building
- 10:35:56because the output is text, like it's a sequence of words.
- 10:36:00Now, it might be possible for a neural network
- 10:36:02to output one word, one word you could represent as a vector of values,
- 10:36:06and you can imagine ways of doing that.
- 10:36:08Next time, we'll talk a little bit more about AI
- 10:36:10as it relates to language and language processing.
- 10:36:13But a sequence of words is much more challenging
- 10:36:15because depending on the image, you might imagine the output
- 10:36:18is a different number of words.
- 10:36:19We could have sequences of different lengths,
- 10:36:22and somehow we still want to be able to generate the appropriate output.
- 10:36:26And so the strategy here is to use a recurrent neural network,
- 10:36:30a neural network that can feed its own output back into itself
- 10:36:34as input for the next time.
- 10:36:36And this allows us to do what we call a one-to-many relationship
- 10:36:40for inputs to outputs, that in vanilla, more traditional neural networks,
- 10:36:43these are what we might consider to be one-to-one neural networks.
- 10:36:47You pass in one set of values as input.
- 10:36:49You get one vector of values as the output.
- 10:36:53But in this case, we want to pass in one value as input, the image,
- 10:36:56and we want to get a sequence, many values as output,
- 10:36:59where each value is like one of these words that
- 10:37:02gets produced by this particular algorithm.
- 10:37:05And so the way we might do this is we might imagine starting
- 10:37:08by providing input, the image, into our neural network.
- 10:37:11And the neural network is going to generate output,
- 10:37:13but the output is not going to be the whole sequence of words,
- 10:37:16because we can't represent the whole sequence of words
- 10:37:18using just a fixed set of neurons.
- 10:37:20Instead, the output is just going to be the first word.
- 10:37:24We're going to train the network to output what the first word of the caption
- 10:37:27should be.
- 10:37:28And you could imagine that Microsoft has trained this
- 10:37:30by running a whole bunch of training samples through the AI,
- 10:37:33giving it a whole bunch of pictures and what the appropriate caption was,
- 10:37:36and having the AI begin to learn from that.
- 10:37:39But now, because the network generates output
- 10:37:42that can be fed back into itself, you could
- 10:37:44imagine the output of the network being fed back into the same network.
- 10:37:47This here looks like a separate network, but it's really
- 10:37:50the same network that's just getting different input,
- 10:37:53that this network's output gets fed back into itself,
- 10:37:57but it's going to generate another output.
- 10:37:59And that other output is going to be the second word in the caption.
- 10:38:04And this recurrent neural network then, this network
- 10:38:06is going to generate other output that can be fed back into itself
- 10:38:09to generate yet another word, fed back into itself
- 10:38:12to generate another word.
- 10:38:13And so recurrent neural networks allow us to represent this one-to-many
- 10:38:18structure.
- 10:38:18You provide one image as input, and the neural network
- 10:38:21can pass data into the next run of the network, and then again and again,
- 10:38:25such that you could run the network multiple times,
- 10:38:28each time generating a different output still based on that original input.
- 10:38:33And this is where recurrent neural networks become particularly useful
- 10:38:37when dealing with sequences of inputs or outputs.
- 10:38:40And my output is a sequence of words, and since I can't very easily
- 10:38:43represent outputting an entire sequence of words,
- 10:38:45I'll instead output that sequence one word at a time
- 10:38:49by allowing my network to pass information about what still
- 10:38:52needs to be said about the photo into the next stage of running the network.
- 10:38:56So you could run the network multiple times, the same network
- 10:38:59with the same weights, just getting different input each time.
- 10:39:02First, getting input from the image, and then getting input from the network
- 10:39:06itself as additional information about what additionally
- 10:39:09needs to be given in a particular caption, for example.
- 10:39:13So this then is a one-to-many relationship inside of a recurrent neural
- 10:39:17network, but it turns out there are other models that we can use,
- 10:39:20other ways we can try and use recurrent neural networks
- 10:39:23to be able to represent data that might be stored in other forms as well.
- 10:39:26We saw how we could use neural networks in order to analyze images
- 10:39:29in the context of convolutional neural networks that take an image,
- 10:39:33figure out various different properties of the image,
- 10:39:35and are able to draw some sort of conclusion based on that.
- 10:39:38But you might imagine that something like YouTube,
- 10:39:40they need to be able to do a lot of learning based on video.
- 10:39:44They need to look through videos to detect if they're like copyright
- 10:39:46violations, or they need to be able to look through videos to maybe identify
- 10:39:50what particular items are inside of the video, for example.
- 10:39:53And video, you might imagine, is much more difficult to put in
- 10:39:56as input to a neural network, because whereas an image, you could just
- 10:40:00treat each pixel as a different value, videos are sequences.
- 10:40:03They're sequences of images, and each sequence might be of different length.
- 10:40:07And so it might be challenging to represent that entire video
- 10:40:10as a single vector of values that you could pass in to a neural network.
- 10:40:15And so here, too, recurrent neural networks
- 10:40:17can be a valuable solution for trying to solve this type of problem.
- 10:40:21Then instead of just passing in a single input into our neural network,
- 10:40:25we could pass in the input one frame at a time, you might imagine.
- 10:40:28First, taking the first frame of the video, passing it into the network,
- 10:40:32and then maybe not having the network output anything at all yet.
- 10:40:35Let it take in another input, and this time, pass it into the network.
- 10:40:40But the network gets information from the last time
- 10:40:43we provided an input into the network.
- 10:40:45Then we pass in a third input, and then a fourth input,
- 10:40:47where each time, what the network gets is it gets the most recent input,
- 10:40:51like each frame of the video.
- 10:40:53But it also gets information the network processed
- 10:40:56from all of the previous iterations.
- 10:40:58So on frame number four, you end up getting the input for frame number four
- 10:41:02plus information the network has calculated from the first three frames.
- 10:41:06And using all of that data combined, this recurrent neural network
- 10:41:10can begin to learn how to extract patterns from a sequence of data
- 10:41:14as well.
- 10:41:14And so you might imagine, if you want to classify a video
- 10:41:17into a number of different genres, like an educational video,
- 10:41:20or a music video, or different types of videos,
- 10:41:22that's a classification task, where you want
- 10:41:24to take as input each of the frames of the video,
- 10:41:27and you want to output something like what it is, what category
- 10:41:31that it happens to belong to.
- 10:41:33And you can imagine doing this sort of thing,
- 10:41:35this sort of many-to-one learning, any time your input is a sequence.
- 10:41:39And so input is a sequence in the context of video.
- 10:41:43It could be in the context of, like, if someone has typed a message
- 10:41:45and you want to be able to categorize that message,
- 10:41:47like if you're trying to take a movie review and trying to classify it
- 10:41:51as, is it a positive review or a negative review?
- 10:41:54That input is a sequence of words, and the output
- 10:41:56is a classification, positive or negative.
- 10:41:59There, too, a recurrent neural network might
- 10:42:01be helpful for analyzing sequences of words.
- 10:42:04And they're quite popular when it comes to dealing with language.
- 10:42:07Could even be used for spoken language as well,
- 10:42:09that spoken language is an audio waveform that
- 10:42:12can be segmented into distinct chunks.
- 10:42:14And each of those could be passed in as an input
- 10:42:17into a recurrent neural network to be able to classify someone's voice,
- 10:42:21for instance.
- 10:42:21If you want to do voice recognition to say, is this one person or is this
- 10:42:24another, here are also cases where you might
- 10:42:27want this many-to-one architecture for a recurrent neural network.
- 10:42:32And then as one final problem, just to take
- 10:42:34a look at in terms of what we can do with these sorts of networks,
- 10:42:37imagine what Google Translate is doing.
- 10:42:39So what Google Translate is doing is it's taking some text written
- 10:42:42in one language and converting it into text written in some other language,
- 10:42:47for example, where now this input is a sequence of data.
- 10:42:50It's a sequence of words.
- 10:42:52And the output is a sequence of words as well.
- 10:42:54It's also a sequence.
- 10:42:55So here we want effectively a many-to-many relationship.
- 10:42:58Our input is a sequence and our output is a sequence as well.
- 10:43:02And it's not quite going to work to just say,
- 10:43:05take each word in the input and translate it into a word in the output.
- 10:43:09Because ultimately, different languages put their words in different orders.
- 10:43:13And maybe one language uses two words for something,
- 10:43:15whereas another language only uses one.
- 10:43:17So we really want some way to take this information, this input,
- 10:43:22encode it somehow, and use that encoding to generate
- 10:43:25what the output ultimately should be.
- 10:43:27And this has been one of the big advancements in automated translation
- 10:43:30technology, is the ability to use the neural networks to do this instead
- 10:43:34of older, more traditional methods.
- 10:43:35And this has improved accuracy dramatically.
- 10:43:37And the way you might imagine doing this is, again,
- 10:43:40using a recurrent neural network with multiple inputs and multiple outputs.
- 10:43:44We start by passing in all the input.
- 10:43:45Input goes into the network.
- 10:43:47Another input, like another word, goes into the network.
- 10:43:49And we do this multiple times, like once for each word in the input
- 10:43:53that I'm trying to translate.
- 10:43:54And only after all of that is done does the network now
- 10:43:58start to generate output, like the first word of the translated sentence,
- 10:44:01and the next word of the translated sentence, so on and so forth,
- 10:44:04where each time the network passes information to itself
- 10:44:08by allowing for this model of giving some sort of state
- 10:44:12from one run in the network to the next run,
- 10:44:15assembling information about all the inputs,
- 10:44:17and then passing in information about which part of the output
- 10:44:20in order to generate next.
- 10:44:22And there are a number of different types of these sorts of recurrent neural
- 10:44:25networks.
- 10:44:26One of the most popular is known as the long short-term memory neural network,
- 10:44:29otherwise known as LSTM.
- 10:44:31But in general, these types of networks can be very, very powerful whenever
- 10:44:35we're dealing with sequences, whether those are sequences of images
- 10:44:38or especially sequences of words when it comes
- 10:44:40towards dealing with natural language.
- 10:44:43And so that then were just some of the different types
- 10:44:46of neural networks that can be used to do all sorts of different computations.
- 10:44:49And these are incredibly versatile tools that
- 10:44:52can be applied to a number of different domains.
- 10:44:54We only looked at a couple of the most popular types of neural networks
- 10:44:57from more traditional feed-forward neural networks, convolutional neural
- 10:45:00networks, and recurrent neural networks.
- 10:45:02But there are other types as well.
- 10:45:04There are adversarial networks where networks compete with each other
- 10:45:07to try and be able to generate new types of data,
- 10:45:10as well as other networks that can solve other tasks based
- 10:45:13on what they happen to be structured and adapted for.
- 10:45:15And these are very powerful tools in machine learning
- 10:45:18from being able to very easily learn based on some set of input data
- 10:45:21and to be able to, therefore, figure out how to calculate some function
- 10:45:25from inputs to outputs, whether it's input to some sort of classification
- 10:45:28like analyzing an image and getting a digit or machine translation
- 10:45:32where the input is in one language and the output is in another.
- 10:45:34These tools have a lot of applications for machine learning more generally.
- 10:45:39Next time, we'll look at machine learning and AI in particular
- 10:45:42in the context of natural language.
- 10:45:44We talked a little bit about this today, but looking at how it is that our AI
- 10:45:47can begin to understand natural language and can
- 10:45:50begin to be able to analyze and do useful tasks with regards
- 10:45:53to human language, which turns out to be a challenging and interesting task.
- 10:45:57So we'll see you next time.
- 10:46:00And welcome back, everybody, to our final class
- 10:46:21in an introduction to artificial intelligence with Python.
- 10:46:24Now, so far in this class, we've been taking problems
- 10:46:26that we want to solve intelligently and framing them
- 10:46:29in ways that computers are going to be able to make sense of.
- 10:46:31We've been taking problems and framing them as search problems
- 10:46:34or constraint satisfaction problems or optimization problems, for example.
- 10:46:38In essence, we have been trying to communicate
- 10:46:40about problems in ways that our computer is going to be able to understand.
- 10:46:45Today, the goal is going to be to get computers
- 10:46:47to understand the way you and I communicate naturally
- 10:46:50via our own natural languages, languages like English.
- 10:46:53But natural language contains a lot of nuance and complexity
- 10:46:57that's going to make it challenging for computers to be able to understand.
- 10:47:00So we'll need to explore some new tools and some new techniques
- 10:47:04to allow computers to make sense of natural language.
- 10:47:07So what is it exactly that we're trying to get computers to do?
- 10:47:10Well, they all fall under this general heading of natural language processing,
- 10:47:14getting computers to work with natural language.
- 10:47:17And these tasks include tasks like automatic summarization.
- 10:47:20Given a long text, can we train the computer
- 10:47:23to be able to come up with a shorter representation of it?
- 10:47:26Information extraction, getting the computer
- 10:47:28to pull out relevant facts or details out of some text.
- 10:47:31Machine translation, like Google Translate,
- 10:47:33translating some text from one language into another language.
- 10:47:36Question answering, if you've ever asked a question to your phone
- 10:47:39or had a conversation with an AI chatbot where you provide some text
- 10:47:43to the computer, the computer is able to understand that text
- 10:47:47and then generate some text in response.
- 10:47:50Text classification, where we provide some text to the computer
- 10:47:53and the computer assigns it a label, positive or negative,
- 10:47:56inbox or spam, for example.
- 10:47:58And there are several other kinds of tasks
- 10:48:00that all fall under this heading of natural language processing.
- 10:48:03But before we take a look at how the computer might
- 10:48:06try to solve these kinds of tasks, it might be useful for us
- 10:48:09to think about language in general.
- 10:48:11What are the kinds of challenges that we might need to deal with
- 10:48:14as we start to think about language and getting a computer
- 10:48:17to be able to understand it?
- 10:48:18So one part of language that we'll need to consider
- 10:48:21is the syntax of language.
- 10:48:22Syntax is all about the structure of language.
- 10:48:25Language is composed of individual words.
- 10:48:27And those words are composed together in some kind of structured whole.
- 10:48:31And if our computer is going to be able to understand language,
- 10:48:33it's going to need to understand something about that structure.
- 10:48:37So let's take a couple of examples.
- 10:48:39Here, for instance, is a sentence.
- 10:48:40Just before 9 o'clock, Sherlock Holmes stepped briskly into the room.
- 10:48:44That sentence is made up of words.
- 10:48:46And those words together form a structured whole.
- 10:48:49This is syntactically valid as a sentence.
- 10:48:52But we could take some of those same words,
- 10:48:55rearrange them, and come up with a sentence that is not syntactically valid.
- 10:48:59Here, for example, just before Sherlock Holmes 9 o'clock stepped briskly
- 10:49:03the room is still composed of valid words.
- 10:49:06But they're not in any kind of logical whole.
- 10:49:08This is not a syntactically well-formed sentence.
- 10:49:12Another interesting challenge is that some sentences will
- 10:49:15have multiple possible valid structures.
- 10:49:18Here's a sentence, for example.
- 10:49:20I saw the man on the mountain with a telescope.
- 10:49:23And here, this is a valid sentence.
- 10:49:25But it actually has two different possible structures
- 10:49:28that lend themselves to two different interpretations
- 10:49:31and two different meanings.
- 10:49:32Maybe I, the one doing the seeing, am the one with the telescope.
- 10:49:36Or maybe the man on the mountain is the one with the telescope.
- 10:49:39And so natural language is ambiguous.
- 10:49:41Sometimes the same sentence can be interpreted in multiple ways.
- 10:49:44And that's something that we'll need to think about as well.
- 10:49:47And this lends itself to another problem within language
- 10:49:50that we'll need to think about, which is semantics.
- 10:49:52While syntax is all about the structure of language,
- 10:49:55semantics is about the meaning of language.
- 10:49:57It's not enough for a computer just to know
- 10:49:59that a sentence is well-structured if it doesn't
- 10:50:02know what that sentence means.
- 10:50:04And so semantics is going to concern itself
- 10:50:06with the meaning of words and the meaning of sentences.
- 10:50:09So if we go back to that same sentence as before,
- 10:50:11just before 9 o'clock, Sherlock Holmes stepped briskly into the room,
- 10:50:16I could come up with another sentence, say the sentence,
- 10:50:19a few minutes before 9, Sherlock Holmes walked quickly into the room.
- 10:50:23And those are two different sentences with some of the words the same
- 10:50:26and some of the words different.
- 10:50:28But the two sentences have essentially the same meaning.
- 10:50:31And so ideally, whatever model we build, we'll
- 10:50:33be able to understand that these two sentences, while different,
- 10:50:36mean something very similar.
- 10:50:38Some syntactically well-formed sentences don't mean anything at all.
- 10:50:42A famous example from linguist Noam Chomsky
- 10:50:44is the sentence, colorless green ideas sleep furiously.
- 10:50:48This is a syntactically, structurally well-formed sentence.
- 10:50:52We've got adjectives modifying a noun, ideas.
- 10:50:55We've got a verb and an adverb in the correct positions.
- 10:50:58But when taken as a whole, the sentence doesn't really mean anything.
- 10:51:01And so if our computers are going to be able to work with natural language
- 10:51:05and perform tasks in natural language processing,
- 10:51:07these are some concerns we'll need to think about.
- 10:51:09We'll need to be thinking about syntax.
- 10:51:11And we'll need to be thinking about semantics.
- 10:51:14So how could we go about trying to teach a computer how
- 10:51:17to understand the structure of natural language?
- 10:51:20Well, one approach we might take is by starting
- 10:51:22by thinking about the rules of natural language.
- 10:51:25Our natural languages have rules.
- 10:51:27In English, for example, nouns tend to come before verbs.
- 10:51:30Nouns can be modified by adjectives, for example.
- 10:51:33And so if only we could formalize those rules,
- 10:51:36then we could give those rules to a computer,
- 10:51:38and the computer would be able to make sense of them and understand them.
- 10:51:41And so let's try to do exactly that.
- 10:51:43We're going to try to define a formal grammar.
- 10:51:46Where a formal grammar is some system of rules
- 10:51:49for generating sentences in a language.
- 10:51:52This is going to be a rule-based approach to natural language processing.
- 10:51:56We're going to give the computer some rules that we know about language
- 10:51:59and have the computer use those rules to make
- 10:52:01sense of the structure of language.
- 10:52:04And there are a number of different types of formal grammars.
- 10:52:06Each one of them has slightly different use cases.
- 10:52:09But today, we're going to focus specifically
- 10:52:11on one kind of grammar known as a context-free grammar.
- 10:52:14So how does the context-free grammar work?
- 10:52:16Well, here is a sentence that we might want a computer to generate.
- 10:52:19She saw the city.
- 10:52:21And we're going to call each of these words a terminal symbol.
- 10:52:24A terminal symbol, because once our computer has generated the word,
- 10:52:27there's nothing else for it to generate.
- 10:52:29Once it's generated the sentence, the computer is done.
- 10:52:32We're going to associate each of these terminal symbols
- 10:52:35with a non-terminal symbol that generates it.
- 10:52:39So here we've got n, which stands for noun, like she or city.
- 10:52:43We've got v as a non-terminal symbol, which stands for a verb.
- 10:52:46And then we have d, which stands for determiner.
- 10:52:48A determiner is a word like the or a or an in English, for example.
- 10:52:52So each of these non-terminal symbols can generate the terminal symbols
- 10:52:57that we ultimately care about generating.
- 10:52:59But how do we know, or how does the computer
- 10:53:01know which non-terminal symbols are associated with which terminal symbols?
- 10:53:05Well, to do that, we need some kind of rule.
- 10:53:08Here are some what we call rewriting rules that
- 10:53:11have a non-terminal symbol on the left-hand side of an arrow.
- 10:53:14And on the right side is what that non-terminal symbol can be replaced with.
- 10:53:18So here we're saying the non-terminal symbol n, again,
- 10:53:21which stands for noun, could be replaced by any of these options separated
- 10:53:25by vertical bars.
- 10:53:26n could be replaced by she or city or car or hairy.
- 10:53:30d for determiner could be replaced by the a or an and so forth.
- 10:53:34Each of these non-terminal symbols could be replaced by any of these words.
- 10:53:40We can also have non-terminal symbols that
- 10:53:42are replaced by other non-terminal symbols.
- 10:53:45Here is an interesting rule, np arrow n bar dn.
- 10:53:50So what does that mean?
- 10:53:52Well, np stands for a noun phrase.
- 10:53:55Sometimes when we have a noun phrase in a sentence,
- 10:53:57it's not just a single word, it could be multiple words.
- 10:54:00And so here we're saying a noun phrase could be just a noun,
- 10:54:04or it could be a determiner followed by a noun.
- 10:54:07So we might have a noun phrase that's just a noun, like she,
- 10:54:11that's a noun phrase.
- 10:54:12Or we could have a noun phrase that's multiple words, something
- 10:54:15like the city also acts as a noun phrase.
- 10:54:18But in this case, it's composed of two words, a determiner, the,
- 10:54:22and a noun city.
- 10:54:24We could do the same for verb phrases.
- 10:54:26A verb phrase, or VP, might be just a verb,
- 10:54:30or it might be a verb followed by a noun phrase.
- 10:54:33So we could have a verb phrase that's just a single word,
- 10:54:35like the word walked, or we could have a verb phrase
- 10:54:38that is an entire phrase, something like saw the city,
- 10:54:42as an entire verb phrase.
- 10:54:45A sentence, meanwhile, we might then define as a noun phrase
- 10:54:48followed by a verb phrase.
- 10:54:50And so this would allow us to generate a sentence like she saw the city,
- 10:54:54an entire sentence made up of a noun phrase, which is just the word she,
- 10:54:59and then a verb phrase, which is saw the city, saw which is a verb,
- 10:55:03and then the city, which itself is also a noun phrase.
- 10:55:07And so if we could give these rules to a computer explaining to it
- 10:55:11what non-terminal symbols could be replaced by what other symbols,
- 10:55:15then a computer could take a sentence and begin
- 10:55:17to understand the structure of that sentence.
- 10:55:20And so let's take a look at an example of how we might do that.
- 10:55:23And to do that, we're going to use a Python library called NLTK,
- 10:55:26or the Natural Language Toolkit, which we'll see a couple of times today.
- 10:55:30It contains a lot of helpful features and functions that we can use
- 10:55:33for trying to deal with and process natural language.
- 10:55:36So here we'll take a look at how we can use NLTK in order
- 10:55:39to parse a context-free grammar.
- 10:55:42So let's go ahead and open up cfg0.py, cfg standing for context-free grammar.
- 10:55:47And what you'll see in this file is that I first import NLTK, the Natural
- 10:55:51Language Toolkit.
- 10:55:53And the first thing I do is define a context-free grammar,
- 10:55:57saying that a sentence is a noun phrase followed by a verb phrase.
- 10:56:00I'm defining what a noun phrase is, defining what a verb phrase is,
- 10:56:03and then giving some examples of what I can
- 10:56:05do with these non-terminal symbols, D for determiner, N for noun,
- 10:56:10and V for verb.
- 10:56:12We're going to use NLTK to parse that grammar.
- 10:56:15Then we'll ask the user for some input in the form of a sentence
- 10:56:18and split it into words.
- 10:56:20And then we'll use this context-free grammar parser
- 10:56:23to try to parse that sentence and print out the resulting syntax tree.
- 10:56:28So let's take a look at an example.
- 10:56:30We'll go ahead and go into my cfg directory, and we'll run cfg0.py.
- 10:56:35And here I'm asked to type in a sentence.
- 10:56:37Let's say I type in she walked.
- 10:56:40And when I do that, I see that she walked is a valid sentence,
- 10:56:43where she is a noun phrase, and walked is the corresponding verb phrase.
- 10:56:49I could try to do this with a more complex sentence too.
- 10:56:52I could do something like she saw the city.
- 10:56:55And here we see that she is the noun phrase,
- 10:56:58and then saw the city is the entire verb phrase that makes up this sentence.
- 10:57:04So that was a very simple grammar.
- 10:57:06Let's take a look at a slightly more complex grammar.
- 10:57:08Here is cfg1.py, where a sentence is still a noun phrase followed
- 10:57:13by a verb phrase, but I've added some other possible non-terminal symbols too.
- 10:57:17I have AP for adjective phrase and PP for prepositional phrase.
- 10:57:22And we specified that we could have an adjective phrase
- 10:57:25before a noun phrase or a prepositional phrase after a noun, for example.
- 10:57:30So lots of additional ways that we might try to structure a sentence
- 10:57:34and interpret and parse one of those resulting sentences.
- 10:57:37So let's see that one in action.
- 10:57:39We'll go ahead and run cfg1.py with this new grammar.
- 10:57:43And we'll try a sentence like she saw the wide street.
- 10:57:48Here, Python's NLTK is able to parse that sentence
- 10:57:51and identify that she saw the wide street has this particular structure,
- 10:57:55a sentence with a noun phrase and a verb phrase,
- 10:57:58where that verb phrase has a noun phrase that within it
- 10:58:00contains an adjective.
- 10:58:02And so it's able to get some sense for what the structure of this language
- 10:58:06actually is.
- 10:58:07Let's try another example.
- 10:58:09Let's say she saw the dog with the binoculars.
- 10:58:14And we'll try that sentence.
- 10:58:16And here, we get one possible syntax tree,
- 10:58:19she saw the dog with the binoculars.
- 10:58:21But notice that this sentence is actually a little bit
- 10:58:24ambiguous in our own natural language.
- 10:58:26Who has the binoculars?
- 10:58:27Is it she who has the binoculars or the dog who has the binoculars?
- 10:58:31And NLTK is able to identify both possible structures for the sentence.
- 10:58:35In this case, the dog with the binoculars
- 10:58:38is an entire noun phrase.
- 10:58:40It's all underneath this NP here.
- 10:58:42So it's the dog that has the binoculars.
- 10:58:45But we also got an alternative parse tree,
- 10:58:48where the dog is just the noun phrase.
- 10:58:52And with the binoculars is a prepositional phrase modifying saw.
- 10:58:57So she saw the dog and she used the binoculars in order
- 10:59:01to see the dog as well.
- 10:59:03So this allows us to get a sense for the structure of natural language.
- 10:59:06But it relies on us writing all of these rules.
- 10:59:08And it would take a lot of effort to write all of the rules for any possible
- 10:59:12sentence that someone might write or say in the English language.
- 10:59:15Language is complicated.
- 10:59:16And as a result, there are going to be some very complex rules.
- 10:59:20So what else might we try?
- 10:59:21We might try to take a statistical lens towards approaching
- 10:59:24this problem of natural language processing.
- 10:59:27If we were able to give the computer a lot of existing data of sentences
- 10:59:31written in the English language, what could we try to learn from that data?
- 10:59:35Well, it might be difficult to try and interpret long pieces of text all
- 10:59:38at once.
- 10:59:39So instead, what we might want to do is break up that longer text
- 10:59:42into smaller pieces of information instead.
- 10:59:45In particular, we might try to create n-grams out of a longer sequence of text.
- 10:59:50An n-gram is just some contiguous sequence of n items from a sample of text.
- 10:59:55It might be n characters in a row or n words in a row, for example.
- 10:59:59So let's take a passage from Sherlock Holmes.
- 11:00:02And let's look for all of the trigrams.
- 11:00:04A trigram is an n-gram where n is equal to 3.
- 11:00:07So in this case, we're looking for sequences of three words in a row.
- 11:00:11So the trigrams here would be phrases like how often have.
- 11:00:15That's three words in a row.
- 11:00:16Often have I is another trigram.
- 11:00:18Have I said, I said to, said to you, to you that.
- 11:00:22These are all trigrams, sequences of three words that appear in sequence.
- 11:00:27And if we could give the computer a large corpus of text
- 11:00:30and have it pull out all of the trigrams in this case,
- 11:00:33it could get a sense for what sequences of three words
- 11:00:36tend to appear next to each other in our own natural language
- 11:00:40and, as a result, get some sense for what the structure of the language
- 11:00:45actually is.
- 11:00:46So let's take a look at an example of that.
- 11:00:48How can we use NLTK to try to get access to information about n-grams?
- 11:00:55So here, we're going to open up ngrams.py.
- 11:00:58And this is a Python program that's going to load a corpus of data, just
- 11:01:02some text files, into our computer's memory.
- 11:01:05And then we're going to use NLTK's ngrams function, which
- 11:01:08is going to go through the corpus of text, pulling out all of the ngrams
- 11:01:12for a particular value of n.
- 11:01:14And then, by using Python's counter class,
- 11:01:17we're going to figure out what are the most common ngrams inside
- 11:01:21of this entire corpus of text.
- 11:01:24And we're going to need a data set in order to do this.
- 11:01:26And I've prepared a data set of some of the stories of Sherlock Holmes.
- 11:01:29So it's just a bunch of text files.
- 11:01:32A lot of words for it to analyze.
- 11:01:33And as a result, we'll get a sense for what sequences of two words or three
- 11:01:38words that tend to be most common in natural language.
- 11:01:42So let's give this a try.
- 11:01:43We'll go into my ngrams directory.
- 11:01:45And we'll run ngrams.py.
- 11:01:47We'll try an n value of 2.
- 11:01:49So we're looking for sequences of two words in a row.
- 11:01:51And we'll use our corpus of stories from Sherlock Holmes.
- 11:01:55And when we run this program, we get a list of the most common ngrams
- 11:01:59where n is equal to 2, otherwise known as a bigram.
- 11:02:02So the most common one is of the.
- 11:02:04That's a sequence of two words that appears quite frequently
- 11:02:07in natural language.
- 11:02:08Then in the.
- 11:02:09And it was.
- 11:02:10These are all common sequences of two words that appear in a row.
- 11:02:14Let's instead now try running ngrams with n equal to 3.
- 11:02:18Let's get all of the trigrams and see what we get.
- 11:02:21And now we see the most common trigrams are it was a.
- 11:02:25One of the.
- 11:02:26I think that.
- 11:02:27These are all sequences of three words that appear quite frequently.
- 11:02:32And we were able to do this essentially via a process known as tokenization.
- 11:02:36Tokenization is the process of splitting a sequence of characters
- 11:02:39into pieces.
- 11:02:40In this case, we're splitting a long sequence of text into individual words
- 11:02:44and then looking at sequences of those words
- 11:02:46to get a sense for the structure of natural language.
- 11:02:49So once we've done this, once we've done the tokenization,
- 11:02:52once we've built up our corpus of ngrams, what
- 11:02:55can we do with that information?
- 11:02:57So the one thing that we might try is we could build a Markov chain,
- 11:03:00which you might recall from when we talked about probability.
- 11:03:02Recall that a Markov chain is some sequence of values
- 11:03:05where we can predict one value based on the values that came before it.
- 11:03:10And as a result, if we know all of the common ngrams in the English language,
- 11:03:14what words tend to be associated with what other words in sequence,
- 11:03:18we can use that to predict what word might come next in a sequence of words.
- 11:03:23And so we could build a Markov chain for language
- 11:03:26in order to try to generate natural language that
- 11:03:28follows the same statistical patterns as some input data.
- 11:03:33So let's take a look at that and build a Markov chain for natural language.
- 11:03:37And as input, I'm going to use the works of William Shakespeare.
- 11:03:41So here I have a file Shakespeare.txt, which
- 11:03:45is just a bunch of the works of William Shakespeare.
- 11:03:48It's a long text file, so plenty of data to analyze.
- 11:03:51And here in generator.py, I'm using a third party Python library
- 11:03:55in order to do this analysis.
- 11:03:57We're going to read in the sample of text,
- 11:04:00and then we're going to train a Markov model based on that text.
- 11:04:03And then we're going to have the Markov chain generate some sentences.
- 11:04:07We're going to generate a sentence that doesn't appear in the original text,
- 11:04:11but that follows the same statistical patterns that's generating it
- 11:04:14based on the ngrams trying to predict what word is likely to come next
- 11:04:19that we would expect based on those statistical patterns.
- 11:04:23So we'll go ahead and go into our Markov directory,
- 11:04:27run this generator with the works of William Shakespeare's input.
- 11:04:31And what we're going to get are five new sentences, where
- 11:04:34these sentences are not necessarily sentences
- 11:04:37from the original input text itself, but just that
- 11:04:39follow the same statistical patterns.
- 11:04:41It's predicting what word is likely to come next based on the input data
- 11:04:45that we've seen and the types of words that
- 11:04:47tend to appear in sequence there too.
- 11:04:50And so we're able to generate these sentences.
- 11:04:53Of course, so far, there's no guarantee that any of the sentences that
- 11:04:56are generated actually mean anything or make any sense.
- 11:04:59They just happen to follow the statistical patterns
- 11:05:01that our computer is already aware of.
- 11:05:04So we'll return to this issue of how to generate text
- 11:05:06in perhaps a more accurate or more meaningful way a little bit later.
- 11:05:09So let's now turn our attention to a slightly different problem,
- 11:05:12and that's the problem of text classification.
- 11:05:15Text classification is the problem where we have some text
- 11:05:18and we want to put that text into some kind of category.
- 11:05:21We want to apply some sort of label to that text.
- 11:05:24And this kind of problem shows up in a wide variety of places.
- 11:05:27A commonplace might be your email inbox, for example.
- 11:05:29You get an email and you want your computer
- 11:05:31to be able to identify whether the email belongs in your inbox
- 11:05:35or whether it should be filtered out into spam.
- 11:05:37So we need to classify the text.
- 11:05:39Is it a good email or is it spam?
- 11:05:42Another common use case is sentiment analysis.
- 11:05:44We might want to know whether the sentiment of some text
- 11:05:47is positive or negative.
- 11:05:50And so how might we do that?
- 11:05:51This comes up in situations like product reviews,
- 11:05:53where we might have a bunch of reviews for a product on some website.
- 11:05:57My grandson loved it so much fun.
- 11:05:58Product broke after a few days.
- 11:06:00One of the best games I've played in a long time and kind of cheap
- 11:06:03and flimsy, not worth it.
- 11:06:05Here's some example sentences that you might see on a product review website.
- 11:06:09And you and I could pretty easily look at this list of product reviews
- 11:06:12and decide which ones are positive and which ones are negative.
- 11:06:15We might say the first one and the third one,
- 11:06:17those seem like positive sentiment messages.
- 11:06:20But the second one and the fourth one seem like negative sentiment messages.
- 11:06:24But how did we know that?
- 11:06:25And how could we train a computer to be able to figure that out as well?
- 11:06:29Well, you might have clued your eye in on particular key words,
- 11:06:32where those particular words tend to mean something positive or negative.
- 11:06:36So you might have identified words like loved and fun and best
- 11:06:40tend to be associated with positive messages.
- 11:06:42And words like broke and cheap and flimsy
- 11:06:45tend to be associated with negative messages.
- 11:06:48So if only we could train a computer to be able to learn
- 11:06:51what words tend to be associated with positive versus negative messages,
- 11:06:55then maybe we could train a computer to do this kind of sentiment analysis
- 11:06:59as well.
- 11:07:00So we're going to try to do just that.
- 11:07:01We're going to use a model known as the bag of words model, which
- 11:07:05is a model that represents text as just an unordered collection of words.
- 11:07:09For the purpose of this model, we're not
- 11:07:11going to worry about the sequence and the ordering of the words,
- 11:07:13which word came first, second, or third.
- 11:07:15We're just going to treat the text as a collection of words
- 11:07:18in no particular order.
- 11:07:19And we're losing information there, right?
- 11:07:21The order of words is important.
- 11:07:22And we'll come back to that a little bit later.
- 11:07:24But for now, to simplify our model, it'll
- 11:07:26help us tremendously just to think about text
- 11:07:29as some unordered collection of words.
- 11:07:32And in particular, we're going to use the bag of words model
- 11:07:35to build something known as a naive Bayes classifier.
- 11:07:38So what is a naive Bayes classifier?
- 11:07:40Well, it's a tool that's going to allow us to classify text based on Bayes
- 11:07:43rule, again, which you might remember from when we talked about probability.
- 11:07:47Bayes rule says that the probability of B given A
- 11:07:51is equal to the probability of A given B multiplied
- 11:07:54by the probability of B divided by the probability of A.
- 11:07:59So how are we going to use this rule to be able to analyze text?
- 11:08:03Well, what are we interested in?
- 11:08:04We're interested in the probability that a message has
- 11:08:07a positive sentiment and the probability that a message has
- 11:08:10a negative sentiment, which I'm here for simplicity
- 11:08:12going to represent just with these emoji, happy face and frown face,
- 11:08:16as positive and negative sentiment.
- 11:08:18And so if I had a review, something like my grandson loved it,
- 11:08:22then what I'm interested in is not just the probability
- 11:08:25that a message has positive sentiment, but the conditional probability
- 11:08:29that a message has positive sentiment given
- 11:08:32that this is the message my grandson loved it.
- 11:08:35But how do I go about calculating this value, the probability
- 11:08:38that the message is positive given that the review is this sequence of words?
- 11:08:42Well, here's where the bag of words model comes in.
- 11:08:45Rather than treat this review as a string of a sequence of words in order,
- 11:08:49we're just going to treat it as an unordered collection of words.
- 11:08:52We're going to try to calculate the probability that the review is positive
- 11:08:56given that all of these words, my grandson loved it,
- 11:08:59are in the review in no particular order, just
- 11:09:02this unordered collection of words.
- 11:09:05And this is a conditional probability, which we can then apply Bayes rule
- 11:09:09to try to make sense of.
- 11:09:11And so according to Bayes rule, this conditional probability is equal to what?
- 11:09:16It's equal to the probability that all of these four words
- 11:09:19are in the review given that the review is positive multiplied
- 11:09:23by the probability that the review is positive divided by the probability
- 11:09:27that all of these words happen to be in the review.
- 11:09:30So this is the value now that we're going to try to calculate.
- 11:09:33Now, one thing you might notice is that the denominator here,
- 11:09:36the probability that all of these words appear in the review,
- 11:09:40doesn't actually depend on whether or not
- 11:09:42we're looking at the positive sentiment or negative sentiment case.
- 11:09:45So we can actually get rid of this denominator.
- 11:09:47We don't need to calculate it.
- 11:09:48We can just say that this probability is proportional to the numerator.
- 11:09:53And then at the end, we're going to need to normalize the probability
- 11:09:56distribution to make sure that all of the values sum up to the value 1.
- 11:10:00So now, how do we calculate this value?
- 11:10:03Well, this is the probability of all of these words given positive times
- 11:10:08probability of positive.
- 11:10:09And that, by the definition of joint probability,
- 11:10:12is just one big joint probability, the probability
- 11:10:15that all of these things are the case, that it's a positive review,
- 11:10:18and that all four of these words are in the review.
- 11:10:22But still, it's not entirely obvious how we calculate that value.
- 11:10:26And here is where we need to make one more assumption.
- 11:10:28And this is where the naive part of naive Bayes comes in.
- 11:10:32We're going to make the assumption that all of the words
- 11:10:34are independent of each other.
- 11:10:36And by that, I mean that if the word grandson is in the review,
- 11:10:40that doesn't change the probability that the word loved is in the review
- 11:10:43or that the word it is in the review, for example.
- 11:10:46And in practice, this assumption might not be true.
- 11:10:48It's almost certainly the case that the probability of words
- 11:10:51do depend on each other.
- 11:10:52But it's going to simplify our analysis and still give us reasonably good
- 11:10:56results just to assume that the words are independent of each other
- 11:10:59and they only depend on whether it's positive or negative.
- 11:11:03You might, for example, expect the word loved
- 11:11:06to appear more often in a positive review than in a negative review.
- 11:11:10So what does that mean?
- 11:11:11Well, if we make this assumption, then we
- 11:11:13can say that this value, the probability we're interested in,
- 11:11:16is not directly proportional to, but it's naively proportional to this value.
- 11:11:22The probability that the review is positive times the probability
- 11:11:26that my is in the review, given that it's positive,
- 11:11:29times the probability that grandson is in the review,
- 11:11:31given that it's positive, and so on for the other two words that
- 11:11:34happen to be in this review.
- 11:11:36And now this value, which looks a little more complex,
- 11:11:39is actually a value that we can calculate pretty easily.
- 11:11:42So how are we going to estimate the probability that the review is positive?
- 11:11:46Well, if we have some training data, some example data of example reviews
- 11:11:50where each one has already been labeled as positive or negative,
- 11:11:53then we can estimate the probability that a review is positive
- 11:11:56just by counting the number of positive samples
- 11:11:58and dividing by the total number of samples that we have in our training
- 11:12:02data.
- 11:12:03And for the conditional probabilities, the probability of loved,
- 11:12:06given that it's positive, well, that's going
- 11:12:08to be the number of positive samples with loved in it
- 11:12:11divided by the total number of positive samples.
- 11:12:15So let's take a look at an actual example to see how
- 11:12:17we could try to calculate these values.
- 11:12:19Here I've put together some sample data.
- 11:12:21The way to interpret the sample data is that based on the training data,
- 11:12:2449% of the reviews are positive, 51% are negative.
- 11:12:29And then over here in this table, we have some conditional probabilities.
- 11:12:33And then we have if the review is positive,
- 11:12:35then there is a 30% chance that my appears in it.
- 11:12:38And if the review is negative, there is a 20% chance that my appears in it.
- 11:12:42And based on our training data among the positive reviews,
- 11:12:451% of them contain the word grandson.
- 11:12:48And among the negative reviews, 2% contain the word grandson.
- 11:12:52So using this data, let's try to calculate this value,
- 11:12:56the value we're interested in.
- 11:12:57And to do that, we'll need to multiply all of these values together.
- 11:13:02The probability of positive, and then all
- 11:13:04of these positive conditional probabilities.
- 11:13:06And when we do that, we get some value.
- 11:13:09And then we can do the same thing for the negative case.
- 11:13:12We're going to do the same thing, take the probability that it's negative,
- 11:13:15multiply it by all of these conditional probabilities,
- 11:13:18and we're going to get some other value.
- 11:13:20And now these values don't sum to one.
- 11:13:22They're not a probability distribution yet.
- 11:13:24But I can normalize them and get some values.
- 11:13:27And that tells me that we're going to predict that my grandson loved it.
- 11:13:31We think there's a 68% chance, probability 0.68,
- 11:13:35that that is a positive sentiment review, and 0.32 probability
- 11:13:40that it's a negative review.
- 11:13:42So what problems might we run into here?
- 11:13:44What could potentially go wrong when doing this kind of analysis
- 11:13:47in order to analyze whether text has a positive or negative sentiment?
- 11:13:51Well, a couple of problems might arise.
- 11:13:53One problem might be, what if the word grandson never
- 11:13:57appears for any of the positive reviews?
- 11:14:00If that were the case, then when we try to calculate the value,
- 11:14:03the probability that we think the review is positive,
- 11:14:06we're going to multiply all these values together,
- 11:14:08and we're just going to get 0 for the positive case,
- 11:14:11because we're all going to ultimately multiply by that 0 value.
- 11:14:14And so we're going to say that we think there is no chance
- 11:14:17that the review is positive because it contains the word grandson.
- 11:14:20And in our training data, we've never seen the word grandson
- 11:14:23appear in a positive sentiment message before.
- 11:14:27And that's probably not the right analysis,
- 11:14:29because in cases of rare words, it might be the case
- 11:14:32that in nowhere in our training data did we ever
- 11:14:34see the word grandson appear in a message that has positive sentiment.
- 11:14:38So what can we do to solve this problem?
- 11:14:40Well, one thing we'll often do is some kind of additive smoothing,
- 11:14:43where we add some value alpha to each value in our distribution
- 11:14:46just to smooth out the data a little bit.
- 11:14:48And a common form of this is Laplace smoothing,
- 11:14:50where we add 1 to each value in our distribution.
- 11:14:53In essence, we pretend we've seen each value one more time
- 11:14:56than we actually have.
- 11:14:58So if we've never seen the word grandson for a positive review,
- 11:15:01we pretend we've seen it once.
- 11:15:02If we've seen it once, we pretend we've seen it twice,
- 11:15:04just to avoid the possibility that we might multiply by 0 and as a result,
- 11:15:09get some results we don't want in our analysis.
- 11:15:12So let's see what this looks like in practice.
- 11:15:14Let's try to do some naive Bayes classification in order
- 11:15:18to classify text as either positive or negative.
- 11:15:22We'll take a look at sentiment.py.
- 11:15:25And what this is going to do is load some sample data into memory,
- 11:15:28some examples of positive reviews and negative reviews.
- 11:15:32And then we're going to train a naive Bayes classifier
- 11:15:35on all of this training data, training data that
- 11:15:39includes all of the words we see in positive reviews
- 11:15:42and all of the words we see in negative reviews.
- 11:15:44And then we're going to try to classify some input.
- 11:15:48And so we're going to do this based on a corpus of data.
- 11:15:50I have some example positive reviews.
- 11:15:52Here are some positive reviews.
- 11:15:53It was great, so much fun, for example.
- 11:15:56And then some negative reviews, not worth it, kind of cheap.
- 11:15:59These are some examples of negative reviews.
- 11:16:02So now let's try to run this classifier and see
- 11:16:04how it would classify particular text as either positive or negative.
- 11:16:09We'll go ahead and run our sentiment analysis on this corpus.
- 11:16:14And we need to provide it with a review.
- 11:16:16So I'll say something like, I enjoyed it.
- 11:16:19And we see that the classifier says there is about a 0.92 probability
- 11:16:23that we think that this particular review is positive.
- 11:16:27Let's try something negative.
- 11:16:28We'll try kind of overpriced.
- 11:16:31And we see that there is a 0.96 probability
- 11:16:34now that we think that this particular review is negative.
- 11:16:37And so our naive Bayes classifier has learned what kinds of words
- 11:16:40tend to appear in positive reviews and what kinds of words
- 11:16:43tend to appear in negative reviews.
- 11:16:45And as a result of that, we've been able to design a classifier that
- 11:16:49can predict whether a particular review is positive or negative.
- 11:16:54And so this definitely is a useful tool that we can use
- 11:16:56to try and make some predictions.
- 11:16:58But we had to make some assumptions in order to get there.
- 11:17:01So what if we want to now try to build some more sophisticated models,
- 11:17:04use some tools from machine learning to try and take
- 11:17:07better advantage of language data to be able to draw
- 11:17:09more accurate conclusions and solve new kinds of tasks
- 11:17:12and new kinds of problems?
- 11:17:13Well, we've seen a couple of times now that when we want to take some data
- 11:17:17and take some input, put it in a way that the computer is
- 11:17:19going to be able to make sense of, it can be helpful to take that data
- 11:17:22and turn it into numbers, ultimately.
- 11:17:25And so what we might want to try to do is come up
- 11:17:27with some word representation, some way to take a word
- 11:17:30and translate its meaning into numbers.
- 11:17:33Because, for example, if we wanted to use a neural network
- 11:17:35to be able to process language, give our language to a neural network
- 11:17:39and have it make some predictions or perform some analysis there,
- 11:17:42a neural network takes its input and produces its output
- 11:17:45a vector of values, a vector of numbers.
- 11:17:48And so what we might want to do is take our data
- 11:17:51and somehow take words and convert them into some kind
- 11:17:54of numeric representation.
- 11:17:56So how might we do that?
- 11:17:57How might we take words and turn them into numbers?
- 11:18:01Let's take a look at an example.
- 11:18:03Here's a sentence, he wrote a book.
- 11:18:05And let's say I wanted to take each of those words
- 11:18:08and turn it into a vector of values.
- 11:18:10Here's one way I might do that.
- 11:18:11We'll say he is going to be a vector that has a 1 in the first position
- 11:18:15and the rest of the values are 0.
- 11:18:17Wrote will have a 1 in the second position and the rest of the values
- 11:18:20are 0.
- 11:18:21A has a 1 in the third position with the rest of the value 0.
- 11:18:24And book has a 1 in the fourth position with the rest of the value 0.
- 11:18:28So each of these words now has a distinct vector representation.
- 11:18:33And this is what we often call a one-hot representation,
- 11:18:36a representation of the meaning of a word as a vector with a single 1
- 11:18:41and all of the rest of the values are 0.
- 11:18:43And so when doing this, we now have a numeric representation for every word
- 11:18:47and we could pass in those vector representations
- 11:18:50into a neural network or other models that
- 11:18:52require some kind of numeric data as input.
- 11:18:55But this one-hot representation actually has a couple of problems
- 11:18:59and it's not ideal for a few reasons.
- 11:19:01One reason is, here we're just looking at four words.
- 11:19:03But if you imagine a vocabulary of thousands of words or more,
- 11:19:07these vectors are going to get quite long in order
- 11:19:09to have a distinct vector for every possible word in a vocabulary.
- 11:19:14And as a result of that, these longer vectors
- 11:19:16are going to be more difficult to deal with, more difficult to train,
- 11:19:19and so forth.
- 11:19:19And so that might be a problem.
- 11:19:21Another problem is a little bit more subtle.
- 11:19:24If we want to represent a word as a vector,
- 11:19:27and in particular the meaning of a word as a vector,
- 11:19:29then ideally it should be the case that words that have similar meanings
- 11:19:33should also have similar vector representations,
- 11:19:36so that they're close to each other together inside a vector space.
- 11:19:40But that's not really going to be the case with these one-hot representations,
- 11:19:44because if we take some similar words, say the word
- 11:19:46wrote and the word authored, which means similar things,
- 11:19:50they have entirely different vector representations.
- 11:19:54Likewise, book and novel, those two words mean somewhat similar things,
- 11:19:57but they have entirely different vector representations
- 11:20:00because they each have a one in some different position.
- 11:20:04And so that's not ideal either.
- 11:20:05So what we might be interested in instead
- 11:20:08is some kind of distributed representation.
- 11:20:10A distributed representation is the representation
- 11:20:13of the meaning of a word distributed across multiple values,
- 11:20:17instead of just being one-hot with a one in one position.
- 11:20:20Here is what a distributed representation of words might be.
- 11:20:25Each word is associated with some vector of values,
- 11:20:28with the meaning distributed across multiple values,
- 11:20:31ideally in such a way that similar words have
- 11:20:34a similar vector representation.
- 11:20:37But how are we going to come up with those values?
- 11:20:39Where do those values come from?
- 11:20:40How can we define the meaning of a word in this distributed sequence of numbers?
- 11:20:45Well, to do that, we're going to draw inspiration
- 11:20:47from a quote from British linguist J.R. Firth, who said,
- 11:20:50you shall know a word by the company it keeps.
- 11:20:54In other words, we're going to define the meaning of a word
- 11:20:56based on the words that appear around it, the context words around it.
- 11:21:01Take, for example, this context, for blank he ate.
- 11:21:05You might wonder, what words could reasonably fill in that blank?
- 11:21:08Well, it might be words like breakfast or lunch or dinner.
- 11:21:11All of those could reasonably fill in that blank.
- 11:21:14And so what we're going to say is because the words breakfast and lunch
- 11:21:17and dinner appear in a similar context, that they must have a similar meaning.
- 11:21:23And that's something our computer could understand and try to learn.
- 11:21:26A computer could look at a big corpus of text,
- 11:21:28look at what words tend to appear in similar context to each other,
- 11:21:32and use that to identify which words have a similar meaning
- 11:21:35and should therefore appear close to each other inside a vector space.
- 11:21:40And so one common model for doing this is known as the word to vec model.
- 11:21:44It's a model for generating word vectors, a vector representation for every word
- 11:21:48by looking at data and looking at the context in which a word appears.
- 11:21:52The idea is going to be this.
- 11:21:54If you start out with all of the words just in some random position in space
- 11:21:58and train it on some training data, what the word to vec model will do
- 11:22:02is start to learn what words appear in similar contexts.
- 11:22:05And it will move these vectors around in such a way
- 11:22:08that hopefully words with similar meanings, breakfast, lunch, and dinner,
- 11:22:12book, memoir, novel, will hopefully appear to be near to each other
- 11:22:17as vectors as well.
- 11:22:19So let's now take a look at what word to vec
- 11:22:21might look like in practice when implemented in code.
- 11:22:24What I have here inside of words.txt is a pre-trained model
- 11:22:29where each of these words has some vector representation
- 11:22:32trained by word to vec.
- 11:22:33Each of these words has some sequence of values representing its meaning,
- 11:22:38hopefully in such a way that similar words are represented by similar vectors.
- 11:22:43I also have this file vectors.py, which is going to open up the words
- 11:22:47and form them into a dictionary.
- 11:22:48And we also define some useful functions like distance
- 11:22:51to get the distance between two word vectors and closest words
- 11:22:55to find which words are nearby in terms of having close vectors to each other.
- 11:23:00And so let's give this a try.
- 11:23:02We'll go ahead and open a Python interpreter.
- 11:23:05And I'm going to import these vectors.
- 11:23:10And we might say, all right, what is the vector representation
- 11:23:13of the word book?
- 11:23:15And we get this big long vector that represents the word book
- 11:23:19as a sequence of values.
- 11:23:21And this sequence of values by itself is not all that meaningful.
- 11:23:24But it is meaningful in the context of comparing it
- 11:23:27to other vectors for other words.
- 11:23:30So we could use this distance function, which
- 11:23:32is going to get us the distance between two word vectors.
- 11:23:35And we might say, what is the distance between the vector
- 11:23:37representation for the word book and the vector representation
- 11:23:42for the word novel?
- 11:23:44And we see that it's 0.34.
- 11:23:46You can kind of interpret 0 as being really close together and 1
- 11:23:49being very far apart.
- 11:23:51And so now, what is the distance between book and, let's say, breakfast?
- 11:23:55Well, book and breakfast are more different from each other
- 11:23:58than book and novel are.
- 11:23:59So I would hopefully expect the distance to be larger.
- 11:24:02And in fact, it is 0.64 approximately.
- 11:24:05These two words are further away from each other.
- 11:24:08And what about now the distance between, let's say, lunch and breakfast?
- 11:24:13Well, that's about 0.2.
- 11:24:15Those are even closer together.
- 11:24:16They have a meaning that is closer to each other.
- 11:24:19Another interesting thing we might do is calculate the closest words.
- 11:24:24We might say, what are the closest words, according to Word2Vec,
- 11:24:28to the word book?
- 11:24:29And let's say, let's get the 10 closest words.
- 11:24:32What are the 10 closest vectors to the vector representation
- 11:24:35for the word book?
- 11:24:37And when we perform that analysis, we get this list of words.
- 11:24:40The closest one is book itself, but we also have books plural,
- 11:24:44and then essay, memoir, essays, novella, anthology, and so on.
- 11:24:48All of these words mean something similar to the word book,
- 11:24:52according to Word2Vec, at least, because they
- 11:24:54have a similar vector representation.
- 11:24:56So it seems like we've done a pretty good job of trying
- 11:24:59to capture this kind of vector representation of word meaning.
- 11:25:03One other interesting side effect of Word2Vec
- 11:25:06is that it's also able to capture something about the relationships
- 11:25:10between words as well.
- 11:25:12Let's take a look at an example.
- 11:25:13Here, for instance, are two words, man and king.
- 11:25:16And these are each represented by Word2Vec as vectors.
- 11:25:20So what might happen if I subtracted one from the other,
- 11:25:23calculated the value king minus man?
- 11:25:27Well, that will be the vector that will take us from man to king,
- 11:25:31somehow represent this relationship between the vector representation
- 11:25:35of the word man and the vector representation of the word king.
- 11:25:38And that's what this value, king minus man, represents.
- 11:25:42So what would happen if I took the vector representation of the word
- 11:25:45woman and added that same value, king minus man, to it?
- 11:25:51What would we get as the closest word to that, for example?
- 11:25:54Well, we could try it.
- 11:25:55Let's go ahead and go back to our Python interpreter and give this a try.
- 11:25:59I could say, what is the closest word to the vector representation
- 11:26:03of the word king minus the representation of the word man
- 11:26:07plus the representation of the word woman?
- 11:26:11And we see that the closest word is the word queen.
- 11:26:14We've somehow been able to capture the relationship between king and man.
- 11:26:17And then when we apply it to the word woman,
- 11:26:19we get, as the result, the word queen.
- 11:26:24So Word2Vec has been able to capture not just the words
- 11:26:27and how they're similar to each other, but also something
- 11:26:29about the relationships between words and how those words are connected
- 11:26:33to each other.
- 11:26:34So now that we have this vector representation of words,
- 11:26:37what can we now do with it?
- 11:26:38Now we can represent words as numbers.
- 11:26:40And so we might try to pass those words as input
- 11:26:43to, say, a neural network.
- 11:26:45Neural networks we've seen are very powerful tools
- 11:26:47for identifying patterns and making predictions.
- 11:26:50Recall that a neural network you can think of as all of these units.
- 11:26:53But really what the neural network is doing
- 11:26:55is taking some input, passing it into the network,
- 11:26:58and then producing some output.
- 11:27:00And by providing the neural network with training data,
- 11:27:02we're able to update the weights inside of the network
- 11:27:05so that the neural network can do a more accurate job of translating
- 11:27:09those inputs into those outputs.
- 11:27:11And now that we can represent words as numbers that
- 11:27:14could be the input or output, you could imagine passing a word in
- 11:27:18as input to a neural network and getting a word as output.
- 11:27:21And so when might that be useful?
- 11:27:23One common use for neural networks is in machine translation,
- 11:27:26when we want to translate text from one language into another,
- 11:27:29say translate English into French by passing English into the neural
- 11:27:33network and getting some French output.
- 11:27:36You might imagine, for instance, that we could take the English word for lamp,
- 11:27:39pass it into the neural network, get the French word for lamp as output.
- 11:27:43But in practice, when we're translating text from one language to another,
- 11:27:48we're usually not just interested in translating
- 11:27:50a single word from one language to another, but a sequence,
- 11:27:53say a sentence or a paragraph of words.
- 11:27:56Here, for example, is another paragraph, again taken
- 11:27:58from Sherlock Holmes, written in English.
- 11:28:00And what I might want to do is take that entire sentence,
- 11:28:03pass it into the neural network, and get as output a French translation
- 11:28:08of the same sentence.
- 11:28:10But recall that a neural network's input and output
- 11:28:12needs to be of some fixed size.
- 11:28:14And a sentence is not a fixed size.
- 11:28:16It's variable.
- 11:28:17You might have shorter sentences, and you might have longer sentences.
- 11:28:20So somehow, we need to solve the problem of translating
- 11:28:23a sequence into another sequence by means of a neural network.
- 11:28:27And that's going to be true not only for machine translation,
- 11:28:30but also for other problems, problems like question answering.
- 11:28:33If I want to pass as input a question, something
- 11:28:36like what is the capital of Massachusetts,
- 11:28:38feed that as input into the neural network,
- 11:28:41I would hope that what I would get as output
- 11:28:43is a sentence like the capital is Boston, again,
- 11:28:46translating some sequence into some other sequence.
- 11:28:50And if you've ever had a conversation with an AI chatbot,
- 11:28:53or have ever asked your phone a question,
- 11:28:55it needs to do something like this.
- 11:28:57It needs to understand the sequence of words that you, the human,
- 11:29:00provided as input.
- 11:29:02And then the computer needs to generate some sequence of words as output.
- 11:29:06So how can we do this?
- 11:29:07Well, one tool that we can use is the recurrent neural network, which
- 11:29:10we took a look at last time, which is a way for us
- 11:29:13to provide a sequence of values to a neural network
- 11:29:16by running the neural network multiple times.
- 11:29:18And each time we run the neural network, what we're going to do
- 11:29:22is we're going to keep track of some hidden state.
- 11:29:25And that hidden state is going to be passed
- 11:29:26from one run of the neural network to the next run of the neural network,
- 11:29:30keeping track of all of the relevant information.
- 11:29:33And so let's take a look at how we can apply that
- 11:29:35to something like this.
- 11:29:36And in particular, we're going to look at an architecture known
- 11:29:39as an encoder-decoder architecture, where
- 11:29:41we're going to encode this question into some kind of hidden state,
- 11:29:46and then use a decoder to decode that hidden state into the output
- 11:29:50that we're interested in.
- 11:29:52So what's that going to look like?
- 11:29:53We'll start with the first word, the word what.
- 11:29:55That goes into our neural network, and it's
- 11:29:58going to produce some hidden state.
- 11:30:00This is some information about the word what that our neural network is
- 11:30:04going to need to keep track of.
- 11:30:06Then when the second word comes along, we're
- 11:30:09going to feed it into that same encoder neural network,
- 11:30:12but it's going to get as input that hidden state as well.
- 11:30:15So we pass in the second word.
- 11:30:17We also get the information about the hidden state,
- 11:30:19and that's going to continue for the other words in the input.
- 11:30:23This is going to produce a new hidden state.
- 11:30:25And so then when we get to the third word, the, that goes into the encoder.
- 11:30:30It also gets access to the hidden state, and then it
- 11:30:32produces a new hidden state that gets passed into the next run
- 11:30:35when we use the word capital.
- 11:30:37And the same thing is going to repeat for the other words
- 11:30:39that appear in the input.
- 11:30:41So of Massachusetts, that produces one final piece of hidden state.
- 11:30:47Now somehow, we need to signal the fact that we're done.
- 11:30:50There's nothing left in the input.
- 11:30:51And we typically do this by passing some kind of special token,
- 11:30:54say an end token, into the neural network.
- 11:30:57And now the decoding process is going to start.
- 11:31:00We're going to generate the word the.
- 11:31:03But in addition to generating the word the,
- 11:31:06this decoder network is also going to generate some kind of hidden state.
- 11:31:11And so what happens the next time?
- 11:31:13Well, to generate the next word, it might
- 11:31:15be helpful to know what the first word was.
- 11:31:18So we might pass the first word the back into the decoder network.
- 11:31:22It's going to get as input this hidden state,
- 11:31:24and it's going to generate the next word capital.
- 11:31:27And that's also going to generate some hidden state.
- 11:31:30And we'll repeat that, passing capital into the network
- 11:31:32to generate the third word is, and then one more time
- 11:31:35in order to get the fourth word Boston.
- 11:31:38And at that point, we're done.
- 11:31:39But how do we know we're done?
- 11:31:40Usually, we'll do this one more time, pass Boston into the decoder network,
- 11:31:45and get an output some end token to indicate that that is the end of our input.
- 11:31:50And so this then is how we could use a recurrent neural network
- 11:31:53to take some input, encode it into some hidden state,
- 11:31:57and then use that hidden state to decode it into the output we're interested in.
- 11:32:01To visualize it in a slightly different way, we have some input sequence.
- 11:32:04This is just some sequence of words.
- 11:32:06That input sequence goes into the encoder, which in this case
- 11:32:10is a recurrent neural network generating these hidden states along the way
- 11:32:14until we generate some final hidden state, at which point
- 11:32:17we start the decoding process.
- 11:32:19Again, using a recurrent neural network, that's
- 11:32:21going to generate the output sequence as well.
- 11:32:23So we've got the encoder, which is encoding the information
- 11:32:26about the input sequence into this hidden state,
- 11:32:29and then the decoder, which takes that hidden state
- 11:32:32and uses it in order to generate the output sequence.
- 11:32:36But there are some problems.
- 11:32:37And for many years, this was the state of the art.
- 11:32:39The recurrent neural network and variance on this approach
- 11:32:42were some of the best ways we knew in order
- 11:32:44to perform tasks in natural language processing.
- 11:32:46But there are some problems that we might want to try to deal with
- 11:32:49and that have been dealt with over the years
- 11:32:51to try and improve upon this kind of model.
- 11:32:54And one problem you might notice happens in this encoder stage.
- 11:32:58We've taken this input sequence, the sequence of words,
- 11:33:01and encoded it all into this final piece of hidden state.
- 11:33:05And that final piece of hidden state needs
- 11:33:07to contain all of the information from the input sequence
- 11:33:10that we need in order to generate the output sequence.
- 11:33:14And while that's possible, it becomes increasingly difficult
- 11:33:18as the sequence gets larger and larger.
- 11:33:20For larger and larger input sequences, it's
- 11:33:22going to become more and more difficult to store
- 11:33:24all of the information we need about the input
- 11:33:27inside this single hidden state piece of context.
- 11:33:30That's a lot of information to pack into just a single value.
- 11:33:33It might be useful for us, when generating output,
- 11:33:36to not just refer to this one value, but to all
- 11:33:40of the previous hidden values that have been generated by the encoder.
- 11:33:44And so that might be useful, but how could we do that?
- 11:33:46We've got a lot of different values.
- 11:33:48We need to combine them somehow.
- 11:33:50So you could imagine adding them together,
- 11:33:52taking the average of them, for example.
- 11:33:54But doing that would assume that all of these pieces of hidden state
- 11:33:57are equally important.
- 11:33:59But that's not necessarily true either.
- 11:34:01Some of these pieces of hidden state are going
- 11:34:03to be more important than others, depending
- 11:34:05on what word they most closely correspond to.
- 11:34:08This piece of hidden state very closely corresponds
- 11:34:11to the first word of the input sequence.
- 11:34:13This one very closely corresponds to the second word of the input sequence,
- 11:34:16for example.
- 11:34:17And some of those are going to be more important than others.
- 11:34:21To make matters more complicated, depending
- 11:34:23on which word of the output sequence we're generating,
- 11:34:26different input words might be more or less important.
- 11:34:30And so what we really want is some way to decide for ourselves
- 11:34:33which of the input values are worth paying attention to,
- 11:34:37at what point in time.
- 11:34:38And this is the key idea behind a mechanism known as attention.
- 11:34:42Attention is all about letting us decide which values
- 11:34:45are important to pay attention to, when generating, in this case,
- 11:34:49the next word in our sequence.
- 11:34:51So let's take a look at an example of that.
- 11:34:54Here's a sentence.
- 11:34:55What is the capital of Massachusetts?
- 11:34:57Same sentence as before.
- 11:34:59And let's imagine that we were trying to answer that question
- 11:35:02by generating tokens of output.
- 11:35:04So what would the output look like?
- 11:35:05Well, it's going to look like something like the capital is.
- 11:35:09And let's say we're now trying to generate this last word here.
- 11:35:12What is that last word?
- 11:35:13How is the computer going to figure it out?
- 11:35:16Well, what it's going to need to do is decide
- 11:35:19which values it's going to pay attention to.
- 11:35:22And so the attention mechanism will allow
- 11:35:24us to calculate some attention scores for each word,
- 11:35:28some value corresponding to each word, determining how relevant
- 11:35:32is it for us to pay attention to that word right now?
- 11:35:36And in this case, when generating the fourth word of the output sequence,
- 11:35:39the most important words to pay attention to
- 11:35:42might be capital and Massachusetts, for example.
- 11:35:46That those words are going to be particularly relevant.
- 11:35:49And there are a number of different mechanisms
- 11:35:50that have been used in order to calculate these attention scores.
- 11:35:53It could be something as simple as a dot product
- 11:35:56to see how similar two vectors are, or we
- 11:35:58could train an entire neural network to calculate these attention scores.
- 11:36:02But the key idea is that during the training process for our neural network,
- 11:36:06we're going to learn how to calculate these attention scores.
- 11:36:09Our model is going to learn what is important to pay attention
- 11:36:12to in order to decide what the next word should be.
- 11:36:17So the result of all of this, calculating these attention scores,
- 11:36:20is that we can calculate some value, some value for each input word,
- 11:36:24determining how important is it for us to pay attention
- 11:36:28to that particular value.
- 11:36:29And recall that each of these input words
- 11:36:32is also associated with one of these hidden state context vectors,
- 11:36:36capturing information about the sentence up to that point,
- 11:36:39but primarily focused on that word in particular.
- 11:36:43And so what we can now do is if we have all of these vectors
- 11:36:46and we have values representing how important is it for us
- 11:36:49to pay attention to those particular vectors,
- 11:36:52is we can take a weighted average.
- 11:36:54We can take all of these vectors, multiply them by their attention scores,
- 11:36:58and add them up to get some new vector value, which
- 11:37:01is going to represent the context from the input,
- 11:37:04but specifically paying attention to the words
- 11:37:07that we think are most important.
- 11:37:09And once we've done that, that context vector
- 11:37:12can be fed into our decoder in order to say
- 11:37:14that the word should be, in this case, Boston.
- 11:37:18So attention is this very powerful tool that
- 11:37:21allows any word when we're trying to decode it
- 11:37:24to decide which words from the input should we pay attention to in order
- 11:37:28to determine what's important for generating the next word of the output.
- 11:37:33And one of the first places this was really used
- 11:37:35was in the field of machine translation.
- 11:37:37Here's an example of a diagram from the paper
- 11:37:39that introduced this idea, which was focused
- 11:37:42on trying to translate English sentences into French sentences.
- 11:37:45So we have an input English sentence up along the top,
- 11:37:48and then along the left side, the output French equivalent
- 11:37:51of that same sentence.
- 11:37:52And what you see in all of these squares are the attention scores
- 11:37:56visualized, where a lighter square indicates a higher attention score.
- 11:38:01And what you'll notice is that there's a strong correspondence
- 11:38:04between the French word and the equivalent English word,
- 11:38:07that the French word for agreement is really
- 11:38:10paying attention to the English word for agreement
- 11:38:12in order to decide what French word should be generated at that point
- 11:38:16in time.
- 11:38:17And sometimes you might pay attention to multiple words
- 11:38:19if you look at the French word for economic.
- 11:38:22That's primarily paying attention to the English word for economic,
- 11:38:25but also paying attention to the English word for European in this case too.
- 11:38:30And so attention scores are very easy to visualize
- 11:38:33to get a sense for what is our machine learning model really
- 11:38:37paying attention to, what information is it using in order
- 11:38:40to determine what's important and what's not in order
- 11:38:42to determine what the ultimate output token should be.
- 11:38:46And so when we combine the attention mechanism
- 11:38:49with a recurrent neural network, we can get very powerful and useful results
- 11:38:52where we're able to generate an output sequence by paying attention
- 11:38:56to the input sequence too.
- 11:38:58But there are other problems with this approach
- 11:39:00of using a recurrent neural network as well.
- 11:39:02In particular, notice that every run of the neural network
- 11:39:05depends on the output of the previous step.
- 11:39:07And that was important for getting a sense
- 11:39:09for the sequence of words and the ordering of those particular words.
- 11:39:12But we can't run this unit of the neural network
- 11:39:15until after we've calculated the hidden state from the run before it
- 11:39:19from the previous input token.
- 11:39:21And what that means is that it's very difficult to parallelize this process.
- 11:39:25That as the input sequence get longer and longer,
- 11:39:28we might want to use parallelism to try and speed up
- 11:39:31this process of training the neural network
- 11:39:33and making sense of all of this language data.
- 11:39:35But it's difficult to do that.
- 11:39:36And it's slow to do that with a recurrent neural network
- 11:39:39because all of it needs to be performed in sequence.
- 11:39:42And that's become an increasing challenge as we've
- 11:39:45started to get larger and larger language models.
- 11:39:47The more language data that we have available to us
- 11:39:50to use to train our machine learning models,
- 11:39:52the more accurate it can be, the better representation of language
- 11:39:55it can have, the better understanding it can have,
- 11:39:58and the better results that we can see.
- 11:40:00And so we've seen this growth of large language models
- 11:40:02that are using larger and larger data sets.
- 11:40:05But as a result, they take longer and longer to train.
- 11:40:08And so this problem that recurrent neural networks
- 11:40:10are not easy to parallelize has become an increasing problem.
- 11:40:15And as a result of that, that was one of the main motivations
- 11:40:18for a different architecture, for thinking about how
- 11:40:20to deal with natural language.
- 11:40:22And that's known as the transformer architecture.
- 11:40:25And this has been a significant milestone in the world of natural language
- 11:40:28processing for really increasing how well we can perform
- 11:40:32these kinds of natural language processing tasks,
- 11:40:34as well as how quickly we can train a machine learning model to be
- 11:40:37able to produce effective results.
- 11:40:39There are a number of different types of transformers
- 11:40:42in terms of how they work.
- 11:40:43But what we're going to take a look at here
- 11:40:45is the basic architecture for how one might work with a transformer
- 11:40:48to get a sense for what's involved and what we're doing.
- 11:40:52So let's start with the model we were looking at before,
- 11:40:54specifically at this encoder part of our encoder-decoder architecture,
- 11:40:59where we used a recurrent neural network to take this input
- 11:41:01sequence and capture all of this information about the hidden state
- 11:41:06and the information we need to know about that input sequence.
- 11:41:09Right now, it all needs to happen in this linear progression.
- 11:41:13But what the transformer is going to allow us to do
- 11:41:15is process each of the words independently in a way that's
- 11:41:18easy to parallelize, rather than have each word wait for some other word.
- 11:41:22Each word is going to go through this same neural network
- 11:41:26and produce some kind of encoded representation
- 11:41:29of that particular input word.
- 11:41:31And all of this is going to happen in parallel.
- 11:41:33Now, it's happening for all of the words at once,
- 11:41:35but we're really just going to focus on what's
- 11:41:37happening for one word to make it clear.
- 11:41:39But know that whatever you're seeing happen for this one word
- 11:41:41is going to happen for all of the other input words, too.
- 11:41:45So what's going on here?
- 11:41:47Well, we start with some input word.
- 11:41:49That input word goes into the neural network.
- 11:41:52And the output is hopefully some encoded representation of the input word,
- 11:41:57the information we need to know about the input word that's
- 11:41:59going to be relevant to us as we're generating the output.
- 11:42:03And because we're doing this each word independently,
- 11:42:06it's easy to parallelize.
- 11:42:07We don't have to wait for the previous word
- 11:42:09before we run this word through the neural network.
- 11:42:12But what did we lose in this process by trying to parallelize this whole thing?
- 11:42:16Well, we've lost all notion of word ordering.
- 11:42:19The order of words is important.
- 11:42:21The sentence, Sherlock Holmes gave the book to Watson,
- 11:42:24has a different meaning than Watson gave the book to Sherlock Holmes.
- 11:42:27And so we want to keep track of that information about word position.
- 11:42:31In the recurrent neural network, that happened for us automatically
- 11:42:34because we could run each word one at a time through the neural network,
- 11:42:37get the hidden state, pass it on to the next run of the neural network.
- 11:42:41But that's not the case here with the transformer,
- 11:42:44where each word is being processed independent of all of the other ones.
- 11:42:49So what are we going to do to try to solve that problem?
- 11:42:51One thing we can do is add some kind of positional encoding to the input word.
- 11:42:57The positional encoding is some vector that
- 11:42:59represents the position of the word in the sentence.
- 11:43:02This is the first word, the second word, the third word, and so forth.
- 11:43:05We're going to add that to the input word.
- 11:43:08And the result of that is going to be a vector
- 11:43:10that captures multiple pieces of information.
- 11:43:12It captures the input word itself as well as where in the sentence it appears.
- 11:43:17The result of that is we can pass the output of that addition,
- 11:43:20the addition of the input word and the positional encoding
- 11:43:23into the neural network.
- 11:43:24That way, the neural network knows the word and where
- 11:43:27it appears in the sentence and can use both of those pieces of information
- 11:43:31to determine how best to represent the meaning of that word
- 11:43:34in the encoded representation at the end of it.
- 11:43:38In addition to what we have here, in addition
- 11:43:40to the positional encoding and this feed forward neural network,
- 11:43:43we're also going to add one additional component, which
- 11:43:47is going to be a self-attention step.
- 11:43:49This is going to be attention where we're paying attention
- 11:43:52to the other input words.
- 11:43:54Because the meaning or interpretation of an input word
- 11:43:57might vary depending on the other words in the input as well.
- 11:44:00And so we're going to allow each word in the input
- 11:44:03to decide what other words in the input it should pay attention
- 11:44:06to in order to decide on its encoded representation.
- 11:44:10And that's going to allow us to get a better encoded representation
- 11:44:13for each word because words are defined by their context,
- 11:44:16by the words around them and how they're used in that particular context.
- 11:44:21This kind of self-attention is so valuable, in fact,
- 11:44:24that oftentimes the transformer will use multiple different self-attention
- 11:44:28layers at the same time to allow for this model
- 11:44:31to be able to pay attention to multiple facets of the input at the same time.
- 11:44:36And we call this multi-headed attention, where each attention head can pay
- 11:44:40attention to something different.
- 11:44:41And as a result, this network can learn to pay attention
- 11:44:45to many different parts of the input for this input word all at the same time.
- 11:44:49And in the spirit of deep learning, these two steps,
- 11:44:52this multi-headed self-attention layer and this neural network layer,
- 11:44:56that itself can be repeated multiple times, too,
- 11:44:59in order to get a deeper representation, in order
- 11:45:01to learn deeper patterns within the input text
- 11:45:04and ultimately get a better representation of language
- 11:45:07in order to get useful encoded representations of all of the input
- 11:45:11words.
- 11:45:12And so this is the process that a transformer might
- 11:45:15use in order to take an input word and get it its encoded representation.
- 11:45:20And the key idea is to really rely on this attention step
- 11:45:23in order to get information that's useful in order
- 11:45:26to determine how to encode that word.
- 11:45:29And that process is going to repeat for all of the input words that
- 11:45:32are in the input sequence.
- 11:45:33We're going to take all of the input words,
- 11:45:35encode them with some kind of positional encoding,
- 11:45:38feed those into these self-attention and feed-forward neural networks
- 11:45:42in order to ultimately get these encoded representations of the words.
- 11:45:46That's the result of the encoder.
- 11:45:48We get all of these encoded representations
- 11:45:51that will be useful to us when it comes time
- 11:45:53then to try to decode all of this information
- 11:45:57into the output sequence we're interested in.
- 11:45:59And again, this might take place in the context of machine translation,
- 11:46:02where the output is going to be the same sentence in a different language,
- 11:46:06or it might be an answer to a question in the case of an AI chatbot,
- 11:46:10for example.
- 11:46:11And so now let's take a look at how that decoder is going to work.
- 11:46:15Ultimately, it's going to have a very similar structure.
- 11:46:19Any time we're trying to generate the next output word,
- 11:46:21we need to know what the previous output word is,
- 11:46:25as well as its positional encoding.
- 11:46:27Where in the output sequence are we?
- 11:46:29And we're going to have these same steps, self-attention,
- 11:46:32because we might want an output word to be
- 11:46:34able to pay attention to other words in that same output,
- 11:46:37as well as a neural network.
- 11:46:39And that might itself repeat multiple times.
- 11:46:42But in this decoder, we're going to add one additional step.
- 11:46:45We're going to add an additional attention step, where
- 11:46:48instead of self-attention, where the output word is going
- 11:46:51to pay attention to other output words, in this step,
- 11:46:55we're going to allow the output word to pay attention
- 11:46:58to the encoded representations.
- 11:47:00So recall that the encoder is taking all of the input words
- 11:47:04and transforming them into these encoded representations
- 11:47:07of all of the input words.
- 11:47:08But it's going to be important for us to be able to decide which
- 11:47:11of those encoded representations we want to pay attention
- 11:47:14to when generating any particular token in the output sequence.
- 11:47:18And that's what this additional attention step is going to allow us to do.
- 11:47:22It's saying that every time we're generating a word of the output,
- 11:47:26we can pay attention to the other words in the output,
- 11:47:28because we might want to know, what are the words we've generated previously?
- 11:47:32And we want to pay attention to some of them
- 11:47:33to decide what word is going to be next in the sequence.
- 11:47:37But we also care about paying attention to the input words, too.
- 11:47:41And we want the ability to decide which of these encoded representations
- 11:47:44of the input words are going to be relevant in order
- 11:47:47for us to generate the next step.
- 11:47:49And so these two pieces combine together.
- 11:47:51We have this encoder that takes all of the input words
- 11:47:55and produces this encoded representation.
- 11:47:57And we have this decoder that is able to take the previous output word,
- 11:48:01pay attention to that encoded input, and then generate the next output word.
- 11:48:06And this is one of the possible architectures
- 11:48:08we could use for a transformer, with the key idea being
- 11:48:12these attention steps that allow words to pay attention to each other.
- 11:48:16During the training process here, we can now much more easily parallelize this,
- 11:48:20because we don't have to wait for all of the words to happen in sequence.
- 11:48:23And we can learn how we should perform these attention steps.
- 11:48:26The model is able to learn what is important to pay attention to,
- 11:48:30what things do I need to pay attention to,
- 11:48:32in order to be more accurate at predicting what the output word is.
- 11:48:37And this has proved to be a tremendously effective model
- 11:48:39for conversational AI agents, for building machine translation systems.
- 11:48:44And there have been many variants proposed on this model, too.
- 11:48:47Some transformers only use an encoder.
- 11:48:49Some only use a decoder.
- 11:48:51Some use some other combination of these different particular features.
- 11:48:54But the key ideas ultimately remain the same,
- 11:48:57this real focus on trying to pay attention to what is most important.
- 11:49:01And the world of natural language processing
- 11:49:04is fast growing and fast evolving.
- 11:49:06Year after year, we keep coming up with new models
- 11:49:08that allow us to do an even better job of performing
- 11:49:11these natural language related tasks, all on the surface
- 11:49:14of solving the tricky problem, which is our own natural language.
- 11:49:18We've seen how the syntax and semantics of our language is ambiguous,
- 11:49:21and it introduces all of these new challenges
- 11:49:24that we need to think about, if we're going
- 11:49:26to be able to design AI agents that are able to work with language
- 11:49:29effectively.
- 11:49:30So as we think about where we've been in this class,
- 11:49:33all of the different types of artificial intelligence we've considered,
- 11:49:36we've looked at artificial intelligence in a wide variety
- 11:49:38of different forms now.
- 11:49:40We started by taking a look at search problems,
- 11:49:42where we looked at how AI can search for solutions, play games,
- 11:49:46and find the optimal decision to make.
- 11:49:48We talked about knowledge, how AI can represent information that it knows
- 11:49:53and use that information to generate new knowledge as well.
- 11:49:57Then we looked at what AI can do when it's less certain,
- 11:49:59when it doesn't know things for sure, and we
- 11:50:01have to represent things in terms of probability.
- 11:50:04We then took a look at optimization problems.
- 11:50:06We saw how a lot of problems in AI can be boiled down
- 11:50:09to trying to maximize or minimize some function.
- 11:50:12And we looked at strategies that AI can use
- 11:50:15in order to do that kind of maximizing and minimizing.
- 11:50:18We then looked at the world of machine learning,
- 11:50:20learning from data in order to figure out some patterns
- 11:50:23and identify how to perform a task by looking at the training data
- 11:50:26that we have available to it.
- 11:50:28And one of the most powerful tools there was the neural network,
- 11:50:31the sequence of units whose weights can be trained in order
- 11:50:34to allow us to really effectively go from input to output
- 11:50:37and predict how to get there by learning these underlying patterns.
- 11:50:41And then today, we took a look at language itself,
- 11:50:44trying to understand how can we train the computer to be
- 11:50:47able to understand our natural language, to be
- 11:50:49able to understand syntax and semantics, make sense of and generate
- 11:50:53natural language, which introduces a number of interesting problems too.
- 11:50:57And we've really just scratched the surface of artificial intelligence.
- 11:51:00There is so much interesting research and interesting new techniques
- 11:51:03and algorithms and ideas being introduced
- 11:51:05to try to solve these types of problems.
- 11:51:07So I hope you enjoyed this exploration into the world
- 11:51:10of artificial intelligence.
- 11:51:11A huge thanks to all of the course's teaching staff and production team
- 11:51:14for making the class possible.
- 11:51:15This was an introduction to artificial intelligence with Python.
About this transcript
This page contains the full transcript of Harvard CS50’s Artificial Intelligence with Python – Full University Course by freeCodeCamp.org, generated from the public captions YouTube serves with the video. The transcript has 144,344 words across 14,299 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.