How to implement KNN from scratch with Python — Transcript
Full transcript
- 0:00the first algorithm we're going to look
- 0:01into is k n or k nearest neighbors
- 0:05how knm works it's basically given a
- 0:08data point you calculate this data point
- 0:10distance from all other data points in
- 0:12your data set
- 0:14and then you get the closest k points so
- 0:16this k is a hyper parameter that the
- 0:19user determines
- 0:21and in regression to get the results you
- 0:23get the average of the values of the k
- 0:26nearest neighbors
- 0:27or in classification you get the label
- 0:30of this data point using the majority
- 0:32vote of the k nearest neighbors
- 0:34so maybe let's see this on an example
- 0:37let's say the green values that we have
- 0:40here are one group and the red values
- 0:43that we see here are another group and
- 0:44then we have a new data point the yellow
- 0:46point
- 0:47what we do is we get the distance of the
- 0:49yellow point to all other data points in
- 0:51our data sets
- 0:53we get the k closest ones so let's say
- 0:55for this example is 3 and in this case
- 0:58it's a classification so we get the
- 1:00majority vote all of them are green that
- 1:02means this point also needs to be green
- 1:05so let's see how we can implement this
- 1:07algorithm in python all right let's
- 1:09start building the canon algorithm so
- 1:10i'm going to make it into a class
- 1:12actually
- 1:15and in the initialization
- 1:18function what i need to pass it is of
- 1:21course self
- 1:22and this is a k nearest neighbor's
- 1:24algorithm and the k is going to be
- 1:25determined when the model is created so
- 1:28that's why i'm also going to have to
- 1:30pass it a k value
- 1:32for now i can say the default value for
- 1:35k is 3
- 1:36and then we create k
- 1:39and this class is going to have a fit
- 1:41function
- 1:43and a
- 1:44predict function in the fit function we
- 1:46don't really need to do much basically
- 1:50what we have to do is to
- 1:53keep the values for the
- 1:55x and y
- 1:57data sets
- 2:02and of course i also need to pass it
- 2:04here
- 2:05to the fit function and the predict
- 2:07function is where we're going to do all
- 2:08the calculations so calculating the
- 2:10distance between this data point and all
- 2:13the other data points and finding the
- 2:15closest ones and then getting the
- 2:16prediction for that to the predict
- 2:18function we're going to be passing the
- 2:20testing data set so the data points that
- 2:22you want the prediction for so what i'm
- 2:24going to do is actually to create a
- 2:27helper function another predict function
- 2:30that will get a single data point value
- 2:34and here what i'm going to do
- 2:36is to say the predictions will be
- 2:40self
- 2:41the helpful function
- 2:43for each of the examples in the data set
- 2:47that is being sent to us
- 2:50and then i can return these predictions
- 2:52and here in this helper function i'm
- 2:54going to calculate the distance of this
- 2:56little x so one single data point uh to
- 3:00all the points in our x train and then
- 3:03return the label
- 3:05based on the three nearest neighbors the
- 3:08main thing that i need to do here is to
- 3:10compute
- 3:11the distances
- 3:13and then i need to
- 3:15get the
- 3:16closest
- 3:18k
- 3:21closest
- 3:24and finally we need to determine the
- 3:27label with majority vote
- 3:30so the computer distance i'm going to be
- 3:32using euclidean distance so
- 3:34let's create a
- 3:36i don't know where this came from
- 3:38but let's create a euclidean distance
- 3:42global function
- 3:46that given
- 3:48to erase will give us a distance between
- 3:51them
- 3:52and numpy square root
- 3:56numpy sum
- 3:58of x one six two
- 4:06of course i also need to
- 4:09import numpy for this
- 4:16distance and then i can return the
- 4:18distance
- 4:22so here i'm going to calculate the
- 4:24distances
- 4:31and the distance is going to be between
- 4:33this x that is passed past this function
- 4:36and each value in x train
- 4:44but self extreme of course
- 4:46from here i'm going to use ark sort from
- 4:49numpy
- 4:52on top of the distances
- 4:55and after it's sorted i'm going to get
- 4:58the first k
- 5:00of these distances of of their indices
- 5:03at least what arcsort does is basically
- 5:05tells you where the original
- 5:08indices of
- 5:10from the previous array from the
- 5:13original array would be after they are
- 5:16sorted so then when you get the first k
- 5:18uh effectively you're getting getting
- 5:20the indices of the closest three
- 5:23neighbors for this data point that we're
- 5:25working with so that would give me the
- 5:27indices
- 5:31and then i will get their labels
- 5:35nearest
- 5:36labels and we can get that from y train
- 5:44for e in
- 5:46the closest indices
- 5:51to get the most common class label i'm
- 5:53going to use
- 5:55from the collections library
- 5:59a counter data structure
- 6:02oops
- 6:05it's just going to make it a bit easier
- 6:06for us
- 6:11i can get the k nearest labels and then
- 6:14i can ask for the most common one
- 6:20and basically all i need to do is to
- 6:22return this most common
- 6:24label so let's see if everything works
- 6:26as intended now i've already imported
- 6:28the iris data set from
- 6:31sklearn from scikit-learn and let's see
- 6:34what the data set looks like first all
- 6:36right so this is what the data set look
- 6:38looks like it looks like there are three
- 6:40separate clusters of labels
- 6:43and the next thing that i want to do is
- 6:46to create a classifier
- 6:52i'll close this
- 6:54with k n but of course i need to import
- 6:56canon here
- 6:58since i just created it
- 7:00from k n we import k n
- 7:05and what we need to pass it is the k
- 7:07value
- 7:08uh let's say okay let's say 5 for now
- 7:10then we call the fit function
- 7:13over the x strain
- 7:15and y train
- 7:17and then we need to do predictions
- 7:22uh why
- 7:24then we send it to x test and that would
- 7:27give me
- 7:28some predictions
- 7:30uh but let's see what these predictions
- 7:32look like first
- 7:37all right so this is one result this is
- 7:40one prediction that we get uh as you
- 7:42remember you might remember we are
- 7:44getting it from the counter the most
- 7:47common function and what it returns is a
- 7:50list
- 7:51of the counts of all instances
- 7:55and uh yeah so how many times it has
- 7:58occurred and what the name of this label
- 8:00is so instead of that of course we need
- 8:02to return only the name of the label and
- 8:04nothing else so that's why i'm going to
- 8:06have to select the first one and the
- 8:08first
- 8:09value inside this tuple also and that is
- 8:12going to give me the labels so let's run
- 8:15this again and see
- 8:18okay now it looks like it's giving me
- 8:20actual labels it's either 0 1 or 2. and
- 8:23now i also want to calculate this
- 8:24accuracy to see if it's working well or
- 8:27not and that is actually quite easy to
- 8:29do i'll just say accuracy
- 8:34count how many times predictions
- 8:38are the same as y test
- 8:40and divide this by
- 8:42number of data points in y test and then
- 8:45we can print this
- 8:49let's see
- 8:520.96 so that's pretty good already for
- 8:55something that we implemented in like
- 8:56what 10 minutes or something like that
- 8:58so that's great that means our k n is
- 9:00working don't forget that you can get
- 9:02this code through our github repository
- 9:04the link is in the description and if
- 9:05you have any questions don't forget to
- 9:07leave a comment i hope you liked this
- 9:08video and i will see you in the next
- 9:10lesson
- 9:11[Music]
- 9:23you
About this transcript
This page contains the full transcript of How to implement KNN from scratch with Python by AssemblyAI, generated from the public captions YouTube serves with the video. The transcript has 1,271 words across 217 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.