YouTube2Text

How to implement KNN from scratch with Python — Transcript

by AssemblyAI · 1,271 words · 217 segments · language en · Watch on YouTube

Full transcript

  1. 0:00the first algorithm we're going to look
  2. 0:01into is k n or k nearest neighbors
  3. 0:05how knm works it's basically given a
  4. 0:08data point you calculate this data point
  5. 0:10distance from all other data points in
  6. 0:12your data set
  7. 0:14and then you get the closest k points so
  8. 0:16this k is a hyper parameter that the
  9. 0:19user determines
  10. 0:21and in regression to get the results you
  11. 0:23get the average of the values of the k
  12. 0:26nearest neighbors
  13. 0:27or in classification you get the label
  14. 0:30of this data point using the majority
  15. 0:32vote of the k nearest neighbors
  16. 0:34so maybe let's see this on an example
  17. 0:37let's say the green values that we have
  18. 0:40here are one group and the red values
  19. 0:43that we see here are another group and
  20. 0:44then we have a new data point the yellow
  21. 0:46point
  22. 0:47what we do is we get the distance of the
  23. 0:49yellow point to all other data points in
  24. 0:51our data sets
  25. 0:53we get the k closest ones so let's say
  26. 0:55for this example is 3 and in this case
  27. 0:58it's a classification so we get the
  28. 1:00majority vote all of them are green that
  29. 1:02means this point also needs to be green
  30. 1:05so let's see how we can implement this
  31. 1:07algorithm in python all right let's
  32. 1:09start building the canon algorithm so
  33. 1:10i'm going to make it into a class
  34. 1:12actually
  35. 1:15and in the initialization
  36. 1:18function what i need to pass it is of
  37. 1:21course self
  38. 1:22and this is a k nearest neighbor's
  39. 1:24algorithm and the k is going to be
  40. 1:25determined when the model is created so
  41. 1:28that's why i'm also going to have to
  42. 1:30pass it a k value
  43. 1:32for now i can say the default value for
  44. 1:35k is 3
  45. 1:36and then we create k
  46. 1:39and this class is going to have a fit
  47. 1:41function
  48. 1:43and a
  49. 1:44predict function in the fit function we
  50. 1:46don't really need to do much basically
  51. 1:50what we have to do is to
  52. 1:53keep the values for the
  53. 1:55x and y
  54. 1:57data sets
  55. 2:02and of course i also need to pass it
  56. 2:04here
  57. 2:05to the fit function and the predict
  58. 2:07function is where we're going to do all
  59. 2:08the calculations so calculating the
  60. 2:10distance between this data point and all
  61. 2:13the other data points and finding the
  62. 2:15closest ones and then getting the
  63. 2:16prediction for that to the predict
  64. 2:18function we're going to be passing the
  65. 2:20testing data set so the data points that
  66. 2:22you want the prediction for so what i'm
  67. 2:24going to do is actually to create a
  68. 2:27helper function another predict function
  69. 2:30that will get a single data point value
  70. 2:34and here what i'm going to do
  71. 2:36is to say the predictions will be
  72. 2:40self
  73. 2:41the helpful function
  74. 2:43for each of the examples in the data set
  75. 2:47that is being sent to us
  76. 2:50and then i can return these predictions
  77. 2:52and here in this helper function i'm
  78. 2:54going to calculate the distance of this
  79. 2:56little x so one single data point uh to
  80. 3:00all the points in our x train and then
  81. 3:03return the label
  82. 3:05based on the three nearest neighbors the
  83. 3:08main thing that i need to do here is to
  84. 3:10compute
  85. 3:11the distances
  86. 3:13and then i need to
  87. 3:15get the
  88. 3:16closest
  89. 3:18k
  90. 3:21closest
  91. 3:24and finally we need to determine the
  92. 3:27label with majority vote
  93. 3:30so the computer distance i'm going to be
  94. 3:32using euclidean distance so
  95. 3:34let's create a
  96. 3:36i don't know where this came from
  97. 3:38but let's create a euclidean distance
  98. 3:42global function
  99. 3:46that given
  100. 3:48to erase will give us a distance between
  101. 3:51them
  102. 3:52and numpy square root
  103. 3:56numpy sum
  104. 3:58of x one six two
  105. 4:06of course i also need to
  106. 4:09import numpy for this
  107. 4:16distance and then i can return the
  108. 4:18distance
  109. 4:22so here i'm going to calculate the
  110. 4:24distances
  111. 4:31and the distance is going to be between
  112. 4:33this x that is passed past this function
  113. 4:36and each value in x train
  114. 4:44but self extreme of course
  115. 4:46from here i'm going to use ark sort from
  116. 4:49numpy
  117. 4:52on top of the distances
  118. 4:55and after it's sorted i'm going to get
  119. 4:58the first k
  120. 5:00of these distances of of their indices
  121. 5:03at least what arcsort does is basically
  122. 5:05tells you where the original
  123. 5:08indices of
  124. 5:10from the previous array from the
  125. 5:13original array would be after they are
  126. 5:16sorted so then when you get the first k
  127. 5:18uh effectively you're getting getting
  128. 5:20the indices of the closest three
  129. 5:23neighbors for this data point that we're
  130. 5:25working with so that would give me the
  131. 5:27indices
  132. 5:31and then i will get their labels
  133. 5:35nearest
  134. 5:36labels and we can get that from y train
  135. 5:44for e in
  136. 5:46the closest indices
  137. 5:51to get the most common class label i'm
  138. 5:53going to use
  139. 5:55from the collections library
  140. 5:59a counter data structure
  141. 6:02oops
  142. 6:05it's just going to make it a bit easier
  143. 6:06for us
  144. 6:11i can get the k nearest labels and then
  145. 6:14i can ask for the most common one
  146. 6:20and basically all i need to do is to
  147. 6:22return this most common
  148. 6:24label so let's see if everything works
  149. 6:26as intended now i've already imported
  150. 6:28the iris data set from
  151. 6:31sklearn from scikit-learn and let's see
  152. 6:34what the data set looks like first all
  153. 6:36right so this is what the data set look
  154. 6:38looks like it looks like there are three
  155. 6:40separate clusters of labels
  156. 6:43and the next thing that i want to do is
  157. 6:46to create a classifier
  158. 6:52i'll close this
  159. 6:54with k n but of course i need to import
  160. 6:56canon here
  161. 6:58since i just created it
  162. 7:00from k n we import k n
  163. 7:05and what we need to pass it is the k
  164. 7:07value
  165. 7:08uh let's say okay let's say 5 for now
  166. 7:10then we call the fit function
  167. 7:13over the x strain
  168. 7:15and y train
  169. 7:17and then we need to do predictions
  170. 7:22uh why
  171. 7:24then we send it to x test and that would
  172. 7:27give me
  173. 7:28some predictions
  174. 7:30uh but let's see what these predictions
  175. 7:32look like first
  176. 7:37all right so this is one result this is
  177. 7:40one prediction that we get uh as you
  178. 7:42remember you might remember we are
  179. 7:44getting it from the counter the most
  180. 7:47common function and what it returns is a
  181. 7:50list
  182. 7:51of the counts of all instances
  183. 7:55and uh yeah so how many times it has
  184. 7:58occurred and what the name of this label
  185. 8:00is so instead of that of course we need
  186. 8:02to return only the name of the label and
  187. 8:04nothing else so that's why i'm going to
  188. 8:06have to select the first one and the
  189. 8:08first
  190. 8:09value inside this tuple also and that is
  191. 8:12going to give me the labels so let's run
  192. 8:15this again and see
  193. 8:18okay now it looks like it's giving me
  194. 8:20actual labels it's either 0 1 or 2. and
  195. 8:23now i also want to calculate this
  196. 8:24accuracy to see if it's working well or
  197. 8:27not and that is actually quite easy to
  198. 8:29do i'll just say accuracy
  199. 8:34count how many times predictions
  200. 8:38are the same as y test
  201. 8:40and divide this by
  202. 8:42number of data points in y test and then
  203. 8:45we can print this
  204. 8:49let's see
  205. 8:520.96 so that's pretty good already for
  206. 8:55something that we implemented in like
  207. 8:56what 10 minutes or something like that
  208. 8:58so that's great that means our k n is
  209. 9:00working don't forget that you can get
  210. 9:02this code through our github repository
  211. 9:04the link is in the description and if
  212. 9:05you have any questions don't forget to
  213. 9:07leave a comment i hope you liked this
  214. 9:08video and i will see you in the next
  215. 9:10lesson
  216. 9:11[Music]
  217. 9:23you

About this transcript

This page contains the full transcript of How to implement KNN from scratch with Python by AssemblyAI, generated from the public captions YouTube serves with the video. The transcript has 1,271 words across 217 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.